Actions4
Agent
AI-generatedOverview
This node operation runs an AI agent to extract data from web pages based on a natural language prompt. It allows users to specify seed URLs for the agent to start from, set a maximum budget in USD for the agent's spend, choose whether to only process the seed URLs without following links, select the LLM model to use, and optionally provide a JSON schema describing the expected structure of the extracted data. The node sends these parameters to the WebCrawlerAPI's agent endpoint, waits for the agent job to complete, and returns the extracted data as JSON. This operation is useful for automated data extraction, web research, and content summarization tasks where users want to leverage AI to interpret and extract specific information from web content.
Use Case Examples
- Extracting product details from a list of e-commerce URLs using a natural language prompt.
- Summarizing news articles from multiple seed URLs with a budget constraint on API usage.
- Using a custom JSON schema to structure extracted data from web pages for further processing.
Properties
| Name | Meaning |
|---|---|
| Prompt | Natural language instruction describing what to extract or find from the web pages. |
| Max Spend (USD) | Maximum budget in USD the agent may spend during its operation. |
| URLs | Comma-separated seed URLs for the agent to start crawling and extracting data from. |
| Seed URLs Only | Boolean flag indicating whether to only process the provided seed URLs without following links on those pages. |
| Model | The large language model (LLM) to use for the agent's data extraction and interpretation. |
| Output Schema (JSON) | Optional JSON Schema describing the expected structure of the extracted data to guide the agent's output formatting. |
Output
JSON
id- Unique identifier of the agent run job.status- Current status of the agent run (e.g., 'done', 'error', 'canceled').result- Extracted data or output produced by the agent according to the prompt and optional schema.error_reason- Reason for failure if the agent run ended with an error.
Dependencies
- WebCrawlerAPI service with API key authentication
Troubleshooting
- Ensure the provided JSON schema is valid JSON to avoid parsing errors.
- Check that the seed URLs are accessible and correctly formatted.
- Monitor the max spend budget to prevent premature termination of the agent run.
- Handle possible timeouts as the agent run can take up to 30 minutes to complete.
- Review error messages returned by the API for specific failure reasons and adjust parameters accordingly.
Links
- WebCrawlerAPI Documentation - Official API documentation for WebCrawlerAPI including agent operation details.