Actions4
Crawl
AI-generatedOverview
This node performs a web crawling operation starting from a specified seed URL. It systematically visits multiple pages up to a defined limit and depth, optionally filtering URLs using whitelist and blacklist regular expressions, and respecting the site's robots.txt rules if enabled. The node can return the crawled content either inline as markdown or as a URL to a combined markdown file. This is useful for scenarios like gathering content from multiple related web pages for analysis, content aggregation, or research.
Use Case Examples
- Crawl a website starting from the homepage URL, limiting to 10 pages and a maximum depth of 2, to collect content for market research.
- Use whitelist and blacklist regex patterns to selectively crawl only certain sections of a website, e.g., only product pages but exclude blog posts.
- Enable respect for robots.txt to ensure the crawl adheres to the website's crawling policies.
Properties
| Name | Meaning |
|---|---|
| URL to Crawl | Seed URL to start crawling from, the initial page where the crawl begins. |
| Items Limit | Maximum number of pages to crawl, limiting the crawl size. |
| Whitelist Regexp | Regex pattern to restrict crawling only to URLs matching this pattern. |
| Blacklist Regexp | Regex pattern to skip URLs matching this pattern during crawling. |
| Respect Robots.txt | Boolean flag to honor the site's robots.txt rules during crawling. |
| Max Depth | Maximum crawl depth from the seed URL, controlling how far links are followed. |
| Output as File | Whether to return a URL to the combined markdown file instead of inline content. |
Output
JSON
job- Metadata and status information about the crawl job.content_url- URL to the combined markdown file if 'Output as File' is enabled.markdown- Inline markdown content of the crawled pages if 'Output as File' is disabled.
Dependencies
- Requires an API key credential for WebCrawlerAPI to authenticate requests.
Troubleshooting
- Common errors include timeouts if the crawl job takes longer than 30 minutes, which can be resolved by increasing timeout settings or reducing crawl size.
- Errors from the API such as 'Crawl job failed' indicate issues with the crawl parameters or site accessibility; verify URL correctness and regex patterns.
- Invalid or missing API credentials will cause authentication errors; ensure the WebCrawlerAPI credential is correctly configured.
Links
- WebCrawlerAPI Documentation - Official documentation for the WebCrawlerAPI service used by this node.