WebCrawlerAPI icon

WebCrawlerAPI

Scrape or crawl web pages with WebCrawlerAPI

Scrape

AI-generated

Overview

This node performs web scraping on a single web page using the WebCrawlerAPI. It sends a request to the API with the specified URL and desired output format, then returns the scraped content in the chosen format. This is useful for extracting content from web pages for further processing or analysis, such as converting web content into markdown, cleaned text, HTML, or extracting links.

Use Case Examples

  1. Scrape a news article URL and get the content in markdown format for use in a content management system.
  2. Extract all links from a product page to analyze outbound links.
  3. Retrieve cleaned text from a blog post for sentiment analysis.

Properties

Name Meaning
URL to Scrape The URL of the web page to scrape content from.
Output Format The format in which the scraped content should be returned. Options include Markdown, Cleaned text, HTML, or Links.

Output

JSON

  • json
    • success - Indicates if the scraping request was successful.
    • status - Status of the scraping operation.
    • data - The scraped content in the requested format.
    • error_message - Error message if the scraping failed.

Dependencies

  • Requires an API key credential for WebCrawlerAPI to authenticate requests.

Troubleshooting

  • Common errors include network issues or invalid URLs causing the API request to fail.
  • If the API returns an error status, check the error_message field for details.
  • Timeouts or rate limits from the API may cause failures; ensure proper API usage limits are respected.

Links

Discussion