Actions19
Scrape → Scrape
AI-generatedOverview
This node performs web scraping on a specified public URL using the ScrapeGraphAI v2 API. It allows users to extract various types of data from web pages, such as HTML content, images, links, JSON data, markdown, screenshots, summaries, and branding information. The node supports different pre-processing modes for HTML content and offers customization options for fetching the page, including cookies, headers, geo-targeting, JavaScript rendering, scrolling, stealth mode, timeouts, and wait times. It is useful for automating data extraction from websites for analysis, monitoring, or integration into workflows.
Use Case Examples
- Extract product details and prices from an e-commerce page by specifying a markdown format with a custom prompt.
- Capture a full-page screenshot of a webpage for visual documentation.
- Retrieve all links from a news article to analyze referenced sources.
- Extract structured JSON data from a webpage using a defined JSON schema.
Properties
| Name | Meaning |
|---|---|
| URL | Public URL of the page to fetch, required for scraping. |
| Formats | Output formats to capture from the page, supporting multiple types such as markdown, HTML, images, JSON, links, screenshot, summary, and branding, each with specific options like mode, prompt, schema, and screenshot settings. |
| Content Type | Override the auto-detected content type of the fetched page, useful for non-HTML content like PDFs. |
| Fetch Config | Controls how the page is fetched before processing, including cookies, country for geo-targeting, custom headers, fetch mode (auto, fast, JS), number of scrolls for infinite scroll content, stealth mode for anti-bot measures, request timeout, and wait time after page load. |
| Output | Shape of the response returned by the operation, with options for simplified, raw, or selected fields output. |
| Fields | Comma-separated list of top-level fields to keep in the output when 'Selected Fields' output mode is chosen. |
Output
JSON
id- Unique identifier of the scrape result.json- Main content of the scrape result in JSON format, structure depends on the requested output formats.
Dependencies
- Requires an API key credential for the ScrapeGraphAI v2 API to authenticate requests.
Troubleshooting
- Common issues include unsupported resource-operation combinations, which result in an error indicating the combination is not supported.
- Errors related to network issues, invalid URLs, or API limits may occur during the fetch process.
- If the output format or fields are incorrectly specified, the node may return incomplete or unexpected data.
- To resolve errors, ensure the URL is publicly accessible, the API key is valid, and the fetch configuration matches the target website's requirements.