Actions19
Extract → Extract
AI-generatedOverview
This node extracts structured data from various content sources using natural language prompts and optional JSON schema constraints. It supports extracting from a URL, raw HTML, or Markdown content. The node is useful for scenarios where you want to pull specific information from web pages or text content, such as extracting product details, article summaries, or any custom data defined by a prompt.
Use Case Examples
- Extract product name and price from an e-commerce product page URL.
- Extract key information from raw HTML content of a news article.
- Extract structured data from Markdown content using a natural language prompt.
Properties
| Name | Meaning |
|---|---|
| Source | Specifies the content source to extract data from, options include URL, raw HTML, or Markdown. |
| URL | Public URL of the page to extract from, required if source is URL. |
| HTML | Raw HTML content to extract from, required if source is HTML. |
| Markdown | Markdown content to extract from, required if source is Markdown. |
| Prompt | Natural-language description of what to extract from the content. |
| Use JSON Schema | Whether to constrain the extraction output to a specified JSON schema. |
| Schema (JSON) | JSON schema describing the desired output shape, used if 'Use JSON Schema' is true. |
| HTML Mode | Pre-processing mode applied to the source HTML, applicable if source is URL or HTML. |
| Fetch Config | Options controlling how the page is fetched when source is URL, including cookies, headers, country for proxy routing, fetch mode, scrolls, stealth mode, timeout, and wait time after load. |
| Output | Shape of the response returned by the extraction operation. |
| Fields | Comma-separated list of top-level fields to keep in the output, used if output mode is 'Selected Fields'. |
Output
JSON
id- Unique identifier of the extraction result.json- Extracted data in JSON format according to the prompt and optional schema.usage- Information about resource usage or API call details.
Dependencies
- Requires an API key credential for ScrapeGraphAI service to perform extraction operations.
Troubleshooting
- Ensure the URL is publicly accessible if using URL source; private or protected pages may fail to fetch.
- If extraction results are empty or incorrect, verify the prompt clearly describes the desired data.
- When using JSON schema, ensure the schema is valid and matches the expected output structure to avoid parsing errors.
- Timeouts or fetch failures can occur if the target page is slow or blocks automated requests; adjust fetch config such as timeout, stealth mode, or fetch mode accordingly.
Links
- ScrapeGraphAI Extract API Documentation - Official API documentation for the extract operation, detailing parameters and response formats.