ScrapegraphAI icon

ScrapegraphAI

Turn any webpage into usable data with the ScrapeGraphAI v2 API — scrape, extract, search, crawl, monitor, history, credits.

Actions19

Extract → Extract

AI-generated

Overview

This node extracts structured data from various content sources using natural language prompts and optional JSON schema constraints. It supports extracting from a URL, raw HTML, or Markdown content. The node is useful for scenarios where you want to pull specific information from web pages or text content, such as extracting product details, article summaries, or any custom data defined by a prompt.

Use Case Examples

  1. Extract product name and price from an e-commerce product page URL.
  2. Extract key information from raw HTML content of a news article.
  3. Extract structured data from Markdown content using a natural language prompt.

Properties

Name Meaning
Source Specifies the content source to extract data from, options include URL, raw HTML, or Markdown.
URL Public URL of the page to extract from, required if source is URL.
HTML Raw HTML content to extract from, required if source is HTML.
Markdown Markdown content to extract from, required if source is Markdown.
Prompt Natural-language description of what to extract from the content.
Use JSON Schema Whether to constrain the extraction output to a specified JSON schema.
Schema (JSON) JSON schema describing the desired output shape, used if 'Use JSON Schema' is true.
HTML Mode Pre-processing mode applied to the source HTML, applicable if source is URL or HTML.
Fetch Config Options controlling how the page is fetched when source is URL, including cookies, headers, country for proxy routing, fetch mode, scrolls, stealth mode, timeout, and wait time after load.
Output Shape of the response returned by the extraction operation.
Fields Comma-separated list of top-level fields to keep in the output, used if output mode is 'Selected Fields'.

Output

JSON

  • id - Unique identifier of the extraction result.
  • json - Extracted data in JSON format according to the prompt and optional schema.
  • usage - Information about resource usage or API call details.

Dependencies

  • Requires an API key credential for ScrapeGraphAI service to perform extraction operations.

Troubleshooting

  • Ensure the URL is publicly accessible if using URL source; private or protected pages may fail to fetch.
  • If extraction results are empty or incorrect, verify the prompt clearly describes the desired data.
  • When using JSON schema, ensure the schema is valid and matches the expected output structure to avoid parsing errors.
  • Timeouts or fetch failures can occur if the target page is slow or blocks automated requests; adjust fetch config such as timeout, stealth mode, or fetch mode accordingly.

Links

Discussion