Extract Data
AI-generatedSummary
Extract structured data from one or more web pages by crawling the specified URL(s) according to the chosen scope and extraction type.
Inputs
- URL (required) — The URL to extract data from
- Extraction Type (required) — Choose the type of data to extract: Contact Info (emails, phones, social media, addresses), Financial Data (currencies, credit cards, IBANs, percentages), or Custom (LLM) for custom extraction using natural language instructions.
- Extraction Instructions (required) — A natural language prompt describing the data to extract. Required only for Custom (LLM) extraction type.
- Schema Fields — Define custom fields specifying name, data type, and description for structured output. Required only for Custom (LLM) extraction type.
- Crawl Scope (required) — Determines how many pages to crawl and extract data from: Single Page, Follow Links (depth 1), or Full Site (depth 3).
- Options — Additional crawling options including cache mode, URL patterns to exclude from crawling (for multi-page crawls), maximum number of pages to crawl, and a CSS selector or JavaScript expression to wait for before extracting content.
Output shape
a list of JSON items each containing the domain, URL(s) crawled, extraction type, extracted data, extraction success flags, number of pages scanned, timestamps, and crawl metrics
The results vary based on extraction type: for Contact Info and Financial Data, data is extracted via regex without LLMs; for Custom (LLM), extraction uses a language model with user-provided instructions and schema, requiring LLM credentials. Multi-page crawls deduplicate URLs. Parse warnings may appear if custom extraction content cannot be parsed as JSON. Supports error handling with optional continuation on failure.
Examples
Example 1: Extract emails and phone numbers from a single webpage
Set URL to target page, Extraction Type to 'Contact Info', Crawl Scope to 'Single Page'. Optional: adjust cache mode and waitFor options.
Example 2: Extract custom product data across linked pages
Set URL to product category page, Extraction Type to 'Custom (LLM)', provide extraction instructions describing desired data, define schema fields for output, and choose Crawl Scope as 'Follow Links' or 'Full Site'. Configure LLM credentials in node settings.