Crawl4AI Plus Advanced icon

Crawl4AI Plus Advanced

Advanced web crawling and extraction with full Crawl4AI API control

LLM Extractor

AI-generated

Summary

Extract structured data from a webpage URL using a large language model (LLM) configured with a user-defined extraction schema and instructions.

Inputs

  • URL (required) — The web page URL from which to extract content.
  • Extraction Instructions (required) — Text instructions guiding the LLM on what specific information to extract from the page content.
  • Schema Input Mode (required) — Choose either 'Simple Fields' for individual field definitions or 'Advanced JSON' to provide a JSON schema defining the data structure to extract.
  • Schema Fields (required) — When using 'Simple Fields' mode, provide one or more fields specifying the name, type, description, and whether each is required for extraction.
  • JSON Schema (required) — When using 'Advanced JSON' mode, provide a valid JSON schema describing the expected data properties and required fields for extraction.
  • LLM Options — Customize LLM input format (Markdown, HTML, or cleaned Markdown), maximum tokens for the response, and temperature controlling randomness of outputs.
  • Options — Controls array handling strategies in extracted data, whether to clean extracted text, include original webpage text, and include metadata with split array items.
  • Browser & Session — Configure browser engine, session details, cookies, JavaScript execution, stealth mode, headers, proxy, viewport size, user agent, headless mode, timeout, and related browser behaviors for fetching the page.
  • Crawl Settings — Additional settings such as anti-bot modes, cache mode, content filtering with CSS selectors or excluded tags, JavaScript execution code before extraction, wait conditions, retry limits, and filtering thresholds.

Output shape

a single structured object matching the user-defined schema with extracted field values from the target webpage, optionally including original text and metadata

The output format respects the extraction schema specified either via simple fields or advanced JSON. Array handling options influence whether arrays are split into multiple items. Original text and metadata inclusion is optional based on configuration. Errors or incomplete extractions occur if required fields are not found or if extraction fails.

Links

Discussion