Crawl4AI Plus Advanced icon

Crawl4AI Plus Advanced

Advanced web crawling and extraction with full Crawl4AI API control

JSON Extractor

AI-generated

Summary

Extract JSON data from a specified URL using configurable extraction methods depending on whether the JSON is directly returned, embedded in a script tag, or structured as JSON-LD; supports targeted extraction with a dot-separated JSON path and browser session customization.

Inputs

  • URL (required) — The URL of the JSON data to extract.
  • JSON Path — Dot-separated path to the JSON data within the response, supporting array indices; leave empty to extract the entire JSON response.
  • Source Type (required) — Specifies where to find the JSON data on the page: direct JSON response, inside a script tag, or JSON-LD structured data.
  • Extraction Schema Type — Schema extraction engine (CSS selector or XPath) used when the source type is Script Tag or JSON-LD.
  • Script Selector — CSS selector or XPath expression to locate the script tag containing JSON data (used only when source type is Script Tag).
  • Options (Headers) — Optional HTTP headers in JSON format to send with the request.
  • Options (Include Full Text) — Whether to include the full crawled text in the output.
  • Browser & Session — Collection of settings to configure the browser environment, including browser type, cookies, JavaScript execution, stealth mode, user agent, headless mode, proxies, viewport size, and other browser behaviors.
  • Crawl Settings — Collection of crawl-related options such as anti-bot handling (magic mode, simulating user, overriding navigator), cache mode, delays, JavaScript execution on page before extraction, wait conditions, content filtering, and retries.

Output shape

a JSON object or array extracted from the target URL according to the configured JSON path and extraction method, optionally including full page text if requested.

  • Extracted JSON data is parsed and filtered by the optional JSON path if provided.
  • For 'direct' source type, expects the URL to return raw JSON.
  • For 'script' or 'jsonld' source types, uses specified extraction schema (CSS or XPath) and script selectors to locate embedded JSON.
  • Browser and crawl options enable fine control of page loading and anti-bot techniques.
  • The output structure matches the extracted JSON, or is null if extraction fails or no valid JSON is found.

Links

Discussion