Crawl4AI Plus Advanced icon

Crawl4AI Plus Advanced

Advanced web crawling and extraction with full Crawl4AI API control

Process Raw HTML

AI-generated

Summary

Process the provided raw HTML content to extract structured data, apply optional content filtering, table extraction, and output formatting without performing any web crawling or browser navigation.

Inputs

  • HTML Content (required) — The raw HTML string to be processed, with a maximum size of 5MB.
  • Base URL — The base URL used to resolve relative links within the HTML content.
  • Crawl Settings — Options to control processing behavior such as anti-bot handling, cache mode, CSS selectors for limiting extraction, JavaScript execution on the page, delay before returning, wait conditions, and filtering of links or tags.
  • Output & Filtering — Settings for content filtering (none, BM25, LLM, pruning), table extraction (none, default, LLM), markdown output format, inclusion of raw HTML, links, media, screenshots, PDFs, and verbosity controls. Includes options for thresholds and instructions for LLM-based filtering.

Output shape

a single structured JSON object containing extracted content, optionally including filtered markdown, raw markdown, extracted tables, links, media, screenshots and PDF data (base64 encoded), and metadata about the processed HTML

The output respects the configuration of content filters and table extraction methods, including the use of LLM-powered or statistical approaches where credentials are provided. Screenshot and PDF options are ignored as no browser navigation occurs. The node does not fetch or navigate to URLs but works entirely with the provided raw HTML input.

Links

Discussion