Crawl4AI Plus Advanced icon

Crawl4AI Plus Advanced

Advanced web crawling and extraction with full Crawl4AI API control

CSS Extractor

AI-generated

Summary

Extract multiple structured data fields from repeating elements on a webpage using CSS selectors by crawling the specified URL with configurable browser and crawl settings.

Inputs

  • URL (required) — The URL of the webpage to extract content from.
  • Base Selector (required) — CSS selector identifying the repeating parent elements containing the data to extract (e.g., product items, article cards).
  • Fields (required) — A list of fields to extract relative to each base selector element. Each field requires a name, a relative CSS selector, a field type (text, HTML, or attribute), and if attribute type is chosen, the attribute name to extract.
  • Options — Options including whether to clean and normalize extracted text and whether to include the full original webpage text in the output.
  • Browser & Session — Configuration for the headless browser to use for crawling, including browser engine, cookies, JavaScript execution, stealth mode, HTTP headers, proxy, viewport size, user agent, session management, and other browser-related settings.
  • Crawl Settings — Advanced crawling options such as anti-bot strategies, caching behavior, respecting robots.txt, JavaScript code injection before extraction, wait conditions before extraction, content exclusion rules, retry limits, and delay settings.

Output shape

a list where each item represents one extracted element matching the base selector, containing the requested fields with extracted values as text, HTML, or attribute values.

The output shape is an array of objects keyed by the specified field names. Text extraction is cleaned and normalized by default but can be configured. The operation uses a headless browser to render the page (with JS enabled by default) to handle dynamic content and applies the given CSS selectors for structured extraction. Optional inclusion of the full original page text is supported. The operation respects advanced crawl and browser options to control spidering behavior, stealth, and session state.

Links

Discussion