Crawl4AI Plus icon

Crawl4AI Plus

Crawl web pages, extract data, and ask questions using Crawl4AI

Extract with CSS Selectors

AI-generated

Summary

Extracts structured data from a web page at a specified URL using CSS selectors by targeting repeating elements and defined fields within them.

Inputs

  • URL (required) — The web page URL from which to extract content
  • Base Selector (required) — CSS selector identifying the repeating element (e.g., product items) containing the desired data
  • Fields (required) — A list of fields to extract for each repeated element, each defined by a name, relative CSS selector, extraction type (text, HTML, or attribute), and optionally an attribute name if extracting an attribute value
  • Options — Optional settings including cache mode, whether to clean extracted text by normalizing whitespace, whether to include the original webpage text in the output, and a CSS selector or JavaScript expression to wait for before extracting content

Output shape

a list of extracted item objects with the specified fields, along with metadata including the domain, URL, item count, success status, fetch timestamp, and crawl metrics

The node returns an array of results corresponding to crawled pages (usually one for single-page extraction). Each result contains the array of extracted items matching the base CSS selector and fields defined. Text cleaning is applied by default but can be disabled. Optionally, the original webpage text can be included in the output. Errors are reported per item, and pagination is handled internally via crawl scope configuration if needed (though this operation is focused on single URL extraction). Cache mode controls usage of caching during crawling.

Links

Discussion