Crawl4AI Plus icon

Crawl4AI Plus

Crawl web pages, extract data, and ask questions using Crawl4AI

Get Page Content

AI-generated

Summary

Crawl the specified URL (and optionally linked pages) to extract and return combined page content as markdown along with metadata such as domain, URLs crawled, links, and crawl timing.

Inputs

  • URL (required) — The URL to crawl and extract content from
  • Crawl Scope — How extensively to crawl from the starting URL: singlePage (only the specified URL), followLinks (discovered links, depth 1), or fullSite (entire site recursively, depth 3). Defaults to singlePage.
  • Options — Additional crawl options including Cache Mode (how to use cache), Content Quality (clean filtered markdown or complete raw markdown), CSS Selector (limit extraction to a specific element), Exclude URL Patterns (patterns to exclude from crawling when multi-page), Include HTML (whether to include raw HTML), Include Links (whether to include structured internal and external links), Max Pages (max pages to crawl for multi-page), and Wait For (CSS selector or JS expression to wait for before extracting content).

Output shape

a single JSON object containing combined markdown content and crawl metadata

Returns an object with domain, url, optionally array of urls crawled, combined markdown content (filtered or complete), optionally included HTML, internal and external links (if enabled), crawl success status, number of pages scanned, fetch timestamp, and crawl metrics. Multi-page results are merged and deduplicated. If no results are returned, an error JSON object is returned instead.

Examples

Example 1: Crawl a single page and get clean filtered content

Set URL to the page, use default Crawl Scope (singlePage), set Content Quality to Clean, and optionally enable Include Links to get page links in output.

Example 2: Crawl entire site to extract full markdown including HTML and links

Set URL, Crawl Scope to Full Site, Content Quality to Complete, enable Include HTML and Include Links, adjust Max Pages and Exclude URL Patterns as needed to control crawl limits and exclude administrative or login pages.

Links

Discussion