Get Page Content
AI-generatedSummary
Crawl the specified URL (and optionally linked pages) to extract and return combined page content as markdown along with metadata such as domain, URLs crawled, links, and crawl timing.
Inputs
- URL (required) — The URL to crawl and extract content from
- Crawl Scope — How extensively to crawl from the starting URL: singlePage (only the specified URL), followLinks (discovered links, depth 1), or fullSite (entire site recursively, depth 3). Defaults to singlePage.
- Options — Additional crawl options including Cache Mode (how to use cache), Content Quality (clean filtered markdown or complete raw markdown), CSS Selector (limit extraction to a specific element), Exclude URL Patterns (patterns to exclude from crawling when multi-page), Include HTML (whether to include raw HTML), Include Links (whether to include structured internal and external links), Max Pages (max pages to crawl for multi-page), and Wait For (CSS selector or JS expression to wait for before extracting content).
Output shape
a single JSON object containing combined markdown content and crawl metadata
Returns an object with domain, url, optionally array of urls crawled, combined markdown content (filtered or complete), optionally included HTML, internal and external links (if enabled), crawl success status, number of pages scanned, fetch timestamp, and crawl metrics. Multi-page results are merged and deduplicated. If no results are returned, an error JSON object is returned instead.
Examples
Example 1: Crawl a single page and get clean filtered content
Set URL to the page, use default Crawl Scope (singlePage), set Content Quality to Clean, and optionally enable Include Links to get page links in output.
Example 2: Crawl entire site to extract full markdown including HTML and links
Set URL, Crawl Scope to Full Site, Content Quality to Complete, enable Include HTML and Include Links, adjust Max Pages and Exclude URL Patterns as needed to control crawl limits and exclude administrative or login pages.