Actions15
Crawl URL
AI-generatedSummary
Crawl a single URL using a configurable headless browser session with advanced options for browser type, JavaScript execution, stealth mode, proxy, and session persistence. Extract content according to customizable crawl settings, apply content filtering and table extraction strategies, and return structured data including markdown, links, tables, media, and optional outputs like screenshots and PDFs.
Inputs
- url (required) — The URL to crawl
- browserSession — Settings that configure the browser engine, cookies, JavaScript, stealth mode, proxy, user agent, viewport size, timeouts, and session persistence
- crawlSettings — Options controlling crawling behavior such as anti-bot measures, cache usage, robots.txt respect, CSS selectors to limit extraction scope, custom JavaScript to run on the page, wait conditions, and retries
- outputFiltering — Settings to control output format, including content filtering strategies (none, pruning, BM25, LLM-based), table extraction options, inclusion of screenshots/PDFs/HTML, and markdown output type
Output shape
a single structured object representing the crawled page data
The output object can include extracted markdown content (raw and filtered), discovered internal and external links, included tables and media, optional screenshot and PDF in base64, SSL certificate info, and verbose debug information if enabled. Outputs are filtered and formatted based on the node parameters for content filtering, table extraction, and output preferences.
Examples
Example 1: Crawl a webpage with stealth browser settings and return filtered markdown and links.
URL set to target page; Browser & Session with stealth mode enabled; Crawl Settings with default anti-bot disabled; Output & Filtering with pruning content filter and include links enabled.