Actions15
Stream Crawl
AI-generatedSummary
Stream crawl multiple URLs using an advanced browser session with customizable settings, retrieving comprehensive crawl data including extracted content, metadata, tables, media, and more, optionally filtered and formatted per user configuration.
Inputs
- URLs (required) — One or more URLs to crawl via the streaming endpoint, provided as newline or comma-separated strings; each URL produces an output item.
- Browser & Session — Collection of browser and session parameters including browser engine type (Chromium, Firefox, WebKit), cookies, JavaScript enabling, stealth mode, custom headers, proxy configuration, user agent, viewport size, session persistence, and other browser behavior controls.
- Crawl Settings — Options to control crawl behavior such as anti-bot features (magic mode, user simulation), cache mode, robots.txt respect, CSS selectors for content scope, delay before returning HTML, exclusion of external links, JavaScript page scripts to run before extraction, wait conditions before extraction, retry logic, and word count filtering.
- Output & Filtering — Configuration for output format and filtering, including content filtering strategies (none, BM25, pruning, LLM), table extraction methods (none, default, LLM), markdown output format, inclusion of screenshots, PDFs, raw HTML, links, media, and verbosity for debugging.
Output shape
a list of output items, one per input URL, each containing crawl results including filtered content in markdown and raw formats, extracted tables and media, screenshots or PDFs if enabled, and metadata about the crawl; possible inclusion of debug info based on verbosity settings.
The streaming crawl aggregates results from a streaming API endpoint, buffering all NDJSON chunks before returning a complete combined result per URL. Parsing errors in stream chunks are surfaced on the last output item to inform users. Optional content filtering and table extraction require proper credentials and configuration. Each URL input produces a separate output item with its own crawl data.