ScrapegraphAI icon

ScrapegraphAI

Turn any webpage into usable data with the ScrapeGraphAI v2 API — scrape, extract, search, crawl, monitor, history, credits.

Actions19

Crawl → Start

AI-generated

Overview

This node operation initiates a web crawling process starting from a specified URL. It is designed to traverse web pages up to a defined depth and number of pages, capturing various types of content formats such as HTML, markdown, images, JSON, links, screenshots, summaries, and branding information. The node supports advanced crawling configurations including link-following rules, content type filtering, and fetch options like cookies, headers, proxy settings, and JavaScript rendering modes. This operation is useful for scenarios like data extraction from websites, content aggregation, SEO analysis, and automated web monitoring.

Use Case Examples

  1. Starting a crawl from a blog homepage to collect all markdown content and images up to 3 levels deep.
  2. Crawling an e-commerce site to extract product details in JSON format while limiting the crawl to 100 pages.
  3. Capturing screenshots of a news website's front page and linked articles for visual monitoring.

Properties

Name Meaning
URL Starting URL to crawl, required to initiate the crawl.
Formats Specifies the output formats to capture from each page, such as markdown, HTML, images, JSON, links, screenshots, summaries, and branding. Each format can have additional options like mode, prompt, schema, full page capture, viewport size, and image quality.
Max Pages Maximum number of pages to crawl during the operation.
Max Depth Maximum levels of links to follow from the starting URL.
Max Links per Page Limit on the number of links expanded per page; zero means unlimited.
Allow External Links Whether to follow links to other domains; defaults to false (same-origin only).
Include Patterns Glob-style URL patterns to include in the crawl.
Exclude Patterns Glob-style URL patterns to exclude from the crawl.
Content Types Optional filter to limit crawled pages to specified MIME types.
Fetch Config Settings controlling how the page is fetched before processing, including cookies, country for geo-targeted proxy, custom headers, fetch mode (auto, fast, JS rendering), scroll count for infinite scroll, stealth mode for anti-bot measures, request timeout, and wait time after page load.

Output

JSON

  • crawlId - Identifier for the initiated crawl process.
  • status - Current status of the crawl operation.
  • pagesCrawled - Number of pages successfully crawled so far.
  • results - Collected data from the crawl in the requested formats.

Dependencies

  • Requires an API key credential for ScrapeGraphAI service to authenticate requests.

Troubleshooting

  • Common issues include invalid or unreachable starting URL, exceeding maximum pages or depth limits, and misconfigured fetch options causing timeouts or incomplete data.
  • Error messages may indicate unsupported resource-operation combinations, network errors, or API quota limits. Verify URL correctness, adjust crawl limits, and check API credentials to resolve.

Links

Discussion