Actions15
Regex Extractor
AI-generatedSummary
Extract content from a specified URL using regular expression patterns, supporting built-in, custom, or LLM-generated regexes for targeted data extraction.
Inputs
- url (required) — The URL to extract content from using regex patterns.
- patternType (required) — Choose pattern source: built-in patterns, custom patterns, LLM-generated patterns, or preset patterns like contact info or financial data.
- builtinPatterns — Select one or more built-in regex patterns for extracting common data types (used when patternType is 'builtin'). Defaults to extracting emails and URLs.
- customPatterns — Define multiple custom labeled regex patterns for extraction (used when patternType is 'custom'). Each requires a label and regex pattern.
- llmLabel — Label to identify the LLM-generated regex pattern (required when patternType is 'llm').
- llmQuery — Natural language description with examples guiding the LLM to generate the regex pattern (required when patternType is 'llm').
- llmSampleUrl — Sample URL containing the data to assist the LLM pattern generation (required when patternType is 'llm').
- options.cleanText — Option to normalize whitespace in all extracted string values.
- options.includeFullText — Option to include the original full webpage text in the output.
- browserSession — Browser session and configuration options such as browser type, cookies, headers, JavaScript enablement, stealth mode, headless mode, proxy, user agent, viewport size, and others to control page loading and retrieval.
- crawlSettings — Advanced crawl settings including anti-bot features, caching behavior, content extraction scope via CSS selector, delays, JavaScript code injection, retries, wait conditions, excluded tags, and word count thresholds.
Output shape
a JSON object containing extracted matches grouped by pattern labels along with metadata; optionally includes the original webpage content if requested.
The output includes keys for each regex pattern label mapping to arrays of extracted strings found on the target URL. If 'includeFullText' is enabled, the original HTML/text content is included in the output. The operation performs an HTTP fetch of the URL with full browser simulation according to the provided session and crawl settings, applies the selected or generated regex patterns to extract matching content, normalizes text if requested, and returns structured matches. Errors during fetching or extraction will be surfaced.
Examples
Example 1: Extract emails and URLs from a webpage using built-in patterns.
Set URL to target page, patternType to 'builtin', select 'Email' and 'URL' built-in patterns; optional cleanText to true.
Example 2: Extract price values from a page using a custom regex.
Set URL, choose patternType 'custom', provide a label and a regex pattern matching prices (e.g., $\s?\d+(?:,\d{3})*(?:.\d{2})?).
Example 3: Generate a regex to extract product IDs using LLM.
Set URL, patternType to 'llm', provide llmLabel (e.g., 'product_id'), descriptive llmQuery explaining format with examples, and sample URL containing the data.