Actions15
Cosine Similarity Extractor
AI-generatedSummary
Extract semantically similar content clusters from a webpage at a specified URL using cosine similarity on sentence embeddings, filtered by user-provided semantic keywords or topics.
Inputs
- URL (required) — The webpage URL to extract and cluster content from.
- Semantic Filter (required) — Keywords or topics defining the semantic focus for content filtering and clustering.
- Clustering Options — Settings to customize the clustering algorithm including linkage method, maximum cluster distance, embedding model name, similarity threshold, number of top clusters to return, verbosity for debugging, and minimum word count per content block.
- Options — Additional output options such as whether to normalize whitespace in extracted text and whether to include the full crawled markdown text in the output.
- Browser & Session — Configuration for the browser instance used to crawl the page, including browser engine, cookies, JavaScript enablement, stealth mode, headers, headless mode, proxy, viewport size, and other browser session details.
- Crawl Settings — Additional crawl parameters such as anti-bot measures, caching behavior, CSS selectors to restrict content extraction region, excluded HTML tags, JavaScript code injection before extraction, wait conditions before extraction, link exclusion, retry settings, and word count filtering.
Output shape
a JSON object containing clusters of semantically similar content extracted from the page, including cluster metadata and optionally full markdown text, depending on options set.
Requires the 'crawl4ai:all' Docker image due to dependency on torch and sentence-transformers. Output includes top K clusters ranked by semantic similarity score above the specified threshold. Optional verbose logging is available for debugging cluster formation. The extraction respects crawl and browser settings provided, and supports filtering and cleaning of extracted data.