Actions18
- Academic Actions
- Bank Statement Actions
- Document Actions
- Identity Document Actions
- Invoice Actions
Document → Extract
AI-generatedOverview
The node extracts structured data from documents such as PDFs, Word, Excel, CSV, and images using AI models. It supports input via a public URL or binary data from a previous node. Users provide an extraction prompt and a JSON schema defining the data structure to extract. This node is useful for automating data extraction from various document types, such as extracting titles, authors, publication years, or other custom data fields from documents.
Use Case Examples
- Extracting metadata like title and authors from academic papers in PDF format.
- Extracting invoice details such as invoice number and total amount from scanned images or PDFs.
- Extracting structured data from bank statements or identity documents for automated processing.
Properties
| Name | Meaning |
|---|---|
| Input Type | How to provide the document, either via a public URL or binary data from a previous node. |
| Document URL | Public URL of the document file (PDF, DOCX, XLSX, CSV, PNG, JPG). Required if Input Type is URL. |
| Input Binary Field | Name of the binary property from a previous node containing the file. Required if Input Type is Binary Data. |
| Model | AI model tier to use for extraction. Higher tiers produce better results but cost more credits. |
| Extraction Prompt | Instructions for what data to extract from the document. Minimum 4 characters. |
| JSON Schema | JSON Schema defining the structure of the data to extract. Used to specify the expected output format. |
| Document ID | Optional identifier for usage tracking. Returned in the response. |
Output
JSON
title- Extracted title from the document as defined in the JSON schema.summary- Extracted summary or other structured data from the document as defined in the JSON schema.documentId- Optional document identifier returned for usage tracking.
Dependencies
- An API key credential for the PDF Vector API
Troubleshooting
- Common issues include providing an invalid or inaccessible document URL, incorrect binary property name, or malformed JSON schema.
- Errors related to API authentication indicate missing or invalid API credentials.
- Extraction prompt too short or unclear may result in incomplete or incorrect data extraction.
- File size or page limits exceeded for the selected AI model tier may cause errors.
Links
- PDF Vector API Reference - Official documentation for all PDF Vector API operations and models.
- JSON Schema Editor - Tool to visually build JSON schemas for defining extraction output structure.