PDF Vector icon

PDF Vector

Convert PDFs, Word, Excel documents, and images to clean markdown, extract structured data with AI, process invoices with specialized parsing, and search millions of academic papers across PubMed, ArXiv, Google Scholar, and more.

Actions18

Document → Extract

AI-generated

Overview

The node extracts structured data from documents such as PDFs, Word, Excel, CSV, and images using AI models. It supports input via a public URL or binary data from a previous node. Users provide an extraction prompt and a JSON schema defining the data structure to extract. This node is useful for automating data extraction from various document types, such as extracting titles, authors, publication years, or other custom data fields from documents.

Use Case Examples

  1. Extracting metadata like title and authors from academic papers in PDF format.
  2. Extracting invoice details such as invoice number and total amount from scanned images or PDFs.
  3. Extracting structured data from bank statements or identity documents for automated processing.

Properties

Name Meaning
Input Type How to provide the document, either via a public URL or binary data from a previous node.
Document URL Public URL of the document file (PDF, DOCX, XLSX, CSV, PNG, JPG). Required if Input Type is URL.
Input Binary Field Name of the binary property from a previous node containing the file. Required if Input Type is Binary Data.
Model AI model tier to use for extraction. Higher tiers produce better results but cost more credits.
Extraction Prompt Instructions for what data to extract from the document. Minimum 4 characters.
JSON Schema JSON Schema defining the structure of the data to extract. Used to specify the expected output format.
Document ID Optional identifier for usage tracking. Returned in the response.

Output

JSON

  • title - Extracted title from the document as defined in the JSON schema.
  • summary - Extracted summary or other structured data from the document as defined in the JSON schema.
  • documentId - Optional document identifier returned for usage tracking.

Dependencies

  • An API key credential for the PDF Vector API

Troubleshooting

  • Common issues include providing an invalid or inaccessible document URL, incorrect binary property name, or malformed JSON schema.
  • Errors related to API authentication indicate missing or invalid API credentials.
  • Extraction prompt too short or unclear may result in incomplete or incorrect data extraction.
  • File size or page limits exceeded for the selected AI model tier may cause errors.

Links

Discussion