PDF Vector icon

PDF Vector

Convert PDFs, Word, Excel documents, and images to clean markdown, extract structured data with AI, process invoices with specialized parsing, and search millions of academic papers across PubMed, ArXiv, Google Scholar, and more.

Actions18

Identity Document → Parse

AI-generated

Overview

This node parses identity documents using the PDF Vector API. It supports input as either a public URL to the document file or binary data from a previous node. The node uses AI models to extract structured data from various document formats such as PDF, DOCX, XLSX, CSV, PNG, and JPG. It is useful for automating identity verification, data extraction from IDs, and document processing workflows.

Use Case Examples

  1. Parsing a scanned passport image provided as binary data to extract personal details.
  2. Providing a URL to a driver's license PDF to automatically extract and structure the identity information.

Properties

Name Meaning
Input Type Specifies how the identity document is provided to the node, either via a public URL or binary data from a previous node.
Identity Document URL The public URL of the identity document file, required if input type is URL.
Input Binary Field The name of the binary property containing the file, required if input type is binary data.
Model AI model tier for parsing the document, with options for automatic selection, standard accuracy, or highest accuracy with HTML output.
Document ID Optional identifier for usage tracking, returned in the response.

Output

JSON

  • documentId - The optional identifier provided for usage tracking, returned in the response.
  • parsedData - The structured data extracted from the identity document by the AI model.
  • modelUsed - The AI model tier used for parsing the document.
  • rawResponse - The full raw response from the PDF Vector API containing detailed parsing results.

Dependencies

  • PDF Vector API with an API key credential

Troubleshooting

  • Ensure the input document URL is publicly accessible if using URL input type.
  • Verify the binary data field name matches the output from the previous node if using binary input.
  • Check API credentials and domain configuration to avoid authentication errors.
  • Common errors include invalid file format, unsupported document type, or exceeding API usage limits.

Links

Discussion