PDF Vector icon

PDF Vector

Convert PDFs, Word, Excel documents, and images to clean markdown, extract structured data with AI, process invoices with specialized parsing, and search millions of academic papers across PubMed, ArXiv, Google Scholar, and more.

Actions18

Identity Document → Extract

AI-generated

Overview

This node extracts structured data from identity documents using AI-powered processing. It supports input via a public URL or binary data from a previous node. Users provide an extraction prompt and a JSON schema defining the desired data structure. The node processes various document formats including PDF, DOCX, XLSX, CSV, PNG, and JPG, making it useful for automating data extraction from identity documents such as passports, ID cards, or driver licenses.

Use Case Examples

  1. Extracting name, date of birth, and document number from a scanned passport image.
  2. Extracting structured identity information from a PDF identity document provided via URL.

Properties

Name Meaning
Input Type How to provide the identity document, either via a public URL or binary data from a previous node.
Identity Document URL Public URL of the identity document file (PDF, DOCX, XLSX, CSV, PNG, JPG). Required if Input Type is URL.
Input Binary Field Name of the binary property from a previous node containing the file. Required if Input Type is Binary Data.
Model AI model tier to use for extraction. Higher tiers produce better results but cost more credits.
Extraction Prompt Instructions for what data to extract from the identity document. Minimum 4 characters.
JSON Schema JSON Schema defining the structure of the data to extract. Used to specify the expected output format.
Document ID Optional identifier for usage tracking. Returned in the response.

Output

JSON

  • title - Extracted title or main identifier from the identity document.
  • summary - Extracted summary or additional structured data from the identity document.

Dependencies

  • An API key credential for the PDF Vector API

Troubleshooting

  • Ensure the input document URL is publicly accessible if using URL input type.
  • Verify the binary data field name matches the output of the previous node if using binary input.
  • The extraction prompt must be at least 4 characters long to be valid.
  • The JSON schema must be correctly formatted to match the expected output structure.
  • API errors may include HTTP status codes and messages; check credentials and API limits if errors occur.

Links

Discussion