Actions10
- Class Actions
- Classification Actions
- Document Actions
- Extraction Actions
- Schema Actions
Extraction → Upload and Extract
AI-generatedSummary
Upload a document provided as a URL, binary file, or base64 string and extract structured data from it using a specified schema in a single operation.
Inputs
- Input Mode (required) — Select how to provide the document: via a public URL, from a binary property of a previous node, or as a base64-encoded string.
- File URL (required) — If 'URL' mode is selected, provide a publicly accessible direct URL to the document file. Supports formats like PDF, PNG, JPEG, TIFF, DOCX, XLSX, CSV, etc.
- Filename — Optional name of the file including extension when using 'URL' or 'Base64' modes. Used for format detection if provided.
- Binary Property (required) — If 'Binary File' mode is selected, specify the name of the binary property of the input item containing the document file, typically 'data'.
- Base64 Content (required) — If 'Base64' mode is selected, provide the document content as a base64-encoded string, which can be referenced from previous nodes.
- Base64 Filename — Optional filename with extension when providing base64 content, useful for certain file types for format detection.
- Schema (required) — Select or enter the ID of a DocuPipe schema defining which fields to extract from the document (e.g. invoice number, amount, date).
- API Version (required) — Choose the extraction engine version to use: V3 (latest, recommended) or V2 (legacy).
- Additional Fields (V3) — Optional parameters for V3 engine such as dataset grouping, effort level (currently only standard), and extraction guidelines.
- Advanced Options (V3) — Optional advanced settings for V3 including metadata JSON to attach and feed into extraction, specific pages to extract, timeout settings, and whether to use metadata in extraction context.
- Additional Fields (V2) — Optional parameters for V2 engine including dataset grouping, display mode, effort level, guidelines, and split mode of the document.
- Advanced Options (V2) — Optional advanced settings for V2 including metadata JSON to attach and feed into extraction, pages to extract, timeout, and metadata usage flag.
Output shape
a single object representing the uploaded document and extraction initiation response
The response JSON includes uploaded document details and extraction operation metadata. The extraction runs asynchronously; subsequent nodes or triggers can fetch extraction results using the document or extraction identifiers.
Examples
Example 1: Uploading a PDF invoice from a public URL and extracting invoice data using a specified schema.
Set Input Mode to 'URL', provide the File URL, select the Schema to extract fields, choose API Version 'V3', optionally configure additional fields for dataset and guidelines, and execute.