Actions18
- Academic Actions
- Bank Statement Actions
- Document Actions
- Identity Document Actions
- Invoice Actions
Document → Parse
AI-generatedOverview
The node parses documents such as PDFs, Word, Excel, CSV, and images to extract clean markdown and structured data using AI models. It supports input via a public URL or binary data from a previous node. Users can select different AI model tiers based on document complexity and size, and optionally include page-separated markdown in the output. This node is useful for automating document processing, data extraction, and content conversion tasks.
Use Case Examples
- Parsing a PDF invoice to extract structured data for accounting automation.
- Converting a Word document into markdown for content management.
- Extracting tables from an Excel file for data analysis.
Properties
| Name | Meaning |
|---|---|
| Input Type | How to provide the document, either via a public URL or binary data from a previous node. |
| Document URL | Public URL of the document file (PDF, DOCX, XLSX, CSV, PNG, JPG). Required if Input Type is URL. |
| Input Binary Field | Name of the binary property from a previous node containing the file. Required if Input Type is Binary Data. |
| Model | AI model tier for parsing. Higher tiers support more complex documents but cost more credits. |
| Include Pages | Whether to include page-separated markdown in the pages array. The full document markdown is still returned. |
| Document ID | Optional identifier for usage tracking. Returned in the response. |
Output
JSON
markdown- The full document content converted to markdown format.pages- An array of page-separated markdown content if 'Include Pages' is enabled.documentId- The optional identifier provided for usage tracking, returned in the response.
Dependencies
- An API key credential for the PDF Vector API
Troubleshooting
- Common issues include invalid or inaccessible document URLs, unsupported file formats, or exceeding size/page limits of the selected AI model tier.
- Error messages may indicate network issues, authentication failures, or input validation errors. Ensure the API key is valid and the input document meets the format and size requirements.
Links
- PDF Vector API Reference - Official documentation for all document operations and API usage.