NextCloud OCR
AI-generatedSummary
Extract text content from a PDF file stored on Nextcloud using OCR (Optical Character Recognition) via either a local Docling server or the cloud-based Mistral OCR API.
Inputs
- filePathMode (required) — Choose how to specify the source PDF file: either select from a list of files filtered by folder or provide a direct file path expression.
- folderPath — Select the folder on Nextcloud to filter for PDF files when choosing from a list.
- filePathFromList — Choose the PDF file from the selected folder when using the list mode.
- filePath — Provide the full Nextcloud file path (supports expressions) of the PDF to process when using path mode.
- ocrProvider (required) — Select the OCR engine to extract text: either 'Docling' for a local server or 'Mistral OCR' for a cloud API.
- doclingEndpoint — Base URL of the locally hosted Docling server (used if Docling OCR is selected).
- doclingApiPath — API endpoint path for Docling server; select according to your Docling server version or specify a custom path.
- doclingApiPathCustom — Custom endpoint path for Docling API, if 'custom' is selected.
- mistralApiKey — API key for authenticating with Mistral OCR cloud service (required if Mistral OCR is selected).
- mistralModel — Identifier of the OCR model to use with Mistral OCR (default is 'mistral-ocr-latest').
- includeRaw — Choose whether to include the full raw OCR engine response in the output for debugging purposes.
Output shape
a list of JSON objects, each corresponding to an input item, containing extracted text and metadata
Each output item includes the PDF file path, OCR provider used, extracted text and markdown content, page count, optionally page details (if provided by OCR engine), and optionally the raw OCR response if requested. The node processes each input item independently and supports error handling via 'continueOnFail'.