AI-API - Documents to text, with the pictures inside a PDF read by OCR (doc_to_text)

AI API: Vectors, Images & LLM

Opensolr API Endpoint: doc_to_text

Overview

The doc_to_text endpoint reads documents into plain text — up to 5 documents in one call — so the words inside a PDF, a Word file or a contract can be indexed and searched like any other text. It is the document counterpart of image_ocr: that one reads pictures, this one reads files.

A PDF is answered page by page, each page under its own Page N line. A PDF is not only its text: the servers look at the pictures inside it — a scanned page, a photographed receipt, a signed and stamped page, a screenshot pasted into a report — and read them by OCR. Their words follow the page's own text, under the same Page N line, so a PDF made only of scans comes back as the OCR reading of every page. The pictures inside a Word or OpenDocument file are read the same way and their words are added at the end, under an Images heading. A document without pictures is read as its text layer alone.

Document What you get back
A supplier invoice as PDF The invoice text, page by page; a stamped or signed page also read by OCR
A scanned contract (pictures only) Every page read by OCR, under Page 1, Page 2, …
A Word file Its paragraphs as plain text, and the words in its pictures under Images
A report with pasted screenshots The report text, and the words in each screenshot, on the page it sits on

This is how Opensolr Mail reads the attachments of a mailbox into its index, and the same endpoint is open to any account.

Text extraction and OCR run on the CPUs of api.opensolr.com — never on the GPU — so a bulk run does not queue behind the embedding and LLM work.


Endpoint URL

https://api.opensolr.com/solr_manager/api/doc_to_text

Supports only POST requests.


Authentication & Core Parameters

Parameter Type Required Description
email string Yes Your Opensolr registration email address.
api_key string Yes Your API key from the Opensolr dashboard.
index_name string Yes Name of your Opensolr index the documents belong to.

Document Parameters

Parameter Type Required Default Description
documents array Yes – 1 to 5 documents, each an object with name (the file name, used only to tell a Word file from other Office files) and data (the file as base64). As a JSON body, or as a JSON string in a form field.

Every document is checked by its content, never by its name: 20 MB ceiling, and the type is decided from the file's own bytes. Read today: PDF, Word (.doc, .docx), RTF, OpenDocument Text (.odt), HTML and plain text. Archives, spreadsheets, presentations and anything nested inside a zip are refused.


Example

curl -s -X POST "https://api.opensolr.com/solr_manager/api/doc_to_text" \
  -H "Content-Type: application/json" \
  -d "{\"email\":\"you@example.com\",\"api_key\":\"YOUR_API_KEY\",\"index_name\":\"your_index\",\"documents\":[{\"name\":\"invoice-2026-08.pdf\",\"data\":\"$(base64 -w0 invoice-2026-08.pdf)\"}]}"

Response

{
  "status": true,
  "charged": 1,
  "results": [
    {
      "status": true,
      "mime": "application/pdf",
      "text": "Page 1\nACME Hosting GmbH\nInvoice 084001059749\nTotal EUR 963.99\n\nPage 2\nTerms of payment ...\n\nPAID\nstamped 08/09/2026\nsignature"
    }
  ]
}
Field Description
results One entry per document, in the order you sent them.
status Per document: true with the text, false with a msg saying why.
text The document as plain text. For a PDF, page by page under Page N lines: the page's text first, then what OCR read in the pictures of that page. Capped at 200,000 characters.
mime The type the servers decided from the bytes.
cached Present and true when the same file had been read before, in which case it is free.
charged How many AI requests this call actually counted (see below).

What it costs

Reading the text of a document costs nothing. Each picture OCR actually reads inside it counts 0.5 of an AI request. A document answered from cache costs nothing, a picture read before costs nothing, and a document that could not be read costs nothing. Readings are kept permanently against the file's and the picture's own fingerprint, so sending the same PDF again is free, on any index of the account.

The endpoint needs a paid plan with vector search on it, like the other AI endpoints, and it counts against the same monthly allowance as every other AI endpoint. What your plan allows is on the API Quota tab of your dashboard.


Refusals

msg Meaning
DOCUMENTS_MUST_BE_A_NON_EMPTY_ARRAY documents was missing or empty.
MAX_5_DOCUMENTS_PER_CALL More than five documents in one call.
INVALID_DOCUMENT data is not valid base64, or is empty.
DOCUMENT_TOO_LARGE Over the 20 MB ceiling.
UNSUPPORTED_DOCUMENT_TYPE The bytes are not one of the types read (an archive, a spreadsheet, a presentation, an executable).
NO_TEXT The document was read and holds no text at all, not even in its pictures. This is a real answer, not a failure.
EXTRACTION_UNAVAILABLE No reader answered. Nothing is charged; try again.
ERROR_NOT_CORE_OWNER The index does not belong to the account making the call.

How the pictures in a PDF are read

  • On every plan.
  • The servers list the pictures embedded in the PDF and take the ones at least 300 × 300 pixels (logos and rules are not text), each picture once even when it is repeated on every page, extracted as they are.
  • Each picture is read in the language of its page, detected automatically among more than 100 languages; when no language is certain, in English, and again in the language of what was read when that is certain.
  • Only confidently read words are kept, grouped in paragraphs, and placed under the page the picture sits on.
  • A page with no text and no picture is left out.

Notes