Opensolr API Endpoint: doc_to_text
Overview
The doc_to_text endpoint reads documents into plain text — up to 5 documents in one call — so the words inside a PDF, a Word file or a contract can be indexed and searched like any other text. It is the document counterpart of image_ocr: that one reads pictures, this one reads files.
A PDF is answered page by page, each page under its own Page N line. A PDF is not only its text: the servers look at the pictures inside it — a scanned page, a photographed receipt, a signed and stamped page, a screenshot pasted into a report — and read them by OCR. Their words follow the page's own text, under the same Page N line, so a PDF made only of scans comes back as the OCR reading of every page. The pictures inside a Word or OpenDocument file are read the same way and their words are added at the end, under an Images heading. A document without pictures is read as its text layer alone.
| Document | What you get back |
|---|---|
| A supplier invoice as PDF | The invoice text, page by page; a stamped or signed page also read by OCR |
| A scanned contract (pictures only) | Every page read by OCR, under Page 1, Page 2, … |
| A Word file | Its paragraphs as plain text, and the words in its pictures under Images |
| A report with pasted screenshots | The report text, and the words in each screenshot, on the page it sits on |
This is how Opensolr Mail reads the attachments of a mailbox into its index, and the same endpoint is open to any account.
Text extraction and OCR run on the CPUs of api.opensolr.com — never on the GPU — so a bulk run does not queue behind the embedding and LLM work.
Endpoint URL
https://api.opensolr.com/solr_manager/api/doc_to_text
Supports only POST requests.
Authentication & Core Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| string | Yes | Your Opensolr registration email address. | |
| api_key | string | Yes | Your API key from the Opensolr dashboard. |
| index_name | string | Yes | Name of your Opensolr index the documents belong to. |
Document Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| documents | array | Yes | – | 1 to 5 documents, each an object with name (the file name, used only to tell a Word file from other Office files) and data (the file as base64). As a JSON body, or as a JSON string in a form field. |
Every document is checked by its content, never by its name: 20 MB ceiling, and the type is decided from the file's own bytes. Read today: PDF, Word (.doc, .docx), RTF, OpenDocument Text (.odt), HTML and plain text. Archives, spreadsheets, presentations and anything nested inside a zip are refused.
Example
curl -s -X POST "https://api.opensolr.com/solr_manager/api/doc_to_text" \ -H "Content-Type: application/json" \ -d "{\"email\":\"you@example.com\",\"api_key\":\"YOUR_API_KEY\",\"index_name\":\"your_index\",\"documents\":[{\"name\":\"invoice-2026-08.pdf\",\"data\":\"$(base64 -w0 invoice-2026-08.pdf)\"}]}"
Response
{ "status": true, "charged": 1, "results": [ { "status": true, "mime": "application/pdf", "text": "Page 1\nACME Hosting GmbH\nInvoice 084001059749\nTotal EUR 963.99\n\nPage 2\nTerms of payment ...\n\nPAID\nstamped 08/09/2026\nsignature" } ] }
| Field | Description |
|---|---|
| results | One entry per document, in the order you sent them. |
| status | Per document: true with the text, false with a msg saying why. |
| text | The document as plain text. For a PDF, page by page under Page N lines: the page's text first, then what OCR read in the pictures of that page. Capped at 200,000 characters. |
| mime | The type the servers decided from the bytes. |
| cached | Present and true when the same file had been read before, in which case it is free. |
| charged | How many AI requests this call actually counted (see below). |
What it costs
Reading the text of a document costs nothing. Each picture OCR actually reads inside it counts 0.5 of an AI request. A document answered from cache costs nothing, a picture read before costs nothing, and a document that could not be read costs nothing. Readings are kept permanently against the file's and the picture's own fingerprint, so sending the same PDF again is free, on any index of the account.
The endpoint needs a paid plan with vector search on it, like the other AI endpoints, and it counts against the same monthly allowance as every other AI endpoint. What your plan allows is on the API Quota tab of your dashboard.
Refusals
msg |
Meaning |
|---|---|
DOCUMENTS_MUST_BE_A_NON_EMPTY_ARRAY |
documents was missing or empty. |
MAX_5_DOCUMENTS_PER_CALL |
More than five documents in one call. |
INVALID_DOCUMENT |
data is not valid base64, or is empty. |
DOCUMENT_TOO_LARGE |
Over the 20 MB ceiling. |
UNSUPPORTED_DOCUMENT_TYPE |
The bytes are not one of the types read (an archive, a spreadsheet, a presentation, an executable). |
NO_TEXT |
The document was read and holds no text at all, not even in its pictures. This is a real answer, not a failure. |
EXTRACTION_UNAVAILABLE |
No reader answered. Nothing is charged; try again. |
ERROR_NOT_CORE_OWNER |
The index does not belong to the account making the call. |
How the pictures in a PDF are read
- On every plan.
- The servers list the pictures embedded in the PDF and take the ones at least 300 × 300 pixels (logos and rules are not text), each picture once even when it is repeated on every page, extracted as they are.
- Each picture is read in the language of its page, detected automatically among more than 100 languages; when no language is certain, in English, and again in the language of what was read when that is certain.
- Only confidently read words are kept, grouped in paragraphs, and placed under the page the picture sits on.
- A page with no text and no picture is left out.
Notes
- Documents are read in memory on a temporary file that is deleted immediately. Nothing is stored except the reading itself, in the cache, against the file's fingerprint.
- Barcodes and QR codes are not decoded by this endpoint.
- See it live: a live search over real PDFs whose words sit only inside pictures.