AI-API - Read the text printed in a picture (image_ocr)

AI API: Vectors, Images & LLM

Opensolr API Endpoint: image_ocr

Overview

The image_ocr endpoint reads the text printed in a picture — up to 10 pictures in one call — and answers the words it read. It is the other half of picture search: image_index and image_clip tell you what a picture shows; image_ocr tells you what is written on it.

That difference is the whole point for paperwork:

Picture What the words endpoints say What image_ocr reads
A gas receipt receipt, paper, text, document CHEVRON STATION 1042, TOTAL $52.86, RECEIPT 004417
A supplier invoice invoice, document, letter The company, the invoice number, the amount, the due date
A shelf label label, product label, price tag The product name and its article number
A screenshot screenshot, website, text Whatever it says
A business card business card, text The name, the company, the phone number

Index that text and a folder of receipts becomes searchable the way a customer actually remembers it: by the gas station, by the total, by the invoice number. This is how Opensolr Photos makes a phone full of paperwork searchable, and the same endpoint is open to any account.

Reading runs with tesseract on the CPUs of api.opensolr.com, one process per picture, every picture of a call at the same time, never on the GPU, so a bulk run does not queue behind the embedding and LLM work.


Endpoint URL

https://api.opensolr.com/solr_manager/api/image_ocr

Supports only POST requests.


Authentication & Core Parameters

Parameter Type Required Description
email string Yes Your Opensolr registration email address.
api_key string Yes Your API key from the Opensolr dashboard.
index_name string Yes Name of your Opensolr index the pictures belong to.

Image Parameters

Parameter Type Required Default Description
images array Yes – 1 to 10 pictures, each as base64. A data: URI prefix is accepted. As a JSON body, or as a JSON string in a form field.

Every picture is checked by its content, never by its name: 20 MB ceiling, and the type is decided from the file's own bytes (JPEG, PNG, WebP, GIF, BMP, TIFF, JP2, HEIC/HEIF, AVIF).


Example

curl -s -X POST "https://api.opensolr.com/solr_manager/api/image_ocr" \
  -H "Content-Type: application/json" \
  -d "{\"email\":\"you@example.com\",\"api_key\":\"YOUR_API_KEY\",\"index_name\":\"my_index\",\"images\":[\"$(base64 -w0 receipt.jpg)\"]}"

Response

{
  "status": true,
  "charged": 1,
  "results": [
    {
      "status": true,
      "text": "CHEVRON STATION 1042\n2400 Main St, Houston TX\nUNLEADED REGULAR\nTOTAL $52.86\nRECEIPT 004417",
      "lines": ["CHEVRON STATION 1042", "2400 Main St, Houston TX", "UNLEADED REGULAR", "TOTAL $52.86", "RECEIPT 004417"],
      "secs": 0.245,
      "lang": "eng+ron"
    }
  ]
}
Field Description
results One entry per picture, in the order you sent them.
status Per picture: true with the text, false with a msg saying why.
text Everything read, in reading order, one line per printed line.
lines The same reading as an array of lines — easier to index or to show.
useful true when the reading holds at least four confidently read words of three letters or more: real text, not the specks OCR finds on any photo.
cached Present and true when the picture had been read before, in which case it is free.
charged How many AI requests this call actually counted (see below).

What it costs

0.5 of an AI request per picture the servers actually read. A picture answered from cache costs nothing, and a picture that could not be read costs nothing. Readings are kept permanently against the picture's own fingerprint, so sending the same receipt again is free.

On an Opensolr Photos index the reading is part of the photo's own tenth of a request and is not charged separately.

The endpoint needs a paid plan with vector search on it, like the other AI endpoints, and it counts against the same monthly allowance as every other AI endpoint. What your plan allows is on the API Quota tab of your dashboard.


Refusals

msg Meaning
IMAGES_MUST_BE_A_NON_EMPTY_ARRAY images was missing or empty.
MAX_10_IMAGES_PER_CALL More than ten pictures in one call.
INVALID_IMAGE / UNSUPPORTED_IMAGE_TYPE_… The bytes are not a picture of a supported type.
IMAGE_TOO_LARGE_MAX_20_MB Over the size ceiling.
OCR_EMPTY The picture was read and carries no usable text. This is a real answer, not a failure.
OCR_UNAVAILABLE No reader answered. Nothing is charged; try again.
ERROR_NOT_CORE_OWNER The index does not belong to the account making the call.

Reading only the pictures worth reading

If you are indexing a whole library, do not send everything: ask image_clip (or image_index) first and look at text_like in its answer. It is true when what the image model saw in the picture is text-like — a label, a document, a receipt, an invoice, a card, a ticket, a menu, a poster, a screenshot, a number — judged on the sentence it wrote and the objects it found. Send those to image_ocr and skip the rest.


Notes

  • Languages read today: English and Romanian (eng+ron), in one pass, with the reading tuned for sparse text — the mode that finds an article number next to a barcode.
  • Pictures are read in memory on a temporary file that is deleted immediately. Nothing is stored except the reading itself, in the cache, against the picture's fingerprint.
  • Barcodes and QR codes are not decoded by this endpoint. image_to_text returns those separately as codes.
  • See it live: a live search over real PDFs whose words sit only inside pictures.