Text in Scans and Pictures (OCR)

Scanned pages and pictures in PDFs, searchable by their words

Text in Scans and Pictures

Opensolr reads the text printed in scanned PDF pages and in the pictures inside your PDFs. A scanned contract, a brochure made of images or a photographed label inside a PDF is found by its words, like any other text.

Every page of a PDF is its own search result, so a visitor lands on the exact page that answers the search.

A PDF page a scan, or pictures with printed text Text read in the language of the page Found that exact page is a result A PDF page a scan, or pictures with text Text read in the language of the page Found that exact page is a result

How the pictures are read

  • Each picture is read in the language of its page. When the page has no text to tell, in the language of the document; when that is unknown too, in English first, then again in the language the text turns out to be.
  • Very small pictures, like logos, icons and bullets, are skipped.
  • Only words read with confidence are kept, so stains and noise do not turn into words.
  • A word read with one wrong letter can still be found: the read text also matches by sound.
  • A picture with no text in it can get a short description of what it shows: Pictures described by what they show.

Where it happens

What it uses from your plan

  • On a crawl, reading pictures is free.
  • On Data Ingestion and by API, each picture read uses part of one AI request of the account that owns the index.
  • A picture read before is free, and the plain text of a document never counts.
  • When the monthly AI allowance runs out, pictures are still read; only the descriptions of pictures without text wait.

More: What site search uses from your plan.

See it live

Search inside scanned public-domain PDFs on the OCR search demo.

Meaning, pictures, answers

Back to Enterprise Site Search