Scanned PDFs and the pictures inside them, searchable.

A photographed label, a scanned form, a product shot, a drawing with no words at all. Opensolr reads every picture inside your PDFs in the language it is written in, describes the ones that carry no text, and puts all of it in your search. Try it live on real PDFs: product photos, shop labels, 1930s posters and century-old newspapers, in six languages.

100+languages, detected page by page
Page Nevery page of the PDF labelled in its text
0.5of an AI request per picture read on the API
0for a picture that was read before

01 — Try it

Search for words that only exist inside a picture.

None of these words are in the text layer of the PDFs. They are printed on a box, a bottle, a shop label or a sign that was photographed and dropped into a PDF, or, in the last two, they are not printed anywhere: the picture was found by what it shows.

01

A supplement box

A photo of the packaging, the brand and the ingredient printed on the front.

ashwagandha
02

A whisky label

The label of a bottle photographed on a shop shelf, read with its age and volume.

single malt whisky
03

A shop label, in Romanian

A label full of product details, read in Romanian because that is the language on the label.

usa lemn crem
04

A sign on a restaurant wall

A notice photographed on the wall of a restaurant, read in Romanian.

dansul nu este permis
05

A photo with no words

A picture of a puppy. Nothing to read, so Opensolr describes what it shows, and the description is what you search.

golden retriever puppy
06

A drawing with no words

A floor plan in the middle of a PDF, found as a floor plan, with the rooms it shows.

floor plan

02 — The test documents

The test PDFs, crawled by the Opensolr Web Crawler.

They sit in one public folder, opensolr.com/ocr-test, and the search above is an ordinary Opensolr Index that crawled it. Open any of them to compare what is on the page with what the search finds. The posters, newspapers and the infographic are in the public domain, from Wikimedia Commons and the CDC.

PDFWhat is insidePages
product-packaging.pdfPhotos of a supplement box, a cigarette pack warning and a whisky label, in English and Romanian.3
product-packaging-copy.pdfThe same file under another name. Its pictures were not read a second time: what was read once is remembered.3
store-labels.pdfPhotographed shop labels, dense with small print, in Romanian.3
signs-and-floor-plan.pdfA sign on a restaurant wall, a floor plan with no words, and a shop label.3
pictures-without-text.pdfTwo photos with nothing written on them, found by what they show.2
all-pictures-mixed.pdfAll of the above in one document, one picture per page.8
text-only-no-pictures.pdfAn ordinary PDF with a text layer and no pictures, indexed as it is.1
wpa-health-posters.pdfUS public health posters from the 1930s: cancer quacks, tuberculosis, pneumonia, diphtheria, nursing the baby. Hand-lettered words inside artwork.8
wpa-safety-posters.pdf1930s safety and nature posters: firecrackers, wildlife, drinking and driving, a child’s eyesight.4
french-war-loan-posters-1917.pdfFrench war loan posters from 1917, read in French.4
german-war-bond-posters-1917.pdfWar bond posters from 1917, one of them printed side by side in Slovenian and German: each language read as it is.2
titanic-front-pages-1912.pdfNewspaper front pages from April 1912 reporting the Titanic, in English, Italian and German. Dense columns of small print.4
la-sera-patria-del-friuli-1917.pdfA scanned Italian newspaper from 1917, exactly as the archive published it.2
cdc-als-get-the-facts.pdfA public health infographic: text and small icons, indexed with its text.1

03 — What happens to a PDF

Every page keeps its place, every picture adds its words.

  • Page by page. The text of each page comes first, under a Page N line, then what was read in the pictures of that page.
  • In the language of the page. The language is detected for every page, in more than 100 languages, and the pictures on it are read in that language.
  • Pictures with no text get a meaning. A photo or a drawing is described in a short sentence, so it can be found by what it shows.
  • Only what is worth keeping. Stray marks, logos and texture that only look like letters are left out, and so are empty pages.
  • Read once. A picture already read, in any PDF, is not read again and not counted again.
  • No waiting on the crawler. A crawled PDF is searchable right away with its own text; the text of its pictures joins it moments later.
Page 1
Anxiety Stress Relief
Ashwagandha KSM-66

Page 2
Fumatul cauzează atacuri
de inimă

The first two pages of product-packaging.pdf as they are stored in the index: nothing on those pages was text before, both are photos.

Page 1
a golden retriever puppy sitting on a white sheet

Page 2
a bunch of colorful tulips with a blue sky in the background

pictures-without-text.pdf: two photos with no words, stored as what they show.

04 — Where it works

On every Opensolr Index. Nothing to switch on.

  • On the Web Crawler, reading pictures is free. Only a description of a picture with no text counts, 0.5 of an AI request.
  • On the API and the modules, each picture read counts 0.5 of an AI request and each picture described another 0.5.
  • Anything read before is free, and the plain text of a document never counts.
  • When the AI allowance runs out, pictures are still read; only the descriptions wait. API Quota

Put your own PDFs in.

Create an Opensolr Index, point the Web Crawler at your site or send your documents to the API, and search what is inside every picture.