Scanned PDFs and the pictures inside them, searchable.
A photographed label, a scanned form, a product shot, a drawing with no words at all. Opensolr reads every picture inside your PDFs in the language it is written in, describes the ones that carry no text, and puts all of it in your search. Try it live on real PDFs: product photos, shop labels, 1930s posters and century-old newspapers, in six languages.
01 — Try it
Search for words that only exist inside a picture.
None of these words are in the text layer of the PDFs. They are printed on a box, a bottle, a shop label or a sign that was photographed and dropped into a PDF, or, in the last two, they are not printed anywhere: the picture was found by what it shows.
A supplement box
A photo of the packaging, the brand and the ingredient printed on the front.
ashwagandhaA whisky label
The label of a bottle photographed on a shop shelf, read with its age and volume.
single malt whiskyA shop label, in Romanian
A label full of product details, read in Romanian because that is the language on the label.
usa lemn cremA sign on a restaurant wall
A notice photographed on the wall of a restaurant, read in Romanian.
dansul nu este permisA photo with no words
A picture of a puppy. Nothing to read, so Opensolr describes what it shows, and the description is what you search.
golden retriever puppyA drawing with no words
A floor plan in the middle of a PDF, found as a floor plan, with the rooms it shows.
floor plan02 — The test documents
The test PDFs, crawled by the Opensolr Web Crawler.
They sit in one public folder, opensolr.com/ocr-test, and the search above is an ordinary Opensolr Index that crawled it. Open any of them to compare what is on the page with what the search finds. The posters, newspapers and the infographic are in the public domain, from Wikimedia Commons and the CDC.
| What is inside | Pages | |
|---|---|---|
| product-packaging.pdf | Photos of a supplement box, a cigarette pack warning and a whisky label, in English and Romanian. | 3 |
| product-packaging-copy.pdf | The same file under another name. Its pictures were not read a second time: what was read once is remembered. | 3 |
| store-labels.pdf | Photographed shop labels, dense with small print, in Romanian. | 3 |
| signs-and-floor-plan.pdf | A sign on a restaurant wall, a floor plan with no words, and a shop label. | 3 |
| pictures-without-text.pdf | Two photos with nothing written on them, found by what they show. | 2 |
| all-pictures-mixed.pdf | All of the above in one document, one picture per page. | 8 |
| text-only-no-pictures.pdf | An ordinary PDF with a text layer and no pictures, indexed as it is. | 1 |
| wpa-health-posters.pdf | US public health posters from the 1930s: cancer quacks, tuberculosis, pneumonia, diphtheria, nursing the baby. Hand-lettered words inside artwork. | 8 |
| wpa-safety-posters.pdf | 1930s safety and nature posters: firecrackers, wildlife, drinking and driving, a child’s eyesight. | 4 |
| french-war-loan-posters-1917.pdf | French war loan posters from 1917, read in French. | 4 |
| german-war-bond-posters-1917.pdf | War bond posters from 1917, one of them printed side by side in Slovenian and German: each language read as it is. | 2 |
| titanic-front-pages-1912.pdf | Newspaper front pages from April 1912 reporting the Titanic, in English, Italian and German. Dense columns of small print. | 4 |
| la-sera-patria-del-friuli-1917.pdf | A scanned Italian newspaper from 1917, exactly as the archive published it. | 2 |
| cdc-als-get-the-facts.pdf | A public health infographic: text and small icons, indexed with its text. | 1 |
03 — What happens to a PDF
Every page keeps its place, every picture adds its words.
- Page by page. The text of each page comes first, under a
Page Nline, then what was read in the pictures of that page. - In the language of the page. The language is detected for every page, in more than 100 languages, and the pictures on it are read in that language.
- Pictures with no text get a meaning. A photo or a drawing is described in a short sentence, so it can be found by what it shows.
- Only what is worth keeping. Stray marks, logos and texture that only look like letters are left out, and so are empty pages.
- Read once. A picture already read, in any PDF, is not read again and not counted again.
- No waiting on the crawler. A crawled PDF is searchable right away with its own text; the text of its pictures joins it moments later.
Page 1 Anxiety Stress Relief Ashwagandha KSM-66 Page 2 Fumatul cauzează atacuri de inimă
The first two pages of product-packaging.pdf as they are stored in the index: nothing on those pages was text before, both are photos.
Page 1 a golden retriever puppy sitting on a white sheet Page 2 a bunch of colorful tulips with a blue sky in the background
pictures-without-text.pdf: two photos with no words, stored as what they show.
04 — Where it works
On every Opensolr Index. Nothing to switch on.
- Web Crawler. Every PDF the crawler finds in HTML + Documents mode, scanned or not. Web Crawler guide
- Data Ingestion API. Every document you send by URL for its text to be extracted. Data Ingestion guide
- Drupal. File attachments sent by the Opensolr Search module. Drupal data ingestion
- WordPress. Media attached to your posts, sent by the Opensolr plugin. WordPress content types
- On the Web Crawler, reading pictures is free. Only a description of a picture with no text counts, 0.5 of an AI request.
- On the API and the modules, each picture read counts 0.5 of an AI request and each picture described another 0.5.
- Anything read before is free, and the plain text of a document never counts.
- When the AI allowance runs out, pictures are still read; only the descriptions wait. API Quota
Put your own PDFs in.
Create an Opensolr Index, point the Web Crawler at your site or send your documents to the API, and search what is inside every picture.