Scanned PDFs and the Pictures Inside Them Are Now Searchable
A lot of what sits in a PDF is not text. A scanned contract, a photographed price tag, a poster, a page of an old newspaper: until now search saw only the words a PDF carried as text and skipped everything that was a picture. Opensolr now reads those pictures, and what they say becomes part of your search.
01 What you get
- The words inside pictures are searchable. Scanned pages and every picture of a useful size inside a PDF are read with OCR, and their text joins the page they sit on.
- In the language of the page. The language is detected page by page, in more than 100 languages, and each page is read in its own.
- Pictures with no words get a meaning. A photo or a drawing is described in a short sentence, so a search for what it shows finds it.
- Page by page. The text of each page comes first, then what was read in its pictures, under a
Page Nline, so a match can be traced to its page. - Read once. A picture that was read before, in any PDF, is not read again.
- No waiting. A crawled PDF is searchable right away with its own text; the text of its pictures joins it moments later.
02 Where it works
- Web Crawler. Every PDF it finds in HTML + Documents mode. Web Crawler guide
- Data Ingestion API. Every document you send by URL for its text to be extracted. Data Ingestion guide
- Drupal and WordPress. The file attachments the Opensolr modules send.
- Document to Text API and Opensolr Mail. Documents and mail attachments are read the same way.
03 What it costs
- On the Web Crawler, reading pictures is free. Only the description of a picture with no words counts, half an AI request.
- On the API and the modules, each picture read counts half an AI request, and each picture described another half.
- Anything read before is free, the plain text of a document never counts, and searching is never charged. API Quota
04 Also: snippets that show the match
- The highlighted words are always in view. On the hosted search page, the text under each result is now cut around the first matched word, so a match deep inside a long paragraph or a dense scanned page is shown highlighted instead of cut away.