Document to Text API: the Pictures Inside a PDF Are Read Too, Page by Page

· API · Improvement · All updates

A PDF is not always text. Scanned contracts, photographed receipts, a diagram with the part numbers printed on it: the Document to Text API used to return the words a PDF carried as text and skip what was only a picture. It now reads those pictures too, and hands the text back page by page.

01 What changed

  • Pictures inside a PDF are read. Every picture of a useful size in the document, up to forty per document, goes through OCR, several at a time, and its text joins the text of the page it sits on. A PDF that is nothing but scans comes back as the text printed on the scans.
  • Page by page. The text comes back in Page 1, Page 2 sections, in the order of the document, so a number found in the answer can be traced to its page.
  • The same request. The endpoint, its parameters and its limits are unchanged: send the document, get the text. Word, Excel, PowerPoint, OpenDocument, RTF, HTML, text and pictures on their own are read as before.
  • Where it shows. Opensolr Mail reads attachments through it, so a scanned invoice in your mail is found by the number printed in it, and anything you build on the endpoint gets the same.
View the full changelog