Document to Text API: the Pictures Inside a PDF Are Read Too, Page by Page
A PDF is not always text. Scanned contracts, photographed receipts, a diagram with the part numbers printed on it: the Document to Text API used to return the words a PDF carried as text and skip what was only a picture. It now reads those pictures too, and hands the text back page by page.
01 What changed
- Pictures inside a PDF are read. Every picture of a useful size in the document, up to forty per document, goes through OCR, several at a time, and its text joins the text of the page it sits on. A PDF that is nothing but scans comes back as the text printed on the scans.
- Page by page. The text comes back in
Page 1,Page 2sections, in the order of the document, so a number found in the answer can be traced to its page. - The same request. The endpoint, its parameters and its limits are unchanged: send the document, get the text. Word, Excel, PowerPoint, OpenDocument, RTF, HTML, text and pictures on their own are read as before.
- Where it shows. Opensolr Mail reads attachments through it, so a scanned invoice in your mail is found by the number printed in it, and anything you build on the endpoint gets the same.