Crawl PDFs, Word, Excel and Other Documents

PDFs page by page, office files, scans and pictures

Documents on your site

PDFs, Word files, spreadsheets and presentations linked from your pages become search results too, with their own text. Long PDFs are split page by page, so a search lands on the right page, not on page one of a 300-page manual.

What is read

PDF

.pdf, with text or scanned

Word

.doc, .docx

Excel

.xls, .xlsx

PowerPoint

.ppt, .pptx

OpenDocument

.odt, .ods, .odp

RTF and CSV

.rtf, .csv

It is on by default: in the crawl settings, Content Types is HTML & Documents (pdf, docx, etc...). Choose HTML Only to leave documents out.

Every PDF page is its own result

Each page of a PDF goes into your Opensolr Index with its own text and its page number, tied to the PDF it belongs to. A search finds the exact page. How the hosted search page shows PDF pages: Web and Documents tabs.

Scans and pictures inside PDFs

Opensolr reads the text printed in scanned pages and in the pictures inside your PDFs, automatically, on every Opensolr Index. The PDF is searchable at once with its own text; the text from its pictures is added a few moments later. A picture with no text in it can get a short description of what it shows.

Left out

  • Picture files linked from your pages (JPG, PNG, WebP and the like) are not crawled as pages.
  • Audio and video files are skipped.
  • A file larger than your plan allows is skipped and listed in Crawl Stats under Oversize Pages.

To read every PDF again, for example after you replaced many files, use Reindex PDFs: Reindex.

Crawl your site: all pages

Back to Enterprise Site Search