Documents on your site
PDFs, Word files, spreadsheets and presentations linked from your pages become search results too, with their own text. Long PDFs are split page by page, so a search lands on the right page, not on page one of a 300-page manual.
What is read
.pdf, with text or scanned
.doc, .docx
.xls, .xlsx
.ppt, .pptx
.odt, .ods, .odp
.rtf, .csv
It is on by default: in the crawl settings, Content Types is HTML & Documents (pdf, docx, etc...). Choose HTML Only to leave documents out.
Every PDF page is its own result
Each page of a PDF goes into your Opensolr Index with its own text and its page number, tied to the PDF it belongs to. A search finds the exact page. How the hosted search page shows PDF pages: Web and Documents tabs.
Scans and pictures inside PDFs
Opensolr reads the text printed in scanned pages and in the pictures inside your PDFs, automatically, on every Opensolr Index. The PDF is searchable at once with its own text; the text from its pictures is added a few moments later. A picture with no text in it can get a short description of what it shows.
Left out
- Picture files linked from your pages (JPG, PNG, WebP and the like) are not crawled as pages.
- Audio and video files are skipped.
- A file larger than your plan allows is skipped and listed in Crawl Stats under Oversize Pages.
To read every PDF again, for example after you replaced many files, use Reindex PDFs: Reindex.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates