Choose What the Data Crawler Indexes

Content types and attached files in the sitemap

The Data Crawler reads the pages listed in the plugin sitemap. You decide what goes into that sitemap: the boxes under Content Types for Crawling, at the top of the Data Crawler tab.

Steps

  1. Open Settings > Opensolr Search, tab Data Crawler.
  2. Tick the content types to index: Posts, Pages, Products and any other public post type of your site. Each shows how many published items it has, and Total adds them up.
  3. Tick Include attached files (PDF, DOCX, etc.) to index your documents too.
  4. Click Save Configuration.
  5. Start a crawl with Reindex Everything so the new list is read: Start, resume and stop the crawl.

What goes into the sitemap

  • Published items of the ticked types. Drafts, private items and the search page itself are never listed.
  • With Include attached files: the document files of your Media Library, PDF, Word, Excel, PowerPoint, OpenDocument text and spreadsheets, plain text and CSV. The crawl also follows links to documents it finds on your pages.

Opensolr reads the text of each document, and a PDF shows up page by page in the results: Documents on your site.

Separate from Data Ingestion

The Data Ingestion tab has its own boxes. What you tick here only decides what the Data Crawler reads.

Data Crawler