Choose What the Data Crawler Reads

Content types, products and attached files for the crawl

Content Types for Crawling

Pick which content types Opensolr reads. Only the published pages of the types you tick are listed in your sitemap, so only those reach your Opensolr Index through the crawl.

Where: Opensolr Search › Data Crawler tab › Content Types for Crawling

Set it

  1. Under Include in Crawler, tick the content types to search. With Drupal Commerce installed, product types are listed too, as Product: <type>.
  2. Tick Include attached files (PDFs, documents) to add the PDF, Word, Excel and other documents attached to those pages.
  3. Watch the counters: each ticked type shows how many published items it has, a total appears below the list, and the Crawl Queue line in Crawl Management counts the addresses in your sitemap, pages and files.
  4. Click Save configuration.

Good to know

  • The same choice decides which pages carry the module's meta tags, the ones Opensolr reads for titles, descriptions, dates and your mapped facet fields. See SEO and meta tags.
  • On a multilingual site, every published translation is listed under its own address. See Multilingual search.
  • For documents, the module recommends Data Ingestion: it keeps the language you set in Drupal, while a crawled document gets the language detected from its own text.
  • Data Ingestion has its own content type list, so you can crawl some types and push others.

Data Crawler