Data Crawler: Let Opensolr Read Your Site

Opensolr reads your Drupal pages and files from the outside, the way a search engine does

The Data Crawler Tab

With the Data Crawler, your Drupal site does no indexing work at all. Opensolr reads your pages from the outside, the way a search engine does, and puts them in your Opensolr Index.

Where: Configuration › Search and metadata › Opensolr Search › Data Crawler tab (/admin/config/search/opensolr/crawler)

How it works

  1. The module lists the published pages of the content types you pick in a sitemap at /opensolr-sitemap.xml, with the attached files too if you want them.
  2. Opensolr reads every address in that sitemap. It does not follow the links inside your pages, so only what the sitemap lists gets in.
  3. Each page and file gets what Opensolr adds to every document it indexes: see What Opensolr adds to every document.
  4. With a crawl schedule running, Opensolr keeps checking for new and changed pages every few minutes.

The tab, part by part

What it needs

  • Your site on HTTPS. Without it the tab shows HTTPS required and no crawl buttons.
  • Opensolr let through. If a firewall or a rate limiter protects your site, allow the crawler server IP shown at the top of Crawl Management.
  • An index connected on the Settings tab.
Crawl, push, or both

The Data Crawler and Data Ingestion produce the same documents in your index. A common setup: Data Ingestion sends each page the moment it is saved, and the crawl schedule runs as a safety net.

Data Crawler