The Data Crawler Tab
With the Data Crawler, your Drupal site does no indexing work at all. Opensolr reads your pages from the outside, the way a search engine does, and puts them in your Opensolr Index.
Where: Configuration › Search and metadata › Opensolr Search › Data Crawler tab (
/admin/config/search/opensolr/crawler)How it works
- The module lists the published pages of the content types you pick in a sitemap at
/opensolr-sitemap.xml, with the attached files too if you want them. - Opensolr reads every address in that sitemap. It does not follow the links inside your pages, so only what the sitemap lists gets in.
- Each page and file gets what Opensolr adds to every document it indexes: see What Opensolr adds to every document.
- With a crawl schedule running, Opensolr keeps checking for new and changed pages every few minutes.
The tab, part by part
- Content Types for Crawling: which content types, products and attached files go in.
- Crawler Settings: how many pages are read at once and the pause between requests.
- Crawl Management: start and stop the crawl schedule, check for changes now, stop a running crawl.
- Reindex Everything, Reindex From Scratch, Reset Index: when to use which.
- Crawl Status: live numbers and the addresses that failed.
What it needs
- Your site on HTTPS. Without it the tab shows HTTPS required and no crawl buttons.
- Opensolr let through. If a firewall or a rate limiter protects your site, allow the crawler server IP shown at the top of Crawl Management.
- An index connected on the Settings tab.
Crawl, push, or both
The Data Crawler and Data Ingestion produce the same documents in your index. A common setup: Data Ingestion sends each page the moment it is saved, and the crawl schedule runs as a safety net.