Data Crawler or Data Ingestion: Which One Should My Drupal Site Use?

Drupal

The Opensolr Search module fills your Opensolr Index in two ways. Both produce the same kind of documents, and you can run them side by side.

01 · Side by side

Data CrawlerData Ingestion
How it worksOpensolr reads your pages from the outside, from the sitemap the module builds at /opensolr-sitemap.xml.Drupal sends each item straight to your index.
New and changed contentFound while a crawl schedule runs; it checks every few minutes.Sent when an editor saves it (Enable real-time sync, on by default) and in search within about a minute.
Your site must bePublic, on HTTPS, and open to the crawler server IP shown in Crawl Management if a firewall protects it.Anything Drupal can load, intranets and sites behind a login included.
What goes inThe published pages of the content types you pick, as visitors see them, and attached files if you tick them.Items of the ticked types that an anonymous visitor may view: title, body, author, image, price, your mapped fields, and attached files if you tick them.
The whole site at onceReindex EverythingIngest All Now

02 · Which one

Data Crawler

Your site is public and on HTTPS, and you want Drupal to do no indexing work at all.

Data Ingestion

Your site is private or behind a firewall, or an edit must show in search within a minute.

Both

Data Ingestion sends each page the moment it is saved, and a crawl schedule runs as a safety net.

03 · More