The Opensolr Search module fills your Opensolr Index in two ways. Both produce the same kind of documents, and you can run them side by side.
01 · Side by side
| Data Crawler | Data Ingestion | |
|---|---|---|
| How it works | Opensolr reads your pages from the outside, from the sitemap the module builds at /opensolr-sitemap.xml. | Drupal sends each item straight to your index. |
| New and changed content | Found while a crawl schedule runs; it checks every few minutes. | Sent when an editor saves it (Enable real-time sync, on by default) and in search within about a minute. |
| Your site must be | Public, on HTTPS, and open to the crawler server IP shown in Crawl Management if a firewall protects it. | Anything Drupal can load, intranets and sites behind a login included. |
| What goes in | The published pages of the content types you pick, as visitors see them, and attached files if you tick them. | Items of the ticked types that an anonymous visitor may view: title, body, author, image, price, your mapped fields, and attached files if you tick them. |
| The whole site at once | Reindex Everything | Ingest All Now |
02 · Which one
Data Crawler
Your site is public and on HTTPS, and you want Drupal to do no indexing work at all.
Data Ingestion
Your site is private or behind a firewall, or an edit must show in search within a minute.
Both
Data Ingestion sends each page the moment it is saved, and a crawl schedule runs as a safety net.
03 · More
The tabs in detail: Data Crawler and Data Ingestion.
Bulk loads and rebuilds: Ingest All Now, Reindex Everything, Reindex From Scratch, Reset Index.
The same two ways for any website: Crawl your site and Push your data.
The whole module: Complete Guide to the Opensolr Search Drupal Module.