Three buttons under Crawl Management, on the Data Crawler tab, start and stop the crawl of your site.
Reindex Everything
Forgets which pages were already read and crawls everything in your sitemap again. WordPress asks you to confirm first. Your documents stay in the index while the pages are read again; each one is replaced as soon as its page has been read. This button also starts the crawl schedule.
Resume Indexing
Continues the crawl where it stopped, without starting over. It also starts the crawl schedule.
Stop Crawl
Stops the running crawl and removes its schedule.
The crawl schedule
While the schedule is on, Opensolr keeps the crawl of your site going in the background, checking it every minute, until you click Stop Crawl. In the Opensolr Control Panel the same index then shows Stop Crawl Schedule: Start, pause, stop, flush.
To crawl your site once, with no schedule, use Reindex From Scratch: Reindex From Scratch or Reset Index.
Every time a crawl starts
- The plugin adds your sitemap to the crawl list again (sites whose address starts with
https://) and sends your Crawler Settings. - Rules you set for this index in the Opensolr Control Panel apply too, such as pages left out of the crawl: Rules per site and Pages left out everywhere.
- The Crawl Status window opens by itself.
When a crawl does not start
The reason appears under the buttons. For example, when the index already holds more pages than the crawl limit of your plan ("The Solr index contains N pages out of the M pages limit"). Remove data first with Reset Index, or move to a bigger plan on https://opensolr.com/pricing.
Data Crawler
- Data Crawler
- What gets crawled
- Crawl speed
- Start, resume, stop
- Reindex or reset
- Crawl status