Reindex All, Reindex From Scratch, Reindex PDFs

Read your site again, all of it, from the start, or only the PDFs

Read your site again

Three buttons in Crawler Setup crawl your site again, once: every page already read, everything from the start, or only the PDFs. None of them deletes your Opensolr Index, so your search keeps working while they run.

The three buttons

Reindex All

Puts every page already crawled back in the queue, the ones left out earlier included, and runs the crawl once. Pages that changed are updated, and pages that now answer with an error leave your index. The right choice after you changed settings or rules, or much of your content.

Reindex From Scratch

Forgets which pages were already read and runs a completely fresh crawl from your start URLs. Nothing is deleted from your index: existing results stay. Use it when the structure of your site changed and new pages should be found again.

Reindex PDFs

Puts only the PDFs already crawled back in the queue and reads them again. Web pages stay as they are. Use it to read every PDF again, the text in its pictures included.

To clear out pages that no longer exist on your site, use Reindex All: Reindex From Scratch only reads the pages it finds, so a page nothing links to any more stays in your index.

Good to know

  • Each button asks you to confirm first.
  • These runs use your saved crawl settings and never create a schedule. While one runs without a schedule, a Stop Crawl button stops it: Start, pause, stop, flush.
  • To read every page again on a regular basis, without clicking, use Keep Index Fresh.
  • By API: start_crawl with requeue=all, or with a list of document types such as requeue=pdf or requeue=pdf,docx, and schedule=no for a single run.

Crawl your site: all pages

Back to Enterprise Site Search