Read your site again
Three buttons in Crawler Setup crawl your site again, once: every page already read, everything from the start, or only the PDFs. None of them deletes your Opensolr Index, so your search keeps working while they run.
The three buttons
Puts every page already crawled back in the queue, the ones left out earlier included, and runs the crawl once. Pages that changed are updated, and pages that now answer with an error leave your index. The right choice after you changed settings or rules, or much of your content.
Forgets which pages were already read and runs a completely fresh crawl from your start URLs. Nothing is deleted from your index: existing results stay. Use it when the structure of your site changed and new pages should be found again.
Puts only the PDFs already crawled back in the queue and reads them again. Web pages stay as they are. Use it to read every PDF again, the text in its pictures included.
To clear out pages that no longer exist on your site, use Reindex All: Reindex From Scratch only reads the pages it finds, so a page nothing links to any more stays in your index.
Good to know
- Each button asks you to confirm first.
- These runs use your saved crawl settings and never create a schedule. While one runs without a schedule, a Stop Crawl button stops it: Start, pause, stop, flush.
- To read every page again on a regular basis, without clicking, use Keep Index Fresh.
- By API: start_crawl with
requeue=all, or with a list of document types such asrequeue=pdforrequeue=pdf,docx, andschedule=nofor a single run.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates