Run the crawl
The Crawler Setup panel of the WebCrawler tab starts and stops the crawl of your site. A crawl schedule keeps it going on its own; Pause, Resume and Flush to Solr cover the moments in between.
The status
The Status line says what is happening now: Running, Idle (a schedule is on and nothing runs at this moment), Stopped (no schedule, nothing runs), Starting... or Stopping.... The round arrow next to it checks again.
The buttons
Starts the crawl now and puts it on a schedule. At regular intervals the schedule calls the crawl again with Resume, so a crawl that stopped picks up where it left off, and pages queued later, by Keep Index Fresh or your CMS, are read. Shown while no schedule is on.
Stops the crawl and removes the schedule. Shown while a schedule is on.
Shown while the crawl runs. Stops it now and keeps the schedule, so the schedule starts it again, with Resume, on its next run. To stop for good, use Stop Crawl Schedule.
Shown while nothing runs. Carries on from where the crawl stopped. Without a schedule it runs once and creates none.
Shown only while a crawl without a schedule runs, like one started by the Reindex buttons. Stops it.
Crawled pages reach your index in batches. This sends the pages waiting in the buffer right away, even while the crawl runs, and shows the progress.
The same panel holds the Reindex buttons and Global Exclusions.
When a crawl does not start
No active URLs found.Every start URL is Inactive: tick one in Crawl URLs and click Activate.
ERROR_PLEASE_VERIFY_YOUR_URLSNo active start URL is verified yet: Prove you own the site.
The Solr index contains ...Your index already holds more documents than your plan lets the crawl fill. Remove documents, or move to a larger plan on https://opensolr.com/pricing.
ERROR_THIS_CRAWLER_IS_NOW_STOPPING_PLEASE_WAITA stop is still in progress. Try again in a moment.
Every button here is also an API call, for scripts and other systems: Control the crawl by API.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates