Watch the crawl
Crawl Stats shows the crawl of your site as it happens: how far it got, what it downloaded, which pages failed and which were left out on purpose. Failed pages can be sent back for another try in one click.
Open it
In the WebCrawler tab, click Crawl Stats in the Quick Access bar. It opens in a new window and refreshes itself: pick Auto-refresh Off, 10 seconds, 30 seconds (the default), 1, 2 or 5 minutes, or click Refresh.
The numbers
Pages discovered, out of the pages your plan allows, with the pages still waiting.
What the crawl downloaded, out of the download limit of your plan.
Pages rendered with JavaScript.
Pages that answered with an error, 4xx and 5xx.
Pages and files larger than your plan allows, skipped.
Pages read for their links and kept out of the index by a Do not index rule, and sitemaps and feeds themselves.
Pages outside the crawl mode.
Pages that redirected outside the allowed domain.
Pages crawled and waiting to be written to your index. While there are some, a Flush to Solr button sends them right away.
How many pages and documents of each type were read.
The lists
Below the numbers, each group has its list of addresses, every one a link: Client Errors (4xx), Server Errors (5xx), Oversize Pages, Noindex Filtered, Out of Scope and Left Domain. A long list shows its first 100 addresses, with the full count next to the title.
Retry URLs, next to Client Errors and Server Errors, puts those pages back in the queue, so the next run of the crawl tries them again. Use it after you fixed the cause on your site.
By API
The same numbers and the last lines of the crawl log are available to your own code: Get live crawl stats and Tail the crawl log.
What people search for on your site is in Query Stats, the button next to Crawl Stats: Search analytics.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates