What your pages tell the crawl
The tags your pages already carry for search engines also steer the crawl of your site. Put the right tag on a page and Opensolr leaves it out, or stops following its links, with no rule to write.
The tags it obeys
<meta name="robots" content="noindex">The page is read for its links, and never indexed. The HTTP header X-Robots-Tag: noindex does the same.
<meta name="robots" content="nofollow">The page is indexed, and none of its links are followed. Also as X-Robots-Tag: nofollow.
content="noindex, nofollow"The page is skipped completely: not indexed, links not followed.
<a href="..." rel="nofollow">That one link is not followed. The other links of the page are.
<link rel="canonical" href="...">When it points to another address, this copy is not indexed; the page it points to is indexed when the crawl reaches it. When it points to the page itself, nothing changes.
robots.txt
Opensolr crawls your own, verified site on your behalf, so it takes its orders from the tags above and from the rules you set in your Opensolr Index, not from robots.txt. To keep the crawl out of part of your site, add a Do not follow these rule for that site, or for every site in Global Exclusions. How: Rules per site.
In your server logs
The crawl visits your site like a recent Chrome browser on Windows, macOS or Linux, and sends a Chrome user agent. If a firewall or a bot filter in front of your site blocks the crawl, its pages come back as errors, and Crawl Stats lists them.
A site that answers with a browser password box needs its user name and password in the rules of that site.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates