Meta Robots, Nofollow, Canonical and robots.txt

What the tags in your pages tell the crawl

What your pages tell the crawl

The tags your pages already carry for search engines also steer the crawl of your site. Put the right tag on a page and Opensolr leaves it out, or stops following its links, with no rule to write.

The tags it obeys

<meta name="robots" content="noindex">

The page is read for its links, and never indexed. The HTTP header X-Robots-Tag: noindex does the same.

<meta name="robots" content="nofollow">

The page is indexed, and none of its links are followed. Also as X-Robots-Tag: nofollow.

content="noindex, nofollow"

The page is skipped completely: not indexed, links not followed.

<a href="..." rel="nofollow">

That one link is not followed. The other links of the page are.

<link rel="canonical" href="...">

When it points to another address, this copy is not indexed; the page it points to is indexed when the crawl reaches it. When it points to the page itself, nothing changes.

robots.txt

Opensolr crawls your own, verified site on your behalf, so it takes its orders from the tags above and from the rules you set in your Opensolr Index, not from robots.txt. To keep the crawl out of part of your site, add a Do not follow these rule for that site, or for every site in Global Exclusions. How: Rules per site.

In your server logs

The crawl visits your site like a recent Chrome browser on Windows, macOS or Linux, and sends a Chrome user agent. If a firewall or a bot filter in front of your site blocks the crawl, its pages come back as errors, and Crawl Stats lists them.

A site that answers with a browser password box needs its user name and password in the rules of that site.

Crawl your site: all pages

Back to Enterprise Site Search