Crawl Rules per Site: Basic Auth, Do Not Index, Do Not Follow

A password, pages kept out, links never followed

Rules for one site

Every start URL has a cog in its Settings column. It opens the rules for that site: a password for a site behind a login box, pages to keep out of search, links never to follow, and how much of each page shapes its meaning.

Open them

In Crawl URLs, click the cog on the row of the site. The dialog Crawl settings for this URL opens. Change what you need and click Save settings. The cog turns to the accent colour when the site has rules of its own.

The four settings

Include detected body in embedder

Off by default. The title, the description and the structured data of a page always shape its meaning, the part hybrid search uses. This adds the rest of the page text. Switch it on when your pages carry details nothing else does, like product specifications. Documents are not affected.

HTTP Basic Auth

Optional. A user name and a password for a site that answers with a browser password box (error 401), like a staging site or an intranet. They are sent with every request to that host, in Chrome mode too. Leave the user empty to remove them.

Do not index these

One rule per line. A matching page is still read and its links are still followed, but it never reaches your index. Good for print versions and paginated listings.

Do not follow these

One rule per line. A matching link is never fetched, so it costs no time and no traffic. Good for baskets, logins, filter combinations and endless calendars.

Do not index: the page is read and its links are followed, but it is never indexed. Do not follow: the link is dropped at once and never indexed. Do not index Do not follow Page is read Links followed Link found Dropped at once Never indexed Never indexed
Do not index: the page is read and its links are followed, but it is never indexed. Do not follow: the link is dropped at once and never indexed. Do not index Do not follow Page is read Links followed Link found Dropped at once Never indexed Never indexed

Both keep a page out of search. Do not index still walks through the page to the pages it links to; Do not follow stops there.

Writing a rule

Each rule is a regular expression with its delimiters, matched against the address of the page in lowercase, without the part after ?. In #/cart#i, the # marks the start and the end and i ignores case: it matches every address with /cart in it.

#/print/#i
#/page/[0-9]+#i
#/cart|/checkout|/login#i

A rule that is not a valid expression is refused when you save, and the dialog names it, so a broken rule never runs silently.

Good to know

  • The rules belong to the host. Several start URLs on the same host are crawled as one site: their Do not index and Do not follow rules are combined, Include detected body in embedder is on as soon as one of them has it on, and they should all carry the same HTTP Basic Auth.
  • They are added to the rules every crawl has and to your Global Exclusions; they never replace them.
  • They apply from the next run of the crawl schedule.
  • Include detected body in embedder applies to pages crawled after you save. To apply it to what is already indexed, use Reindex From Scratch.
  • Pages kept out by Do not index are counted in Crawl Stats under Noindex Filtered.

Crawl your site: all pages

Back to Enterprise Site Search