Global Exclusions: Pages Left Out of Every Crawl

Rules for every site of the index, and the ones built in

Pages left out everywhere

Global Exclusions holds rules for every site your Opensolr Index crawls. Use it for what is about content rather than about one site: login pages, downloads, a section every site of yours has.

Set them

  1. In Crawler Setup, click Global Exclusions. The dialog Crawl settings for this index opens.
  2. Fill in Do not index these and Do not follow these, one rule per line, written the same way as the rules of one site.
  3. Click Save settings. The button turns to the accent colour while the index has global rules.

The rules add up: a page is left out when the built-in rules, your Global Exclusions or the rules of its site say so. A rule can only leave more out, never bring a page back.

What every crawl leaves out on its own

You never need a rule for these.

Never fetched

  • Files a search has no use for: stylesheets, scripts, audio and video, archives and installers, fonts.
  • Picture files linked from your pages. The pictures inside your PDFs are still read: Documents on your site.
  • Links that are not web pages: javascript:, mailto:, tel:, data:.
  • Addresses more than ten levels deep, a sign of a loop rather than a page.
  • Documents, when Content Types is set to HTML Only.

Read for their links, never indexed

  • Feed pages, with /rss, /feed, /feeds or /atom in the address, and sitemaps and feeds themselves.
  • Print versions under /print/.

Crawl your site: all pages

Back to Enterprise Site Search