Pages left out everywhere
Global Exclusions holds rules for every site your Opensolr Index crawls. Use it for what is about content rather than about one site: login pages, downloads, a section every site of yours has.
Set them
- In Crawler Setup, click Global Exclusions. The dialog Crawl settings for this index opens.
- Fill in Do not index these and Do not follow these, one rule per line, written the same way as the rules of one site.
- Click Save settings. The button turns to the accent colour while the index has global rules.
The rules add up: a page is left out when the built-in rules, your Global Exclusions or the rules of its site say so. A rule can only leave more out, never bring a page back.
What every crawl leaves out on its own
You never need a rule for these.
Never fetched
- Files a search has no use for: stylesheets, scripts, audio and video, archives and installers, fonts.
- Picture files linked from your pages. The pictures inside your PDFs are still read: Documents on your site.
- Links that are not web pages:
javascript:,mailto:,tel:,data:. - Addresses more than ten levels deep, a sign of a loop rather than a page.
- Documents, when Content Types is set to HTML Only.
Read for their links, never indexed
- Feed pages, with
/rss,/feed,/feedsor/atomin the address, and sitemaps and feeds themselves. - Print versions under
/print/.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates