What a Page Needs to Be Indexed by the Crawl

The checks every page passes on its way in

What a page needs

A page reaches your Opensolr Index when it passes a few simple checks. When a page you expected is missing from your search, this list tells you where to look.

The checks

The crawl can reach it

It is linked from a page the crawl reads, listed in your sitemap or feed, or is a start URL, and it sits inside the crawl mode. It is on the public internet.

It answers well

An answer below 400: a page that loads, or a redirect. A redirected page is kept under the address it lands on. An answer of 400 or more, like 404 or 500, is not indexed.

It is a page or a document

A web page, or one of the documents Opensolr reads when Content Types includes documents. Sitemaps and feeds are read for their links, never indexed.

It has a title or a description

The title comes from the page itself. A page without one gets the first words of its text as a title. A page with no title, no description and no text is skipped.

It fits your plan

A page or file larger than your plan allows is skipped and listed in Crawl Stats under Oversize Pages.

Nothing leaves it out

No noindex tag or header and no canonical link to another address (tags), no Do not index or Do not follow rule (rules), and not a "checking your browser" page.

A page that passed before and now answers 400 or more is removed from your index the next time the crawl reads it.

Still missing?

  1. Open Crawl Stats and look for the address in the lists: errors, Oversize Pages, Noindex Filtered, Out of Scope, Left Domain.
  2. Check that the address you search for is the one the crawl keeps: URLs and duplicates.
  3. For a page built with JavaScript, switch the renderer: Sites built with JavaScript.

Crawl your site: all pages

Back to Enterprise Site Search