What a page needs
A page reaches your Opensolr Index when it passes a few simple checks. When a page you expected is missing from your search, this list tells you where to look.
The checks
It is linked from a page the crawl reads, listed in your sitemap or feed, or is a start URL, and it sits inside the crawl mode. It is on the public internet.
An answer below 400: a page that loads, or a redirect. A redirected page is kept under the address it lands on. An answer of 400 or more, like 404 or 500, is not indexed.
A web page, or one of the documents Opensolr reads when Content Types includes documents. Sitemaps and feeds are read for their links, never indexed.
The title comes from the page itself. A page without one gets the first words of its text as a title. A page with no title, no description and no text is skipped.
A page or file larger than your plan allows is skipped and listed in Crawl Stats under Oversize Pages.
A page that passed before and now answers 400 or more is removed from your index the next time the crawl reads it.
Still missing?
- Open Crawl Stats and look for the address in the lists: errors, Oversize Pages, Noindex Filtered, Out of Scope, Left Domain.
- Check that the address you search for is the one the crawl keeps: URLs and duplicates.
- For a page built with JavaScript, switch the renderer: Sites built with JavaScript.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates