One page, one address
The same page often answers at several addresses: with and without www., with tracking parameters, with a slash at the end. Opensolr keeps every page under one clean address, so one page is not indexed twice under different addresses.
How an address is cleaned
https://example.com/shoes?sort=pricehttps://example.com/shoeshttps://example.com/shoes#reviewshttps://example.com/shoeshttps://www.example.com/shoeshttps://example.com/shoeshttps://example.com/Shoes/https://example.com/shoes- The query string, everything after
?, and the#part are removed. Only sitemap addresses keep their query string. www.is dropped when it is the only part before your domain:www.example.comandexample.comare one site,www.shop.example.comstays as it is.- The host and the folders are written in lowercase; a file name with an extension, like
Report.pdf, keeps its case. The slash at the end is removed.
Redirects and canonical links
- A page that redirects is kept under the address it lands on. The old address is remembered as a redirect and not read again.
- A page that redirects to another site, outside the crawl mode, is not indexed and is listed in Crawl Stats under Left Domain.
- A page whose canonical link points to another address is not indexed; that other address is indexed when the crawl reaches it.
Read again later, a page keeps its address, so it is updated in your index, never added twice.
What it means for your site
Pages that differ only by their query string, like /product?id=1 and /product?id=2, become one address and only one of them is kept, even when your sitemap lists them all. Give each page its own path, like /product/1, and each one becomes its own result.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates