Crawled URLs: Query Strings, www, Redirects, Duplicates

One page, one address

One page, one address

The same page often answers at several addresses: with and without www., with tracking parameters, with a slash at the end. Opensolr keeps every page under one clean address, so one page is not indexed twice under different addresses.

How an address is cleaned

FoundKept as
https://example.com/shoes?sort=pricehttps://example.com/shoes
https://example.com/shoes#reviewshttps://example.com/shoes
https://www.example.com/shoeshttps://example.com/shoes
https://example.com/Shoes/https://example.com/shoes
  • The query string, everything after ?, and the # part are removed. Only sitemap addresses keep their query string.
  • www. is dropped when it is the only part before your domain: www.example.com and example.com are one site, www.shop.example.com stays as it is.
  • The host and the folders are written in lowercase; a file name with an extension, like Report.pdf, keeps its case. The slash at the end is removed.

Redirects and canonical links

  • A page that redirects is kept under the address it lands on. The old address is remembered as a redirect and not read again.
  • A page that redirects to another site, outside the crawl mode, is not indexed and is listed in Crawl Stats under Left Domain.
  • A page whose canonical link points to another address is not indexed; that other address is indexed when the crawl reaches it.

Read again later, a page keeps its address, so it is updated in your index, never added twice.

What it means for your site

Pages that differ only by their query string, like /product?id=1 and /product?id=2, become one address and only one of them is kept, even when your sitemap lists them all. Give each page its own path, like /product/1, and each one becomes its own result.

Crawl your site: all pages

Back to Enterprise Site Search