The Plugin Sitemap

The list of pages the Data Crawler reads

The plugin builds its own sitemap at https://your-site/opensolr-sitemap.xml. It is the list of pages the Data Crawler reads, so what is in it is what gets crawled.

What it lists

  • The published items of the content types ticked under Content Types for Crawling on the Data Crawler tab: What gets crawled.
  • With Include attached files: the document files of your Media Library (PDF, Word, Excel, PowerPoint, OpenDocument, text and CSV).
  • For each address, the date it was last modified.
  • On a site translated with WPML or Polylang, the addresses of the translations of each post.

The search page itself is never listed.

Large sites

Each sitemap file holds up to 1,000 addresses. Up to 1,000, /opensolr-sitemap.xml lists them directly. Above that, it becomes an index of files named /opensolr-sitemap-1.xml, /opensolr-sitemap-2.xml and so on, so every request stays light for your server whatever the size of your site.

How Opensolr knows about it

The plugin adds the sitemap to the crawl list of your index when you connect your account and every time a crawl starts, when your site address starts with https://. The request is signed with your API key, so Opensolr accepts it without a verification file. The Data Crawler tab shows how many addresses the sitemap holds: "N URLs in sitemap".

A site without HTTPS can add the sitemap address by hand in the Opensolr Control Panel: Add start URLs and Prove you own the site.

SEO & Sitemap