Recrawl a Page the Moment It Changes in Your CMS

Changed pages read again on the next run

A changed page, read again on the next run

Your CMS knows the moment a page changes. It can tell Opensolr, so the next run of the crawl schedule reads that page again, instead of waiting for Keep Index Fresh.

With Drupal

The Opensolr Search module does it for you, with no setup:

  • It adds your sitemap as a start URL with a call signed by your API key, so the site is verified at once, with no file to upload.
  • When you save a page, it tells Opensolr to forget that page; the next run of the schedule reads it again from your sitemap, with the new content.
  • When you unpublish or delete a page, it removes the page from your Opensolr Index.

More: Drupal: Opensolr Search module.

With WordPress

The Opensolr Search plugin adds your sitemap as a start URL the same signed way, so no verification file is needed: WordPress: Opensolr Search plugin.

With any other CMS

Two API calls do the same from your own code. Both carry a signature: HMAC-SHA256 of the page address followed by the index name, made with your API key.

add_crawl_url_signed

Adds an https:// address, usually your sitemap, as a start URL, already verified. Refused when the address is not https, not on the public internet, or over the number of start URLs of your plan. Reference.

delete_crawler_url

Makes the crawl forget one page, so the next run reads it again from your sitemap. The page stays in your index until then, and the Crawl URLs list does not change. Reference.

Both rely on a running crawl schedule: the page is read on its next run. To send a page to your index when you save it, without waiting for a crawl run, push it instead: Push your data.

Crawl your site: all pages

Back to Enterprise Site Search