Crawl Your Site Into Your Opensolr Index

Your pages, documents and pictures in your Opensolr Index, with no code

Your site, read for you

Give Opensolr the address of your site. It reads your pages, your PDFs and office documents, and the text printed in the pictures inside your PDFs, and puts all of it into your Opensolr Index. Your visitors then search it on a ready search page, ask the chat bot about it, or your own code queries it. Nothing to install on your site, no code to write, and it can keep itself up to date.

Your site, its pages and documents, are read by Opensolr into your Opensolr Index, which is ready for a search page, a chat bot and the API Your site Opensolr reads Your Opensolr Index Ready to use pages, sitemap PDFs, documents text and links text in pictures kept up to date search page chat bot, API
Your site, its pages and documents, are read by Opensolr into your Opensolr Index, which is ready for a search page, a chat bot and the API Your site Opensolr reads Your Opensolr Index Ready to use pages, sitemap PDFs, documents text and links text in pictures kept up to date search page chat bot, API

You give the address. Opensolr does the reading, and your index is ready for search, chat and your own code.

Before you start

Your Opensolr Index has to be in a region with site search: only there does its Control Panel show the WebCrawler tab. Everything the index needs: What your index needs.

Fill your index in five steps

  1. Open the Control Panel of your Opensolr Index and click the WebCrawler tab.
  2. Under Crawl URLs, type the address of your site, starting with https://, and click Add URL: Add start URLs.
  3. Put the verification file on your site: Prove you own the site.
  4. Optional: click Settings to choose how far and how fast the crawl goes: Crawl settings.
  5. Click Start Crawl Schedule, then watch it in Crawl Stats.

Pages reach your index in batches while the crawl runs. The schedule keeps the crawl going on its own, and Keep Index Fresh reads every page again every few days.

Set it up

Add start URLs

Your home page, sitemap or feed: where the crawl begins.

Prove you own the site

One small file on your site, checked automatically.

Crawl settings

Threads, pause between pages, resume or start over, and what your plan sets.

How far the crawl goes

Whole domain, one host or one folder, deep or shallow.

Sites built with JavaScript

Read pages in a real browser, the way visitors see them.

Documents on your site

PDFs page by page, office files, scans and pictures.

Choose what is left out

Rules per site

A password for a protected site, pages kept out, links never followed.

Pages left out everywhere

Rules for every site of the index, and the ones built in.

Meta robots, nofollow, canonical

What the tags in your pages tell the crawl.

Run it and keep it fresh

Start, pause, stop, flush

The buttons that run the crawl, and the schedule.

Crawl Stats

Progress, errors and skipped pages, live.

Reindex

Read your site again: all of it, from the start, or only the PDFs.

Keep your index fresh

Every page read again on a schedule, dead pages removed.

Recrawl from your CMS

A page read again the moment it changes.

How pages are read

What a page needs

The checks every page passes on its way into your index.

URLs and duplicates

One page, one address: query strings, www, redirects.

Other ways to fill your index

Your data does not live on web pages, or you want it sent to your index when you save it? Push your data by API, or use the Drupal module or the WordPress plugin. Every way fills the same kind of Opensolr Index.

Back to Enterprise Site Search