Your site, read for you
Give Opensolr the address of your site. It reads your pages, your PDFs and office documents, and the text printed in the pictures inside your PDFs, and puts all of it into your Opensolr Index. Your visitors then search it on a ready search page, ask the chat bot about it, or your own code queries it. Nothing to install on your site, no code to write, and it can keep itself up to date.
You give the address. Opensolr does the reading, and your index is ready for search, chat and your own code.
Before you start
Your Opensolr Index has to be in a region with site search: only there does its Control Panel show the WebCrawler tab. Everything the index needs: What your index needs.
Fill your index in five steps
- Open the Control Panel of your Opensolr Index and click the WebCrawler tab.
- Under Crawl URLs, type the address of your site, starting with
https://, and click Add URL: Add start URLs. - Put the verification file on your site: Prove you own the site.
- Optional: click Settings to choose how far and how fast the crawl goes: Crawl settings.
- Click Start Crawl Schedule, then watch it in Crawl Stats.
Pages reach your index in batches while the crawl runs. The schedule keeps the crawl going on its own, and Keep Index Fresh reads every page again every few days.
Set it up
Your home page, sitemap or feed: where the crawl begins.
One small file on your site, checked automatically.
Threads, pause between pages, resume or start over, and what your plan sets.
Whole domain, one host or one folder, deep or shallow.
Read pages in a real browser, the way visitors see them.
PDFs page by page, office files, scans and pictures.
Choose what is left out
A password for a protected site, pages kept out, links never followed.
Rules for every site of the index, and the ones built in.
What the tags in your pages tell the crawl.
Run it and keep it fresh
The buttons that run the crawl, and the schedule.
Progress, errors and skipped pages, live.
Read your site again: all of it, from the start, or only the PDFs.
Every page read again on a schedule, dead pages removed.
A page read again the moment it changes.
How pages are read
The checks every page passes on its way into your index.
One page, one address: query strings, www, redirects.
Other ways to fill your index
Your data does not live on web pages, or you want it sent to your index when you save it? Push your data by API, or use the Drupal module or the WordPress plugin. Every way fills the same kind of Opensolr Index.