Crawl settings
One dialog holds the settings of the crawl: how many pages are read at the same time, how far the crawl goes, what it reads, and how gently it treats your server. The defaults suit most sites, so you can start without touching them.
Open them
In the WebCrawler tab, click Settings in the Quick Access bar, or Edit on the summary line of Crawler Setup. The Index Settings dialog opens; the settings are under Crawler Configuration.
The settings
How many pages are read at the same time, from 1 up to the most your plan allows. More threads finish sooner and put more load on your server.
How far the crawl goes: every host of your domain, one host or one folder, deep or shallow. The default is Follow HOST Links. The six modes.
HTML & Documents (pdf, docx, etc...), the default, or HTML Only. Documents on your site.
Curl (Fast), the default, or Chrome (JS Rendering) for pages built with JavaScript. Sites built with JavaScript.
Resume Crawl, shown with the date of the last crawl, carries on where the crawl stopped. Start Over forgets which pages were already read and begins again from your start URLs. Your Opensolr Index keeps its pages either way. Before the first crawl, only Start Over is offered.
How long the crawl waits between two pages: 0.1 to 0.9 seconds, or 1 to 7 seconds. A longer pause is gentler on a busy server.
The summary line in Crawler Setup shows the current values: Threads, Mode, Renderer, Content, Resume, Pause and Refresh.
Save them
- Click Save Settings.
- If the crawl is already running, click Stop Crawl Schedule, then Start Crawl Schedule, so the new settings take effect.
The crawl schedule always carries on with Resume. Start Over applies when you start a crawl yourself.
What your plan sets
Your plan decides how many pages one Opensolr Index can take from the crawl, how much the crawl may download, the largest page or file it reads, and the most CPU threads. When a crawl reaches the page limit or the download limit, it stops. Crawl Stats shows where you stand, and the plans are on https://opensolr.com/pricing.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates