Where the crawl begins
A start URL is the address the crawl of your site begins from. One per site is enough: Opensolr follows the links from there, and every page it reaches within the crawl mode goes into your Opensolr Index.
Add one
- In the WebCrawler tab of your Opensolr Index, find the Crawl URLs panel.
- Type the address, like
https://example.comorhttps://example.com/sitemap.xml. - Click Add URL.
- Prove you own the site, once.
Only https:// addresses are accepted, on the public internet. Your plan sets how many start URLs you can have across all your Opensolr Indexes: https://opensolr.com/pricing.
What to add
The surest way to reach every page. A sitemap index is followed to every sitemap it lists, also sitemaps split into pages, in every crawl mode.
The links are followed from there, as far as the crawl mode allows.
An RSS or Atom feed: every page it lists is crawled. Handy for new articles nothing links to yet. The feed itself is never a search result.
Add one start URL per site. Start URLs on the same host are crawled as one site and share the rules of that site.
The list
Each start URL is a row with four columns: URL, State (Active or Inactive), Verified (Yes or No) and Settings, a cog that opens the rules of that site.
- Tick rows, then click Delete, Activate or Deactivate. An inactive start URL stays in the list and is left out of the next crawl.
- URL Verification shows where to put your verification file.
A new start URL is Active and not yet Verified. A crawl reads only the start URLs that are both Active and Verified.
Crawl your site: all pages
- Crawl your site overview
- Add start URLs
- Prove you own the site
- Crawl settings
- How far the crawl goes
- Sites built with JavaScript
- Documents on your site
- Rules per site
- Pages left out everywhere
- Meta robots, nofollow, canonical
- Start, pause, stop, flush
- Crawl Stats
- Reindex
- Keep your index fresh
- Recrawl from your CMS
- What a page needs
- URLs and duplicates