Reindexing feeds the new index from the same source that fed the old one. Then you adjust the application, validate, and switch.
This is the part of the migration we cannot do for you. Reindexing means feeding documents into the new index from the same source that originally fed the old one. Concretely:
- If you use the site search crawler: start a crawl on the new index (crawl your site). Same start URL, same mode, same threads. The crawler will populate it from scratch — it does not need the old index at all.
- If you use the Drupal Opensolr Search module: in the new index settings, click Save & Connect to bind the module to the new index, then run Ingest All Now from the Data Ingestion tab.
- If you use the WordPress Opensolr Search plugin: same pattern — Save & Connect against the new index, then bulk-ingest from the Data Ingestion tab.
- If you use Drupal Search API Solr: upload the new server config zip, point the Search API server at the new index, then
drush search-api:reset-tracker+drush search-api:index. - If you have a custom indexer (cron job, message queue consumer, ETL script): change the connection string and run it against the new index. The shape of the docs you push is unchanged.
- If you push via the Data Ingestion API: change the target index name in your client; the API endpoint URL stays the same.
Fallback: dump from old index, replay into new index
Only consider this if your source-of-truth is genuinely gone (or it would take longer to rebuild than the retirement deadline allows). The script below paginates the old index via cursorMark, strips two internal fields that must not be re-imported (_version_, _root_), and POSTs each batch into the new index. Requires jq.
Caveats — read before using this script:
- Only fields where
stored="true"on the OLD schema can be dumped. Index-only fields are gone. - Fields the new schema does not declare are silently dropped — or rejected, depending on
schemalesssettings. - Computed-at-index-time fields (copyField targets, language detection, sentiment, embeddings, custom analyzers) cannot be reproduced this way — they will be missing or stale.
- Block-join parent/child relationships are lost (the
_root_field is stripped). - Vector embeddings from any old
embeddingsfield will not be regenerated — if you want hybrid search you still need to re-embed from text.
A real reindex from source-of-truth fixes all of this. The dump script does not.
# OPTIONAL FALLBACK ONLY - if you have NO source-of-truth. # Read the warnings in <a href="https://opensolr.com/learn/solr-migrations/345/reindex-and-cut-over-to-the-new-solr-index" style="color:#c05520 !important;font-weight:600 !important;text-decoration:underline !important;text-underline-offset:3px !important">how to reindex</a> first. Recommended path is to re-ingest from your CMS / DB / file store. OLD="https://old-cluster.example.com/solr/OLD_CORE" NEW="https://new-cluster.example.com/solr/NEW_CORE" AUTH_OLD="opensolr:OLD_INDEX_API_KEY" AUTH_NEW="opensolr:NEW_INDEX_API_KEY" BATCH=2000 CURSOR="*" while : ; do RESP=$(curl -s -u "$AUTH_OLD" \ --data-urlencode "q=*:*" \ --data-urlencode "rows=$BATCH" \ --data-urlencode "sort=id asc" \ --data-urlencode "fl=*" \ --data-urlencode "wt=json" \ --data-urlencode "cursorMark=$CURSOR" \ "$OLD/select") DOCS=$(echo "$RESP" | jq -c '.response.docs | map(del(._version_, ._root_))') [ "$DOCS" = "[]" ] && break curl -s -u "$AUTH_NEW" -H "Content-Type: application/json" \ -X POST -d "$DOCS" \ "$NEW/update?commitWithin=10000" NEXT=$(echo "$RESP" | jq -r .nextCursorMark) [ "$NEXT" = "$CURSOR" ] && break CURSOR="$NEXT" done # Final hard commit on the new core curl -s -u "$AUTH_NEW" "$NEW/update?commit=true"
Why strip _version_: Solr uses it for optimistic concurrency control. Carrying old values forward causes spurious 409 conflicts on the new index. Let Solr assign fresh ones.
Why strip _root_: Block-join parent linkage that is meaningless without simultaneously reindexing the children in the same payload.
For very large indexes (50M+ docs): split the cursor loop into N parallel workers by partitioning on the first character of id. Each worker runs the same script with an extra fq=id:0* filter. Throughput scales close to linearly until the destination Solr write side saturates.
Application-side changes
- SolrJ: bump to
9.xor8.11.xmatching the server. Mixed major versions over HTTP work for basic ops but fail on streaming expressions and SolrCloud collection ops. - Drupal Search API Solr: regenerate the config zip via
drush search-api-solr:get-server-config <server>, upload via the Configuration tab on the new index. Search API knows about leader/follower terminology and emits the right configset for the target Solr major. - Drupal Opensolr Search module: point at the new index name in Settings → Save & Connect.
- WordPress Opensolr Search plugin: same — new index name in Settings, Save & Connect.
- Custom
/selectqueries withqt=: replace with calls to/HANDLERdirectly —handleSelectdispatch is going away. - Custom queries with
defaultOperator: passq.op=ANDper request, or set in handler defaults. - Custom queries that reference
defaultSearchField: passdf=fieldnameper request. - Range queries on numerics: PointField range syntax (
field:[1 TO 10]) is identical at query time. No client change needed.
Validation and cutover
- Doc count parity. Hit
/select?q=*:*&rows=0on both indexes.numFoundshould match exactly. Differences are usually new docs added to the old index during the parallel period — rerun the reindex if you need parity. - Spot-check sample documents. Pull 20–50 random IDs and compare every field side by side.
/select?q=id:DOC_ID&wt=json&fl=*on both, diff the JSON. - Top-N relevance check. Pick 10–20 of your most common production queries (from analytics, or just the ones you care about). Hit both indexes, compare top 10 results. Order may differ slightly because of BM25 tweaks between versions — that is normal.
- Faceting smoke test. Run any production query that uses faceting against the new index. If a numeric or date facet returns zero buckets where it used to return many, you forgot
docValues="true"on a PointField — fix in schema, reload core, reindex. - Cut over. Point your application at the new index. Keep the old index running and read-only for a quiet observation window (24–48h is plenty for most apps), then delete from the control panel. Old server slot is freed.
- Hard deadline. The retirement date communicated in your email is the absolute cutoff. Old index stops serving on that date regardless — plan accordingly.
Solr migration
Reindex, application changes and cutover (this page)