Indexing into Solr Without a Memory Spike

Account Resources

Indexing problems are almost always commit problems. A commit per document is a new searcher per document.

Send documents in batches, not one at a time

A few hundred to a few thousand documents per request is the working range. One request per document spends most of its time on connection and commit overhead.

Commit on a schedule, not per write

Configure autoCommit with openSearcher=false for durability and autoSoftCommit for visibility. Explicit commits from the client are what turn a bulk load into an hour of searcher churn.

Do not optimize a live index out of habit

A forced merge rewrites the entire index into one segment. It is IO, disk and page cache for the duration, and Solr handles merging by itself. Run it after a bulk load if at all, never on a schedule.

Keep stored fields to what you return

Fields that are only searched should be stored="false". It shrinks the index, the documentCache entries and the response all at once — see how to save disk space for your index.

Solr best practices

Indexing without a memory spike (this page)