Solr docValues and the TextField Trap

Account Resources

Faceting, sorting and grouping need a column of values per field. docValues keep that column on disk; without them Solr rebuilds it in the heap.

Faceting, sorting, grouping and function queries all need the same thing: the value of a field for a given document. The inverted index answers the opposite question — which documents contain a given value — so Solr has to get there somehow. It has two ways, and they differ by where the data ends up.

Facet, sort orgroup on a fieldField has docValuescolumn stored on diskMemory mapped,nothing on the heapField has noneSolr un-inverts itThe whole columnsits on the heapMeasured on one 5 million document index: 357.9 MB of un-inverted field data, 149 entries.

docValues is a column written at index time and read straight off disk. Without it, Solr rebuilds that column in the heap at query time.

The measurement in that diagram is real. One 5 million document index on our fleet, whose fields were declared without docValues, was holding 357.9 MB of un-inverted field data on the heap across 149 entries, the largest single entry being 161.7 MB for one field on one segment. None of that memory is reclaimable while the searcher is open. The same index answers the first facet request on such a field in 285 ms and the second in 53 ms — the 232 ms difference is the un-inverting, paid again after every commit that opens a new searcher.

Declare docValues on everything you facet, sort or group by

It moves the column out of the heap and onto disk, where the operating system caches it and can also release it. It costs index size, which is far cheaper than heap.

docValues cannot be added in place

Changing the docValues setting of an existing field requires a reindex — Solr will refuse to open an index whose segments disagree. See Cannot change DocValues type.

Do not put it on a full text field

Faceting on analyzed text is almost never what you want anyway: you would be faceting on stems and tokens, not on values.

That last point is also a hard limit, and it is the single most common mistake in a schema. A field whose type is text_general — that is, solr.TextField — cannot carry docValues="true". Solr does not warn about it and does not ignore it: the core refuses to load, with This field type does not support doc values. Copy the text into a string field instead, or use solr.SortableTextField, which is a text field that keeps a docValues copy of the original value alongside the analyzed one.

FIELD DECLARATIONtype="text_general"docValues="true"The core will not loadThis field type does not support doc valuesTHE FIXcopyField the text intoa string fielddocValues works thereand so do facets, sorting and grouping

The declaration on the left does not degrade gracefully. It stops the core from starting.

<!-- WRONG: solr.TextField does not support doc values, the core will not load -->
<field name="brand" type="text_general" indexed="true" stored="true" docValues="true"/>

<!-- RIGHT, option 1: search the analyzed field, facet and sort on a string copy -->
<field     name="brand"       type="text_general" indexed="true"  stored="true"/>
<field     name="brand_facet" type="string"       indexed="true"  stored="false" docValues="true"/>
<copyField source="brand"     dest="brand_facet"/>

<!-- RIGHT, option 2: one field that does both, docValues capped at maxCharsForDocValues -->
<fieldType name="text_sortable" class="solr.SortableTextField" positionIncrementGap="100"
           maxCharsForDocValues="1024">
  <analyzer>
    <tokenizer class="solr.StandardTokenizerFactory"/>
    <filter class="solr.LowerCaseFilterFactory"/>
  </analyzer>
</fieldType>
<field name="title" type="text_sortable" indexed="true" stored="true" docValues="true"/>
Highlighting is the exception that does not want docValues. The UnifiedHighlighter, the default in Solr 9, only needs offsets, and the cheap way to give it those is storeOffsetsWithPositions="true" on the stored field. Term vectors work too, but they add a second copy of the term data to the index and cost far more disk and IO than the highlighter needs.
<!-- enough for the default highlighter, without the weight of term vectors -->
<field name="body" type="text_general" indexed="true" stored="true"
       storeOffsetsWithPositions="true"/>

Solr best practices

docValues and the TextField trap (this page)