Faceting, sorting and grouping need a column of values per field. docValues keep that column on disk; without them Solr rebuilds it in the heap.
Faceting, sorting, grouping and function queries all need the same thing: the value of a field for a given document. The inverted index answers the opposite question — which documents contain a given value — so Solr has to get there somehow. It has two ways, and they differ by where the data ends up.
docValues is a column written at index time and read straight off disk. Without it, Solr rebuilds that column in the heap at query time.
The measurement in that diagram is real. One 5 million document index on our fleet, whose fields were declared without docValues, was holding 357.9 MB of un-inverted field data on the heap across 149 entries, the largest single entry being 161.7 MB for one field on one segment. None of that memory is reclaimable while the searcher is open. The same index answers the first facet request on such a field in 285 ms and the second in 53 ms — the 232 ms difference is the un-inverting, paid again after every commit that opens a new searcher.
Declare docValues on everything you facet, sort or group by
It moves the column out of the heap and onto disk, where the operating system caches it and can also release it. It costs index size, which is far cheaper than heap.
docValues cannot be added in place
Changing the docValues setting of an existing field requires a reindex — Solr will refuse to open an index whose segments disagree. See Cannot change DocValues type.
Do not put it on a full text field
Faceting on analyzed text is almost never what you want anyway: you would be faceting on stems and tokens, not on values.
That last point is also a hard limit, and it is the single most common mistake in a schema. A field whose type is text_general — that is, solr.TextField — cannot carry docValues="true". Solr does not warn about it and does not ignore it: the core refuses to load, with This field type does not support doc values. Copy the text into a string field instead, or use solr.SortableTextField, which is a text field that keeps a docValues copy of the original value alongside the analyzed one.
The declaration on the left does not degrade gracefully. It stops the core from starting.
<!-- WRONG: solr.TextField does not support doc values, the core will not load --> <field name="brand" type="text_general" indexed="true" stored="true" docValues="true"/> <!-- RIGHT, option 1: search the analyzed field, facet and sort on a string copy --> <field name="brand" type="text_general" indexed="true" stored="true"/> <field name="brand_facet" type="string" indexed="true" stored="false" docValues="true"/> <copyField source="brand" dest="brand_facet"/> <!-- RIGHT, option 2: one field that does both, docValues capped at maxCharsForDocValues --> <fieldType name="text_sortable" class="solr.SortableTextField" positionIncrementGap="100" maxCharsForDocValues="1024"> <analyzer> <tokenizer class="solr.StandardTokenizerFactory"/> <filter class="solr.LowerCaseFilterFactory"/> </analyzer> </fieldType> <field name="title" type="text_sortable" indexed="true" stored="true" docValues="true"/>
storeOffsetsWithPositions="true" on the stored field. Term vectors work too, but they add a second copy of the term data to the index and cost far more disk and IO than the highlighter needs.<!-- enough for the default highlighter, without the weight of term vectors --> <field name="body" type="text_general" indexed="true" stored="true" storeOffsetsWithPositions="true"/>
Solr best practices
docValues and the TextField trap (this page)