How Text Is Analysed - Stemming, Accents, Autocomplete

Why running finds run and cafe finds café

How text is analysed

A text field is never matched against the raw value. When a document is indexed, its text is turned into a list of cleaned-up terms; when someone searches, their words go through almost the same steps, and the two lists are compared. That is why running finds run, café finds cafe, and a synonym you add works at once.

text_general: the main search type

Used by uri, title, description, text, author and every *_t field. The steps at indexing time, in order:

  1. 01HTMLStripCharFilterRemoves any HTML tags left in the value.
  2. 02MappingCharFilterFolds accents with mapping-ISOLatin1Accent.txt: é becomes e, ß becomes ss.
  3. 03ICUTokenizerSplits the text into words with Unicode rules, so Chinese, Japanese, Thai and mixed scripts split correctly.
  4. 04CJKWidthFilterMakes full-width and half-width CJK characters the same.
  5. 05EnglishPossessiveFilterOpensolr's becomes Opensolr.
  6. 06ASCIIFoldingFilterA second accent pass for anything the mapping file missed.
  7. 07StopFilterDrops the words listed in stopwords.txt.
  8. 08WordDelimiterGraphFilterWi-Fi becomes wi, fi and wifi; iPhone15 becomes iphone, 15 and iphone15. The original word is kept too; words in protwords.txt are left whole.
  9. 09LowerCaseFilterEverything in lower case.
  10. 10SnowballPorterFilterEnglish stemming: running, runs and runner share the stem run. Words in protwords.txt are not stemmed.
  11. 11LengthFilterDrops terms longer than 500 characters (encoded blobs, broken markup).
  12. 12RemoveDuplicatesTokenFilterRemoves terms that ended up identical at the same position.

At search time the steps are the same with one more: SynonymGraphFilter adds the synonyms from synonyms.txt to the words searched. Because synonyms apply only at search time, a synonym you save works on the next search, with no reindex. All the word lists are edited in Configuration.

Autocomplete: edgy_text_kw and edgy_text_ws

At indexing time every term is also kept as all its beginnings, 1 to 25 characters long: runner becomes r, ru, run, runn, runne, runner. At search time what you type is matched as a beginning. The difference between the two:

edgy_text_kw
The whole value is one phrase, so it matches the start of a full title. Fields: tags, title_tags.
edgy_text_ws
Each word is matched on its own, so typing run finds Trail Runner on its second word. Fields: tags_ws, title_tags_ws.

The suggestions of the hosted search page, while visitors type, search these four fields.

Spelling: textSpell

Deliberately simple: words split and put in lower case, nothing else. No stemming and no stopwords, so the dictionary holds the real words of your content. "Did you mean" reads the spell field, allows up to 2 changed letters between a typed word and a suggestion, and leaves alone a word found in more than half of your documents: a word that common is taken as spelled right.

Sounds like: text_general_phonetic

Used by any phonetic_* field. Next to each term it keeps a sound code (Double Metaphone), so Smyth matches Smith. Opensolr fills phonetic_text on PDF pages whose text was partly read from pictures, so a word misread by one letter still matches.

The meaning vector: knn_vector

A vector of 1,024 numbers compared by cosine similarity. The size is fixed by the model Opensolr uses to make the vectors; a vector of any other length is refused. It is searched with {!knn f=embeddings topK=N} or, on Opensolr, with the {!hybrid} parser that ranks words and meaning together: Hybrid search.

Plain types

TypeSolr classWhat it holds
stringStrFieldOne exact value, upper and lower case kept; documents without it sort last.
booleanBoolFieldtrue or false; 1, t and T also count as true.
pint, plong, pfloat, pdoublePoint fieldsNumbers with fast range queries. The old names int, tint, float and the like map to the same classes, for old configurations.
pdateDatePointFieldDates in ISO 8601 UTC, with date math: [NOW-7DAYS TO NOW], NOW/DAY.
locationLatLonPointSpatialField"lat,lon" points. {!geofilt sfield=coords pt=44.43,26.10 d=25} keeps what is within 25 km.
ignoredStrField, not indexed, not storedAccepted and thrown away, for the ignored_* and *_nlp fields.

More examples of stemming, synonyms and accent folding: How text analysis works.

Schema reference pages

Back to Enterprise Site Search