How text is analysed
A text field is never matched against the raw value. When a document is indexed, its text is turned into a list of cleaned-up terms; when someone searches, their words go through almost the same steps, and the two lists are compared. That is why running finds run, café finds cafe, and a synonym you add works at once.
text_general: the main search type
Used by uri, title, description, text, author and every *_t field. The steps at indexing time, in order:
- 01
HTMLStripCharFilterRemoves any HTML tags left in the value. - 02
MappingCharFilterFolds accents withmapping-ISOLatin1Accent.txt: é becomes e, ß becomes ss. - 03
ICUTokenizerSplits the text into words with Unicode rules, so Chinese, Japanese, Thai and mixed scripts split correctly. - 04
CJKWidthFilterMakes full-width and half-width CJK characters the same. - 05
EnglishPossessiveFilterOpensolr's becomes Opensolr. - 06
ASCIIFoldingFilterA second accent pass for anything the mapping file missed. - 07
StopFilterDrops the words listed instopwords.txt. - 08
WordDelimiterGraphFilterWi-Fi becomes wi, fi and wifi; iPhone15 becomes iphone, 15 and iphone15. The original word is kept too; words inprotwords.txtare left whole. - 09
LowerCaseFilterEverything in lower case. - 10
SnowballPorterFilterEnglish stemming: running, runs and runner share the stem run. Words inprotwords.txtare not stemmed. - 11
LengthFilterDrops terms longer than 500 characters (encoded blobs, broken markup). - 12
RemoveDuplicatesTokenFilterRemoves terms that ended up identical at the same position.
At search time the steps are the same with one more: SynonymGraphFilter adds the synonyms from synonyms.txt to the words searched. Because synonyms apply only at search time, a synonym you save works on the next search, with no reindex. All the word lists are edited in Configuration.
Autocomplete: edgy_text_kw and edgy_text_ws
At indexing time every term is also kept as all its beginnings, 1 to 25 characters long: runner becomes r, ru, run, runn, runne, runner. At search time what you type is matched as a beginning. The difference between the two:
edgy_text_kwtags, title_tags.edgy_text_wstags_ws, title_tags_ws.The suggestions of the hosted search page, while visitors type, search these four fields.
Spelling: textSpell
Deliberately simple: words split and put in lower case, nothing else. No stemming and no stopwords, so the dictionary holds the real words of your content. "Did you mean" reads the spell field, allows up to 2 changed letters between a typed word and a suggestion, and leaves alone a word found in more than half of your documents: a word that common is taken as spelled right.
Sounds like: text_general_phonetic
Used by any phonetic_* field. Next to each term it keeps a sound code (Double Metaphone), so Smyth matches Smith. Opensolr fills phonetic_text on PDF pages whose text was partly read from pictures, so a word misread by one letter still matches.
The meaning vector: knn_vector
A vector of 1,024 numbers compared by cosine similarity. The size is fixed by the model Opensolr uses to make the vectors; a vector of any other length is refused. It is searched with {!knn f=embeddings topK=N} or, on Opensolr, with the {!hybrid} parser that ranks words and meaning together: Hybrid search.
Plain types
| Type | Solr class | What it holds |
|---|---|---|
string | StrField | One exact value, upper and lower case kept; documents without it sort last. |
boolean | BoolField | true or false; 1, t and T also count as true. |
pint, plong, pfloat, pdouble | Point fields | Numbers with fast range queries. The old names int, tint, float and the like map to the same classes, for old configurations. |
pdate | DatePointField | Dates in ISO 8601 UTC, with date math: [NOW-7DAYS TO NOW], NOW/DAY. |
location | LatLonPointSpatialField | "lat,lon" points. {!geofilt sfield=coords pt=44.43,26.10 d=25} keeps what is within 25 km. |
ignored | StrField, not indexed, not stored | Accepted and thrown away, for the ignored_* and *_nlp fields. |
More examples of stemming, synonyms and accent folding: How text analysis works.