Fields Opensolr adds to every document
A pushed document needs four fields: uri, title, description and text. Everything else in the table below is worked out by Opensolr before the document reaches your index, the same way for pushed documents and for pages read by the crawl. You never have to send these fields.
The fields, and whose value wins
| Field | What it holds | Your value |
|---|---|---|
id | MD5 of the uri with its trailing slash removed, so sending the same uri again updates the document instead of adding a second one. Documents sent by the Opensolr Search module for Drupal keep their entity id. | Ignored |
uri_s, title_s, author_s | Exact-value copies of uri, title and author, for facets and exact filters. | Overwritten |
tags, title_tags | The cleaned words of the title, the description and the first 3,000 characters of the text (title_tags: the title alone). They feed the suggestions while typing. | Overwritten |
spell | The same words as tags, the dictionary behind "Did you mean". | Overwritten |
embeddings | The meaning vector of the title, the description and the first whole sentences of the text (of text_t when it has one). Left out when the plan of the index owner has no AI, or the monthly AI allowance is used up. | Ignored |
sent_pos, sent_neu, sent_neg, sent_com | The sentiment of the same text. | Overwritten |
meta_detected_language, meta_language | The language you send when it is a real language code; otherwise the one detected from the title and description, or the one in the URL path. | Kept if valid |
meta_og_locale | The locale you send when it names a real country, otherwise the one in the URL path; left out when neither is known. | Kept if valid |
creation_date, meta_creation_date | Taken from timestamp when you send it (Unix time or any readable date). | Kept |
size | The size of the title, description and text: bytes for a pushed document, kilobytes for a page read by the crawl. | Overwritten |
meta_domain | The host of the uri, without www. | Kept if sent |
content_type | text/html by default; the real file type for a file sent with rtf:true. | Kept |
category, og_image, meta_icon | Set to an empty value when missing, so the search page never meets an absent field. | Kept |
page_i, uri_id_s | A PDF read page by page becomes one document per page: page_i is the page number and uri_id_s the id of the PDF, the same on all its pages (fq=uri_id_s:"<id>" returns them all). A page's id is the MD5 of the PDF's id and the page number. Its text, size and meaning vector are the page's own; every other field, the language included, is the PDF's. | Overwritten |
phonetic_text | On a PDF page whose text was partly read from its pictures: a sounds-like copy of the page's text, so a word misread by one letter still matches. | Overwritten |
quality_f | Pages read by the crawl only: a 0 to 1 score of how rich the content is, used by the quality boost. A document without it counts as 0.5. | Kept if sent |
The page fields need the *_i, *_s and phonetic_* dynamic fields of the Opensolr schema. A custom schema that sends unknown fields to an ignored catch-all drops them silently, and the pages lose their number and their link to the PDF.
Any other field in your document reaches the index unchanged. That is why the type suffixes work: nothing is mapped or renamed, the name decides the type (Add any field with a suffix).