A PDF you send as a file ("rtf": true, or "file") is not stored as one big document. Opensolr writes one document per page. Each page is searched, ranked, highlighted and given a meaning vector on its own, so a search lands on the page that answers it instead of a 200-page file that mentions it somewhere.
What each page holds
| Field | On a page |
|---|---|
id | Its own id: the MD5 of the PDF's id, #page= and the page number, like md5("<PDF id>#page=7"). |
uri_id_s | The PDF's id (the one in doc_ids when you pushed), the same on every page. This is what ties the pages of one PDF together. |
page_i | The page number, from 1. |
text | The text of that page only: its own text, then the words read in its pictures (or what a picture without words shows). |
| Language | The PDF's language, the same on every page. |
| Meaning vector | Made from your title and description plus the page's text. |
phonetic_text | Only on a page with words read by OCR: a sound-alike copy of its text, so a word OCR misread by one letter still matches. |
| Everything else | Copied from your document on every page: uri, title, description, dates and your own fields. |
A page with no text at all is left out. Only PDFs read from their file are split. Every other file (Word, Excel, slides) and every document whose text you send stays one document. A PDF with no readable page is indexed as one document with the text you sent.
Find the pages of a PDF
Filter on uri_id_s with the PDF's id. Your index's address and its user and password: Connection URL.
curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select?q=*:*&fq=uri_id_s:%22PDF_ID%22&fl=id,page_i,title&sort=page_i%20asc&rows=100&wt=json"
One page of it: add fq=page_i:7. The hosted search page shows Page N on each such result and opens the PDF at that page. To show each PDF once with its best pages: One result per PDF.
Send a PDF again
Before the new pages are written, every earlier page of that PDF and any earlier one-document version of it are removed. A PDF that now has fewer pages leaves nothing behind. If that removal fails, the PDF is not written and the job reports it as failed.
What it uses from your plan
Every page gets its own meaning vector and counts against your monthly AI allowance like one document; a page whose words were vectorised before is free. Pictures read with OCR or described count too: What it uses from your plan.
page_i, uri_id_s and phonetic_text are stored through the *_i, *_s and phonetic_* fields of the Opensolr schema. A schema of your own without them refuses the pages or drops those fields, and then pages cannot be found together or replaced when the PDF comes again (Dynamic fields).