PDFs Indexed Page by Page

Every page of a PDF is a result of its own

A PDF you send as a file ("rtf": true, or "file") is not stored as one big document. Opensolr writes one document per page. Each page is searched, ranked, highlighted and given a meaning vector on its own, so a search lands on the page that answers it instead of a 200-page file that mentions it somewhere.

What each page holds

FieldOn a page
idIts own id: the MD5 of the PDF's id, #page= and the page number, like md5("<PDF id>#page=7").
uri_id_sThe PDF's id (the one in doc_ids when you pushed), the same on every page. This is what ties the pages of one PDF together.
page_iThe page number, from 1.
textThe text of that page only: its own text, then the words read in its pictures (or what a picture without words shows).
LanguageThe PDF's language, the same on every page.
Meaning vectorMade from your title and description plus the page's text.
phonetic_textOnly on a page with words read by OCR: a sound-alike copy of its text, so a word OCR misread by one letter still matches.
Everything elseCopied from your document on every page: uri, title, description, dates and your own fields.

A page with no text at all is left out. Only PDFs read from their file are split. Every other file (Word, Excel, slides) and every document whose text you send stays one document. A PDF with no readable page is indexed as one document with the text you sent.

Find the pages of a PDF

Filter on uri_id_s with the PDF's id. Your index's address and its user and password: Connection URL.

curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select?q=*:*&fq=uri_id_s:%22PDF_ID%22&fl=id,page_i,title&sort=page_i%20asc&rows=100&wt=json"

One page of it: add fq=page_i:7. The hosted search page shows Page N on each such result and opens the PDF at that page. To show each PDF once with its best pages: One result per PDF.

Send a PDF again

Before the new pages are written, every earlier page of that PDF and any earlier one-document version of it are removed. A PDF that now has fewer pages leaves nothing behind. If that removal fails, the PDF is not written and the job reports it as failed.

What it uses from your plan

Every page gets its own meaning vector and counts against your monthly AI allowance like one document; a page whose words were vectorised before is free. Pictures read with OCR or described count too: What it uses from your plan.

Keep the Opensolr schema

page_i, uri_id_s and phonetic_text are stored through the *_i, *_s and phonetic_* fields of the Opensolr schema. A schema of your own without them refuses the pages or drops those fields, and then pages cannot be found together or replaced when the PDF comes again (Dynamic fields).

Push your data pages

Back to Enterprise Site Search