One Search Result per PDF

Each PDF once, with its best pages under it

Every page of a PDF is its own document (PDFs page by page), so one search can return several pages of the same PDF. To show each PDF once, with its best pages under it, group the results on uri_id_s. Solr does it natively, and the Documents tab of the hosted search page already does it for you.

This page is for your own search code. The examples query your index directly; its address and its user and password are on the Connection URL page.

Collapse and expand (recommended)

The results stay one flat list with one entry per PDF: its best page. An expanded section holds the next best pages of each PDF, keyed by uri_id_s. numFound, paging with start and rows, and facet counts all count PDFs. This is what the hosted search page uses.

curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select" \
  --data-urlencode 'q=how big are meteors' \
  --data-urlencode 'defType=edismax' \
  --data-urlencode 'qf=title^3 text' \
  --data-urlencode 'fq={!collapse field=uri_id_s nullPolicy=expand sort=$qgsort}' \
  --data-urlencode 'qgsort=score desc,page_i asc' \
  --data-urlencode 'expand=true' \
  --data-urlencode 'expand.rows=4' \
  --data-urlencode 'expand.sort=score desc,page_i asc' \
  --data-urlencode 'fl=id,uri,title,page_i,score' \
  --data-urlencode 'rows=10' \
  --data-urlencode 'wt=json'
ParameterWhat it does
nullPolicy=expandEvery document without uri_id_s (a web page, a Word file) stays a result of its own.
sort=$qgsort, expand.sortWhich page stands for the PDF, and the order of its other pages: the most relevant first, the page number breaking ties. With a sort of your own (by date, by price), put it first: timestamp desc,page_i asc. Without a query (q=*:*), page_i asc lists the pages in order.
expand.rows=4Up to four more pages per PDF, five in all.

With hybrid search ({!hybrid}) the same parameters work. Keep the collapse sort as a reference (sort=$qgsort), as above: then hybrid search ranks every page first and the collapse groups them afterwards. With the sort written inside the filter, the pages would be grouped before ranking and the expanded section would come back nearly empty.

All the pages of one PDF, to list more than five: the same query without the collapse and expand parameters, with fq=uri_id_s:"PDF_ID" and sort=score desc,page_i asc, paged with start and rows.

Grouping

The answer is a list of groups, one per PDF, each with its matching pages in order of relevance and how many pages matched. group.ngroups=true adds how many PDFs matched.

curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select" \
  --data-urlencode 'q=how big are meteors' \
  --data-urlencode 'defType=edismax' \
  --data-urlencode 'qf=title^3 text' \
  --data-urlencode 'group=true' \
  --data-urlencode 'group.field=uri_id_s' \
  --data-urlencode 'group.limit=5' \
  --data-urlencode 'group.ngroups=true' \
  --data-urlencode 'fl=id,uri,title,page_i,score' \
  --data-urlencode 'rows=10' \
  --data-urlencode 'wt=json'

With grouping, every document without uri_id_s lands in one shared group whose groupValue is null; collapse with nullPolicy=expand keeps them as separate results instead. For hybrid search, use collapse and expand.

See it on real PDFs: OCR search demo.

Push your data pages

Back to Enterprise Site Search