Solr Deep Paging and cursorMark

Account Resources

Deep paging is the most expensive request a search page can receive, and the easiest one for a bot to send.

Asking for start=500000&rows=10 does not skip half a million documents. Solr has to find, score and rank every one of them to know which ten come next, and it holds a queue of 500,010 entries while it does. Every shard does this independently. The cost grows with the offset, and it grows on both CPU and heap.

QTIME IN MILLISECONDS02000400060005362,8605,98862 ms at every depthstart plus rowscursorMark, same page size010,000100,000500,0001,000,000start offset, rows=1000, sorted by id

Measured on a 5,029,434 document index, Solr 9. Both series use the same page size, the same sort and the same field list. The cursorMark series was measured across 100 consecutive pages, 100,000 documents deep.

Offsetstart plus rowscursorMark
074 ms62 ms
10,000111 ms62 ms
100,000536 ms62 ms
500,0002,860 msunchanged by depth
1,000,0005,988 msunchanged by depth

Two rules follow from that curve. For a user interface, keep start under 50,000 — nobody browses to page five thousand, and the requests that do are almost always crawlers walking your paginator. For exporting or re-processing an entire result set, never page with start at all: use cursorMark, which carries the position inside the sort instead of counting from the beginning.

GET /solr/mycore/select?q=category:shoes&sort=id asc&rows=1000&fl=id,title&cursorMark=*

# the response carries nextCursorMark; send it back as cursorMark for the next page,
# and stop when nextCursorMark comes back identical to the one you sent.
# the sort must end in a unique field, id asc being the usual choice.
This is not theoretical. One index on our platform received 144 identical requests for start=950500 in eight minutes, each taking between 17 and 72 seconds, each one also faceting twelve fields with facet.limit=-1. That was 3,850 CPU seconds on four cores and a load average of 38. The origin was not an attack — it was the site's own application retrying a page deep in a public catalogue after the first request timed out. Opensolr now watches for exactly this pattern and mails the index owner when it appears.

Solr best practices

Deep paging and cursorMark (this page)