Data Ingestion Limits and Refusals

What is checked before a batch is queued

Before a batch is queued, Opensolr checks it against your index and your plan. A refused request queues nothing: status is false and msg says why. Fix it and send again.

What is checked

  • Batch. 1 to 50 documents per request.
    Answer: ERROR_DOCUMENTS_MUST_BE_NON_EMPTY_ARRAY, ERROR_BATCH_LIMIT_50_DOCUMENTS_MAX
  • Request size. The documents must fit the largest request the server takes.
    Answer: ERROR_PAYLOAD_TOO_LARGE_MAX_<size>, with the size in the message
  • Index size. Your index plus the batch must stay within the index size of your plan.
    Answer: ERROR_DISK_QUOTA_EXCEEDED, with current_mb and max_mb
  • Bandwidth. The index must not have used up its monthly bandwidth.
    Answer: ERROR_BANDWIDTH_LIMIT_EXCEEDED, with used_mb and max_mb
  • Document size. The title, description and text of one document together fit the document size of your plan (4.8 MB on the free plan). A file sent with "rtf": true is stopped at the same size when it is downloaded.
    Answer: VALIDATION_ERRORS: "exceeds your plan's max page size"
  • Fields. uri is an http(s) URL; title and description are there; text is there unless "rtf": true; file comes with "rtf": true and a job id from file_extract.
    Answer: VALIDATION_ERRORS, with an errors list such as Document [3]: description is required
  • Index. The index is yours and sits on a site search server.
    Answer: ERROR_NOT_CORE_OWNER, ERROR_CORE_NOT_WEBCRAWLER_ENABLED, ERROR_INVALID_CORE_NAME
  • Speed. By default every account may make 120 API requests a minute and 1,200 an hour.
    Answer: HTTP 429, ERROR_RATE_LIMIT_PER_MINUTE or ERROR_RATE_LIMIT_PER_HOUR
  • Address. Data Ingestion answers on api.opensolr.com only.
    Answer: WRONG_API_HOST, with the right address
  • Queue. The queue takes the batch.
    Answer: ERROR_QUEUE_UNAVAILABLE_TRY_AGAIN: send the same batch again a little later

The sizes of each plan: plans. Your index's use: Disk space and Bandwidth. Rate limits in full: API rate limits.

Not refusals

  • Duplicates. A document already waiting in the queue for this index, or sent twice in one batch, is skipped and the rest is queued. When every document was already waiting, the answer is All documents already queued with their number in dupes_skipped.
  • Your AI allowance. Past your monthly AI allowance, documents are still indexed and found by their words, only without meaning vectors.

Problems found while a job runs (a file that cannot be downloaded, a field your index refuses) do not refuse the request. They show in the job's result, with the reasons (The Data Ingestion queue).

Push your data pages

Back to Enterprise Site Search