API - Push documents (ingest)

Data Ingestion

Sends up to 50 documents to your Opensolr Index in one call. The batch goes into the Data Ingestion queue and is written in the background, each document enriched the same way as a crawled page.

Endpointhttps://api.opensolr.com/solr_manager/api/ingest
MethodPOST with a JSON body (or form fields)
Authemail and api_key of your account, as GET or POST parameters: Authentication.
NeedsAn index in a region with site search: What your index needs.

01 · Parameters

core_nameRequired

The name of your index (index_name works too).

documentsRequired

1 to 50 documents: a JSON array, or a JSON string in a form field.

payload_fileOptional

Instead of documents: a JSON file sent as multipart, holding the documents, or the whole request: Send a JSON file.

02 · A document

uriRequired

The http or https address of the document. It identifies it: the document id is made from it, and sending the same uri again replaces the document.

titleRequired

The title.

descriptionRequired

A short summary.

textRequired unless rtf

The full text.

rtfOptional

true: Opensolr downloads uri and reads the text itself, PDF, Word and more included: Files and documents.

More fields, such as prices, dates and your own fields: Document fields.

03 · Example

curl -s -X POST "https://api.opensolr.com/solr_manager/api/ingest" \
  -H "Content-Type: application/json" \
  -d '{
    "email": "YOUR_EMAIL",
    "api_key": "YOUR_API_KEY",
    "core_name": "my_index",
    "documents": [
      {"uri": "https://www.example.com/products/blue-chair", "title": "Blue chair",
       "description": "A solid oak chair in blue.", "text": "The full product text..."},
      {"uri": "https://www.example.com/manuals/blue-chair.pdf", "title": "Blue chair manual",
       "description": "Assembly and care.", "rtf": true}
    ]
  }'

04 · Answer

{
    "status": true,
    "msg": "QUEUED",
    "job_id": "9f2c41d0b7e84a1c8e3f5a6b7c8d9e0f",
    "total_docs": 2,
    "doc_ids": [
        "...",
        "..."
    ]
}

Keep job_id to follow the job: ingest_status. A document whose uri is already waiting in the queue for this index is skipped; when every document was, the answer is "msg": "All documents already queued" with dupes_skipped.

05 · Errors

VALIDATION_ERRORSHTTP 200

One or more documents break the rules above. errors lists each one by its position.

ERROR_DOCUMENTS_MUST_BE_NON_EMPTY_ARRAYHTTP 200

No documents.

ERROR_BATCH_LIMIT_50_DOCUMENTS_MAXHTTP 200

More than 50 documents.

ERROR_PAYLOAD_TOO_LARGE_MAX_...HTTP 200

The batch is larger than the server accepts. The code ends with that size, for example ERROR_PAYLOAD_TOO_LARGE_MAX_128MB: send fewer documents per call.

ERROR_DISK_QUOTA_EXCEEDEDHTTP 200

The index would go over its disk limit (current_mb, max_mb).

ERROR_BANDWIDTH_LIMIT_EXCEEDEDHTTP 200

The index used its traffic for this month (used_mb, max_mb).

ERROR_CORE_NOT_WEBCRAWLER_ENABLEDHTTP 200

The index is not in a region with site search. The message names the regions that have it.

ERROR_NOT_CORE_OWNERHTTP 200

The index is not yours.

ERROR_QUEUE_UNAVAILABLE_TRY_AGAINHTTP 200

The queue did not answer. Send the batch again.

WRONG_API_HOSTHTTP 404

Called on opensolr.com. The answer carries the right URL.

ERROR_AUTHENTICATION_FAILEDHTTP 403

The email and API key do not match.

ERROR_RATE_LIMIT_PER_MINUTEHTTP 429

Too many calls. Wait for the Retry-After seconds: Rate limits.

All the limits, per document and per plan: Limits and refusals. What the enrichment and the vectors use from your plan: What it uses from your plan.

06 · Related

The queue: ingest_queue. Update and remove documents: Update and remove.

Every error code and HTTP status of the API: API errors. Calls per minute and per hour: Rate limits.