Data Ingestion API - Push Documents to Solr

Push documents directly — no crawler needed

Data Ingestion API

The Data Ingestion API is part of the Web Crawler site search product: it lets you push documents into a Web Crawler index from your own application, database, or script. Instead of crawling a website, you send your content as structured JSON — and Opensolr takes care of the rest, including AI enrichment with embeddings, sentiment analysis, and language detection. For the complete API reference with code examples in multiple languages, see the Data Ingestion API guide.

Web Crawler indexes only

Data Ingestion works only on indexes created on our Web Crawler servers, with the Web Crawler config set. Any other index answers ERROR_CORE_NOT_WEBCRAWLER_ENABLED. A regular Opensolr Index takes documents through the standard Solr API (/update), like any Solr server.

On Drupal or WordPress? Start with the module instead.

On Drupal, install Opensolr Search or Search API Opensolr; on WordPress, the Opensolr Search plugin. You never call this API yourself: Opensolr Search for Drupal and the WordPress plugin push content on save, update and delete (translations, files and structured fields included), and Search API Opensolr indexes through Search API Solr. From your own code, the Laravel, LangChain, LlamaIndex, Haystack and MCP packages do it for you as well. This page is for custom apps and scripts on a Web Crawler index.

How Data Ingestion Works

POST /api/ { "title": "..." "text": "..." } Your Application JSON API Opensolr Receives & Validates Opensolr API AI Enrichment Embeddings (1024-dim) Sentiment Analysis Language Detection AI Enrichment Searchable Opensolr Index Solr Index Send JSON documents via API — AI enrichment happens automatically before indexing.

When to Use Data Ingestion vs. Web Crawler

Opensolr gives you two ways to get content into your index. Here is how to pick the right one:

Data Ingestion API

Best for: databases, product catalogs, CMS content, custom applications, programmatic content, or anything that is not a public website. You control exactly what data goes into the index.

Web Crawler

Best for: public websites, blogs, documentation sites, or any content that is accessible via a URL. The crawler discovers pages automatically — no coding required.

You can use both!

Data Ingestion and the Web Crawler can feed the same index. For example, you might crawl your public website AND push product data from your database via the API. Both end up searchable in the same index.

API Endpoint

To push documents, send a POST request to the Data Ingestion endpoint. Here is an example using curl:

# Push a batch of documents to your Opensolr index curl -X POST "https://api.opensolr.com/v1/ingest/YOUR_INDEX_NAME" \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "documents": [ { "uri": "https://example.com/product/123", "title": "Wireless Bluetooth Headphones", "description": "Noise-cancelling over-ear headphones with 30-hour battery", "text": "Full product details and specifications go here...", "category": "Electronics", "price_f": 79.99, "currency_s": "USD" } ] }'

Replace YOUR_INDEX_NAME with your Opensolr index name and YOUR_API_KEY with your API key (found in your Opensolr dashboard). Need to index PDFs, Word documents, or other binary files? See Document Extraction (extractOnly) for how to use the rtf=true parameter. The file type is detected from the file itself (PDF, Word, Excel, PowerPoint, OpenDocument, RTF, text, HTML).

Scanned PDFs and pictures inside documents are read too. Every picture inside a PDF is read with OCR in the language of its page, detected automatically, in more than 100 languages. A picture with no text in it, such as a photo or a chart, gets a short description of what it shows, so it can be found by its meaning. A PDF read this way is indexed one Solr document per page (PDFs are indexed page by page); each page holds its own text first, then what was read in its pictures. Each picture read costs 0.5 of an AI request and each picture described another 0.5; anything read before is free (API Quota). See it live on 33 test PDFs.

PDFs Are Indexed Page by Page

A PDF sent with rtf:true is not stored as one big document. Opensolr reads it and writes one Solr document per page. Each page is searched, ranked, highlighted and vectorised on its own, so a query finds the page that answers it instead of a 200-page file that mentions it somewhere. Every other document (Word, Excel, HTML, text, and any document you send with its text already filled in) is still one document, exactly as you sent it.

What each page document holds

FieldValue on a page
idThe MD5 of the document's id followed by #page= and the page number: md5("<document id>#page=7"). The document's id is the id you sent, or the MD5 of the normalised uri when you sent none.
uri_id_sThe document's id, the same on every page of that PDF. This is the field that ties the pages of one PDF together.
page_iThe page number in the whole file, starting at 1.
textThe text of that page only: its own text, then the text read by OCR in its pictures (or the description of a picture without text).
meta_detected_language, meta_languageThe language detected on that page. When the page is too short to be sure, the document's language: the one you sent, else the one detected on the document.
embeddingsA vector of the page: your title and description plus the page's text.
phonetic_textOnly on a page with text read by OCR: a phonetic copy of the page's text, so a word OCR misread by a letter still matches.
sizeThe page's size: bytes of title, description and page text.
Every other fieldCopied from your document: uri, title, description, dates and every custom field you sent are on every page.

Finding the pages of a PDF

The document id you sent is no longer a document in the index: its pages are. To work with one PDF, filter on uri_id_s:

curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select?q=*:*&fq=uri_id_s:%22DOCUMENT_ID%22&fl=id,page_i,title&rows=1000&wt=json"

One page of it: fq=uri_id_s:"DOCUMENT_ID" and fq=page_i:7. A search returns pages; the hosted search page shows Page N on each and opens the PDF at that page, and on its Documents tab every PDF is one result, with its best pages under it.

One result per PDF, with its best pages

Since every page is a document, a search can return several pages of the same PDF. To show one result per PDF with its best pages under it, group on uri_id_s. Solr does it natively, two ways.

Collapse and expand. The results list stays flat: one entry per PDF (its best page), and an expanded section holds the other matching pages of each PDF, keyed by uri_id_s, with how many of them matched in its numFound. numFound, paging with start and rows, and facet counts all count PDFs. This is what the hosted search page uses.

curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select" \ --data-urlencode 'q=how big are meteors' \ --data-urlencode 'defType=edismax' \ --data-urlencode 'qf=title^3 text' \ --data-urlencode 'fq={!collapse field=uri_id_s nullPolicy=expand sort=$qgsort}' \ --data-urlencode 'qgsort=score desc,page_i asc' \ --data-urlencode 'expand=true' \ --data-urlencode 'expand.rows=4' \ --data-urlencode 'expand.sort=score desc,page_i asc' \ --data-urlencode 'fl=id,uri,title,page_i,score' \ --data-urlencode 'rows=10' \ --data-urlencode 'wt=json'
ParameterWhat it does
nullPolicy=expandEvery document without uri_id_s (web pages, Word files, PDFs that are not paged) stays its own result.
sort=$qgsort, expand.sortWhich page stands for the PDF and in what order its other pages come: the most relevant first, the page number breaking ties. With a sort of your own (by date, by price), put it first: timestamp desc,page_i asc. Without a query (q=*:*), use page_i asc to list the pages in order.
expand.rows=4Up to four more pages per PDF, so five pages in all.

On a hybrid query. The same parameters work with {!hybrid}. Give the collapse sort by reference, as above: {!hybrid} picks its candidates with your filter queries but leaves out any filter that refers to a parameter named q…, so it ranks every page and the collapse groups them afterwards. With the sort written inside the filter (sort='score desc'), the candidates would already be collapsed and the expanded section would come back nearly empty.

All the pages of one PDF, in the same order, to list more than five: the same query without the collapse and expand parameters, with fq=uri_id_s:"DOCUMENT_ID" and sort=score desc,page_i asc, paged with start and rows.

Grouping. The response is a list of groups, one per PDF, each with its matching pages in order of relevance and the total number of pages that matched. group.ngroups=true adds the number of PDFs that matched.

curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/select" \ --data-urlencode 'q=how big are meteors' \ --data-urlencode 'defType=edismax' \ --data-urlencode 'qf=title^3 text' \ --data-urlencode 'group=true' \ --data-urlencode 'group.field=uri_id_s' \ --data-urlencode 'group.limit=5' \ --data-urlencode 'group.ngroups=true' \ --data-urlencode 'fl=id,uri,title,page_i,score' \ --data-urlencode 'rows=10' \ --data-urlencode 'wt=json'

With grouping, every document without uri_id_s lands in one shared group whose groupValue is null; collapse with nullPolicy=expand keeps them as separate results instead.

Grouping works on keyword (edismax) queries; on hybrid queries use collapse and expand. To read a PDF from its first page to its last: fq=uri_id_s:"DOCUMENT_ID" with sort=page_i asc. See it on the live OCR demo.

Sending a PDF again

Before a PDF's pages are written, every earlier page of that PDF and any earlier whole-document version of it are removed (id or uri_id_s equal to the document's id). Re-sending a PDF that now has fewer pages leaves no pages behind. A PDF that cannot be read, or has no page with text, is indexed as one document with the text you sent.

What it costs

Each page gets its own vector: one AI request per page vectorised; a page whose content was vectorised before is free. Pictures read by OCR count 0.5 each, as before (API Quota).

Your schema must keep these fields

page_i and uri_id_s are stored through the *_i and *_s dynamic fields of the Opensolr schema, and phonetic_text through phonetic_*. On an index with its own schema that lacks them, and that sends unknown fields to an ignored catch-all (<dynamicField name="*" type="ignored"/>, as the Drupal Search API schema does), the pages are still written but these fields are dropped: no page number, no way to find all pages of a PDF, and no removal of earlier pages when the PDF is sent again with fewer pages. Add the three dynamic fields to such a schema before sending PDFs with rtf:true.

Document Structure

Each document you push is a JSON object with fields. Some fields are required, others are optional. Here is the full reference:

Required Fields

Every document must include these fields:

Field Type Description
uri String A unique identifier for the document. Usually a URL or a unique ID like product-123. This is also used for deduplication.
title String The title of the document. Displayed in search results as the clickable heading.
description String A short summary or excerpt. Displayed below the title in search results.
text String The full text content of the document. This is the main body that gets searched. Can be as long as needed.

Optional Fields

Add these fields to enrich your documents with extra metadata:

Field Type Description
category String A category label (e.g., "Electronics", "Blog", "FAQ"). Useful for faceted filtering.
author String The author name. Useful for filtering or displaying in search results.
og_image String (URL) An image URL to display as a thumbnail in search results.
price_f Float A numeric price value. The _f suffix tells Opensolr this is a float (decimal number).
currency_s String Currency code (e.g., "USD", "EUR"). The _s suffix marks it as a string field.
meta_* Varies Any field starting with meta_ is stored as custom metadata. Use appropriate type suffixes.

Field Type Suffixes

When you create custom fields, add a suffix to tell Opensolr what type of data the field contains. This ensures proper indexing and enables features like sorting and filtering.

Suffix Data Type Example _s String (text) brand_s: "Nike" _ss Multi-value string tags_ss: ["red","xl"] _f Float (decimal) price_f: 29.99 _i Integer (whole number) stock_i: 142 _dt Date/Time published_dt: "2026-01-15T10:30:00Z"
Dynamic fields are automatic

You do not need to define these fields in your schema ahead of time. Just use the right suffix and Opensolr automatically handles the field type. For example, sending rating_f: 4.5 automatically creates a float field called rating_f. For every suffix, what it enables (facet, sort, range, full-text) and the rules behind them, see the Vector Search Schema Reference; the older Index Field Reference lists the same fields in FAQ form.

Batch Limit: 50 Documents Per Request

Each API request can contain up to 50 documents in the documents array. If you have more than 50 documents to push, split them into multiple requests. There is no limit on how many requests you can make — just keep each batch at 50 or fewer.

AI Enrichment Pipeline

Every document you push through the Data Ingestion API is automatically enriched by AI before it hits your index. This happens behind the scenes — you do not need to do anything extra.

Raw Document Your JSON input [0.23, -0.87, 0.15, 0.42, -0.31, 0.78, ... 1024 dimensions 0.56, 0.12, -0.94] Embeddings 1024-dimensional vector Positive / Negative / Neutral scoring Sentiment Tone analysis EN DE ... FR ES ... Auto-detected Language Auto-detection

Here is what each enrichment step does:

  • Embeddings — your document's text is converted into a 1024-dimensional numerical vector. This enables vector/semantic search — finding documents by meaning, not just keywords.
  • Sentiment Analysis — the text is analyzed to determine its overall tone: positive, negative, or neutral. Useful for filtering reviews, feedback, or social media content.
  • Language Detection — the language of the text is automatically identified. This enables language-aware search features and helps with multilingual indexes.

Ingestion Queue Dashboard

When you push documents, they enter an ingestion queue where they are processed one batch at a time. The dashboard shows you exactly what is happening with each batch.

Job States

Each ingestion job goes through these states:

Pending Processing Completed or Failed Queued AI enriching In your index Needs attention

Jobs can also be in a Paused state if you manually pause them from the dashboard.

Available Actions

Run

Start processing a pending or paused job immediately.

Pause

Temporarily halt a running job. It will stay in the queue and can be resumed.

Resume

Continue processing a paused job from where it left off.

Retry

Re-process a failed job. Useful if the failure was caused by a temporary issue.

Delete

Remove a job from the queue entirely. The documents will not be indexed.

Progress Monitoring

While a job is processing, the dashboard shows a progress bar and counts: how many documents have been enriched and indexed out of the total. You can watch it in real time. For detailed documentation on all queue actions including pause, resume, retry, and monitoring, see Ingestion Queue Management.

Automatic Deduplication

Opensolr automatically prevents duplicate documents using your uri field. Here is how it works:

uri = "example.com/page/1" md5(uri) a1b2c3d4e5f6... Unique Document ID Same URI = Same Doc If you push a document with a URI that already exists, the old version is replaced — not duplicated. This means you can safely re-push updated content without worrying about duplicates.
Safe to re-push

You can push the same document as many times as you want. As long as the uri stays the same, Opensolr will update the existing document instead of creating a duplicate. This makes it safe to run your ingestion script repeatedly.

A PDF sent with rtf:true follows the same rule one level down: its pages carry ids derived from the document's id and the page number, so re-pushing the PDF replaces its pages (PDFs are indexed page by page).

Step-by-Step: Pushing Your First Document

Here is a quick walkthrough to get your first document into your index via the API:

  1. Get your API key — find it in your Opensolr Dashboard under your account settings.
  2. Know your index name — this is the name you gave your index when you created it (e.g., my-products).
  3. Prepare your document as a JSON object with at least the four required fields: uri, title, description, and text.
  4. Send a POST request to https://api.opensolr.com/v1/ingest/YOUR_INDEX_NAME with your API key in the Authorization header and your document in the request body (see the code example above).
  5. Check the response — a successful push returns a confirmation with the job ID. You can track the job in the Ingestion Queue Dashboard.
  6. Search for your document — once the job completes (usually within seconds), your document is live and searchable in your index.

Related Documentation