Document Extraction (rtf:true)
Index PDFs, Word documents, spreadsheets, presentations, and other rich document formats through the Data Ingestion API. Just add rtf:true to any document in your payload and point uri at the file. Opensolr fetches it, detects the real type from the file's bytes, extracts the text, reads the pictures and scanned pages inside it with OCR, and indexes it with all the same enrichment as any other document.
How It Works
Your App
POST withrtf:true
+ document URI
Validate
Check URI, MIME type, file size
Fetch
Download file from URI
Extract
Detect format, extract plaintext
Enrich
Embeddings, sentiment, language, derived fields
Index
Searchable in Solr
The extracted text fills the text field automatically. You provide the title, description, and any other metadata you have. The full enrichment pipeline (embeddings, sentiment, language detection, derived fields) runs on the extracted text just like any other document.
Supported Formats
Example Payload
// Mix regular docs and RTF docs in the same batch:
{
"email": "you@example.com",
"api_key": "your_api_key",
"core_name": "my_index",
"documents": [
{
// Regular document — you provide the text
"title": "Product Announcement",
"description": "New features for Q1 2026",
"text": "We are excited to announce...",
"uri": "https://example.com/blog/announcement"
},
{
// RTF document — text extracted from the PDF automatically
"rtf": true,
"title": "2025 Annual Report",
"description": "Company financials and key metrics",
"uri": "https://example.com/docs/annual-report-2025.pdf",
"timestamp": 1735689600,
"category": "Reports"
},
{
// RTF document — Word file from an internal server
"rtf": true,
"title": "Employee Handbook",
"description": "Company policies and procedures",
"uri": "https://intranet.example.com/hr/handbook.docx",
"og_image": "https://example.com/img/handbook-cover.png"
}
]
}
What Happens
- The
rtf:trueflag is detected on the document - The file at
uriis fetched over HTTP/HTTPS - The file type is detected from the actual content, never from the URL or the file name: a Word file served as a generic download, a spreadsheet with no extension, a PDF behind a document ID all work
- Text is extracted with a reader for that format: every page of a PDF; the paragraphs, tables, text boxes, headers and footers of a Word file; every sheet of a spreadsheet; every slide and its notes
- Pictures inside the document, scanned pages included, are read with OCR and their words are added to the text
- The extracted text populates the
textfield - All other fields you provided (
title,description,timestamp, etc.) are kept as-is - The full enrichment pipeline runs: embeddings, sentiment, language detection, derived fields
- The document is pushed to your Solr index
rtf:true documents in the same batch. Regular docs use the text you provide. RTF docs extract text from the file. Both go through the same enrichment pipeline.
Scanned Documents and Images (OCR)
On every plan, the extraction reads more than the text layer of a file:
- A PDF is indexed one document per page: each page document has the PDF's
uri,title,descriptionand every field you sent, its owntext(the page's own text first, then what was read in its pictures), its own language and its own vector, pluspage_i(the page number) anduri_id_s(the id of the PDF, the same on all its pages). A search finds the very page that answers it, and the hosted search page links straight to that page of the PDF. A scanned PDF (pages that are only pictures) gives the reading of every page - Sending the same PDF again replaces all of its pages: the earlier pages, and an earlier whole-document version, are removed before the new pages are written
- Each page's
idismd5("<document id>#page=N"), where the document id is theidyou sent (or the MD5 of the normaliseduri). The document id itself is no longer a document in the index: its pages are, anduri_id_sholds it on each of them. All pages of one PDF:fq=uri_id_s:"<document id>"; one page: addfq=page_i:N - On each page, the language is detected on that page's text; when it cannot be, it is the language you sent, else the one detected on the document
- These fields are kept by the
*_i,*_sandphonetic_*dynamic fields of the Opensolr schema. A custom schema that sends unknown fields to anignoredcatch-all (the Drupal Search API schema does) drops them: the pages are written, but without their page number and without the link to their PDF. Add those dynamic fields to such a schema first - Each picture of a PDF is read in the language of its page, detected automatically, among more than 100 languages; when no language is certain, English
- A picture of a PDF with no text in it (a photo, a chart without labels) gets a short description of what it shows, in English, so it can be found by its meaning
- The pictures inside a Word, Excel, PowerPoint or OpenDocument file are extracted and read, and their words are added at the end of
textunder anImagesheading - An image file (PNG, JPEG, GIF, WEBP, TIFF, BMP) sent as the document is read whole
- Pictures smaller than 300 pixels on a side are skipped (icons, bullets, logos), and a picture repeated on every page of a PDF is read once
text: indexed, highlighted and vectorised exactly like typed text, no separate field to query. A page read by OCR is also indexed phonetically (phonetic_text), so on the hosted search page and the search API a word OCR misread by a letter still finds it.
Requirements for RTF Documents
rtfmust be set totrue(boolean, not a string)urimust be a validhttp://orhttps://URL pointing to the document file- The file must be publicly accessible (or accessible from the Opensolr server)
- Maximum file size: the document size limit of your plan (the same limit as the web crawler's page size)
titleanddescriptionare recommended but the text field will be extracted from the document, so at minimum you needrtf:trueanduri
If Extraction Fails
If the file cannot be fetched (bad URL, unreachable server, over the size limit of your plan), that specific document is marked as failed in the job results with a descriptive error message. Other documents in the same batch continue processing normally. A file that arrives but cannot be read (an unsupported format such as a zip archive, an empty or corrupted file) is not a failure: the document is indexed with the text you sent (empty if you sent none), with the title, description and every other field you sent, and the detected type in content_type.
Check the job status via the API or the Ingestion Queue page to see per-document results.
See it live: a live search over real PDFs whose words sit only inside pictures.
Need to index a large document library? Combine rtf:true with batch uploads for maximum efficiency.
Full API Docs