You do not have to turn your files into text first. Send a PDF, a Word or Excel file, a slide deck or a scan, and Opensolr reads it for you: its text, the words inside its pictures, and a short description of pictures that have no words.
Files Opensolr reads: PDF, Word (DOC, DOCX), Excel (XLS, XLSX), PowerPoint (PPT, PPTX), OpenDocument (ODT, ODS, ODP), RTF, CSV, HTML, plain text and pictures (PNG, JPEG, GIF, WebP, TIFF, BMP). The type is read from the file's content, never from its name.
A file with a web address
- Put the file's address in
uri. - Add
"rtf": true("true"and1work too). - Send
titleanddescriptionas usual. Leave outtext: it comes from the file.
{
"uri": "https://example.com/files/annual-report-2025.pdf",
"title": "Annual report 2025",
"description": "Results, outlook and the full financial statements.",
"rtf": true
}
Opensolr downloads the file when the job runs. It must be reachable on the internet at a public address, and no larger than the document size of your plan (Limits and refusals).
- Pictures and scans: pictures inside the document and scanned pages are read with OCR (in a PDF, in the language of each page). While your AI allowance lasts, a picture with no words gets a short description of what it shows. See Text in scans and pictures and Pictures described by what they show.
- PDFs: each page becomes a result of its own (PDFs page by page).
- The file cannot be downloaded (wrong address, refused, too large): that document fails, with the reason in the job's result. The other documents of the job go on.
- The file arrives but has no readable text (an archive, an empty or damaged file): the document is still indexed, with the text you sent (or none), its title, description and every other field.
A file with no web address
A file on your disk, in an e-mail or in private storage has no public address to download. Send the file itself first, then push the document. Three calls, all on api.opensolr.com:
- Send the file to
file_extractwithindex_name, the file'ssha1, itssizein bytes, itsname,offset=0and the file aspart. A large file goes in parts, each with itsoffset; every answer givesmax_part_bytes, the largest part one call takes. Keep thejob_idof the answer. - Wait until it is read:
file_extract_statuswithjob_ids(up to 20 at once) answersqueued,processing, thendoneorfailed. - Push the document with
"rtf": trueand"file": "<job_id>", plustitle,descriptionand auri. Theurican be any http(s) address that names the document: it stays its identity.
# 1. Send the file and keep its job_id
SHA1=$(shasum msa.pdf | cut -d' ' -f1); SIZE=$(wc -c < msa.pdf | tr -d ' ')
curl -s -F part=@msa.pdf "https://api.opensolr.com/solr_manager/api/file_extract?email=you@example.com&api_key=YOUR_API_KEY&index_name=YOUR_INDEX_NAME&sha1=$SHA1&size=$SIZE&name=msa.pdf&offset=0&text=0"
# 2. Ask until the state is done
curl -s "https://api.opensolr.com/solr_manager/api/file_extract_status?email=you@example.com&api_key=YOUR_API_KEY&index_name=YOUR_INDEX_NAME&job_ids=YOUR_FILE_JOB_ID&text=0"
# 3. Push the document
curl -s -X POST "https://api.opensolr.com/solr_manager/api/ingest" -H "Content-Type: application/json" \
-d '{"email":"you@example.com","api_key":"YOUR_API_KEY","core_name":"YOUR_INDEX_NAME","documents":[{"uri":"https://example.com/contracts/msa.pdf","title":"Master services agreement","description":"Signed contract","rtf":true,"file":"YOUR_FILE_JOB_ID"}]}'
A file you sent before, with the same content, is not read again. A document pushed while its file is still being read, or with a file that could not be read, fails with a message that says so. The LangChain, LlamaIndex, Haystack and MCP packages do these three steps for a whole folder (Developer packages).
What it uses from your plan
Pictures read with OCR and pictures described count against your monthly AI allowance; a picture read before is free. What each one counts: What it uses from your plan.
Technical reference: Document extraction. See it live on real PDFs whose words sit only in pictures: OCR search demo.