Push Files and Documents: PDF, Word, Scans

By their web address, or the file itself

You do not have to turn your files into text first. Send a PDF, a Word or Excel file, a slide deck or a scan, and Opensolr reads it for you: its text, the words inside its pictures, and a short description of pictures that have no words.

Files Opensolr reads: PDF, Word (DOC, DOCX), Excel (XLS, XLSX), PowerPoint (PPT, PPTX), OpenDocument (ODT, ODS, ODP), RTF, CSV, HTML, plain text and pictures (PNG, JPEG, GIF, WebP, TIFF, BMP). The type is read from the file's content, never from its name.

A file with a web address

  1. Put the file's address in uri.
  2. Add "rtf": true ("true" and 1 work too).
  3. Send title and description as usual. Leave out text: it comes from the file.
{
  "uri": "https://example.com/files/annual-report-2025.pdf",
  "title": "Annual report 2025",
  "description": "Results, outlook and the full financial statements.",
  "rtf": true
}

Opensolr downloads the file when the job runs. It must be reachable on the internet at a public address, and no larger than the document size of your plan (Limits and refusals).

  • Pictures and scans: pictures inside the document and scanned pages are read with OCR (in a PDF, in the language of each page). While your AI allowance lasts, a picture with no words gets a short description of what it shows. See Text in scans and pictures and Pictures described by what they show.
  • PDFs: each page becomes a result of its own (PDFs page by page).
  • The file cannot be downloaded (wrong address, refused, too large): that document fails, with the reason in the job's result. The other documents of the job go on.
  • The file arrives but has no readable text (an archive, an empty or damaged file): the document is still indexed, with the text you sent (or none), its title, description and every other field.

A file with no web address

A file on your disk, in an e-mail or in private storage has no public address to download. Send the file itself first, then push the document. Three calls, all on api.opensolr.com:

  1. Send the file to file_extract with index_name, the file's sha1, its size in bytes, its name, offset=0 and the file as part. A large file goes in parts, each with its offset; every answer gives max_part_bytes, the largest part one call takes. Keep the job_id of the answer.
  2. Wait until it is read: file_extract_status with job_ids (up to 20 at once) answers queued, processing, then done or failed.
  3. Push the document with "rtf": true and "file": "<job_id>", plus title, description and a uri. The uri can be any http(s) address that names the document: it stays its identity.
# 1. Send the file and keep its job_id
SHA1=$(shasum msa.pdf | cut -d' ' -f1); SIZE=$(wc -c < msa.pdf | tr -d ' ')
curl -s -F part=@msa.pdf "https://api.opensolr.com/solr_manager/api/file_extract?email=you@example.com&api_key=YOUR_API_KEY&index_name=YOUR_INDEX_NAME&sha1=$SHA1&size=$SIZE&name=msa.pdf&offset=0&text=0"

# 2. Ask until the state is done
curl -s "https://api.opensolr.com/solr_manager/api/file_extract_status?email=you@example.com&api_key=YOUR_API_KEY&index_name=YOUR_INDEX_NAME&job_ids=YOUR_FILE_JOB_ID&text=0"

# 3. Push the document
curl -s -X POST "https://api.opensolr.com/solr_manager/api/ingest" -H "Content-Type: application/json" \
  -d '{"email":"you@example.com","api_key":"YOUR_API_KEY","core_name":"YOUR_INDEX_NAME","documents":[{"uri":"https://example.com/contracts/msa.pdf","title":"Master services agreement","description":"Signed contract","rtf":true,"file":"YOUR_FILE_JOB_ID"}]}'

A file you sent before, with the same content, is not read again. A document pushed while its file is still being read, or with a file that could not be read, fails with a message that says so. The LangChain, LlamaIndex, Haystack and MCP packages do these three steps for a whole folder (Developer packages).

What it uses from your plan

Pictures read with OCR and pictures described count against your monthly AI allowance; a picture read before is free. What each one counts: What it uses from your plan.

Technical reference: Document extraction. See it live on real PDFs whose words sit only in pictures: OCR search demo.

Push your data pages

Back to Enterprise Site Search