Sends Opensolr a file that has no public address, such as a document stored in your CMS, and gets its text back: PDF pages one by one, scanned pages and pictures read by OCR.
https://api.opensolr.com/solr_manager/api/file_extractemail and api_key of your account, as GET or POST parameters: Authentication.01 · Parameters
index_nameRequiredThe index the file is for.
sha1RequiredThe SHA-1 of the whole file, 40 hex characters.
sizeRequiredThe size of the whole file, in bytes. At most the per-document size of your plan.
nameOptionalThe file name.
partOptionalThe bytes of the file from offset, as a multipart file, at most max_part_bytes.
offsetOptionalWhere part starts in the file, in bytes. Default 0.
pagesOptional1: a PDF comes back as its pages (page, text, is_ocr, lang).
textOptional0: the state alone, without the text.
02 · How it goes
Ask first with sha1 and size only. A file read before comes back at once, state done, with its text.
Otherwise send the file in parts: each answer says how many bytes were received; send the next part from there.
After the last part the job is queued, then processing, then done. Follow it with file_extract_status.
To index the text, send a document with rtf: true and file set to the job_id to ingest.
03 · Example
SHA=$(shasum -a 1 manual.pdf | cut -d' ' -f1); SIZE=$(wc -c < manual.pdf | tr -d ' ') curl -s -X POST "https://api.opensolr.com/solr_manager/api/file_extract" \ -F "email=YOUR_EMAIL" -F "api_key=YOUR_API_KEY" -F "index_name=my_index" \ -F "sha1=$SHA" -F "size=$SIZE" -F "name=manual.pdf" -F "offset=0" -F "part=@manual.pdf"
04 · Answer
{ "status": true, "job_id": "2fd4e1c67a2d28fced849ee1bb76e7391b93eb12-od", "max_part_bytes": 67108864, "state": "queued" }
Reading the text is free. Scanned pages and pictures read by OCR, and pictures described by what they show, count against the AI allowance of the index owner: What it uses from your plan.
05 · Errors
ERROR_INVALID_SHA1HTTP 200sha1 is not 40 hex characters.
ERROR_INVALID_SIZEHTTP 200No size.
ERROR_FILE_TOO_LARGEHTTP 200Larger than your plan allows (max_bytes).
ERROR_PART_TOO_LARGEHTTP 200The part is larger than max_part_bytes.
ERROR_INVALID_OFFSETHTTP 200The part does not fit at this offset.
ERROR_UPLOAD_FAILEDHTTP 200The part did not arrive whole. Send it again.
ERROR_EXTRACTION_QUEUE_FULLHTTP 200The extraction queue is full. Try again later.
ERROR_EXTRACTION_QUEUE_UNAVAILABLEHTTP 200The queue did not answer. Try again.
ERROR_NOT_CORE_OWNERHTTP 200The index is not yours.
WRONG_API_HOSTHTTP 404Called on opensolr.com. The answer carries the right URL.
ERROR_AUTHENTICATION_FAILEDHTTP 403The email and API key do not match.
ERROR_RATE_LIMIT_PER_MINUTEHTTP 429Too many calls. Wait for the Retry-After seconds: Rate limits.
06 · Related
PDFs page by page in your index: PDFs page by page. OCR: Text in scans and pictures.
Every error code and HTTP status of the API: API errors. Calls per minute and per hour: Rate limits.