API - Send a file for text extraction (file_extract)

Data Ingestion

Sends Opensolr a file that has no public address, such as a document stored in your CMS, and gets its text back: PDF pages one by one, scanned pages and pictures read by OCR.

Endpointhttps://api.opensolr.com/solr_manager/api/file_extract
MethodPOST, multipart/form-data when a part is sent
Authemail and api_key of your account, as GET or POST parameters: Authentication.

01 · Parameters

index_nameRequired

The index the file is for.

sha1Required

The SHA-1 of the whole file, 40 hex characters.

sizeRequired

The size of the whole file, in bytes. At most the per-document size of your plan.

nameOptional

The file name.

partOptional

The bytes of the file from offset, as a multipart file, at most max_part_bytes.

offsetOptional

Where part starts in the file, in bytes. Default 0.

pagesOptional

1: a PDF comes back as its pages (page, text, is_ocr, lang).

textOptional

0: the state alone, without the text.

02 · How it goes

Ask first with sha1 and size only. A file read before comes back at once, state done, with its text.

Otherwise send the file in parts: each answer says how many bytes were received; send the next part from there.

After the last part the job is queued, then processing, then done. Follow it with file_extract_status.

To index the text, send a document with rtf: true and file set to the job_id to ingest.

03 · Example

SHA=$(shasum -a 1 manual.pdf | cut -d' ' -f1); SIZE=$(wc -c < manual.pdf | tr -d ' ')
curl -s -X POST "https://api.opensolr.com/solr_manager/api/file_extract" \
  -F "email=YOUR_EMAIL" -F "api_key=YOUR_API_KEY" -F "index_name=my_index" \
  -F "sha1=$SHA" -F "size=$SIZE" -F "name=manual.pdf" -F "offset=0" -F "part=@manual.pdf"

04 · Answer

{
    "status": true,
    "job_id": "2fd4e1c67a2d28fced849ee1bb76e7391b93eb12-od",
    "max_part_bytes": 67108864,
    "state": "queued"
}

Reading the text is free. Scanned pages and pictures read by OCR, and pictures described by what they show, count against the AI allowance of the index owner: What it uses from your plan.

05 · Errors

ERROR_INVALID_SHA1HTTP 200

sha1 is not 40 hex characters.

ERROR_INVALID_SIZEHTTP 200

No size.

ERROR_FILE_TOO_LARGEHTTP 200

Larger than your plan allows (max_bytes).

ERROR_PART_TOO_LARGEHTTP 200

The part is larger than max_part_bytes.

ERROR_INVALID_OFFSETHTTP 200

The part does not fit at this offset.

ERROR_UPLOAD_FAILEDHTTP 200

The part did not arrive whole. Send it again.

ERROR_EXTRACTION_QUEUE_FULLHTTP 200

The extraction queue is full. Try again later.

ERROR_EXTRACTION_QUEUE_UNAVAILABLEHTTP 200

The queue did not answer. Try again.

ERROR_NOT_CORE_OWNERHTTP 200

The index is not yours.

WRONG_API_HOSTHTTP 404

Called on opensolr.com. The answer carries the right URL.

ERROR_AUTHENTICATION_FAILEDHTTP 403

The email and API key do not match.

ERROR_RATE_LIMIT_PER_MINUTEHTTP 429

Too many calls. Wait for the Retry-After seconds: Rate limits.

06 · Related

PDFs page by page in your index: PDFs page by page. OCR: Text in scans and pictures.

Every error code and HTTP status of the API: API errors. Calls per minute and per hour: Rate limits.