To change a document, send it again with the same uri. To take it out of your Opensolr Index, delete it by its id. There is nothing else to keep in sync.
Update: send it again
- Same
uri, same document. The id is made from theuri, so the new version replaces the old one. - The whole document is replaced. It is not a partial update: a field you leave out is gone. Always send every field, even to change one.
- Pages read from your site too. A page your crawl indexed is replaced when you push the same address. The next crawl of that page writes it again from your site.
- A PDF sent again replaces all its pages (PDFs page by page).
A document that a job has not finished yet is skipped when you send it again, so it is never queued twice. When every document of a request was already waiting, the answer is All documents already queued with their number in dupes_skipped. Send the new version once that job is finished, or Pause the job, change it and Save Payload, then Resume (The Data Ingestion queue).
Remove: delete by id
Data Ingestion has no delete call: you delete from the index itself. The id of a document is in doc_ids when you push it; it is the MD5 of its uri (without a trailing slash). One request deletes the document and, for a PDF, all its pages:
curl -u USER:PASS "https://YOUR_SERVER/solr/YOUR_INDEX_NAME/update?commitWithin=1000&wt=json" \
-H "Content-Type: application/json" \
-d '{"delete": {"query": "id:\"DOC_ID\" OR uri_id_s:\"DOC_ID\""}}'
Your index's address and its user and password: Connection URL. Without code, in the Control Panel of the index: Tools > Delete By Query (Delete By Query).
On Drupal and WordPress, with real-time sync on, the Opensolr Search module and plugin remove a page from the index when you delete or unpublish it.