Sync & Re-Sync

One algorithm for both: compare the ids on the phone with the ids in the index, and act on the difference.

There is one algorithm. A first Sync and a Re-Sync differ only in what the phone's own copy of the index already holds: the comparison is between your folders and that copy, on the phone, and a first sync compares against an empty copy, so everything is added.

EVERY SYNC, FORCED OR SCHEDULED 01 Make sure the index is thereReuse it, create it, or recreate it if it vanished. The schema is checked and uploaded ifmissing. 02 Scan the chosen foldersThrough MediaStore. Every photo gets its id: md5 of its absolute path. 03 Download the index's ids and sizesq=*:*, fl=id,size_bytes, sort=id asc, 1000 per page with cursorMark, compared as each pagearrives. 04 Delete what left the phoneIds in the index but not on the phone, deleted 500 at a time. 05 Send what is new or changedFive photos per photos_ingest call; the server reads them and writes the documents,searchable within ten seconds. Only one sync ever runs at a time. Force Re-Sync while one is running tells you to be patientinstead of starting another.
EVERY SYNC, FORCED ORSCHEDULED 01 Make sure the index isthereReuse it, create it, orrecreate it if itvanished. The schema ischecked and uploaded ifmissing. 02 Scan the chosenfoldersThrough MediaStore. Everyphoto gets its id: md5 ofits absolute path. 03 Download the index'sids and sizesq=*:*, fl=id,size_bytes,sort=id asc, 1000 perpage with cursorMark,compared as each pagearrives. 04 Delete what left thephoneIds in the index but noton the phone, deleted 500at a time. 05 Send what is new orchangedFive photos perphotos_ingest call; theserver reads them andwrites the documents,searchable within tenseconds. Only one sync ever runs at atime. Force Re-Sync while oneis running tells you to bepatient instead of startinganother.

Figure 1 — the steps of every sync, forced, scheduled or started by pulling the photo grid down. The comparison happens on the phone; only the deletes, the photos to be read and the words to be carried up leave it.

00 · The phone's copy of the index

The phone keeps a copy of every document its index holds: the id, the path, the file name, the size, the dates, the camera, the EXIF, the place, the words the photo was read into, the printed text read out of it, the people, your own tags and the md5 of the file. It does not hold the search vector and it does not hold the duplicate keys.

The copy is read out of the index once, at install or reinstall. From then on every write keeps it in step, so it is never read again.

What this buys you

A sync with nothing to do makes no requests at all. For a library of 10,000 photos version 2.4 made about 21 requests and downloaded about 1.8 MB on every single sync, only to find out that nothing had changed.

01 · The steps
  1. Make sure the index is there. Reuse it, create it, or create it again if it vanished; check it carries the app's current configuration. See the phone's index.
  2. Scan the chosen folders through MediaStore and give every photo its id, the md5 of its path inside the phone's storage. MediaStore already leaves out files still being written and files in the trash.
  3. Compare, on the phone. The ids and file sizes found in the folders are held against the phone's copy of the index. No request is made, nothing is paged, nothing is downloaded: both sides of the comparison are already on the phone.
  4. Delete every id that is in the copy but no longer on the phone, 500 per request, and drop it from the copy.
  5. Hand to Opensolr every photo that is on the phone but not in the copy, every photo whose file size differs from the one the copy holds (a photo edited in place keeps its id and is read again from scratch), every photo you asked to re-sync from the grid, and every photo indexed without words while the plan had no AI or while the AI server could not answer. A photo that comes back without words again waits a little longer each time (1 hour, 4 hours, 16 hours, then once a day) before it is asked again, so it can never keep the sync busy. A missing place never sends a photo again: the place is looked up once, when the photo is handed over, and a position with no known place is simply indexed without one. Five photos per call, to photos_ingest.
  6. Carry up the words you changed since the last sync — tags, people, wording — through photos_words, 50 photos per call, with no pictures attached, because only the words changed. A sync may consist of nothing else.

A sync therefore has two kinds of work, and either can be empty: photos that must be read, five per call with a picture attached, and words to be carried up, fifty per call with no picture. After the last call the app makes a hard commit, updates its copy of the index with everything it just wrote, and reads the plan usage again.

The phone's copy and the index are matched to the file by its md5. That is what lets a later pass that cannot read a photo — a plan without vector search, or the monthly AI allowance used up — keep the printed text and the words read out of it earlier instead of clearing them, as long as it is the same file.

A sync that actually wrote something is also what throws away the cached filter lists, so they are asked for again; a sync that wrote nothing leaves them as they are.

ON THE PHONEa1f3… IMG_0001.jpgb72c… IMG_0002.jpgc9d0… IMG_0003.jpge41b… IMG_0005.jpgfresh scan of your folders IN THE INDEXa1f3…b72c…d5e6…id and size, 1000 at atime RE-SYNC DOESa1f3…keepb72c…keepc9d0…adde41b…addd5e6…delete id = md5 of the photo's absolute path, built by one function for the phone and for the index.A photo whose size changed is read again, as if it were new. A first Sync compares against an emptyindex.
ON THE PHONEa1f3… IMG_0001.jpgb72c… IMG_0002.jpgc9d0… IMG_0003.jpge41b… IMG_0005.jpgfresh scan of your folders IN THE INDEXa1f3…b72c…d5e6…id and size, 1000 at a time RE-SYNC DOESa1f3…keepb72c…keepc9d0…adde41b…addd5e6…delete id = md5 of the photo'sabsolute path, built by onefunction for the phone and forthe index.A photo whose size changed isread again, as if it were new.A first Sync compares againstan empty index.

Figure 2 — the comparison, both sides on the phone: keep what is in both, add what is only in the folders, delete what is only in the phone's copy of the index.

02 · One call of five photos, one call of fifty
  1. For each photo that has to be read, the app reads its camera details on the phone and makes an upright 1024 px JPEG copy carrying the original's EXIF, adding the md5 of the original file, and your tags, your wording and the people you named on that photo when this phone holds them. That copy is the only picture that ever leaves the phone, and it is made only on this path: on a plan without vector search there is nothing to read, so no copy is made and no picture is sent.
  2. One photos_ingest call takes up to ten of those. The server does the rest: it reads the EXIF, asks the image model what the photo shows, has the text printed in the photo read when it carries any, turns the words into the search vector, resolves the GPS position into a place (city, region, province, community, country), takes the names of the people from the phone or keeps the ones already in the index, keeps the tags and wording already in the index for photos the phone did not speak for, writes the duplicate keys, builds the complete document and writes it into your index.
  3. One photos_words call takes up to fifty photos whose words alone changed — tags, people, wording. No picture is attached and no photo is read again: only those fields are written, and everything else the document holds stays as it is. On a plan with vector search the server makes the new search vector out of those words, one embedding for the whole batch.
  4. Both writes carry commitWithin=10000, so a photo turns up in search within about ten seconds while the next call is already on its way. Nothing is written back into your photos; what comes back is the document the server has just written, and that is what the phone stores in its copy of the index, so the two stay in step.

A photo the monthly allowance cannot cover is indexed without words and without a vector, and read again at a later sync. If it had already been read before, the printed text and the words read out of it are kept, not blanked, as long as it is the same file — the md5 says so.

03 · What the phone keeps

The phone keeps two things in a private SQLite database: its copy of the index, described above, and your edits that have not gone up yet — the tags, the people and the wording you gave a photo. Both are written the moment you save.

Tagging is local-first. Saving tags, names or wording is finished on the phone, right then; the sync that starts straight after carries the change up, fifty photos per call, with no picture attached. You do not wait for it and nothing is lost if it has to be tried again.

Nothing is ever written to the photo files themselves: your words live in the phone's copy and in your index.

What this means in practice

A reinstall is the one moment the whole index is read back, to rebuild the phone's copy — the only time the app reads the index's contents wholesale. After that nothing is re-indexed: an unchanged photo is already in the index and is not touched, a changed one is read again, and the tags and wording in the index come back with its document. Opensolr keeps its own answers, so a picture it has already read is served from its cache instead of being read again.

04 · When a run stops early
A photo cannot be decoded, or is refused

Skips it, counts it as skipped, carries on. A photo the phone cannot read is not tried again until the file changes or you pick it for Re-sync, and it is listed on the Photos screen under the red icon; a photo refused by the server is tried again at the next sync.

Per-minute or per-hour rate limit

The app paces its calls under your account's limits by itself. If the server still asks it to slow down, the run stops and is started again a little later, rather than holding the phone awake waiting.

Monthly AI requests used up

Carries on. From the first refused request the rest of the run writes photos without words and without vectors (date, camera, place, file name and your tags only), counts them, and says so on the Sync screen and in a notification. Those photos carry no clip_model, so the first sync after the allowance resets reads them into words. A photo that had already been read keeps the printed text and the words it was read into, matched to the file by its md5.

A lot of photos to read at once

More than 500 photos waiting to be read waits for the charger, so a long first sync or a reset does not empty the battery.

Index over disk space or bandwidth

Stops and posts a notification that opens your Opensolr account.

Plan without vector search

Nothing is sent to be read at all: no 1024 px copy is made and no picture leaves the phone. Photos are indexed by date, camera, place, file name and the owner's tags, found by those words, and search runs on words only, as Opensolr's site search does on such a plan. A photo read on an earlier plan keeps the printed text and the words it was read into, matched to the file by its md5. The Sync screen says how many photos were indexed this way, and the account screen says why.

Index password changed

Reads the new one from the account and retries.

API key refused

Stops and asks you to sign in again.

Anything else

Commits what was written and stops; the next sync continues.

A run never has to start over. The next Re-Sync compares ids again, against the phone's copy, and only adds what is still missing. Pulling the photo grid down reloads it and starts a sync as well; a sync that is running is stopped first and started again.

05 · Re-read all photos

Re-read on the Sync screen sends every photo through Opensolr again, as when it was first indexed, so it gets whatever Opensolr does better now: the words it is read into, the place, the people and its search vector.

  • Nothing is emptied. Search keeps working the whole time; each photo is simply written again.
  • Your tags, the people you named and your own wording are kept, taken from the phone, from the index as it is, or from words an older version of the app left in the photo files.
  • Your photos are not touched.
  • The progress shows the real total from the first photo. If the run stops, for a low battery or because you pressed Stop, the next sync carries on where it was: photos already read again are not sent twice.
  • The phone's copy of the index is rewritten photo by photo as the writes go up, so the copy and the index are never out of step.

Use it after an update that improves how photos are read. Reset does the same by emptying the index first, which is only needed when the app says the index configuration changed.

06 · When the app's configuration moves on

Every configuration carries a version number, inside the index itself (/opensolr-photos-config answers it) and inside the app. A new version of the app can ship a newer schema.xml or solrconfig.xml; when the index is on an older one, nothing happens until you agree:

Your index will be reset

This version of Opensolr Photos comes with a new index configuration. Your index will now be reset and every photo synced again. — Later keeps searching on the old configuration and pauses syncing; Reset and re-sync starts it.

  1. Nothing of yours is lost. The tags, the people and the wording — including any that came from a reinstall or from another phone — are already on the phone, in its copy of the index, and they are read out of it before anything is emptied.
  2. The index is emptied, and only then is the new configuration uploaded: no document ever sits in the index under a schema that is not its own.
  3. Every photo is handed to Opensolr again, five per call, exactly as in an ordinary sync, with your tags and wording carried along.
  4. The phone's copy is emptied and rebuilt along with the index, write by write, so the two end the reset in step.

The photo grid shows Rebuilding your index with the count meanwhile, and search keeps working while the photos come back, ten seconds behind each write. An index set up by a newer app than the one installed asks for an app update instead, and is left alone.

Opensolr Photos is open source and MIT licensed. Questions about your Opensolr account, index or plan go to opensolr.com/contact; questions about the app itself belong on GitHub.

Opensolr Photos Documentation