8 checks that catch offline-first sync bugs before your users do
8 checks that catch offline-first sync bugs before your users do Sync bugs are the worst class you can ship. They don't crash and they don't log. They reproduce under one specific interleaving of edits, retries and rec
8 checks that catch offline-first sync bugs before your users do
Sync bugs are the worst class you can ship. They don't crash and they don't log. They reproduce under one specific interleaving of edits, retries and reconnects, usually on a phone in a basement with one bar of signal. By the time someone reports a duplicate invoice or a vanished order, the data is already wrong and nobody can tell when it broke.
I audit offline-first sync layers. These are the eight properties I check, in order. For each: the symptom a user sees, the test that exposes it, and the fix that usually holds.
1. Idempotency (retry-safe writes)
- Symptom: duplicate records after a flaky connection. "I only created one order."
- Why: the ack got lost, the client retried, the server wrote a second row.
- Test: send the same write twice with the same operation id; assert exactly one row.
- Fix: every write carries a client-generated op id. The server dedupes on it. Retrying is always safe.
2. Durable outbox (the queue survives a kill)
- Symptom: a change made offline never arrives, even after reconnecting.
- Why: the queue lives in memory and dies with the app process, or the drain is fire-and-forget.
- Test: create an op offline, force-kill the app, reopen, reconnect; assert it lands exactly once.
- Fix: persist the outbox in the local database. Drain in order. Mark an op done only after a real ack.
3. Tombstones (deletes that stick)
- Symptom: a deleted item comes back on the next sync.
- Why: the delete removed the row locally, the server still had the record, and the pull re-created it.
- Test: delete offline, sync, sync again; assert it stays gone.
- Fix: a delete writes a tombstone (id + deleted_at), it does not remove the row. Tombstones replicate and win over stale copies.
4. A total conflict rule (a deterministic winner)
- Symptom: the same record differs on two devices forever, or one person's edit silently disappears.
- Why: last-write-wins by device clock, and device clocks drift.
- Test: edit the same field on two devices while both are offline, then sync both; assert one deterministic outcome every run.
- Fix: choose a rule (server timestamp, version vector, per-field LWW) and make it total and testable. "Whichever arrives last" is not a rule.
5. Client-generated IDs (UUID/ULID at creation)
- Symptom: duplicate parents and children, orphaned rows, broken foreign keys after an offline create.
- Why: ids are assigned by the server, so offline rows have no stable identity yet.
- Test: create a parent and a child offline, sync; assert one parent and a correctly linked child.
- Fix: mint a UUID/ULID on the client at creation time. The server accepts it as the primary key.
6. Atomic change sets (no partial apply)
- Symptom: half a change lands and the rest doesn't. An invoice without its lines. Broken invariants.
- Why: a multi-op change is applied op-by-op and one op fails in the middle.
- Test: force a failure on the second of three batch ops; assert no partial state is visible.
- Fix: apply a change set inside a transaction, all-or-nothing, then ack. Never ack a half-applied write.
7. Ordering (causal backpressure)
- Symptom: an update arrives before its create, or a delete before the insert.
- Why: retries and parallel connections reorder ops.
- Test: shuffle the delivery order of create/update/delete for one record; assert the final state equals the latest intent.
- Fix: per-record ordering, or explicit causal dependencies. The server buffers out-of-order ops instead of applying them blind.
8. Schema migration during sync
- Symptom: sync breaks after an app update, or an old device corrupts data for everyone.
- Why: the local schema changed but queued ops were written for the old shape.
- Test: queue an op on schema v1, update to v2, reconnect; assert a clean apply or a defined conversion.
- Fix: version every op and the local store. Migrate queued ops on update, before you drain them.
The thread running through all eight
Every one of these is a question about what happens when the ack is lost, the order changes, or two writers disagree. If you can force those three conditions in a test, you can find these bugs before your users do. If you can't, you're debugging by anecdote forever.
The cheapest version of this is a checklist plus one deterministic replay test that you actually run. Pick the two properties your layer is most likely to violate, write the test for each, and see which one goes red. That is often enough to name the bug that has been hiding for weeks.
Want a second pair of eyes
I do a fixed-scope audit: I read your sync layer, run these eight checks, and hand you one page with what stands, what breaks, and the failing test that reproduces the worst bug. One deliverable, one price, no retainer.
- Price: $20, delivered in 48h. Email me: [email protected]
- If you only want the second opinion, this article is already the free version. Run check 1 and check 4 first; they account for most of the duplicates and the "my edit disappeared" reports.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.