Skip to content

Origin sync: cross-collection deltas from a single Lite client cannot apply (one document per database vs per collection) #220

Description

@farhan-syah

Version / build tested against

origin/main @ 1d4a0a8fc

Deployment mode

Origin — single node (local), with a NodeDB-Lite embedded client syncing to it.

Engine(s) involved

Document (schemaless), Key-Value — any two collections on one sync stream reproduce it.

Summary

A NodeDB-Lite client keeps one Loro document for the whole database and exports each delta as an incremental slice of that shared oplog. Origin keeps one document per collection. A delta for collection A whose causal predecessors were routed into collection B's document therefore arrives at Origin without its history, and cannot be applied. Since #208 this is refused loudly and retryably instead of being silently dropped, so no acknowledged write is lost — but the write still never lands. Any workload interleaving writes to two collections from a single Lite client is affected; the KV deferred-flush path in #208 was the first reported trigger, not a special case.

Steps to reproduce

-- On a NodeDB-Lite client, from a fresh Origin data dir:
--   CREATE COLLECTION probe WITH (bitemporal=true)
--   start sync (SyncClient + run_sync_loop), wait Connected
--
--   3x document_put into `probe` with distinct ids,
--   interleaved with a write to a SECOND collection after each one
--   (e.g. kv_put("signals", <id>, <bytes>) + kv_flush()).
--
-- Then on Origin over pgwire:

SELECT count(*) FROM probe;   -- 1   <- WRONG, expected 3

Writing to probe only, with no second collection interleaved, returns 3 — a single-collection workload happens to produce a causally contiguous operation chain.

Expected behavior

3 rows. A delta a client emits for a collection should be self-contained, so Origin can apply it into that collection's document regardless of what the client wrote to other collections in between.

Actual behavior

1 row. The first delta applies; every later probe delta depends on operations that were routed into the other collection's document on Origin, so it is refused as retryable (crdt_pending_dependencies) with the high-water-mark held. Loud and counted since #208, but the rows never materialize.

What actually happened?

  • Core functionality is broken with no acceptable workaround

(Acknowledged data is no longer lost — #208 closed that half. A workaround exists only in the trivial sense of never writing to two collections from one client.)

Proposed severity

SEV-2 — High: major functionality broken or silently-wrong results; stored data intact

Reproducibility

Always — every attempt

Last known-good version / commit (if a regression)

Not a regression — this has never worked.

Environment & logs

Linux x86_64. Origin single-node, trust auth.

Origin refuses each affected delta with:

crdt sync apply refused: delta depends on operations absent from this
collection's document; nothing applied, high-water-mark held

Direction (for triage)

Lite moves to one CrdtState per collection, matching Origin, so every delta is self-contained. This removes the asymmetry at the source rather than detecting it after the fact.

Known scope:

  • CrdtEngine.state: CrdtState becomes a per-collection map. Deferred accumulation (upsert_deferred / flush_deltas) and batch_upsert need per-collection tracking rather than one shared deferred_version.
  • SyncDelegate::import_remote takes no collection parameter today, so the inbound path cannot route an update to the right document — the signature needs one.
  • Persistence: the Loro snapshot and vector clock are currently single-document; both become per-collection.

The inverse — Origin adopting one document per tenant — was considered and rejected: per-collection isolation on Origin is load-bearing for per-collection snapshot export/import, purge_collection, per-collection frontier digests, and transaction rollback.

Related findings (each deserves its own issue, not folded in here)

  1. Origin fans out CRDT deltas to Lite as DeltaPush frames, but Lite has no inbound handler for that message type — the dispatcher logs "unexpected frame type from Origin" and discards them, so the server-to-client CRDT delta path is inert.
  2. Those outbound frames are stamped producer_id: 0, epoch: 0, seq: 0 — the ungated sentinel — so the server-to-client direction has no sequence gate at all.

Activity

  1. added
    type:bugA defect — broken, incorrect, or lost data
    sev:2-highMajor functionality broken; no acceptable workaround
    status:needs-triageAwaiting maintainer triage (severity + priority)
    engine:documentDocument engine (schemaless + strict)
    and removed
    status:needs-triageAwaiting maintainer triage (severity + priority)
    on Jul 25, 2026
  2. farhan-syah commented on Jul 26, 2026

    @farhan-syah
    MemberAuthor

    Fixed. The repro now returns 3 rows.

    What changed

    Lite now keeps one Loro document per collection, matching Origin, so every exported delta is self-contained.

    The blocker was the deferred KV fast path: upsert_deferred accumulated N writes and exported one delta, but Origin binds one surrogate per (collection, document_id) and rejects a delta whose write-set names more than its frame target. Partitioning by collection alone did not fix that — 50 KV rows in one collection is still 50 rows in one delta.

    Resolved with a bounded range export rather than by giving up batching or by reworking cross-engine identity:

    • loro exposes ExportMode::UpdatesInRange { spans }, which nodedb-crdt was not using — it only had export_updates_since ("everything after a version"). Added CrdtState::local_op_counter() and CrdtState::export_local_range(from, to).
    • Deferred writes now record the exact [from, to) counter range their operations occupy. Because deferred ops are sequential local writes, each row's operations are a contiguous range, so a flush slices out one self-contained delta per row.
    • This also removes the original performance motive for coalescing: the fast path existed because export_updates_since re-exports everything after a version, whereas a bounded export costs only its span. Batching is preserved.

    Also fixed along the way:

    • Each collection's document derives a distinct Loro peer id. Previously every collection shared the node peer id, so unrelated writes in different collections minted identical (peer, counter) operation ids — anything merging two collections into one document silently dropped a row. Lite now mirrors Origin's derivation exactly.
    • Snapshots are persisted per collection (loro_snapshot:<collection>); a corrupt snapshot now discards only that collection instead of resetting the whole engine.
    • SyncDelegate::import_remote takes the collection — update bytes alone do not identify their target document.
    • Server-originated RowPush frames are now admission-gated per (peer_id, collection) by their sequence. There was no server-to-client equivalent of sync_admit, so a re-delivered frame could resurrect a deleted row and an out-of-order pair could leave the older post-image winning. Unsequenced frames (sequence == 0) still always apply, mirroring how producer_id == 0 is treated on the inbound side.

    Related findings — updated

    1. Fixed. The inert DeltaPush fan-out is gone. Origin was reusing a client-to-server message (whose delta field is Loro update bytes) to carry a MessagePack row post-image, and Lite had no handler, so every frame was discarded as an unexpected type. Server-originated writes now travel as a dedicated RowPush message with an explicit Upsert/Delete op, and Lite applies them.
    2. Still open. Those frames carried producer_id: 0, epoch: 0, seq: 0. RowPush replaced that with a real sequence and the admit gate above, so the practical exposure is closed, but the server-to-client direction still has no epoch/producer fencing equivalent to sync_admit.

    Verification

    • New interop test carries both variants of the original report side by side: documents-only, and documents interleaved with kv_put + kv_flush into a second collection. The interleaved case asserts Origin sees exactly 3 rows.
    • Engine-level regression test pins the mechanism: deltas for one collection import cleanly into a fresh document even when writes to another collection were interleaved between them.
    • nodedb-lite 705 lib tests, 111 interop/e2e/integration; nodedb + nodedb-crdt + nodedb-types 5520 lib tests. All green, clippy -D warnings clean in both repos.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:crdt-syncCRDT, edge-to-cloud syncengine:documentDocument engine (schemaless + strict)engine:kvKey-Value enginesev:2-highMajor functionality broken; no acceptable workaroundstatus:confirmedReproduced by a maintainertype:bugA defect — broken, incorrect, or lost data

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions