Repository navigation
0.4.0: trust-mode sync sessions leave collections owned by non-durable 'sync-client' — restart fails catalog sanity, data dir unbootable #207
Description
Activity
Triage proposal (no label rights):
type:bugsev:1-criticalarea:crdt-sync— data-dir loss on restart from normal sync usage, no workaround beyond discarding the dir. Per the sev-1 convention this one's yours, Farhan — evidence dir + boot log preserved if you want them.- addedtype:bugA defect — broken, incorrect, or lost dataA defect — broken, incorrect, or lost datasev:1-criticalData loss, corruption, security, or crash; no workaroundData loss, corruption, security, or crash; no workaroundarea:crdt-syncCRDT, edge-to-cloud syncCRDT, edge-to-cloud sync
on Jul 22, 2026 Thanks for the detailed repro — the mechanism you describe is real, but it was fixed before the
v0.4.0release, so it doesn't reproduce on the tagged build.The version reference in the report is inconsistent: the issue says
v0.4.0 (tag 38bfc3084), but those point at different commits.v0.4.0tags 26ac75c.38bfc3084(fix(tenant): reject duplicate CREATE TENANT names) is an ancestor of thev0.4.0tag — an earlier pre-release commit. That's the tree you actually built against, and yes, at that commit the trust-mode sync handshake fabricates the ephemeralsync-clientidentity exactly as you quote. So the bug was genuinely present at38bfc3084.
Between
38bfc3084and the release,74febcf8(fix(auth): materialize trust superuser identity, inv0.4.0) replaced that path. On thev0.4.0tag the handshake resolves the empty trust token to the configured durable principal instead of a pseudo-identity:// Trust mode: an empty token resolves to the configured durable // principal. Never fabricate an identity that cannot own catalog data. if msg.jwt_token.is_empty() { let Some(identity) = configured_trust_identity(state) else { /* reject */ }; self.identity = Some(identity);
sync-clientno longer exists anywhere in the tree. Sync-created collections are now owned by the durable trust principal, so the catalog sanity check has no dangling owner ref to trip on at restart — this is effectively suggested fix #2.Could you pull latest from
main/HEAD (or rebuild against the current head), re-run the repro, and confirm boot survives a restart? I'm confident it will, so I'm removing the triage labels and closing this as already fixed — please reopen with a fresh boot log if it still reproduces on a clean HEAD build.The one part not covered by the above is suggestion #3 (making the catalog checker repair/warn-and-continue on dangling owner refs as defense-in-depth). That's a legitimate hardening item, but it's a separate durability concern rather than a sync bug — I'll track it on its own if we decide to pursue it.
- removedtype:bugA defect — broken, incorrect, or lost dataA defect — broken, incorrect, or lost datasev:1-criticalData loss, corruption, security, or crash; no workaroundData loss, corruption, security, or crash; no workaroundarea:crdt-syncCRDT, edge-to-cloud syncCRDT, edge-to-cloud sync
on Jul 22, 2026 Confirmed fixed — re-ran the exact repro against a fresh
origin/main @ 81169d3c7release build: sync-created collection (3 docs via the announce, all materialized) → restart on the same data dir → boot survives,catalog sanity check passed, data intact (count(*) = 3post-restart). No integrity violations, no dangling refs.You were right on every count including the stale tag ref on my side —
git fetchhad kept a pre-releasev0.4.0pointer locally and I tested/filed against38bfc3084without re-checking main. That's fixed in my process (fetch--tags --force+ main-only verification + the issue template going forward). Thanks for the patient breakdown.One disclosure on the verification build: it carries the one-line listener-address env override from #209 so the instance could bind its sync port beside the production Origin on this box — it touches only the bind address, none of the identity/catalog logic verified here.
Summary
A trust-mode NodeDB-Lite sync session that creates a collection via the outbound
CollectionSchemaannounce leaves that collection owned by the ephemeralsync-clientidentity, which is never materialized as a durable catalog principal. Everything works at runtime — but on the next server restart, the catalog sanity check finds a dangling owner reference, refuses to repair it, and fails boot. The data directory is permanently unbootable.Same failure family as #195 (ephemeral identity + fail-closed catalog checker, no repair path): PR #198 materialized the configured trust superuser as a durable principal, but the sync session identity was missed.
Environment
v0.4.0(tag 38bfc30), release build, Linux x86_64, single nodeauth.mode = "trust", standalone instance, fresh data dirmain(c913b6b) built against the v0.4.0 workspace crates, connecting with an empty JWT (trust posture)Reproduction
SyncClient+run_sync_loop), create a collection on lite only (e.g.CREATE COLLECTION entries WITH (bitemporal=true)orcreate_collection),document_putsome docs.CollectionSchemaannounce registers the collection on Origin (PutCollectionIfAbsent). Sync +SELECTon Origin work. So far so good.There is no repair path; recovery = discard the data dir. (Bricked dir + full boot log preserved and available on request.)
Mechanism
nodedb/src/control/server/sync/session/handshake.rs:83— the trust-mode sync handshake builds its session identity withusername: "sync-client"— an in-memory identity, never installed in the credential/catalog store.CollectionSchemaannounce path proposesPutCollectionIfAbsentunder that session identity, so the created collection's catalog entry recordssync-clientas owner.owner → user(sync-client)with no such durable user and aborts boot.Suggested fixes (any one suffices)
Notes