Skip to content

0.4.0: trust-mode sync sessions leave collections owned by non-durable 'sync-client' — restart fails catalog sanity, data dir unbootable #207

Description

@emanzx

Summary

A trust-mode NodeDB-Lite sync session that creates a collection via the outbound CollectionSchema announce leaves that collection owned by the ephemeral sync-client identity, which is never materialized as a durable catalog principal. Everything works at runtime — but on the next server restart, the catalog sanity check finds a dangling owner reference, refuses to repair it, and fails boot. The data directory is permanently unbootable.

Same failure family as #195 (ephemeral identity + fail-closed catalog checker, no repair path): PR #198 materialized the configured trust superuser as a durable principal, but the sync session identity was missed.

Environment

  • nodedb v0.4.0 (tag 38bfc30), release build, Linux x86_64, single node
  • auth.mode = "trust", standalone instance, fresh data dir
  • Client: nodedb-lite @ main (c913b6b) built against the v0.4.0 workspace crates, connecting with an empty JWT (trust posture)

Reproduction

  1. Boot a fresh single-node 0.4.0 Origin in trust mode (fresh data dir). Boots clean.
  2. Connect a lite replica (SyncClient + run_sync_loop), create a collection on lite only (e.g. CREATE COLLECTION entries WITH (bitemporal=true) or create_collection), document_put some docs.
  3. The outbound CollectionSchema announce registers the collection on Origin (PutCollectionIfAbsent). Sync + SELECT on Origin work. So far so good.
  4. Restart the server (clean shutdown or kill — same result).
  5. Boot fails:
Error: catalog sanity check failed: catalog_sanity: applied_index_ok=true gap=0 integrity_violations=1 integrity_repaired=0 registry_divergences=0 ...
  integrity: dangling reference owner(collection:0:1:entries) → user(sync-client) not found

There is no repair path; recovery = discard the data dir. (Bricked dir + full boot log preserved and available on request.)

Mechanism

  • nodedb/src/control/server/sync/session/handshake.rs:83 — the trust-mode sync handshake builds its session identity with username: "sync-client" — an in-memory identity, never installed in the credential/catalog store.
  • The CollectionSchema announce path proposes PutCollectionIfAbsent under that session identity, so the created collection's catalog entry records sync-client as owner.
  • On restart, the 0.4.0 catalog sanity checker (fail-closed, detect-but-never-repair — same checker as 0.4.0: CREATE TENANT permanently locks out trust-mode auth and leaves the data directory unbootable #195) sees owner → user(sync-client) with no such durable user and aborts boot.

Suggested fixes (any one suffices)

  1. Materialize the sync session identity as a durable principal before it can own catalog objects (the fix(auth): materialize trust superuser identity #198 approach, extended to sync).
  2. Attribute sync-created collections to the configured trust superuser (already durable after fix(auth): materialize trust superuser identity #198) instead of the session pseudo-identity.
  3. Teach the sanity checker to repair (or warn-and-continue on) dangling owner refs instead of failing boot — 0.4.0: CREATE TENANT permanently locks out trust-mode auth and leaves the data directory unbootable #195 suggested this as defense-in-depth; it would also cover future identity gaps of this class.

Notes

Activity

  1. emanzx commented on Jul 22, 2026

    @emanzx
    ContributorAuthor

    Triage proposal (no label rights): type:bug sev:1-critical area:crdt-sync — data-dir loss on restart from normal sync usage, no workaround beyond discarding the dir. Per the sev-1 convention this one's yours, Farhan — evidence dir + boot log preserved if you want them.

  2. added
    type:bugA defect — broken, incorrect, or lost data
    sev:1-criticalData loss, corruption, security, or crash; no workaround
    on Jul 22, 2026
  3. self-assigned this
    on Jul 22, 2026
  4. farhan-syah commented on Jul 22, 2026

    @farhan-syah
    Member

    Thanks for the detailed repro — the mechanism you describe is real, but it was fixed before the v0.4.0 release, so it doesn't reproduce on the tagged build.

    The version reference in the report is inconsistent: the issue says v0.4.0 (tag 38bfc3084), but those point at different commits.

    • v0.4.0 tags 26ac75c.
    • 38bfc3084 (fix(tenant): reject duplicate CREATE TENANT names) is an ancestor of the v0.4.0 tag — an earlier pre-release commit. That's the tree you actually built against, and yes, at that commit the trust-mode sync handshake fabricates the ephemeral sync-client identity exactly as you quote. So the bug was genuinely present at 38bfc3084.

    Between 38bfc3084 and the release, 74febcf8 (fix(auth): materialize trust superuser identity, in v0.4.0) replaced that path. On the v0.4.0 tag the handshake resolves the empty trust token to the configured durable principal instead of a pseudo-identity:

    // Trust mode: an empty token resolves to the configured durable
    // principal. Never fabricate an identity that cannot own catalog data.
    if msg.jwt_token.is_empty() {
        let Some(identity) = configured_trust_identity(state) else { /* reject */ };
        self.identity = Some(identity);

    sync-client no longer exists anywhere in the tree. Sync-created collections are now owned by the durable trust principal, so the catalog sanity check has no dangling owner ref to trip on at restart — this is effectively suggested fix #2.

    Could you pull latest from main/HEAD (or rebuild against the current head), re-run the repro, and confirm boot survives a restart? I'm confident it will, so I'm removing the triage labels and closing this as already fixed — please reopen with a fresh boot log if it still reproduces on a clean HEAD build.

    The one part not covered by the above is suggestion #3 (making the catalog checker repair/warn-and-continue on dangling owner refs as defense-in-depth). That's a legitimate hardening item, but it's a separate durability concern rather than a sync bug — I'll track it on its own if we decide to pursue it.

  5. removed
    type:bugA defect — broken, incorrect, or lost data
    sev:1-criticalData loss, corruption, security, or crash; no workaround
    on Jul 22, 2026
  6. emanzx commented on Jul 22, 2026

    @emanzx
    ContributorAuthor

    Confirmed fixed — re-ran the exact repro against a fresh origin/main @ 81169d3c7 release build: sync-created collection (3 docs via the announce, all materialized) → restart on the same data dir → boot survives, catalog sanity check passed, data intact (count(*) = 3 post-restart). No integrity violations, no dangling refs.

    You were right on every count including the stale tag ref on my side — git fetch had kept a pre-release v0.4.0 pointer locally and I tested/filed against 38bfc3084 without re-checking main. That's fixed in my process (fetch --tags --force + main-only verification + the issue template going forward). Thanks for the patient breakdown.

    One disclosure on the verification build: it carries the one-line listener-address env override from #209 so the instance could bind its sync port beside the production Origin on this box — it touches only the bind address, none of the identity/catalog logic verified here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions