Skip to content

0.4.0: CREATE TENANT permanently locks out trust-mode auth and leaves the data directory unbootable #195

Description

@emanzx

Summary

On a standalone 0.4.0 node in auth.mode = "trust", a single CREATE TENANT statement:

  1. reports success,
  2. immediately locks out all authentication (FATAL: trust auth: user '<superuser>' does not exist), and
  3. leaves the data directory permanently unbootable — the new catalog sanity check detects the resulting dangling references but does not repair them, and fails startup closed.

There is no recovery path short of discarding the data directory.

Reproduced deterministically from a fresh data dir, on both single_node_calvin = true (default) and false.

Environment

  • nodedb 0.4.0, git commit 38bfc3084 (tag v0.4.0), release build
  • Single standalone node, fresh data dir, Linux x86_64, rustc 1.95.0
  • Client: psql 16.14 over pgwire
  • Config: minimal — [server] host/data_dir/memory_limit/ports + [auth] mode = "trust", superuser_name = "nodedb"

Reproduction

-- fresh data dir, server starts clean; ordinary DDL/DML works:
CREATE COLLECTION smoke (id INT PRIMARY KEY, note TEXT);
INSERT INTO smoke VALUES (1, 'ok');
SELECT count(*) FROM smoke;          -- 1
BEGIN; INSERT INTO smoke VALUES (2,'a'); INSERT INTO smoke VALUES (3,'b'); COMMIT;  -- commits

-- the trigger:
CREATE TENANT alpha;                 -- returns: CREATE TENANT   (apparent success)

-- the very next connection, same credentials:
SELECT 1;
-- psql: FATAL: trust auth: user 'nodedb' does not exist

Every subsequent connection is refused. Restarting the server does not recover it:

ERROR nodedb::control::startup::startup_sequencer: StartupSequencer transitioned to Failed
  error=subsystem 'catalog-sanity-check' failed during CatalogSanityCheck:
  catalog sanity check failed: catalog_sanity: applied_index_ok=true gap=0
  integrity_violations=3 integrity_repaired=0 registry_divergences=0 repaired=0
  all_repairs_ok=true elapsed=197.037µs
    integrity: dangling reference owner(collection:0:1:g0cov)    → user(nodedb) not found
    integrity: dangling reference owner(collection:0:1:g0sl)     → user(nodedb) not found
    integrity: dangling reference owner(collection:0:1:smoke040) → user(nodedb) not found
Error: catalog sanity check failed: ...

The process exits; the data directory cannot be opened again.

Root cause

The three pieces compose into a trap:

1. In trust mode the configured superuser is never materialised as a stored user.
bootstrap_superuser() — nodedb/src/bootstrap/credentials.rs:47-64. resolve_superuser_password() returns Ok(None) under trust mode, so the Ok(Some(password)) arm that calls credentials.bootstrap_superuser(&config.auth.superuser_name, …) never runs. The Ok(None) arm only prints the trust-mode warning banner. No user row is created.

2. Trust auth therefore succeeds only via an "empty store" escape hatch.
nodedb/src/control/server/pgwire/handler/trust_auth.rs:38-56:

if let Some(identity) = stored_user_identity(&self.state, &username, AuthMethod::Trust) {
    return Ok(identity);
}
if self.state.credentials.is_empty() {
    return Ok(trust_identity(&self.state, &username));   // <- the only reason trust mode works
}
… Err(FATAL 28000 "trust auth: user '{username}' does not exist")

On a fresh trust-mode node the credentials store is empty, so every connection takes the second branch.

3. CREATE TENANT installs a user into that store.
nodedb/src/control/catalog_entry/post_apply/tenant.rs:38-39:

pub fn put_with_admin(tenant: StoredTenant, admin: StoredUser, shared: Arc<SharedState>) {
    shared.credentials.install_replicated_user(&admin, None);
    put(tenant, shared);
}

The store is no longer empty ⇒ the escape hatch closes ⇒ the configured superuser (which step 1 never created) fails stored_user_identity ⇒ FATAL for every connection, including the superuser's own.

4. The unbootable state follows from the same missing row. Collections created before the tenant were stamped owner = user(nodedb). Since no such user row ever existed, they are dangling references, and catalog-sanity-check fails startup. Note the counters: integrity_violations=3 but integrity_repaired=0 with all_repairs_ok=true — the checker classifies these violations as detected-but-not-repairable, then fails closed, so the node can never start again.

Not a consensus fault

Worth stating explicitly, since the first hypothesis was raft-related. On the default path the log also shows, from startup:

ERROR nodedb::control::distributed_applier::applier: leader-change no-op committed at index
  where a proposer was waiting; surfacing RetryableLeaderChange … group_id=3 log_index=1
… (groups 1, 3, 4 at log_index 1–2)
WARN nodedb::control::surrogate::assign::core::flush: surrogate hwm raft propose failed;
  followers may lag hwm=1 error=configuration error: surrogate_alloc propose timed out
  waiting for log index 4      (then 14, 20, 34 — climbing)

However, re-running the identical repro with [server] single_node_calvin = false produces zero RetryableLeaderChange and zero surrogate_alloc timeouts — and CREATE TENANT still locks the node out identically. So the lockout is independent of the Calvin/raft path.

The single-node raft flapping above may still be a genuine separate issue (surrogate_alloc proposals never committing on a standalone node with single_node_calvin = true) — happy to file that separately if useful.

Impact

  • Any standalone trust-mode deployment (the documented local-dev/CI posture) is one CREATE TENANT away from an unrecoverable node.
  • The statement reports success, so there is no signal at the point of damage.
  • Data loss is total for that directory: no repair, no fsck path (nodedb migrate/repair/fsck are listed as reserved/not-yet-implemented).

Suggested directions

  1. Materialise the superuser in trust mode too — install auth.superuser_name as a stored user at bootstrap regardless of mode, so the identity exists independently of the store being empty. This alone closes the lockout.
  2. Make the empty-store escape hatch not silently revocable — e.g. resolve the configured superuser name explicitly rather than depending on credentials.is_empty(), which any feature that installs a user can flip as a side effect.
  3. Let the sanity check repair a dangling owner reference (re-point to the configured superuser, or quarantine the entry) instead of failing closed forever — an unbootable directory is a harsher outcome than the violation warrants.
  4. Possibly: have CREATE TENANT fail loudly up front when the caller's own identity cannot survive the operation.

Artifacts

Three bricked data directories are preserved locally and can be uploaded or inspected on request:

  • default path (single_node_calvin = true)
  • Calvin disabled (single_node_calvin = false) — the control proving the lockout is not consensus-related
  • the original occurrence, with the full server log

Happy to attach logs, run a variant, or test a patch against this setup.

Activity

  1. added
    type:bugA defect — broken, incorrect, or lost data
    sev:1-criticalData loss, corruption, security, or crash; no workaround
    priority:P0Drop everything — fix now
    area:pgwirePostgreSQL wire protocol / client compat
    on Jul 21, 2026
  2. self-assigned this
    on Jul 21, 2026
  3. farhan-syah commented on Jul 21, 2026

    @farhan-syah
    Member

    Fixed by #198 and merged into main.

    The underlying problem was broader than CREATE TENANT: trust-mode authority was represented by an ephemeral identity and, in pgwire, depended on the credential catalog being empty. Persisting any principal changed later authentication behavior, while objects created by the ephemeral identity could retain an owner that did not exist in the catalog.

    The merged fix now:

    • materializes the configured trust superuser as a durable catalog principal before network listeners start;
    • resolves trust authentication through durable identities across pgwire, native, HTTP/WebSocket, and sync;
    • rejects unknown or unavailable trust identities instead of fabricating user_id = 0 authority;
    • preserves the configured principal's stable ID across restarts and trust-to-password transitions;
    • atomically persists a new bootstrap user together with the next-user-ID counter; and
    • covers tenant/user/service-account creation, reconnects, restart ownership integrity, protocol parity, and persistence rollback with regression tests.

    As a result, creating a tenant or another principal no longer locks out the configured trust superuser, and catalog ownership remains valid after restart.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:pgwirePostgreSQL wire protocol / client compatsev:1-criticalData loss, corruption, security, or crash; no workaroundstatus:confirmedReproduced by a maintainertype:bugA defect — broken, incorrect, or lost data

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions