Summary
On a standalone 0.4.0 node in auth.mode = "trust", a single CREATE TENANT statement:
- reports success,
- immediately locks out all authentication (
FATAL: trust auth: user '<superuser>' does not exist), and
- leaves the data directory permanently unbootable — the new catalog sanity check detects the resulting dangling references but does not repair them, and fails startup closed.
There is no recovery path short of discarding the data directory.
Reproduced deterministically from a fresh data dir, on both single_node_calvin = true (default) and false.
Environment
nodedb 0.4.0, git commit 38bfc3084 (tag v0.4.0), release build
- Single standalone node, fresh data dir, Linux x86_64, rustc 1.95.0
- Client:
psql 16.14 over pgwire
- Config: minimal —
[server] host/data_dir/memory_limit/ports + [auth] mode = "trust", superuser_name = "nodedb"
Reproduction
-- fresh data dir, server starts clean; ordinary DDL/DML works:
CREATE COLLECTION smoke (id INT PRIMARY KEY, note TEXT);
INSERT INTO smoke VALUES (1, 'ok');
SELECT count(*) FROM smoke; -- 1
BEGIN; INSERT INTO smoke VALUES (2,'a'); INSERT INTO smoke VALUES (3,'b'); COMMIT; -- commits
-- the trigger:
CREATE TENANT alpha; -- returns: CREATE TENANT (apparent success)
-- the very next connection, same credentials:
SELECT 1;
-- psql: FATAL: trust auth: user 'nodedb' does not exist
Every subsequent connection is refused. Restarting the server does not recover it:
ERROR nodedb::control::startup::startup_sequencer: StartupSequencer transitioned to Failed
error=subsystem 'catalog-sanity-check' failed during CatalogSanityCheck:
catalog sanity check failed: catalog_sanity: applied_index_ok=true gap=0
integrity_violations=3 integrity_repaired=0 registry_divergences=0 repaired=0
all_repairs_ok=true elapsed=197.037µs
integrity: dangling reference owner(collection:0:1:g0cov) → user(nodedb) not found
integrity: dangling reference owner(collection:0:1:g0sl) → user(nodedb) not found
integrity: dangling reference owner(collection:0:1:smoke040) → user(nodedb) not found
Error: catalog sanity check failed: ...
The process exits; the data directory cannot be opened again.
Root cause
The three pieces compose into a trap:
1. In trust mode the configured superuser is never materialised as a stored user.
bootstrap_superuser() — nodedb/src/bootstrap/credentials.rs:47-64. resolve_superuser_password() returns Ok(None) under trust mode, so the Ok(Some(password)) arm that calls credentials.bootstrap_superuser(&config.auth.superuser_name, …) never runs. The Ok(None) arm only prints the trust-mode warning banner. No user row is created.
2. Trust auth therefore succeeds only via an "empty store" escape hatch.
nodedb/src/control/server/pgwire/handler/trust_auth.rs:38-56:
if let Some(identity) = stored_user_identity(&self.state, &username, AuthMethod::Trust) {
return Ok(identity);
}
if self.state.credentials.is_empty() {
return Ok(trust_identity(&self.state, &username)); // <- the only reason trust mode works
}
… Err(FATAL 28000 "trust auth: user '{username}' does not exist")
On a fresh trust-mode node the credentials store is empty, so every connection takes the second branch.
3. CREATE TENANT installs a user into that store.
nodedb/src/control/catalog_entry/post_apply/tenant.rs:38-39:
pub fn put_with_admin(tenant: StoredTenant, admin: StoredUser, shared: Arc<SharedState>) {
shared.credentials.install_replicated_user(&admin, None);
put(tenant, shared);
}
The store is no longer empty ⇒ the escape hatch closes ⇒ the configured superuser (which step 1 never created) fails stored_user_identity ⇒ FATAL for every connection, including the superuser's own.
4. The unbootable state follows from the same missing row. Collections created before the tenant were stamped owner = user(nodedb). Since no such user row ever existed, they are dangling references, and catalog-sanity-check fails startup. Note the counters: integrity_violations=3 but integrity_repaired=0 with all_repairs_ok=true — the checker classifies these violations as detected-but-not-repairable, then fails closed, so the node can never start again.
Not a consensus fault
Worth stating explicitly, since the first hypothesis was raft-related. On the default path the log also shows, from startup:
ERROR nodedb::control::distributed_applier::applier: leader-change no-op committed at index
where a proposer was waiting; surfacing RetryableLeaderChange … group_id=3 log_index=1
… (groups 1, 3, 4 at log_index 1–2)
WARN nodedb::control::surrogate::assign::core::flush: surrogate hwm raft propose failed;
followers may lag hwm=1 error=configuration error: surrogate_alloc propose timed out
waiting for log index 4 (then 14, 20, 34 — climbing)
However, re-running the identical repro with [server] single_node_calvin = false produces zero RetryableLeaderChange and zero surrogate_alloc timeouts — and CREATE TENANT still locks the node out identically. So the lockout is independent of the Calvin/raft path.
The single-node raft flapping above may still be a genuine separate issue (surrogate_alloc proposals never committing on a standalone node with single_node_calvin = true) — happy to file that separately if useful.
Impact
- Any standalone trust-mode deployment (the documented local-dev/CI posture) is one
CREATE TENANT away from an unrecoverable node.
- The statement reports success, so there is no signal at the point of damage.
- Data loss is total for that directory: no repair, no fsck path (
nodedb migrate/repair/fsck are listed as reserved/not-yet-implemented).
Suggested directions
- Materialise the superuser in trust mode too — install
auth.superuser_name as a stored user at bootstrap regardless of mode, so the identity exists independently of the store being empty. This alone closes the lockout.
- Make the empty-store escape hatch not silently revocable — e.g. resolve the configured superuser name explicitly rather than depending on
credentials.is_empty(), which any feature that installs a user can flip as a side effect.
- Let the sanity check repair a dangling owner reference (re-point to the configured superuser, or quarantine the entry) instead of failing closed forever — an unbootable directory is a harsher outcome than the violation warrants.
- Possibly: have
CREATE TENANT fail loudly up front when the caller's own identity cannot survive the operation.
Artifacts
Three bricked data directories are preserved locally and can be uploaded or inspected on request:
- default path (
single_node_calvin = true)
- Calvin disabled (
single_node_calvin = false) — the control proving the lockout is not consensus-related
- the original occurrence, with the full server log
Happy to attach logs, run a variant, or test a patch against this setup.
Summary
On a standalone 0.4.0 node in
auth.mode = "trust", a singleCREATE TENANTstatement:FATAL: trust auth: user '<superuser>' does not exist), andThere is no recovery path short of discarding the data directory.
Reproduced deterministically from a fresh data dir, on both
single_node_calvin = true(default) andfalse.Environment
nodedb 0.4.0,git commit 38bfc3084(tagv0.4.0), release buildpsql16.14 over pgwire[server]host/data_dir/memory_limit/ports +[auth] mode = "trust", superuser_name = "nodedb"Reproduction
Every subsequent connection is refused. Restarting the server does not recover it:
The process exits; the data directory cannot be opened again.
Root cause
The three pieces compose into a trap:
1. In trust mode the configured superuser is never materialised as a stored user.
bootstrap_superuser()—nodedb/src/bootstrap/credentials.rs:47-64.resolve_superuser_password()returnsOk(None)under trust mode, so theOk(Some(password))arm that callscredentials.bootstrap_superuser(&config.auth.superuser_name, …)never runs. TheOk(None)arm only prints the trust-mode warning banner. No user row is created.2. Trust auth therefore succeeds only via an "empty store" escape hatch.
nodedb/src/control/server/pgwire/handler/trust_auth.rs:38-56:On a fresh trust-mode node the credentials store is empty, so every connection takes the second branch.
3.
CREATE TENANTinstalls a user into that store.nodedb/src/control/catalog_entry/post_apply/tenant.rs:38-39:The store is no longer empty ⇒ the escape hatch closes ⇒ the configured superuser (which step 1 never created) fails
stored_user_identity⇒ FATAL for every connection, including the superuser's own.4. The unbootable state follows from the same missing row. Collections created before the tenant were stamped
owner = user(nodedb). Since no such user row ever existed, they are dangling references, andcatalog-sanity-checkfails startup. Note the counters:integrity_violations=3butintegrity_repaired=0withall_repairs_ok=true— the checker classifies these violations as detected-but-not-repairable, then fails closed, so the node can never start again.Not a consensus fault
Worth stating explicitly, since the first hypothesis was raft-related. On the default path the log also shows, from startup:
However, re-running the identical repro with
[server] single_node_calvin = falseproduces zeroRetryableLeaderChangeand zerosurrogate_alloctimeouts — andCREATE TENANTstill locks the node out identically. So the lockout is independent of the Calvin/raft path.The single-node raft flapping above may still be a genuine separate issue (
surrogate_allocproposals never committing on a standalone node withsingle_node_calvin = true) — happy to file that separately if useful.Impact
CREATE TENANTaway from an unrecoverable node.nodedb migrate/repair/fsckare listed as reserved/not-yet-implemented).Suggested directions
auth.superuser_nameas a stored user at bootstrap regardless of mode, so the identity exists independently of the store being empty. This alone closes the lockout.credentials.is_empty(), which any feature that installs a user can flip as a side effect.CREATE TENANTfail loudly up front when the caller's own identity cannot survive the operation.Artifacts
Three bricked data directories are preserved locally and can be uploaded or inspected on request:
single_node_calvin = true)single_node_calvin = false) — the control proving the lockout is not consensus-relatedHappy to attach logs, run a variant, or test a patch against this setup.