Repository navigation
Native protocol: a row committed inside BEGIN/COMMIT is invisible to PK point lookups and filtered COUNT(*) — full scans see it #193
Description
Activity
- addedstatus:needs-triageAwaiting maintainer triage (severity + priority)Awaiting maintainer triage (severity + priority)type:bugA defect — broken, incorrect, or lost dataA defect — broken, incorrect, or lost datasev:2-highMajor functionality broken; no acceptable workaroundMajor functionality broken; no acceptable workaroundpriority:P1Fix in the current milestoneFix in the current milestonestatus:confirmedReproduced by a maintainerReproduced by a maintainerarea:txnTransactions, isolation, MVCCTransactions, isolation, MVCCengine:documentDocument engine (schemaless + strict)Document engine (schemaless + strict)and removedstatus:needs-triageAwaiting maintainer triage (severity + priority)Awaiting maintainer triage (severity + priority)
on Jul 20, 2026 Maintainer triage: SEV-2 / P1, confirmed.
A committed native-protocol transaction becomes durably visible to scans but remains absent from PK and filtered read paths across reconnects. That violates the core commit/index-visibility invariant and produces silently inconsistent reads. Avoiding explicit transactions or switching transports is not an acceptable workaround for the native transaction surface.
I traced this through the code and found the root cause. It is not the index code and not the plan cache: the native COMMIT sends the committed data to the wrong vShard.
Root cause
Both pgwire and native use the same shared commit code (
run_commit→dispatch_single_shardinnodedb/src/control/server/shared/session/commit.rs). That code builds the commit tasks (MetaOp::ResolveTxn+MetaOp::TransactionBatch) and puts the correctvshard_idon them. Each transport then sends the tasks through its ownTxnDataPlane::dispatch_no_wal. That last step is where they differ:- pgwire (
pgwire/handler/transaction_cmds/commit.rs) sends the task straight to the data plane usingtask.vshard_id. Correct. - native (
native/dispatch/transaction.rs:42-93) sends it through the gateway instead:gw.execute(&gw_ctx, task.plan). The gateway context only carries tenant/trace/database/txn ids —task.vshard_idis thrown away. The gateway then tries to work out the route from the plan itself. ButMetaOpplans have no collection name in them, so the router falls back to vShard 0 (control/gateway/router.rs:268-275— the doc comment even says: "Falls back to vShard 0 for plans that have no named collection (Meta ops).").
So the whole commit batch (row write + index updates,
data/executor/handlers/transaction/batch.rs) is written durably to vShard 0's core, not to the vShard that owns the collection.Why this matches every symptom
What you saw Why Full scan finds the row Scans read every core's document heap — including vShard 0, where the row wrongly landed PK point lookup finds nothing Point lookups go to the collection's real owning vShard. The row is not there Filtered count(*)is 0The index lookup reads the INDEXES table on the owning vShard. The index entry was written on vShard 0 Stays wrong after reconnect The row is durably committed — just on the wrong core. No cache involved Autocommit works Autocommit sends the real DocumentOpwrite plan, which does have a collection name, so the gateway routes it to the right vShardpgwire works Its commit path never uses the gateway; it keeps task.vshard_idAbout the plan-cache guess in the report: that v0.4.0 fix (
3eaa4987) only touched pgwire, and the native SQL path has no plan cache at all (native/dispatch/sql.rsre-plans every statement). The similar symptom was a coincidence.Also note: the native seam's non-gateway fallback branch (
transaction.rs:72-90) passesvshard_idcorrectly — only the gateway branch is broken.Proposed fix
- Fix: make
NativeTxnDp::dispatch_no_walsend commit tasks through the path that keepstask.vshard_id(the same way pgwire and the native fallback branch already do), instead ofgw.execute, which cannot routeMetaOpplans. - Guardrail: in the gateway router, a Meta write plan with no route target should return an error instead of silently going to vShard 0. That silent fallback is what turned a routing gap into wrong-shard writes.
- Tests: a native-protocol integration test (BEGIN op / INSERT / COMMIT op → PK point lookup + filtered
count(*)on a fresh connection), plus the same test over pgwire. The multi-shard commit branch ofclassify_dispatchneeds the same check too.
One warning for triage: rows already committed through native transactions are sitting on vShard 0. The fix stops new cases but does not move the old rows — they stay scan-visible but lookup-invisible. That likely needs a separate repair issue.
Happy to open a PR with the fix + tests.
- pgwire (
- added a commit that references this issue
on Jul 21, 2026 - added a commit that references this issue
on Jul 21, 2026 - added a commit that references this issue
on Jul 21, 2026 - added a commit that references this issue
on Jul 21, 2026 - removedpriority:P1Fix in the current milestoneFix in the current milestone
on Jul 21, 2026
Version / build tested against
origin/main @ eea86b279Deployment mode
Origin — single node (local)
Engine(s) involved
Document (strict)
Summary
Over the native binary protocol (
:6433), a row inserted inside an explicit transaction (BEGIN op → SQL INSERT → COMMIT op) is durably stored — full scans return it and it survives reconnects — but a PK point lookup (WHERE id = '<pk>') returns 0 rows, on the same connection and on a fresh one, and a filteredcount(*)misses it too. The identical statement sequence over pgwire (:6432) is correct, and an autocommit insert over native is also correct — the trigger is specifically the transactional write path on the native transport. Looks adjacent to the session plan cache fixed in v0.4.0 for pgwire point reads (that fix did not help the native path).Steps to reproduce
Native-protocol session against
:6433(requests are the documented MessagePack frames[op, seq, [0x01, {field => value}]];0x20= SQL,0x40/0x41= BEGIN/COMMIT ops). SQL payloads and returned rows:The same SQL sequence over pgwire on
:6432(psql,BEGIN;/COMMIT;as statements) returnsa1from the point lookup and1from the filtered count.Expected behavior
Committed transactional writes are visible to every read path on the native transport — point lookups and filtered aggregates — identical to full scans, to autocommit writes, and to pgwire.
Actual behavior
PK point lookups and filtered
count(*)permanently return 0 rows/0 for rows written through BEGIN/COMMIT over native, on the writing connection and all later ones, while full scans return the row. No error anywhere.What actually happened? (check all that are true)
Proposed severity
SEV-2 — High: silently-wrong reads of durably stored data; the stored data itself is intact (scans prove it). "Workaround" is switching transports or avoiding transactions, not a rewrite within the native path.
Reproducibility
Always — every attempt
Last known-good version / commit (if a regression)
(blank — first observed on
3eaa49873; unknown whether older builds were affected)Environment & logs
Linux x86_64, release build. No server-side errors or warnings logged during the sequence. Can attach a packet capture of the native session if useful.
Before submitting
mainbuild (not a stale local branch).