Skip to content

Native protocol: a row committed inside BEGIN/COMMIT is invisible to PK point lookups and filtered COUNT(*) — full scans see it #193

Description

@mkhairi

Version / build tested against

origin/main @ eea86b279

Deployment mode

Origin — single node (local)

Engine(s) involved

Document (strict)

Summary

Over the native binary protocol (:6433), a row inserted inside an explicit transaction (BEGIN op → SQL INSERT → COMMIT op) is durably stored — full scans return it and it survives reconnects — but a PK point lookup (WHERE id = '<pk>') returns 0 rows, on the same connection and on a fresh one, and a filtered count(*) misses it too. The identical statement sequence over pgwire (:6432) is correct, and an autocommit insert over native is also correct — the trigger is specifically the transactional write path on the native transport. Looks adjacent to the session plan cache fixed in v0.4.0 for pgwire point reads (that fix did not help the native path).

Steps to reproduce

Native-protocol session against :6433 (requests are the documented MessagePack frames [op, seq, [0x01, {field => value}]]; 0x20 = SQL, 0x40/0x41 = BEGIN/COMMIT ops). SQL payloads and returned rows:

-- fresh data directory; authenticated native session
CREATE COLLECTION n_probe (id TEXT PRIMARY KEY, name TEXT) WITH (engine='document_strict');
-- status OK

-- op BEGIN (0x40)
INSERT INTO n_probe (id, name) VALUES ('a1', 'alpha');   -- rows_affected 1
-- op COMMIT (0x41)

SELECT id FROM n_probe WHERE id = 'a1';                  -- (0 rows)   <- WRONG
SELECT id FROM n_probe;                                  -- a1 (1 row) -- durably stored
SELECT count(*) FROM n_probe WHERE name = 'alpha';       -- 0          <- WRONG

-- fresh native connection:
SELECT id FROM n_probe WHERE id = 'a1';                  -- (0 rows)   <- WRONG, persists across sessions

-- control, same session, autocommit (no BEGIN/COMMIT ops):
INSERT INTO n_probe (id, name) VALUES ('b2', 'beta');
SELECT id FROM n_probe WHERE id = 'b2';                  -- b2 (1 row) -- correct

The same SQL sequence over pgwire on :6432 (psql, BEGIN;/COMMIT; as statements) returns a1 from the point lookup and 1 from the filtered count.

Expected behavior

Committed transactional writes are visible to every read path on the native transport — point lookups and filtered aggregates — identical to full scans, to autocommit writes, and to pgwire.

Actual behavior

PK point lookups and filtered count(*) permanently return 0 rows/0 for rows written through BEGIN/COMMIT over native, on the writing connection and all later ones, while full scans return the row. No error anywhere.

What actually happened? (check all that are true)

  • Core functionality is broken with no acceptable workaround (on the native transport; pgwire unaffected)

Proposed severity

SEV-2 — High: silently-wrong reads of durably stored data; the stored data itself is intact (scans prove it). "Workaround" is switching transports or avoiding transactions, not a rewrite within the native path.

Reproducibility

Always — every attempt

Last known-good version / commit (if a regression)

(blank — first observed on 3eaa49873; unknown whether older builds were affected)

Environment & logs

Linux x86_64, release build. No server-side errors or warnings logged during the sequence. Can attach a packet capture of the native session if useful.

Before submitting

  • I searched existing issues and this is not a duplicate.
  • I reproduced this on a released tag or a current main build (not a stale local branch).
  • This is not a security vulnerability (those go to a private advisory).

Activity

  1. added
    status:needs-triageAwaiting maintainer triage (severity + priority)
    type:bugA defect — broken, incorrect, or lost data
    sev:2-highMajor functionality broken; no acceptable workaround
    priority:P1Fix in the current milestone
    area:txnTransactions, isolation, MVCC
    engine:documentDocument engine (schemaless + strict)
    and removed
    status:needs-triageAwaiting maintainer triage (severity + priority)
    on Jul 20, 2026
  2. farhan-syah commented on Jul 20, 2026

    @farhan-syah
    Member

    Maintainer triage: SEV-2 / P1, confirmed.

    A committed native-protocol transaction becomes durably visible to scans but remains absent from PK and filtered read paths across reconnects. That violates the core commit/index-visibility invariant and produces silently inconsistent reads. Avoiding explicit transactions or switching transports is not an acceptable workaround for the native transaction surface.

  3. laksamanakeris commented on Jul 21, 2026

    @laksamanakeris
    Contributor

    I traced this through the code and found the root cause. It is not the index code and not the plan cache: the native COMMIT sends the committed data to the wrong vShard.

    Root cause

    Both pgwire and native use the same shared commit code (run_commit → dispatch_single_shard in nodedb/src/control/server/shared/session/commit.rs). That code builds the commit tasks (MetaOp::ResolveTxn + MetaOp::TransactionBatch) and puts the correct vshard_id on them. Each transport then sends the tasks through its own TxnDataPlane::dispatch_no_wal. That last step is where they differ:

    • pgwire (pgwire/handler/transaction_cmds/commit.rs) sends the task straight to the data plane using task.vshard_id. Correct.
    • native (native/dispatch/transaction.rs:42-93) sends it through the gateway instead: gw.execute(&gw_ctx, task.plan). The gateway context only carries tenant/trace/database/txn ids — task.vshard_id is thrown away. The gateway then tries to work out the route from the plan itself. But MetaOp plans have no collection name in them, so the router falls back to vShard 0 (control/gateway/router.rs:268-275 — the doc comment even says: "Falls back to vShard 0 for plans that have no named collection (Meta ops).").

    So the whole commit batch (row write + index updates, data/executor/handlers/transaction/batch.rs) is written durably to vShard 0's core, not to the vShard that owns the collection.

    Why this matches every symptom

    What you saw Why
    Full scan finds the row Scans read every core's document heap — including vShard 0, where the row wrongly landed
    PK point lookup finds nothing Point lookups go to the collection's real owning vShard. The row is not there
    Filtered count(*) is 0 The index lookup reads the INDEXES table on the owning vShard. The index entry was written on vShard 0
    Stays wrong after reconnect The row is durably committed — just on the wrong core. No cache involved
    Autocommit works Autocommit sends the real DocumentOp write plan, which does have a collection name, so the gateway routes it to the right vShard
    pgwire works Its commit path never uses the gateway; it keeps task.vshard_id

    About the plan-cache guess in the report: that v0.4.0 fix (3eaa4987) only touched pgwire, and the native SQL path has no plan cache at all (native/dispatch/sql.rs re-plans every statement). The similar symptom was a coincidence.

    Also note: the native seam's non-gateway fallback branch (transaction.rs:72-90) passes vshard_id correctly — only the gateway branch is broken.

    Proposed fix

    1. Fix: make NativeTxnDp::dispatch_no_wal send commit tasks through the path that keeps task.vshard_id (the same way pgwire and the native fallback branch already do), instead of gw.execute, which cannot route MetaOp plans.
    2. Guardrail: in the gateway router, a Meta write plan with no route target should return an error instead of silently going to vShard 0. That silent fallback is what turned a routing gap into wrong-shard writes.
    3. Tests: a native-protocol integration test (BEGIN op / INSERT / COMMIT op → PK point lookup + filtered count(*) on a fresh connection), plus the same test over pgwire. The multi-shard commit branch of classify_dispatch needs the same check too.

    One warning for triage: rows already committed through native transactions are sitting on vShard 0. The fix stops new cases but does not move the old rows — they stay scan-visible but lookup-invisible. That likely needs a separate repair issue.

    Happy to open a PR with the fix + tests.

  4. self-assigned this
    on Jul 21, 2026
  5. added a commit that references this issue on Jul 21, 2026
  6. added a commit that references this issue on Jul 21, 2026
  7. added a commit that references this issue on Jul 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:txnTransactions, isolation, MVCCengine:documentDocument engine (schemaless + strict)sev:2-highMajor functionality broken; no acceptable workaroundstatus:confirmedReproduced by a maintainertype:bugA defect — broken, incorrect, or lost data

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions