Skip to content

perf(gfql): native chain seeds resolve through the resident indexes; named patterns served (#2027) - #2037

Merged
lmeyerov merged 4 commits into
masterfrom
perf/gfql-native-seed-resolution
Sep 5, 2026
Merged

lmeyerov merged 4 commits into
masterfrom
perf/gfql-native-seed-resolution

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #2035 → #2031 → #2030. Closes the "adapter gap" (plan T2.10.3): the SNB benchmark lane runs GFQL as native op lists (n(..., name=..), e_forward, rows, select), not Cypher text, so the Cypher fast paths never applied there, and the native seeded paths scanned the node table for a seed predicate on a property that is not the node binding.

Changes

  • Named single-node ops are served on the chain fast path (they always ran the full two-pass chain); the seed resolves through the node-id index, the node-property index, or the canonical filter, and the alias flag column is placed after the node id like the full path.
  • The native seeded typed hop resolves its seed with the shared resolver (property index for a predicate off the binding column).
  • The "defer to the index" decline for named patterns is removed: by the time the fast path runs, the indexed kernel has already declined the middle (_handle_boundary_calls), so the deferral only cost the scan.
  • Alias tagging matches the full path's column layout and RangeIndex, so bindings-table suffixes see identical frames.

Evidence (local, not published; lane's own conformance suite on the SF0.1 fixture, pandas): seed-lookup 45.8 → 9.2 ms, message-content 28 → 5.5 ms, message-creator 51.5 → 7.7 ms. On dgx the same lane at the parent head measured 33.4 / 10.7 / 28.8 ms while the Cypher sentinel measured 1.9 / 1.0 ms on the same host; the re-measure at this head follows. Polars' native chain has its own index consult and is unchanged here (#2033).

Tests: test_native_seed_resolution_2027.py (lane-shaped graph: synthetic node key, predicate on id, label__X columns; four shapes served with value parity on pandas and cuDF; property-index seam engagement; alias layout), test_chain.py deferral pin rewritten to the new contract; chain/hop/fast-path/index/semantics suites 1753 passed (3 pre-existing local cuBLAS failures). Guards green.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh1i1

Fixes #2027.

@lmeyerov

lmeyerov commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

READY (9b190d2): 75/75 + RTD incl. tck-gfql; pins on both sides (policy off / stale index / IsIn+GT seeds → scan with parity); real-GPU: cuDF suites in the RAPIDS 26.02 image pass (stack tip run 958 passed; the 3 cuDF-26-only chain divergences are pre-existing on master and tracked in #2043/#2044). Stacked on #2035.

@lmeyerov
lmeyerov force-pushed the perf/gfql-native-seed-resolution branch 7 times, most recently from ab3a32e to ff818e1 Compare September 5, 2026 15:42
@lmeyerov
lmeyerov changed the base branch from perf/gfql-seeded-lookup-fastpaths to master September 5, 2026 16:32
lmeyerov and others added 3 commits September 5, 2026 09:32
…named patterns served on the chain fast path

- single-node op: named ops served; seed via node-id index, node-property index or the
  canonical filter; alias flag column placed after the node id like the full path
- seeded typed hop helper resolves the seed with the shared _seed_node_rows (property
  index for a predicate off the node binding)
- the 'defer to the index' decline for named patterns is removed: the indexed kernel has
  already declined the middle by the time the fast path runs, so it served nothing
- alias tagging matches the full path's column layout and RangeIndex
- pins: lane-shaped graph (synthetic node key, predicate on id, label__ columns) served with
  value parity on pandas/cuDF; property-index seam engagement; alias layout

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh1i1
…ale indexes, and non-scalar seeds

Every lane shape keeps full-path parity and reports no index use under
index_policy="off" and after a rebind that stales every resident index;
IsIn and GT seed predicates never take the native seed lanes and keep parity.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01QztW7jYsDd66e8rb8pJNQA
@lmeyerov
lmeyerov force-pushed the perf/gfql-native-seed-resolution branch from ff818e1 to a45374b Compare September 5, 2026 16:33
…eated row

Round-003 amplification of the native seed lanes: with a duplicated node
row the served lookup answers one row per matching table row (as the
polars lanes do) while the pandas/cuDF full path self-joins the
duplicates into eight rows. Both sides are recorded so a change to
either flips the pin.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01QztW7jYsDd66e8rb8pJNQA
@lmeyerov

lmeyerov commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

READY (4b91e9f) — restacked on master 9226829 after #2035 landed (patch-id identical to the previous head; base retargeted to master), plus one round-003 pin.

Round-003 amplification of the native seed lanes (owner ask: reflect on the review flow before this lands). A 51-shape × pandas/polars/cuDF differential probe (default route vs the native lanes forced to decline, plus the pandas full path as cross-engine oracle) was run at the #2038 head (which contains this PR) and at master: 140 parity / 3 both-raise (same exception class both routes) / 5 fixture errors that are identical on master / 3 cross-engine divergences that are identical on master. One stack-introduced divergence, pinned here:

  • a node table with a repeated key row: the served native lookup answers one row per matching table row (the polars answer on both routes), while the pandas/cuDF full path self-joins the duplicates into 8 rows. The full-path blow-up predates this stack (master answers 8 on both routes because nothing is served); the new pin records both sides so a change to either is visible.

Pre-existing findings surfaced by the probe, not in this diff (for triage): gfql_index_all(engine="polars") on an edges-only graph raises AttributeError: 'DataFrame' object has no attribute 'get_column'; gfql_index_all(engine="cudf") raises cupy does not support object / category for string node keys and categorical props; polars resolves alias/column collisions with _right suffixes where pandas overwrites.

Cost receipt (local box, medians of 7, ms; [n({"id": k}, name="p")] and the seeded typed hop; indexed lanes flat in table size where master was linear, unindexed unchanged):

engine shape indexed master 10k / 100k / 1M this stack 10k / 100k / 1M
pandas node-lookup yes 4.73 / 5.26 / 10.90 1.05 / 1.08 / 1.05
pandas node-lookup no 4.66 / 5.15 / 11.00 0.96 / 1.02 / 1.28
pandas seeded-hop yes 17.07 / 15.91 / 45.65 2.40 / 2.53 / 2.49
pandas seeded-hop no 1.70 / 2.98 / 11.41 2.33 / 3.65 / 12.28

The one cost this PR adds is on the unindexed pandas seeded hop: about 0.6 ms at 10k rows for the gate that consults the resident indexes; it does not grow with table size.

Receipts at 4b91e9f

@lmeyerov

lmeyerov commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

Addendum (owner ask: the receipt quoted pandas only). This PR's product diff touches four functions on the pandas/cuDF chain entry (_try_chain_fast_path, _seeded_typed_hop_pandas_cudf, _seed_node_rows, _tag_fast_path_aliases) and no polars code path (the polars lane is #2038), so polars is unchanged by construction; measured anyway, same ladder as above at master 9226829 vs this head 4b91e9f (local box, medians of 7, ms; polars 1.42, cudf 25.10 on an RTX 3080 Ti):

engine shape indexed master 10k / 100k / 1M this head 10k / 100k / 1M
polars node-lookup yes 0.22 / 0.35 / 0.47 0.25 / 0.34 / 0.42
polars node-lookup no 0.24 / 0.28 / 0.43 0.22 / 0.23 / 0.41
polars seeded-hop yes 4.35 / 6.48 / 16.30 4.10 / 6.39 / 15.70
polars seeded-hop no 4.87 / 8.18 / 26.01 4.86 / 8.02 / 27.84
cudf node-lookup yes 9.18 / 11.07 / 13.17 3.03 / 5.06 / 4.98
cudf node-lookup no 9.11 / 11.02 / 12.72 2.00 / 2.17 / 2.47
cudf seeded-hop yes 34.76 / 41.81 / 44.27 8.74 / 14.53 / 14.50
cudf seeded-hop no 7.34 / 8.82 / 9.35 7.73 / 9.24 / 9.61
  • polars: every cell within noise of master (unchanged, as the diff predicts).
  • cuDF: every cell faster than master (lookups 3–4.5×, indexed hop 3–4×), and master's inversion (indexed hop 35–44 ms slower than the unindexed 7–9 ms) is gone.
  • Honest loss that stays: on cuDF the indexed seeded hop (8.7 / 14.5 / 14.5) is still slower than the unindexed one (7.7 / 9.2 / 9.6) and is not flat in table size the way pandas is. The pandas/cuDF lane pays device-side fixed costs on the index path (id-array conversions in _index_node_rows / _ids_to_key_array, and the alias tagging) that a plain cuDF filter does not. Not a regression against master and not this PR's scope to close; recorded in the plan's step-9 attribution as the cuDF lever, with the local-GPU caveat (dgx receipts for cuDF come from the master re-measure).

@lmeyerov
lmeyerov merged commit 4c5e252 into master Sep 5, 2026
78 checks passed
@lmeyerov
lmeyerov deleted the perf/gfql-native-seed-resolution branch September 5, 2026 18:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(gfql): ~25 ms fixed floor per point-lookup query on the chain/Cypher path (SNB seed lookup on 1.5k persons)

1 participant