Repository navigation
perf(gfql): native op-lists seeded on a membership set take the resident index (#2116 3a) - #2127
Conversation
…ent index (3a)
[n({"id": is_in(seeds)}), e_forward(), n()] ran the isin scan with gfql_explain recording
nothing while the scalar seed and the Cypher twin took the index. The native lanes resolve
a membership set on the node-id key through the kernel's own _membership_seed_ids, look the
ids up in the node-id index, and re-apply the canonical filter on the hits.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…engaged) The routes-off replay does not collect chain_specializations/ today; marking keeps the contract if it ever does: result pins stay unmarked, used_index/seam assertions are marked. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…yping imports Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
… pins + polars-lane registration) into perf/gfql-native-membership-seed-3a
…es serve it On the real GPU the cuDF engagement pins saw no usable index: the graph was built from pandas frames and run with engine='cudf', so the resident indexes matched the pandas frames. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
|
Real-GPU receipt (dgx-spark, |
… too On the SNB sentinel fixture the node-id lookup is unavailable and the scalar seed is served through the property index on id; the membership path skipped that index and fell to the scan (12.2 ms vs 1.0 ms GFQL-only, pyg-bench arms a3/b3). The property lookup is array-based, so a tuple of ids slots in; non-integral or empty member sets still decline. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…3.9 mypy Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…e native lane The SNB sentinel binds the node column under another name and seeds on the id property; the scalar seed is served through the property index on id, and a membership set on that column must be too. The lane admits membership sets on any seed column: the node-id index serves the binding key, a resident property index serves the rest, the scan keeps the rows. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…redicate still is not The #2027 pin asserted that no non-scalar seed takes the native lane; parity keeps holding for both, and is_in now takes the lane by design (#2127). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…the lane too Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
…n is a route_engaged test Parity stays a result pin for both predicates; served/not-served is an engagement claim and skips under the routes-off replay like the rest of the file's engagement pins. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
Read-only review (parallel session) — native membership seeds + pyg-bench #285Nothing here touches the branch; findings only, fixes deferred. Target: Verdict in one lineNo blockers. Parity holds on 83 admitted/declined shapes × 4 policies (0 divergences); the 50-seed shape is 4× faster (8.4 → 2.1 ms at 100k/500k pandas); two IMPORTANT items (a CHANGELOG/explain claim the single-node lane does not honour; the real-GPU receipt predates the five commits that extend the lane to property columns) and a handful of suggestions. Findings — PR #2127IMPORTANTF3 — CHANGELOG/PR say both native lanes record a decline; only the hop lane does. g = graphistry.edges(edges, "src", "dst").nodes(nodes, "id")
g.gfql_explain([n({"id": is_in(seeds)})], engine="pandas", index_policy="use")
# → used_index=False, decision_code=None, steps=[] (matrix cells int_noidx/single_isin, uint_idx/single_isin)Fix either way: record F9 — GPU evidence is stale for the cuDF-reachable half of the change. The only real-GPU receipt (PR comment: dgx-spark, RAPIDS 26.02, at SUGGESTIONF4 — No cost gate on the native membership lane (measured).
F1 — F2 — F5 — Decline reason is a constant. F6 — The new test file is never replayed by the routes-off CI lanes (pre-existing gap, not a regression): Test gaps (pin the cheap ones): no end-to-end tests for empty Style: Verified OK (no action)
Perf table (pandas, 100k nodes / 500k edges,
|
| shape (policy=use) | master 88b99cc | PR 7d12173 | Δ |
|---|---|---|---|
[n({"id": is_in(50 ids)}), e_forward(), n()] |
8.37 ms | 2.09 ms | 4.0× |
| same with a plain list of 50 ids | 8.77 ms | 2.06 ms | 4.3× |
[n({"id": is_in(50 ids)})] |
1.71 ms | 1.26 ms | 1.4× |
[n({"id": 5}), e_forward(), n()] (single seed) |
1.24 ms | 1.14 ms | no regression |
[n(), e_forward(), n()] (unseeded) |
28.8 ms | 27.4 ms | no regression |
is_in(50% of nodes) hop |
63.8 ms | 79.1 ms | −24% (F4) |
Policy off, 50-seed hop: master 8.35 → PR 6.60 ms (the lane's isin scan branch beats the general path). |
Findings — pyg-bench #285
- Checks:
aggregate-contractpass,smokepass; MERGEABLE/CLEAN. - The new
native-hop-memberspoint measures exactly the perf(gfql): native op-lists seeded on a membership set take the resident index (#2116 3a) #2127 shape (3-op native list,is_inof 10 Messageids +label__Message, HAS_CREATOR, Person) and, because the SNB fixture binds the node column under another name, it is served through the property index onidthat the bench builds — i.e. the0811b3dd5extension, not the node-id path. Good sentinel; the earlier arms' 12.9 ms plateau is explained (members drawn from the binding column). - Thresholds are derived, not typed: every pandas
max_median_msinthresholds-master.jsonequals the receipt median × 1.5 (checked all 18 pandas rows + both new polars rows); older polars bounds are retained round numbers (pre-existing, declared in_comment). Baseline receiptb18d8089fb1dis a master commit (0.59.1 merge), dgx-spark, 8 warmups / 63 repeats. - Candidate file pins
pandas/native-hop-membersat 22.006 ms (= master scan 14.671 × 1.5),served: true,property_index_served: false(gate enforces onlytrue,gate_snb_point_latency.py:91-93). Consequence: [FEA] typecheck invalid api=... value #285 passes under both master and perf(gfql): native op-lists seeded on a membership set take the resident index (#2116 3a) #2127 trees and does not functionally require perf(gfql): native op-lists seeded on a membership set take the resident index (#2116 3a) #2127 to land first; merging after perf(gfql): native op-lists seeded on a membership set take the resident index (#2116 3a) #2127 is still correct so the next baseline swap can tighten the point (≈3 ms,property_index_served: true). Agree with the PR's stated order. - SUGGESTION
native_points._index_served:except Exception: return None— narrow, or record the exception type so a broken explain is distinguishable from "not served". - SUGGESTION
scripts/snb_native_seed_explain_probe.py::_shape_keyclassifies via"IsIn" in repr(...)— substring control flow in a diagnostic script. - SUGGESTION
member_idsloops the whole Messageidcolumn in Python (setup, untimed);ids[ids > message_id].nsmallest(count - 1)is the vectorized form. - Note
polars/native-hop-memberspinsserved: false— the gate will demand an update when polars gains the lane (by design). - Runner change (
dgx_spark_runner.py):results/now ships only*baseline*dirs;PGBENCH_UPLOAD_KIBPSbudget — tests updated (tests/test_dgx_spark_runner.py), fine.
Merge recommendations
PR #2127 — approve after two small follow-ups, no blockers. The change is a real extension of #2117 (any column + property index for native op-lists), correctness is solid (0/83 parity divergences incl. the nasty corners), the claimed win reproduces (4× on the target shape) with no regression on the scalar-seed and unseeded shapes, CI is fully green, and route pins are correctly marked. Before merging: (F3) make the CHANGELOG/explain claim true for the single-node lane or narrow the wording, and (F9) post a real-GPU receipt at head since the last one predates the property-index commits that cuDF can reach. F4 (cost gate, ≤10% regression only for seed sets ≥20% of nodes) and the dead-code/unused-param cleanups (F1, F2, F5) can land here or as a follow-up.
pyg-bench #285 — approve; merge after #2127 as stated. The sentinel measures the right shape, bounds are mechanically derived from a clean master receipt on the same host, the gate semantics make the candidate arm pass on either tree, and the harness fixes are covered by tests. The three suggestions are hygiene in untimed/diagnostic code. The one thing to remember is operational: the post-#2127 baseline swap is what makes native-hop-members a tight 3a sentinel (property_index_served: true, ≈3 ms); until then it only guards against regressing past the scan.
🤖 Generated with Claude Code
…hop lane Review found the CHANGELOG claim that both native lanes report a declined index held only for the hop lane: with no usable index resident the single-node lane returned scan rows and recorded no step at all, so gfql_explain said nothing rather than index_path_unavailable. Pinned on both engines, served and declined. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
|
Both items addressed at fdcfd16. F3 — CHANGELOG says both lanes record a decline; only the hop lane did. Confirmed at F9 — GPU receipt predates the property-column commits. Confirmed: the receipt is at On the suggestions, for the record rather than as changes here: F4's cost gate is a real crossover at roughly 20% of the node table and I would rather gate it in its own PR with the measurement than add an untested threshold now; F1's unused Separately, my own review pass flagged something on the companion pyg-bench PR that I am treating as a blocker there, not here: the swapped baseline's polars medians are 2.4 to 3.4 times the other master runs and that run's own gate-report calls them drift failures. I am re-running the baseline before asking for #285 to be merged. |
The new pin asserts the lane's own explain step, so it must skip when that lane is declined: without the marker it failed the all-routes-off replay. Also merges master, which the branch was 20 commits behind. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp
|
F9 is answered with a receipt at the real head. dgx-spark, GB10, A first pass at the pre-merge tree (fdcfd16) failed two cuDF cells, The run also surfaced an unrelated product limitation I am pinning on #2122 rather than here: |
Summary
Stacked on #2117 (retarget to
masteronce it lands). #2116 item 3a:[n({"id": is_in(seeds)}), e_forward(), n()](or a plain list of ids) returned the right rows but ran the isin scan, andgfql_explainrecorded nothing (used_index=False,decision_code=None,steps=[]) while the scalar seedn({"id": 5})and the Cypher twin (WHERE a.id IN [...], #2117) took the index.Mechanism:
chain._try_chain_fast_pathadmits the shape, but_seeded_scalar_filtersbails on theIsInvalue, so the native seeded-hop lane returns None and the general chain records nothing for the seed.Fix, mirroring #2117 rather than forking:
_seeded_seed_filtersresolves a membership set on the node-id key through the bindings kernel's own_membership_seed_ids;_seed_node_rows_from_indexlooks the ids up in the node-id index (it already took a list) and skips the property index for them; the scan fallback usesisin;_indexed_kernel_admitscounts the tuple as seeded-on-binding. The native seeded-hop and single-node lanes take it; the Cypher RETURN-destination lane stays scalar-only so Cypher membership seeds keep going to the bindings kernel exactly as #2117 pins. The native lane's scan branch now records a decline (index_path_unavailable) instead of leaving explain silent.Measured at 20k nodes / 100k edges, 50 seeds, pandas (local, direction only):
is_in6.4 ms → 1.4 ms, list 24.4 ms → 1.4 ms;gfql_explainrecordsnative_seeded_hop/native_seed_lookupwithindex_selectedunder auto/use/force,policy_offunder off. Residual filters on the seed node and the destination still apply (135 / 127 rows == truth); non-integral or boolean members and membership on a non-id column keep the scan and its rows. Polars already took the index for this shape via its own route (unchanged).Test plan
graphistry/tests/compute/chain_specializations/test_native_membership_seed_3a.py: 7 pandas pass; the 4 cuDF twins hit this box's missinglibnvrtc(GPU lane decides); the 4 behaviour pins fail at perf(gfql/cypher): multi-seedWHERE a.id IN [...]hop 4,004 ms → 10 ms (vectorized IN, index reaches the row pipeline, IN seeds the pattern, kernel takes a seed set) #2117's head-k "not cudf"), perf(gfql/cypher): multi-seedWHERE a.id IN [...]hop 4,004 ms → 10 ms (vectorized IN, index reaches the row pipeline, IN seeds the pattern, kernel takes a seed set) #2117's seam pins includedbin/lint.sh, mypy clean🤖 Generated with Claude Code
https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp