Repository navigation
GFQL: Cypher multi-seed hop (WHERE a.id IN [...]) does not use the resident adjacency index; single-seed does #2116
Description
Activity
Chased this on master
37d7d9b6a; the index gap is real but it was not where the time went. Same query shape at 100k nodes / 500k edges, pandas, 50 seeds:MATCH (a)-[e]->(b) WHERE a.id IN [50 ids] RETURN b: 4,004 ms, identical underindex_policyuse / off / force. The native[n({"id": is_in(seeds)}), e_forward(), n()]chain: 7.6 ms.g.hop(seed_df): 0.8 ms.- cProfile: 15.6 of 15.7 s (under the profiler) in
row/pipeline.py::_gfql_eval_in_expr—INwas evaluated per row × per list element in Python (_gfql_cypher_value_equalcalled 5,005,803 times). Cost model ≈ 400 ms + 75 ms per list element at 100k rows, which is whyIN [2]was already 566 ms. - Why single-seed looked indexed and multi-seed not:
WHERE a.id IN [...]lowers to an alias prefilter +where_rows, not onto the pattern'sfilter_dictthe waya.id = xdoes, soconnected_bindingsdeclines at its first gate (prefilters present) and the hop runs inside the row pipeline — wheremaybe_index_hopmissed the registry's identity guard because the pipeline tags the edge frame with__gfql_edge_ident_0__without migrating the index.forceonly "worked" by rebuilding the CSR per query. - Also: with only
edge_out_adjcreated (as in the repro), the single-seedindex_selectedis theseeded_typed_hopfast path step, which the trace recorder labelspath: "index"; the real index seam (destination_return) reportsindex_missinguntilnode_idis also built. The docs example should create both.
#2117 fixes the three layers (vectorized
INlane 4,004 → 101 ms; index migrates onto the pipeline's edge frame → 67 ms with the hop step nowindex;IN [literals]seeds the pattern asis_in→ 19.7 ms with indexes, 62 ms without). Still open after it: the indexed bindings kernel declines a membership seed, so the unseeded middle runs once before the bindings path — that is the piece that closes this issue as filed, and it is next. Pre-existing and separate:WHERE NOT (b.id IN [...])raises "AST evaluator unsupported".Attribution correction, from the #2117 investigation: the observable here (index_path_unavailable → scan for
WHERE a.id IN [...]) was accurate, but the index miss was only one of four layers, and not the dominant cost. The row pipeline evaluatedINper row × per list element in Python (~5M calls for 50 seeds on 100k nodes / 500k edges, ~4 s), the IN predicate lowered as a post-join prefilter instead of seeding the pattern, the row pipeline's edge-frame copy broke the index registry's identity guard, and the bindings kernel declined membership seeds.#2117 fixes all four (4,004 ms → 10.5 ms indexed / 62 ms unindexed on pandas; cuDF 19 ms, polars 5.4 ms). Leaving this open for #2117 to close.
Status of the three follow-ups (2026-10-03), each with its own evidence trail in the PR:
- 3b
WHERE NOT (b.id IN [...])→ "AST evaluator unsupported": reproduced and fixed in fix(gfql): row-values series carry the table index so NOT IN keeps its rows under the index path #2121. It fires only when the NOT excludes every surviving row — the evaluator built its mask from an empty list (float64/object instead of bool). Same root also made cuDF raise on anINthat matches nothing (GFQL cuDF:WHERE alias.col IN [...]that matches no row raises "cudf does not support mixed types" instead of returning no rows #2125) and pandas raise on an empty list<.where_rowsnow validates then short-circuits on an empty frame; tri-valued results are boolean even when empty. - 3a native
[n({"id": is_in(seeds)}), e_forward(), n()]not taking the index / explain silent: perf(gfql): native op-lists seeded on a membership set take the resident index (#2116 3a) #2127 (stacked on perf(gfql/cypher): multi-seedWHERE a.id IN [...]hop 4,004 ms → 10 ms (vectorized IN, index reaches the row pipeline, IN seeds the pattern, kernel takes a seed set) #2117). The native lanes resolve a membership seed through the kernel's own_membership_seed_ids;gfql_explainrecordsnative_seeded_hop/native_seed_lookupwithindex_selected, and the scan branch now records its decline instead of nothing. 1.4 ms vs 6–24 ms at 20k/100k locally. Cypher membership seeds keep going to the bindings kernel as perf(gfql/cypher): multi-seedWHERE a.id IN [...]hop 4,004 ms → 10 ms (vectorized IN, index reaches the row pipeline, IN seeds the pattern, kernel takes a seed set) #2117 pins. TheGRAPH { ... WHERE a.id IN [...] }form takes the index with it too. - 3c unseeded
RETURN b LIMIT 5and "cuDF IN ... OR": profiled first. TheIN ... ORcost is not cuDF-specific and not theIN— a cross-aliasORdisables the prefilter and the string equality then paid ~900 ms of type sniffing (regex over every value, both operands); perf(gfql): row-evaluator text sniffers decide on a sample before scanning a column (#2116 3c) #2128 makes the sniffers decide on a sample (exact, 176-series equivalence). TheLIMIT 5case is the full binding table built beforeLIMIT— filed as GFQL Cypher: an unseeded MATCH ... RETURN b LIMIT k builds the full binding table before LIMIT (81 ms vs 3.6 ms native at 100k edges) #2129 (design item).
- 3b
Landed in #2117 (merge ceb095f). The multi-seed hop is served by the resident indexes end to end —
WHERE a.id IN [...]lowers onto the pattern asis_in, the bindings kernel accepts the membership seed, and the row pipeline's hop reaches the adjacency index; 100k nodes / 500k edges, 50 seeds: 4,004 ms → 10.5 ms (pandas), cuDF 19 ms, polars 5.4 ms. The shipped GraphBench/SNB board was re-measured at that tree and republished (graphistry/pyg-bench#283), so the docs drift budget is reset without a waiver. Remaining, separate: the native[n({"id": is_in(...)}), e_forward(), n()]chain still takes the isin-scan fast path and records nothing ingfql_explain;WHERE NOT (x IN [...])raises in the AST evaluator.
Summary
With a resident
edge_out_adjindex and the defaultindex_policy='use', a Cypher seeded hop uses the index when the seed is a single id but falls back to the O(E) scan when the seeds are an IN list. The multi-seed form is the one the index exists for ("the neighbors of these 50 accounts"), so the fast path is missing exactly where it matters.Repro (graphistry 0.59.0+320, pandas 2.3.3)
WHERE a.id = 'a'behaves like the inline-property form (index_selected).UNWIND [...] AS sid MATCH (a {id: sid})raises the known multi-source row-lowering limit (#1273), so there is currently no Cypher spelling of a multi-seed indexed hop.Answers are identical with
index_policy='off'in every case; this is a performance gap, not a correctness one.Related, not duplicates
node_idindex (chain node filter). This issue is the hop path with a seed set.connected_bindingsseam reportsunsupported_shapehere.[n({"id": is_in(seeds)}), e_forward(), n()]reportsused_index=Falsewithdecision_code=Noneunder bothuseandforce— i.e. explain records no decision at all for it. Whether that chain is served by the index or merely unreported is a separate question worth a pin; the docs currently state it is used automatically.Where found
While verifying a Cypher example for the
index_adjacencydocs page actually engages the index before publishing it. The docs will lead with the single-seed form that is proven to engage.🤖 Generated with Claude Code
https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp