Repository navigation
gfql: rows(table=nodes, source=alias) multiplies rows for duplicate node ids and joins null ids to each other #2034
Description
Activity
Disposition (0.60 landing phase): deferred behind the stacks. The seeded fast paths (#2035/#2037/#2038) return the intended one-row-per-node answer for duplicate and null ids and pin it; the full-path chain combine still self-joins on the id column. This is the same chain-combine seam as the cuDF 26.2 duplicate-id divergence in #2043, so both get fixed together right after the stacks land (positional node-frame lookups; null ids never link). Not release-blocking on its own: the shapes that reach the full path with duplicate ids are the ones the fast paths already decline, and the differential harness excludes duplicate/null ids until then.
Evidence update (round-003 differential sweep of the native op-list lanes, 2026-09-05): with a repeated node row, the native single-node lookup in #2037 answers one row per matching table row (as polars does on both routes) while the pandas/cuDF full path still self-joins to 8; #2037 now carries a two-sided pin (test_native_seed_resolution_2027.py::test_duplicate_node_rows_are_answered_once_each_on_the_native_lookup) so the full-path fix flips it. Disposition unchanged: fix the chain-combine seam together with #2043 right after the stacks land.
- added a commit that references this issue
on Sep 6, 2026 Status after the routes-off replay (#2054/#2061) and the dtype fixes (#2062): unchanged. With every hot path declined,
MATCH (p:Person {id: 7}) RETURN p.score AS son a node table carrying twoid = 7rows still returns 8 rows on pandas (the combine's node-id self-join multiplies the duplicates through forward/reverse/combine), while the seeded lanes return 2; the all-off ledger liststest_node_lookup_returns_each_duplicate_id_row_onceas the one result divergence outside #2058/#2059.One contract question to settle before the fix, since two pins disagree by shape today:
test_chain.py::test_fast_path_dedups_duplicate_node_ids_on_hoprequires a 1-hop chain to COLLAPSE duplicate node-id rows (the closure step'sdrop_duplicates(subset=[node]), which #2062 keeps), whereas this issue's pin requires a node lookup to return EACH duplicate row once. Proposal: node-frame lookups (rows(table=nodes, source=alias)and the seeded lanes) key on row position and keep one output row per source row; hop/chain results keep the documented collapse; null ids never link on either. The fix then lives incombine_steps' node self-join (positional key) andframe_ops.rows(source=...), with the two pins stated side by side.- added 8 commits that reference this issue
on Sep 6, 2026
Found while pinning the seeded node-lookup fast path (parity harness on the branch for the seeded fast paths, 2026-09-04).
id = 7,MATCH (p {id: 7}) RETURN p.age AS aon the full path returns 8 rows on pandas and cuDF (2 duplicate rows → 2^3 through the chain's forward/reverse/combine self-joins); polars returns 2. The expected answer is one row per matching node row (2).idis null,MATCH (p {name: 'n3'}) RETURN p.id, p.agereturns both null-id rows on the full path (NaN-keyed merge matches NaN to NaN) instead of the one row that matches the predicate. A seeded typed hop whose seed row has a null id raisesGFQLTypeErroron the full path (merge object vs float64) instead of returning no rows.The seeded fast paths (
_execute_seeded_node_lookup_fast_path,_execute_seeded_typed_hop_fast_path) return the intended answer for these inputs;graphistry/tests/compute/gfql/test_seeded_node_lookup_fastpath.py::test_node_lookup_returns_each_duplicate_id_row_oncepins the fast answer on all three engines. The full path should agree; until it does, the differential parity tests exclude duplicate and null ids.Where:
graphistry/compute/chain.py_chain_implcombine step (self-join on the node id) andgraphistry/compute/gfql/row/frame_ops.pyrows(source=...); the merges should key on row position for node-frame lookups, and null ids must never link.