Found on dgx-spark (GB10, cudf 26.02.01, cudf-polars 26.02.01) during the 2026-08 GPU validation. Root-caused as engine-independent — polars-gpu is the surface that currently exposes it, not the cause.
Repro
MATCH (p {kind:'P'})-[]->(c {kind:'C'}) RETURN c.city AS city, count(*) AS n ORDER BY n DESC, city ASC
on a fixture with duplicate node-id rows (_dup_start_node_rows_data):
| engine |
result |
| pandas |
LA=6, NY=2, SF=1 |
| polars |
LA=6, NY=2, SF=1 |
| cudf |
LA=6, NY=2, SF=1 |
| polars-gpu |
LA=10, NY=3, SF=1 |
Repro script: dgx-spark:/home/lmeyerov/gpuval-repro/probe_dup.py.
Why polars-gpu differs — and why that is not the real bug
The OLAP fused fast path is currently inert on polars-gpu: cudf_polars 26.02 cannot execute its plan (NotImplementedError: Unhandled map function unnest, No GPU support for DataFrameScan(... 'node_type': Null ...), and a type_dispatcher Invalid type_id). The raise escapes _execute_single_hop_grouped_aggregate_fast_path before its eager twin can serve, so polars-gpu lands on the generic route while every other engine lands on the fast path.
The two routes disagree. Confirmed by stubbing the fast path to None (probe_generic.py): with the fast path disabled, pandas, polars and cuDF all answer 10/3 too. So this is a pre-existing fast-path-vs-generic-route multiplicity divergence on duplicate node-id rows, and polars-gpu merely makes it user-visible today.
Which answer is right
Under openCypher bag semantics, duplicate node-id rows in the source table should each contribute, so the generic route's higher count may be the correct one and the fast path may be the side that silently dedupes. That needs deciding — closely related to #1994 (whole-entity projection multiplicity), which is the same "the fast path answers a set where a bag is required" family.
Expected
The fast path and the generic route must agree, whichever answer is correct. Today a user gets a different count purely from which engine they selected, with no warning.
Note on discovery
This is only observable with a working GPU. There is no CI lane that runs cuDF or polars-gpu (bin/ci_gpu_gate_audit.py reports "37 gates across 23 test files / workflows setting TEST_CUDF: NONE"), so nothing in CI can catch this class.
Found on dgx-spark (GB10, cudf 26.02.01, cudf-polars 26.02.01) during the 2026-08 GPU validation. Root-caused as engine-independent — polars-gpu is the surface that currently exposes it, not the cause.
Repro
on a fixture with duplicate node-id rows (
_dup_start_node_rows_data):Repro script:
dgx-spark:/home/lmeyerov/gpuval-repro/probe_dup.py.Why polars-gpu differs — and why that is not the real bug
The OLAP fused fast path is currently inert on polars-gpu: cudf_polars 26.02 cannot execute its plan (
NotImplementedError: Unhandled map function unnest,No GPU support for DataFrameScan(... 'node_type': Null ...), and atype_dispatcherInvalid type_id). The raise escapes_execute_single_hop_grouped_aggregate_fast_pathbefore its eager twin can serve, so polars-gpu lands on the generic route while every other engine lands on the fast path.The two routes disagree. Confirmed by stubbing the fast path to
None(probe_generic.py): with the fast path disabled, pandas, polars and cuDF all answer 10/3 too. So this is a pre-existing fast-path-vs-generic-route multiplicity divergence on duplicate node-id rows, and polars-gpu merely makes it user-visible today.Which answer is right
Under openCypher bag semantics, duplicate node-id rows in the source table should each contribute, so the generic route's higher count may be the correct one and the fast path may be the side that silently dedupes. That needs deciding — closely related to #1994 (whole-entity projection multiplicity), which is the same "the fast path answers a set where a bag is required" family.
Expected
The fast path and the generic route must agree, whichever answer is correct. Today a user gets a different count purely from which engine they selected, with no warning.
Note on discovery
This is only observable with a working GPU. There is no CI lane that runs cuDF or polars-gpu (
bin/ci_gpu_gate_audit.pyreports "37 gates across 23 test files / workflows setting TEST_CUDF: NONE"), so nothing in CI can catch this class.