Skip to content

GFQL: polars-gpu answers a different count() than every other engine on duplicate node-id rows (fast-path vs generic-route divergence) #1996

Description

@lmeyerov

Found on dgx-spark (GB10, cudf 26.02.01, cudf-polars 26.02.01) during the 2026-08 GPU validation. Root-caused as engine-independent — polars-gpu is the surface that currently exposes it, not the cause.

Repro

MATCH (p {kind:'P'})-[]->(c {kind:'C'}) RETURN c.city AS city, count(*) AS n ORDER BY n DESC, city ASC

on a fixture with duplicate node-id rows (_dup_start_node_rows_data):

engine result
pandas LA=6, NY=2, SF=1
polars LA=6, NY=2, SF=1
cudf LA=6, NY=2, SF=1
polars-gpu LA=10, NY=3, SF=1

Repro script: dgx-spark:/home/lmeyerov/gpuval-repro/probe_dup.py.

Why polars-gpu differs — and why that is not the real bug

The OLAP fused fast path is currently inert on polars-gpu: cudf_polars 26.02 cannot execute its plan (NotImplementedError: Unhandled map function unnest, No GPU support for DataFrameScan(... 'node_type': Null ...), and a type_dispatcher Invalid type_id). The raise escapes _execute_single_hop_grouped_aggregate_fast_path before its eager twin can serve, so polars-gpu lands on the generic route while every other engine lands on the fast path.

The two routes disagree. Confirmed by stubbing the fast path to None (probe_generic.py): with the fast path disabled, pandas, polars and cuDF all answer 10/3 too. So this is a pre-existing fast-path-vs-generic-route multiplicity divergence on duplicate node-id rows, and polars-gpu merely makes it user-visible today.

Which answer is right

Under openCypher bag semantics, duplicate node-id rows in the source table should each contribute, so the generic route's higher count may be the correct one and the fast path may be the side that silently dedupes. That needs deciding — closely related to #1994 (whole-entity projection multiplicity), which is the same "the fast path answers a set where a bag is required" family.

Expected

The fast path and the generic route must agree, whichever answer is correct. Today a user gets a different count purely from which engine they selected, with no warning.

Note on discovery

This is only observable with a working GPU. There is no CI lane that runs cuDF or polars-gpu (bin/ci_gpu_gate_audit.py reports "37 gates across 23 test files / workflows setting TEST_CUDF: NONE"), so nothing in CI can catch this class.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions