Two related polars-gpu findings from the 2026-08 dgx-spark validation (GB10, cudf 26.02.01, cudf-polars 26.02.01, 26.02-gfql-polars image). Neither is a wrong query answer in GFQL's own logic; both are worth tracking before we make polars-gpu performance claims.
1. cudf_polars 26.02 returns None where CPU returns 0 — and does NOT raise
Grouped sum() over a Boolean column, with a widening signed cast on the aggregate result, diverges between CPU and GPU collect. Critically GPUEngine(raise_on_fail=True) does not raise — it silently answers differently.
Pure-polars repro, no graphistry involved:
import polars as pl
base = pl.LazyFrame({"city": ["LA","LA","NYC","NYC"],
"flag": pl.Series([None, None, True, False], dtype=pl.Boolean)})
p = base.group_by("city").agg(pl.col("flag").sum().cast(pl.Int64).alias("s")).sort("city")
p.collect() # LA -> 0 correct
p.collect(engine=pl.GPUEngine(raise_on_fail=True)) # LA -> None WRONG, no raise
Characterization sweep — it is only grouped sum() over Boolean with a widening signed cast on the result:
| expression |
GPU vs CPU |
sum().cast(Int64) |
DIVERGES |
sum() (no cast) |
agree |
sum().cast(UInt32) (same width) |
agree |
cast(Int64).sum() (cast the input) |
agree |
sum().fill_null(0).cast(Int64) |
agree |
mean(), count().cast(Int64), min()/max() |
agree |
Our trigger: polars_agg_result_cast in graphistry/compute/gfql/agg_types.py casts the result. Any of the three agreeing forms above would sidestep it. Blast radius is confined to the OLAP fused fast path — the generic row pipeline answers correctly on polars-gpu.
Surfaced as test_aggregate_type_contract.py::test_fast_path_boolean_aggregate_follows_the_same_contract[polars-gpu] (assert rows[0]["s"] == 0 → None == 0), the single failure in an otherwise clean 493-passed run with all 141 polars-gpu params executing.
This is an upstream bug; the actionable part on our side is choosing a cast form that does not trip it, plus deciding whether to report it upstream.
2. The fused OLAP fast path is inert on polars-gpu
86 params in cypher/test_grouped_aggregate_fused_polars.py fail their engagement canary (assert calls == [True] / "fused lane must serve on polars-gpu"). These are not value comparisons — the answers are correct, served by the generic route.
Causes, all cudf_polars 26.02 gaps:
NotImplementedError: Unhandled map function unnest (the low-cardinality pure count(*) value_counts() + UNNEST plan)
NotImplementedError: No GPU support for DataFrameScan(... 'node_type': Null ...) with a Null column dtype
RuntimeError: CUDF failure ... type_dispatcher.hpp:556: Invalid type_id
#1979's no-silent-CPU contract HOLDS — the lane declines loudly and re-runs on the GPU target rather than quietly serving CPU under a GPU label. So this is a capability/perf gap, not a correctness one.
It does have a correctness consequence indirectly: because polars-gpu takes the generic route where other engines take the fast path, it exposes the multiplicity divergence tracked in #1996.
Test-matrix gaps found alongside
test_strictness_levels.py has no polars-gpu arm — ENGINES = ("pandas", "polars", "cudf"), and the tests resolve the engine from frame type, which cannot express polars-gpu. An ad-hoc variant found no polars-gpu strictness divergence, so this is a coverage gap rather than a behavior gap.
test_endpoint_closure_matrix.py's xfail marks are applied via _engines(polars=...), i.e. polars only — polars-gpu inherits the same divergence unmarked.
Two related polars-gpu findings from the 2026-08 dgx-spark validation (GB10, cudf 26.02.01, cudf-polars 26.02.01,
26.02-gfql-polarsimage). Neither is a wrong query answer in GFQL's own logic; both are worth tracking before we make polars-gpu performance claims.1. cudf_polars 26.02 returns
Nonewhere CPU returns0— and does NOT raiseGrouped
sum()over a Boolean column, with a widening signed cast on the aggregate result, diverges between CPU and GPU collect. CriticallyGPUEngine(raise_on_fail=True)does not raise — it silently answers differently.Pure-polars repro, no graphistry involved:
Characterization sweep — it is only grouped
sum()over Boolean with a widening signed cast on the result:sum().cast(Int64)sum()(no cast)sum().cast(UInt32)(same width)cast(Int64).sum()(cast the input)sum().fill_null(0).cast(Int64)mean(),count().cast(Int64),min()/max()Our trigger:
polars_agg_result_castingraphistry/compute/gfql/agg_types.pycasts the result. Any of the three agreeing forms above would sidestep it. Blast radius is confined to the OLAP fused fast path — the generic row pipeline answers correctly on polars-gpu.Surfaced as
test_aggregate_type_contract.py::test_fast_path_boolean_aggregate_follows_the_same_contract[polars-gpu](assert rows[0]["s"] == 0→None == 0), the single failure in an otherwise clean 493-passed run with all 141 polars-gpu params executing.This is an upstream bug; the actionable part on our side is choosing a cast form that does not trip it, plus deciding whether to report it upstream.
2. The fused OLAP fast path is inert on polars-gpu
86 params in
cypher/test_grouped_aggregate_fused_polars.pyfail their engagement canary (assert calls == [True]/ "fused lane must serve on polars-gpu"). These are not value comparisons — the answers are correct, served by the generic route.Causes, all cudf_polars 26.02 gaps:
NotImplementedError: Unhandled map function unnest(the low-cardinality purecount(*)value_counts()+ UNNEST plan)NotImplementedError: No GPU support for DataFrameScan(... 'node_type': Null ...)with a Null column dtypeRuntimeError: CUDF failure ... type_dispatcher.hpp:556: Invalid type_id#1979's no-silent-CPU contract HOLDS — the lane declines loudly and re-runs on the GPU target rather than quietly serving CPU under a GPU label. So this is a capability/perf gap, not a correctness one.
It does have a correctness consequence indirectly: because polars-gpu takes the generic route where other engines take the fast path, it exposes the multiplicity divergence tracked in #1996.
Test-matrix gaps found alongside
test_strictness_levels.pyhas no polars-gpu arm —ENGINES = ("pandas", "polars", "cudf"), and the tests resolve the engine from frame type, which cannot express polars-gpu. An ad-hoc variant found no polars-gpu strictness divergence, so this is a coverage gap rather than a behavior gap.test_endpoint_closure_matrix.py's xfail marks are applied via_engines(polars=...), i.e. polars only — polars-gpu inherits the same divergence unmarked.