Skip to content

GFQL: cuDF cross-engine result divergences (list-literal order, toString(float), min_hops seed hop-label, group_by Series-truthiness) #1663

Description

@lmeyerov

Summary

Four cuDF-vs-pandas result divergences in GFQL, all found while building the native polars-engine
differential-conformance matrix (graphistry/tests/compute/gfql/test_engine_polars_conformance_matrix.py).
All are cuDF-engine issues: g.gfql(query, engine='cudf') differs from engine='pandas' (the oracle).
They are orthogonal to the polars engine (polars is parity-or-honest-NIE on each) and are currently
scoped out of the 4-engine conformance _assert_invariant (dedicated pandas-vs-polars tests cover the
polars intent), so they don't block the polars work — but each is a real cuDF correctness gap.

Repro pattern: run the same query on engine='pandas' vs engine='cudf' and compare value-level.


1. cuDF reorders list-literal [a, b, c] elements vs pandas

A cypher list literal materialized into a column (e.g. a row-pipeline expr building [n.num, n.num+1, 99])
comes back with the list ELEMENTS permuted under cuDF; pandas preserves construction order.

  • Expected: element order matches pandas (construction order). Severity: wrong-answer for list-valued projections.

2. cuDF formats toString(float) differently than pandas

toString(n.f) over a float column yields a different string representation under cuDF than pandas
(precision / trailing zeros / exponent style).

  • Expected: match pandas str(float) formatting (or document a canonical format). polars declines this as
    honest-NIE (it also can't match pandas float-repr), so cuDF is the silent-divergence here.

3. cuDF multi-hop min_hops>1 labels the SEED node's hop wrong

For e_forward(min_hops=2, max_hops=3) etc., the SEED node appears with __gfql_output_node_hop__ =
max_hops under cuDF but None/NaN under pandas. (Secondary: num comes back int under cuDF vs float
under pandas for the same result.)

  • Repro: [n({"id":[0]}), e_forward(min_hops=2, max_hops=3), n()] on a small attributed graph; compare the
    seed row's __gfql_output_node_hop__. Found via the NA-hardened conformance signature.

4. cuDF group_by row-op raises "truth value of a Series is ambiguous"

call("group_by", {"keys":[k], "aggregations":[("c","count"),("s","sum",col)]}) on a row table that carries
EXTRA non-key/non-aggregated columns (e.g. float f + string name alongside grouped flag/num) raises:
GFQLTypeError: [invalid-node-reference] Error executing 'group_by': The truth value of a Series is ambiguous.

  • Repro: a 5-col node frame (id/num/f/name/flag) + the group_by above; pandas+polars return [flag,c,s],
    cuDF raises. A minimal 3-col graph does NOT trigger it — the extra columns drive a if <series>: path
    (should be .any()/.all()). Likely in the GFQL group_by handler graphistry/compute/gfql/row/pipeline.py.

Findings 1–2 were known from earlier sessions; 3–4 were found 2026-06-30. None blocks the polars engine PRs.
🤖 Generated with Claude Code

Activity

  1. lmeyerov commented on Aug 18, 2026

    @lmeyerov
    ContributorAuthor

    Verified fixed on master e6625ed28 (cudf 25.10) — all three filed divergences now match hand-computed oracles; recommending close after review, with one small dtype footnote.

    1. List-literal order — MATCH (n) RETURN n.id AS id, [n.num, n.num + 1, 99] AS lst over num=[5,7]:
      pandas and cudf both return [(0, [5, 6, 99]), (1, [7, 8, 99])] (hand-computed construction order preserved).

    2. toString(float) — toString(n.f) over [1.5, 2.0, 0.1]:
      pandas and cudf both return ['1.5', '2.0', '0.1'] (= str(float)).

    3. min_hops seed hop label — [n({"id":[0]}), e_forward(min_hops=2, max_hops=3), n()] on chain 0→1→2→3→4:
      cudf seed row's __gfql_output_node_hop__ is now NaN (was mislabeled with max_hops), matching pandas None; hop labels 1/2/3 agree across engines.

    4. (Bonus, same-issue family) group_by row op with extra non-key/non-agg columns no longer crashes on cudf: both engines return flag=True → c=3, s=8; flag=False → c=2, s=7 (hand-computed).

    Residual footnote (arguably not a bug): in case 3, cudf keeps num as int64 while pandas upcasts to float64 via the NaN-introducing join — values are equal and cudf is the more faithful dtype. If that matters it can be split into its own cosmetic issue; it doesn't block closing this one.

  2. lmeyerov commented on Aug 18, 2026

    @lmeyerov
    ContributorAuthor

    Closing: verified fixed on master e6625ed28 by the 2026-08 stack — empirical repro + hand-computed oracle in the verification comment above. Residual sub-items, where any, are tracked in the successor issues named there (#1916, #1908, #1906, #1934-#1938).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions