Skip to content

fix(test): route IC6 cuDF assertion through Arrow-safe helper - #1437

Merged
lmeyerov merged 1 commit into
masterfrom
fix/ic6-cudf-test-rapids-25.02-segfault-1433
May 15, 2026
Merged

lmeyerov merged 1 commit into
masterfrom
fix/ic6-cudf-test-rapids-25.02-segfault-1433

Conversation

@lmeyerov

Copy link
Copy Markdown
Contributor

Summary

Fixes a numba_cuda segfault during test_issue_1396_issue_1415_tag_cooccurrence_join_aggregation_counts_on_cudf on the RAPIDS 25.02-cuda12.8 aarch64 image. The test bypassed the existing _to_pandas_df helper (graphistry/tests/compute/gfql/cypher/test_lowering.py:851) whose own comment cites this exact failure mode ("RAPIDS 25.02 can segfault on some direct to_pandas() paths; prefer Arrow."). Switched to the helper, which routes cuDF DataFrames through to_arrow().to_pandas() and avoids numba_cuda's CUDA driver context init path.

Surfaced while validating graphistry/pyg-bench#12 (benchmark-side closeout of #1415) — the segfault is preexisting infrastructure brittleness on 25.02 + ARM, not a regression from #1427 or #1396/#1426, but it blocks the IC6 RAPIDS 25.02 receipt.

Test plan

  • DGX docker/test-rapids-official-local.sh with RAPIDS_IMAGE=nvcr.io/nvidia/rapidsai/base:25.02-cuda12.8-py3.12 — both test_issue_1396_issue_1415_tag_cooccurrence_join_aggregation_counts and _on_cudf PASS (previously the _on_cudf variant segfaulted in numba_cuda during cudf.to_pandas()).
  • DGX docker/test-rapids-official-local.sh with RAPIDS_IMAGE=nvcr.io/nvidia/rapidsai/base:26.02-cuda13-py3.13 CUDA_VARIANT=cuda13 — both variants PASS.
  • Confirmed via sanity check (test_graph_constructor_cudf_support on RAPIDS 25.02-cuda12.8) that the broader image isn't broken — only the direct .to_pandas() path triggers the segfault.

Out of scope

  • Other direct .to_pandas().to_dict() callsites in test_lowering.py. They don't currently segfault but are latently vulnerable to the same numba_cuda path. A sweep to use _to_pandas_df everywhere is a follow-up.
  • pyg-bench's scripts/run_dgx_spark_suite.py defaulting to the broken 25.02-cuda12.8 image. Addressed in a separate pyg-bench PR.

🤖 Generated with Claude Code

…#1415, #880)

`test_issue_1396_issue_1415_tag_cooccurrence_join_aggregation_counts_on_cudf`
called `result._nodes.to_pandas().to_dict(orient="records")` directly, which
segfaults inside `numba_cuda`'s CUDA driver context init on the RAPIDS
25.02-cuda12.8 aarch64 image (crash chain: `cudf.core.column.numerical.to_pandas`
→ `values_host` → `data_array_view` → `as_cuda_array` → `_require_cuda_context`
→ `safe_cuda_api_call`).

The file already defines `_to_pandas_df` (line 851) precisely for this
case — its comment explicitly cites the RAPIDS 25.02 segfault path and
prefers Arrow conversion. The IC6 test now routes through that helper.

Validated on DGX RAPIDS 25.02 (`nvcr.io/nvidia/rapidsai/base:25.02-cuda12.8-py3.12`)
and 26.02 (`nvcr.io/nvidia/rapidsai/base:26.02-cuda13-py3.13`) via
`docker/test-rapids-official-local.sh`: both pandas and cuDF variants
PASS on both images after the fix.

Other direct `.to_pandas().to_dict(...)` callsites in this file are
left alone for now — a sweep is a separate concern.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant