Repository navigation
Investigate RAPIDS 26.02 string-ID graph-build regression with pure cuDF/cuGraph repro #977
Description
Activity
Related: GFQL SIGSEGV in RAPIDS 25.02 on SNB sf1 (2026-04-05)
Context: pyg-bench branch
perf/rebaseline-2026-04-05, pincd707f0a(0.54.0), RAPIDS containernvcr.io/nvidia/rapidsai/base:25.02-cuda12.8-py3.12, dgx-spark (CUDA 13.0 driver).While attempting GPU rebaseline, the RAPIDS 25.02 container produces
returncode 139(SIGSEGV) on any GFQL execution withengine="cudf", including the simplest seed-lookup query. Triage:- Container starts: ✅
pip install pyg-bench: ✅import graphistry: ✅import cudf; cudf.read_csv(sf1_path): ✅graph.gfql(query, engine="cudf"): ❌ SIGSEGV
This may overlap with the string-ID regression tracked here. Crash occurs on sf1 (scale factor 1, ~10k persons) and even on the smallest seed-lookup pattern. Prior GPU baseline (530 ms) was measured with an older pygraphistry pin + older RAPIDS — unclear if the SIGSEGV is a RAPIDS 25.02 API change or a recent pygraphistry regression in the cudf code path.
Impact: GPU baseline entirely blocked. Prior 530 ms warmed GPU baseline (IC8) cannot be reproduced until resolved.
- added a commit that references this issue
on Apr 6, 2026 PR #1066 opened to fix the SIGSEGV sub-issue (RAPIDS 25.02, any
engine="cudf"query). Root cause:Series.apply(lambda)andSeries.map(dict)on cudf Series trigger numba JIT which SIGSEGVs. Primary site:filter_by_dict.py:_label_series_contains(). Secondary:hop.py,chain.py,df_executor.py. Fix: replace with vectorized/pandas-bridge equivalents.- added 7 commits that reference this issue
on Apr 6, 2026 Fix landed in PR #1066. Root cause confirmed via Python
faulthandler:Series.to_pandas()on numerical cudf columns →numba_cuda.as_cuda_array→ SIGSEGV on RAPIDS 25.02. Fixed withsafe_map_series()inEngine.pythat uses merge+to_arrow()to stay on GPU, applied at 7 call sites. Validated on both 25.02 and 26.02 containers on DGX.- added 13 commits that reference this issue
on Apr 6, 2026
Summary
We found a GPU performance regression when upgrading the official RAPIDS base image from
25.02-cuda12.8to26.02-cuda13on theto_cugraph()/from_cudf_edgelist()path used by our GFQL benchmarks.The important update is that this reproduces with a pure cuDF/cuGraph script and no PyGraphistry imports, and the repro has now been revalidated from current
masterondgx-spark.Where we first saw it naturally
In the PyGraphistry GFQL
gplusGPU benchmark,25.02-cuda12.8 -> 26.02-cuda13regressed on the warm path:pipeline_total:1.7019s -> 2.1780s(+27.97%)pagerank stage:0.3240s -> 0.5108s(+57.65%)Direct
to_cugraph()+pageranktiming on the same workload was worse:total:0.8171s -> 1.3695s(+67.60%)build:0.7630s -> 1.3215s(+73.20%)pagerank:0.0537s -> 0.0487s(-9.31%)That suggested graph build / renumber, not the PageRank kernel.
Pure RAPIDS reproducer
Tracked artifact:
benchmarks/gfql/filter_pagerank/pure_rapids_string_build_repro.pyPrimary repro shape:
synthetic_string_gplus_shape10,000,000edges107,614unique verticesRepresentative commands:
Fresh revalidation on current
masterRevalidated from branch
feat/rapids-string-id-cugraph-reproondgx-spark:25.02-cuda12.8build:0.1866spagerank:0.0074stotal:0.1941s26.02-cuda13build:0.3130spagerank:0.0034stotal:0.3170sDelta:
build:+67.74%total:+63.32%So the pure-RAPIDS repro still holds on current
master, and the kernel is still faster on26.02.Controls
Sparse integer control,
synthetic_offset,10,000,000edges:25.02-cuda12.8:0.2224s26.02-cuda13:0.2173s-2.29%High-cardinality string control,
synthetic_string_offset,10,000,000edges:25.02-cuda12.8:0.6864s26.02-cuda13:0.8303s+20.96%Store-transposed check on the main repro shape:
25.02-cuda12.8:0.2006s26.02-cuda13:0.3166s+57.83%So:
store_transposeddoes not materially change the storyWhy filing here first
This still looks upstream-facing, but the original natural failure was in our benchmark suite and our wrapper is where the investigation started. Filing here first gives us a stable local record and a place to decide whether to escalate directly to RAPIDS/cuGraph.
Requested outcome