Skip to content

Investigate RAPIDS 26.02 string-ID graph-build regression with pure cuDF/cuGraph repro #977

Description

@lmeyerov

Summary

We found a GPU performance regression when upgrading the official RAPIDS base image from 25.02-cuda12.8 to 26.02-cuda13 on the to_cugraph() / from_cudf_edgelist() path used by our GFQL benchmarks.

The important update is that this reproduces with a pure cuDF/cuGraph script and no PyGraphistry imports, and the repro has now been revalidated from current master on dgx-spark.

Where we first saw it naturally

In the PyGraphistry GFQL gplus GPU benchmark, 25.02-cuda12.8 -> 26.02-cuda13 regressed on the warm path:

  • pipeline_total: 1.7019s -> 2.1780s (+27.97%)
  • pagerank stage: 0.3240s -> 0.5108s (+57.65%)

Direct to_cugraph()+pagerank timing on the same workload was worse:

  • total: 0.8171s -> 1.3695s (+67.60%)
  • build: 0.7630s -> 1.3215s (+73.20%)
  • pagerank: 0.0537s -> 0.0487s (-9.31%)

That suggested graph build / renumber, not the PageRank kernel.

Pure RAPIDS reproducer

Tracked artifact:

  • benchmarks/gfql/filter_pagerank/pure_rapids_string_build_repro.py

Primary repro shape:

  • synthetic_string_gplus_shape
  • 10,000,000 edges
  • 107,614 unique vertices
  • repeated low-cardinality string/object IDs

Representative commands:

docker run --gpus all --rm -v "$WORKTREE:/workspace" -w /workspace \
  nvcr.io/nvidia/rapidsai/base:25.02-cuda12.8-py3.12 \
  python -W ignore benchmarks/gfql/filter_pagerank/pure_rapids_string_build_repro.py \
    --graph synthetic_string_gplus_shape --synthetic-edges 10000000 --runs 3 --warmup 1

docker run --gpus all --rm -v "$WORKTREE:/workspace" -w /workspace \
  nvcr.io/nvidia/rapidsai/base:26.02-cuda13-py3.13 \
  python -W ignore benchmarks/gfql/filter_pagerank/pure_rapids_string_build_repro.py \
    --graph synthetic_string_gplus_shape --synthetic-edges 10000000 --runs 3 --warmup 1

Fresh revalidation on current master

Revalidated from branch feat/rapids-string-id-cugraph-repro on dgx-spark:

  • 25.02-cuda12.8
    • build: 0.1866s
    • pagerank: 0.0074s
    • total: 0.1941s
  • 26.02-cuda13
    • build: 0.3130s
    • pagerank: 0.0034s
    • total: 0.3170s

Delta:

  • build: +67.74%
  • total: +63.32%

So the pure-RAPIDS repro still holds on current master, and the kernel is still faster on 26.02.

Controls

Sparse integer control, synthetic_offset, 10,000,000 edges:

  • 25.02-cuda12.8: 0.2224s
  • 26.02-cuda13: 0.2173s
  • delta: -2.29%

High-cardinality string control, synthetic_string_offset, 10,000,000 edges:

  • 25.02-cuda12.8: 0.6864s
  • 26.02-cuda13: 0.8303s
  • delta: +20.96%

Store-transposed check on the main repro shape:

  • 25.02-cuda12.8: 0.2006s
  • 26.02-cuda13: 0.3166s
  • delta: +57.83%

So:

  • this is not an integer sparse-ID problem
  • string/object IDs are implicated
  • repeated low-cardinality string/object IDs are still the strongest repro shape
  • store_transposed does not materially change the story

Why filing here first

This still looks upstream-facing, but the original natural failure was in our benchmark suite and our wrapper is where the investigation started. Filing here first gives us a stable local record and a place to decide whether to escalate directly to RAPIDS/cuGraph.

Requested outcome

  1. Decide whether this should remain a pure upstream-facing repro package.
  2. If yes, use this script plus the natural GFQL benchmark evidence as the filing package.
  3. If no, document what narrow local mitigation we do or do not want in PyGraphistry.

Activity

  1. lmeyerov commented on Apr 6, 2026

    @lmeyerov
    ContributorAuthor

    Related: GFQL SIGSEGV in RAPIDS 25.02 on SNB sf1 (2026-04-05)

    Context: pyg-bench branch perf/rebaseline-2026-04-05, pin cd707f0a (0.54.0), RAPIDS container nvcr.io/nvidia/rapidsai/base:25.02-cuda12.8-py3.12, dgx-spark (CUDA 13.0 driver).

    While attempting GPU rebaseline, the RAPIDS 25.02 container produces returncode 139 (SIGSEGV) on any GFQL execution with engine="cudf", including the simplest seed-lookup query. Triage:

    • Container starts: ✅
    • pip install pyg-bench: ✅
    • import graphistry: ✅
    • import cudf; cudf.read_csv(sf1_path): ✅
    • graph.gfql(query, engine="cudf"): ❌ SIGSEGV

    This may overlap with the string-ID regression tracked here. Crash occurs on sf1 (scale factor 1, ~10k persons) and even on the smallest seed-lookup pattern. Prior GPU baseline (530 ms) was measured with an older pygraphistry pin + older RAPIDS — unclear if the SIGSEGV is a RAPIDS 25.02 API change or a recent pygraphistry regression in the cudf code path.

    Impact: GPU baseline entirely blocked. Prior 530 ms warmed GPU baseline (IC8) cannot be reproduced until resolved.

  2. added a commit that references this issue on Apr 6, 2026
  3. lmeyerov commented on Apr 6, 2026

    @lmeyerov
    ContributorAuthor

    PR #1066 opened to fix the SIGSEGV sub-issue (RAPIDS 25.02, any engine="cudf" query). Root cause: Series.apply(lambda) and Series.map(dict) on cudf Series trigger numba JIT which SIGSEGVs. Primary site: filter_by_dict.py:_label_series_contains(). Secondary: hop.py, chain.py, df_executor.py. Fix: replace with vectorized/pandas-bridge equivalents.

  4. lmeyerov commented on Apr 6, 2026

    @lmeyerov
    ContributorAuthor

    Fix landed in PR #1066. Root cause confirmed via Python faulthandler: Series.to_pandas() on numerical cudf columns → numba_cuda.as_cuda_array → SIGSEGV on RAPIDS 25.02. Fixed with safe_map_series() in Engine.py that uses merge+to_arrow() to stay on GPU, applied at 7 call sites. Validated on both 25.02 and 26.02 containers on DGX.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions