Skip to content

fix(gfql): resolve SIGSEGV with engine="cudf" on RAPIDS 25.02 (#977) - #1066

Merged
lmeyerov merged 20 commits into
masterfrom
feat/issue-977-cudf-sigsegv
Apr 6, 2026
Merged

lmeyerov merged 20 commits into
masterfrom
feat/issue-977-cudf-sigsegv

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Apr 6, 2026 •

Copy link
Copy Markdown
Contributor

Closes #977.

Problem

graph.gfql(query, engine="cudf") crashes with SIGSEGV (returncode 139) on RAPIDS 25.02. Root cause traced to two distinct numba JIT trigger paths:

  1. Series.map(dict/Series) — triggers numba JIT on RAPIDS 25.02
  2. Series.to_pandas() on numerical columns — triggers numba_cuda.as_cuda_array → safe_cuda_api_call → SIGSEGV

(A third path, Series.apply(lambda), was fixed in #1061 via marker.notna().)

Root cause confirmed via Python faulthandler:

safe_map_series → series.to_pandas() → cudf.core.column.numerical.to_pandas →
values_host → data_array_view → numba_cuda.as_cuda_array → safe_cuda_api_call → SIGSEGV

Fix

Introduce safe_map_series(series, mapping) in Engine.py that for cudf:

  • Builds a lookup DataFrame from the mapping using to_arrow().to_pylist() (bypasses numba path)
  • Uses a GPU-native merge-based lookup that stays entirely on GPU; cudf left-merge preserves row order (verified on DGX 25.02)
  • Falls back to native .map() for pandas

Applied at 8 call sites: hop.py (4×), chain.py (1×), df_executor.py (2×), pipeline.py (1×).

The filter_by_dict.py cudf guard (lines 75–81) was already safe — converts to pandas before .apply().

Validation

Tests added

  • test_safe_map_series.py — 16 tests: pandas/cudf paths for dict, pd.Series, cudf.Series, cudf+pd.Series, missing keys, empty mapping, non-default index, NaN in series, mixed value types
  • test_hop.py — 9 new label tests: reverse, undirected, node-only, edge-only, empty result, empty result with edge hops, multi-hop ordering, missing-mask guard (hop.py:973), _mk_abc_chain() helper DRY
  • test_chain.py — chain integration test for label_node_hops propagation through combine_steps
  • test_lowering.py — label-filter regression, single-hop regression, multi-column ORDER BY test
  • graphistry/tests/compute/gfql/cypher/_dgx_977_smoke.py — DGX smoke script (pandas + cudf paths)

lmeyerov and others added 15 commits April 6, 2026 01:16
…t) with pandas bridge (#977)

cudf Series.map(dict/Series) triggers numba JIT which SIGSEGVs on RAPIDS 25.02.
Introduce safe_map_series() in Engine.py that bridges through pandas for cudf.
Apply at all 7 affected call sites: hop.py (4), chain.py (1), df_executor.py (2), pipeline.py (1).
Includes DGX smoke test + unit tests for safe_map_series + cudf guard regression tests.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
edge_map is only assigned inside the edge_hop_col branch; line 974
used it unconditionally which raised UnboundLocalError when
edge_hop_col is None. Previously masked by SIGSEGV on 25.02.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
@lmeyerov
lmeyerov force-pushed the feat/issue-977-cudf-sigsegv branch from d41a24c to 527ace2 Compare April 6, 2026 08:18
lmeyerov and others added 5 commits April 6, 2026 01:27
<<<<<<< HEAD line was left from incomplete conflict resolution
during rebase; caused SyntaxError in CI test collection.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Gap tests:
- safe_map_series: cudf + pd.Series mapping path (Engine.py:399-400)
- hop: missing_mask guard when edge_hop_col=None (hop.py:973)
- hop: empty edge_map_df else branch (hop.py:961-965)

DRY: extract _mk_abc_chain() helper in test_hop.py — 5 label tests refactored

CHANGELOG: fix call-site count (8, not 7); add #977 Tests entry (23 tests);
correct file list (no gfql_unified.py — that fix was in #1061)

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
New tests:
- output_max_hops path (hop.py:834-891 fallback_map try/except + line-1011 post-filter)
- label_seeds=True with implicit starting_nodes (hop.py:843-859 both-branches path)
- numeric (int) node IDs through safe_map_series int→int hop labeling (hop.py:938)
- undirected label_seeds=False clears seed hop (hop.py:987-1000)

DRY: extract _mk_abcd_chain() helper — multi_hop_ordering + output_max_hops refactored

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
…hs (#977)

New smoke tests cover:
- safe_map_series + pd.Series mapping (Engine.py:399-400)
- output_max_hops filtering (hop.py:834-891 fallback_map + line-1011 post-filter)
- label_seeds=True seed rows from node_hop_records (hop.py:843-859)
- label_seeds=False + undirected clears seed hop (hop.py:987-1000)

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
…umns

to_pandas() on integer cudf Series SIGSEGVs on RAPIDS 25.02 (same numba
path as the original #977 bug). String columns are safe. Switch all numeric
hop value reads to to_arrow().to_pylist() to avoid the numba JIT trigger.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
@lmeyerov
lmeyerov merged commit 081187c into master Apr 6, 2026
101 checks passed
@lmeyerov
lmeyerov deleted the feat/issue-977-cudf-sigsegv branch April 6, 2026 18:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Investigate RAPIDS 26.02 string-ID graph-build regression with pure cuDF/cuGraph repro

1 participant