Skip to content

Bulk g.hop() regression: LJ 1-hop +27%, 2-hop +6%, disjoint ranges; entered during the 0.60 stack, not the perf PRs #2102

Description

@lmeyerov

What

Bulk traversal via g.hop() regressed measurably between e951e9a2c (2026-09-04) and current master. Found while re-measuring the GraphFrames ladder, whose staleness had been hiding it — no board in CI or the release suite exercises this shape.

LiveJournal (34.7M edges), polars, 50 seeds, warmups=2 iters=5, via run_streaming_rung_dgx.sh (perf lock, load gate, safe_run, breach classifier). 3 reps per arm, interleaved, one session, one container. Result sizes identical in all 12 runs.

task e951e9a2c f283a305e (shipped) ranges delta
filter 24.2, 25.8, 24.5 24.4, 24.7, 24.9 overlap flat
hop1 351.8, 259.5, 333.4 403.9, 423.2, 426.1 DISJOINT (351.8 < 403.9) +26.9%
hop2 7211.0, 6895.5, 6915.9 7350.1, 7366.4, 7347.5 DISJOINT (7211.0 < 7347.5) +6.3%

filter is the built-in control: same frames, same container, same I/O, no traversal — and it is flat. Both traversal tasks regressed and only the traversal tasks regressed.

NOT caused by the 0.60 performance PRs

Bisect round 1, f44707bb6 (the tree immediately before #2084): 445.0, 444.0, 430.3 — already in the regressed band.

So it entered in the 137 compute commits between e951e9a2c and f44707bb6, i.e. before #2084/#2086/#2087/#2088/#2090. Those five are exonerated.

Scope: q1-q9 and SNB are unaffected

q1-q9 at 100k, same two trees, 2 reps, rows identical:

q OLD NEW q OLD NEW
q1 23.30 22.38 q6 8.11 7.40
q2a 23.33 23.16 q7 5.25 5.28
q3 8.86 8.68 q8 13.10 12.98
q4 7.16 7.04 q9 36.85 36.27
q5 6.75 6.68

No regression on any published cell; NEW is slightly faster on q1 and q6. The defect is in the bulk-hop path (g.hop() from many seeds over tens of millions of edges), not in the two-star/chain fast paths the release boards use.

Lead for whoever picks this up

The new code is more CONSISTENT and slower. hop2 spread is 0.3% (7347.5-7366.4) on the new tree versus 4.6% (6895.5-7211.0) on the old. That is the signature of work that is now always paid rather than sometimes skipped — look for a guard, cache, or fast-path admission that stopped firing, not for newly added computation.

Bisect is continuing; midpoint under test is b74d2f530 "perf(gfql): avoid repeated Polars schema reads".

Reproduce

bash benchmarks/graphframes_streaming/run_streaming_rung_dgx.sh \
  <out> <bench_dir> <pygraphistry_tree> lj polars \
  /home/lmeyerov/data/snap/com-lj.ungraph.txt.gz.parquet 34681189 --skip-tasks pagerank

Compare hop1 medians across trees; allow load to settle below 1.2 between runs or the runner's own gate will (correctly) refuse.

Also worth fixing alongside

The ladder's provenance pins only the image tag (graphistry/test-rapids-official:26.02-gfql-polars), not a sha256 digest — unlike the GraphBench board, which pins sha256:74544b1b.... Published ladder numbers do not reproduce today even on identical code (219.7 → ~340, +55%), which a rebuilt mutable tag would explain. Pin the digest so ladder numbers are reproducible.

🤖 Generated with Claude Code

https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp

Activity

  1. lmeyerov commented on Sep 20, 2026

    @lmeyerov
    ContributorAuthor

    Correction — two claims in the original report were wrong

    1. "Published ladder numbers do not reproduce (219.7 → ~340, +55%)" — WRONG.

    They reproduce to within 0.2%. The published lane ran with cpu_streaming: true; my runs had it false. I compared two execution modes and called the gap environment drift. Re-measured with the published lane's flag:

    run hop1
    e951e9a2c with --cpu-streaming 220.2
    published 2026-09-04 219.7

    The mutable-image-tag hypothesis is also dead: graphistry/test-rapids-official:26.02-gfql-polars still resolves to 74544b1b30fa, the same digest the GraphBench board pins. Nothing drifted.

    2. "The stale-waived ladder was hiding a live regression" — WRONG.

    In streaming mode there is no regression: e951e9a2c 220.2 vs fixed tree 226.6, a 2.9% gap inside noise. The regression lives only in the default in-memory path. The ladder's published cells are cpu_streaming: true, so re-measuring the ladder would never have caught this. I found it because I ran the ladder's workload in a mode the ladder does not publish.

    What is unchanged

    The regression itself, and everything used to establish it. Every OLD-vs-NEW comparison used the same flags on both arms, one box, one session, identical result sizes:

    • hop1 +26.9%, hop2 +6.3%, ranges disjoint
    • bisect to a59990981: parent 351.1/328.0/348.9 clean, commit 421.8/438.9/468.3 regressed, disjoint
    • fix returns hop1 to 332.8/269.1/280.9, inside the clean band

    Those were never compared against the published cell.

    What this strengthens

    The systemic point is now sharper, not weaker: no board watches the affected path at all — not q1-q9, not SNB, not CI, and not even the ladder in its published mode. A 27% in-memory bulk-hop regression was invisible to every gate that exists.

    The floors added in graphistry/pyg-bench#278 are measured in the affected mode, so they guard the right thing; they are being updated to record cpu_streaming explicitly so nobody compares them against a streaming lane and is misled the way I was.

    🤖 Generated with Claude Code

    https://claude.ai/code/session_012Me1E7ZdDuGqJGu3mMEzhp

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions