Skip to content

ci: test-gfql-core (3.14) with coverage takes 9+ min and gates downstream jobs #2042

Description

@lmeyerov

Observed on the 2026-09-05 PR runs: the test-gfql-core (3.14) lane (the coverage-audited one) takes 9+ minutes while the other Python versions of the same suite finish in about 5, and several downstream jobs wait on it.

Candidates: run coverage on one version only and the plain suite elsewhere (already the case?) but with pytest-xdist (-n auto is used in test-gfql-core per the repo notes) — check whether the coverage lane lost xdist or runs --cov with branch coverage on 3.14 only; split the coverage audit into its own job that is not in the critical path of the merge gate; or shard the suite by directory.

Filed from the landing session; not part of the 0.60 stacks.

Activity

  1. lmeyerov commented on Sep 5, 2026

    @lmeyerov
    ContributorAuthor

    Same pattern on test-polars (3.12): the coverage-profiled variant takes 9+ min, 2x+ the other Python versions of the same suite (owner observation, 2026-09-05). Both lanes point at the same fix: keep coverage off the critical path (separate non-gating job, or nightly/master-only), and keep xdist on the gating runs.

  2. lmeyerov commented on Oct 3, 2026

    @lmeyerov
    ContributorAuthor

    Re-measured on 2026-10-03 (PR #2121's green run at 11edb3c, full matrix):

    • test-gfql-core (3.14) now takes 5.2 min — the 9+ min complaint no longer holds; the coverage-audited 3.12 lane is the 10.3-min one.
    • What actually gates: test-polars (3.12) and every gfql-routes-off cell list test-gfql-core in needs, so they start only when it finishes (07:12→07:22 gap, timestamps in the PR). gfql-routes-off (all-off) is the slowest lane (19 min) and ends the run at 07:41; wall-clock 30.7 min.
    • Fix: ci: let test-polars and gfql-routes-off start without waiting on test-gfql-core (#2042) #2126 drops that edge from the two lanes (they only consume the lockfiles artifact); changed-line-coverage keeps it for the coverage artifact. Expected wall-clock ≈ 21 min. The routes-off cells themselves (13–19 min each) are the remaining floor and would need sharding or xdist to go lower — out of scope for this change.
  3. added a commit that references this issue on Oct 4, 2026
  4. lmeyerov commented on Oct 4, 2026

    @lmeyerov
    ContributorAuthor

    Half of this is done on master, half is not, so I am leaving it open rather than closing it.

    Done (#2126, merge 28d8082): test-polars and gfql-routes-off no longer wait on test-gfql-core. Measured on the run that prompted it, test-polars (3.12) now starts at 2.9 min instead of 11.8. gfql-routes-off keeps a lint gate via python-lint-types directly, so the 11 cells cannot run on a red lint.

    Not done — the issue's actual complaint. The coverage-audited cells still cost roughly twice the plain ones, and after the decoupling they are the critical path. Measured on run 37151747760:

    lane minutes
    gfql-routes-off, 11 cells 9.3 to 19.2 each, ~175 runner-min total
    test-gfql-core (3.12) 11.2
    test-polars (3.12) 10.8

    I audited that lane while landing the 2026-10 queue, and the measured levers are:

    • Parallelism. Every other test lane in ci.yml runs -n auto and pytest-xdist is already in the test extra, but bin/test-routes-off.sh runs single-process. Locally, same tree and routes declined: 725 s at 1 process, 241 s at 4 (3.0x), 157 s at 8 (4.6x). The 1-process counts match the CI all-off cell exactly (13,419 passed / 3,538 skipped / 405 xfailed), so it is the same work.
    • Prerequisite, and it is a real defect. The suite is not parallel-clean: 171 failures at -n 4, 148 of them in test_grouped_aggregate_fused_polars.py, 19 in test_grouped_aggregate_lowcard_count.py, 4 in test_fast_path_engagement.py, every one an assert [] == [True] spy that never fired. --dist loadfile makes it worse (211). Root cause is ordering, not parallelism: that file alone under routes-off fails 188 of 234, but passes inside the full single-process replay, so its engagement assertions depend on state other tests leave in the process. An armed probe confirmed the route switch itself never leaks (1,501 checks, zero regressions), so the dependency is on something else the earlier tests warm.
    • Scope. Instrumenting every declinable gate shows 112 of 133 timed files actually consult one; the other 21 are 149 s of the 610 s replay set (24% of every cell) and cannot be affected by declining a route. The two largest are test_differential_grammar_parity.py (78.3 s) and test_seed_rediscovery_2023.py (34.1 s). Dropping those from the replay is measured, not argued, and loses no signal.
    • Wrong repo. test_hop_scaling_pin.py, test_gfql_latency_contract.py and one block in test_seeded_typed_hop_fastpath.py are the only wall-clock assertions in tests/compute; by the repo's own rule those belong in pyg-bench.

    So the remaining work is: make those three files self-sufficient, then turn on -n auto for the replay, scope it to the files that can engage, and move the timing pins out. That is worth about 3x on the longest lane in the pipeline.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions