Repository navigation
ci: test-gfql-core (3.14) with coverage takes 9+ min and gates downstream jobs #2042
Description
Activity
Same pattern on
test-polars (3.12): the coverage-profiled variant takes 9+ min, 2x+ the other Python versions of the same suite (owner observation, 2026-09-05). Both lanes point at the same fix: keep coverage off the critical path (separate non-gating job, or nightly/master-only), and keep xdist on the gating runs.Re-measured on 2026-10-03 (PR #2121's green run at 11edb3c, full matrix):
test-gfql-core (3.14)now takes 5.2 min — the 9+ min complaint no longer holds; the coverage-audited 3.12 lane is the 10.3-min one.- What actually gates:
test-polars (3.12)and everygfql-routes-offcell listtest-gfql-coreinneeds, so they start only when it finishes (07:12→07:22 gap, timestamps in the PR).gfql-routes-off (all-off)is the slowest lane (19 min) and ends the run at 07:41; wall-clock 30.7 min. - Fix: ci: let test-polars and gfql-routes-off start without waiting on test-gfql-core (#2042) #2126 drops that edge from the two lanes (they only consume the
lockfilesartifact);changed-line-coveragekeeps it for the coverage artifact. Expected wall-clock ≈ 21 min. The routes-off cells themselves (13–19 min each) are the remaining floor and would need sharding or xdist to go lower — out of scope for this change.
- added a commit that references this issue
on Oct 4, 2026 Half of this is done on master, half is not, so I am leaving it open rather than closing it.
Done (#2126, merge 28d8082):
test-polarsandgfql-routes-offno longer wait ontest-gfql-core. Measured on the run that prompted it,test-polars (3.12)now starts at 2.9 min instead of 11.8.gfql-routes-offkeeps a lint gate viapython-lint-typesdirectly, so the 11 cells cannot run on a red lint.Not done — the issue's actual complaint. The coverage-audited cells still cost roughly twice the plain ones, and after the decoupling they are the critical path. Measured on run 37151747760:
lane minutes gfql-routes-off, 11 cells 9.3 to 19.2 each, ~175 runner-min total test-gfql-core (3.12) 11.2 test-polars (3.12) 10.8 I audited that lane while landing the 2026-10 queue, and the measured levers are:
- Parallelism. Every other test lane in
ci.ymlruns-n autoandpytest-xdistis already in the test extra, butbin/test-routes-off.shruns single-process. Locally, same tree and routes declined: 725 s at 1 process, 241 s at 4 (3.0x), 157 s at 8 (4.6x). The 1-process counts match the CI all-off cell exactly (13,419 passed / 3,538 skipped / 405 xfailed), so it is the same work. - Prerequisite, and it is a real defect. The suite is not parallel-clean: 171 failures at
-n 4, 148 of them intest_grouped_aggregate_fused_polars.py, 19 intest_grouped_aggregate_lowcard_count.py, 4 intest_fast_path_engagement.py, every one anassert [] == [True]spy that never fired.--dist loadfilemakes it worse (211). Root cause is ordering, not parallelism: that file alone under routes-off fails 188 of 234, but passes inside the full single-process replay, so its engagement assertions depend on state other tests leave in the process. An armed probe confirmed the route switch itself never leaks (1,501 checks, zero regressions), so the dependency is on something else the earlier tests warm. - Scope. Instrumenting every declinable gate shows 112 of 133 timed files actually consult one; the other 21 are 149 s of the 610 s replay set (24% of every cell) and cannot be affected by declining a route. The two largest are
test_differential_grammar_parity.py(78.3 s) andtest_seed_rediscovery_2023.py(34.1 s). Dropping those from the replay is measured, not argued, and loses no signal. - Wrong repo.
test_hop_scaling_pin.py,test_gfql_latency_contract.pyand one block intest_seeded_typed_hop_fastpath.pyare the only wall-clock assertions intests/compute; by the repo's own rule those belong in pyg-bench.
So the remaining work is: make those three files self-sufficient, then turn on
-n autofor the replay, scope it to the files that can engage, and move the timing pins out. That is worth about 3x on the longest lane in the pipeline.- Parallelism. Every other test lane in
Observed on the 2026-09-05 PR runs: the
test-gfql-core (3.14)lane (the coverage-audited one) takes 9+ minutes while the other Python versions of the same suite finish in about 5, and several downstream jobs wait on it.Candidates: run coverage on one version only and the plain suite elsewhere (already the case?) but with
pytest-xdist(-n autois used intest-gfql-coreper the repo notes) — check whether the coverage lane lost xdist or runs--covwith branch coverage on 3.14 only; split the coverage audit into its own job that is not in the critical path of the merge gate; or shard the suite by directory.Filed from the landing session; not part of the 0.60 stacks.