Skip to content

Add benchmarks for coverage.py performance paths - #2238

Open
KRRT7 wants to merge 12 commits into
coveragepy:mainfrom
KRRT7:asv-bench
Open

KRRT7 wants to merge 12 commits into
coveragepy:mainfrom
KRRT7:asv-bench

Conversation

@KRRT7

@KRRT7 KRRT7 commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add an opt-in pytest-benchmark suite for coverage.py's tracing, data handling, and reporting costs, alongside a pinned Jinja2 test-suite workload. Each case validates the work performed outside the timed region.

The suite covers:

  • Line and branch tracing with Python, C, and sys.monitoring cores across loops, calls, imports, generators, and async execution.
  • Dynamic contexts, dense and sparse database writes, parallel-data combination, path remapping, unused-source discovery, and multiprocessing.
  • Analysis and text, HTML, XML, JSON, and LCOV reporting over mixed coverage data, including context filtering and incremental HTML updates.
  • Tokenization of frozen source snapshots, with hashes and workload versions for comparable results.
  • Jinja2 3.1.6's 909 tests, uninstrumented and under each coverage core, plus separate reporting and fresh-process peak-memory measurements.

Assertions check core selection, context membership, combined data, worker coverage, regenerated HTML pages, and matching real-project test identities.

Usage

make bench
make bench ARGS='-k html --benchmark-autosave'

# Validate component workloads without collecting timings:
python3 -m pytest -n0 --benchmarks --bench-smoke \
    -m 'benchmark and not real_project' tests/benchmarks

# Prepare and run the real-project workload with Python 3.14:
make bench-prepare
make bench-real

Benchmarks require explicit opt-in and serial execution. Real-project preparation downloads verified inputs and installs separately pinned dependencies before measurement. Setup-dependent cases use ten measured rounds, or five for slow cases, adjustable with --bench-rounds.

See the benchmark README for workload sizes, timing boundaries, and comparison guidance.

CI and validation

Relevant PRs run component timings on Linux Python 3.14 and smoke checks on Linux Python 3.10/3.15 and macOS Python 3.14. A weekly/manual job runs the real-project and slow workloads. JSON artifacts and summaries report timing variability, sample counts, workload metadata, and peak RSS where available. Results are informational, without a regression threshold.

Validation completed:

  • Local timing runs: 64 component cases and 16 extended cases passed.
  • Component and real-project smoke runs passed, including verification of all 909 Jinja2 test identities.
  • Targeted regression and testing/core checks: 59 passed, 1 skipped. Workload regressions also passed under metacoverage with both C and Python tracers.
  • Local type, lint, formatting, workflow, manifest, and package-build checks passed.
  • PR component timings and all three compatibility smoke jobs passed.

Comparisons require matching workload versions and machine/Python configurations. Jinja2 provides one real application workload; additional projects would broaden the results.

@nedbat

nedbat commented Jul 26, 2026

Copy link
Copy Markdown
Member

Thanks. I tried this out and got some failures:

% asv run
· Creating environments
· Discovering benchmarks
· Running 21 total benchmarks (1 commits * 1 environments * 21 benchmarks)
[ 0.00%] · For coverage.py commit d37859cd <main>:
[ 0.00%] ·· Benchmarking uv-py3.14
[ 2.38%] ··· Running (measurement.TimeCombine.time_combine_many_files_line_only--)................
[40.48%] ··· Running (reporting.TimeReport.time_json_report--).....
[52.38%] ··· measurement.TimeCombine.time_combine_many_files_line_only                                                                                  178±0.2ms
[54.76%] ··· measurement.TimeCombine.time_combine_many_files_with_contexts                                                                                178±1ms
[57.14%] ··· measurement.TimeDynamicContexts.time_collect_test_function_contexts                                                                       53.5±0.4ms
[59.52%] ··· measurement.TimeMeasurement.time_collect                                                                                                          ok
[59.52%] ··· ========= ============ ============
             --                  branch
             --------- -------------------------
                core      False         True
             ========= ============ ============
              pytrace   162±0.4ms    184±0.9ms
               ctrace   57.3±0.4ms   73.3±0.5ms
               sysmon   43.5±0.7ms   51.2±0.6ms
             ========= ============ ============

[61.90%] ··· measurement.TimeMultiprocessing.time_multiprocessing_run                                                                                     225±2ms
[64.29%] ··· measurement.TimeMultiprocessing.time_multiprocessing_run_and_combine                                                                         303±5ms
[66.67%] ··· measurement.TimePersistence.time_add_arcs                                                                                                 4.72±0.1ms
[69.05%] ··· measurement.TimePersistence.time_add_lines                                                                                               1.55±0.02ms
[71.43%] ··· measurement.TimeSourceTree.time_collect_with_large_unused_source_tree                                                                     51.9±0.5ms
[73.81%] ··· reporting.TimeAnalysis.time_analysis2_all_files                                                                                               failed
[76.19%] ··· reporting.TimeHtmlContexts.time_html_report_with_contexts                                                                                     failed
[78.57%] ··· reporting.TimeHtmlContexts.time_html_report_with_filtered_contexts                                                                            failed
[80.95%] ··· reporting.TimeHtmlIncremental.time_html_report_single_source_change                                                                           failed
[83.33%] ··· reporting.TimeHtmlIncremental.time_html_report_unchanged                                                                                      failed
[85.71%] ··· reporting.TimeLargeHtml.time_html_report_large_module                                                                                         failed
[88.10%] ··· reporting.TimeReport.time_html_report                                                                                                         failed
[90.48%] ··· reporting.TimeReport.time_json_report                                                                                                         failed
[92.86%] ··· reporting.TimeReport.time_lcov_report                                                                                                         failed
[95.24%] ··· reporting.TimeReport.time_report_text                                                                                                         failed
[97.62%] ··· reporting.TimeReport.time_xml_report                                                                                                          failed
[100.00%] ··· reporting.TimeReportReuse.time_report_then_html_same_process                                                                                  failed

Using --verbose, I see this exception:

               File "/private/tmp/covasv/coveragepy/benchmarks/reporting.py", line 82, in time_report_then_html_same_process
                 self.cov._clear_analysis_caches()
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
             AttributeError: 'Coverage' object has no attribute '_clear_analysis_caches'

Is this benchmark layered on top of one of the other changes?

@KRRT7

KRRT7 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor Author

Is this benchmark layered on top of one of the other changes?

it was.

@KRRT7 KRRT7 changed the title Add ASV benchmarks for coverage.py performance paths Add benchmarks for coverage.py performance paths Aug 27, 2026
ASV brought its own environment manager, runner and result store,
duplicating infrastructure we already have in pytest and tox. The
benchmark content itself is framework-agnostic, so keep all of it and
drive it with pytest-benchmark instead.

Move benchmarks/ to tests/benchmarks/, turn the Time* classes into
test functions, setup/teardown into fixtures, and params into
parametrize. Two things needed fixing rather than porting:

- ASV re-runs setup() before every repeat; pytest-benchmark does not.
  Reporting into a populated htmlcov takes the incremental path, and
  writing into an existing SQLite database is an update rather than a
  create, so rounds after the first measured a different code path.
  Everything non-idempotent now uses benchmark.pedantic(setup=...).

- core_available("sysmon") was wrong: core.py refuses sysmon when
  branch is on and the version can't measure branches, so the "sysmon
  arcs" benchmark silently measured ctrace, and the no-sysmon warning
  that would have said so is suppressed by tests/conftest.py. Skip
  that combination, and assert the core actually in use. Select the
  core with the run:core option instead of COVERAGE_CORE, which igor
  sets process-wide.

The benchmarks are kept out of the test suite by two layers, because
pytest-benchmark's disabled modes still execute the function:
--benchmarks gates collection, and -m "not benchmark" deselects.

Add a Benchmarks workflow that records results as an artifact. It
deliberately doesn't fail on a regression: shared runners swing wider
than most real regressions do.
Two CI checks caught things the local run had missed.

mypy: pytest-benchmark 5.3.0 ships py.typed but leaves BenchmarkFixture
unannotated, so --strict rejects every benchmark.pedantic() call. (The
local mypy env was stale and had neither the new pin nor mypy 2.3, so
the annotation quietly fell back to Any.) Describe the two entry points
we use as a Protocol instead: the call sites stay checked, and there are
no type: ignore comments to go stale under warn_unused_ignores when
upstream annotates.

zizmor: the workflow_dispatch input was interpolated straight into the
run block, and the pseudo-ternary it used was unsound. Since '' is
falsy, `inputs.slow && '' || ' and not slow'` always chose the right
side, so asking for the slow benchmarks would never have included them.
Pass the input through the environment and branch in the shell.
An external review of the conversion turned up two things.

pytest doesn't consult pytest_ignore_collect for a path given as an
initial command-line argument, so `pytest -m benchmark
tests/benchmarks/test_thing.py` collected and ran benchmarks with no
--benchmarks flag. igor's no-path invocation was never at risk, but the
gate is documented as mandatory, and a gate with a hole is worse than
none because people rely on it. Deselect in pytest_collection_modifyitems
too, which sees items however the path was named.

Also hoist the subprocess environment out of the multiprocessing
benchmarks' timed region. Copying and scrubbing os.environ is harness
work: measured at 257us against a ~450ms subprocess even with a padded
288-variable environment, so it was well inside the noise, but a
benchmark shouldn't be timing its own harness.
@KRRT7
KRRT7 marked this pull request as ready for review September 14, 2026 08:35
@nedbat

nedbat commented Sep 14, 2026

Copy link
Copy Markdown
Member

Let me know when you are ready to discuss the future of this. I would prefer it was in its own repo, but there might be a reason why that wouldn't work.

@KRRT7

KRRT7 commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

Hi, i was traveling for the past few days but now i'm back, the reason they're in repo is because of my experience as a perf engineer.

When I optimize a path, the workload often needs to change in the same PR: a new scenario to expose the hot spot, a validation tweak because the measured work moved, or a workload version bump. In-repo that is one commit and one review. In a separate repo it becomes two PRs that have to be kept in step.

additionally, With the suite in-tree, git bisect run make bench ARGS='-k html' walks history and every checkout has the benchmarks that match its own internals. An external suite has to guess which coverage revisions it still runs against.

finally, The suite reuses the test infrastructure. It uses tests/conftest.py, testenv.py, and the same core selection the ordinary tests use, and the smoke mode asserts the workload did the intended work (cores, contexts, combined data, regenerated HTML pages). it depends on test helpers that a separate repo would have to copy and keep in sync.

@nedbat

nedbat commented Sep 26, 2026

Copy link
Copy Markdown
Member

I appreciate those factors, but I just don't feel comfortable adding this much code to this repo for performance testing. Let's work on a way to have it live it its own.

@KRRT7

KRRT7 commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

@nedbat putting it all in https://github.com/coveragepy/benchmark is an option but since it's so home-made, so to speak, I'm not sure I feel ok replacing it, maybe a new repo? and then we can install codspeed there?

@nedbat

nedbat commented Sep 27, 2026

Copy link
Copy Markdown
Member

Yes, a new repo makes sense.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants