Repository navigation
Conversation
|
Thanks. I tried this out and got some failures: Using --verbose, I see this exception: File "/private/tmp/covasv/coveragepy/benchmarks/reporting.py", line 82, in time_report_then_html_same_process
self.cov._clear_analysis_caches()
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AttributeError: 'Coverage' object has no attribute '_clear_analysis_caches'Is this benchmark layered on top of one of the other changes? |
it was. |
ASV brought its own environment manager, runner and result store,
duplicating infrastructure we already have in pytest and tox. The
benchmark content itself is framework-agnostic, so keep all of it and
drive it with pytest-benchmark instead.
Move benchmarks/ to tests/benchmarks/, turn the Time* classes into
test functions, setup/teardown into fixtures, and params into
parametrize. Two things needed fixing rather than porting:
- ASV re-runs setup() before every repeat; pytest-benchmark does not.
Reporting into a populated htmlcov takes the incremental path, and
writing into an existing SQLite database is an update rather than a
create, so rounds after the first measured a different code path.
Everything non-idempotent now uses benchmark.pedantic(setup=...).
- core_available("sysmon") was wrong: core.py refuses sysmon when
branch is on and the version can't measure branches, so the "sysmon
arcs" benchmark silently measured ctrace, and the no-sysmon warning
that would have said so is suppressed by tests/conftest.py. Skip
that combination, and assert the core actually in use. Select the
core with the run:core option instead of COVERAGE_CORE, which igor
sets process-wide.
The benchmarks are kept out of the test suite by two layers, because
pytest-benchmark's disabled modes still execute the function:
--benchmarks gates collection, and -m "not benchmark" deselects.
Add a Benchmarks workflow that records results as an artifact. It
deliberately doesn't fail on a regression: shared runners swing wider
than most real regressions do.
Two CI checks caught things the local run had missed. mypy: pytest-benchmark 5.3.0 ships py.typed but leaves BenchmarkFixture unannotated, so --strict rejects every benchmark.pedantic() call. (The local mypy env was stale and had neither the new pin nor mypy 2.3, so the annotation quietly fell back to Any.) Describe the two entry points we use as a Protocol instead: the call sites stay checked, and there are no type: ignore comments to go stale under warn_unused_ignores when upstream annotates. zizmor: the workflow_dispatch input was interpolated straight into the run block, and the pseudo-ternary it used was unsound. Since '' is falsy, `inputs.slow && '' || ' and not slow'` always chose the right side, so asking for the slow benchmarks would never have included them. Pass the input through the environment and branch in the shell.
An external review of the conversion turned up two things. pytest doesn't consult pytest_ignore_collect for a path given as an initial command-line argument, so `pytest -m benchmark tests/benchmarks/test_thing.py` collected and ran benchmarks with no --benchmarks flag. igor's no-path invocation was never at risk, but the gate is documented as mandatory, and a gate with a hole is worse than none because people rely on it. Deselect in pytest_collection_modifyitems too, which sees items however the path was named. Also hoist the subprocess environment out of the multiprocessing benchmarks' timed region. Copying and scrubbing os.environ is harness work: measured at 257us against a ~450ms subprocess even with a padded 288-variable environment, so it was well inside the noise, but a benchmark shouldn't be timing its own harness.
|
Let me know when you are ready to discuss the future of this. I would prefer it was in its own repo, but there might be a reason why that wouldn't work. |
|
Hi, i was traveling for the past few days but now i'm back, the reason they're in repo is because of my experience as a perf engineer. When I optimize a path, the workload often needs to change in the same PR: a new scenario to expose the hot spot, a validation tweak because the measured work moved, or a workload version bump. In-repo that is one commit and one review. In a separate repo it becomes two PRs that have to be kept in step. additionally, With the suite in-tree, finally, The suite reuses the test infrastructure. It uses |
|
I appreciate those factors, but I just don't feel comfortable adding this much code to this repo for performance testing. Let's work on a way to have it live it its own. |
|
@nedbat putting it all in https://github.com/coveragepy/benchmark is an option but since it's so home-made, so to speak, I'm not sure I feel ok replacing it, maybe a new repo? and then we can install codspeed there? |
|
Yes, a new repo makes sense. |
Summary
Add an opt-in pytest-benchmark suite for coverage.py's tracing, data handling, and reporting costs, alongside a pinned Jinja2 test-suite workload. Each case validates the work performed outside the timed region.
The suite covers:
Assertions check core selection, context membership, combined data, worker coverage, regenerated HTML pages, and matching real-project test identities.
Usage
Benchmarks require explicit opt-in and serial execution. Real-project preparation downloads verified inputs and installs separately pinned dependencies before measurement. Setup-dependent cases use ten measured rounds, or five for slow cases, adjustable with
--bench-rounds.See the benchmark README for workload sizes, timing boundaries, and comparison guidance.
CI and validation
Relevant PRs run component timings on Linux Python 3.14 and smoke checks on Linux Python 3.10/3.15 and macOS Python 3.14. A weekly/manual job runs the real-project and slow workloads. JSON artifacts and summaries report timing variability, sample counts, workload metadata, and peak RSS where available. Results are informational, without a regression threshold.
Validation completed:
Comparisons require matching workload versions and machine/Python configurations. Jinja2 provides one real application workload; additional projects would broaden the results.