Repository navigation
perf(ci): build Rust coverage as a Nix derivation - #1266
Conversation
The coverage step ran `cargo llvm-cov` inside the dev shell, compiling the whole workspace into `rust/target` on every CI run. Only `/nix` is kept on the per-job Blacksmith sticky disk, so this target dir was never cached and the coverage build was the one Rust compile that stayed cold each run. Move it into a crane `cargoLlvmCov` package (`ccusage-coverage`) that reuses the shared `cargoArtifacts` closure, so the deps stay warm on the sticky disk and an unchanged source tree returns the cached cobertura report instantly. CI now `nix build`s the package and copies its `$out` report. The build sandbox has no $HOME, so a preBuild seeds the empty Claude data directories the workspace tests resolve, matching what the dev-shell test job created.
📝 WalkthroughWalkthroughThis PR adds a Nix module that builds Rust coverage as a package, imports it into the flake, and changes the CI test job to build ChangesRust Coverage as Nix Package
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
ccusage-guide | 16f8a74 | Commit Preview URL Branch Preview URL |
Jun 11 2026, 02:36 PM |
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — moves Rust coverage from a dev-shell cargo llvm-cov invocation to a crane cargoLlvmCov Nix derivation, so it shares the warm cargoArtifacts cache instead of recompiling the whole workspace cold on every CI run.
- New
nix/coverage.nix— definesccusage-coverageviacraneLib.cargoLlvmCov, reusingcargoArtifactsandcommonArgsfrom the existingccusagepackage's passthru. Usesnix-filterto exclude noise (node_modules,target,dist,coverage), setssourceRoot = "source/rust"to scope cargo to the workspace, and seeds a writable$HOMEinpreBuildfor test sandbox requirements. flake.niximport — adds./nix/coverage.nixto the flake-parts import list.- CI workflow swap — replaces
env -u CFLAGS -u CXXFLAGS nix develop --command cargo llvm-covwithnix build .#ccusage-coverage --print-build-logs --out-link "$RUNNER_TEMP/coverage-report"followed by acpto stage the cobertura XML for upload.
Big Pickle (free via Pullfrog for OSS) | 𝕏
ccusage
@ccusage/ccusage-darwin-arm64
@ccusage/ccusage-linux-arm64
@ccusage/ccusage-linux-x64
@ccusage/ccusage-win32-x64
commit: |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
The crane coverage build is hermetic and has no system zoneinfo database, so jiff could not resolve named time zones (e.g. Asia/Tokyo) and fell back to UTC. That shifted two timezone-dependent tests by a day and failed the build, even though they pass in the dev shell where the runner's zoneinfo is present. Set TZDIR to the nixpkgs tzdata so jiff resolves real offsets; referencing the store path also pulls it into the sandbox as a build input.
DiagnosisCheck suite Failing tests:
FixThe PR author already committed the fix in All 253 tests pass with Task list (5/5 completed)
|
Code Coverage OverviewLanguages: Rust Rust / code-coverage/cargo-llvm-covThe overall coverage in the Show a code coverage summary of the most impacted files.
Updated |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |

Summary
Moves the Rust coverage report from a dev-shell
cargo llvm-covinvocationto a crane
cargoLlvmCovNix package, so it shares the warmcargoArtifactscache instead of recompiling the whole workspace cold on every CI run.
What changed
nix/coverage.nix(new): accusage-coveragepackage built with crane'scargoLlvmCov, reusing the sharedcargoArtifactsclosure and emitting thecobertura report directly at
$out. ApreBuildseeds the empty Claude datadirectories under a writable
$HOME, since the build sandbox has none.flake.nix: import the new module..github/workflows/ci.yaml: the coverage step nownix builds thepackage and copies its report, replacing the
nix develop --command cargo llvm-covcall.Why
cargo llvm-covcompiled the entire workspace intorust/target, but only/nixis persisted on the per-job Blacksmith sticky disk, so that target dirwas never cached. It was the one Rust compile that stayed cold on every run.
As a Nix derivation the dependencies stay warm on the sticky disk, and an
unchanged source tree returns the cached report instantly — bringing coverage
onto the same crane caching path as the rest of the Rust builds.
Testing
nix build .#ccusage-coveragelocally (aarch64-darwin): all workspace testspass in the sandbox and a valid cobertura XML is produced (line-rate 77.4%).
Need help on this PR? Tag
/codesmithwith what you need. Autofix is disabled.Summary by cubic
Build Rust coverage as a Nix derivation using crane to reuse the shared
cargoArtifactscache, eliminating cold rebuilds and speeding up CI. CI now buildsccusage-coverageand uploads its Cobertura report.Refactors
nix/coverage.nixdefiningccusage-coverageviacrane.cargoLlvmCov; reusescargoArtifacts, writes Cobertura to$out, and seeds a writable$HOMEfor tests.flake.nix.cargo llvm-covwithnix build .#ccusage-coverageand copy the report.Bug Fixes
TZDIRfrom nixpkgstzdatain the coverage build so timezone-dependent tests resolve real offsets in the hermetic sandbox.Written for commit 16f8a74. Summary will update on new commits.
Summary by CodeRabbit