Repository navigation
perf(ci): pin crane deps in a Nix profile to persist across runs - #1279
Conversation
The check, test, and native-build jobs recompiled the entire crate dependency set (~3 min on the arm runner) on nearly every run, even though the deps derivation is deterministic. The Blacksmith sticky disk persists the store but trims it to GC roots on commit, and crane's cargoArtifacts is an unrooted intermediate (referenced only at build time), so it was dropped and rebuilt. Pin the deps into a Nix profile after they are built. Profiles are the one root that survives the trim (confirmed by inspecting a restored disk and a CI A/B: `nix flake check` dropped from ~184s to ~38s with zero dependency recompilation on the warm run). - new `pin-nix-deps` composite action: `nix-env --set` the cargoArtifacts path into /nix/var/nix/profiles/ccusage-deps. - check + test jobs pin the glibc deps (`.#ccusage.cargoArtifacts`). - the Linux native build pins the musl deps (`.#ccusage-static.cargoArtifacts`); re-expose them via passthru. Each job has its own sticky disk, so each pins its own deps.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
📝 WalkthroughWalkthroughThis PR exposes static package cargoArtifacts, adds a composite ChangesNix dependency pinning for CI caching
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
ccusage-guide | cd795c8 | Commit Preview URL Branch Preview URL |
Jun 11 2026, 07:07 PM |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/ci.yaml:
- Around line 33-35: The pinning steps that use the
./.github/actions/pin-nix-deps action (the steps with `attr:
.#ccusage.cargoArtifacts` and the similar step around lines 68-70) currently run
only on success and are skipped on failures; update those job steps to run
unconditionally by adding `if: always()` (or the equivalent unconditional
condition) so pinning occurs even when prior steps fail, ensuring built Nix
artifacts are rooted and cached for subsequent CI runs.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: eec8b55c-9858-4bb7-8b84-59dba540e8bb
📒 Files selected for processing (4)
.github/actions/build-linux-native-package/action.yaml.github/actions/pin-nix-deps/action.yaml.github/workflows/ci.yamlnix/static-package.nix
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — adds a pin-nix-deps composite action that roots crane cargoArtifacts in a Nix profile so the Blacksmith sticky disk's GC trim preserves them across CI runs, and wires it into the check, test, and build-linux-native-package jobs. The root cause is well-investigated (confirmed by inspecting sticky disk state + A/B validation showing 184s→38s on nix flake check), and the fix is minimal and correct.
- Add
.github/actions/pin-nix-deps/action.yaml— new composite action that runsnix build --no-linkto verify the path is realized, thennix-env --setinto/nix/var/nix/profiles/ccusage-depsto create a GC root that survives the sticky disk trim. - Wire pinning into
checkandtestjobs — pin.#ccusage.cargoArtifactsafter their respective build steps (nix flake check/nix build .#ccusage-tests). - Wire pinning into
build-linux-native-package— pin.#ccusage-static.cargoArtifactsafter the static binary build. - Expose static cargoArtifacts via passthru in
nix/static-package.nix— mirrors the existing pattern inpackage.nix:62-68.
Big Pickle (free via Pullfrog for OSS) | 𝕏
DiagnosisThe Evidence
ConclusionNo fix is needed for the PR. The Windows failure is an environmental/infrastructure issue that has been seen as a pre-existing flake (see learnings L32 — pre-existing Nix packaging/platform issues). The PR's Linux Nix profile pinning is correct and the Windows failure is unrelated to the change. Task list (3/5 completed)
|
There was a problem hiding this comment.
2 issues found across 4 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
If `nix flake check` or `nix build .#ccusage-tests` fails, the default
`success()` condition skipped the pin-nix-deps step, leaving any deps
built during the failing run unrooted. They were then trimmed on commit
and the next run had to recompile them, defeating the warm-cache goal.
Run the pin step with `if: ${{ !cancelled() }}` so failed runs still
root the deps for the next attempt, while skipping on cancel (where
there is nothing useful to pin).
Co-authored-by: Codesmith <[email protected]>
ccusage
@ccusage/ccusage-darwin-arm64
@ccusage/ccusage-linux-arm64
@ccusage/ccusage-linux-x64
@ccusage/ccusage-win32-x64
commit: |
The Linux native build's `pin-nix-deps` step invokes `nix build` and `nix eval` to root the static-build deps. On PR runs from forks the preceding build step uses `NIX_CONFIG="access-tokens ="` so Nix cannot use the GitHub token for fetches; the pin step inherited the default config and bypassed that hardening. Forward the same gate via an `env:` on the composite step so the pin's Nix invocations honor the caller's `nix-github-token` input. Co-authored-by: Codesmith <[email protected]>
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — two incremental commits adding PR token hardening to the linux native pin step and making the check/test pin steps resilient to job failure (but skipped on cancellation).
- Token-hardened the linux-native pin step — mirrors the existing
NIX_CONFIGgate from the build step so pin runs correctly onpull_requestevents. - Pin deps even on build/test failure — the
checkandtestjobs now useif: ${{ !cancelled() }}, replacing the implicitsuccess()default so deps are always rooted after building regardless of outcome (but still skipped on cancellation).
Big Pickle (free via Pullfrog for OSS) | 𝕏
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |

Summary
check,test, and the native build jobs recompiled the entire cratedependency set (~3 min on the arm runner) on nearly every run, even though the
deps derivation is deterministic.
Root cause (confirmed by inspection + A/B)
The Blacksmith sticky disk persists the Nix store but trims it to GC roots on
commit. crane's
cargoArtifacts(deps-only) is an unrooted intermediate —referenced only at build time — so it was dropped on every commit and rebuilt.
I confirmed this by attaching the build-native sticky disk and finding the deps
output path missing on restore while ~7GB of other paths persisted, and the
only surviving GC root was a Nix profile. An A/B on the
checkjob then showednix flake checkdrop from ~184s → ~38s with zero dependencyrecompilation once the deps were pinned in a profile.
(Notes from the investigation: the GHA action-cache approach used on macOS does
not help here — even on a cache hit the deps rebuilt, because
gc-max-store-sizetrims unrooted paths before saving, the same failure mode. A binary cache would
also fix it but needs extra infra; this keeps the fast Blacksmith disk.)
Change
pin-nix-depscomposite action:nix-env --setthe cargoArtifacts pathinto
/nix/var/nix/profiles/ccusage-deps(profiles survive the trim).check+testpin the glibc deps (.#ccusage.cargoArtifacts)..#ccusage-static.cargoArtifacts);re-exposed via
passthru.Each job has its own sticky disk, so each pins its own deps.
Expected effect
Warm runs:
check184s→38s, Linux native build deps reused (no ~190s rebuild),testRust deps reused. First run after merge pins; the run after that benefits.Summary by CodeRabbit
Summary by cubic
Pin Rust dependency artifacts from
craneinto a Nix profile in CI so the Blacksmith sticky disk keeps them across runs, removing repeated recompiles and speedingcheck,test, and native builds. Native builds are now serialized to protect the pinned cache across overlapping runs; warmnix flake checkruns dropped from ~184s to ~38s and the native musl build avoids ~190s of dep rebuilds../.github/actions/pin-nix-depsto rootcargoArtifactsin/nix/var/nix/profiles/ccusage-depsvianix-env --set; first run pins, later runs reuse..#ccusage.cargoArtifactsincheckandtest(runs on failure; skips on cancel)..#ccusage-static.cargoArtifactsin the Linux native build; exposed viapassthruon.#ccusage-staticand hardened on PRs withNIX_CONFIG="access-tokens ="whennix-github-tokenisfalse.build-native-packagesper matrix entry withconcurrencyso overlapping runs don’t overwrite each other’s pinned deps.Written for commit cd795c8. Summary will update on new commits.
Need help on this PR? Tag
/codesmithwith what you need. Autofix is enabled.