Skip to content

fix(pricing): stabilise models.dev snapshot generation - #1304

Merged
ryoppippi merged 7 commits into
mainfrom
codex/fix-models-dev-pricing-determinism
Jun 12, 2026
Merged

ryoppippi merged 7 commits into
mainfrom
codex/fix-models-dev-pricing-determinism

Conversation

@ryoppippi

@ryoppippi ryoppippi commented Jun 12, 2026 •

Copy link
Copy Markdown
Member

Summary

Stabilises the models.dev pricing snapshot generator so duplicate pricing keys no longer depend on upstream provider iteration order.

The generator now sorts provider/model entries and ranks duplicate candidates, preferring Anthropic provider rows and complete cache/context metadata before falling back to deterministic source ids. The committed snapshot was regenerated with the fixed generator.

Testing

  • pnpm exec vitest run nix/models-dev-compact.test.ts
  • direnv exec . cargo test --manifest-path rust/Cargo.toml -p ccusage pricing::tests::
  • direnv exec . just gen-models-dev-pricing
  • direnv exec . just test-vitest
  • pre-push hook via direnv exec . git push -u origin codex/fix-models-dev-pricing-determinism

View with Codesmith
Need help on this PR? Tag /codesmith with what you need. Autofix is enabled.


Summary by cubic

Stabilizes models.dev pricing snapshot generation with bytewise string sorting and deterministic duplicate resolution, so output no longer varies by provider order or locale. Regenerates the snapshot to restore canonical Anthropic cache/context values, including cache_read/write and context limits.

  • Bug Fixes
    • Sort provider and model entries using bytewise comparison; resolve duplicate keys by preferring anthropic, then entries with explicit cache read/write and context, then stable provider/model id tie-break.
    • Remove unused candidate field; add tests for Anthropic-priority selection and stable alias tie-breaks; regenerate snapshot and apply treefmt formatting in Nix/Rust tests.

Written for commit a41584e. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes

    • Updated pricing and context limits for several Claude and Anthropic models; normalized cache-related cost values.
  • Tests

    • Added unit tests ensuring deterministic and preferred-provider selection when duplicate pricing entries occur.
  • Chores

    • Improved duplicate-entry resolution to produce more stable, predictable pricing snapshots.
  • Style

    • Minor test comment and formatting cleanup.

Sort models.dev provider and model entries before compacting the pinned catalog so generator output no longer depends on upstream glob iteration order.

Resolve duplicate pricing keys with an explicit candidate ranking: prefer the Anthropic provider, keep entries with complete cache and context metadata, and use source identifiers as a deterministic tie-break. This prevents aliased provider rows from randomly winning over canonical Anthropic rows.

Add regression coverage for Anthropic-priority selection and stable duplicate tie-breaking.
Regenerate the pinned models.dev pricing snapshot after making duplicate key selection deterministic.

The updated snapshot restores canonical Anthropic cache and context metadata for rows that previously depended on provider iteration order. Keeping the generated output separate makes the generator logic and data update independently revertable.
Apply the treefmt change required by the pre-push hook in the Codex fallback loader test.

Keeping this as a separate commit isolates the hook-driven formatting change from the models.dev pricing fix.
@coderabbitai

coderabbitai Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Refactors models.dev pricing selection to build typed candidates with provenance, deterministically resolve duplicate pricing keys via a comparison predicate, updates the generator to use the predicate and sorted iteration, adds unit tests for replacement logic, and adjusts pricing JSON entries. Minor Rust test comment/formatting edits included.

Changes

Pricing Snapshot Selection Refactor

Layer / File(s) Summary
Pricing candidate type and comparison logic
nix/models-dev-compact.ts
Introduces ModelsDevPricingCandidate type capturing pricing key, source identifiers, and feature flags; exports shouldReplaceModelsDevPricingCandidate predicate and implements internal comparison helpers ordering candidates by provider priority, cache/context capabilities, and string identifiers.
Models.dev generation with typed candidate selection
nix/models-dev-gen.ts
Extends Provider type with optional id field; refactors the snapshot builder to construct ModelsDevPricingCandidate objects capturing provenance and features, use deterministic iteration via sortedEntries, and apply the replacement predicate to resolve duplicate pricing keys rather than skipping them immediately.
Candidate selection test coverage
nix/models-dev-compact.test.ts
Validates that the replacement predicate prefers Anthropic provider candidates when pricing keys match, and establishes stable tie-break ordering when candidates are otherwise equivalent.
Pricing snapshot data updates
rust/crates/ccusage/src/models-dev-pricing.json
Updates context limits and normalizes cache cost values across Claude Opus, Sonnet, Haiku, and Thinking variants; replaces cost and limit fields for specific models per the new selection logic output.

Rust Test Cleanup

Layer / File(s) Summary
Test comment and assertion formatting
rust/crates/ccusage/src/adapter/codex/loader.rs
Corrects test comment wording ("unparseable" → "unparsable") and reformats an assertion closure predicate without functional changes.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Suggested reviewers

  • pullfrog

Poem

🐇 I sniff the candidates, nose twitching bright,
I pick Anthropic when keys match just right,
I hop in order, ties broken neat and fair,
The snapshot sorted, provenance laid bare,
A rabbit cheers the code with joyful flair 🥕

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 10.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely summarizes the main change: stabilizing the models.dev pricing snapshot generation to eliminate non-deterministic behavior caused by provider iteration order.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/fix-models-dev-pricing-determinism

Comment @coderabbitai help to get the list of available commands and usage tips.

@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide a41584e Commit Preview URL

Branch Preview URL
Jun 12 2026, 03:11 PM

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes — Fixes non-determinism in the models.dev pricing snapshot generator by sorting provider/model iteration and resolving duplicate pricing keys via a well-defined comparator rather than first-wins.

  • Deterministic duplicate resolution — compareModelsDevPricingCandidates prefers Anthropic provider entries, then entries with cache/context metadata, then stable string tie-breaks.
  • Sorted iteration — sortedEntries in nix/models-dev-gen.ts ensures provider and model iteration order is deterministic regardless of upstream object key ordering.
  • Tests for comparator — two new unit tests covering Anthropic priority and stable string tie-break.
  • Regenerated snapshot — multiple models regained cache_read/cache_write fields that were silently dropped before; some context limits corrected.
  • Formatting fix — cosmetic rustfmt + spelling tweak in loader.rs.

Pullfrog  | View workflow run | Using Big Pickle (free via Pullfrog for OSS) | 𝕏

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
nix/models-dev-compact.ts (1)

60-62: 💤 Low value

compareStringAscending has inverted semantics compared to standard comparators.

The function returns 1 when left < right, but standard ascending comparators (like localeCompare or sort callbacks) return negative values in that case. The current implementation effectively makes "lexicographically smaller" strings rank higher in priority, which appears intentional for the tie-break behavior, but the name is misleading.

Consider renaming to compareStringPreferSmaller or inverting the return values to match conventional comparator semantics.

Option A: Rename for clarity
-function compareStringAscending(left: string, right: string): number {
+function compareStringPreferSmaller(left: string, right: string): number {
 	return left === right ? 0 : left < right ? 1 : -1;
 }
Option B: Use standard semantics
-function compareStringAscending(left: string, right: string): number {
-	return left === right ? 0 : left < right ? 1 : -1;
+function compareStringDescending(left: string, right: string): number {
+	return left.localeCompare(right) * -1;
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@nix/models-dev-compact.ts` around lines 60 - 62, The comparator
compareStringAscending has inverted semantics (returns 1 when left < right);
rename it to compareStringPreferSmaller (or another name reflecting that smaller
strings are preferred) and update every reference/call site to use the new name,
and update its JSDoc/comment to document the non-standard sign convention;
alternatively, if standard ascending behavior is desired, invert the return
values to return -1 when left < right and 1 when left > right and adjust any
dependent tie-break logic accordingly.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@nix/models-dev-compact.ts`:
- Around line 60-62: The comparator compareStringAscending has inverted
semantics (returns 1 when left < right); rename it to compareStringPreferSmaller
(or another name reflecting that smaller strings are preferred) and update every
reference/call site to use the new name, and update its JSDoc/comment to
document the non-standard sign convention; alternatively, if standard ascending
behavior is desired, invert the return values to return -1 when left < right and
1 when left > right and adjust any dependent tie-break logic accordingly.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: db842957-31a9-49b0-9363-b25002072d97

📥 Commits

Reviewing files that changed from the base of the PR and between bfd28e0 and 1c02a04.

📒 Files selected for processing (5)
  • nix/models-dev-compact.test.ts
  • nix/models-dev-compact.ts
  • nix/models-dev-gen.ts
  • rust/crates/ccusage/src/adapter/codex/loader.rs
  • rust/crates/ccusage/src/models-dev-pricing.json

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 5 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

Comment thread nix/models-dev-gen.ts Outdated
Comment thread nix/models-dev-compact.ts Outdated
@pkg-pr-new

pkg-pr-new Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ccusage@1304

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-arm64@1304

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-x64@1304

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/@ccusage/ccusage-linux-arm64@1304

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/@ccusage/ccusage-linux-x64@1304

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/@ccusage/ccusage-win32-x64@1304

commit: a41584e

Use locale-independent string comparison for sortedEntries so snapshot
generation does not depend on the runner's locale, and remove the unused
pricingKey field from ModelsDevPricingCandidate (the key is tracked
externally as the selected map key).

Co-authored-by: Codesmith <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 1c02a047c33d
Base SHA: bfd28e031e0d

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 956.5ms 932.4ms 61.6ms 3
PR pkg.pr.new 1c02a04 845.2ms 984.3ms 67.2ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: 1c02a04. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 882.4ms 899.7ms 0.98x 736.00 MiB 696.25 MiB 0.95x 1.14 GiB/s 1.12 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 260.3ms 233.1ms 1.12x 91.25 MiB 91.75 MiB 1.01x 3.87 GiB/s 4.32 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 896.5ms 1.12 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 769.9ms 1.31 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 177.7ms 5.66 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 132.0ms 7.63 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 46.6ms 38.1ms 1.22x 44.00 MiB 44.00 MiB 1.00x 0.03 MiB/s 0.04 MiB/s
claude session --offline --json 0.00 MiB 37.1ms 36.7ms 1.01x 44.00 MiB 44.25 MiB 1.01x 0.04 MiB/s 0.04 MiB/s
codex daily --offline --json 0.00 MiB 36.7ms 40.7ms 0.90x 44.25 MiB 44.00 MiB 0.99x 0.02 MiB/s 0.02 MiB/s
codex session --offline --json 0.00 MiB 39.3ms 39.6ms 0.99x 44.25 MiB 44.00 MiB 0.99x 0.02 MiB/s 0.02 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 759.7ms 923.3ms 0.82x 734.25 MiB 749.50 MiB 1.02x 1.33 GiB/s 1.09 GiB/s
codex --offline --json 1.01 GiB 156.5ms 160.3ms 0.98x 92.75 MiB 90.50 MiB 0.98x 6.43 GiB/s 6.28 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 1c02a047c33d
Base SHA: bfd28e031e0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 682.2ms 900.1ms 63.0ms 3
PR pkg.pr.new 1c02a04 985.8ms 699.1ms 71.5ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: 1c02a04. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 931.0ms 984.0ms 0.95x 732.25 MiB 724.50 MiB 0.99x 1.08 GiB/s 1.02 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 215.6ms 218.1ms 0.99x 87.75 MiB 91.75 MiB 1.05x 4.67 GiB/s 4.62 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 962.9ms 1.05 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 849.8ms 1.18 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 179.3ms 5.61 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 136.0ms 7.40 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 47.7ms 6.8ms 6.99x 44.25 MiB 2.75 MiB 0.06x 0.03 MiB/s 0.23 MiB/s
claude session --offline --json 0.00 MiB 43.2ms 6.2ms 7.01x 44.00 MiB 2.75 MiB 0.06x 0.04 MiB/s 0.25 MiB/s
codex daily --offline --json 0.00 MiB 41.0ms 5.7ms 7.25x 44.25 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.15 MiB/s
codex session --offline --json 0.00 MiB 38.2ms 6.6ms 5.77x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.13 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 929.6ms 923.8ms 1.01x 704.25 MiB 735.25 MiB 1.04x 1.08 GiB/s 1.09 GiB/s
codex --offline --json 1.01 GiB 190.0ms 148.2ms 1.28x 91.25 MiB 92.50 MiB 1.01x 5.30 GiB/s 6.79 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes — incremental review of commit c4260d3 on top of the prior pullfrog review (at 1c02a04). Drops an unused field and replaces locale-dependent localeCompare with locale-independent JS comparison for deterministic snapshot generation.

  • Removed unused pricingKey from ModelsDevPricingCandidate — the field was written but never read; clean removal from type, construction site, and test fixtures.
  • Replaced localeCompare with plain JS comparison in sortedEntries — localeCompare output varies by runtime locale; the ternary comparison uses UTF-16 code unit ordering (deterministic across machines for the ASCII keys used here).

Pullfrog  | View workflow run | Using Big Pickle (free via Pullfrog for OSS) | 𝕏

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: c4260d3be0b3
Base SHA: bfd28e031e0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 1.059s 768.1ms 55.2ms 3
PR pkg.pr.new c4260d3 1.087s 786.6ms 54.8ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: c4260d3. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 862.3ms 893.1ms 0.97x 733.25 MiB 728.75 MiB 0.99x 1.17 GiB/s 1.13 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 183.5ms 181.0ms 1.01x 92.25 MiB 93.75 MiB 1.02x 5.49 GiB/s 5.56 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 856.7ms 1.18 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 767.8ms 1.31 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 154.3ms 6.52 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 120.8ms 8.33 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 38.0ms 4.5ms 8.41x 44.00 MiB 2.75 MiB 0.06x 0.04 MiB/s 0.34 MiB/s
claude session --offline --json 0.00 MiB 47.8ms 4.5ms 10.58x 44.25 MiB 2.75 MiB 0.06x 0.03 MiB/s 0.34 MiB/s
codex daily --offline --json 0.00 MiB 36.8ms 4.2ms 8.86x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.21 MiB/s
codex session --offline --json 0.00 MiB 38.5ms 4.3ms 9.01x 44.25 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.20 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 804.4ms 806.6ms 1.00x 740.50 MiB 735.75 MiB 0.99x 1.25 GiB/s 1.25 GiB/s
codex --offline --json 1.01 GiB 155.1ms 121.7ms 1.27x 92.75 MiB 92.25 MiB 0.99x 6.49 GiB/s 8.27 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: c4260d3be0b3
Base SHA: bfd28e031e0d

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 817.5ms 863.3ms 67.3ms 3
PR pkg.pr.new c4260d3 796.7ms 906.6ms 64.4ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: c4260d3. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 858.7ms 917.9ms 0.94x 732.50 MiB 738.50 MiB 1.01x 1.17 GiB/s 1.10 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 203.6ms 204.2ms 1.00x 89.50 MiB 91.25 MiB 1.02x 4.95 GiB/s 4.93 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 929.9ms 1.08 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 935.8ms 1.08 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 173.7ms 5.80 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 139.9ms 7.19 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 46.0ms 41.0ms 1.12x - 44.00 MiB - 0.03 MiB/s 0.04 MiB/s
claude session --offline --json 0.00 MiB 44.5ms 45.9ms 0.97x 44.25 MiB 44.00 MiB 0.99x 0.03 MiB/s 0.03 MiB/s
codex daily --offline --json 0.00 MiB 47.3ms 44.1ms 1.07x 44.25 MiB 44.00 MiB 0.99x 0.02 MiB/s 0.02 MiB/s
codex session --offline --json 0.00 MiB 45.3ms 44.7ms 1.01x 44.25 MiB 44.00 MiB 0.99x 0.02 MiB/s 0.02 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 874.5ms 923.9ms 0.95x 736.75 MiB 737.00 MiB 1.00x 1.15 GiB/s 1.09 GiB/s
codex --offline --json 1.01 GiB 182.8ms 181.9ms 1.01x 90.25 MiB 91.50 MiB 1.01x 5.51 GiB/s 5.53 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

Remove unused candidate state flagged during review and rename the string tie-break helper so its priority semantics are explicit.

Replace localeCompare in the models.dev generator with bytewise string comparison to keep snapshot ordering independent of locale or ICU behaviour.
@pullfrog

pullfrog Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

The failed check suite (ID 73773779627) ran on commit c4260d3 — an older commit before the treefmt formatting fix was applied. Commits 1bde8b8 and 7dbd001 (already on the branch) contain the fix for nix/models-dev-gen.ts that treefmt was complaining about. Pushed an empty commit (a41584e) to trigger a fresh CI run on the latest tip.

Task list (6/6 completed)

Pullfrog  | View workflow run | via Pullfrog | Using Big Pickle (free via Pullfrog for OSS) | 𝕏

@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 1bde8b86fb2c
Base SHA: bfd28e031e0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 737.7ms 798.6ms 51.5ms 3
PR pkg.pr.new 1bde8b8 984.2ms 915.6ms 52.2ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: 1bde8b8. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 903.9ms 774.7ms 1.17x 736.50 MiB 739.25 MiB 1.00x 1.11 GiB/s 1.30 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 172.1ms 174.0ms 0.99x 88.25 MiB 92.00 MiB 1.04x 5.85 GiB/s 5.79 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 767.5ms 1.31 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 866.9ms 1.16 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 154.3ms 6.52 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 117.7ms 8.56 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 38.4ms 4.6ms 8.43x 44.25 MiB 2.75 MiB 0.06x 0.04 MiB/s 0.34 MiB/s
claude session --offline --json 0.00 MiB 37.0ms 4.5ms 8.22x 44.00 MiB 2.75 MiB 0.06x 0.04 MiB/s 0.34 MiB/s
codex daily --offline --json 0.00 MiB 37.4ms 4.2ms 8.84x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.20 MiB/s
codex session --offline --json 0.00 MiB 36.8ms 4.2ms 8.77x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.20 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 797.9ms 709.1ms 1.13x 742.00 MiB 733.75 MiB 0.99x 1.26 GiB/s 1.42 GiB/s
codex --offline --json 1.01 GiB 155.7ms 120.9ms 1.29x 90.50 MiB 87.75 MiB 0.97x 6.47 GiB/s 8.33 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 1bde8b86fb2c
Base SHA: bfd28e031e0d

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 861.4ms 1.118s 88.2ms 3
PR pkg.pr.new 1bde8b8 821.0ms 1.040s 70.4ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: 1bde8b8. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 997.6ms 1.009s 0.99x 740.75 MiB 726.50 MiB 0.98x 1.01 GiB/s 1021.46 MiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 317.1ms 248.5ms 1.28x 92.25 MiB 92.00 MiB 1.00x 3.18 GiB/s 4.05 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 786.7ms 1.28 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 926.3ms 1.09 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 262.1ms 3.84 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 216.5ms 4.65 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 55.9ms 44.7ms 1.25x 44.00 MiB 44.25 MiB 1.01x 0.03 MiB/s 0.03 MiB/s
claude session --offline --json 0.00 MiB 56.1ms 56.1ms 1.00x 44.00 MiB 44.25 MiB 1.01x 0.03 MiB/s 0.03 MiB/s
codex daily --offline --json 0.00 MiB 49.4ms 67.5ms 0.73x 44.25 MiB 44.25 MiB 1.00x 0.02 MiB/s 0.01 MiB/s
codex session --offline --json 0.00 MiB 79.9ms 60.1ms 1.33x 44.00 MiB 44.25 MiB 1.01x 0.01 MiB/s 0.01 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 848.7ms 1.025s 0.83x 719.25 MiB 735.00 MiB 1.02x 1.19 GiB/s 1005.98 MiB/s
codex --offline --json 1.01 GiB 261.4ms 293.0ms 0.89x 95.75 MiB 92.50 MiB 0.97x 3.85 GiB/s 3.44 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 7dbd001bdf86
Base SHA: bfd28e031e0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 815.7ms 1.001s 66.0ms 3
PR pkg.pr.new 7dbd001 888.4ms 990.3ms 63.6ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: 7dbd001. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 998.0ms 959.6ms 1.04x 741.75 MiB 725.50 MiB 0.98x 1.01 GiB/s 1.05 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 225.7ms 211.0ms 1.07x 91.00 MiB 92.50 MiB 1.02x 4.46 GiB/s 4.77 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 909.8ms 1.11 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 894.6ms 1.13 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 202.4ms 4.97 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 142.5ms 7.06 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 43.3ms 6.4ms 6.77x 44.00 MiB 3.00 MiB 0.07x 0.04 MiB/s 0.24 MiB/s
claude session --offline --json 0.00 MiB 44.0ms 6.8ms 6.51x 44.00 MiB 2.75 MiB 0.06x 0.04 MiB/s 0.23 MiB/s
codex daily --offline --json 0.00 MiB 41.5ms 5.7ms 7.30x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.15 MiB/s
codex session --offline --json 0.00 MiB 40.1ms 5.4ms 7.36x 44.25 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.16 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 817.5ms 778.2ms 1.05x 728.75 MiB 738.25 MiB 1.01x 1.23 GiB/s 1.29 GiB/s
codex --offline --json 1.01 GiB 171.2ms 144.5ms 1.18x 93.25 MiB 87.25 MiB 0.94x 5.88 GiB/s 6.97 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 7dbd001bdf86
Base SHA: bfd28e031e0d

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 812.7ms 729.0ms 71.4ms 3
PR pkg.pr.new 7dbd001 1.353s 816.3ms 70.8ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: 7dbd001. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 992.7ms 1.065s 0.93x 730.50 MiB 726.25 MiB 0.99x 1.01 GiB/s 967.79 MiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 224.2ms 229.4ms 0.98x 92.50 MiB 92.00 MiB 0.99x 4.49 GiB/s 4.39 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 1.005s 1.00 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 989.6ms 1.02 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 196.3ms 5.13 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 150.4ms 6.69 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 50.8ms 49.5ms 1.03x 44.25 MiB 44.00 MiB 0.99x 0.03 MiB/s 0.03 MiB/s
claude session --offline --json 0.00 MiB 48.3ms 53.0ms 0.91x 44.25 MiB 44.00 MiB 0.99x 0.03 MiB/s 0.03 MiB/s
codex daily --offline --json 0.00 MiB 44.2ms 45.5ms 0.97x 44.25 MiB 44.00 MiB 0.99x 0.02 MiB/s 0.02 MiB/s
codex session --offline --json 0.00 MiB 47.7ms 48.9ms 0.98x 44.00 MiB 44.00 MiB 1.00x 0.02 MiB/s 0.02 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 892.6ms 954.5ms 0.94x 740.50 MiB 738.25 MiB 1.00x 1.13 GiB/s 1.05 GiB/s
codex --offline --json 1.01 GiB 186.1ms 188.7ms 0.99x 93.00 MiB 95.00 MiB 1.02x 5.41 GiB/s 5.33 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: a41584ec32a4
Base SHA: bfd28e031e0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 785.9ms 837.0ms 62.0ms 3
PR pkg.pr.new a41584e 883.6ms 868.3ms 66.6ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: a41584e. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 901.5ms 989.4ms 0.91x 729.50 MiB 731.75 MiB 1.00x 1.12 GiB/s 1.02 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 199.1ms 203.6ms 0.98x 89.00 MiB 88.25 MiB 0.99x 5.06 GiB/s 4.95 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 970.8ms 1.04 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 842.7ms 1.19 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 177.6ms 5.67 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 131.1ms 7.68 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 43.7ms 6.4ms 6.85x 44.00 MiB 3.00 MiB 0.07x 0.04 MiB/s 0.24 MiB/s
claude session --offline --json 0.00 MiB 44.5ms 6.0ms 7.44x 44.25 MiB 2.75 MiB 0.06x 0.03 MiB/s 0.26 MiB/s
codex daily --offline --json 0.00 MiB 39.8ms 5.4ms 7.44x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.16 MiB/s
codex session --offline --json 0.00 MiB 45.0ms 6.0ms 7.50x 44.00 MiB 2.75 MiB 0.06x 0.02 MiB/s 0.14 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 950.3ms 883.0ms 1.08x 723.50 MiB 743.50 MiB 1.03x 1.06 GiB/s 1.14 GiB/s
codex --offline --json 1.01 GiB 176.5ms 133.4ms 1.32x 90.75 MiB 91.50 MiB 1.01x 5.71 GiB/s 7.55 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: a41584ec32a4
Base SHA: bfd28e031e0d

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new bfd28e031e0d 813.0ms 885.4ms 65.0ms 3
PR pkg.pr.new a41584e 1.205s 889.0ms 65.7ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: bfd28e031e0d; PR package: a41584e. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 966.1ms 926.5ms 1.04x 729.50 MiB 737.75 MiB 1.01x 1.04 GiB/s 1.09 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 203.7ms 213.3ms 0.95x 89.25 MiB 89.50 MiB 1.00x 4.94 GiB/s 4.72 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 954.4ms 1.05 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 960.7ms 1.05 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 177.9ms 5.66 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 139.2ms 7.23 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 48.7ms 45.1ms 1.08x 44.25 MiB 44.00 MiB 0.99x 0.03 MiB/s 0.03 MiB/s
claude session --offline --json 0.00 MiB 42.6ms 45.2ms 0.94x 44.25 MiB 44.25 MiB 1.00x 0.04 MiB/s 0.03 MiB/s
codex daily --offline --json 0.00 MiB 43.1ms 43.6ms 0.99x 44.00 MiB 44.25 MiB 1.01x 0.02 MiB/s 0.02 MiB/s
codex session --offline --json 0.00 MiB 41.9ms 44.7ms 0.94x 44.25 MiB 44.00 MiB 0.99x 0.02 MiB/s 0.02 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 938.9ms 902.5ms 1.04x 739.75 MiB 739.75 MiB 1.00x 1.07 GiB/s 1.12 GiB/s
codex --offline --json 1.01 GiB 176.2ms 184.7ms 0.95x - - - 5.72 GiB/s 5.45 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 17.32 KiB 17.32 KiB +0.00 KiB 1.00x
installed native package binary 3328.59 KiB 3328.59 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@ryoppippi
ryoppippi merged commit 54d318f into main Jun 12, 2026
36 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant