Skip to content

perf(all): stream default agent aggregations - #1158

Merged
ryoppippi merged 2 commits into
mainfrom
codex/rust-perf-memory
May 25, 2026
Merged

ryoppippi merged 2 commits into
mainfrom
codex/rust-perf-memory

Conversation

@ryoppippi

@ryoppippi ryoppippi commented May 25, 2026 •

Copy link
Copy Markdown
Member

Optimizes the default all-agent reports without adding dependencies or increasing the release binary size.

What changed:

  • Reuses Claude daily summaries for all-agent date reports instead of loading full Claude entry vectors first.
  • Reuses Codex grouped aggregation for unfiltered all-agent reports so large Codex sessions stream directly into report groups instead of materializing an intermediate event vector.

This helps the memory class discussed in #1124 for the default all-agent path; direct Codex reports already stream session JSONL files on current main.

Performance:

  • Synthetic 50 MB Codex JSONL fixture, isolated HOME/XDG/CODEX_HOME, after checking background process load.
  • main: 259.8 ms mean, 82,132,992 byte max RSS.
  • head: 240.9 ms mean, 5,505,024 byte max RSS.
  • Release binary size: unchanged at 2,746,368 bytes versus origin/main.

Testing:

  • pnpm run format
  • pnpm run test
  • cargo fmt --manifest-path rust/Cargo.toml --all --check
  • cargo clippy --manifest-path rust/Cargo.toml --workspace --all-targets -- -D warnings
  • Fixture JSON parity for all-agent, Claude, and Codex daily reports against origin/main
  • Synthetic 50 MB Codex all-agent benchmark and RSS check

Summary by cubic

Streamed default all‑agent aggregations to cut peak memory without changing the binary size. Fixed Codex grouped dedupe so events from different sessions aren’t merged.

  • Refactors

    • Claude all‑agent date path reads daily summaries instead of materializing full entry vectors.
    • Codex all‑agent (no date filters) uses grouped aggregation to stream session JSONL directly into report groups; falls back to events when filters are set.
    • Added compact date‑bound filtering for Claude daily summaries; exposed load_groups as pub(crate).
    • Perf (50 MB Codex fixture): 259.8 ms → 240.9 ms; max RSS 82 MB → 5.5 MB; release binary size unchanged.
  • Bug Fixes

    • Codex grouped dedupe includes session id (hash + length) in the key to keep distinct sessions separate; added regression test.

Written for commit 2676671. Summary will update on new commits. Review in cubic

Summary by CodeRabbit

  • Bug Fixes

    • Optimized Codex data loading to short-circuit when no date bounds are specified.
    • Improved Claude agent row loading logic with enhanced session handling and daily summary filtering.
  • Tests

    • Added unit tests for daily summary date filtering validation.

Review Change Stack

Reuse the specialized Claude daily summary loader for all-agent date reports so the default all-agent path does not build full Claude loaded-entry vectors before summarizing.

Reuse the Codex grouped aggregation path for unfiltered all-agent reports so large Codex session histories stream directly into report groups instead of materializing and deduplicating an intermediate event vector. This reduces peak memory for large Codex JSONL inputs while keeping binary size unchanged.

Validation: pnpm run format; pnpm run test; cargo fmt --manifest-path rust/Cargo.toml --all --check; cargo clippy --manifest-path rust/Cargo.toml --workspace --all-targets -- -D warnings; release JSON parity against fixtures; synthetic 50 MB Codex all-agent benchmark.
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented May 25, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR refactors agent-row loading in the ccusage adapter. Claude loading now routes through a new load_claude_rows helper that preserves session handling but switches non-session cases to load daily summaries with custom date filtering. Codex loading short-circuits to return precomputed group rows when no date bounds are provided. The Codex module's load_groups function is widened to pub(crate) visibility and added to module re-exports.

Changes

Agent loader refactoring and Codex export adjustment

Layer / File(s) Summary
Claude agent loading with daily summaries and date filtering
rust/crates/ccusage/src/adapter/all/loader.rs
New load_claude_rows helper dispatches to session-capable loader or daily-summary loading with filter_daily_summaries_by_date for date bounds. Updated load_codex_rows early-returns precomputed codex::load_groups(...) when no bounds are set. Unit tests validate daily summary filtering within compact date windows.
Codex load_groups visibility and re-export refactoring
rust/crates/ccusage/src/adapter/codex/aggregate.rs, rust/crates/ccusage/src/adapter/codex/mod.rs
load_groups visibility widened from pub(super) to pub(crate) in aggregate.rs. Module re-exports updated to include load_groups alongside existing exports. Duplicate local import removed in favor of the re-export pattern.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • ryoppippi/ccusage#1131: Codex row loading refactoring and early-return behavior directly relate to Codex-events loader surface changes.
  • ryoppippi/ccusage#1137: Both PRs modify Codex aggregation entrypoint and re-export patterns from adapter/codex/mod.rs.
  • ryoppippi/ccusage#1130: Related work refactoring Claude agent loading path and daily-summary handling through dedicated adapter.

Suggested labels

enhancement

Poem

🐰 Agent loaders now split their ways,
Claude takes daily summaries through filtered days,
Codex shortcuts when bounds aren't found,
Visibility widened, exports unwound!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 11.11% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately reflects the main objective of optimizing all-agent aggregations through streaming, which is the primary goal mentioned in the PR objectives.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/rust-perf-memory

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai

coderabbitai Bot commented May 25, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented May 25, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide 2676671 Commit Preview URL

Branch Preview URL
May 25 2026, 07:24 PM

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 3 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="rust/crates/ccusage/src/adapter/all/loader.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/all/loader.rs:418">
P1: The new unfiltered Codex fast path uses a dedupe key that omits `session_id`, which can merge usage from distinct sessions and undercount all-agent totals.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

pricing: &PricingMap,
) -> Result<AgentRows> {
if shared.since.is_none() && shared.until.is_none() {
let groups = codex::load_groups(shared, kind)?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: The new unfiltered Codex fast path uses a dedupe key that omits session_id, which can merge usage from distinct sessions and undercount all-agent totals.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/all/loader.rs, line 418:

<comment>The new unfiltered Codex fast path uses a dedupe key that omits `session_id`, which can merge usage from distinct sessions and undercount all-agent totals.</comment>

<file context>
@@ -386,11 +378,55 @@ fn load_session_capable_summary_agent_rows(
     pricing: &PricingMap,
 ) -> Result<AgentRows> {
+    if shared.since.is_none() && shared.until.is_none() {
+        let groups = codex::load_groups(shared, kind)?;
+        let detected = !groups.is_empty();
+        let speed = codex::resolve_codex_speed(CodexSpeed::Auto);
</file context>

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@cubic-dev-ai @coderabbitai Fixed in 2676671 by adding the Codex session id hash and length to the grouped aggregation dedupe key, plus a regression test for matching usage events from distinct session files. Validation: cargo fmt --manifest-path rust/Cargo.toml --all --check; cargo test --manifest-path rust/Cargo.toml -p ccusage; cargo clippy --manifest-path rust/Cargo.toml --workspace --all-targets -- -D warnings; fixture JSON parity; synthetic 50 MB Codex benchmark.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@ryoppippi, thanks for the thorough fix and validation! Adding both the session ID hash and length to the dedupe key is a solid approach to prevent cross-session merging, and the regression test for distinct session files is a great addition. The fix looks complete.

(っ◔◡◔)っ ♥


If you found this review helpful, would you consider giving us a shout-out on X?

Thank you for using CodeRabbit!

Include the Codex session id hash and length in the grouped aggregation dedupe key so events with matching timestamp, model, and token counts from different sessions are not collapsed.

This fixes the issue identified by cubic after the all-agent fast path started reusing grouped Codex aggregation. Added a regression test covering matching usage events from distinct session files.

Validation: cargo fmt --manifest-path rust/Cargo.toml --all --check; cargo test --manifest-path rust/Cargo.toml -p ccusage; cargo clippy --manifest-path rust/Cargo.toml --workspace --all-targets -- -D warnings; fixture JSON parity; synthetic 50 MB Codex benchmark.
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented May 25, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@pkg-pr-new

pkg-pr-new Bot commented May 25, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ryoppippi/ccusage@1158

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-darwin-arm64@1158

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-darwin-x64@1158

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-linux-arm64@1158

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-linux-x64@1158

@ccusage/ccusage-win32-arm64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-win32-arm64@1158

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-win32-x64@1158

commit: 2676671

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 45d122b46290
Base SHA: 63cf17e9656a

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 63cf17e9656a 538.9ms 812.2ms 33.1ms 3
PR pkg.pr.new 45d122b46290 455.6ms 558.7ms 34.7ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 63cf17e9656a; PR package: 45d122b46290. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 492.2ms 494.1ms 1.00x 261.83 MiB 286.95 MiB 1.10x 2.05 GiB/s 2.04 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 377.2ms 373.9ms 1.01x 62.70 MiB 63.08 MiB 1.01x 2.67 GiB/s 2.69 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 490.9ms 2.05 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 460.9ms 2.18 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 380.7ms 2.64 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 327.0ms 3.08 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 30.7ms 4.1ms 7.48x 43.48 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.38 MiB/s
claude session --offline --json 0.00 MiB 30.7ms 4.2ms 7.39x 43.73 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.37 MiB/s
codex daily --offline --json 0.00 MiB 31.4ms 3.7ms 8.37x 43.61 MiB 2.83 MiB 0.06x 0.03 MiB/s 0.23 MiB/s
codex session --offline --json 0.00 MiB 30.7ms 3.7ms 8.37x 43.61 MiB 2.83 MiB 0.06x 0.03 MiB/s 0.23 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 488.4ms 474.4ms 1.03x 258.20 MiB 264.83 MiB 1.03x 2.06 GiB/s 2.12 GiB/s
codex --offline --json 1.01 GiB 365.8ms 324.7ms 1.13x 63.95 MiB 55.95 MiB 0.87x 2.75 GiB/s 3.10 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 14.25 KiB 14.25 KiB +0.00 KiB 1.00x
installed native package binary 3289.49 KiB 3289.49 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 45d122b46290
Base SHA: 63cf17e9656a

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 63cf17e9656a 654.1ms 576.7ms 33.1ms 3
PR pkg.pr.new 45d122b46290 475.8ms 579.3ms 34.0ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 63cf17e9656a; PR package: 45d122b46290. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 496.7ms 503.5ms 0.99x 272.58 MiB 265.45 MiB 0.97x 2.03 GiB/s 2.00 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 363.7ms 358.3ms 1.01x 65.20 MiB 59.95 MiB 0.92x 2.77 GiB/s 2.81 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 485.1ms 2.08 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 459.7ms 2.19 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 350.1ms 2.88 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 324.6ms 3.10 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 30.6ms 30.3ms 1.01x 43.48 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 30.7ms 30.2ms 1.02x 43.48 MiB 43.73 MiB 1.01x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 30.3ms 30.7ms 0.99x 43.61 MiB 43.73 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 30.1ms 30.0ms 1.00x 43.48 MiB 43.61 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 501.8ms 490.3ms 1.02x 264.33 MiB 264.70 MiB 1.00x 2.01 GiB/s 2.05 GiB/s
codex --offline --json 1.01 GiB 361.2ms 348.4ms 1.04x 51.83 MiB 52.95 MiB 1.02x 2.79 GiB/s 2.89 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 14.25 KiB 14.25 KiB +0.00 KiB 1.00x
installed native package binary 3289.49 KiB 3289.49 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 2676671f61d3
Base SHA: 63cf17e9656a

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 63cf17e9656a 530.9ms 447.3ms 34.2ms 3
PR pkg.pr.new 2676671f61d3 655.7ms 427.4ms 34.7ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 63cf17e9656a; PR package: 2676671f61d3. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 506.8ms 522.7ms 0.97x 248.33 MiB 258.70 MiB 1.04x 1.99 GiB/s 1.93 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 364.1ms 370.3ms 0.98x 50.95 MiB 73.08 MiB 1.43x 2.76 GiB/s 2.72 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 522.9ms 1.93 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 455.6ms 2.21 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 370.7ms 2.72 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 343.8ms 2.93 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 31.3ms 4.2ms 7.53x 43.61 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.37 MiB/s
claude session --offline --json 0.00 MiB 31.7ms 4.2ms 7.63x 43.48 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.37 MiB/s
codex daily --offline --json 0.00 MiB 31.5ms 3.9ms 8.16x 43.61 MiB 2.83 MiB 0.06x 0.03 MiB/s 0.22 MiB/s
codex session --offline --json 0.00 MiB 30.7ms 3.8ms 8.18x 43.61 MiB 2.83 MiB 0.06x 0.03 MiB/s 0.23 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 496.9ms 459.1ms 1.08x 258.83 MiB 257.58 MiB 1.00x 2.03 GiB/s 2.19 GiB/s
codex --offline --json 1.01 GiB 362.4ms 342.8ms 1.06x 64.08 MiB 81.76 MiB 1.28x 2.78 GiB/s 2.94 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 14.25 KiB 14.25 KiB +0.00 KiB 1.00x
installed native package binary 3289.49 KiB 3289.49 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 2676671f61d3
Base SHA: 63cf17e9656a

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 63cf17e9656a 861.9ms 475.9ms 33.8ms 3
PR pkg.pr.new 2676671f61d3 408.3ms 483.7ms 33.8ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 63cf17e9656a; PR package: 2676671f61d3. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 513.8ms 507.5ms 1.01x 267.20 MiB 254.33 MiB 0.95x 1.96 GiB/s 1.98 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 370.9ms 382.3ms 0.97x 55.08 MiB 78.58 MiB 1.43x 2.71 GiB/s 2.63 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 496.9ms 2.03 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 471.6ms 2.13 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 378.7ms 2.66 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 347.0ms 2.90 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 31.2ms 30.7ms 1.02x 43.61 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 30.5ms 30.5ms 1.00x 43.48 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 30.8ms 30.6ms 1.01x 43.61 MiB 43.48 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 30.2ms 30.7ms 0.98x 43.48 MiB 43.48 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 496.3ms 492.3ms 1.01x 262.20 MiB 272.08 MiB 1.04x 2.03 GiB/s 2.05 GiB/s
codex --offline --json 1.01 GiB 359.0ms 369.7ms 0.97x 57.20 MiB 78.58 MiB 1.37x 2.80 GiB/s 2.72 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 14.25 KiB 14.25 KiB +0.00 KiB 1.00x
installed native package binary 3289.49 KiB 3289.49 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@ryoppippi
ryoppippi merged commit eef2543 into main May 25, 2026
39 checks passed
@ryoppippi
ryoppippi deleted the codex/rust-perf-memory branch May 25, 2026 21:22
jay-tau added a commit to jay-tau/ccusage that referenced this pull request Jun 8, 2026
…omment

Opus 4.7-xhigh (seq-18 r1f) caught a stale doc-comment value: the
header for `accepts_fractional_premium_request_cost`
(`parser.rs:1194`) says the fixture mixes "a fractional row
(Opus 4.7, cost 7.5) with an integer-cost sibling (Sonnet, cost 4)",
but the actual fixture (`parser.rs:1221`) sets the Sonnet sibling
to cost 1 ("1 request × 1× multiplier = cost 1"), the inline
comment at `parser.rs:1219-1220` says "1× multiplier ... so 1
request × 1× = cost 1", and the assertion at `parser.rs:1240`
checks `Some(1.0)`.

Reviewer's git archaeology: the fixture was introduced as `cost: 4`
in `20b20ca` together with the matching "Sonnet, cost 4" header.
A later commit (`cc48cad`) renumbered the fixture to `cost: 1`
and added the "1× multiplier" inline comment, but the older
header line was not updated alongside it. Both commits are in this
PR's history.

## Fix

Reworded the header parenthetical from "(Sonnet, cost 4)" to
"(Sonnet 4.5, cost 1 — 1 request × 1× multiplier)". This:

- Matches the fixture (cost 1).
- Matches the inline comment's multiplier explanation.
- Names the specific Sonnet variant (4.5) so a reader doesn't
  have to scroll to the fixture to disambiguate.

## Validation at this tip

- `cd rust && cargo fmt --check` — clean
- `cd rust && cargo clippy --workspace --all-targets -- -D warnings` — clean
- `cd rust && cargo test --workspace` — 368 tests pass (unchanged;
  doc-comment-only change)

## Out-of-scope finding from same round (NOT addressed by this commit)

GPT-5.5 (seq-18 r1f) flagged a HIGH about `codex/aggregate.rs:347`
double-counting branch-history events because the aggregate dedupe
key includes `session_id` while the loader dedupe intentionally
removed it. `git blame` confirms that file was last touched by
upstream PRs ccusage#1158 (eef2543) and ccusage#1137 (41c0c6f) and has NOT been
touched by PR ccusage#1209. `git log upstream/main..HEAD -- rust/crates/ccusage/src/adapter/codex/`
returns zero commits from this PR.

That bug is pre-existing in the Codex adapter and unrelated to PR
ccusage#1209's scope (Copilot adapter + cross-source --all credits
handling). Per repo guidelines ("Don't fix pre-existing issues
unrelated to your task ... unless they're tightly coupled to the
code you're changing"), it's out of scope here and should be filed
as a separate issue against the Codex adapter or fixed in a
follow-up PR. The seq-17 / seq-18 review framework happened to
surface it, but addressing it in this branch would expand scope.

Refs ccusage#1174.

Co-authored-by: Copilot <[email protected]>
jay-tau added a commit to jay-tau/ccusage that referenced this pull request Jun 9, 2026
…omment

Opus 4.7-xhigh (seq-18 r1f) caught a stale doc-comment value: the
header for `accepts_fractional_premium_request_cost`
(`parser.rs:1194`) says the fixture mixes "a fractional row
(Opus 4.7, cost 7.5) with an integer-cost sibling (Sonnet, cost 4)",
but the actual fixture (`parser.rs:1221`) sets the Sonnet sibling
to cost 1 ("1 request × 1× multiplier = cost 1"), the inline
comment at `parser.rs:1219-1220` says "1× multiplier ... so 1
request × 1× = cost 1", and the assertion at `parser.rs:1240`
checks `Some(1.0)`.

Reviewer's git archaeology: the fixture was introduced as `cost: 4`
in `20b20ca` together with the matching "Sonnet, cost 4" header.
A later commit (`cc48cad`) renumbered the fixture to `cost: 1`
and added the "1× multiplier" inline comment, but the older
header line was not updated alongside it. Both commits are in this
PR's history.

## Fix

Reworded the header parenthetical from "(Sonnet, cost 4)" to
"(Sonnet 4.5, cost 1 — 1 request × 1× multiplier)". This:

- Matches the fixture (cost 1).
- Matches the inline comment's multiplier explanation.
- Names the specific Sonnet variant (4.5) so a reader doesn't
  have to scroll to the fixture to disambiguate.

## Validation at this tip

- `cd rust && cargo fmt --check` — clean
- `cd rust && cargo clippy --workspace --all-targets -- -D warnings` — clean
- `cd rust && cargo test --workspace` — 368 tests pass (unchanged;
  doc-comment-only change)

## Out-of-scope finding from same round (NOT addressed by this commit)

GPT-5.5 (seq-18 r1f) flagged a HIGH about `codex/aggregate.rs:347`
double-counting branch-history events because the aggregate dedupe
key includes `session_id` while the loader dedupe intentionally
removed it. `git blame` confirms that file was last touched by
upstream PRs ccusage#1158 (eef2543) and ccusage#1137 (41c0c6f) and has NOT been
touched by PR ccusage#1209. `git log upstream/main..HEAD -- rust/crates/ccusage/src/adapter/codex/`
returns zero commits from this PR.

That bug is pre-existing in the Codex adapter and unrelated to PR
ccusage#1209's scope (Copilot adapter + cross-source --all credits
handling). Per repo guidelines ("Don't fix pre-existing issues
unrelated to your task ... unless they're tightly coupled to the
code you're changing"), it's out of scope here and should be filed
as a separate issue against the Codex adapter or fixed in a
follow-up PR. The seq-17 / seq-18 review framework happened to
surface it, but addressing it in this branch would expand scope.

Refs ccusage#1174.

Co-authored-by: Copilot <[email protected]>
jay-tau added a commit to jay-tau/ccusage that referenced this pull request Jun 9, 2026
…omment

Opus 4.7-xhigh (seq-18 r1f) caught a stale doc-comment value: the
header for `accepts_fractional_premium_request_cost`
(`parser.rs:1194`) says the fixture mixes "a fractional row
(Opus 4.7, cost 7.5) with an integer-cost sibling (Sonnet, cost 4)",
but the actual fixture (`parser.rs:1221`) sets the Sonnet sibling
to cost 1 ("1 request × 1× multiplier = cost 1"), the inline
comment at `parser.rs:1219-1220` says "1× multiplier ... so 1
request × 1× = cost 1", and the assertion at `parser.rs:1240`
checks `Some(1.0)`.

Reviewer's git archaeology: the fixture was introduced as `cost: 4`
in `20b20ca` together with the matching "Sonnet, cost 4" header.
A later commit (`cc48cad`) renumbered the fixture to `cost: 1`
and added the "1× multiplier" inline comment, but the older
header line was not updated alongside it. Both commits are in this
PR's history.

## Fix

Reworded the header parenthetical from "(Sonnet, cost 4)" to
"(Sonnet 4.5, cost 1 — 1 request × 1× multiplier)". This:

- Matches the fixture (cost 1).
- Matches the inline comment's multiplier explanation.
- Names the specific Sonnet variant (4.5) so a reader doesn't
  have to scroll to the fixture to disambiguate.

## Validation at this tip

- `cd rust && cargo fmt --check` — clean
- `cd rust && cargo clippy --workspace --all-targets -- -D warnings` — clean
- `cd rust && cargo test --workspace` — 368 tests pass (unchanged;
  doc-comment-only change)

## Out-of-scope finding from same round (NOT addressed by this commit)

GPT-5.5 (seq-18 r1f) flagged a HIGH about `codex/aggregate.rs:347`
double-counting branch-history events because the aggregate dedupe
key includes `session_id` while the loader dedupe intentionally
removed it. `git blame` confirms that file was last touched by
upstream PRs ccusage#1158 (eef2543) and ccusage#1137 (41c0c6f) and has NOT been
touched by PR ccusage#1209. `git log upstream/main..HEAD -- rust/crates/ccusage/src/adapter/codex/`
returns zero commits from this PR.

That bug is pre-existing in the Codex adapter and unrelated to PR
ccusage#1209's scope (Copilot adapter + cross-source --all credits
handling). Per repo guidelines ("Don't fix pre-existing issues
unrelated to your task ... unless they're tightly coupled to the
code you're changing"), it's out of scope here and should be filed
as a separate issue against the Codex adapter or fixed in a
follow-up PR. The seq-17 / seq-18 review framework happened to
surface it, but addressing it in this branch would expand scope.

Refs ccusage#1174.

Co-authored-by: Copilot <[email protected]>
jay-tau added a commit to jay-tau/ccusage that referenced this pull request Jun 10, 2026
…omment

Opus 4.7-xhigh (seq-18 r1f) caught a stale doc-comment value: the
header for `accepts_fractional_premium_request_cost`
(`parser.rs:1194`) says the fixture mixes "a fractional row
(Opus 4.7, cost 7.5) with an integer-cost sibling (Sonnet, cost 4)",
but the actual fixture (`parser.rs:1221`) sets the Sonnet sibling
to cost 1 ("1 request × 1× multiplier = cost 1"), the inline
comment at `parser.rs:1219-1220` says "1× multiplier ... so 1
request × 1× = cost 1", and the assertion at `parser.rs:1240`
checks `Some(1.0)`.

Reviewer's git archaeology: the fixture was introduced as `cost: 4`
in `20b20ca` together with the matching "Sonnet, cost 4" header.
A later commit (`cc48cad`) renumbered the fixture to `cost: 1`
and added the "1× multiplier" inline comment, but the older
header line was not updated alongside it. Both commits are in this
PR's history.

## Fix

Reworded the header parenthetical from "(Sonnet, cost 4)" to
"(Sonnet 4.5, cost 1 — 1 request × 1× multiplier)". This:

- Matches the fixture (cost 1).
- Matches the inline comment's multiplier explanation.
- Names the specific Sonnet variant (4.5) so a reader doesn't
  have to scroll to the fixture to disambiguate.

## Validation at this tip

- `cd rust && cargo fmt --check` — clean
- `cd rust && cargo clippy --workspace --all-targets -- -D warnings` — clean
- `cd rust && cargo test --workspace` — 368 tests pass (unchanged;
  doc-comment-only change)

## Out-of-scope finding from same round (NOT addressed by this commit)

GPT-5.5 (seq-18 r1f) flagged a HIGH about `codex/aggregate.rs:347`
double-counting branch-history events because the aggregate dedupe
key includes `session_id` while the loader dedupe intentionally
removed it. `git blame` confirms that file was last touched by
upstream PRs ccusage#1158 (eef2543) and ccusage#1137 (41c0c6f) and has NOT been
touched by PR ccusage#1209. `git log upstream/main..HEAD -- rust/crates/ccusage/src/adapter/codex/`
returns zero commits from this PR.

That bug is pre-existing in the Codex adapter and unrelated to PR
ccusage#1209's scope (Copilot adapter + cross-source --all credits
handling). Per repo guidelines ("Don't fix pre-existing issues
unrelated to your task ... unless they're tightly coupled to the
code you're changing"), it's out of scope here and should be filed
as a separate issue against the Codex adapter or fixed in a
follow-up PR. The seq-17 / seq-18 review framework happened to
surface it, but addressing it in this branch would expand scope.

Refs ccusage#1174.

Co-authored-by: Copilot <[email protected]>
jay-tau added a commit to jay-tau/ccusage that referenced this pull request Jun 11, 2026
…omment

Opus 4.7-xhigh (seq-18 r1f) caught a stale doc-comment value: the
header for `accepts_fractional_premium_request_cost`
(`parser.rs:1194`) says the fixture mixes "a fractional row
(Opus 4.7, cost 7.5) with an integer-cost sibling (Sonnet, cost 4)",
but the actual fixture (`parser.rs:1221`) sets the Sonnet sibling
to cost 1 ("1 request × 1× multiplier = cost 1"), the inline
comment at `parser.rs:1219-1220` says "1× multiplier ... so 1
request × 1× = cost 1", and the assertion at `parser.rs:1240`
checks `Some(1.0)`.

Reviewer's git archaeology: the fixture was introduced as `cost: 4`
in `20b20ca` together with the matching "Sonnet, cost 4" header.
A later commit (`cc48cad`) renumbered the fixture to `cost: 1`
and added the "1× multiplier" inline comment, but the older
header line was not updated alongside it. Both commits are in this
PR's history.

## Fix

Reworded the header parenthetical from "(Sonnet, cost 4)" to
"(Sonnet 4.5, cost 1 — 1 request × 1× multiplier)". This:

- Matches the fixture (cost 1).
- Matches the inline comment's multiplier explanation.
- Names the specific Sonnet variant (4.5) so a reader doesn't
  have to scroll to the fixture to disambiguate.

## Validation at this tip

- `cd rust && cargo fmt --check` — clean
- `cd rust && cargo clippy --workspace --all-targets -- -D warnings` — clean
- `cd rust && cargo test --workspace` — 368 tests pass (unchanged;
  doc-comment-only change)

## Out-of-scope finding from same round (NOT addressed by this commit)

GPT-5.5 (seq-18 r1f) flagged a HIGH about `codex/aggregate.rs:347`
double-counting branch-history events because the aggregate dedupe
key includes `session_id` while the loader dedupe intentionally
removed it. `git blame` confirms that file was last touched by
upstream PRs ccusage#1158 (eef2543) and ccusage#1137 (41c0c6f) and has NOT been
touched by PR ccusage#1209. `git log upstream/main..HEAD -- rust/crates/ccusage/src/adapter/codex/`
returns zero commits from this PR.

That bug is pre-existing in the Codex adapter and unrelated to PR
ccusage#1209's scope (Copilot adapter + cross-source --all credits
handling). Per repo guidelines ("Don't fix pre-existing issues
unrelated to your task ... unless they're tightly coupled to the
code you're changing"), it's out of scope here and should be filed
as a separate issue against the Codex adapter or fixed in a
follow-up PR. The seq-17 / seq-18 review framework happened to
surface it, but addressing it in this branch would expand scope.

Refs ccusage#1174.

Co-authored-by: Copilot <[email protected]>
jay-tau added a commit to jay-tau/ccusage that referenced this pull request Jun 11, 2026
…omment

Opus 4.7-xhigh (seq-18 r1f) caught a stale doc-comment value: the
header for `accepts_fractional_premium_request_cost`
(`parser.rs:1194`) says the fixture mixes "a fractional row
(Opus 4.7, cost 7.5) with an integer-cost sibling (Sonnet, cost 4)",
but the actual fixture (`parser.rs:1221`) sets the Sonnet sibling
to cost 1 ("1 request × 1× multiplier = cost 1"), the inline
comment at `parser.rs:1219-1220` says "1× multiplier ... so 1
request × 1× = cost 1", and the assertion at `parser.rs:1240`
checks `Some(1.0)`.

Reviewer's git archaeology: the fixture was introduced as `cost: 4`
in `20b20ca` together with the matching "Sonnet, cost 4" header.
A later commit (`cc48cad`) renumbered the fixture to `cost: 1`
and added the "1× multiplier" inline comment, but the older
header line was not updated alongside it. Both commits are in this
PR's history.

## Fix

Reworded the header parenthetical from "(Sonnet, cost 4)" to
"(Sonnet 4.5, cost 1 — 1 request × 1× multiplier)". This:

- Matches the fixture (cost 1).
- Matches the inline comment's multiplier explanation.
- Names the specific Sonnet variant (4.5) so a reader doesn't
  have to scroll to the fixture to disambiguate.

## Validation at this tip

- `cd rust && cargo fmt --check` — clean
- `cd rust && cargo clippy --workspace --all-targets -- -D warnings` — clean
- `cd rust && cargo test --workspace` — 368 tests pass (unchanged;
  doc-comment-only change)

## Out-of-scope finding from same round (NOT addressed by this commit)

GPT-5.5 (seq-18 r1f) flagged a HIGH about `codex/aggregate.rs:347`
double-counting branch-history events because the aggregate dedupe
key includes `session_id` while the loader dedupe intentionally
removed it. `git blame` confirms that file was last touched by
upstream PRs ccusage#1158 (eef2543) and ccusage#1137 (41c0c6f) and has NOT been
touched by PR ccusage#1209. `git log upstream/main..HEAD -- rust/crates/ccusage/src/adapter/codex/`
returns zero commits from this PR.

That bug is pre-existing in the Codex adapter and unrelated to PR
ccusage#1209's scope (Copilot adapter + cross-source --all credits
handling). Per repo guidelines ("Don't fix pre-existing issues
unrelated to your task ... unless they're tightly coupled to the
code you're changing"), it's out of scope here and should be filed
as a separate issue against the Codex adapter or fixed in a
follow-up PR. The seq-17 / seq-18 review framework happened to
surface it, but addressing it in this branch would expand scope.

Refs ccusage#1174.

Co-authored-by: Copilot <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant