Skip to content

fix(codex): dedupe goal rollout token events - #1237

Merged
ryoppippi merged 3 commits into
mainfrom
codex/goal-dedupe-codex-rollouts
Jun 8, 2026
Merged

ryoppippi merged 3 commits into
mainfrom
codex/goal-dedupe-codex-rollouts

Conversation

@ryoppippi

@ryoppippi ryoppippi commented Jun 8, 2026 •

Copy link
Copy Markdown
Member

Deduplicates copied Codex token_count rows when rollout or goal session files contain the same usage event as another session file.

The streaming group aggregation path now uses the same token tuple based dedupe shape as the event loader, so table reports do not count copied session history once per rollout file.

Testing:

  • direnv exec . cargo test -p ccusage adapter::codex -- --nocapture
  • direnv exec . pnpm run test
  • pre-push hook: clippy, treefmt, gitleaks, cargo test

Summary by cubic

Fixes double-counting of Codex token usage when rollout or goal session files copy the same events. The streaming aggregator now uses the event loader’s token-tuple key and scopes it by report kind, so daily/weekly/monthly dedupe across files while session reports keep distinct groups.

  • Bug Fixes
    • Use a token-tuple dedupe key; drop session_id for daily/weekly/monthly to dedupe across session files, but include it for session reports to avoid collapsing separate sessions. Tests cover both single-thread and parallel paths.

Written for commit c7ffd05. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes

    • Improved token-usage deduplication: identical events across multiple session files are now deduplicated for aggregated totals, while session-specific reports remain separated so per-session counts are preserved.
  • Tests

    • Expanded unit tests to verify deduplication across distinct session files (single- and multi-threaded) and to confirm session-kind reports are kept separate.

Align the streaming Codex group aggregation dedupe key with the event loader so copied token_count rows from goal or rollout session files are counted once instead of once per session file.

This prevents daily and monthly table reports from double-counting copied Codex history while preserving the same token tuple based duplicate detection already used by JSON/all-agent loading paths.
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai review\n\nPlease review the latest push. This fixes Codex goal/rollout token_count dedupe in the streaming aggregation path.

@ryoppippi

Copy link
Copy Markdown
Member Author

@cubic-dev-ai review\n\nPlease review the latest push as well.

@cubic-dev-ai

cubic-dev-ai Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review\n\nPlease review the latest push as well.

@ryoppippi I have started the AI code review. It will take a few minutes to complete.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Jun 8, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide c7ffd05 Commit Preview URL

Branch Preview URL
Jun 08 2026, 10:28 PM

@coderabbitai

coderabbitai Bot commented Jun 8, 2026 •

Copy link
Copy Markdown
🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Fetch PR details and diff
gh pr view 1237 --repo ryoppippi/ccusage --json title,body,files,commits,headRefName 2>/dev/null | jq .

echo "=== PR DIFF ==="
gh pr diff 1237 --repo ryoppippi/ccusage 2>/dev/null

Length of output: 6051



@ryoppippi: I've reviewed the latest push. Here's the full analysis:


Review of fix(codex): dedupe copied rollout token events

Summary

The root cause was that CodexEventKey included session_id (hash + length), so identical token-count events copied into goal/rollout session files were treated as distinct because they carried a different session_id. Removing those two fields from the dedupe key means events with the same (timestamp, model, all 5 token counts) are now collapsed to one entry regardless of which session file they came from. This mirrors what the event-loader path already did.


aggregate.rs — Key changes

Old key (10 fields):

type CodexEventKey = (
    u64,        // hash_text(&event.session_id)
    usize,      // event.session_id.len()
    TimestampMs,
    u64,        // hash_text(model)
    usize,      // model.len()
    u64, u64, u64, u64, u64,  // token counts
);

New key (8 fields):

type CodexEventKey = (TimestampMs, u64, usize, u64, u64, u64, u64, u64);
//                   timestamp  model_hash  model_len  5×tokens

The tuple ordering in codex_event_key() and the type alias are consistent. ✅


mod.rs — Test rename + value correction

The existing test was asserting doubled counts (200, 100, 300), which was documenting the buggy behavior. The rename from keeps_matching_... → dedupes_matching_... and the corrected assertions (100, 50, 150) are correct. ✅


New test

dedupes_copied_token_usage_across_session_files precisely reproduces the reported scenario (same event line in sessions/root.jsonl and sessions/goal.jsonl) and asserts single-count values. Good targeted regression test. ✅


Minor observation (non-blocking)

The new key relies on millisecond-precision timestamp + identical model + identical token counts to identify a duplicate. There is a theoretical (negligible in practice) edge case: two genuinely independent LLM calls that finish at the exact same millisecond for the same model and produce bit-identical token counts would be incorrectly collapsed. Given that Codex sessions are typically sequential and token counts vary with each call, this is not a practical concern — but worth being aware of if the dedup strategy is ever revisited.


Verdict

The fix is correct, minimal, and well-tested. The streaming aggregation path is now consistent with the event-loader dedupe shape. No issues to block merge.

✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Jun 8, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 6672939f-b9b7-4a41-a756-040e4793600c

📥 Commits

Reviewing files that changed from the base of the PR and between d319394 and c7ffd05.

📒 Files selected for processing (1)
  • rust/crates/ccusage/src/adapter/codex/aggregate.rs

📝 Walkthrough

Walkthrough

Make Codex deduplication key kind-aware by passing AgentReportKind into codex_event_key/insert_event_key; session-derived key components are only used for AgentReportKind::Session. Update aggregation call sites and add tests validating dedupe across session files and session-isolated grouping.

Changes

Codex Event Deduplication by Report Kind

Layer / File(s) Summary
Kind-aware key and insertion calls
rust/crates/ccusage/src/adapter/codex/aggregate.rs
insert_event_key now accepts AgentReportKind; parallel and local aggregation call sites pass kind. codex_event_key(..., kind) includes session-derived hash/length only for AgentReportKind::Session and uses (0,0) otherwise.
Tests and expectations
rust/crates/ccusage/src/adapter/codex/aggregate.rs, rust/crates/ccusage/src/adapter/codex/mod.rs
Add dedupes_copied_token_usage_across_session_files and keeps_matching_token_usage_in_distinct_session_groups; update test name and expected aggregated token totals to reflect deduplication across session files.

Sequence Diagram

sequenceDiagram
  participant SessionFile
  participant Aggregator
  participant insert_event_key
  participant codex_event_key
  SessionFile->>Aggregator: read token-count event
  Aggregator->>insert_event_key: insert_event_key(event, kind)
  insert_event_key->>codex_event_key: codex_event_key(event, kind)
  codex_event_key->>insert_event_key: CodexEventKey(kind-aware)
  insert_event_key->>Aggregator: returns whether key is new (dedupe decision)
  Aggregator->>Aggregator: group/sum tokens based on date + CodexEventKey
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • ryoppippi/ccusage#1122: Modifies the Codex token-usage deduplication key and associated aggregation expectations.
  • ryoppippi/ccusage#1162: Updates the Codex grouped-token deduplication unit test and session-file setup for dedupe verification.
  • ryoppippi/ccusage#1156: Adjusts Codex deduplication semantics related to session-derived fields used in event fingerprinting.

Suggested labels

enhancement

Poem

🐰 I sniff the keys in session rows,
One kind keeps tracks, the other knows,
Copies fold into single sums,
While session-kind keeps separate drums,
Hoppity hop — the tokens show!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 55.56% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically identifies the main change: deduplicating token events in Codex aggregation, which aligns with the primary objective of removing duplicate token-count rows across session files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/goal-dedupe-codex-rollouts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@ryoppippi ryoppippi changed the title fix(codex): goal直した rollout token dedupe fix(codex): dedupe goal rollout token events Jun 8, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
rust/crates/ccusage/src/adapter/codex/aggregate.rs (1)

477-493: ⚡ Quick win

Cover both aggregation threading paths in this regression test.

Line 477 currently validates dedupe in a single runtime mode. Looping over single_thread in this test would lock the contract for both local and parallel aggregation paths.

♻️ Suggested test update
-        let shared = SharedArgs {
-            timezone: Some("UTC".to_string()),
-            ..SharedArgs::default()
-        };
-
-        let groups =
-            load_groups_from_directory(&fixture.path("sessions"), &shared, AgentReportKind::Daily)
-                .unwrap();
-
-        assert_eq!(groups.len(), 1);
-        let group = groups.get("2026-05-29").unwrap();
-        assert_eq!(group.input_tokens, 1_000);
-        assert_eq!(group.cached_input_tokens, 100);
-        assert_eq!(group.output_tokens, 200);
-        assert_eq!(group.reasoning_output_tokens, 20);
-        assert_eq!(group.total_tokens, 1_200);
+        for single_thread in [true, false] {
+            let shared = SharedArgs {
+                single_thread,
+                timezone: Some("UTC".to_string()),
+                ..SharedArgs::default()
+            };
+
+            let groups = load_groups_from_directory(
+                &fixture.path("sessions"),
+                &shared,
+                AgentReportKind::Daily,
+            )
+            .unwrap();
+
+            assert_eq!(groups.len(), 1);
+            let group = groups.get("2026-05-29").unwrap();
+            assert_eq!(group.input_tokens, 1_000);
+            assert_eq!(group.cached_input_tokens, 100);
+            assert_eq!(group.output_tokens, 200);
+            assert_eq!(group.reasoning_output_tokens, 20);
+            assert_eq!(group.total_tokens, 1_200);
+        }

As per coding guidelines, "Add Rust fixture-backed tests for path discovery, parser behavior, aggregation totals, and important legacy compatibility."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust/crates/ccusage/src/adapter/codex/aggregate.rs` around lines 477 - 493,
The test currently only exercises one aggregation threading mode; update it to
loop over both local and parallel aggregation paths by iterating over SharedArgs
{ single_thread: true/false, timezone: Some("UTC".into()),
..SharedArgs::default() } (or setting single_thread on the existing `shared`),
and for each iteration call
`load_groups_from_directory(&fixture.path("sessions"), &shared,
AgentReportKind::Daily)` and run the same assertions on the resulting `groups`
and `group` (checks for `input_tokens`, `cached_input_tokens`, `output_tokens`,
`reasoning_output_tokens`, and `total_tokens`) so both aggregation threads are
validated.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@rust/crates/ccusage/src/adapter/codex/aggregate.rs`:
- Around line 477-493: The test currently only exercises one aggregation
threading mode; update it to loop over both local and parallel aggregation paths
by iterating over SharedArgs { single_thread: true/false, timezone:
Some("UTC".into()), ..SharedArgs::default() } (or setting single_thread on the
existing `shared`), and for each iteration call
`load_groups_from_directory(&fixture.path("sessions"), &shared,
AgentReportKind::Daily)` and run the same assertions on the resulting `groups`
and `group` (checks for `input_tokens`, `cached_input_tokens`, `output_tokens`,
`reasoning_output_tokens`, and `total_tokens`) so both aggregation threads are
validated.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 66fe19dd-9d96-4609-964e-72d5ba2dbfc6

📥 Commits

Reviewing files that changed from the base of the PR and between e36a3c2 and 88ee6e1.

📒 Files selected for processing (2)
  • rust/crates/ccusage/src/adapter/codex/aggregate.rs
  • rust/crates/ccusage/src/adapter/codex/mod.rs

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes — removes session_id from the streaming aggregation dedupe key so identical token_count rows copied across root and goal Codex session files are counted once instead of once per file.

  • Align codex_event_key with loader dedupe shape — the streaming CodexEventKey tuple in aggregate.rs now starts with (timestamp, hash_text(model), ...) instead of (hash_text(&event.session_id), event.session_id.len(), timestamp, ...). The loader in loader.rs already deduped on (timestamp, model, input_tokens, ...) without session_id; the streaming path was the outlier.

  • Update existing test and add root/goal dedupe test — keeps_matching_grouped_codex_usage_events_from_distinct_sessions in mod.rs is renamed to dedupes_... and its assertions updated from 2x to 1x totals. New test dedupes_copied_token_usage_across_session_files in aggregate.rs places the same event in root.jsonl and goal.jsonl and asserts deduped single-event totals.

Pullfrog  | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

Extend the copied rollout token regression test across both local and parallel Codex aggregation paths.

This addresses CodeRabbit feedback and keeps the dedupe contract covered for each SharedArgs single_thread mode.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Re-trigger cubic

@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai Addressed your test coverage feedback in d319394 by exercising the copied rollout token dedupe regression across both single-threaded and parallel aggregation paths. Validation: direnv exec . cargo test -p ccusage adapter::codex -- --nocapture; pre-push clippy/treefmt/gitleaks/cargo test passed. Please review the latest push.

@ryoppippi

Copy link
Copy Markdown
Member Author

@cubic-dev-ai Please review the latest push d319394 as well.

@cubic-dev-ai

cubic-dev-ai Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai Please review the latest push d319394 as well.

@ryoppippi I have started the AI code review. It will take a few minutes to complete.

@coderabbitai

coderabbitai Bot commented Jun 8, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Fetch the latest commit details
gh pr view 1237 --repo ryoppippi/ccusage --json commits | jq '.commits[-3:]'

echo "=== LATEST COMMIT DIFF ==="
gh api repos/ryoppippi/ccusage/commits/d3193940 --jq '.files[] | "--- \(.filename) ---\n\(.patch)"'

Length of output: 3513



@ryoppippi: Reviewed commit d3193940 — looks great.

test(codex): cover rollout dedupe threading modes

The change wraps the entire dedupes_copied_token_usage_across_session_files test body in a for single_thread in [true, false] loop and threads the flag through SharedArgs, so the dedupe contract is now exercised against both the sequential path (single_thread = true) and the rayon parallel path (single_thread = false). All five token-count assertions remain at their correct single-count values (1_000 / 100 / 200 / 20 / 1_200) for each iteration. The change is minimal, clean, and exactly what was needed. ✅

No issues — ready to merge.

(๑˃ᴗ˂)ﻭ

@pkg-pr-new

pkg-pr-new Bot commented Jun 8, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ccusage@1237

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-arm64@1237

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-x64@1237

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/@ccusage/ccusage-linux-arm64@1237

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/@ccusage/ccusage-linux-x64@1237

@ccusage/ccusage-win32-arm64

npx https://pkg.pr.new/@ccusage/ccusage-win32-arm64@1237

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/@ccusage/ccusage-win32-x64@1237

commit: c7ffd05

@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 88ee6e1e1dc3
Base SHA: e36a3c22db70

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new e36a3c22db70 832.9ms 655.4ms 33.6ms 3
PR pkg.pr.new 88ee6e1 504.0ms 628.6ms 32.7ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: e36a3c22db70; PR package: 88ee6e1. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 575.0ms 586.0ms 0.98x 331.20 MiB 320.45 MiB 0.97x 1.75 GiB/s 1.72 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 378.3ms 354.0ms 1.07x 79.70 MiB 57.20 MiB 0.72x 2.66 GiB/s 2.84 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 565.5ms 1.78 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 525.3ms 1.92 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 351.5ms 2.86 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 329.9ms 3.05 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 31.1ms 4.2ms 7.42x 43.61 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.37 MiB/s
claude session --offline --json 0.00 MiB 30.5ms 4.3ms 7.16x 43.48 MiB 2.83 MiB 0.07x 0.05 MiB/s 0.36 MiB/s
codex daily --offline --json 0.00 MiB 30.5ms 3.9ms 7.91x 43.61 MiB 2.70 MiB 0.06x 0.03 MiB/s 0.22 MiB/s
codex session --offline --json 0.00 MiB 29.9ms 3.9ms 7.73x 43.48 MiB 2.70 MiB 0.06x 0.03 MiB/s 0.22 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 551.4ms 539.4ms 1.02x 334.58 MiB 326.45 MiB 0.98x 1.83 GiB/s 1.87 GiB/s
codex --offline --json 1.01 GiB 379.2ms 325.9ms 1.16x - 55.83 MiB - 2.65 GiB/s 3.09 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 16.83 KiB 16.83 KiB +0.00 KiB 1.00x
installed native package binary 3353.74 KiB 3353.74 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 88ee6e1e1dc3
Base SHA: e36a3c22db70

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new e36a3c22db70 646.3ms 540.4ms 29.1ms 3
PR pkg.pr.new 88ee6e1 630.4ms 651.2ms 30.4ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: e36a3c22db70; PR package: 88ee6e1. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 552.7ms 544.2ms 1.02x 318.33 MiB 331.83 MiB 1.04x 1.82 GiB/s 1.85 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 371.1ms 352.8ms 1.05x 72.33 MiB 67.58 MiB 0.93x 2.71 GiB/s 2.85 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 533.3ms 1.89 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 518.0ms 1.94 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 345.1ms 2.92 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 324.9ms 3.10 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 28.4ms 28.2ms 1.01x 43.61 MiB 43.73 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 28.8ms 29.1ms 0.99x 43.48 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 28.8ms 28.4ms 1.01x 43.48 MiB 43.73 MiB 1.01x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 28.7ms 28.0ms 1.02x 43.61 MiB 43.61 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 551.2ms 543.7ms 1.01x 323.95 MiB 322.95 MiB 1.00x 1.83 GiB/s 1.85 GiB/s
codex --offline --json 1.01 GiB 365.0ms 346.0ms 1.06x 77.33 MiB 61.70 MiB 0.80x 2.76 GiB/s 2.91 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 16.83 KiB 16.83 KiB +0.00 KiB 1.00x
installed native package binary 3353.74 KiB 3353.74 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes — wraps the new rollout dedupe test in for single_thread in [true, false] so both sequential and parallel aggregation paths are validated.

  • Cover both threading modes in rollout dedupe test — dedupes_copied_token_usage_across_session_files now iterates over single_thread: true and single_thread: false, passing single_thread into SharedArgs so the dedupe contract is validated for both the local and parallel aggregation paths.

Pullfrog  | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 2 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread rust/crates/ccusage/src/adapter/codex/aggregate.rs Outdated
@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: d31939405495
Base SHA: e36a3c22db70

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new e36a3c22db70 531.8ms 896.6ms 31.9ms 3
PR pkg.pr.new d319394 1.424s 834.5ms 32.1ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: e36a3c22db70; PR package: d319394. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 564.1ms 561.5ms 1.00x 316.58 MiB 319.83 MiB 1.01x 1.78 GiB/s 1.79 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 371.0ms 358.5ms 1.04x 82.83 MiB 59.70 MiB 0.72x 2.71 GiB/s 2.81 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 560.0ms 1.80 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 518.9ms 1.94 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 353.8ms 2.85 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 324.1ms 3.11 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.4ms 3.9ms 7.54x 43.61 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.40 MiB/s
claude session --offline --json 0.00 MiB 29.4ms 4.0ms 7.35x 43.61 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.39 MiB/s
codex daily --offline --json 0.00 MiB 29.5ms 3.7ms 8.01x 43.48 MiB 2.70 MiB 0.06x 0.03 MiB/s 0.23 MiB/s
codex session --offline --json 0.00 MiB 29.5ms 3.8ms 7.77x 43.48 MiB 2.70 MiB 0.06x 0.03 MiB/s 0.23 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 551.5ms 534.7ms 1.03x 337.45 MiB 331.95 MiB 0.98x 1.83 GiB/s 1.88 GiB/s
codex --offline --json 1.01 GiB 375.0ms 330.6ms 1.13x 81.20 MiB 66.83 MiB 0.82x 2.68 GiB/s 3.04 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 16.83 KiB 16.83 KiB +0.00 KiB 1.00x
installed native package binary 3353.74 KiB 3353.74 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: d31939405495
Base SHA: e36a3c22db70

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new e36a3c22db70 576.2ms 749.3ms 30.0ms 3
PR pkg.pr.new d319394 663.9ms 907.0ms 31.6ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: e36a3c22db70; PR package: d319394. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 567.2ms 561.8ms 1.01x 321.45 MiB 316.08 MiB 0.98x 1.78 GiB/s 1.79 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 373.5ms 361.4ms 1.03x 82.20 MiB 54.70 MiB 0.67x 2.70 GiB/s 2.79 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 556.2ms 1.81 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 518.8ms 1.94 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 358.7ms 2.81 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 335.0ms 3.01 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.8ms 29.2ms 1.02x 43.73 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 30.0ms 29.5ms 1.02x 43.48 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 29.5ms 29.7ms 0.99x 43.48 MiB 43.48 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 29.4ms 29.5ms 1.00x 43.61 MiB 43.61 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 564.3ms 542.7ms 1.04x 285.45 MiB 324.33 MiB 1.14x 1.78 GiB/s 1.86 GiB/s
codex --offline --json 1.01 GiB 377.8ms 355.6ms 1.06x 82.33 MiB 52.20 MiB 0.63x 2.66 GiB/s 2.83 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 16.83 KiB 16.83 KiB +0.00 KiB 1.00x
installed native package binary 3353.74 KiB 3353.74 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

Keep Codex rollout token dedupe across files for daily, weekly, and monthly reports while restoring session-scoped dedupe keys for session reports.

The previous dedupe key intentionally removed session_id to collapse copied root/goal rollout token_count rows, but that same global key ran before session grouping and could drop a distinct session bucket when two session files had identical token usage rows. Session reports now include the session hash in the key, while aggregate period reports use a zero session component to retain cross-session dedupe.

Adds a regression test that identical token usage rows in two session files remain visible as two session groups across both single-threaded and parallel aggregation.
@ryoppippi

Copy link
Copy Markdown
Member Author

@cubic-dev-ai Please re-review c7ffd05. I addressed your session-report undercount finding by preserving session_id only for AgentReportKind::Session dedupe keys and added coverage for identical token rows in distinct session groups across both aggregation threading modes.

@cubic-dev-ai

cubic-dev-ai Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai Please re-review c7ffd05. I addressed your session-report undercount finding by preserving session_id only for AgentReportKind::Session dedupe keys and added coverage for identical token rows in distinct session groups across both aggregation threading modes.

@ryoppippi I have started the AI code review. It will take a few minutes to complete.

@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai review

Pushed c7ffd05 to address the session-report undercount finding: session reports now keep session-scoped dedupe keys, while daily/weekly/monthly reports still dedupe copied rollout token events across session files. Validation passed locally: direnv exec . cargo test --manifest-path rust/Cargo.toml -p ccusage adapter::codex -- --nocapture; direnv exec . pnpm run format; direnv exec . pnpm run test; and pre-push clippy/treefmt/gitleaks/cargo test.

@coderabbitai

coderabbitai Bot commented Jun 8, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Fetch the latest commit details
gh pr view 1237 --repo ryoppippi/ccusage --json commits | jq '.commits[-3:]'

echo "=== LATEST COMMIT DIFF ==="
gh api repos/ryoppippi/ccusage/commits/c7ffd05 --jq '.files[] | "--- \(.filename) ---\n\(.patch)"'

Length of output: 6865



@ryoppippi: Reviewed commit c7ffd05 — clean and correct.

fix(codex): preserve session report dedupe scope

Key design:

Report kind session_hash / session_len in key Effect
Session real hash_text(session_id) / session_id.len() Distinct session buckets are preserved even with identical token rows
Daily / Weekly / Monthly 0 / 0 (sentinel) Cross-file dedup from the previous fix is retained

No collision risk from the zero-sentinel: a real session always has len > 0, the sentinel always has len == 0. ✅

Both aggregation paths updated consistently:

  • Parallel path: insert_event_key(event, timestamp, model, kind, seen) ✅
  • Local path: codex_event_key(event, timestamp, model, kind) ✅

New regression test keeps_matching_token_usage_in_distinct_session_groups:

  • Two session files with the same token-usage line, loaded with AgentReportKind::Session
  • Asserts groups.len() == 2 and each group has the correct input_tokens == 1_000 (not collapsed)
  • Covers both single_thread = true and false ✅

The fix is minimal, the sentinel approach is elegant, and the test suite now covers both the cross-session-dedup path (non-session reports) and the session-preservation path (session reports). No blockers — ready to merge.

(˶˃ᵕ˂˶) ✨

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Re-trigger cubic

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes — the new commit preserves session-report granularity by adding session_id back to the dedupe key for AgentReportKind::Session, while Daily/Weekly/Monthly reports continue deduplicating identical token events across session files.

  • Conditional session-scoped dedupe key — codex_event_key() now takes AgentReportKind and prepends (session_hash, session_len) only for Session kind; non-session kinds use (0, 0) sentinel values. insert_event_key(), add_event_to_groups(), and add_event_to_groups_local() thread kind through to preserve both behaviors.
  • New session-granularity test — keeps_matching_token_usage_in_distinct_session_groups places identical events in root.jsonl and goal.jsonl with AgentReportKind::Session, asserts 2 distinct groups each with full token totals, and loops over both threading modes.

Pullfrog  | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: c7ffd05c10da
Base SHA: e36a3c22db70

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new e36a3c22db70 598.7ms 529.7ms 32.8ms 3
PR pkg.pr.new c7ffd05 1.006s 553.6ms 32.6ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: e36a3c22db70; PR package: c7ffd05. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 556.6ms 578.5ms 0.96x 315.95 MiB 325.33 MiB 1.03x 1.81 GiB/s 1.74 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 385.9ms 374.8ms 1.03x 80.20 MiB 70.20 MiB 0.88x 2.61 GiB/s 2.69 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 566.3ms 1.78 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 529.3ms 1.90 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 371.9ms 2.71 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 341.5ms 2.95 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.9ms 4.1ms 7.21x 43.48 MiB 2.70 MiB 0.06x 0.05 MiB/s 0.37 MiB/s
claude session --offline --json 0.00 MiB 30.2ms 4.3ms 7.11x 43.61 MiB 2.83 MiB 0.06x 0.05 MiB/s 0.36 MiB/s
codex daily --offline --json 0.00 MiB 30.3ms 4.2ms 7.27x 43.48 MiB 2.70 MiB 0.06x 0.03 MiB/s 0.21 MiB/s
codex session --offline --json 0.00 MiB 29.5ms 3.8ms 7.75x 43.61 MiB 2.70 MiB 0.06x 0.03 MiB/s 0.23 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 577.0ms 541.8ms 1.07x - 316.83 MiB - 1.74 GiB/s 1.86 GiB/s
codex --offline --json 1.01 GiB 377.4ms 346.6ms 1.09x 79.83 MiB 70.08 MiB 0.88x 2.67 GiB/s 2.90 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 16.83 KiB 16.83 KiB +0.00 KiB 1.00x
installed native package binary 3353.74 KiB 3353.74 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: c7ffd05c10da
Base SHA: e36a3c22db70

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new e36a3c22db70 386.9ms 433.3ms 33.3ms 3
PR pkg.pr.new c7ffd05 732.4ms 498.2ms 32.5ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: e36a3c22db70; PR package: c7ffd05. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 573.3ms 577.3ms 0.99x 298.70 MiB 335.08 MiB 1.12x 1.76 GiB/s 1.74 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 395.7ms 382.4ms 1.03x 81.20 MiB 75.70 MiB 0.93x 2.54 GiB/s 2.63 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 590.1ms 1.71 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 562.7ms 1.79 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 386.4ms 2.61 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 352.9ms 2.85 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 30.3ms 30.4ms 1.00x 43.73 MiB 43.61 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 30.2ms 29.8ms 1.01x 43.61 MiB - - 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 30.5ms 30.3ms 1.01x 43.48 MiB 43.73 MiB 1.01x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 29.7ms 30.3ms 0.98x 43.48 MiB 43.61 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 650.5ms 568.4ms 1.14x 329.83 MiB 313.08 MiB 0.95x 1.55 GiB/s 1.77 GiB/s
codex --offline --json 1.01 GiB 377.9ms 374.7ms 1.01x 79.95 MiB 78.45 MiB 0.98x 2.66 GiB/s 2.69 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 16.83 KiB 16.83 KiB +0.00 KiB 1.00x
installed native package binary 3353.74 KiB 3353.74 KiB +0.00 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@ryoppippi
ryoppippi merged commit 8a52395 into main Jun 8, 2026
40 checks passed
@ryoppippi
ryoppippi deleted the codex/goal-dedupe-codex-rollouts branch June 8, 2026 22:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant