Skip to content

feat(grok): add Grok Build CLI usage adapter - #1593

Merged
ryoppippi merged 17 commits into
mainfrom
fix/grok-adapter
Aug 10, 2026
Merged

ryoppippi merged 17 commits into
mainfrom
fix/grok-adapter

Conversation

@ryoppippi

@ryoppippi ryoppippi commented Aug 10, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds the Grok Build CLI as a data source: ccusage grok daily|monthly|session, plus Grok rows in the unified reports.

This builds on @LeeShunEE's work in #1520 (branch kept intact in the history, then merged with main) and corrects three things that only showed up when the output was checked against real ~/.grok logs.

Closes #1513. Supersedes #1520, #1447, #1202.

What changed on top of #1520

Cost now comes from Grok's own costUsdTicks. Every completed turn records it as fixed-point USD, one tick being 1e-10 USD. The adapter discarded it and recomputed from the pricing table. Grok bills each API request separately, but a turn_completed row only carries the sum over the requests in that turn, so recomputation cannot place the long-context tier boundary where Grok did — it lands on either side of the real figure depending on how the turn split:

Date Recomputed Actual
2026-07-19 $5.2822 $5.9794
2026-07-20 $1.4215 $1.2639
2026-08-04 $2.7632 $5.1217
2026-08-09 $2.0994 $1.8515
2026-08-10 $3.6111 $3.1858
Total $15.1774 $17.4023

display and the default auto now report the invoice figure. calculate still recomputes, and auto falls back to the pricing table for turns that recorded no ticks (Grok only started writing usage on 2026-07-19).

Reasoning tokens are no longer counted twice. They were added to extra_total_tokens, the bucket for tokens not already covered by the usage sum. Grok reports totalTokens == inputTokens + outputTokens and its reasoning count never exceeds output, so reasoning is a subset of output. Over 85 local turns this inflated the grand total by 113,556 tokens.

cacheCreationTokens is read. Grok added the field around CLI 0.2.118 and still writes it in 1.0.0; the adapter hardcoded cache_creation_input_tokens: 0. Every observed turn reports zero, so there is no sample proving whether it sits inside inputTokens or beside it — it is carved out of the same total, matching the proven behaviour of cachedReadTokens and keeping the parts summing back to inputTokens.

Two smaller items: the merge with main left the adapter-count array in agent_commands_are_exposed_by_independent_crates at 14 while there are now 15, and the turn_line test helper exceeded the clippy argument limit, so its token counts moved into a struct.

Testing

  • cargo test --workspace — 573 passed, 0 failed
  • cargo clippy --workspace --all-targets -- -D warnings — clean
  • treefmt — clean
  • End-to-end against a real ~/.grok (85 turns across Grok CLI 0.2.103 through 1.0.0):
Raw logs ccusage grok daily
Total tokens 32,371,143 32,371,143
Input (uncached) 1,548,986 1,548,986
Cache read 30,639,232 30,639,232
Output 182,925 182,925
Cost $17.4023 $17.4023

The tick scale was pinned on a single-request turn, where no tiering applies: 7,180 uncached x $2.00/M + 11,264 cached x $0.30/M + 130 output x $6.00/M = $0.0185192, against a recorded 185,192,000 ticks. All 58 turns from CLI 1.0.0 reproduce the xAI list price for xai/grok-4.5 exactly, which also confirms that grok-4.5-build bills at grok-4.5 rates rather than the separately-listed grok-build-0.1.

Known limits, documented rather than fixed

  • A session killed mid-turn never writes turn_completed, so its usage cannot be reported. logs/unified.jsonl records the underlying requests but carries no per-request model id, so it cannot be priced or attributed and is not used as a source.
  • Turns written before 2026-07-19 carry no usage at all.

View with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is enabled.


Summary by cubic

Adds Grok Build CLI as a first-class data source with daily, monthly, and session reports in unified views. Uses Grok’s recorded costUsdTicks for invoice-accurate costs and decodes UTF‑8 project paths.

  • New Features

    • New adapter ccusage-adapter-grok parsing sessions/**/updates.jsonl under GROK_HOME or ~/.grok (env-only, no path flag).
    • New commands: ccusage grok daily, monthly, session; Grok also appears in unified daily|monthly|session.
    • Costs use costUsdTicks; auto falls back when ticks are missing; calculate still recomputes.
    • Reads cacheCreationTokens and keeps parts summing to input tokens.
    • CLI/parser, grok config schema, and docs added.
  • Bug Fixes

    • Decode percent-encoded UTF‑8 in project paths; invalid bytes degrade to the replacement character; session reports show correct paths.
    • Cross-file dedupe now uses only eventId; entries without one are left as-is so distinct turns aren’t collapsed.
    • Stop double counting reasoning tokens; session reports include first/last activity and project path.
    • Docs clarify GROK_HOME is a single root (no comma-separated paths).

Written for commit 0a5ab84. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • New Features

    • Added Grok Build CLI as a supported data source.
    • Added daily, monthly, session, and unified usage reports.
    • Added token, cache, model, and cost reporting with pricing options.
    • Added automatic session-log discovery through GROK_HOME or the default data directory.
    • Added Grok-specific configuration and command-line support, including date and recent-usage filtering.
  • Documentation

    • Added setup, configuration, environment-variable, troubleshooting, and reporting guidance for Grok Build CLI.

LeeShunEE and others added 11 commits July 28, 2026 18:02
Parse completed Grok Build turns, map token usage and LiteLLM pricing, and expose focused and unified reports. Discover data from non-empty GROK_HOME or ~/.grok, add schema and CLI integration, and document the source.
Session reports grouped by key then moved the group key into session_id,
which left first_activity, last_activity and project_path unset. Accumulate
with SessionAccumulator like qwen/opencode/antigravity, and include session
metadata in the JSON report.
Add fixture-backed tests for path discovery edge cases, token/pricing
mapping, timestamp and summary metadata, cross-session dedupe, corrupt
file isolation, and session aggregation. Synthesize updates.jsonl rather
than committing real session trees.
…usage-adapter

# Conflicts:
#	apps/ccusage/README.md
#	docs/guide/index.md
#	docs/guide/source-support-qa.md
#	rust/crates/ccusage-adapter-all/src/loader.rs
#	rust/crates/ccusage-cli-parser/src/parser.rs
#	rust/crates/ccusage-cli-parser/src/snapshots/ccusage_cli_parser__tests__root_help.snap
#	rust/crates/ccusage-cli-parser/src/tests.rs
#	rust/crates/ccusage-config/src/config_schema.rs
#	rust/crates/ccusage-config/src/snapshots/ccusage_config__config_schema__tests__snapshots_schema_agent_specific_option_edges.snap
#	rust/crates/ccusage-core/src/lib.rs
#	rust/crates/ccusage/src/cli/last_window.rs
…usage-adapter

Rebase onto current upstream main (24cbb82). Resolve conflicts so Grok stays
wired while respecting the Antigravity revert (#1569).
…tals

Two corrections to the Grok adapter, both found by comparing its output
against real ~/.grok logs.

Cost: every completed turn records `costUsdTicks`, a fixed-point USD amount
where one tick is 1e-10 USD. The adapter discarded it and recomputed from the
pricing table instead. Grok bills each API request separately, but a
`turn_completed` row only carries the sum over the requests in that turn, so a
recomputation cannot place the long-context tier boundary where Grok did and
lands on either side of the real figure. Across 85 turns of local data the
recomputed total is $15.18 against an actual $17.40, and per-day it ranges from
-46% to +12%. Feeding the ticks through `cost_usd` reproduces the invoice
exactly in `display` and `auto`, while `calculate` keeps recomputing and `auto`
still falls back for turns that recorded no ticks.

Tokens: `reasoningTokens` were added to `extra_total_tokens`, which is the
bucket for tokens *not* already covered by the usage sum. Grok reports
`totalTokens == inputTokens + outputTokens` and its reasoning count never
exceeds output, so reasoning is a subset of output and was being counted twice.
That inflated the grand total by 113,556 tokens over the same 85 turns.

The `turn_line` test helper moves its four token counts into a struct so it
stays under the clippy argument limit.
Merging main brought the adapter count to 15 while the array annotation in the
`agent_commands_are_exposed_by_independent_crates` test still said 14, so the
test crate stopped compiling. Also runs treefmt over the two files the grok
branch left unformatted, which the clippy and treefmt checks require.
Grok started writing `cacheCreationTokens` alongside `cachedReadTokens` around
CLI 0.2.118 and still writes it in 1.0.0. The adapter hardcoded
`cache_creation_input_tokens: 0`, so the field would be dropped the moment it
went non-zero.

Every turn observed so far reports it as zero, so there is no sample proving
whether it sits inside `inputTokens` or beside it. `cachedReadTokens` is
provably inside — session totals match `logs/unified.jsonl`, where
`cached_prompt_tokens` is part of `prompt_tokens` — so the sibling field is
carved out of the same total. That keeps the three parts summing back to
`inputTokens`, which is the invariant the adapter already relies on.

Also documents two limits that surfaced while checking real logs: a session
killed mid-turn never writes `turn_completed`, so its usage cannot be reported
at all, and `logs/unified.jsonl` cannot stand in for it because it records no
per-request model id.
@coderabbitai

coderabbitai Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds Grok Build CLI as a supported ccusage source. The change includes session discovery, completed-turn parsing, token and cost accounting, report generation, CLI and configuration integration, unified loading, tests, and documentation.

Changes

Grok support

Layer / File(s) Summary
Grok session ingestion and reports
rust/adapters/grok/*
Adds environment-based discovery, JSONL parsing, deduplication, token accounting, pricing fallback, cost handling, and daily, weekly, monthly, and session reports.
Grok CLI and configuration contracts
rust/crates/ccusage-cli-parser/*, rust/crates/ccusage-cli/src/types.rs, rust/crates/ccusage-config/*, rust/crates/ccusage-core/src/lib.rs, apps/ccusage/config-schema.json
Adds the grok command, report routes, Command::Grok, shared configuration, command-specific configuration, and built-in agent registration.
Unified loading and executable wiring
rust/crates/ccusage-adapter-all/*, rust/crates/ccusage/*, rust/Cargo.toml, rust/crates/*/Cargo.toml, nix/cargo-artifacts.nix
Registers the adapter in workspace dependencies, unified loading, report labeling, application dispatch, --last handling, build artifacts, and integration tests.
Grok support documentation
docs/guide/*, docs/.vitepress/config.ts, apps/ccusage/README.md, rust/adapters/grok/README.md, docs/superpowers/specs/*
Documents commands, data paths, GROK_HOME, configuration, report behavior, troubleshooting, adapter behavior, and supported-source status.

Estimated code review effort: 5 (Critical) | ~120 minutes

Sequence Diagram(s)

sequenceDiagram
  participant User
  participant CcusageCLI
  participant GrokAdapter
  participant GrokSessionFiles
  participant UnifiedReports
  User->>CcusageCLI: run ccusage grok daily/monthly/session
  CcusageCLI->>GrokAdapter: dispatch Command::Grok
  GrokAdapter->>GrokSessionFiles: discover updates.jsonl files
  GrokSessionFiles-->>GrokAdapter: return session files
  GrokAdapter->>GrokAdapter: parse completed turns and calculate costs
  GrokAdapter-->>CcusageCLI: return formatted or JSON report
  UnifiedReports->>GrokAdapter: load Grok entries
  GrokAdapter-->>UnifiedReports: return summarized usage rows
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 74.11% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The changes implement the requested Grok commands, unified reports, parsing, configuration, cost handling, integration, tests, and documentation for issue #1513.
Out of Scope Changes check ✅ Passed The changes are focused on Grok Build CLI support and its required adapter, integration, tests, configuration, documentation, and design details.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely identifies the addition of the Grok Build CLI usage adapter, which is the pull request's main change.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/grok-adapter

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide 0a5ab84 Commit Preview URL

Branch Preview URL
Aug 10 2026, 03:35 PM

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • Grok adapter crate — rust/adapters/grok/: token mapping (inputTokens split into uncached/cache-read/cache-creation), costUsdTicks invoice-cost path with display/auto/calculate modes, and pricing candidate resolution (grok-4.5-build → xai/grok-4.5, etc.)
  • CLI wiring — standard agent pattern: Command::Grok variant in types.rs, parse_basic_agent_command in parser.rs, root-help and parse-shape snapshots, --last and config dispatch
  • Adapter-all integration — AgentLoadSpec at index 15, agent_label("grok") → "Grok", unified-row tests with GROK_HOME fixture
  • Config schema — GrokConfig/GrokCommandsConfig structs, schema JSON regeneration, source-specific option isolation tests
  • Docs — new docs/guide/grok/index.md, updated agent lists in 6 other doc pages, GROK_HOME in env-variables table, source-support-qa moved Grok from unsupported to supported, env-only path discovery design doc

Pullfrog  | ⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to @v0 or keep the SHA fresh with Dependabot | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (2)
rust/adapters/grok/src/loader.rs (1)

74-76: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick win

has_data does not short-circuit.

discover_session_files walks the whole sessions tree, collects every .jsonl path, filters the list, and sorts it. has_data then only checks whether the result is non-empty. On a large Grok home this reads far more directory entries than needed.

Add a detection helper in rust/adapters/grok/src/paths.rs that returns as soon as it finds the first updates.jsonl, and call it here.

♻️ Proposed fix
// rust/adapters/grok/src/paths.rs
/// Return `true` as soon as one `sessions/**/updates.jsonl` exists.
pub(super) fn has_session_file() -> bool {
    let Some(root) = resolve_root() else {
        return false;
    };
    let sessions = root.join("sessions");
    // Walk lazily and stop at the first match instead of collecting every path.
    // ...
}
 pub fn has_data() -> bool {
-    discover_session_files().is_ok_and(|files| !files.is_empty())
+    super::paths::has_session_file()
 }

As per coding guidelines: "Implement fast detection that short-circuits once a usable source file is found."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust/adapters/grok/src/loader.rs` around lines 74 - 76, Replace the
`discover_session_files` call in `has_data` with a new `has_session_file` helper
from `paths.rs`. Implement `has_session_file` to resolve the root, walk
`sessions` lazily, and return true immediately upon finding any `updates.jsonl`,
returning false when the root is unavailable or no match exists.

Source: Coding guidelines

rust/adapters/grok/src/parser.rs (1)

320-320: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Remove the no-op read of total_tokens.

let _ = model_usage.total_tokens; has no effect. model_usage_rows already populates the field from usage.total_tokens, so the statement is not needed to suppress a dead-field warning. If the field is intentionally unused for accounting, state that in a comment instead.

♻️ Proposed cleanup
-            let _ = model_usage.total_tokens;
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust/adapters/grok/src/parser.rs` at line 320, Remove the no-op
`model_usage.total_tokens` read near the model usage handling;
`model_usage_rows` already populates this field, so no replacement is needed
unless an intentional accounting decision must be documented with a comment.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/guide/environment-variables.md`:
- Line 26: Update the directory-variable guidance in
docs/guide/environment-variables.md:26-26 to explicitly exclude GROK_HOME from
comma-separated multi-root behavior. Also revise the custom-path guidance in
docs/guide/getting-started.md:200-200 so GROK_HOME is not included in the
multi-directory statement; preserve its single-root contract.

In `@rust/adapters/grok/README.md`:
- Around line 63-69: Update the “Public surface” list in the README to remove
report::report_from_rows, leaving only the publicly accessible symbols
loader::has_data, loader::load_entries, report::summarize_entries, and run.

In `@rust/adapters/grok/src/loader.rs`:
- Around line 48-69: Align the fallback deduplication key in the loader’s
entries.retain block with parse_session_files by using reasoning tokens as its
final discriminator instead of entry.extra_total_tokens. If reasoning is not
available on LoadedEntry, add and populate that field through the loader path,
then use it consistently in both keys while preserving the existing event-ID key
behavior.

In `@rust/adapters/grok/src/parser.rs`:
- Around line 498-516: Update url_decode_lightweight to accumulate decoded
percent bytes and raw UTF-8 bytes in a byte buffer, then construct the result
once with String::from_utf8_lossy (or reuse an established percent-decoding
crate) so UTF-8 project paths remain intact. Add a test covering a UTF-8 project
segment such as %C3%A9 and verify it produces é rather than mojibake.

In `@rust/adapters/grok/src/report.rs`:
- Around line 83-126: Update the Grok test helper entry to set
extra_total_tokens to zero, matching parse_session_files’ production contract
where reasoning tokens are included in outputTokens. Adjust the affected
daily-total and session assertions to expect 120 and 0 respectively, and prefer
fixture-backed parser/loader coverage for these Rust tests where applicable.

---

Nitpick comments:
In `@rust/adapters/grok/src/loader.rs`:
- Around line 74-76: Replace the `discover_session_files` call in `has_data`
with a new `has_session_file` helper from `paths.rs`. Implement
`has_session_file` to resolve the root, walk `sessions` lazily, and return true
immediately upon finding any `updates.jsonl`, returning false when the root is
unavailable or no match exists.

In `@rust/adapters/grok/src/parser.rs`:
- Line 320: Remove the no-op `model_usage.total_tokens` read near the model
usage handling; `model_usage_rows` already populates this field, so no
replacement is needed unless an intentional accounting decision must be
documented with a comment.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 6a9d153e-c3e7-497c-8499-cdaae4f5ccb8

📥 Commits

Reviewing files that changed from the base of the PR and between a6ea444 and 93c7bef.

⛔ Files ignored due to path filters (4)
  • rust/Cargo.lock is excluded by !**/*.lock
  • rust/crates/ccusage-cli-parser/src/snapshots/ccusage_cli_parser__tests__root_help.snap is excluded by !**/*.snap
  • rust/crates/ccusage-cli-parser/src/snapshots/ccusage_cli_parser__tests__snapshots_representative_cli_parse_shapes.snap is excluded by !**/*.snap
  • rust/crates/ccusage-config/src/snapshots/ccusage_config__config_schema__tests__snapshots_schema_agent_specific_option_edges.snap is excluded by !**/*.snap
📒 Files selected for processing (38)
  • apps/ccusage/README.md
  • apps/ccusage/config-schema.json
  • docs/.vitepress/config.ts
  • docs/guide/all-reports.md
  • docs/guide/config-files.md
  • docs/guide/environment-variables.md
  • docs/guide/getting-started.md
  • docs/guide/grok/index.md
  • docs/guide/index.md
  • docs/guide/source-support-qa.md
  • docs/superpowers/specs/2026-07-28-grok-env-only-path-design.md
  • rust/Cargo.toml
  • rust/adapters/grok/Cargo.toml
  • rust/adapters/grok/README.md
  • rust/adapters/grok/src/lib.rs
  • rust/adapters/grok/src/loader.rs
  • rust/adapters/grok/src/parser.rs
  • rust/adapters/grok/src/paths.rs
  • rust/adapters/grok/src/report.rs
  • rust/crates/ccusage-adapter-all/Cargo.toml
  • rust/crates/ccusage-adapter-all/README.md
  • rust/crates/ccusage-adapter-all/src/lib.rs
  • rust/crates/ccusage-adapter-all/src/loader.rs
  • rust/crates/ccusage-adapter-all/src/report.rs
  • rust/crates/ccusage-adapter-all/src/tests.rs
  • rust/crates/ccusage-cli-parser/src/cli-commands.json
  • rust/crates/ccusage-cli-parser/src/parser.rs
  • rust/crates/ccusage-cli-parser/src/tests.rs
  • rust/crates/ccusage-cli/README.md
  • rust/crates/ccusage-cli/src/types.rs
  • rust/crates/ccusage-config/src/config.rs
  • rust/crates/ccusage-config/src/config_schema.rs
  • rust/crates/ccusage-core/src/lib.rs
  • rust/crates/ccusage/Cargo.toml
  • rust/crates/ccusage/README.md
  • rust/crates/ccusage/src/adapter/mod.rs
  • rust/crates/ccusage/src/cli/last_window.rs
  • rust/crates/ccusage/src/main.rs

Comment thread docs/guide/environment-variables.md
Comment thread rust/adapters/grok/README.md
Comment thread rust/adapters/grok/src/loader.rs
Comment thread rust/adapters/grok/src/parser.rs
Comment thread rust/adapters/grok/src/report.rs
@pullfrog

pullfrog Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Diagnosis

The check suite (ID 85172972129) failed on commit 93c7bef because nix/cargo-artifacts.nix was missing "grok" from the agentNames array. The Nix source filter excluded adapters/grok/Cargo.toml, causing cargo build to fail when resolving the ccusage-adapter-all crate's dependency on ccusage-adapter-grok.

Fix

Commit 462ab69 ("fix(nix): include grok adapter in cargo artifact layer") added "grok" to agentNames (line 43 of nix/cargo-artifacts.nix). This is the only change needed — the adapterCrates, adapterDirNames, crateDir, and crateSource functions all derive from agentNames automatically. No other Nix files require changes.

Status

The fix is already pushed to the fix/grok-adapter branch. The check suite that failed ran on the pre-fix commit. A new CI run on 462ab69 should pass. I do not have push permission to force-push or add a trivial commit to re-trigger CI.

Task list (3/6 completed)
  • Checkout PR branch and fetch check suite logs
  • Verify failure was introduced by this PR
  • Diagnose and fix CI failures
  • Verify fix by running exact CI commands
  • Review diff, commit, and push
  • Report progress

Pullfrog  | ⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to @v0 or keep the SHA fresh with Dependabot | View workflow run | via Pullfrog | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

…root contract

Decode percent triplets into bytes so multi-byte cwd segments survive, drop the
always-zero discriminator from the loader dedupe key, match the report test helper
to the parser's reasoning contract, and correct the README surface list plus the
GROK_HOME multi-root docs.

Co-authored-by: Codesmith <[email protected]>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3 issues found and verified against the latest diff

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="rust/adapters/grok/src/parser.rs">

<violation number="1" location="rust/adapters/grok/src/parser.rs:441">
P2: Distinct no-`eventId` turns can be dropped when only `cacheCreationTokens` differs, undercounting token and cost reports. Include cache-creation tokens in this fallback key (and the matching cross-file fallback key in `loader.rs`).</violation>

<violation number="2" location="rust/adapters/grok/src/parser.rs:498">
P3: Fallback project paths are mojibake for non-ASCII CWDs when `summary.json` is absent. Decode percent escapes into bytes, then construct the UTF-8 string once.</violation>
</file>

<file name="docs/guide/environment-variables.md">

<violation number="1" location="docs/guide/environment-variables.md:26">
P3: The new `GROK_HOME` row is listed in the environment-variables table, but the intro right above it promises that 'Directory variables can be one directory or a comma-separated list of directories'. `GROK_HOME` is a single root — the Grok adapter (resolve_root) does not split on commas and instead treats the whole, non-directory value as missing and falls back to `~/.grok`, so a comma-separated `GROK_HOME` will silently not do what the intro implies. This same PR tightened the language in `guide/index.md` to 'variables that support multiple roots can contain comma-separated directories', which is the right framing. Consider applying the same clarification to the `environment-variables.md` intro (e.g. noting that most directory variables support comma-separated lists, with `GROK_HOME` being a single root) so readers don't configure it as a list.</violation>
</file>

Tip: cubic can generate docs of your entire codebase and keep them up to date. Try it here.

Re-trigger cubic

Comment thread rust/adapters/grok/src/parser.rs Outdated
return format!("{event_id}|{model}");
}
format!(
"{session_id}|{}|{model}|{}|{}|{}|{reasoning}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Distinct no-eventId turns can be dropped when only cacheCreationTokens differs, undercounting token and cost reports. Include cache-creation tokens in this fallback key (and the matching cross-file fallback key in loader.rs).

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/adapters/grok/src/parser.rs, line 441:

<comment>Distinct no-`eventId` turns can be dropped when only `cacheCreationTokens` differs, undercounting token and cost reports. Include cache-creation tokens in this fallback key (and the matching cross-file fallback key in `loader.rs`).</comment>

<file context>
@@ -0,0 +1,1051 @@
+        return format!("{event_id}|{model}");
+    }
+    format!(
+        "{session_id}|{}|{model}|{}|{}|{}|{reasoning}",
+        timestamp.as_millis(),
+        usage.input_tokens,
</file context>

)
}

fn url_decode_lightweight(value: &str) -> String {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: Fallback project paths are mojibake for non-ASCII CWDs when summary.json is absent. Decode percent escapes into bytes, then construct the UTF-8 string once.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/adapters/grok/src/parser.rs, line 498:

<comment>Fallback project paths are mojibake for non-ASCII CWDs when `summary.json` is absent. Decode percent escapes into bytes, then construct the UTF-8 string once.</comment>

<file context>
@@ -0,0 +1,1051 @@
+    )
+}
+
+fn url_decode_lightweight(value: &str) -> String {
+    // Session parents are URL-encoded cwd paths (e.g. `D%3A%5Cproj`).
+    let bytes = value.as_bytes();
</file context>

| `QWEN_DATA_DIR` | Qwen | `~/.qwen` |
| `COPILOT_OTEL_FILE_EXPORTER_PATH` | Copilot CLI | Explicit `.jsonl` file |
| `GEMINI_DATA_DIR` | Gemini CLI | `~/.gemini/tmp` |
| `GROK_HOME` | Grok Build CLI | `~/.grok` |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The new GROK_HOME row is listed in the environment-variables table, but the intro right above it promises that 'Directory variables can be one directory or a comma-separated list of directories'. GROK_HOME is a single root — the Grok adapter (resolve_root) does not split on commas and instead treats the whole, non-directory value as missing and falls back to ~/.grok, so a comma-separated GROK_HOME will silently not do what the intro implies. This same PR tightened the language in guide/index.md to 'variables that support multiple roots can contain comma-separated directories', which is the right framing. Consider applying the same clarification to the environment-variables.md intro (e.g. noting that most directory variables support comma-separated lists, with GROK_HOME being a single root) so readers don't configure it as a list.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At docs/guide/environment-variables.md, line 26:

<comment>The new `GROK_HOME` row is listed in the environment-variables table, but the intro right above it promises that 'Directory variables can be one directory or a comma-separated list of directories'. `GROK_HOME` is a single root — the Grok adapter (resolve_root) does not split on commas and instead treats the whole, non-directory value as missing and falls back to `~/.grok`, so a comma-separated `GROK_HOME` will silently not do what the intro implies. This same PR tightened the language in `guide/index.md` to 'variables that support multiple roots can contain comma-separated directories', which is the right framing. Consider applying the same clarification to the `environment-variables.md` intro (e.g. noting that most directory variables support comma-separated lists, with `GROK_HOME` being a single root) so readers don't configure it as a list.</comment>

<file context>
@@ -6,23 +6,24 @@ ccusage supports several environment variables for configuration and customizati
+| `QWEN_DATA_DIR`                   | Qwen           | `~/.qwen`                          |
+| `COPILOT_OTEL_FILE_EXPORTER_PATH` | Copilot CLI    | Explicit `.jsonl` file             |
+| `GEMINI_DATA_DIR`                 | Gemini CLI     | `~/.gemini/tmp`                    |
+| `GROK_HOME`                       | Grok Build CLI | `~/.grok`                          |
 
 Example:
</file context>

@pkg-pr-new

pkg-pr-new Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ccusage@1593

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-arm64@1593

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-x64@1593

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/@ccusage/ccusage-linux-arm64@1593

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/@ccusage/ccusage-linux-x64@1593

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/@ccusage/ccusage-win32-x64@1593

commit: 0a5ab84

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • UTF-8 project path decoding — url_decode_lightweight now accumulates percent-decoded bytes into a Vec<u8> and constructs the result with String::from_utf8_lossy, fixing mojibake for multi-byte UTF-8 path segments like %C3%A9 (é). A new test url_decodes_multi_byte_project_segment covers this.
  • Dedupe key simplification — Removed extra_total_tokens from the loader's fallback dedupe key since it is always 0 for Grok entries (reasoning stays inside output_tokens). The key now uses session, timestamp, model, input, output, and cache-read — consistent with the parser's key sans reasoning.
  • GROK_HOME single-root docs — Clarified in docs/guide/environment-variables.md, docs/guide/getting-started.md, and docs/guide/grok/index.md that GROK_HOME accepts a single root only, unlike comma-separated agent directories. Added GROK_HOME to the env-var example and grep command.
  • Test assertion alignments — Updated report and parser tests to expect extra_total_tokens: 0 and totalTokens: 120 (reasoning was previously double-counted at 130), matching the production contract.

Pullfrog  | ⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to @v0 or keep the SHA fresh with Dependabot | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 462ab693f5a2
Base SHA: a6ea44439d0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 354.4ms 2.84 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 329.6ms 3.05 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 115.4ms 8.72 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 95.1ms 10.59 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 31.6ms 5.4ms 5.89x 54.75 MiB 12.46 MiB 0.23x 0.05 MiB/s 0.29 MiB/s
claude session --offline --json 0.00 MiB 28.1ms 3.4ms 8.18x 55.00 MiB 12.71 MiB 0.23x 0.05 MiB/s 0.45 MiB/s
codex daily --offline --json 0.00 MiB 25.9ms 2.4ms 10.99x 55.00 MiB 10.71 MiB 0.19x 0.03 MiB/s 0.36 MiB/s
codex session --offline --json 0.00 MiB 25.5ms 2.4ms 10.58x 55.00 MiB 10.71 MiB 0.19x 0.03 MiB/s 0.36 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 369.7ms 334.1ms 1.11x 950.82 MiB 952.59 MiB 1.00x 2.72 GiB/s 3.01 GiB/s
codex --offline --json 1.01 GiB 120.6ms 97.7ms 1.23x 422.89 MiB 402.91 MiB 0.95x 8.34 GiB/s 10.30 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.41 KiB +62.31 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 462ab693f5a2
Base SHA: a6ea44439d0d

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 355.0ms 2.84 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 319.6ms 3.15 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 124.9ms 8.06 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 109.3ms 9.21 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 28.9ms 24.4ms 1.18x 55.25 MiB 55.00 MiB 1.00x 0.05 MiB/s 0.06 MiB/s
claude session --offline --json 0.00 MiB 25.5ms 25.2ms 1.01x 55.00 MiB 54.75 MiB 1.00x 0.06 MiB/s 0.06 MiB/s
codex daily --offline --json 0.00 MiB 27.0ms 25.0ms 1.08x 54.75 MiB 55.00 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 26.8ms 24.3ms 1.10x 55.25 MiB 55.25 MiB 1.00x 0.03 MiB/s 0.04 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 373.5ms 363.2ms 1.03x 948.57 MiB 954.59 MiB 1.01x 2.70 GiB/s 2.77 GiB/s
codex --offline --json 1.01 GiB 126.9ms 118.6ms 1.07x 412.89 MiB 404.91 MiB 0.98x 7.94 GiB/s 8.49 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.41 KiB +62.31 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: ef3a80e22e8d
Base SHA: a6ea44439d0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 344.5ms 2.92 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 315.7ms 3.19 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 120.0ms 8.39 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 93.2ms 10.80 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 28.4ms 5.3ms 5.37x 55.00 MiB 12.46 MiB 0.23x 0.05 MiB/s 0.29 MiB/s
claude session --offline --json 0.00 MiB 25.9ms 2.8ms 9.16x 55.00 MiB 12.71 MiB 0.23x 0.06 MiB/s 0.55 MiB/s
codex daily --offline --json 0.00 MiB 23.3ms 2.3ms 10.11x 55.00 MiB 10.70 MiB 0.19x 0.04 MiB/s 0.37 MiB/s
codex session --offline --json 0.00 MiB 24.8ms 2.3ms 10.86x 55.00 MiB 10.70 MiB 0.19x 0.03 MiB/s 0.38 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 367.0ms 304.1ms 1.21x 954.58 MiB 894.58 MiB 0.94x 2.74 GiB/s 3.31 GiB/s
codex --offline --json 1.01 GiB 118.4ms 96.3ms 1.23x 432.90 MiB 416.90 MiB 0.96x 8.50 GiB/s 10.46 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.47 KiB +62.38 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: ef3a80e22e8d
Base SHA: a6ea44439d0d

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 376.0ms 2.68 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 322.7ms 3.12 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 117.7ms 8.55 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 94.5ms 10.66 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 28.2ms 26.0ms 1.08x 54.75 MiB 55.00 MiB 1.00x 0.05 MiB/s 0.06 MiB/s
claude session --offline --json 0.00 MiB 23.6ms 24.2ms 0.98x 54.75 MiB 55.00 MiB 1.00x 0.07 MiB/s 0.06 MiB/s
codex daily --offline --json 0.00 MiB 23.0ms 25.2ms 0.91x 55.00 MiB 55.00 MiB 1.00x 0.04 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 22.9ms 22.8ms 1.00x 55.25 MiB 55.00 MiB 1.00x 0.04 MiB/s 0.04 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 386.2ms 358.5ms 1.08x 956.58 MiB 948.58 MiB 0.99x 2.61 GiB/s 2.81 GiB/s
codex --offline --json 1.01 GiB 141.2ms 121.2ms 1.17x 422.90 MiB 422.91 MiB 1.00x 7.13 GiB/s 8.31 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.47 KiB +62.38 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

codesmith-bot and others added 3 commits August 10, 2026 15:33
The existing case decodes a two-byte scalar end-to-end. This adds a direct unit
test for a three-triplet CJK segment, and pins the behaviour for a percent
triplet that is not valid UTF-8: it degrades to the replacement character rather
than panicking or dropping the rest of the path.
The global dedupe rebuilt a content key for entries without an `eventId`, but
that key can only be built from what `LoadedEntry` carries, and reasoning tokens
are not among them. `parse_session_files` deduped its own file using the full
record including reasoning, so two eventId-less turns in the same second with
identical input, output and cache counts but different reasoning were kept by
the parser and then collapsed here, dropping one turn entirely.

Cross-file dedupe is about the same server event appearing in more than one
session export, which is what `eventId` identifies. Entries without one are now
left alone rather than matched on token counts.

Reported by an adversarial review pass. Every `turn_completed` in the logs
checked so far carries an `eventId`, so this was not reachable in practice.
@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 02319b32444a
Base SHA: a6ea44439d0d

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 465.9ms 2.16 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 418.4ms 2.41 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 158.9ms 6.34 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 121.9ms 8.26 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 44.2ms 33.1ms 1.34x 55.00 MiB 55.50 MiB 1.01x 0.03 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 38.2ms 37.4ms 1.02x 55.25 MiB 55.25 MiB 1.00x 0.04 MiB/s 0.04 MiB/s
codex daily --offline --json 0.00 MiB 32.6ms 36.2ms 0.90x 55.25 MiB 54.75 MiB 0.99x 0.03 MiB/s 0.02 MiB/s
codex session --offline --json 0.00 MiB 34.3ms 37.3ms 0.92x 55.00 MiB 55.00 MiB 1.00x 0.03 MiB/s 0.02 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 516.9ms 454.3ms 1.14x 960.57 MiB 942.59 MiB 0.98x 1.95 GiB/s 2.22 GiB/s
codex --offline --json 1.01 GiB 162.8ms 160.0ms 1.02x 418.90 MiB 420.91 MiB 1.00x 6.19 GiB/s 6.29 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.47 KiB +62.38 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 02319b32444a
Base SHA: a6ea44439d0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 367.5ms 2.74 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 333.0ms 3.02 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 122.7ms 8.21 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 97.9ms 10.28 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.8ms 3.6ms 8.23x 55.00 MiB 12.45 MiB 0.23x 0.05 MiB/s 0.43 MiB/s
claude session --offline --json 0.00 MiB 25.3ms 4.5ms 5.63x 55.25 MiB 12.71 MiB 0.23x 0.06 MiB/s 0.34 MiB/s
codex daily --offline --json 0.00 MiB 24.8ms 2.7ms 9.32x 55.00 MiB 10.71 MiB 0.19x 0.03 MiB/s 0.32 MiB/s
codex session --offline --json 0.00 MiB 26.1ms 2.5ms 10.62x 55.25 MiB 10.70 MiB 0.19x 0.03 MiB/s 0.35 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 378.4ms 368.7ms 1.03x 946.58 MiB 956.59 MiB 1.01x 2.66 GiB/s 2.73 GiB/s
codex --offline --json 1.01 GiB 118.4ms 98.8ms 1.20x 402.90 MiB 400.90 MiB 1.00x 8.50 GiB/s 10.19 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.47 KiB +62.38 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • Global cross-file dedupe narrowed to eventId only — entries without an eventId are always kept. The single-file parser already deduped them with the full record (including reasoning tokens, which LoadedEntry does not carry), so a coarser fallback key could only collapse distinct turns the parser deliberately kept apart.
  • Regression test keeps_event_id_less_turns_that_differ_only_in_reasoning — validates that two eventId-less turns differing only in their reasoning count both survive the global dedupe.
  • url_decode_lightweight edge-case coverage — new test exercises wide CJK scalars, Windows-style drive-letter paths, and invalid UTF-8 byte triples that must degrade to the replacement character without panicking or truncation.
  • Module-local helpers made private — split_tokens, pricing_candidates, and resolve_root changed from pub(super) to plain fn for hawk compliance.

Pullfrog  | ⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to @v0 or keep the SHA fresh with Dependabot | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: e476d4cbb704
Base SHA: a6ea44439d0d

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 338.7ms 2.97 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 302.9ms 3.32 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 119.0ms 8.46 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 92.9ms 10.84 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.0ms 24.3ms 1.19x 55.00 MiB 55.25 MiB 1.00x 0.05 MiB/s 0.06 MiB/s
claude session --offline --json 0.00 MiB 23.8ms 24.3ms 0.98x 55.00 MiB 55.25 MiB 1.00x 0.06 MiB/s 0.06 MiB/s
codex daily --offline --json 0.00 MiB 24.2ms 23.1ms 1.05x 55.00 MiB 55.00 MiB 1.00x 0.04 MiB/s 0.04 MiB/s
codex session --offline --json 0.00 MiB 22.3ms 24.5ms 0.91x 55.50 MiB 55.00 MiB 0.99x 0.04 MiB/s 0.04 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 378.2ms 353.4ms 1.07x 948.58 MiB 950.59 MiB 1.00x 2.66 GiB/s 2.85 GiB/s
codex --offline --json 1.01 GiB 120.8ms 115.9ms 1.04x 422.89 MiB 420.90 MiB 1.00x 8.33 GiB/s 8.68 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.47 KiB +62.38 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: e476d4cbb704
Base SHA: a6ea44439d0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 357.0ms 2.82 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 304.9ms 3.30 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 121.8ms 8.27 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 94.2ms 10.69 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 27.9ms 3.2ms 8.70x 55.25 MiB 12.46 MiB 0.23x 0.06 MiB/s 0.48 MiB/s
claude session --offline --json 0.00 MiB 24.8ms 4.2ms 5.86x 55.00 MiB 12.70 MiB 0.23x 0.06 MiB/s 0.37 MiB/s
codex daily --offline --json 0.00 MiB 23.7ms 2.4ms 9.87x 55.00 MiB 10.70 MiB 0.19x 0.04 MiB/s 0.36 MiB/s
codex session --offline --json 0.00 MiB 23.3ms 2.3ms 10.04x 55.00 MiB 10.70 MiB 0.19x 0.04 MiB/s 0.37 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 383.0ms 319.7ms 1.20x 960.58 MiB 952.59 MiB 0.99x 2.63 GiB/s 3.15 GiB/s
codex --offline --json 1.01 GiB 123.2ms 97.9ms 1.26x 402.89 MiB 404.91 MiB 1.01x 8.17 GiB/s 10.28 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.47 KiB +62.38 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 0a5ab844a3d3
Base SHA: a6ea44439d0d

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 363.6ms 2.77 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 312.9ms 3.22 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 120.1ms 8.38 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 93.2ms 10.80 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.2ms 5.5ms 5.34x 55.00 MiB 12.70 MiB 0.23x 0.05 MiB/s 0.28 MiB/s
claude session --offline --json 0.00 MiB 26.2ms 3.0ms 8.84x 55.00 MiB 12.70 MiB 0.23x 0.06 MiB/s 0.52 MiB/s
codex daily --offline --json 0.00 MiB 23.2ms 2.4ms 9.70x 55.00 MiB 10.71 MiB 0.19x 0.04 MiB/s 0.36 MiB/s
codex session --offline --json 0.00 MiB 24.5ms 2.3ms 10.47x 55.00 MiB 10.71 MiB 0.19x 0.04 MiB/s 0.37 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 379.0ms 323.1ms 1.17x 960.59 MiB 966.84 MiB 1.01x 2.66 GiB/s 3.12 GiB/s
codex --offline --json 1.01 GiB 121.7ms 95.8ms 1.27x 414.90 MiB 420.91 MiB 1.01x 8.27 GiB/s 10.51 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.22 KiB +62.13 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 0a5ab844a3d3
Base SHA: a6ea44439d0d

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 364.6ms 2.76 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 312.6ms 3.22 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 113.9ms 8.84 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 91.1ms 11.06 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 28.8ms 26.7ms 1.08x 55.00 MiB 55.00 MiB 1.00x 0.05 MiB/s 0.06 MiB/s
claude session --offline --json 0.00 MiB 23.4ms 24.3ms 0.96x 55.00 MiB 55.00 MiB 1.00x 0.07 MiB/s 0.06 MiB/s
codex daily --offline --json 0.00 MiB 25.1ms 24.6ms 1.02x 55.00 MiB 55.00 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 23.7ms 24.8ms 0.96x 55.00 MiB 55.00 MiB 1.00x 0.04 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 368.1ms 350.4ms 1.05x 960.58 MiB 956.84 MiB 1.00x 2.74 GiB/s 2.87 GiB/s
codex --offline --json 1.01 GiB 116.7ms 120.0ms 0.97x 428.90 MiB 416.91 MiB 0.97x 8.63 GiB/s 8.39 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.78 KiB 19.09 KiB +0.32 KiB 0.98x
installed native package binary 4163.09 KiB 4225.22 KiB +62.13 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@ryoppippi
ryoppippi merged commit 215c61c into main Aug 10, 2026
42 checks passed
@ryoppippi
ryoppippi deleted the fix/grok-adapter branch August 10, 2026 18:37
ryoppippi added a commit that referenced this pull request Aug 10, 2026
Resolves the models.dev snapshot conflict by regenerating it from the merged
pin rather than taking either side: main's copy was produced by the hourly
job under the old selection rules, and this branch's predates the pin bump.
The snapshot goes to 602 entries.

Grok Build (#1593) landed on main meanwhile, which is what makes the three
adapter test adjustments below necessary.

`grok-4.5` is now in the snapshot, so `find("grok-4.5-build")` resolves on the
grok adapter's *first* pricing candidate instead of falling through to the
`xai/grok-4.5` form. The two cost-mode tests overrode the key the fallback
used to land on, so they now override `grok-4.5`; the candidate-order test
needs a model no table prices at all, or the first candidate resolves it and
the ordering is never exercised.

Both keys carry xAI's list rate ($2/$6 per Mtok) in every table, so reported
costs are unchanged. What does change is that a user pricing override keyed
`xai/grok-4.5` no longer reaches `grok-4.5-build`, while one keyed `grok-4.5`
now does. Flagged on the PR rather than papered over: the adapter's candidate
order is advisory once `PricingMap::find` matches fuzzily, and restoring the
old precedence would mean an exact-first lookup in the adapter.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add Grok Build CLI usage reporting

3 participants