Skip to content

fix(ccusage): parallelise all-agent loading - #1066

Merged
ryoppippi merged 4 commits into
mainfrom
codex/parallel-all-agent-loads
May 19, 2026
Merged

ryoppippi merged 4 commits into
mainfrom
codex/parallel-all-agent-loads

Conversation

@ryoppippi

@ryoppippi ryoppippi commented May 19, 2026 •

Copy link
Copy Markdown
Member

Summary:

  • Parallelises unified all-agent source loading with scoped threads.
  • Preserves the existing output order before session sorting or all-row aggregation.
  • Adds a regression test that fails unless independent source loaders overlap.

Testing:

  • pnpm run format
  • pnpm typecheck
  • pnpm run test
  • pre-push clippy

@coderabbitai please review this PR.


Summary by cubic

Parallelized all-agent source loading to speed up unified reports while keeping output order stable. The concurrent set now includes Hermes and Kilo, and a regression test enforces overlapping loaders to prevent a return to serial execution.

  • Refactors
    • Explicit load-kind mapping: weekly/monthly load daily rows before aggregation.
    • Introduced load_agent_rows_parallel (scoped threads + channel) with per-agent progress; panic-safe.
    • Sort by original agent index for stable JSON/table output; strip metadata_agents for Session.
    • Deterministic multi-agent errors: return the lowest-index failure; tests use distinct progress agents.

Written for commit 9f7d1e7. Summary will update on new commits. Review in cubic

Summary by CodeRabbit

  • Performance

    • Agent loading now runs concurrently, improving report generation speed and responsiveness.
    • Progress for each agent load is tracked, offering better visibility during long operations.
  • Reliability

    • Improved error handling and ordering guarantees when loading multiple agents concurrently.
  • Tests

    • Added tests to validate concurrent agent loading and ensure reliable, ordered results.

Review Change Stack

All-agent reports previously loaded each coding-agent source one after another. That made the unified reports wait for every source serially even though the loaders operate on independent local data directories.

Run the source loaders through scoped threads and collect their results over a channel so progress can update as each source finishes. The final rows are sorted back into the existing agent order before aggregation, preserving JSON and table output while reducing wall-clock time for multi-source reports.

Add a regression test that requires two agent loaders to overlap, which guards against accidentally returning this path to serial execution.
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai please review this PR. The key change is that all-agent source loaders now run concurrently, while final row ordering remains deterministic.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented May 19, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide 9f7d1e7 Commit Preview URL

Branch Preview URL
May 19 2026, 03:59 PM

@coderabbitai

coderabbitai Bot commented May 19, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Fetch PR details
gh pr view 1066 --repo ryoppippi/ccusage --json title,body,files,additions,deletions,changedFiles

# Get the diff
gh pr diff 1066 --repo ryoppippi/ccusage

Length of output: 12073


@ryoppippi: This is a clean, well-structured parallelisation. The use of thread::scope + mpsc::channel is idiomatic Rust for bounded concurrent work, and the regression test with AtomicUsize + spin-wait is an excellent way to assert that loaders actually overlap rather than serialise. A few observations below.


🟡 load_kind is an identity copy of kind

let load_kind = if kind == AgentReportKind::Session {
    AgentReportKind::Session
} else {
    AgentReportKind::Daily
};

This is just kind — the if-else returns the same variant you started with. If AgentReportKind: Copy, this can be replaced with:

let load_kind = kind;

or, if the only reason for the binding is to allow multiple closures to capture it without moving kind (which is also used later in if kind == AgentReportKind::Session and aggregate_rows(rows, kind)), the naming is fine but the if-else branch is redundant.


🟡 Test reuses UsageLoadAgent::Claude for both specs

AgentLoadSpec {
    index,
    agent,
    progress_agent: crate::progress::UsageLoadAgent::Claude, // ← both specs
    ...
}

load_agent_rows_parallel calls progress.start(spec.progress_agent) for each spec, so start(Claude) is called twice. Depending on the internals of UsageLoadProgress, this could silently swallow a double-start, mask a bug, or produce misleading terminal output during the test. Using distinct variants (e.g. Claude and Codex) would make the test more faithful.


🟢 Order preservation is correct

loaded.sort_by_key(|loaded| loaded.index) inside thread::scope after the channel drains guarantees output order matches the original specs slice regardless of which thread finishes first. ✓


🟢 Panic safety is correct

Because drop(sender) is called before the receiver loop, the receiver ends exactly when all spawned senders are dropped (either sent or panicked). The subsequent join loop catches panics via handle.join().is_err() and calls progress.fail() — each agent gets exactly one progress update in all paths. ✓


🟡 Error determinism: first-by-arrival, not first-by-index

When multiple agents fail, first_error captures whichever error arrives first on the channel (non-deterministic). If users expect the error message to consistently name the "first" agent (by declaration order), you might want to sort the errors by index before selecting. Low priority since one error is usually enough context, but worth noting.


Overall the change is a solid improvement — eliminates ~100 lines of duplicated sequential code, reduces wall-clock load time proportionally to the number of agents, and keeps output ordering deterministic. The two minor points above are easy to address but not blocking.

@coderabbitai

coderabbitai Bot commented May 19, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Agent loading is refactored: shared pricing is loaded once, per-agent AgentLoadSpecs are built and executed concurrently via load_agent_rows_parallel, results are merged in spec order, and session-specific row post-processing is preserved. A concurrency test verifies overlapping execution and ordering.

Changes

Concurrent Agent Loading

Layer / File(s) Summary
Imports and loader spec/result types
rust/crates/ccusage/src/adapter/all.rs
Standard library imports are grouped; adds AgentLoadSpec and LoadedAgentRows to represent per-agent loader closures and indexed results.
Main load_rows integration
rust/crates/ccusage/src/adapter/all.rs
load_rows computes shared pricing and load_kind, constructs per-agent AgentLoadSpec entries with boxed loader closures, invokes load_agent_rows_parallel, then appends loaded rows and applies session vs non-session post-processing (clear metadata for sessions, aggregate otherwise).
Parallel agent loader implementation
rust/crates/ccusage/src/adapter/all.rs
Adds load_agent_rows_parallel which spawns a scoped thread per spec, sends (index, agent, progress_agent, result) via mpsc, updates per-agent progress, joins threads (captures panics), sorts by spec index, and returns ordered LoadedAgentRows or the first error.
Concurrency validation test
rust/crates/ccusage/src/adapter/all.rs
Adds loads_agent_rows_concurrently to assert two loader closures overlap in execution and to verify returned loaded results preserve spec ordering.

Sequence Diagram(s)

sequenceDiagram
  participant Main
  participant LoaderThread
  participant Channel
  participant Progress as ProgressAgent
  Main->>LoaderThread: spawn scoped thread with AgentLoadSpec.loader()
  LoaderThread->>Channel: send (index, agent, progress_agent, result)
  Main->>Channel: receive messages
  Main->>Progress: mark success/failure per agent
  Main->>Main: join threads, detect panics
  Main->>Main: sort by spec index and return results or first error
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested labels

enhancement

Poem

🐇 I built two threads to fetch in race,

They hopped and met in the same shared space,
Channels hummed, the results aligned,
Ordered rows returned on time,
A happy rabbit watches code embrace.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title 'fix(ccusage): parallelise all-agent loading' directly and clearly summarizes the main change: introducing parallelised concurrent loading for all agents in the ccusage adapter.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/parallel-all-agent-loads

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Clarify why unified weekly and monthly reports still load daily source rows before aggregation by spelling out the load-kind mapping with an explicit match.

Make the concurrency regression test use distinct progress agents, matching production behaviour more closely. Also make multi-agent error reporting deterministic by selecting the lowest declaration index after all worker results have been collected.
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai addressed the review notes in d0596c5.

  • Rewrote the all-report load-kind mapping as an explicit match so weekly/monthly daily-row aggregation is clear.
  • Updated the concurrency test to use distinct progress agents.
  • Made multi-loader error selection deterministic by declaration index.

Validation after the fix:

  • cargo test --manifest-path rust/Cargo.toml loads_agent_rows_concurrently
  • pnpm typecheck
  • pnpm run test
  • pre-push clippy

Please review again.

@coderabbitai

coderabbitai Bot commented May 19, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Fetch the diff for the updated commit
gh pr diff 1066 --repo ryoppippi/ccusage

Length of output: 11548


@ryoppippi: All three fixes look correct — verified against the diff for d0596c5.

  1. Load-kind mapping — the explicit match now clearly documents that Weekly and Monthly share the Daily row-loading path. ✓

  2. Distinct progress agents in test — Claude and Codex are now used as separate entries, making the test properly faithful to production behaviour. ✓

  3. Deterministic error selection — errors.sort_by_key(|(index, _)| *index) before .into_iter().next() guarantees the first-by-declaration-order error is always surfaced. ✓

The implementation is clean and solid. Happy to approve.

(´。• ᵕ •。`) ♡

[approve]

Bring in the latest main branch so the PR is mergeable.

Resolve the all-agent adapter conflict by keeping the parallel loader path and adding the newly introduced Kilo source to the same concurrent load set. This preserves the new source support from main while keeping unified reports parallel.

Validation after resolving the merge:

- pnpm run format

- pnpm typecheck

- pnpm run test
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai merged latest origin/main and resolved the conflict by keeping all-agent loading parallel while adding the new Kilo source to the concurrent loader set.

Validation after the merge:

  • pnpm run format
  • pnpm typecheck
  • pnpm run test
  • pre-push clippy

Please review the updated branch again.

@coderabbitai

coderabbitai Bot commented May 19, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Fetch the current diff for the updated PR branch
gh pr diff 1066 --repo ryoppippi/ccusage

Length of output: 12215


@ryoppippi: The merge conflict resolution is clean. Here's what I verified:

  1. Kilo correctly wired in — load_kilo_rows(load_kind, shared, &pricing) at index 5 is consistent with the other pricing-aware loaders (codex, amp, copilot, gemini). The load_kind mapping correctly passes AgentReportKind::Session for session reports and AgentReportKind::Daily for everything else, matching the semantics of the old sequential code. ✓

  2. Agent ordering preserved — Indices 0–7 (claude → codex → opencode → amp → pi → kilo → copilot → gemini) are contiguous and consistent; the post-join sort-by-index will produce deterministic output order. ✓

  3. Previously approved fixes intact — explicit match for load_kind, distinct Claude/Codex progress agents in the concurrency test, and deterministic error selection by index are all still in place. ✓

No issues introduced by the merge.

(´。• ᵕ •。`) ♡

[approve]

@pkg-pr-new

pkg-pr-new Bot commented May 19, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ryoppippi/ccusage@1066

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-darwin-arm64@1066

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-darwin-x64@1066

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-linux-arm64@1066

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-linux-x64@1066

@ccusage/ccusage-win32-arm64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-win32-arm64@1066

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/ryoppippi/ccusage/@ccusage/ccusage-win32-x64@1066

commit: 9f7d1e7

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: a6a7e1c66d46
Base SHA: 35cb05e38459

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 35cb05e38459 633.5ms 502.3ms 32.5ms 3
PR pkg.pr.new a6a7e1c66d46 1.020s 607.9ms 33.8ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 35cb05e38459; PR package: a6a7e1c66d46. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 692.1ms 686.6ms 1.01x 1.45 GiB/s 1.47 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 840.6ms 829.2ms 1.01x 1.20 GiB/s 1.21 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 679.2ms 1.48 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 643.9ms 1.56 GiB/s 1
claude --offline --json Workspace release binary 1.01 GiB 710.5ms 1.42 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 818.3ms 1.23 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 794.9ms 1.27 GiB/s 1
codex --offline --json Workspace release binary 1.01 GiB 811.8ms 1.24 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude daily --offline --json 0.00 MiB 31.0ms 4.6ms 6.68x 0.05 MiB/s 0.33 MiB/s
claude session --offline --json 0.00 MiB 31.1ms 4.6ms 6.75x 0.05 MiB/s 0.34 MiB/s
codex daily --offline --json 0.00 MiB 31.6ms 4.3ms 7.33x 0.03 MiB/s 0.20 MiB/s
codex session --offline --json 0.00 MiB 31.1ms 4.3ms 7.24x 0.03 MiB/s 0.20 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude --offline --json 1.01 GiB 682.2ms 714.8ms 0.95x 1.48 GiB/s 1.41 GiB/s
codex --offline --json 1.01 GiB 836.7ms 820.3ms 1.02x 1.20 GiB/s 1.23 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 9.00 KiB 9.00 KiB -0.00 KiB 1.00x
installed native package binary 3160.24 KiB 3160.24 KiB +0.00 KiB 1.00x
Rust release binary rust/target/release/ccusage - 2827.68 KiB - -

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: a6a7e1c66d46
Base SHA: 35cb05e38459

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 35cb05e38459 577.0ms 879.5ms 33.4ms 3
PR pkg.pr.new a6a7e1c66d46 766.2ms 486.3ms 33.1ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 35cb05e38459; PR package: a6a7e1c66d46. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 683.8ms 682.0ms 1.00x 1.47 GiB/s 1.48 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 825.7ms 819.0ms 1.01x 1.22 GiB/s 1.23 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 669.8ms 1.50 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 645.8ms 1.56 GiB/s 1
claude --offline --json Workspace release binary 1.01 GiB 707.4ms 1.42 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 803.5ms 1.25 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 781.9ms 1.29 GiB/s 1
codex --offline --json Workspace release binary 1.01 GiB 799.9ms 1.26 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.8ms 29.9ms 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 30.3ms 30.3ms 1.00x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 30.6ms 29.8ms 1.03x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 29.5ms 30.1ms 0.98x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude --offline --json 1.01 GiB 685.5ms 672.4ms 1.02x 1.47 GiB/s 1.50 GiB/s
codex --offline --json 1.01 GiB 817.8ms 814.3ms 1.00x 1.23 GiB/s 1.24 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 9.00 KiB 9.00 KiB -0.00 KiB 1.00x
installed native package binary 3160.24 KiB 3160.24 KiB +0.00 KiB 1.00x
Rust release binary rust/target/release/ccusage - 2827.68 KiB - -

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

Bring in the latest main branch so the PR remains mergeable.

Resolve the all-agent adapter conflict by preserving the parallel loader and adding the newly introduced Hermes source to the concurrent load set alongside Kilo. This keeps current source support from main while preserving the all-agent parallelisation.

Validation after resolving the merge:

- pnpm run format

- pnpm typecheck

- pnpm run test
@ryoppippi

Copy link
Copy Markdown
Member Author

@coderabbitai merged latest origin/main again and resolved the new Hermes conflict by keeping the all-agent loader parallel and adding Hermes to the concurrent load set.

Validation after this merge:

  • pnpm run format
  • pnpm typecheck
  • pnpm run test
  • pre-push clippy

Please review the updated branch again.

@coderabbitai

coderabbitai Bot commented May 19, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
gh pr diff 1066 --repo ryoppippi/ccusage

Length of output: 12898


@ryoppippi: The Hermes merge resolution is clean. Here's what I verified:

  1. Hermes correctly wired in — load_hermes_rows(load_kind, shared, &pricing) at index 4 is consistent with the other pricing-aware loaders (amp, kilo, copilot, gemini). ✓

  2. Agent ordering preserved — Indices 0–8 (claude → codex → opencode → amp → hermes → pi → kilo → copilot → gemini) are contiguous with no gaps; the post-join sort-by-index produces deterministic output order. ✓

  3. Previously approved fixes intact — explicit match for load_kind, distinct Claude/Codex progress agents in the concurrency test, and deterministic error selection by index are all still in place. ✓

No issues introduced by this merge.

(´。• ᵕ •。`) ♡

[approve]

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 9f7d1e71ea97
Base SHA: 34f9c68d7c61

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 34f9c68d7c61 2.497s 432.4ms 33.5ms 3
PR pkg.pr.new 9f7d1e71ea97 954.5ms 382.2ms 33.8ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 34f9c68d7c61; PR package: 9f7d1e71ea97. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 690.5ms 677.6ms 1.02x 1.46 GiB/s 1.49 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 865.5ms 901.0ms 0.96x 1.16 GiB/s 1.12 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 683.4ms 1.47 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 651.9ms 1.54 GiB/s 1
claude --offline --json Workspace release binary 1.01 GiB 718.8ms 1.40 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 843.1ms 1.19 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 824.8ms 1.22 GiB/s 1
codex --offline --json Workspace release binary 1.01 GiB 811.5ms 1.24 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude daily --offline --json 0.00 MiB 30.9ms 4.7ms 6.64x 0.05 MiB/s 0.33 MiB/s
claude session --offline --json 0.00 MiB 32.1ms 4.7ms 6.85x 0.05 MiB/s 0.33 MiB/s
codex daily --offline --json 0.00 MiB 31.2ms 4.3ms 7.24x 0.03 MiB/s 0.20 MiB/s
codex session --offline --json 0.00 MiB 31.6ms 4.2ms 7.48x 0.03 MiB/s 0.20 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs rust/target/release/ccusage directly. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude --offline --json 1.01 GiB 686.9ms 717.6ms 0.96x 1.47 GiB/s 1.40 GiB/s
codex --offline --json 1.01 GiB 857.8ms 809.8ms 1.06x 1.17 GiB/s 1.24 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 9.06 KiB 9.06 KiB +0.00 KiB 1.00x
installed native package binary 3160.24 KiB 3160.24 KiB +0.00 KiB 1.00x
Rust release binary rust/target/release/ccusage - 2827.68 KiB - -

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 9f7d1e71ea97
Base SHA: 34f9c68d7c61

This compares the PR package against the configured base package on the same CI runner.

Package runner startup

Execution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one bunx -p <url> ccusage --version run with an empty Bun install cache. Warm reuses that cache and reports the median of repeated runs.

Package SHA Execution setup Bunx temp cache Bunx warm median Warm samples
Base pkg.pr.new 34f9c68d7c61 2.947s 780.4ms 35.3ms 3
PR pkg.pr.new 9f7d1e71ea97 869.5ms 1.033s 33.9ms 3

Cached bunx execution performance

Runs the same large fixture through bunx -p <pkg.pr.new URL> ccusage after the Bun install cache has already been populated by the startup measurement. This separates cached package-runner execution from first-fetch package materialization.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base package: 34f9c68d7c61; PR package: 9f7d1e71ea97. Both run through bunx -p <pkg.pr.new URL> ccusage using the warmed Bun install cache from package runner startup, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
bunx -p <pkg> ccusage claude --offline --json 1.01 GiB 687.6ms 686.2ms 1.00x 1.46 GiB/s 1.47 GiB/s
bunx -p <pkg> ccusage codex --offline --json 1.01 GiB 916.7ms 849.3ms 1.08x 1.10 GiB/s 1.19 GiB/s

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 680.5ms 1.48 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 653.3ms 1.54 GiB/s 1
claude --offline --json Workspace release binary 1.01 GiB 716.9ms 1.40 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 843.6ms 1.19 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 814.4ms 1.24 GiB/s 1
codex --offline --json Workspace release binary 1.01 GiB 815.4ms 1.23 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude daily --offline --json 0.00 MiB 33.3ms 32.2ms 1.04x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 33.1ms 31.7ms 1.04x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 32.4ms 32.9ms 0.98x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 32.7ms 32.2ms 1.02x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/work/_temp/ccusage-large-fixture (1.01 GiB, 2,597 files), Codex /home/runner/work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2,597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.

Command Input Base median PR median PR vs base Base throughput PR throughput
claude --offline --json 1.01 GiB 707.5ms 680.9ms 1.04x 1.42 GiB/s 1.48 GiB/s
codex --offline --json 1.01 GiB 854.9ms 850.1ms 1.01x 1.18 GiB/s 1.18 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 9.06 KiB 9.06 KiB +0.00 KiB 1.00x
installed native package binary 3160.24 KiB 3160.24 KiB +0.00 KiB 1.00x
Rust release binary rust/target/release/ccusage - 2827.68 KiB - -

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@ryoppippi
ryoppippi merged commit ba9ad76 into main May 19, 2026
38 checks passed
@ryoppippi
ryoppippi deleted the codex/parallel-all-agent-loads branch May 19, 2026 16:16
@coderabbitai coderabbitai Bot mentioned this pull request Jul 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant