Skip to content

perf(adapter): parallelize file/DB reads across all agent loaders - #1332

Merged
ryoppippi merged 4 commits into
mainfrom
perf-opencode-parallel-reads
Jun 15, 2026
Merged

ryoppippi merged 4 commits into
mainfrom
perf-opencode-parallel-reads

Conversation

@ryoppippi

@ryoppippi ryoppippi commented Jun 15, 2026 •

Copy link
Copy Markdown
Member

Summary

claude and codex were the only loaders that read their source files in parallel; every other agent read them one at a time, so large histories were bottlenecked on serial I/O. This PR extracts the proven Claude pattern into one shared helper and adopts it across all agent loaders, with byte-identical output.

What changed

  • New shared helper adapter::read_files_parallel — reads files on a pool sized to available_parallelism (sequential fallback when --single-thread or a single worker), balances by byte size via chunk_file_indexes_by_size, and reassembles results in original file order. That ordering guarantee lets every caller keep its existing sequential dedup unchanged, so parallelism never changes which duplicate survives or the final ordering.
  • opencode also skips message files the SQLite DB already covers (file stem = message id), avoiding the read entirely.
  • File-based loaders now parallelized: amp, codebuff, droid, gemini, kimi, openclaw, pi, qwen.
  • OTEL / database loaders now parallelized across multiple sources: copilot, goose, hermes, kilo (each DB on its own read-only connection; within-DB row scan stays sequential — inherent to one connection).
  • Per-file read errors are logged and skipped instead of aborting the whole load (matches the Claude loader).

Why

Addresses slow ccusage <agent> on large histories. The cost is I/O (opening/reading many small files), not parsing — parsing was already optimized by #1326/#1327.

Testing

  • cargo test (151 adapter tests) green; clippy clean; just fmt applied.
  • New helper tests: order preservation, single/multi-thread equivalence, empty input.
  • Byte-identical output verified against the previous release for daily/monthly/weekly/session (JSON + table), plus --single-thread == multi-thread:
    • Real local data: amp, gemini, pi, opencode, copilot.
    • Synthesized fixtures (80–120 files/rows, with duplicates and multi-DB overlap): codebuff, droid, kimi, openclaw, qwen, goose, hermes, kilo.
  • Benchmarks (hyperfine --warmup 3 --runs 10 --shell none, synthetic fixtures):
    • opencode 50k files, DB covers 90%: ~5.36× faster (1.41s → 263ms).
    • opencode 50k files, no DB: ~1.14× faster.

Notes

Squash-merge per repo convention. The first commit (opencode parallel + DB-skip) is kept distinct; the helper commit then unifies opencode onto the shared path.

Summary by CodeRabbit

  • Performance
    • Parallelized loading of multiple chat/session files across adapters while preserving deterministic output ordering.
    • Reduced redundant JSON reads when matching cached/database-covered message IDs.
  • Bug Fixes
    • Improved robustness: individual file read/parse failures are now logged and skipped instead of aborting the entire load.
  • Tests
    • Added coverage ensuring deduplication output remains identical between single-thread and multi-thread loading.

The opencode loader read every storage/message/*.json file serially and
then discarded any whose id the SQLite DB pass had already contributed.
On large message trees this serial I/O dominates wall-clock, and on
DB-heavy histories most of those reads are pure waste.

Two I/O-focused changes, both preserving byte-identical output:

- Parallelize the file pass by reusing the Claude loader's proven
  pattern (chunk_file_indexes_by_size + thread::scope +
  available_parallelism). Results are reassembled by original file index
  before the sequential dedup pass, so parallelism never changes which
  duplicate survives or the final ordering. Honors shared.single_thread.
  Per-file read errors now skip+debug_log instead of aborting the whole
  load, matching the Claude loader's swallow-and-continue behaviour.

- Skip files the DB pass already covered. Message files are stored as
  storage/message/<sessionID>/<messageID>.json, so the file stem is the
  message id used for dedup. When the DB already contributed that id, the
  file would be discarded anyway, so it is filtered out before any read.
  Files whose stem is not a known id are still read and parsed normally.

Benchmarks (50k synthetic message files, hyperfine --warmup 3 --runs 10):
~1.14x faster with no DB, ~5.36x faster (1.41s -> 263ms) when the DB
covers 90% of messages. JSON and table output stay byte-identical across
daily/monthly/weekly/session, --since/--until, and single-thread.

Adds tests skips_message_files_already_covered_by_database and
dedup_is_stable_across_thread_counts.
@pullfrog

pullfrog Bot commented Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

no API key found — this repo is configured to use deepseek/deepseek-v4-pro, which needs DEEPSEEK_API_KEY, but the runner has no key for it.

To fix: add the key as a GitHub Actions secret (referenced from your workflow's env: block) or as a Pullfrog secret in the console — or switch this repo to a different model (free models need no key).

Open repo secrets → · Configure model → · Setup docs → · Ask in Discord →

Pullfrog  | Rerun failed job ➔ | View workflow run | via Pullfrog | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

@coderabbitai

coderabbitai Bot commented Jun 15, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: edfe72d4-1092-4eeb-be89-97cf337e64f5

📥 Commits

Reviewing files that changed from the base of the PR and between 2360db6 and 8569c13.

📒 Files selected for processing (15)
  • rust/crates/ccusage/src/adapter/amp/loader.rs
  • rust/crates/ccusage/src/adapter/codebuff/loader.rs
  • rust/crates/ccusage/src/adapter/copilot/loader.rs
  • rust/crates/ccusage/src/adapter/droid/loader.rs
  • rust/crates/ccusage/src/adapter/gemini/loader.rs
  • rust/crates/ccusage/src/adapter/goose/loader.rs
  • rust/crates/ccusage/src/adapter/hermes/loader.rs
  • rust/crates/ccusage/src/adapter/kilo/loader.rs
  • rust/crates/ccusage/src/adapter/kimi/loader.rs
  • rust/crates/ccusage/src/adapter/mod.rs
  • rust/crates/ccusage/src/adapter/openclaw/loader.rs
  • rust/crates/ccusage/src/adapter/opencode/loader.rs
  • rust/crates/ccusage/src/adapter/pi/loader.rs
  • rust/crates/ccusage/src/adapter/qwen/parser.rs
  • rust/crates/ccusage/src/main.rs
🚧 Files skipped from review as they are similar to previous changes (1)
  • rust/crates/ccusage/src/adapter/opencode/loader.rs

📝 Walkthrough

Walkthrough

The PR introduces read_files_parallel, a configurable parallel file-reading helper with deterministic output ordering, and applies it across 11 adapter loaders (Amp, Codebuff, Copilot, Droid, Gemini, Goose, Hermes, Kilo, Kimi, OpenClaw, Pi, Qwen). All adapters now read files concurrently (or sequentially when forced), log per-file errors instead of propagating them, and preserve existing dedup/sort semantics. OpenCode additionally filters JSON message files against SQLite DB results before parallel read and extends tests for DB-based skipping and thread-mode stability.

Changes

Parallel File Reading Across Adapters

Layer / File(s) Summary
Parallel file reading infrastructure
rust/crates/ccusage/src/adapter/mod.rs, rust/crates/ccusage/src/main.rs
Adds read_files_parallel function with configurable worker threads, byte-size-based chunking, scoped spawning, and deterministic output-order preservation. Includes unit tests for order preservation and single-thread equivalence. Re-exported in main.rs.
Adapter migrations to parallel file reading
rust/crates/ccusage/src/adapter/{amp,codebuff,copilot,droid,gemini,goose,hermes,kilo,kimi,openclaw,pi,qwen}/loader.rs
All 11 adapters import read_files_parallel and debug_log, replace sequential per-file loops with parallel reads, and change error handling from propagation to logging-then-skip. Dedup and sort logic iterate results in original order to preserve deterministic duplicate-selection behavior.
OpenCode: DB-seen filter and enhanced testing
rust/crates/ccusage/src/adapter/opencode/loader.rs
Reworks read_message_file to accept shared and return Option<LoadedEntry> with error logging. load_entries_from_directory filters JSON message files to skip stems already covered by SQLite DB before parallel read. New tests verify DB-based file skipping and identical dedup output across single-thread and multi-thread modes.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~75 minutes

Possibly Related PRs

  • ccusage/ccusage#1327: Overlaps with OpenCode loader message-file handling, specifically how JSON messages are read and deserialized and how read_message_file is used.

Poem

🐇 A parallel warren of threads now hops,
Reading files in chunks—no sequential stops!
The database knows which ones we've seen,
Skipping the JSON that's already been.
Errors logged gently, results intact and clean! 🌟

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: parallelizing file and database reads across agent loaders. It is concise, specific, and clearly reflects the core optimization work described in the changeset.
Docstring Coverage ✅ Passed Docstring coverage is 84.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf-opencode-parallel-reads

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Jun 15, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide 8569c13 Commit Preview URL

Branch Preview URL
Jun 15 2026, 09:38 PM

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
rust/crates/ccusage/src/adapter/opencode/loader.rs (1)

277-277: 💤 Low value

Consider logging JSON parse failures for debuggability.

File read errors are logged at lines 267-273, but JSON parse failures here are silently swallowed via .ok()?. While this is consistent with how the DB loader handles invalid JSON (line 235), the docstring at line 110-111 states "Per-file read failures are logged" which could be misread to include parse failures.

Either adjust the docstring to clarify that only I/O errors are logged, or add debug logging here for parity:

Suggested change (optional)
-    let value = serde_json::from_slice::<OpenCodeMessage>(&content).ok()?;
+    let value = match serde_json::from_slice::<OpenCodeMessage>(&content) {
+        Ok(value) => value,
+        Err(error) => {
+            debug_log(
+                shared,
+                format!(
+                    "Failed to parse OpenCode message file {}: {error}",
+                    path.display()
+                ),
+            );
+            return None;
+        }
+    };
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust/crates/ccusage/src/adapter/opencode/loader.rs` at line 277, The JSON
parse failure at the serde_json::from_slice call is silently swallowed via
.ok()?. The docstring states "Per-file read failures are logged" but currently
only file read I/O errors are logged (lines 267-273). Either add debug logging
at the parse failure point for parity with how other parse failures are handled,
or update the docstring around lines 110-111 to clarify that only I/O errors are
logged and JSON parse failures are silently skipped.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@rust/crates/ccusage/src/adapter/opencode/loader.rs`:
- Line 277: The JSON parse failure at the serde_json::from_slice call is
silently swallowed via .ok()?. The docstring states "Per-file read failures are
logged" but currently only file read I/O errors are logged (lines 267-273).
Either add debug logging at the parse failure point for parity with how other
parse failures are handled, or update the docstring around lines 110-111 to
clarify that only I/O errors are logged and JSON parse failures are silently
skipped.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: e8b8876c-ce66-4e36-abf7-2e18a1b8fb6f

📥 Commits

Reviewing files that changed from the base of the PR and between 6744254 and 2360db6.

📒 Files selected for processing (1)
  • rust/crates/ccusage/src/adapter/opencode/loader.rs

@pkg-pr-new

pkg-pr-new Bot commented Jun 15, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ccusage@1332

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-arm64@1332

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-x64@1332

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/@ccusage/ccusage-linux-arm64@1332

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/@ccusage/ccusage-linux-x64@1332

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/@ccusage/ccusage-win32-x64@1332

commit: 8569c13

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 1 file

Re-trigger cubic

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 2360db67738f
Base SHA: 6744254d695f

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 272.6ms 3.69 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 258.5ms 3.89 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 114.6ms 8.79 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 88.7ms 11.35 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 32.1ms 33.5ms 0.96x 53.75 MiB 53.50 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
claude session --offline --json 0.00 MiB 33.5ms 34.8ms 0.96x 53.75 MiB 53.75 MiB 1.00x 0.05 MiB/s 0.04 MiB/s
codex daily --offline --json 0.00 MiB 31.3ms 30.0ms 1.04x 53.50 MiB 53.75 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 28.0ms 32.4ms 0.86x 53.75 MiB 53.50 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 332.8ms 309.2ms 1.08x 950.58 MiB 952.32 MiB 1.00x 3.02 GiB/s 3.26 GiB/s
codex --offline --json 1.01 GiB 119.3ms 121.7ms 0.98x 415.27 MiB 411.29 MiB 0.99x 8.44 GiB/s 8.28 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.08 KiB 18.08 KiB +0.00 KiB 1.00x
installed native package binary 3975.50 KiB 3979.62 KiB +4.12 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 2360db67738f
Base SHA: 6744254d695f

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 330.3ms 3.05 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 238.0ms 4.23 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 111.6ms 9.02 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 86.2ms 11.68 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 36.2ms 6.9ms 5.22x 53.50 MiB 10.20 MiB 0.19x 0.04 MiB/s 0.22 MiB/s
claude session --offline --json 0.00 MiB 35.6ms 3.7ms 9.53x 53.75 MiB 10.20 MiB 0.19x 0.04 MiB/s 0.41 MiB/s
codex daily --offline --json 0.00 MiB 32.0ms 3.0ms 10.53x 53.75 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.28 MiB/s
codex session --offline --json 0.00 MiB 29.5ms 2.7ms 10.73x 53.75 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.31 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 324.5ms 263.3ms 1.23x 950.58 MiB 938.33 MiB 0.99x 3.10 GiB/s 3.82 GiB/s
codex --offline --json 1.01 GiB 151.5ms 109.8ms 1.38x 399.28 MiB 421.29 MiB 1.06x 6.64 GiB/s 9.17 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.08 KiB 18.08 KiB +0.00 KiB 1.00x
installed native package binary 3975.50 KiB 3979.62 KiB +4.12 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

…ncode

Extract the Claude/opencode parallel-file-read pattern into a single
generic helper, adapter::read_files_parallel, so every loader can share
one audited implementation instead of re-deriving thread::scope plumbing.

The helper reads files on a pool sized to available_parallelism (falling
back to sequential when single_thread is set or only one worker would
run), balances files across workers by byte size via
chunk_file_indexes_by_size, and reassembles results in the original file
order. That ordering guarantee is what lets each caller keep its existing
sequential dedup pass unchanged: parallelism can never change which
duplicate survives or the final ordering.

opencode's bespoke read_message_files is removed in favour of the helper;
behaviour (including the DB-covered-file skip and id dedup) is unchanged.

Adds helper tests covering order preservation, single/multi-thread
equivalence, and empty input.
Adopt read_files_parallel in the loaders that read many small per-session
files serially: amp, codebuff, droid, gemini, kimi, openclaw, pi, and
qwen. These were the I/O-bound loaders where the read loop dominates
wall-clock on large histories; parsing was already optimized by the
shared JSONL helpers.

Each loader keeps its existing dedup semantics by running the dedup pass
sequentially over the parallel results in their original file order:
- last-wins HashMap (codebuff),
- reverse latest-wins per session (droid),
- first-wins HashSet (kimi, openclaw, pi, qwen),
- plain accumulate + stable sort (amp, gemini).

Per-file read errors are now logged and skipped instead of aborting the
whole load, matching the Claude loader's swallow-and-continue behavior.
Output stays byte-identical: verified against the previous release for
every report mode on real and synthesized fixtures, with single-thread
and multi-thread runs producing identical results.
Adopt read_files_parallel in the remaining loaders that iterate over
multiple sources: copilot (OTEL files) and the SQLite-backed goose,
hermes, and kilo loaders, which can each see more than one database when
multiple data dirs / homes / channels are configured.

Each database is opened on its own read-only connection inside the
worker, and the sequential dedup pass runs over the parallel results in
their original path order, so the surviving record per key is identical
to the single-threaded read. Within a single database the row scan stays
sequential (inherent to one SQLite connection); the win is overlapping
multiple sources and keeping all loaders on one shared code path.

Output verified byte-identical to the previous release across report
modes on synthesized multi-database fixtures with overlapping ids, with
single-thread and multi-thread runs matching.
@ryoppippi ryoppippi changed the title perf(opencode): parallelize message-file reads and skip DB-covered files perf(adapter): parallelize file/DB reads across all agent loaders Jun 15, 2026

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

7 issues found across 15 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="rust/crates/ccusage/src/adapter/droid/loader.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/droid/loader.rs:28">
P2: This swallows all per-file errors, not just read failures, so malformed Droid settings are silently skipped and data can be lost without surfacing an error.</violation>
</file>

<file name="rust/crates/ccusage/src/adapter/goose/loader.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/goose/loader.rs:41">
P3: Added `unwrap_or_else` error branch is currently unreachable and adds misleading error-handling noise. Simplify by making the loader return `Vec<LoadedEntry>` (or otherwise centralize error handling in one layer).</violation>
</file>

<file name="rust/crates/ccusage/src/adapter/pi/loader.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/pi/loader.rs:34">
P2: This change increases peak memory by buffering every file’s parsed entries before dedup runs. Large PI histories with many duplicate records can see significant temporary RAM growth versus the prior streaming loop.</violation>
</file>

<file name="rust/crates/ccusage/src/adapter/openclaw/loader.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/openclaw/loader.rs:37">
P2: Parallel read now materializes all file-entry vectors before deduplication. Peak memory scales with total parsed entries, risking large-memory regressions on big histories.</violation>

<violation number="2" location="rust/crates/ccusage/src/adapter/openclaw/loader.rs:38">
P2: File read/parse errors are swallowed and replaced with empty results. This can silently drop OpenClaw usage data in normal (non-debug) runs.</violation>
</file>

<file name="rust/crates/ccusage/src/adapter/codebuff/loader.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/codebuff/loader.rs:29">
P2: This change materializes the entire parsed dataset before dedup, increasing peak memory substantially on large inputs. In high-volume histories this can negate the I/O speedup with memory pressure or OOM risk.</violation>
</file>

<file name="rust/crates/ccusage/src/adapter/qwen/parser.rs">

<violation number="1" location="rust/crates/ccusage/src/adapter/qwen/parser.rs:72">
P2: Determinism claim is not fully preserved on timestamp-fallback error paths under parallel reads. Thread scheduling can change fallback timestamps, affecting sort/dedup behavior for malformed records.</violation>
</file>

Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

// the subsequent stable sort and reverse latest-wins dedup pick the same
// snapshot per session as the single-threaded read.
let loaded = read_files_parallel(&files, shared.single_thread, |file| {
load_settings_file(file).unwrap_or_else(|error| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This swallows all per-file errors, not just read failures, so malformed Droid settings are silently skipped and data can be lost without surfacing an error.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/droid/loader.rs, line 28:

<comment>This swallows all per-file errors, not just read failures, so malformed Droid settings are silently skipped and data can be lost without surfacing an error.</comment>

<file context>
@@ -21,12 +21,22 @@ fn load_entries_inner(shared: &SharedArgs, pricing: &PricingMap) -> Result<Vec<L
+    // the subsequent stable sort and reverse latest-wins dedup pick the same
+    // snapshot per session as the single-threaded read.
+    let loaded = read_files_parallel(&files, shared.single_thread, |file| {
+        load_settings_file(file).unwrap_or_else(|error| {
+            debug_log(
+                shared,
</file context>
Fix with cubic

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Error-tolerant skip-and-log is the deliberate, codebase-wide convention for read_files_parallel (qwen, amp, copilot, gemini, goose, kimi, openclaw, opencode all do this); the error is logged via debug_log, and skipping one malformed file rather than aborting the whole report is the intended behavior. This PR brings droid into parity rather than regressing it.

// Read session files in parallel; the first-wins dedup runs sequentially
// over the original file order so the surviving record per id matches the
// single-threaded read.
let loaded = read_files_parallel(&files, shared.single_thread, |file| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This change increases peak memory by buffering every file’s parsed entries before dedup runs. Large PI histories with many duplicate records can see significant temporary RAM growth versus the prior streaming loop.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/pi/loader.rs, line 34:

<comment>This change increases peak memory by buffering every file’s parsed entries before dedup runs. Large PI histories with many duplicate records can see significant temporary RAM growth versus the prior streaming loop.</comment>

<file context>
@@ -27,8 +28,22 @@ fn load_entries_inner(
+        // Read session files in parallel; the first-wins dedup runs sequentially
+        // over the original file order so the surviving record per id matches the
+        // single-threaded read.
+        let loaded = read_files_parallel(&files, shared.single_thread, |file| {
+            parser::read_session_file(file, tz.as_ref(), shared.mode, pricing).unwrap_or_else(
+                |error| {
</file context>
Fix with cubic

// Read session files in parallel; the first-wins dedup runs sequentially
// over the original file order so the surviving record per id is the
// same as the single-threaded read.
let loaded = read_files_parallel(&files, shared.single_thread, |file| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Parallel read now materializes all file-entry vectors before deduplication. Peak memory scales with total parsed entries, risking large-memory regressions on big histories.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/openclaw/loader.rs, line 37:

<comment>Parallel read now materializes all file-entry vectors before deduplication. Peak memory scales with total parsed entries, risking large-memory regressions on big histories.</comment>

<file context>
@@ -28,8 +30,24 @@ fn load_entries_inner(
+        // Read session files in parallel; the first-wins dedup runs sequentially
+        // over the original file order so the surviving record per id is the
+        // same as the single-threaded read.
+        let loaded = read_files_parallel(&files, shared.single_thread, |file| {
+            parse_session_file(file, tz.as_ref(), shared.mode, pricing).unwrap_or_else(|error| {
+                debug_log(
</file context>
Fix with cubic

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Buffering file-entry vectors before dedup is the same intentional, bounded tradeoff used by every parallel loader, and the PR's own benchmarks show peak RSS flat (~1.00x) on the 1.01 GiB / 2597-file fixtures, so there is no measured memory regression to address.

// over the original file order so the surviving record per id is the
// same as the single-threaded read.
let loaded = read_files_parallel(&files, shared.single_thread, |file| {
parse_session_file(file, tz.as_ref(), shared.mode, pricing).unwrap_or_else(|error| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: File read/parse errors are swallowed and replaced with empty results. This can silently drop OpenClaw usage data in normal (non-debug) runs.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/openclaw/loader.rs, line 38:

<comment>File read/parse errors are swallowed and replaced with empty results. This can silently drop OpenClaw usage data in normal (non-debug) runs.</comment>

<file context>
@@ -28,8 +30,24 @@ fn load_entries_inner(
+        // over the original file order so the surviving record per id is the
+        // same as the single-threaded read.
+        let loaded = read_files_parallel(&files, shared.single_thread, |file| {
+            parse_session_file(file, tz.as_ref(), shared.mode, pricing).unwrap_or_else(|error| {
+                debug_log(
+                    shared,
</file context>
Fix with cubic

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skip-and-log per-file error handling is the established convention across all read_files_parallel loaders (qwen, amp, copilot, gemini, goose, kimi); errors are surfaced via debug_log and skipping one bad file beats aborting the whole report. This PR brings openclaw into parity, not a regression.

// Read files in parallel but apply the last-wins dedup sequentially over the
// original (sorted) file order, so the surviving entry per dedup key is
// identical to the single-threaded read.
let loaded = read_files_parallel(&files, shared.single_thread, |file| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This change materializes the entire parsed dataset before dedup, increasing peak memory substantially on large inputs. In high-volume histories this can negate the I/O speedup with memory pressure or OOM risk.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/codebuff/loader.rs, line 29:

<comment>This change materializes the entire parsed dataset before dedup, increasing peak memory substantially on large inputs. In high-volume histories this can negate the I/O speedup with memory pressure or OOM risk.</comment>

<file context>
@@ -23,9 +23,24 @@ fn load_entries_inner(shared: &SharedArgs, pricing: &PricingMap) -> Result<Vec<L
+    // Read files in parallel but apply the last-wins dedup sequentially over the
+    // original (sorted) file order, so the surviving entry per dedup key is
+    // identical to the single-threaded read.
+    let loaded = read_files_parallel(&files, shared.single_thread, |file| {
+        load_chat_file(file).unwrap_or_else(|error| {
+            debug_log(
</file context>
Fix with cubic

// Read chat files in parallel; the first-wins dedup runs sequentially over
// the original discovery order so the surviving record per id matches the
// single-threaded read.
let loaded = read_files_parallel(&files, shared.single_thread, |file| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Determinism claim is not fully preserved on timestamp-fallback error paths under parallel reads. Thread scheduling can change fallback timestamps, affecting sort/dedup behavior for malformed records.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/qwen/parser.rs, line 72:

<comment>Determinism claim is not fully preserved on timestamp-fallback error paths under parallel reads. Thread scheduling can change fallback timestamps, affecting sort/dedup behavior for malformed records.</comment>

<file context>
@@ -64,10 +65,25 @@ pub(super) fn load_entries(shared: &SharedArgs) -> Result<Vec<LoadedEntry>> {
+    // Read chat files in parallel; the first-wins dedup runs sequentially over
+    // the original discovery order so the surviving record per id matches the
+    // single-threaded read.
+    let loaded = read_files_parallel(&files, shared.single_thread, |file| {
+        read_chat_file(file, tz.as_ref(), shared.mode, pricing.as_ref(), shared).unwrap_or_else(
+            |error| {
</file context>
Fix with cubic

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fallback uses file mtime (deterministic, independent of thread scheduling); SystemTime::now() is only hit when mtime is unreadable AND the record lacks a timestamp, a path already non-deterministic in single-threaded mode. Dedup is first-wins over discovery order applied sequentially, so survivor selection stays deterministic. No new non-determinism introduced.

// run the sequential per-db dedup over the original path order so the
// surviving session per key matches the single-threaded read.
let loaded = read_files_parallel(&db_paths, shared.single_thread, |db_path| {
load_entries_from_db(db_path, tz.as_ref(), pricing, shared).unwrap_or_else(|error| {

@cubic-dev-ai cubic-dev-ai Bot Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: Added unwrap_or_else error branch is currently unreachable and adds misleading error-handling noise. Simplify by making the loader return Vec<LoadedEntry> (or otherwise centralize error handling in one layer).

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/adapter/goose/loader.rs, line 41:

<comment>Added `unwrap_or_else` error branch is currently unreachable and adds misleading error-handling noise. Simplify by making the loader return `Vec<LoadedEntry>` (or otherwise centralize error handling in one layer).</comment>

<file context>
@@ -31,10 +33,26 @@ pub(crate) fn load_entries(shared: &SharedArgs, pricing: &PricingMap) -> Result<
+    // run the sequential per-db dedup over the original path order so the
+    // surviving session per key matches the single-threaded read.
+    let loaded = read_files_parallel(&db_paths, shared.single_thread, |db_path| {
+        load_entries_from_db(db_path, tz.as_ref(), pricing, shared).unwrap_or_else(|error| {
+            debug_log(
+                shared,
</file context>
Fix with cubic

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The unreachable error branch is pre-existing: load_entries_from_db already returned a never-Err Result called with ? before this PR, which only swapped ? for unwrap_or_else to match the shared parallel-loader closure shape. Keeping goose uniform with every other loader outweighs a P3 cosmetic signature refactor in a perf PR.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 8569c139d613
Base SHA: 6744254d695f

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 291.2ms 3.46 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 262.1ms 3.84 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 146.1ms 6.89 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 98.0ms 10.28 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 35.1ms 6.9ms 5.08x 53.50 MiB 10.20 MiB 0.19x 0.04 MiB/s 0.22 MiB/s
claude session --offline --json 0.00 MiB 28.4ms 3.3ms 8.55x 53.75 MiB 10.19 MiB 0.19x 0.05 MiB/s 0.47 MiB/s
codex daily --offline --json 0.00 MiB 28.7ms 2.9ms 9.95x 53.75 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.30 MiB/s
codex session --offline --json 0.00 MiB 29.7ms 2.9ms 10.10x 53.50 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.29 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 335.0ms 268.0ms 1.25x 958.33 MiB 942.32 MiB 0.98x 3.01 GiB/s 3.76 GiB/s
codex --offline --json 1.01 GiB 114.1ms 92.7ms 1.23x 425.29 MiB 409.29 MiB 0.96x 8.82 GiB/s 10.86 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.08 KiB 18.08 KiB -0.00 KiB 1.00x
installed native package binary 3975.50 KiB 4027.00 KiB +51.50 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 8569c139d613
Base SHA: 6744254d695f

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 304.0ms 3.31 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 245.6ms 4.10 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 117.4ms 8.57 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 90.1ms 11.18 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 37.8ms 35.3ms 1.07x 53.50 MiB 53.75 MiB 1.00x 0.04 MiB/s 0.04 MiB/s
claude session --offline --json 0.00 MiB 33.3ms 32.3ms 1.03x 54.00 MiB 53.75 MiB 1.00x 0.05 MiB/s 0.05 MiB/s
codex daily --offline --json 0.00 MiB 30.1ms 29.8ms 1.01x 53.50 MiB 53.75 MiB 1.00x 0.03 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 29.8ms 30.1ms 0.99x 53.50 MiB 53.75 MiB 1.00x 0.03 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 337.3ms 302.4ms 1.12x 948.57 MiB 920.33 MiB 0.97x 2.98 GiB/s 3.33 GiB/s
codex --offline --json 1.01 GiB 118.1ms 125.1ms 0.94x 415.27 MiB 401.28 MiB 0.97x 8.53 GiB/s 8.05 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.08 KiB 18.08 KiB -0.00 KiB 1.00x
installed native package binary 3975.50 KiB 4027.00 KiB +51.50 KiB 0.99x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant