Repository navigation
perf(adapter): unify JSONL prefilter and byte_lines across agents - #1326
Conversation
Migrate the pi, kimi, qwen, openclaw, and copilot adapters to the same fast line-scanning path the claude and codex adapters already use: read the file as bytes, iterate newline-delimited slices with `byte_lines`, and skip non-usage lines with a reusable `memmem` substring prefilter before paying for a full `serde_json` parse. Add a shared `LinePrefilter` helper to `fast.rs` so the per-line skip check is built once (owned `Finder` needles) and reused across every line via the SIMD-accelerated `memmem` path, instead of allocating a fresh searcher on each `str::contains` call. The helper supports both "contains all markers" and "contains any marker" modes to cover the existing per-adapter filter semantics. Behavioral notes: - pi, kimi, openclaw, copilot keep their exact existing marker substrings; only the scanning mechanism changes. - qwen previously parsed every line; it now skips lines without `"usageMetadata"`, which is a required field for any emitted entry, so the result set is unchanged. - Reading bytes instead of a UTF-8 string makes a single invalid line fail to parse and be skipped rather than aborting the whole file, matching the claude/codex behavior. Verified byte-identical `daily --json --offline` output against ccusage@latest for every agent with local logs (pi, copilot, claude, codex, gemini, opencode, amp), with logs capped at a fixed date.
|
no API key found — this repo is configured to use To fix: add the key as a GitHub Actions secret (referenced from your workflow's Open repo secrets → · Configure model → · Setup docs → · Ask in Discord →
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (6)
🚧 Files skipped from review as they are similar to previous changes (4)
📝 WalkthroughWalkthrough
ChangesByte-oriented LinePrefilter rollout
Estimated code review effort🎯 2 (Simple) | ⏱️ ~12 minutes Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
|
@coderabbitai review This PR unifies the JSONL parsing fast-path ( |
The pi, kimi, qwen, openclaw, and copilot parsers all repeated the same
three-step boilerplate after the previous commit: iterate `byte_lines`,
skip lines rejected by the `LinePrefilter`, and `serde_json::from_slice`
each surviving line into a `Value`.
Move that loop into a single `prefiltered_json_values` helper in
`fast.rs` that yields the parsed `Value`s, and have every line-delimited
adapter consume it. This removes the duplicated scan/parse glue and
leaves each adapter with only its own per-record logic.
No behavior change: the helper performs the identical filter-then-parse
sequence and drops unparseable lines exactly as the inline `let Ok(..)
else { continue }` did.
|
@coderabbitai review Pushed a follow-up refactor ( |
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
ccusage-guide | 519a249 | Commit Preview URL Branch Preview URL |
Jun 15 2026, 03:30 PM |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
# Get the PR details and files changed
gh pr view 1326 --repo ccusage/ccusage --json headRefName,baseRefName,files | jq .
# Read fast.rs
cat rust/crates/ccusage/src/fast.rs
echo "=== copilot parser ==="
cat rust/crates/ccusage/src/adapter/copilot/parser.rs
echo "=== kimi parser ==="
cat rust/crates/ccusage/src/adapter/kimi/parser.rs
echo "=== openclaw parser ==="
cat rust/crates/ccusage/src/adapter/openclaw/parser.rs
echo "=== pi parser ==="
cat rust/crates/ccusage/src/adapter/pi/parser.rs
echo "=== qwen parser ==="
cat rust/crates/ccusage/src/adapter/qwen/parser.rsLength of output: 66099
|
| Adapter | Mode | Notes |
|---|---|---|
| copilot | all([b"\"attributes\""]) |
Correctly collects into Vec<Map<..>> for the two-pass trace-context algorithm — no way around this, and it's clear why |
| kimi | all([b"\"StatusUpdate\"", b"\"token_usage\""]) |
Single-pass .filter_map(...).collect() — minimal and clean |
| openclaw | any([b"\"model_change\"", b"\"model-snapshot\"", b"\"usage\""]) |
The value.as_object() guard in the loop correctly handles any non-object Value that sneaks through; any semantics match the original "at least one of" filter |
| pi | all([b"\"usage\"", b"\"message\""]) |
Retains the is_pi_message_usage post-filter for the role/type checks that the prefilter can't express — correct layering |
| qwen | all([b"\"usageMetadata\""]) |
The newly added prefilter is safe: parse_line always checks for usageMetadata anyway, so no valid records are dropped |
Correctness of the marker byte strings
All markers use the quoted form (e.g. br#""usage""# = the bytes "usage" with surrounding quotes), so there are no false positives from fields like "usageFoo" or "usageMetadata" triggering the "usage" marker.
Test coverage
The new prefiltered_json_values_skips_filtered_and_malformed_lines test in fast.rs covers the key cases: prefilter rejection, JSON parse failure (silent drop), and successful emission. The existing line_prefilter_all/any tests cover the two modes. Good shape.
Overall: the abstraction is well-placed, the boilerplate is gone, each adapter retains only its unique logic, and no behavioral regressions are visible. LGTM. 🐇
[approve]
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
echo "=== fast.rs ==="
cat rust/crates/ccusage/src/fast.rs
echo ""
echo "=== copilot parser ==="
cat rust/crates/ccusage/src/adapter/copilot/parser.rs
echo ""
echo "=== kimi parser ==="
cat rust/crates/ccusage/src/adapter/kimi/parser.rs
echo ""
echo "=== openclaw parser ==="
cat rust/crates/ccusage/src/adapter/openclaw/parser.rs
echo ""
echo "=== pi parser ==="
cat rust/crates/ccusage/src/adapter/pi/parser.rs
echo ""
echo "=== qwen parser ==="
cat rust/crates/ccusage/src/adapter/qwen/parser.rsLength of output: 65046
|
| Adapter | Mode | Notes |
|---|---|---|
| copilot | all([b"\"attributes\""]) |
Correctly collects into Vec<Map<..>> — the two-pass collect_trace_contexts → to_candidate algorithm requires all records upfront. No alternative. |
| kimi | all([b"\"StatusUpdate\"", b"\"token_usage\""]) |
Single-pass .filter_map(wire_line_to_entry).collect() — minimal and clean. |
| openclaw | any([b"\"model_change\"", b"\"model-snapshot\"", b"\"usage\""]) |
any correctly matches the original "at least one of" semantics. The value.as_object() guard in the loop handles any non-object Value that sneaks through. |
| pi | all([b"\"usage\"", b"\"message\""]) |
Retains the is_pi_message_usage post-filter for role/type checks the prefilter can't express — correct layering. |
| qwen | all([b"\"usageMetadata\""]) |
Safe new addition: parse_line re-checks usageMetadata anyway, so no valid records are dropped. |
Error-handling improvement
Switching from BufReader + reader.lines() to fs::read + byte_lines + from_slice().ok() aligns pi, kimi, qwen, openclaw, and copilot with the claude/codex behavior: a single malformed or non-UTF-8 line is silently skipped rather than aborting the entire file. This is strictly better.
Test coverage
The new prefiltered_json_values_skips_filtered_and_malformed_lines test in fast.rs covers the key cases: prefilter rejection, JSON parse failure (silent drop), and successful emission. The line_prefilter_all/line_prefilter_any tests cover both modes. Each adapter has existing unit tests. Good shape overall.
Overall: the implementation is correct, the marker strings are safe, each adapter retains its exact behavioral semantics, and the error handling is uniformly improved. LGTM. 🐇
[approve]
@ccusage/ccusage-darwin-arm64
@ccusage/ccusage-darwin-x64
@ccusage/ccusage-linux-arm64
@ccusage/ccusage-linux-x64
@ccusage/ccusage-win32-x64
ccusage
commit: |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
…refilter Adopt the `LinePrefilter` helper (originally proposed in #1326) so the typed `jsonl::records` path can prefilter on multiple markers with AND/OR semantics instead of a single substring. `records` now takes `Option<&LinePrefilter>`. This restores the precise prefilters that single-marker filtering had to drop: - pi: require both `"usage"` and `"message"` (was `"usage"` only) - kimi: require both `"StatusUpdate"` and `"token_usage"` (was `"token_usage"` only) - openclaw: admit any of `"model_change"` / `"model-snapshot"` / `"usage"` (previously had no prefilter at all) - qwen / copilot: unchanged markers, now via LinePrefilter `LinePrefilter` builds the `memmem` needles once and reuses them across every line. Output remains byte-for-byte identical (verified against ccusage@latest for all agents with local logs). Supersedes #1326 by folding its prefilter into the typed-parsing refactor.
…d helper (#1327) * refactor(adapter): parse JSONL via typed structs and shared helper Convert the agent adapters from dynamic `serde_json::Value` parsing plus hand-written `Value::get` navigation to typed `#[derive(Deserialize)]` structs parsed through the shared `adapter::jsonl` helper. Previously each adapter built a full `Value` tree per line (and several allocated a `String` per line via `BufReader::lines()`); they now read the file once, prefilter lines with `memmem`, and deserialize only the fields ccusage consumes. Converted: amp, codebuff, copilot, gemini, goose, kilo, kimi, openclaw, opencode, pi, qwen. Claude and Codex already used typed parsing. Behavior is preserved exactly: the lenient `deserialize_with` helpers match the old `Value` accessor semantics, prefilter markers are chosen to never drop an accepted line (or `None` where no always-present substring exists), and parse/non-object handling keeps the same skip-vs-error outcomes. Adapters with no line-delimited JSON hot path (single-object or SQLite sources) reuse only the leniency helpers. Droid is intentionally left on its existing path: it has no JSONL hot path, so typing it would add edge divergences with no performance benefit. Verified output parity (daily/monthly/weekly/session, --json) against both the previous main build and `ccusage@latest` for every agent with local logs (claude, codex, amp, copilot, gemini, opencode, pi): byte-for-byte identical. * perf(adapter): prefilter JSONL lines with a shared multi-marker LinePrefilter Adopt the `LinePrefilter` helper (originally proposed in #1326) so the typed `jsonl::records` path can prefilter on multiple markers with AND/OR semantics instead of a single substring. `records` now takes `Option<&LinePrefilter>`. This restores the precise prefilters that single-marker filtering had to drop: - pi: require both `"usage"` and `"message"` (was `"usage"` only) - kimi: require both `"StatusUpdate"` and `"token_usage"` (was `"token_usage"` only) - openclaw: admit any of `"model_change"` / `"model-snapshot"` / `"usage"` (previously had no prefilter at all) - qwen / copilot: unchanged markers, now via LinePrefilter `LinePrefilter` builds the `memmem` needles once and reuses them across every line. Output remains byte-for-byte identical (verified against ccusage@latest for all agents with local logs). Supersedes #1326 by folding its prefilter into the typed-parsing refactor. * fix(adapter): deserialize nested JSONL objects leniently A malformed non-object nested field (e.g. "cache": 5) made the typed struct fail to deserialize and silently dropped the whole usage record via records().ok(). The pre-refactor Value navigation treated such a field as absent and kept the otherwise usable tokens/cost. Add a shared jsonl::lenient_object helper and apply it to the nested tokens/time/cache fields in the opencode and kilo parsers to restore that behavior. Co-authored-by: Codesmith <[email protected]> * fix(gemini): compare the type discriminator untrimmed The typed refactor read the gemini record discriminator with the trimming non_empty_string helper, but the pre-refactor code compared it with a raw Value::as_str (record.get("type").and_then(Value::as_str) == Some("gemini")). Trimming let padded values like " gemini " spuriously match the "gemini" discriminator. Use the local untrimmed lenient_str helper, which mirrors Value::as_str: strings verbatim, non-strings to None without failing the line, so records still fall through to stats parsing. Co-authored-by: Codesmith <[email protected]> * fix(pi): deserialize usage.cost leniently to keep records A non-object `cost` payload made `PiUsage` deserialization fail, which dropped the whole usage record via `jsonl::records().ok()`. The pre-refactor `Value` navigation treated a non-object cost as absent display cost while keeping the record's tokens. Apply `jsonl::lenient_object` to the nested `cost` field to restore that behavior. Addresses a CodeRabbit review finding on PR #1327. * fix(adapter): keep amp and pi records on malformed nested shapes The typed refactor regressed three lenient-navigation behaviors: - amp messages were strict-deserialized as Vec<AmpMessage>, so a single non-object element (or a non-array messages field) dropped the entire thread. Parse them with the new jsonl::lenient_vec so bad elements are skipped, matching the old Value::as_array navigation. - amp ledger precedence fired on any usageLedger object (including {}), so threads with no usable events array stopped falling back to message usage. Track events as Option<Vec<_>> via jsonl::lenient_array and only take the ledger branch when the events array is present, and read usage_ledger via jsonl::lenient_object so a non-object usageLedger no longer fails the whole thread. - pi usage.cost was strict-typed, so a non-object cost dropped an otherwise usable record; read it via jsonl::lenient_object. Adds jsonl::lenient_array / jsonl::lenient_vec helpers plus regression tests for each path. Co-authored-by: Codesmith <[email protected]> * fix(openclaw): deserialize message leniently to keep model state A non-object message field made OpenClawLine deserialization fail, so jsonl::records().ok() dropped the whole line. For a model_change or model-snapshot record that also carried a malformed message, this lost the model/provider state update that subsequent usage entries rely on. The pre-refactor Value navigation treated a non-object message as no message while keeping the line, so apply jsonl::lenient_object to restore that behavior. Co-authored-by: Codesmith <[email protected]> * style(claude): satisfy treefmt on null-field match arms `nix flake check`'s treefmt wants the first `matches!` arm in `is_unsupported_nullable_field` split across lines like the rest. Apply the pinned formatter so the check passes; no behavior change. --------- Co-authored-by: Codesmith <[email protected]>

Summary
Unifies the JSONL parsing fast-path across agent adapters. The
claudeandcodexadapters already read file bytes, iterate newline-delimited slices viabyte_lines, and use a reusablememmemsubstring prefilter to skip non-usage lines before runningserde_json. This PR brings the remaining line-delimited adapters — pi, kimi, qwen, openclaw, copilot — onto the same path.A new shared
LinePrefilterhelper infast.rsbuilds thememmem::Finderneedles once (owned, reused across every line) instead of allocating a fresh searcher on eachstr::containscall. It supports both "contains all markers" and "contains any marker" modes to match the existing per-adapter filter semantics.What changed
read_to_string().lines()+str::containsfs::read+byte_lines+LinePrefilter::allread_to_string().lines()+str::containsfs::read+byte_lines+LinePrefilter::allBufReader::lines(), no prefilterfs::read+byte_lines+LinePrefilter::all([usageMetadata])BufReader::lines()+str::containsfs::read+byte_lines+LinePrefilter::anyread_to_string().lines()+str::containsfs::read+byte_lines+LinePrefilter::allParsing logic is otherwise unchanged (still
serde_json::Value-based).Behavioral notes
"usageMetadata", which is a required field for any emitted entry (record.get("usageMetadata")?), so the result set is unchanged — it just gains a prefilter it didn't have.Scope notes
memmemand streams withBufReader::read_until; converting it to whole-filebyte_linescould regress very large session logs.serde_json::Valueto fully typed structs is a larger, higher-risk follow-up and is intentionally not included here.Validation
cargo test -p ccusage: 271 passed (incl. newLinePrefiltertests and all adapter loader-level tests).cargo clippy --all-targets: clean.treefmt/just fmt: clean.ccusage@latest(v20.0.13): ran<agent> daily --json --offline --until 20260614for every agent with local logs and confirmed byte-identical output:pi(4 rows),copilot(1 row) ✅fast.rshot path:claude(155 rows),codex(110 rows) ✅gemini,opencode,ampkimi/qwen/openclawhad no local logs (skipped).Need help on this PR? Tag
/codesmithwith what you need. Autofix is enabled.Summary by cubic
Unifies the JSONL fast path across agents with
byte_lines, a reusableLinePrefilter, and a sharedprefiltered_json_valueshelper to skip non-usage lines beforeserde_json. Migratespi,kimi,qwen,openclaw, andcopilotto the same path asclaude/codex, improving scan speed and resilience to invalid lines with no output changes.LinePrefilterandprefiltered_json_valuesinfast.rs(all/any markers) built onmemchr::memmem::Finder.fs::read+prefiltered_json_valueswith the same marker substrings as before.qwen: now prefilters on"usageMetadata"; emitted results stay the same.codexandgeminipaths remain unchanged.Written for commit 519a249. Summary will update on new commits.
Summary by CodeRabbit