Skip to content

feat(widgets): add cost breakdown by category (input/output/cache-write/cache-read) - #587

Open
kruttik-lab49 wants to merge 3 commits into
sirmalloc:mainfrom
kruttik-lab49:fixes-and-cost-breakdown
Open

kruttik-lab49 wants to merge 3 commits into
sirmalloc:mainfrom
kruttik-lab49:fixes-and-cost-breakdown

Conversation

@kruttik-lab49

Copy link
Copy Markdown

What's missing

ccstatusline shows one total cost number, read straight from Claude Code's
cost.total_cost_usd. There's no visibility into why that number is what it
is — no split by token category, and no per-model pricing table of its own.

What this adds

Four widgets, following the existing single-purpose pattern (TokensInput,
CacheRead, etc.) rather than one composite widget:

  • cost-input, cost-output, cost-cache-write, cost-cache-read

Each turn is priced by its own model (message.model), so a mid-session model
switch prices historical turns correctly. The per-model table only supplies
the relative weight between categories — getCostBreakdown scales the
estimate to sum to Claude Code's own total_cost_usd rather than trusting the
estimate's absolute value, so the displayed total always matches SessionCost
exactly, regardless of any billing detail (e.g. a provider markup) I can't
independently verify from outside Claude Code.

Credits — this PR includes two other fixes, not just mine

1. Token/speed-metric dedup, from #549 (@rborkow, independently confirmed
by @bmihaila-bd in the issue thread). Claude Code logs one JSONL entry per
content block, so summing without dedup over-reports tokens by ~1.4-2.5x —
unevenly across categories, since only output_tokens grows across the
duplicates while the prompt-side fields repeat unchanged. I needed this fixed
first: an uneven overcount would throw off this PR's category split even
after scaling to the real total. My commit implements the exact fix #549
recommends (dedup by message.id, keep the entry with the highest
output_tokens) and extends it to the previously-undeduped speed-metrics
path. All credit for finding and diagnosing the bug goes to #549.

2. Cache-timer real-expiry fix, from #576 (@durandom, open PR:
#576). Unrelated to cost, but I
was testing this cost feature locally alongside it and it's included here
as-is (unmodified, full credit to @durandom) rather than left out and
re-tested separately. If #576 merges first, this commit becomes a no-op here.

Real example

From a live session transcript:

Cost: $8.42
In $0.00  Out $0.88  CacheW $1.96  CacheR $5.58

The four categories sum to exactly $8.42, matching SessionCost.

Testing

  • bun test: 2224 pass, 0 fail (35 new)
  • bun run lint: clean (tsc + eslint)
  • bun run build: succeeds, verified against a real transcript + real
    total_cost_usd, output cross-checked by hand

🤖 Generated with Claude Code

k504866430 and others added 3 commits September 13, 2026 14:16
Claude Code writes one JSONL entry per content block (thinking / text /
each tool_use) within a single API response. All entries for one call
share a message.id and repeat identical prompt-side usage; only
output_tokens grows across the group, and the last entry carries the
complete value.

The old stop_reason-based filter assumed only one entry per call could
carry a non-null stop_reason, but every content-block entry does, so
none were filtered and totals over-reported by ~1.4-2.5x depending on
how many content blocks a response produces (see sirmalloc#549). The speed
metrics path (collectSpeedMetricRecord) had no dedup at all, inflating
requestCount and outTps the same way.

Both collectors now group consecutive entries by message.id, keeping
the entry with the highest output_tokens, and flush the group when the
id changes or the stream ends. This subsumes the old streaming-partial
case and also handles duplicate rows that repeat an already-complete,
byte-identical usage snapshot.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Applies sirmalloc#576 (upstream, open, not authored by us)
for local testing ahead of our own cost-breakdown widget, which reads
the same prompt_cache status field. CacheTimer previously inferred its
countdown from the last transcript timestamp, which reports HOT on a
resumed session that has actually gone cold. This reads
prompt_cache.expires_at when Claude Code provides it, falling back to
the old inference on older Claude Code versions.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
…te/cache-read)

ccstatusline shows one total cost number read directly from Claude
Code's cost.total_cost_usd, with no visibility into why that number is
what it is and no per-model pricing table of its own.

Adds four widgets (Cost Input, Cost Output, Cost Cache Write, Cost
Cache Read) that split the session cost by token category, following
the existing single-purpose widget pattern (TokensInput/CacheRead/etc.)
rather than one composite widget.

Each turn is priced by its own model (message.model), so a mid-session
model switch prices historical turns correctly, and duplicate
content-block rows are collapsed by message.id before pricing (reusing
the dedup fix from the previous commit) rather than after, so a
multi-block response isn't priced twice.

The per-model price table only supplies the *relative* weight between
categories -- getCostBreakdown (widgets/shared/cost-metrics.ts) scales
the four estimates to sum to Claude Code's own total_cost_usd rather
than trusting their absolute value, so the displayed total always
matches SessionCost exactly regardless of any billing adjustment we
can't independently verify (e.g. a provider-specific rate).

Gated behind a new includeCostEstimate scan option, set only when a
cost-breakdown widget is on the layout, so sessions without one pay no
extra cost to compute it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token widgets over-report ~1.84x: one JSONL entry per content block is counted per API call

2 participants