Repository navigation
feat(widgets): add cost breakdown by category (input/output/cache-write/cache-read) - #587
Open
kruttik-lab49 wants to merge 3 commits into
Open
kruttik-lab49 wants to merge 3 commits into
kruttik-lab49 wants to merge 3 commits into
Conversation
Claude Code writes one JSONL entry per content block (thinking / text / each tool_use) within a single API response. All entries for one call share a message.id and repeat identical prompt-side usage; only output_tokens grows across the group, and the last entry carries the complete value. The old stop_reason-based filter assumed only one entry per call could carry a non-null stop_reason, but every content-block entry does, so none were filtered and totals over-reported by ~1.4-2.5x depending on how many content blocks a response produces (see sirmalloc#549). The speed metrics path (collectSpeedMetricRecord) had no dedup at all, inflating requestCount and outTps the same way. Both collectors now group consecutive entries by message.id, keeping the entry with the highest output_tokens, and flush the group when the id changes or the stream ends. This subsumes the old streaming-partial case and also handles duplicate rows that repeat an already-complete, byte-identical usage snapshot. Co-Authored-By: Claude Sonnet 5 <[email protected]>
Applies sirmalloc#576 (upstream, open, not authored by us) for local testing ahead of our own cost-breakdown widget, which reads the same prompt_cache status field. CacheTimer previously inferred its countdown from the last transcript timestamp, which reports HOT on a resumed session that has actually gone cold. This reads prompt_cache.expires_at when Claude Code provides it, falling back to the old inference on older Claude Code versions. Co-Authored-By: Claude Sonnet 5 <[email protected]>
…te/cache-read) ccstatusline shows one total cost number read directly from Claude Code's cost.total_cost_usd, with no visibility into why that number is what it is and no per-model pricing table of its own. Adds four widgets (Cost Input, Cost Output, Cost Cache Write, Cost Cache Read) that split the session cost by token category, following the existing single-purpose widget pattern (TokensInput/CacheRead/etc.) rather than one composite widget. Each turn is priced by its own model (message.model), so a mid-session model switch prices historical turns correctly, and duplicate content-block rows are collapsed by message.id before pricing (reusing the dedup fix from the previous commit) rather than after, so a multi-block response isn't priced twice. The per-model price table only supplies the *relative* weight between categories -- getCostBreakdown (widgets/shared/cost-metrics.ts) scales the four estimates to sum to Claude Code's own total_cost_usd rather than trusting their absolute value, so the displayed total always matches SessionCost exactly regardless of any billing adjustment we can't independently verify (e.g. a provider-specific rate). Gated behind a new includeCostEstimate scan option, set only when a cost-breakdown widget is on the layout, so sessions without one pay no extra cost to compute it. Co-Authored-By: Claude Sonnet 5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What's missing
ccstatuslineshows one total cost number, read straight from Claude Code'scost.total_cost_usd. There's no visibility into why that number is what itis — no split by token category, and no per-model pricing table of its own.
What this adds
Four widgets, following the existing single-purpose pattern (
TokensInput,CacheRead, etc.) rather than one composite widget:cost-input,cost-output,cost-cache-write,cost-cache-readEach turn is priced by its own model (
message.model), so a mid-session modelswitch prices historical turns correctly. The per-model table only supplies
the relative weight between categories —
getCostBreakdownscales theestimate to sum to Claude Code's own
total_cost_usdrather than trusting theestimate's absolute value, so the displayed total always matches
SessionCostexactly, regardless of any billing detail (e.g. a provider markup) I can't
independently verify from outside Claude Code.
Credits — this PR includes two other fixes, not just mine
1. Token/speed-metric dedup, from #549 (@rborkow, independently confirmed
by @bmihaila-bd in the issue thread). Claude Code logs one JSONL entry per
content block, so summing without dedup over-reports tokens by ~1.4-2.5x —
unevenly across categories, since only
output_tokensgrows across theduplicates while the prompt-side fields repeat unchanged. I needed this fixed
first: an uneven overcount would throw off this PR's category split even
after scaling to the real total. My commit implements the exact fix #549
recommends (dedup by
message.id, keep the entry with the highestoutput_tokens) and extends it to the previously-undeduped speed-metricspath. All credit for finding and diagnosing the bug goes to #549.
2. Cache-timer real-expiry fix, from #576 (@durandom, open PR:
#576). Unrelated to cost, but I
was testing this cost feature locally alongside it and it's included here
as-is (unmodified, full credit to @durandom) rather than left out and
re-tested separately. If #576 merges first, this commit becomes a no-op here.
Real example
From a live session transcript:
The four categories sum to exactly $8.42, matching
SessionCost.Testing
bun test: 2224 pass, 0 fail (35 new)bun run lint: clean (tsc + eslint)bun run build: succeeds, verified against a real transcript + realtotal_cost_usd, output cross-checked by hand🤖 Generated with Claude Code