Repository navigation
Bug: Massive token overcounting for Codex subagent sessions (91x inflation) #950
Description
Activity
@pullfrog could you work on it?
Reacted by pullfrog and Tengs FamImplemented fix for #950: detect thread_spawn subagent sessions and skip replayed parent token history to prevent 91x token overcounting.
PR: #1218
Approach:
- Pre-scan files for
thread_spawnbyte pattern to detect subagent sessions - For subagent sessions, pre-scan token_count entries to find the replay timestamp pattern (≥2 entries with the same second confirms replay)
- In the main parse loop, skip all token_count entries matching the confirmed replay timestamp
All 215 existing tests pass, plus 2 new fixture-based tests covering single-subagent and multi-subagent scenarios.
Task list (6/6 completed)
- Explore codebase: understand Codex agent source parsing, JSONL record handling, and token counting
- Design fix: detect subagent sessions via thread_spawn, skip replayed entries by timestamp threshold
- Implement fix in parser.rs: subagent detection + replay skipping
- Add tests for subagent replay deduplication in loader.rs
- Self-review via reviewfrog subagent
- Commit, push, create PR
| View workflow run | via Pullfrog | UsingDeepSeek Pro(free via Pullfrog for OSS) | 𝕏- Pre-scan files for
- added a commit that references this issue
on Jun 6, 2026 - added a commit that references this issue
on Jun 8, 2026 - added 2 commits that reference this issue
on Jul 21, 2026 - added a commit that references this issue
on Jul 25, 2026 - added a commit that references this issue
on Jul 27, 2026 I still see this problem at ccusage 20.0.18, it shows 54,035,907 input and 6,638,218 output with 26 subagents. That is impossible in a day of single thread usage
same. i resumed an old chat, and it created a bunch of new subagents and reused previous ones, and my days input cache tokens jumped up by 3.5 billion. looking at sessions it seems to come from subagents.
unless openai has gotten really generous i kinda doubt that i was able to rack up 1,7kUSD usage in 10% of my weekly limit lol


Summary
When OpenAI Codex spawns subagent threads via
thread_spawn, the subagent's rollout JSONL file contains a full replay of the parent thread's token usage history, re-timestamped to the subagent's creation time.@ccusage/codextreats all these replayedtoken_countevents as real API calls, inflating reported usage by up to 91x.Environment
@ccusage/codexversion: latest (as of 2026-04-18)Reproduction
thread_spawn(e.g., worker/explorer agents)npx @ccusage/codex dailyon the date the subagents were createdRoot Cause (triple inflation)
Layer 1: Parent history replay in subagent rollout files
When a subagent is forked from a parent thread, its rollout file includes two
session_metaentries:The subagent file then replays the parent's entire conversation history — all
event_msgentries withtoken_countpayloads — with timestamps all set to the subagent's creation time (within the same second). For example, a subagent created atTwill have thousands oftoken_countentries all timestampedT.xxx.The replay accounts for 99.8% of the token data in each subagent's rollout file. The subagent's own actual API calls are only ~0.2%.
Layer 2: Duplicate logging
Within each rollout file, roughly 47% of entries are exact duplicates — same timestamp (to the millisecond), same
last_token_usage, sametotal_token_usage, samerate_limits. This causes ccusage to count each replayed call twice (sincelast_token_usageis non-null in both copies).Layer 3: Multiple subagents replaying the same history
The parent spawned 12 subagents for the same task, and each independently replayed the parent's full history. So the parent's token usage was counted 12 times.
Impact
Metric | ccusage reported | Actual -- | -- | -- Input tokens | ~20.6B | ~226M Cost | ~$9,041 | ~$100 Inflation factor | | 91xSuggested Fix
ccusage could detect and skip replayed entries by:
session_metaentries withsource.subagent.thread_spawn. If present, the session is a forked subagent.session_meta(with the parent's ID) marks the start of replayed history. Entries before the first "live" event (where timestamp meaningfully differs) should be treated as context, not new usage.(timestamp, last_token_usage.input_tokens, last_token_usage.output_tokens)should be deduplicated to prevent double-counting.Minimal test case indicators
A subagent rollout file affected by this bug will have:
session_metaentries (self + parent)source.subagent.thread_spawnin the firstsession_metatoken_countentries all within 1 second of the file's creation timetotal_token_usageprogression as the parent's rollout file