Skip to content

Bug: Massive token overcounting for Codex subagent sessions (91x inflation) #950

Description

@fakechris

Summary

When OpenAI Codex spawns subagent threads via thread_spawn, the subagent's rollout JSONL file contains a full replay of the parent thread's token usage history, re-timestamped to the subagent's creation time. @ccusage/codex treats all these replayed token_count events as real API calls, inflating reported usage by up to 91x.

Environment

  • @ccusage/codex version: latest (as of 2026-04-18)
  • Codex CLI version: 0.120.0
  • Platform: macOS

Reproduction

  1. Have a long-running Codex session (parent thread) that accumulates significant token usage over days
  2. The parent spawns multiple subagents via thread_spawn (e.g., worker/explorer agents)
  3. Run npx @ccusage/codex daily on the date the subagents were created
  4. Observe astronomically high token counts

Root Cause (triple inflation)

Layer 1: Parent history replay in subagent rollout files

When a subagent is forked from a parent thread, its rollout file includes two session_meta entries:

{"type": "session_meta", "payload": {"id": "<subagent-id>", "source": {"subagent": {"thread_spawn": {"parent_thread_id": "<parent-id>"}}}}}
{"type": "session_meta", "payload": {"id": "<parent-id>"}}

The subagent file then replays the parent's entire conversation history — all event_msg entries with token_count payloads — with timestamps all set to the subagent's creation time (within the same second). For example, a subagent created at T will have thousands of token_count entries all timestamped T.xxx.

The replay accounts for 99.8% of the token data in each subagent's rollout file. The subagent's own actual API calls are only ~0.2%.

Layer 2: Duplicate logging

Within each rollout file, roughly 47% of entries are exact duplicates — same timestamp (to the millisecond), same last_token_usage, same total_token_usage, same rate_limits. This causes ccusage to count each replayed call twice (since last_token_usage is non-null in both copies).

Layer 3: Multiple subagents replaying the same history

The parent spawned 12 subagents for the same task, and each independently replayed the parent's full history. So the parent's token usage was counted 12 times.

Impact

Metric | ccusage reported | Actual -- | -- | -- Input tokens | ~20.6B | ~226M Cost | ~$9,041 | ~$100 Inflation factor |   | 91x

Suggested Fix

ccusage could detect and skip replayed entries by:

  1. Detect subagent sessions: Check for session_meta entries with source.subagent.thread_spawn. If present, the session is a forked subagent.
  2. Identify replay boundary: In subagent files, the replayed entries all share the same creation timestamp (first ~N thousand entries). Real entries have timestamps that differ by minutes. The replay can be detected by finding where timestamps diverge from the session creation time.
  3. Alternatively, skip parent history: The second session_meta (with the parent's ID) marks the start of replayed history. Entries before the first "live" event (where timestamp meaningfully differs) should be treated as context, not new usage.
  4. Deduplicate: Even for non-replay entries, entries with identical (timestamp, last_token_usage.input_tokens, last_token_usage.output_tokens) should be deduplicated to prevent double-counting.

Minimal test case indicators

A subagent rollout file affected by this bug will have:

  • Two session_meta entries (self + parent)
  • source.subagent.thread_spawn in the first session_meta
  • Thousands of token_count entries all within 1 second of the file's creation time
  • The same total_token_usage progression as the parent's rollout file
  • ~47% exact-duplicate entries (identical timestamp + token values)

Activity

  1. Tengs-Fan commented on Jun 6, 2026

    @Tengs-Fan

    I meet the same problem, I was scared when ccusage shows that I spent $5500 on token in one day

    Image
  2. ryoppippi commented on Jun 6, 2026

    @ryoppippi
    Member

    @pullfrog could you work on it?

  3. pullfrog commented on Jun 6, 2026

    @pullfrog
    Contributor

    Implemented fix for #950: detect thread_spawn subagent sessions and skip replayed parent token history to prevent 91x token overcounting.

    PR: #1218

    Approach:

    • Pre-scan files for thread_spawn byte pattern to detect subagent sessions
    • For subagent sessions, pre-scan token_count entries to find the replay timestamp pattern (≥2 entries with the same second confirms replay)
    • In the main parse loop, skip all token_count entries matching the confirmed replay timestamp

    All 215 existing tests pass, plus 2 new fixture-based tests covering single-subagent and multi-subagent scenarios.

    Task list (6/6 completed)
    • Explore codebase: understand Codex agent source parsing, JSONL record handling, and token counting
    • Design fix: detect subagent sessions via thread_spawn, skip replayed entries by timestamp threshold
    • Implement fix in parser.rs: subagent detection + replay skipping
    • Add tests for subagent replay deduplication in loader.rs
    • Self-review via reviewfrog subagent
    • Commit, push, create PR

    Pullfrog  | View workflow run | via Pullfrog | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏

  4. added a commit that references this issue on Jun 6, 2026
    38c883b
  5. added a commit that references this issue on Jun 8, 2026
    22e5944
  6. weathon commented on Aug 23, 2026

    @weathon

    I still see this problem at ccusage 20.0.18, it shows 54,035,907 input and 6,638,218 output with 26 subagents. That is impossible in a day of single thread usage

  7. retrievaaaa commented on Sep 5, 2026

    @retrievaaaa

    same. i resumed an old chat, and it created a bunch of new subagents and reused previous ones, and my days input cache tokens jumped up by 3.5 billion. looking at sessions it seems to come from subagents.

    unless openai has gotten really generous i kinda doubt that i was able to rack up 1,7kUSD usage in 10% of my weekly limit lol

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions