Skip to content

Codex: Cache Create always 0 and GPT-5.6 cost underestimated (cache writes now billed at 1.25x via cache_write_tokens) #1422

Description

@zengbo

Summary

For Codex sessions, Cache Create is always reported as 0 and cache-write tokens are never priced. With the GPT-5.6 family this is no longer just a cosmetic gap: OpenAI now bills prompt-cache writes at 1.25× the uncached input rate for GPT-5.6 models and later (earlier models keep free cache writes), and reports them in usage as cache_write_tokens (usage.input_tokens_details.cache_write_tokens in the Responses API, usage.prompt_tokens_details.cache_write_tokens in Chat Completions). See the prompt caching guide and GPT-5.6 announcement.

So for GPT-5.6 usage, ccusage currently:

  1. shows Cache Create = 0 in the unified report (ccusage -b), and
  2. underestimates cost: written-to-cache tokens are billed by OpenAI at 1.25× input rate, while ccusage prices all non-cached input at 1×.

Repro

$ bunx ccusage -b -s 20260709
│ ... │ - Codex │ - gpt-5.6-s… │ 2,484,357 │ 208,845 │ 0 │ 62,561,536 │ ... │
                                                       ^ Cache Create always 0

Root cause (as of 654d80f)

The codex adapter never parses or prices a cache-write bucket:

  • rust/crates/ccusage/src/adapter/codex/types.rs — CodexRawUsageFields has aliases for cached-input (cached_input_tokens / cache_read_input_tokens / cached_tokens) but no cache_write_tokens field, and CodexRawUsage has no slot for it.
  • rust/crates/ccusage/src/adapter/codex/report.rs — "cacheCreationTokens": 0 is hardcoded in group_json, model_usage_json, and totals_json.
  • rust/crates/ccusage/src/adapter/codex/report.rs — calculate_codex_model_cost() has no cache-write term (the Pricing struct already carries cache_create, which even defaults to input * 1.25 — matching OpenAI's new rule — but it is unused for codex).
  • rust/crates/ccusage/src/adapter/all/loader.rs — codex_group_row() hardcodes cache_creation_tokens: 0.

Upstream caveat

Codex CLI ≤ 0.144.1 does not persist cache_write_tokens in rollout token_count events (its protocol TokenUsage only has input_tokens / cached_input_tokens / output_tokens / reasoning_output_tokens / total_tokens), so for sessions recorded by today's codex CLI the value genuinely isn't recoverable. Still, ccusage should parse and price the field when present, so that:

  • newer codex CLI versions that start emitting it are picked up automatically,
  • headless/exec-format logs that embed raw API usage objects benefit immediately.

That upstream gap probably deserves its own issue on openai/codex; happy to file it too.

Proposed fix

Parse cache_write_tokens (plus a cache_creation_input_tokens alias) into CodexRawUsage, treat it as a subset of input_tokens alongside cached reads (non-cached, non-written input = input − cached − written), plumb it through the cumulative-total delta logic and aggregation, price it with Pricing::cache_create, and report it as cacheCreationTokens in the codex and unified reports.

I have a PR ready for this — will link it here.

Activity

  1. github-actions commented on Jul 10, 2026

    @github-actions
    Contributor

    This issue was auto-closed. Issues from new contributors are auto-closed by default.

    Maintainers review auto-closed issues and reopen worthwhile ones. Issues that do not meet the quality bar in CONTRIBUTING.md may not be reopened or receive a reply.

    Keep the issue short, concrete, and written in your own voice.

    If a maintainer replies lgtmi, your future issues will stay open. If a maintainer replies lgtm, your future issues and PRs will stay open.

    See CONTRIBUTING.md.

  2. zengbo commented on Jul 10, 2026

    @zengbo
    Author

    Fix is ready on my fork: zengbo/ccusage@main...zengbo:ccusage:fix/codex-cache-write-tokens (single commit, 10 files, +278/-20).

    What it does:

    • parses cache_write_tokens into CodexRawUsage (including the cumulative-total delta path), clamped to the non-cached input subset
    • aggregates it per group/model with a long-context bucket, mirroring the existing cached-input handling
    • prices writes with Pricing::cache_create / cache_create_above_200k — the pricing side already carried the GPT-5.6 rates, only the codex usage plumbing was missing
    • reports it as cacheCreationTokens in codex JSON + the unified report, subtracts writes from plain input to avoid double counting, and adds a Cache Create column to the codex focused table
    • docs: updated the codex cost-formula bullet; 5 new tests (parse, cumulative delta, flat cost, embedded gpt-5.6 short+long tier cost, report JSON)

    Verified: cargo test --workspace all green; a synthetic gpt-5.6-sol session with 1M input / 400K cached / 200K written prices to $9.15, matching hand-computed long-context rates; on real Codex CLI logs (which don't emit the field yet) output is byte-identical to v20.0.16.

    Since PRs from unapproved contributors are auto-closed, I'll open the PR once this issue gets an lgtm — or feel free to pull the branch directly.

  3. darthShadow commented on Jul 20, 2026

    @darthShadow

    Sorry for the ping @ryoppippi, but would it be possible to fix this or merge the linked changeset?

  4. ryoppippi commented on Aug 31, 2026

    @ryoppippi
    Member

    Historical audit: this discussion was auto-closed by the legacy contributor gate. That closure did not assess technical importance.

    Audit result: resolved. A later merged change or the current main implementation covers this request. This item is kept for history and does not need to be reopened.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    triage:resolvedResolved by a later change or current implementation.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions