Repository navigation
fix(codex): follow a rewritten replay burst past a second tick - #1501
Conversation
When a forked session's parent log is not in the scanned set, the parser cannot subtract the exact replayed prefix and falls back to skipping the burst Codex rewrote to the fork instant. That fallback bucketed events by their recorded second, so any burst written across a second tick was truncated and its remainder counted as the child's own usage. One real subagent log on hand reports 47,175,282 tokens through the fallback against 19,820,638 when the parent log is present, a 2.4x over-count from 315 replayed records that landed in the second after the tick. Follow the run instead of the second: skip while successive events stay within a second of each other. Measured across the fork logs on hand, a rewritten burst spans 10 to 40ms and the child's own first turn follows a pause of 5.8 to 15.3 seconds, so a second sits two orders of magnitude above the burst and well below the pause. Verified against 210 real fork sessions whose parent log could be located: the fallback now returns the same totals as exact prefix subtraction for every one of them, and a full scan is unchanged down to the cost figure. `skips_missing_parent_replay_when_duplicate_snapshot_is_suppressed` placed the child's own turn 800ms after the burst, which only read as the child's own under second bucketing. Its subject is snapshot suppression, so the fixture now uses a pause a real log would show and keeps testing that. Closes the gap left by #1457.
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
ccusage-guide | afca9d3 | Commit Preview URL Branch Preview URL |
Jul 27 2026, 02:25 PM |
📝 WalkthroughWalkthroughCodex replay detection now uses millisecond timestamp bursts instead of same-second anchoring. Loader tests cover multi-second replay bursts, preserved fork-local usage, and duplicate snapshot suppression. ChangesCodex replay detection
Estimated code review effort: 3 (Moderate) | ~20 minutes Possibly related issues
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — replaces second-bucketed replay detection with time-continuity, following the rewritten burst past tick boundaries instead of splitting it.
detect_rewritten_burstreplacesdetect_replay_second— checks the first two usage events are withinCODEX_REWRITTEN_BURST_PAUSE_MS(1000ms) rather than sharing a second-prefix byte string.SkippingRewrittenBurst(TimestampMs)replacesSkippingSecond([u8; 19])— carries the last skipped timestamp so successive events are gated by gap rather than second match, consuming any burst that stays within the threshold.- Two new tests cover the straddle case and multi-turn fork-local usage; an existing test fixture was adjusted to use a realistic post-burst pause (8s instead of 800ms) since the old gap no longer registers as "past the burst."
@v0 or keep the SHA fresh with Dependabot | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏
ccusage
@ccusage/ccusage-darwin-arm64
@ccusage/ccusage-darwin-x64
@ccusage/ccusage-linux-arm64
@ccusage/ccusage-linux-x64
@ccusage/ccusage-win32-x64
commit: |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
rust/adapters/codex/src/loader.rs (1)
1725-1749: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winConsider a regression test for the documented residual limitation.
The PR notes the fallback can still skip a real child turn that begins within one second of the replay burst. None of the added tests pin this documented edge case down, so a future change to the pause threshold or state machine could silently alter that behavior without a failing test.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@rust/adapters/codex/src/loader.rs` around lines 1725 - 1749, Add a regression test alongside skips_a_rewritten_burst_that_straddles_a_second_boundary covering a genuine child turn that begins within one second of the replay burst. Assert the documented fallback behavior still skips that turn, so future changes to the pause threshold or state machine cannot alter this residual limitation unnoticed.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@rust/adapters/codex/src/loader.rs`:
- Around line 1725-1749: Add a regression test alongside
skips_a_rewritten_burst_that_straddles_a_second_boundary covering a genuine
child turn that begins within one second of the replay burst. Assert the
documented fallback behavior still skips that turn, so future changes to the
pause threshold or state machine cannot alter this residual limitation
unnoticed.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: a87cd26d-8d68-4871-b105-7a487c94eb53
📒 Files selected for processing (2)
rust/adapters/codex/src/loader.rsrust/adapters/codex/src/parser.rs
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
There was a problem hiding this comment.
All reported issues were addressed across 2 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
…urst Co-authored-by: Codesmith <[email protected]>
There was a problem hiding this comment.
Reviewed changes — tightens the burst gap check from <= 1000 to a 0..=1000 range, requiring monotonic timestamps so wrapping subtraction of a non-monotonic pair cannot slip through.
CODEX_REWRITTEN_BURST_PAUSE_MSgap check indetect_rewritten_burstandSkippingRewrittenBurst—gap <= 1000→(0..=1000).contains(&gap). Sinceas_millis()returnsi64, non-monotonic timestamps wrap to a negative value that the old<=check accepted but the range check correctly rejects.
@v0 or keep the SHA fresh with Dependabot | View workflow run | Using DeepSeek Pro (free via Pullfrog for OSS) | 𝕏
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |

Closes the gap left by #1457, which I closed as superseded — the one-second heuristic it rewrote no longer exists, but the residual bug it pointed at was real.
Problem
When a forked session's parent log is in the scanned set,
CodexReplayPlansubtracts the exact replayed prefix and the accounting is correct. When it is not — the parent was deleted, archived elsewhere, or lives outside the scanned directory — the parser falls back to skipping the burst Codex rewrote to the fork instant.That fallback bucketed events by their recorded second (
[u8; 19]prefix compare). Codex writes the replayed history in a few milliseconds, so whenever a fork lands late in a second the burst straddles the tick — and everything after the tick was counted as the child's own usage.One real subagent log in a local
~/.codex/sessions:A 2.4x over-count, from 315 replayed records that landed in the second after the tick. Its burst is 458 records at
00:36:00plus 315 at00:36:01, all written between.985and.000— about 15ms of wall clock spread across two second strings.Fix
Follow the run rather than the second: skip while successive usage events stay within a second of each other, carrying the last skipped timestamp instead of a fixed second.
The threshold is picked from measurement, not taste. Across the 229 fork sessions in that log directory, the 9 with a leading burst show:
One second sits two orders of magnitude above the burst and roughly 6x below the shortest real pause, so the two populations are cleanly separated.
I also checked #1457's own signal —
task_started.started_atmatching the record second — and did not adopt it. Across 1,833 realtask_startedrecords in fork sessions, 135 havestarted_atdiffering from their record second, including differences of exactly 1 second on native records that merely crossed a tick. Strict equality there misclassifies them.Verification
210 real fork sessions whose parent log could be located, each scanned twice, once alone and once beside its parent:
The fallback now agrees with exact prefix subtraction on every one of them.
A full scan of the same directory is unchanged down to the cost figure:
totalTokens 23,772,515,519,costUSD 12927.978162399995before and after. The normal path is untouched.just testgreen,just fmt0 changed,cargo clippy --all-targetsno warnings.Test change to flag
skips_missing_parent_replay_when_duplicate_snapshot_is_suppressedplaced the child's own turn 800ms after the burst. That only read as the child's own usage under second bucketing — by the measurements above an 800ms gap is inside a burst, not after one. The test's subject is snapshot suppression, so its fixture now uses a pause a real log would show and it keeps testing exactly that.Known limits
A fork whose own first turn begins within a second of the replayed burst would still be skipped. The shortest real pause observed is 5.8 seconds, and the previous code carried the same class of exposure, so this narrows the window rather than closing it. The exact path via the parent log has no such limit; this is only the fallback.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is enabled.Summary by cubic
Fixes inflated token usage when a forked Codex session’s parent log is missing by following the rewritten replay burst across second boundaries and requiring monotonic timestamps. Aligns the fallback with exact-prefix subtraction and prevents counting replayed history as the child’s usage.
CODEX_REWRITTEN_BURST_PAUSE_MS,detect_rewritten_burst, and a new stateSkippingRewrittenBurst; useparse_ts_timestampand real timestamps instead of[u8; 19]second prefixes.Written for commit afca9d3. Summary will update on new commits.
Summary by CodeRabbit
Bug Fixes
Tests