Repository navigation
Codex date-bounded reports scan all historical session files #1598
Description
Activity
pullfrog commented
on Aug 13, 2026 pullfrogboton Aug 13, 2026 – with PullfrogContributorMore actionsEvaluation: actionable. The issue matches the current implementation: Codex file discovery recursively collects every
*.jsonl(rust/adapters/codex/src/paths.rs:74-79),load_codex_eventsloads every selected file without usingSharedArgs.since/until(rust/adapters/codex/src/loader.rs:22-79), and the date check happens only after events are parsed (rust/adapters/codex/src/aggregate.rs:318-340,629-650). The unified loader also derives Codex detection from the loaded event vector (rust/crates/ccusage-adapter-all/src/loader.rs:625-654), so an optimized empty-window load could incorrectly omit Codex fromDetected.Implementation plan
- Add a Codex date-window abstraction built from
since,until, and the reporting timezone. Use it only as a conservative file candidate filter, with unbounded or non-resolvable bounds retaining the current all-files path. Keepcodex_period_forandfilter_events_by_dateas the authoritative exact event-level checks. - Push the lower-bound candidate filter into Codex path/file loading using cheap file or canonical session-path metadata, with a conservative boundary window. Unknown layouts, unavailable metadata, and files that could still contain an in-range event must remain eligible. Preserve active/archived source deduplication and ensure long-running sessions crossing the boundary are not excluded.
- Preserve replay correctness by ensuring the replay plan can still resolve parent-session dependencies for retained child files, without re-reading full historical event streams unnecessarily. Apply the same candidate selection to focused Codex reports and unified reports, while leaving unbounded reports on the existing behavior.
- Separate source detection from in-range row loading. Unified reports should mark Codex as detected when any Codex source file exists, even when the selected window produces no Codex rows; focused reports should continue returning an empty result for that window.
- Add fixture-backed tests for old and in-range files, path/metadata boundary cases, long-running sessions, timezone behavior, replayed parent sessions, active versus archived duplicates, single-threaded and parallel loading, and focused/unified parity. Add a regression test where all Codex usage is outside the window but unified output still lists Codex as detected. Verify unbounded output remains unchanged and profile a large historical fixture to confirm the scan reduction. This is an internal performance and correctness change, so no documentation update is expected unless the final implementation changes observable CLI semantics.
Task list (3/3 completed)
- Inspect issue details and repository guidance for Codex date filtering
- Trace Codex report file discovery and assess scope and test coverage
- Post an actionable implementation plan or issue analysis
|⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to@v0or keep the SHA fresh with Dependabot | Implement plan ➔ | View workflow run | via Pullfrog | UsingGPT Luna(free via Pullfrog for OSS) | 𝕏- Add a Codex date-window abstraction built from
github-actions commented
on Aug 13, 2026 on Aug 13, 2026 – with GitHub ActionsContributorMore actionsThis issue was auto-closed. Issues from new contributors are auto-closed by default.
Maintainers review auto-closed issues and reopen worthwhile ones. Issues that do not meet the quality bar in CONTRIBUTING.md may not be reopened or receive a reply.
Keep the issue short, concrete, and written in your own voice.
If a maintainer replies
lgtmi, your future issues will stay open. If a maintainer replieslgtm, your future issues and PRs will stay open.See CONTRIBUTING.md.
@ryoppippi @ubugeeei Could you take a look at this auto-closed issue?
ccusage codex daily --last 1 --offlinestill walks and parses the full Codex session history before applying--since/--until. On a large store, a one-day report takes seconds even though only recent files can contribute rows.I would like to implement a conservative file-level prefilter: skip historical session files before parsing, keep event-level date filtering as the source of truth, leave unbounded reports unchanged, and still detect Codex in unified reports when the selected window has no Codex rows.
I opened #1599 before reading that new-contributor PRs should wait for
lgtm— happy to treat this issue as the request and iterate if you want it reopened.@ryoppippi @ubugeeei I'd like to ask for a reopen of this issue — I hit the same problem independently and can add measurements plus a self-contained reproduction that anyone can run without real data.
I maintain a dashboard tool that shells out to
ccusage daily --since <day> --until <day>. One of my machines has a ~30 GB~/.codex, and there a single one-day query takes long enough to blow through the tool's 30 s subprocess timeout. The cost is proportional to the whole history, not to the queried window.Measurements (ccusage 20.0.19, macOS arm64)
Real
~/.codex/sessions: 4.7 GB, 3653 files, of which only 22 were modified near the query window.CODEX_HOME wall CPU totals full directory 1.47 s 5.64 s totalTokens=1477099, totalCost=1.79991 mirror containing only the 22 recent files 0.04 s 0.11 s identical So for a date-bounded query, ~98% of the work goes into files that cannot contribute any row (~1.2 s CPU per GB of history on this machine). That is where the "30 GB takes minutes" experience comes from, and re-runs don't help because there is no cache of parsed history.
Reproduction (synthetic data, stdlib-only Python)
Script (~250 lines, stdlib only): https://gist.github.com/tacogips/3716de633a59016bb962546e633097f5
It has two subcommands:
generatebuilds a fakeCODEX_HOMEwith rollout JSONL spread over N days (each file's mtime set to its own day, like the real layout), andbenchruns the same one-day query twice — full directory vs. a hard-linked mirror that keeps only files whose mtime falls inside the window plus 2 days of slack — and asserts the totals match.$ curl -sLO https://gist.githubusercontent.com/tacogips/3716de633a59016bb962546e633097f5/raw/codex_since_bench.py $ python3 codex_since_bench.py generate --out /tmp/codex-fixture --days 730 --files-per-day 5 --filler-lines 40 generated 3650 files / 2555 MB under /tmp/codex-fixture/sessions newest day: 2026-08-24 $ python3 codex_since_bench.py bench --codex-home /tmp/codex-fixture query: ccusage daily --json --offline --since 2026-08-24 --until 2026-08-24 files: 3650 total, 15 modified inside the window (+2d slack) full directory : wall 0.96s cpu 2.92s totalTokens=133750 totalCost=0.25334375 pruned mirror : wall 0.01s cpu 0.03s totalTokens=133750 totalCost=0.25334375 totals are identical -> every skipped file was irrelevant to the query
(Real rollouts parse slower than the synthetic filler, so real directories are hit harder per GB — see the table above.
bench --codex-home ~/.codexruns the same comparison against a real directory.)Notes
- The mirror in the benchmark applies exactly the conservative rule an in-tree prefilter could use before parsing: skip files whose mtime is older than
sinceminus a couple of days of slack. A file's mtime is an upper bound on its newest event timestamp, so this cannot drop in-range rows;filter_events_by_datestays the source of truth and unbounded queries keep the current path. - On current
mainthe gap is still there:collect_codex_usage_filesgathers every*.jsonl(rust/adapters/codex/src/paths.rs),load_codex_eventsnever looks atSharedArgs.since/until(rust/adapters/codex/src/loader.rs), and date filtering happens only after parsing (rust/adapters/codex/src/aggregate.rs). - Two things a fix needs to keep intact, as the plan above already lists: replay-plan parent lookups for forked sessions (a retained child may reference a pruned parent), and unified reports still marking Codex as detected when the window has no Codex rows.
I'm happy to implement this and send a PR if the issue is reopened and you're okay with the approach (I've read CONTRIBUTING.md, so I'm asking here first instead of opening a PR).
- The mirror in the benchmark applies exactly the conservative rule an in-tree prefilter could use before parsing: skip files whose mtime is older than
@pullfrog Triage this issue using the repository standing issue-triage instructions. Inspect the full issue body, comments, current main branch, and related issues or pull requests. Choose exactly one priority label. If the issue is clearly unresolved and is priority:critical or priority:high with safe and well-scoped requirements, implement it and open a focused pull request. Otherwise, only triage it. Keep the explanation short and use simple English.
Reacted by pullfrog- addedpriority:highHigh user impact or an important regression.High user impact or an important regression.
on Aug 27, 2026 pullfrog commented
on Aug 27, 2026 pullfrogboton Aug 27, 2026 – with PullfrogContributorMore actionsPriority: high.
This is reproducible and can make date-bounded Codex reports exceed timeouts on large histories. Main still parses all session files. I implemented a focused fix that prunes old files, keeps exact event filtering, preserves replay parents, handles stale mtimes, and keeps Codex detection when the window has no rows.
The fix is validated, but I could not open a PR because pushing was denied for
github-actions[bot]with HTTP 403. |⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to@v0or keep the SHA fresh with Dependabot | View workflow run | via Pullfrog | UsingGPT Luna(free via Pullfrog for OSS) | 𝕏github-actions commented
on Aug 28, 2026 on Aug 28, 2026 – with GitHub ActionsContributorMore actionsPullfrog triage: priority:high
A maintainer explicitly requested an implementation attempt. Clear, reproducible repository-scoped performance bug with concrete acceptance criteria. Related PR #1599 was auto-closed and no event shows this issue was resolved or merged.
Decision: kept open; Pullfrog will attempt a focused implementation PR.
- addedtriage:resolvedResolved by a later change or current implementation.Resolved by a later change or current implementation.
on Aug 31, 2026 Historical audit: this discussion was auto-closed by the legacy contributor gate. That closure did not assess technical importance.
Audit result: resolved. A later merged change or the current
mainimplementation covers this request. This item is kept for history and does not need to be reopened.

ccusage codex daily --last 1 --offlinestill walks and parses the complete Codex session history before applying the requested date range. On large stores this makes a one-day report take seconds even though only recent files can contribute rows.Push the lower date bound into Codex file loading, retain exact event-level filtering, and preserve Codex source detection when the selected window has no rows.
Acceptance criteria: