Skip to content

fix(usage): fingerprint the usage cache by refresh token so a routine token refresh is not read as an account switch - #536

Merged
sirmalloc merged 3 commits into
sirmalloc:mainfrom
emotionalboySY:fix/stale-cache-through-token-refresh
Sep 17, 2026
Merged

sirmalloc merged 3 commits into
sirmalloc:mainfrom
emotionalboySY:fix/stale-cache-through-token-refresh

Conversation

@emotionalboySY

Copy link
Copy Markdown
Contributor

Symptom

A server-side 429 with Retry-After: 3600 left every usage widget rendering [Rate limited] for the full hour, even though usage.json held a complete, usable snapshot the entire time. The stale-cache fallback that exists for exactly this case never fired.

7d: [Rate limited] | 7d-F: [Rate limited]
// ~/.cache/ccstatusline/usage.json, written 2h38m earlier
{"sessionUsage":0,"weeklyUsage":8,"weeklyResetAt":"2026-08-07T12:00:00Z",
 "weeklySonnetUsage":0,"weeklyOpusUsage":0,"fableUsage":16,
 "fableResetAt":"2026-08-07T12:00:00Z","tokenHash":"fe98a0af7b2e6fe5"}

// ~/.cache/ccstatusline/usage.lock
{"blockedUntil":1785803730,"error":"rate-limited"}

Root cause

The cache fingerprint hashes claudeAiOauth.accessToken. Per the comment on fingerprintUsageToken, its job is to catch a login switch:

persisted alongside the cache so a login switch (e.g. enterprise<->personal, a different token) invalidates the cache immediately instead of waiting out the TTL

But the access token is reissued on a short cycle, so an ordinary refresh of the very same account changes the fingerprint too, and readStaleUsageCache() discards the snapshot as if the account had switched. On my machine, minutes apart:

value
fingerprint stored in the cache fe98a0af7b2e6fe5
fingerprint of the live access token 25a70170398d1eb1
expiresAt of that access token ~8h out, reissued after the cache was written
refreshTokenExpiresAt ~13 days out

Same account, same subscription — the session token simply rotated.

This is worst exactly when the cache is load-bearing. Once a backoff lock is active, every render goes through getStaleUsageOrError(), finds the fingerprint mismatched, drops a valid snapshot and returns error text until the lock expires — up to the server's Retry-After.

In practice the widgets left visibly broken are the per-model ones (weekly-opus-usage, fable-weekly-usage), since session / weekly_all can still be served from the stdin rate_limits payload.

Fix

Fingerprint the refresh token instead. It identifies the login rather than the session — roughly two orders of magnitude longer-lived (~13 days vs ~8 hours in the values above) — so the account-switch guard keeps working without churning every few hours. It falls back to the access token when absent, so older credential files and any keychain entry that stores only an access token behave exactly as before.

getUsageToken() keeps its signature and still returns the access token used for the API call; the credentials record it is now built from carries the refresh token purely for fingerprinting.

What I did not do

My first attempt was to relax the fallback instead — serve the fingerprint-mismatched cache while a transient error (rate-limited / timeout / api-error) is in effect. Two existing tests deliberately pin the opposite behavior:

  • does not serve a mismatched account cache during an active lock
  • does not serve a mismatched account cache during a rate-limit backoff

Those assertions are right — a mismatched fingerprint should not be served. The bug is that a refresh produces a mismatch at all, so this fixes the fingerprint rather than weakening the guard. Both tests still pass unmodified.

Compatibility

Caches written by earlier versions carry an access-token hash, so the first run after upgrading sees one mismatch and refetches once, then stores the new fingerprint. No user-visible effect beyond a single extra request.

Tests

Three added to the fetchUsageData error handling probe suite:

  • serves a fresh cache after the access token was refreshed (same login) — refresh token unchanged, access token reissued; cache is served with requestCount 0.
  • serves a stale cache through a rate-limit backoff after the access token was refreshed — the reported scenario: cache past CACHE_MAX_AGE, active rate-limited lock, refreshed access token. Asserts the snapshot renders and no request is made (the backoff is still honored).
  • refetches when the refresh token changes (account switch) — a different login still invalidates immediately.

bun test src/utils/__tests__/: 791 pass / 3 fail, against a 788 pass / 3 fail baseline on an unmodified checkout — same three failures, all environment-dependent on my Windows machine (saves through a symlinked settings file..., silences child stderr on best-effort probes..., and preserves root errors within a process... timing out at 5000 ms). bun run lint (tsc + eslint) clean.

Related

#534 — different root cause (placeholder parsing of an unused model-scoped quota), same user-visible text. That one is about a field never reaching the cache; this one is about a cache that has the field being thrown away. They are independent and do not conflict.

🤖 Generated with Claude Code

The cache fingerprint hashes the OAuth access token, which is reissued on
a short cycle (observed ~8h expiry), so an ordinary refresh of the very
same account changes it and the cached snapshot is discarded as if the
account had switched.

That is worst exactly when the cache is load-bearing. While a fetch
failure is being backed off, getStaleUsageOrError finds the fingerprint
mismatched, drops an otherwise valid snapshot and returns error text for
the whole window - observed as [Rate limited] on every usage widget for a
server-issued Retry-After of 3600s, with complete usage data sitting in
usage.json the entire time.

Fingerprint the refresh token instead: it identifies the login rather than
the session (observed ~13 day expiry, two orders of magnitude longer), so
the account-switch guard stays intact without churning on every refresh.
It falls back to the access token when absent, keeping older credential
files and access-token-only keychain entries working as before.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@sirmalloc
sirmalloc merged commit 282c5a3 into sirmalloc:main Sep 17, 2026
3 checks passed
pcvelz added a commit to pcvelz/ccstatusline-usage that referenced this pull request Sep 25, 2026
Upstream: TUI kept off the render path (sirmalloc#575), flex mode default full (sirmalloc#590), usage cache fingerprinted by refresh token (sirmalloc#536), CLAUDE_CONFIG_DIR keychain credential first (sirmalloc#573), llms.txt (sirmalloc#527), faster terminal width probing (sirmalloc#501), git/jj symbol slots (sirmalloc#574), model-scoped 0% quota as real zero (sirmalloc#534), custom-command output cache + timeout (sirmalloc#539), usage-percent widgets on a shared module (sirmalloc#545), hideable reset-timer placeholders (sirmalloc#542), git command timeouts (sirmalloc#559, sirmalloc#585)

Hand edits outside conflicts:
- src/widgets/shared/usage-percent-widget.ts: compat fix - pass RenderContext to getUsageProgressBarWidth (fork narrow/medium bar widths) and add fork short labels WS:/WO: that the extracted Sonnet/Opus widgets used to render; point the fable-weekly kind at the fork field weeklyFableUsage / resolveWeeklyFableUsageWindow (upstream's fableUsage / resolveFableUsageWindow do not exist in the fork and crashed the render)
- .fork-keep-deleted: drop llms.txt (points agents at the upstream package and at docs the fork deletes)

Conflict resolutions that deviate from upstream on purpose:
- src/types/Settings.ts: keep fork default flexMode full-minus-40 (upstream sirmalloc#590 switched to full)
- src/utils/terminal.ts: keep fork tmux $TMUX_PANE width probe, ported to execFileSync; upstream's CCSTATUSLINE_WIDTH override in getTerminalWidth replaces the fork copy
- src/ccstatusline.ts: keep getTerminalWidth() without the per-session width cache options (fork invariant); add upstream customCommandCacheTtlSeconds
- src/widgets/WeeklyFableUsage.ts (+ test): keep fork widget (Fable:/F: labels, 0% for accounts without Fable, weeklyFableUsage field) instead of upstream's shared-module FableWeeklyUsage
- src/utils/__tests__/usage-fetch.test.ts: adopt upstream sirmalloc#534 real-zero semantics on the fork field weeklyFableUsage
- docs/test-retirements.md: ledger entries for four tests upstream renamed or replaced (sirmalloc#542, sirmalloc#534) and one duplicate upstream test dropped
- src/utils/__tests__/usage-fetch.test.ts: compat - upstream's model-scoped real-zero test expects the fork field weeklyFableUsage
- src/widgets/__tests__/WeeklyFableUsage.test.ts: compat - set the shared suite's new required expectedWholePercentTime
- src/tui/components/__tests__/ImportPreviewDialog.test.ts: fork deviation - default flexMode is full-minus-40, so the non-default side is full
- src/utils/__tests__/usage-fetch.test.ts: drop upstream's duplicate "missing fable window" test that used the upstream-only fableUsage field
thariman added a commit to thariman/ccstatusline that referenced this pull request Sep 27, 2026
Conflict resolutions:
- usage-fetch: adopt upstream's CLAUDE_CONFIG_DIR keychain service
  lookup (sirmalloc#573), which supersedes the fork's getConfigDirKeychainService;
  keep the fork's per-config-dir usage cache/lock paths, now keyed on
  upstream's cacheIdentity fingerprint (sirmalloc#536)
- usage tests: take upstream's usage-token tests; point upstream's new
  cache-seeding helpers at the fork's per-config-dir cache filenames
- package.json: keep @thariman/ccstatusline, bump to 2.2.32
- README: keep fork notes alongside upstream's v2.2.29-v2.2.30 notes

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01EchK7KtaafjsNhMkNamXqj
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants