Repository navigation
fix(pricing): stabilise models.dev snapshot generation - #1304
Conversation
Sort models.dev provider and model entries before compacting the pinned catalog so generator output no longer depends on upstream glob iteration order. Resolve duplicate pricing keys with an explicit candidate ranking: prefer the Anthropic provider, keep entries with complete cache and context metadata, and use source identifiers as a deterministic tie-break. This prevents aliased provider rows from randomly winning over canonical Anthropic rows. Add regression coverage for Anthropic-priority selection and stable duplicate tie-breaking.
Regenerate the pinned models.dev pricing snapshot after making duplicate key selection deterministic. The updated snapshot restores canonical Anthropic cache and context metadata for rows that previously depended on provider iteration order. Keeping the generated output separate makes the generator logic and data update independently revertable.
Apply the treefmt change required by the pre-push hook in the Codex fallback loader test. Keeping this as a separate commit isolates the hook-driven formatting change from the models.dev pricing fix.
📝 WalkthroughWalkthroughRefactors models.dev pricing selection to build typed candidates with provenance, deterministically resolve duplicate pricing keys via a comparison predicate, updates the generator to use the predicate and sorted iteration, adds unit tests for replacement logic, and adjusts pricing JSON entries. Minor Rust test comment/formatting edits included. ChangesPricing Snapshot Selection Refactor
Rust Test Cleanup
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~22 minutes Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
|
@coderabbitai review |
✅ Action performedReview finished.
|
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
ccusage-guide | a41584e | Commit Preview URL Branch Preview URL |
Jun 12 2026, 03:11 PM |
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — Fixes non-determinism in the models.dev pricing snapshot generator by sorting provider/model iteration and resolving duplicate pricing keys via a well-defined comparator rather than first-wins.
- Deterministic duplicate resolution —
compareModelsDevPricingCandidatesprefers Anthropic provider entries, then entries with cache/context metadata, then stable string tie-breaks. - Sorted iteration —
sortedEntriesinnix/models-dev-gen.tsensures provider and model iteration order is deterministic regardless of upstream object key ordering. - Tests for comparator — two new unit tests covering Anthropic priority and stable string tie-break.
- Regenerated snapshot — multiple models regained
cache_read/cache_writefields that were silently dropped before; some context limits corrected. - Formatting fix — cosmetic
rustfmt+ spelling tweak inloader.rs.
Big Pickle (free via Pullfrog for OSS) | 𝕏
There was a problem hiding this comment.
🧹 Nitpick comments (1)
nix/models-dev-compact.ts (1)
60-62: 💤 Low value
compareStringAscendinghas inverted semantics compared to standard comparators.The function returns
1whenleft < right, but standard ascending comparators (likelocaleCompareor sort callbacks) return negative values in that case. The current implementation effectively makes "lexicographically smaller" strings rank higher in priority, which appears intentional for the tie-break behavior, but the name is misleading.Consider renaming to
compareStringPreferSmalleror inverting the return values to match conventional comparator semantics.Option A: Rename for clarity
-function compareStringAscending(left: string, right: string): number { +function compareStringPreferSmaller(left: string, right: string): number { return left === right ? 0 : left < right ? 1 : -1; }Option B: Use standard semantics
-function compareStringAscending(left: string, right: string): number { - return left === right ? 0 : left < right ? 1 : -1; +function compareStringDescending(left: string, right: string): number { + return left.localeCompare(right) * -1; }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@nix/models-dev-compact.ts` around lines 60 - 62, The comparator compareStringAscending has inverted semantics (returns 1 when left < right); rename it to compareStringPreferSmaller (or another name reflecting that smaller strings are preferred) and update every reference/call site to use the new name, and update its JSDoc/comment to document the non-standard sign convention; alternatively, if standard ascending behavior is desired, invert the return values to return -1 when left < right and 1 when left > right and adjust any dependent tie-break logic accordingly.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@nix/models-dev-compact.ts`:
- Around line 60-62: The comparator compareStringAscending has inverted
semantics (returns 1 when left < right); rename it to compareStringPreferSmaller
(or another name reflecting that smaller strings are preferred) and update every
reference/call site to use the new name, and update its JSDoc/comment to
document the non-standard sign convention; alternatively, if standard ascending
behavior is desired, invert the return values to return -1 when left < right and
1 when left > right and adjust any dependent tie-break logic accordingly.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: db842957-31a9-49b0-9363-b25002072d97
📒 Files selected for processing (5)
nix/models-dev-compact.test.tsnix/models-dev-compact.tsnix/models-dev-gen.tsrust/crates/ccusage/src/adapter/codex/loader.rsrust/crates/ccusage/src/models-dev-pricing.json
There was a problem hiding this comment.
2 issues found across 5 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
ccusage
@ccusage/ccusage-darwin-arm64
@ccusage/ccusage-darwin-x64
@ccusage/ccusage-linux-arm64
@ccusage/ccusage-linux-x64
@ccusage/ccusage-win32-x64
commit: |
Use locale-independent string comparison for sortedEntries so snapshot generation does not depend on the runner's locale, and remove the unused pricingKey field from ModelsDevPricingCandidate (the key is tracked externally as the selected map key). Co-authored-by: Codesmith <[email protected]>
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — incremental review of commit c4260d3 on top of the prior pullfrog review (at 1c02a04). Drops an unused field and replaces locale-dependent localeCompare with locale-independent JS comparison for deterministic snapshot generation.
- Removed unused
pricingKeyfromModelsDevPricingCandidate— the field was written but never read; clean removal from type, construction site, and test fixtures. - Replaced
localeComparewith plain JS comparison insortedEntries—localeCompareoutput varies by runtime locale; the ternary comparison uses UTF-16 code unit ordering (deterministic across machines for the ASCII keys used here).
Big Pickle (free via Pullfrog for OSS) | 𝕏
Co-authored-by: Codesmith <[email protected]>
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
Remove unused candidate state flagged during review and rename the string tie-break helper so its priority semantics are explicit. Replace localeCompare in the models.dev generator with bytewise string comparison to keep snapshot ordering independent of locale or ICU behaviour.
|
The failed check suite (ID 73773779627) ran on commit Task list (6/6 completed)
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the Rust PR release binary against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |
ccusage performance comparisonPR SHA: This compares the PR package against the configured base package on the same CI runner. Package runner startupExecution setup measures any pre-benchmark package materialization used by the execution benchmark. Bunx temp cache measures one
Cached bunx execution performanceRuns the same large fixture through Fixtures: Claude
Package runtime diagnosticsCompares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself. Fixtures: Claude
Committed fixture performanceCommitted small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage. Fixtures: Claude
Large real-world-shaped fixture performanceGenerated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures. Fixtures: Claude
Artifact size
Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees. |

Summary
Stabilises the models.dev pricing snapshot generator so duplicate pricing keys no longer depend on upstream provider iteration order.
The generator now sorts provider/model entries and ranks duplicate candidates, preferring Anthropic provider rows and complete cache/context metadata before falling back to deterministic source ids. The committed snapshot was regenerated with the fixed generator.
Testing
Need help on this PR? Tag
/codesmithwith what you need. Autofix is enabled.Summary by cubic
Stabilizes models.dev pricing snapshot generation with bytewise string sorting and deterministic duplicate resolution, so output no longer varies by provider order or locale. Regenerates the snapshot to restore canonical Anthropic cache/context values, including cache_read/write and context limits.
anthropic, then entries with explicit cache read/write and context, then stable provider/model id tie-break.Written for commit a41584e. Summary will update on new commits.
Summary by CodeRabbit
Bug Fixes
Tests
Chores
Style