Repository navigation
A ledger rule's explained set is printed but never checked, so a diff can shrink underneath it and every run stays green #452
Description
Activity
- added a commit that references this issue
on Aug 29, 2026 A sibling proposal from #451's review: a reach ceiling
#451 landed the ban on fields-only rules (#453), and its review surfaced a check that belongs with this issue rather than beside it.
The ban closed the shape, not the property.
validate_rulesnow requires aname_regex, but that bounds nothing by itself: the only width check is the sentinel probe, which rejects a pattern only when it matches all four of_SENTINELS. Measured over the 1090-name corpus:name_regex = "[a-z]" passes validation, reaches 941 / 1090 name_regex = " " passes validation, reaches 1027 / 1090 name_regex = "[^ -鿿]" passes validation, reaches 1063 / 1090Any of those plus a
fieldslist is the retired catch-all with a fig leaf. What #451 bought is that such a rule now carries a_CORPUS_CLAIMSreach and digest — so its breadth is visible once, to whoever reviews the recorded number, rather than never. And_CORPUS_CLAIMSis by its own docstring "inert for a brand-new rule, whose author simply records whatever number it produces". So the reviewer is the check.Why here and not its own issue. This issue is "nothing counts what a rule explains". The ceiling is "nothing bounds what a rule reaches". They are the two halves of the same gap — a rule can be too wide on either axis, and neither has a mechanism — and whoever implements the declared-reach modifier will already be in
compare.py's validation and_CORPUS_CLAIMS. Splitting them across two issues would mean two designs for one question.Sketch. No rule may claim more than N% of the corpus without an explicit roster opt-in. Measured today, the maximum is 25.6% — the two comma rules at 279/1090 — with the next highest at 9.9%. A 30% ceiling costs nothing now and is the check that is actually scoped to see this:
expected_since_1.4.0.toml fix(comma-precomma-family) … 279 25.6% expected_since_1.4.0.toml fix(comma-family) lone post-… 279 25.6% expected_since_1.4.0.toml fix(#271/#272/#298) native-… 108 9.9%Open in the design, not decided:
- Ceiling on reach, or on reach × fields? A rule reaching 279 names but narrowing to two roles is not the hazard; one reaching 279 across five roles is.
_Claimalready records both. - Where it lives.
validate_rulesfails at startup and knows nothing about the corpus;test_ledger_guards.pyhas the corpus but is a unit test. The corpus floors precedent (_CORPUS_FLOORS) puts a corpus-shaped constant incompare.py. - The opt-in's shape. The two comma rules would need one on day one, so the roster is not hypothetical — and an opt-out that nobody has to justify is the failure
dormantwas designed against.
Not urgent: no rule is near the ceiling today, and the reviewer-reads-the-number path is a real improvement over what #451 replaced. Filed so the second half of the gap is written down beside the first.
- Ceiling on reach, or on reach × fields? A rule reaching 279 names but narrowing to two roles is not the hazard; one reaching 279 across five roles is.
- added a commit that references this issue
on Aug 29, 2026 Reopening: the explained-reach half landed in #455 (merged
6ac31f4), but this issue also carries the reach ceiling proposal, which has not been built. GitHub auto-closed on aCloses #452in a commit message that was scoped in its own next clause — my bookkeeping error, not a decision.Landed:
compare.pynow fails the run when a rule's declaredfieldsdo not equal the union of the diffs it explains. Fourteen over-declared rules narrowed first (3 at 1.4.0, 5 at 2.0.0, 6 at 2.1.0), so the check was silent on arrival.Still open here: the reach ceiling — no rule may claim more than N% of the corpus without an explicit opt-in. That bounds what a rule reaches; the landed check bounds what it explains. Measurements are in the comment above: today's maximum is 25.6% (the two comma rules at 279/1090), next highest 9.9%, so a 30% ceiling costs nothing now.
- added a commit that references this issue
on Aug 29, 2026 - added 10 commits that reference this issue
on Aug 29, 2026 The reach ceiling, re-measured: two things it predates
Leaving this open — the ceiling is still an unbuilt proposal. But the sketch above dates from before #468/#482 tiered the corpora and before the differential's run cost was ever measured, and both change how it should be argued. Everything below is measured 2026-09-03 on
claude/differential-stale-claims(#497), and every figure is recomputed here rather than copied.1. The run cost, since a placement argument is about to be made on it
Measured over the whole corpus,
time (uv run python tools/differential/compare.py --baseline B >/dev/null):baseline 1.4.0 0.34s wall worker pass 0.10s baseline 2.0.0 0.46s worker pass 0.24s baseline 2.1.0 0.47s worker pass 0.26s baseline 2.2.0 0.55s worker pass 0.29s all four back to back 1.97sThe worker-pass shares are timed inside
main()— wrapcompare._run_workerin a timer and callmain(), not_run_workeron the loaded entries directly, which aborts at a 1.4.0 baseline because the seven order-bearing shape-4/5 entries have to be dropped by the baseline-minimum skip first.Nothing in this issue actually rests on that cost, and that is worth saying rather than assuming. I went looking for the argument and it is not here: the "Where it lives" bullet weighs
validate_rules(startup, no corpus) againsttest_ledger_guards.py(has the corpus, is a unit test) againstcompare.py(the_CORPUS_FLOORSprecedent) on capability, and "a 30% ceiling costs nothing now" is about headroom, not seconds. The re-derivation still lands, in two places:- The ceiling needs no worker at all. Reach is
name_regexagainst the corpus files; no wheel, no subprocess. The full four-ledger census printed below runs in 0.25s with nothing installed. So placement is free on cost either way, and the choice has to be made on something else. - That something else is coverage, and it points the other way from the
_CORPUS_FLOORSprecedent. A check incompare.pysees one ledger — the one the run was invoked with — and only when somebody runs the tool; a static one covers a rule from the moment it is written, including a rule nobody has run a wheel against. That is now the whole of the stated reason for the neighbouring decline on run-time contest detection, after the cost half of that argument turned out never to have been measured (docs/design/decisions.md, the rule-order arc'sDeclined:bullet, and the recorded-shapes arc above it). The same reasoning is available to the ceiling and was not before.
2. #482's tier split inverts the ranking the sketch ranks by
The sketch measures the ceiling as a share of the whole corpus — "25.6%, the two comma rules at 279/1090, next highest 9.9%". Since #468/#482 the corpora are two tiers and a radar diff cannot fail the run, so a rule reaching many radar names is not the hazard a rule reaching many contract names is.
docs/design/decisions.md#differential-ledger, the recorded-shapes arccarries a bullet on exactly this.Recomputed today. Tiering reads
compare._CORPUS_TIERSrather than a hand-built file→tier map —corpus.jsonlis the largest corpus at 486 names and it is radar, demoted by #468 as a v1-test-bank scrape, and reconstructing that map by hand has already produced a published wrong split in this repo. Dedup follows the run: contract files first, first-seen wins.corpus 1116 distinct names: 326 contract, 790 radar ledger max by TOTAL reach max by CONTRACT reach expected_since_1.4.0.toml 288 = 25.8% (64 contract = 19.6%) 108 = 9.7% (80 contract = 24.5%) fix(comma-precomma-family) / fix(comma-family) fix(#271/#272/#298) native-script CJK expected_since_2.0.0.toml 108 = 9.7% (80 contract = 24.5%) same rule fix(#271/#272/#298) expected_since_2.1.0.toml 27 = 2.4% (16 contract = 4.9%) same rule fix(#385/#402) expected_since_2.2.0.toml 18 = 1.6% (1 contract = 0.3%) same rule (only one name_regex rule in it) fix(#462)It inverts, in the ledger the sketch's figures come from, and it reorders in the other two that have room to. At 1.4.0 the two comma rules are the outlier by total reach (25.8%, up from 279/1090 as the corpus grew) and
fix(#271/#272/#298)is the outlier by contract reach — 9.7% of the corpus, which the sketch lists as the next highest and therefore as safely under any ceiling, but 24.5% of the contract tier, more than the comma rules' 19.6%. A ceiling on total reach ranks the wrong rule as the biggest hazard.The top rule does not change in the three 2.x ledgers — I checked rather than assumed, and 2.2.0 has one
name_regexrule so the question is degenerate there. But the same reordering happens one rank down in both of the others. The sharpest case isfix(#462)(the facade keeping an initial-shaped conjunction letter): it reaches 18 names, tied second by total reach at 2.1.0, and exactly 1 of them is contract (0.3%), whilefix(#379)reaches 13 total and 11 contract (3.4%). On the axis that can fail a run they are an order of magnitude apart, and total reach has them the wrong way round.Two consequences for the open questions:
- Two denominators, not one. Whatever N is, contract reach is the axis a run can fail on. Total reach is still worth seeing — a rule sweeping 25.8% of everything is a review question — but it is a reporting number, not the ceiling.
- The "reach, or reach × fields?" bullet. On total reach × declared roles the comma rules still lead (288×3 and 288×2 against the CJK rule's 108×3). On contract reach × roles they do not: 80×3 = 240 against 64×3 = 192. So that bullet's answer depends on which denominator is chosen first — the denominator is the prior question.
RECOMPUTE (all of the above): load
tools/differential/compare.pyby path, read_CORPUS_TIERSfrom it, build the name→tier map by loadingcorpus*.jsonlwith the contract files first, and for each[[change]]carrying aname_regexcountre.searchhits over the name list and over its contract subset.- The ceiling needs no worker at all. Reach is
- added 6 commits that reference this issue
on Oct 8, 2026
compare.pycomputesby_issue-- the exact set of names each rule explained -- prints it (tools/differential/compare.py:868), and then passes onlyset(by_issue)todormant_rules. Membership is checked; the counts never are. The only property any run asserts about a rule's reach is "it explained at least one name."The shape
classify()takes the first rule whose declaredfieldsare a superset of the observed diff. A rule therefore keeps matching when the diff beneath it shrinks -- and shrinking is the common direction, since most parser fixes move fewer fields, not more.Freiherr von Richthofen Vis the case that already happened. #410 narrowed its diff from{given, family, suffix}to{family, suffix}; thefix(#424) a title-led chain before the numeral is the one name piecerule declared all three, so it kept claiming the name and no run named the movement.docs/design/decisions.md#H1records it:That rule's
fieldsare["family", "suffix"]today, narrowed by hand during #410, and it explains exactly 1 name. The instance is closed. The mechanism that let it hide is not.Why no existing guard sees it
dormant(Make a ledger rule that explains nothing say why (#372) #373) covers one point on the scale: zero. A rule that explained 8 names and now explains 3 is not dormant, and nothing reports it._CORPUS_CLAIMSmeasures what a rule'sname_regexreaches, not what it explains -- the two differ by thefieldstest and by rule order. It is a unit test with no baseline worker, so it cannot measure explanation even in principle: the diff set only exists inside acompare.pyrun._CROSS_RULE_WINNERSpins which rule wins for a hand-picked list of contested names. A wall around known arguments, not a census.This is the gap #451 reports from the other side. #451's rule hides growth because it has no
name_regex;fix(#424)hid a shrink despite having one. Both are "a rule broader than the diff it explains", whichdecisions.md#H1already calls "the lesson worth keeping rather than either fix."Sketch
A per-rule declared reach in the ledger, checked by
compare.pyagainst the run's ownby_issue:Open in the design, not decided here:
dormant.dormantisexplains = 0carrying a reason. Whether the keys merge or sit side by side is part of the design._Claimalready documents why a bare count is weak -- "swapping feat(Support smart quotes (“Jack”), guillemets («Petit»), and CJK brackets as nickname delimiters #273)'s delimiter class for a single accented letter holds the count at 6 while claiming six entirely different names." A digest closes that, at the cost of an unreadable field in a file that is otherwise all prose.corpus_issues.jsonlis append-only from the tracker,corpus_rules.jsonlis generated fromrules.md, so one harvest moves numbers across many rules at once._CORPUS_CLAIMSand_CORPUS_FLOORSboth already accept a version of this cost, and the discipline is set: "Re-measure rather than adjust them if a parser change moves one: a diff shape that shifted is a finding, not a number to update."compare.py, nottests/v2/test_ledger_guards.py-- the guards file has no baseline worker.Relationship to #451
Independent and deliberately unbundled. #451's bundle retires the fields-only rule and lands without this; this closes the half of the shape a
name_regexdoes not protect against. Bundling them would make #451's commits unbisectable.