Repository navigation
expected_changes.toml promises the harness will fail on a Ph. D. regression; it won't #328
Description
Activity
Re-verified against
masteratd910709(2026-08-09): still valid, and the exposure is larger than filed.Escalation: the shape is in the shipped corpus now
The original repro used a scratch three-name corpus. It no longer needs one —
'John Smith, Ph. D.'is in bothcorpus.jsonlandcorpus_issues.jsonltoday, and'John Smith Ph. D.'is incorpus.jsonl.Running master's
classify()against master's ledger:'John Smith, Jr. Ph. D.' + a {suffix} diff -> fix(comma-family) (the original repro) 'John Smith, Ph. D.' + a {suffix} diff -> fix(comma-family) (now shipped) 'John Smith Ph. D.' + a {suffix} diff -> fix(suffix-routing) (now shipped)So a real regression on the exact shape the ledger comment protects would be absorbed by a rule whose prose is about lone post-comma pieces, and the run would exit 0. The promise at
expected_since_1.4.0.toml:393— "if it ever starts diffing, the harness must fail" — is false for names the harness actually parses on every run.The mechanism is unchanged
_sorted_rules(compare.py:102) still keys on exactly one boolean — has aname_regexor not — andclassify(compare.py:462) still returns the first match. Within thename_regextier, file order still settles everything, which is what the issue identified.The recent ledger guards do NOT cover this
#333 and #350 added guards to
tests/v2/test_ledger_guards.pythat bound what each rule may claim, including_CORPUS_CLAIMS. Worth stating plainly that none of them helps here, because the resemblance is misleading:_CORPUS_CLAIMSrecords the number of corpus names each rule's regex reaches — 236 of 751 forfix(comma-family), and'John Smith, Ph. D.'is already among them. A regression that makes an already-reached name start diffing changes that number by zero.Those guards bound what a rule can claim. This issue is about what a rule does claim in a run, and nothing measures that.
Stale references in the description
Three things moved after filing; the analysis is unaffected:
tools/differential/expected_changes.tomlwas renamed toexpected_since_1.4.0.toml(48819a7). The cited lines301-311are now ~385-401.fieldsmoved to theRolevocabulary, sofix(comma-family)now readsfields = ["given", "title", "suffix"], not["first", ...].compare.py:40-48/:69are now:462/:102.
Half the shape did get a rule
The leading case is handled:
fix(leading-credential) a split 'Ph. D.' before the name stays one unitnow exists, deliberately anchored to the start, with a comment saying "widening this regex would mask a regression". It is the trailing case that remains unguarded by the harness.The issue's "Interim" note still holds —
suffix_comma_split_phd_after_another_suffixintests/v2/cases.pypins the divergence, so nothing is unguarded today. What is wrong is the ledger's claim about its own behaviour, plus the general absorption hazard.On the four candidate directions
The fourth ("assert the promise directly") is now considerably cheaper than when filed:
tests/v2/test_ledger_guards.pyalready loads every ledger and the full corpus as module-level constants, so a test asserting that no rule claims the trailing-Ph. D.shape is a few lines and needs no harness redesign. It does not fix the structural hazard, but it would make this comment honest.- added a commit that references this issue
on Aug 16, 2026
tools/differential/expected_changes.toml:301-311deliberately leaves trailingPh. D.healing unclassified, and says why:It does not. Reproduced — a three-name corpus containing
John Smith, Jr. Ph. D.:The divergence is real — 1.4.0 gives suffix
"Ph. D., Jr.", 2.x gives"Jr. Ph. D."— and it is absorbed silently by a rule whose prose is about something else entirely.Why
fix(comma-family)is declared as:Any suffix-only diff on a comma-bearing name is a subset of that, so the rule claims it.
classify()(compare.py:40-48) returns the first matching rule, andcompare.py:69sorts on only two tiers:—
name_regexrules ahead of fields-only rules, stable within tier. Twelve of the thirteen rules are in thename_regextier, so for almost every rule precedence is file order.That is what makes this hard to fix locally. A narrower rule for the
Ph. D.shape could only win by being written earlier in the file — which makes file position load-bearing again, the exact thing the sort's own comment says it exists to prevent ("sort is stable, so rules within a tier keep the order they were written").Scope, not just this shape
The problem is structural.
,and/are among the declaredname_regexvalues, so any rule sitting above a narrower one can absorb its cases. Nothing detects it: a rule that swallows more than it describes producesunexplained: 0and looks like a pass. ThePh. D.case surfaced only because a reviewer probed the file's own written promise.Candidate directions
classify()a real specificity order — e.g. rank byname_regexlength or by an explicitprioritykey — so a narrow rule beats a broad one regardless of position.Ph. D.shape to a scratch corpus and requires it to come back unexplained. Narrow, but it makes this comment honest without a harness redesign.Interim
The divergence itself is now pinned by
suffix_comma_split_phd_after_another_suffixintests/v2/cases.py(classificationfix(credential-pair-order)), and the toml comment records that the case table — not the harness — is what guards it. So nothing is unguarded today; what is wrong is the file's claim about its own behavior, and the general absorption hazard behind it.Found during the #319 review (PR #327).