Skip to content

expected_changes.toml promises the harness will fail on a Ph. D. regression; it won't #328

Description

@derek73

tools/differential/expected_changes.toml:301-311 deliberately leaves trailing Ph. D. healing unclassified, and says why:

Adding a suppression rule for it would risk masking a real regression in this exact shape, so it is intentionally left unclassified: if it ever starts diffing, the harness must fail.

It does not. Reproduced — a three-name corpus containing John Smith, Jr. Ph. D.:

corpus: 296 names; intentional diffs: 103; unexplained: 0
## fix(comma-family) lone post-comma piece routes to suffix/title, not first (8)
  'John Smith, Jr. Ph. D.'

The divergence is real — 1.4.0 gives suffix "Ph. D., Jr.", 2.x gives "Jr. Ph. D." — and it is absorbed silently by a rule whose prose is about something else entirely.

Why

fix(comma-family) is declared as:

name_regex = ","
fields = ["first", "title", "suffix"]

Any suffix-only diff on a comma-bearing name is a subset of that, so the rule claims it. classify() (compare.py:40-48) returns the first matching rule, and compare.py:69 sorts on only two tiers:

rules.sort(key=lambda r: not isinstance(r.get("name_regex"), str))

— name_regex rules ahead of fields-only rules, stable within tier. Twelve of the thirteen rules are in the name_regex tier, so for almost every rule precedence is file order.

That is what makes this hard to fix locally. A narrower rule for the Ph. D. shape could only win by being written earlier in the file — which makes file position load-bearing again, the exact thing the sort's own comment says it exists to prevent ("sort is stable, so rules within a tier keep the order they were written").

Scope, not just this shape

The problem is structural. , and / are among the declared name_regex values, so any rule sitting above a narrower one can absorb its cases. Nothing detects it: a rule that swallows more than it describes produces unexplained: 0 and looks like a pass. The Ph. D. case surfaced only because a reviewer probed the file's own written promise.

Candidate directions

  • Give classify() a real specificity order — e.g. rank by name_regex length or by an explicit priority key — so a narrow rule beats a broad one regardless of position.
  • Report the runner-up. If two rules match, say so. Silent absorption is the failure mode; visibility may be enough without reordering anything.
  • Let a rule declare itself non-absorbing — a flag meaning "only claim a diff no other rule matches."
  • Assert the promise directly — a test that adds the Ph. D. shape to a scratch corpus and requires it to come back unexplained. Narrow, but it makes this comment honest without a harness redesign.

Interim

The divergence itself is now pinned by suffix_comma_split_phd_after_another_suffix in tests/v2/cases.py (classification fix(credential-pair-order)), and the toml comment records that the case table — not the harness — is what guards it. So nothing is unguarded today; what is wrong is the file's claim about its own behavior, and the general absorption hazard behind it.

Found during the #319 review (PR #327).

Activity

  1. derek73 commented on Aug 9, 2026

    @derek73
    OwnerAuthor

    Re-verified against master at d910709 (2026-08-09): still valid, and the exposure is larger than filed.

    Escalation: the shape is in the shipped corpus now

    The original repro used a scratch three-name corpus. It no longer needs one — 'John Smith, Ph. D.' is in both corpus.jsonl and corpus_issues.jsonl today, and 'John Smith Ph. D.' is in corpus.jsonl.

    Running master's classify() against master's ledger:

    'John Smith, Jr. Ph. D.'  + a {suffix} diff  ->  fix(comma-family)      (the original repro)
    'John Smith, Ph. D.'      + a {suffix} diff  ->  fix(comma-family)      (now shipped)
    'John Smith Ph. D.'       + a {suffix} diff  ->  fix(suffix-routing)    (now shipped)
    

    So a real regression on the exact shape the ledger comment protects would be absorbed by a rule whose prose is about lone post-comma pieces, and the run would exit 0. The promise at expected_since_1.4.0.toml:393 — "if it ever starts diffing, the harness must fail" — is false for names the harness actually parses on every run.

    The mechanism is unchanged

    _sorted_rules (compare.py:102) still keys on exactly one boolean — has a name_regex or not — and classify (compare.py:462) still returns the first match. Within the name_regex tier, file order still settles everything, which is what the issue identified.

    The recent ledger guards do NOT cover this

    #333 and #350 added guards to tests/v2/test_ledger_guards.py that bound what each rule may claim, including _CORPUS_CLAIMS. Worth stating plainly that none of them helps here, because the resemblance is misleading:

    _CORPUS_CLAIMS records the number of corpus names each rule's regex reaches — 236 of 751 for fix(comma-family), and 'John Smith, Ph. D.' is already among them. A regression that makes an already-reached name start diffing changes that number by zero.

    Those guards bound what a rule can claim. This issue is about what a rule does claim in a run, and nothing measures that.

    Stale references in the description

    Three things moved after filing; the analysis is unaffected:

    • tools/differential/expected_changes.toml was renamed to expected_since_1.4.0.toml (48819a7). The cited lines 301-311 are now ~385-401.
    • fields moved to the Role vocabulary, so fix(comma-family) now reads fields = ["given", "title", "suffix"], not ["first", ...].
    • compare.py:40-48 / :69 are now :462 / :102.

    Half the shape did get a rule

    The leading case is handled: fix(leading-credential) a split 'Ph. D.' before the name stays one unit now exists, deliberately anchored to the start, with a comment saying "widening this regex would mask a regression". It is the trailing case that remains unguarded by the harness.

    The issue's "Interim" note still holds — suffix_comma_split_phd_after_another_suffix in tests/v2/cases.py pins the divergence, so nothing is unguarded today. What is wrong is the ledger's claim about its own behaviour, plus the general absorption hazard.

    On the four candidate directions

    The fourth ("assert the promise directly") is now considerably cheaper than when filed: tests/v2/test_ledger_guards.py already loads every ledger and the full corpus as module-level constants, so a test asserting that no rule claims the trailing-Ph. D. shape is a few lines and needs no harness redesign. It does not fix the structural hazard, but it would make this comment honest.

  2. added this to the v2.2 milestone on Aug 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions