Skip to content

Six contested diffs in the 2.x ledgers have no pin, and the empty rosters read as if none existed (MD, PHD goes to whichever rule comes first) #501

Description

@derek73

Rationale

_CROSS_RULE_WINNERS pins which rule classifies a contested name, because nothing
else in the suite can see a name change hands: per-rule rosters measure a rule
alone, and the gate total is per-corpus
(mechanisms.md#CROSS-RULE-OUTCOME-PINS). A pure file reorder in the 1.4 ledger
fails that roster and nothing else.

Both 2.x ledger sections are empty, and until now that read as "nothing here
is contested". It is not. Measured 2026-09-03 against the real pinned wheels,
six diffs are admitted by more than one rule, with file order alone deciding
the winner and nothing recording it:

ledger name measured diff rules admitting
2.0.0 田中さん 様. {family, given, suffix} 2
2.0.0 김민준 박사님 {family, given, suffix} 2
2.0.0 선생님 {family, given} 2
2.0.0 田中さん, 様. {family, given, suffix} 2
2.0.0 MD, PHD {suffix, title} 2
2.1.0 MD, PHD {suffix, title} 2

For contrast, the 1.4.0 ledger has 48 such diffs and 31 pinned rows; 2.2.0 has
none and is correctly empty.

How this surfaced

#497 moved the recorded diff shapes into tools/differential/compare.py as
_RECORDED_DIFFS so a run could verify them. Verifying them found that all four
2.x rows recorded shapes no run produces — 'Nguyen, Van' diffs at no baseline
at all, and 'Jane née and Jones Smith' diffs four roles where three were
recorded. Both were removable rather than correctable, because neither name is
contested at a shape any run makes, so the rows adjudicated nothing.

Deleting them left both 2.x sections empty, which is what exposed this: the
sections were never a statement that nothing is contested, and nobody had
measured whether anything was.

What pinning these needs

Not mechanical. Each winner has to be checked against the winning rule's prose
before it is recorded — a rule that classifies a diff it does not describe is
#372, and it stays green everywhere else. That is the same per-pair judgement
#382's eleven precedes_narrower exemptions each needed, and the CJK four sit in
the glued-honorific neighbourhood that took several review rounds to get right
there.

Note also that a recorded row is only as good as the shape beside it. #497's
whole finding is that a shape nothing measures agrees with itself forever, so
each of these six should be recorded with a shape taken from a run, and
_RECORDED_DIFFS is now the place that holds it and the run is what checks it.

Recompute

Replay main()'s load for the baseline (glob → tier order → shape-to-order →
dedupe on (name, order) → baseline-minimum skip), run compare._run_worker
over the surviving entries, compute each name's diff the way main() does, then
count the rules whose name_regex reaches the name and whose fields are a
superset of that diff. More than one is a contest.

Not in scope

Whether file order picks the RIGHT winner for these six. That is the question
#382 answered for the 1.4 ledger's pairs, and it is a separate argument from
recording what wins today.

Activity

  1. added this to the v2.3 milestone on Sep 3, 2026
  2. self-assigned this
    on Sep 3, 2026
  3. derek73 commented on Sep 3, 2026

    @derek73
    OwnerAuthor

    Framing correction — the empty sections are a stated position, not an oversight,
    and this issue should be read as asking whether that position should change.

    The issue body above calls the empty 2.x rosters a gap. The roster's own comment
    takes a narrower and more defensible line, and it is worth quoting since it is
    the thing this issue is actually questioning:

    None of those boundaries is argued about in this file, and this roster pins the
    arguments this file makes — so a row is owed when someone argues one, not
    before.

    That is a real position. _CROSS_RULE_WINNERS is not an inventory of every
    contested name — the same comment records that hundreds of corpus names are
    reachable by two or more rules in the same tier (301 in the 1.4 tier as of #414,
    up from 42), and pins 31. It pins the boundaries the file argues about, because
    a pin is only as good as someone having reasoned about which rule should win.

    So the question here is not "why were these six left out". It is: now that the
    six are measured rather than merely unexamined, does that change what is owed?

    Two honest answers:

    The comment now cites this issue at that sentence, so a reader meets the question
    where the position is stated. Nothing in the tree presumes the answer.

    Also correcting two claims I made while filing this, both wrong and both caught
    by measurement:

    • I said the 2.x ledgers had no contested name. They have six; that is this
      issue.
    • I justified deleting the four rows as "they pinned a one-rule race". Measured,
      13 of the 31 surviving 1.4.0 rows are also one-runner races and are kept
      deliberately — they pin shapes runs actually make, so widening a fields or
      moving a rule hands the name over and the guard says so. The property that
      distinguished the four deleted rows is narrower: no run makes their shapes at
      all. The roster comment states it that way, with a warning not to follow the
      looser reasoning.
  4. derek73 commented on Sep 5, 2026

    @derek73
    OwnerAuthor

    Scoping down, not closing — the six contested diffs now carry measured shapes, and no winner.

    PR #504 adds a second roster beside _RECORDED_DIFFS: compare._WATCHED_DIFFS, for a diff shape recorded with no winner pinned. The six contests this issue measured are in it:

    ledger name recorded shape severity
    2.0.0 田中さん 様. {family, given, suffix} fatal (contract tier)
    2.0.0 김민준 박사님 {family, given, suffix} fatal (contract tier)
    2.0.0 선생님 {family, given} fatal (contract tier)
    2.0.0 田中さん, 様. {family, given, suffix} printed, non-fatal (radar tier)
    2.0.0 MD, PHD {suffix, title} printed, non-fatal (radar tier)
    2.1.0 MD, PHD {suffix, title} printed, non-fatal (radar tier)

    Every run now checks each against the diff it measures, so a shape that moves is reported (MOVED SHAPE / MOVED SHAPE (radar)). What this does not guard is the thing this issue is about: a handover — file order still picks the winner, and only a _CROSS_RULE_WINNERS pin sees a name change hands. Recording a shape without a winner was chosen precisely so the roster does not pin an argument nobody has made (the framing comment above), and the _CROSS_RULE_WINNERS 2.x sections stay empty as a stated position.

    So the question here is unchanged: whether a measured-but-unargued contest should be pinned. decisions.md's watched-shapes arc records the decline of pinning them in this bundle and why. The day one of these boundaries is argued, the row moves from _WATCHED_DIFFS to _RECORDED_DIFFS with its winner.

  5. derek73 commented on Sep 5, 2026

    @derek73
    OwnerAuthor

    Correction to the table above, after PR #505: 田中さん 様. is now radar tier — its trailing ASCII period put it in the same listing-artifact class as its comma twin, and the 2026-09-01 sweep's criterion had missed the class. Its watched row prints MOVED SHAPE (radar) and cannot fail the run. Two of the four 2.0.0 contests remain contract-tier and fatal: 김민준 박사님 and 선생님.

  6. derek73 commented on Sep 5, 2026

    @derek73
    OwnerAuthor

    Closed by PR #506: all six pinned, one after a narrowing, and the position changed.

    The six were argued per name against each admitting rule's prose and the measured field values at the pinned wheel. Three were already argued — in expected_since_2.0.0.toml's own comment, which names the names — and one of those arguments was wrong: the glued-peel rule won 김민준 박사님 by matching 님 inside 박사님, while measured nothing is peeled at all. That rule's name_regex is narrowed in both ledgers that ship it ((?<!박사)(?<!선생)(?<!교수)), one name changes hands per ledger, totals byte-identical, and the spaced rule — whose title is what happens — is pinned. The 様. pair pins the glued rule (both mechanisms fire; it describes the family token shrinking and names both strings); 선생님 pins the order rule (the whole diff, no peel); MD, PHD pins the credential-only rule at both baselines, the same pair and winner as its 1.4.0 row.

    So the answer to "does measuring change what is owed": no — but finding the argument does, and the criterion was always "argued anywhere", not "argued in the guard module"; a ledger comment's claim about which rule wins is precisely what nothing but a roster row can check. mechanisms.md#CROSS-RULE-OUTCOME-PINS now says so.

    Two things the adjudication surfaced, recorded in decisions.md: the narrowing dissolved two of the six contests (박사님, 선생님 are one-runner rows now, kept pinned as 13 of the 1.4.0 rows are); and a regex accident in a narrow-first or non-nested pair has no declaration site — precedes_narrower is wide-first-only — which is how a false label sat on a contract-tier name and stayed green for a month. #498's non-nested blindness is one face of that shape.

  7. added a commit that references this issue on Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions