Skip to content

Case-table bookkeeping: 29 CJK honorific rows are classified single-issue where they depend on #271 too #324

Description

@derek73

tools/differential/expected_since_1.4.0.toml describes the CJK honorific rules as compounding with the #271 order flip. tests/v2/cases.py labels the same rows single-issue (fix(#307), fix(#308)).

Compound notation exists and is already used — feat(#273) + fix(#271), twice, in the nickname block for precisely this shape (one issue enables, #271 orders). So the honorific block is inconsistent with both the toml and its own table's precedent.

Rows like ko_honorific_ssi (김민준 씨) depend on #271's segmentation and family-first order to reach their asserted fields, and say fix(#307) alone. fix(#271) + fix(#307) would be strictly more honest.

This is a block-wide sweep rather than a one-row edit, which is why #320 left it alone rather than fixing one row into a different inconsistency.

Found during the #320 review. Does not affect behavior; it misleads a reader tracing why a row lands where it does.

Re-verified 2026-08-09 against master (d910709)

The count in the title is still exact, and the reasoning is now better supported than when filed.

  • 29 rows still classified single-issue: 10 × fix(#307), 19 × fix(#308).
  • The compound precedent still exists, still exactly twice: feat(#273) + fix(#271).
  • Both ledgers still use the compounding language the labels contradict.
  • ko_honorific_ssi asserts {family: 김, given: 민준, suffix: 씨}, and its own note now states the dependency the label omits — "the name still segments (suffix classification runs after the script_segment stage, which only ever saw 김민준)". Reaching family: 김 / given: 민준 requires Unspaced Chinese/Korean names: surnames constants + longest-match segmentation #271's hangul segmentation. The row documents in prose what its classification= field leaves out.

Nothing enforces the correspondence between cases.py classifications and ledger issue strings — tools/differential/README.md describes it as a convention, and no test checks it. So the sweep cannot break a ledger guard, and it is independent of #328 (which is about a ledger rule's reach exceeding its prose, a latent safety hole, rather than a label understating its dependencies).

What was removed from this issue

A second item — ko_honorific_after_comma's note naming is_suffix_lenient where the real gate is _group._is_suffix_piece — has been fixed and is dropped. The note now reads:

group's _is_suffix_piece diverts this one because 씨 is a single-token piece carrying vocab:suffix. NOT the lenient comma gate, which an earlier note named: measured, lenient_comma_suffixes=False leaves this row unchanged.

It names the right gate, records that an earlier note was wrong, and measures the counterfactual. (is_suffix_lenient does still exist, at nameparser/_pipeline/_assign.py:351 — the old note named a real function, just not the one that decides this row.)

Activity

  1. added
    docsDocumentation fixes and updates
    on Aug 2, 2026
  2. derek73 commented on Aug 5, 2026

    @derek73
    OwnerAuthor

    Part 2 (the wrong gate name) is fixed on master in 88ca0b3. Part 1 remains, and the issue is retitled to cover only it.

    Recording that part 2's repro no longer reproduces. The issue argued the gate was misnamed because "before #320, is_suffix_lenient("씨.") was already True and 씨. still went to given". #320 has since landed and 김민준, 씨. now gives suffix 씨., so anyone trying that demonstration will not see it.

    The conclusion was right anyway, by a negative control instead:

    김민준, 씨                      →  family 김민준, suffix 씨
    김민준, 태호                     →  family 김민준, given 태호      ← default post-comma reading
    lenient_comma_suffixes=False   →  row unchanged                ← so the lenient gate is not what admits it
    

    씨 classifies with vocab:suffix + vocab:suffix-word, the structure is FAMILY_COMMA, and _group._is_suffix_piece (_group.py:72-79) diverts the single-token piece out of the given-name position. lenient_comma_suffixes is not on the path.

    Worth keeping in mind for part 1: the wrong name was plausible because both gates are real, both concern suffixes, and both sit on the post-comma side — the note described the mechanism that genuinely handles the neighbouring Latin case (John Ingram, V). Checking these by reading will not separate them; flipping the policy the claim depends on will.

    Scope of what's left. 29 rows carry classification="fix(#307)" or "fix(#308)"; 2 rows in the whole table use compound notation today, both in the nickname block. So this is not 29 typos — it is a convention that never propagated out of the block where it was invented. It wants a stated rule (every row whose asserted fields depend on #271's segmentation and its order flip gets fix(#271) + ...) applied as a sweep, with the dependency measured per row rather than assumed from the row's shape. Fixing rows one at a time creates a third inconsistency, which is why #320 declined to.

  3. changed the title [-]Case-table bookkeeping: compound classifications and one note naming the wrong gate[/-] [+]Case-table bookkeeping: 29 CJK honorific rows are classified single-issue where they depend on #271 too[/+] on Aug 5, 2026
  4. added this to the v2.2 milestone on Aug 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    docsDocumentation fixes and updates

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions