Skip to content

Glued honorific: "田中さん, V." does not peel where "田中さん, PhD" does #319

Description

@derek73

Same credential, three spellings, two answers:

parse("田中さん, PhD")     # title PhD, family 田中, suffix さん      ← peels
parse("田中さん, V.")      # family '田中さん', given 'V.'            ← does not
parse("田中さん, Ph. D.")  # family '田中さん', suffix 'Ph. D.'       ← does not

Not a regression — all three are byte-identical to 1.4.0 and to pre-#312 — but #312 fixed the first and left its siblings, so which behavior you get depends on whether the post-comma token passes an initial veto.

Why

#312 scoped the peel's site to the name-bearing runs: segments[:2] under FAMILY_COMMA, segments[0] otherwise. That rests on segments[1] being name text under a family comma. It usually is — but not always.

_segment.py picks SUFFIX_COMMA only when suffixy(groups[1]) and len(groups[0]) > 1. When the post-comma part is entirely suffix-shaped and the pre-comma part is a single word, the structure falls through to FAMILY_COMMA with segments[1] holding post-nominals rather than given-name text. _assign agrees with that reading — it runs _peel_leading_titles over that run and routes it to title/suffix.

So the peel walks into the one kind of run it was scoped away from, and the gap it was scoped against bites:

  • segment() admits a post-comma run on is_suffix_lenient
  • _is_post_nominal asks is_suffix_strict
  • the initial-shaped suffix words fall between them — measured, V., V, I are lenient-True and strict-False

An initial-shaped token therefore becomes the peel site, ends in no listed tail, and the peel is silently abandoned. Adding a word before the comma (Dr 田中さん, V.) flips the structure to SUFFIX_COMMA and it peels correctly.

Why it was not fixed in #318

The clean condition is "include segments[1] only when it is actually name text", which is what segment()'s own suffixy already decides. But suffixy is a closure inside segment(), and its comment warns it must stay in sync with group's _PH/_D merge — so a copy in _script_segment would be a third definition of one rule.

The cheap substitute was measured and rejected: all(is_suffix_lenient) fixes V. and Jr. but not Ph. D. (because Ph. alone is not lenient-suffix), leaving PhD and Ph. D. inconsistent — trading one arbitrary split for another.

田中さん, V. is pinned as-is by a parity case row in #318 so the limit is recorded rather than latent.

What the fix probably looks like

Extract suffixy from segment() into _pipeline/_vocab.py alongside is_suffix_lenient/is_suffix_strict, so segment, group and script_segment share one definition of "this run is post-nominals, not name". Then the peel's scope becomes: include the second run under a family comma unless suffixy says it is a suffix run.

Care needed on the Ph./D. merge — suffixy treats an adjacent pair as one unit, and group has a matching merge that the comment says must not drift.

Worth doing alongside anything else that wants the same predicate; it is the third caller that makes the extraction pay.

Found by whole-branch review on #318 (#312).

Activity

  1. added this to the v2.1 milestone on Aug 2, 2026
  2. self-assigned this
    on Aug 2, 2026
  3. derek73 commented on Aug 2, 2026

    @derek73
    OwnerAuthor

    Updated by #320, and measured

    Three things have changed since this was written.

    The gap this issue is built on has shrunk, and what's left is correct. The body says V., V, I fall between is_suffix_lenient and is_suffix_strict. Still true — but #320 removed the other half. Measured across all 652 default vocabulary entries in every case variant, the two predicates now disagree on exactly seven forms:

    2.   I   I.   V   V.   i.   v.
    

    Before #320 the set also held 様. 殿. 氏. 군. 님. 씨. 양., and those were a bug. What remains is roman I/V and a digit — precisely where the initial veto is doing its intended job, since Smith, I. really should read as an initial.

    So the framing shifts. The gap is not the defect and closing it is not the fix. The defect is that the peel re-derives suffix-ness token by token when the question it needs answered — is this run name text? — has already been decided by segment(). That makes the "extract suffixy" plan below the right one, for a slightly different reason than originally stated.

    The fix is measured, not predicted. Prototyped against current master:

    input before after
    田中さん, PhD family 田中, suffix さん unchanged
    田中さん, V. family 田中さん, given V. family 田中, suffix さん, given V.
    田中さん, Ph. D. family 田中さん, suffix Ph. D. family 田中, suffix さん, Ph. D.
    김, 민준씨 family 김, given 민준, suffix 씨 unchanged

    Exactly one case row fails — ja_honorific_glued_family_comma_suffixy_second_run, the row #318 added to record this as a known limit. The fix moves precisely the parked behavior and nothing else. Differential stays unexplained: 0.

    One correction to the body: "it is the third caller that makes the extraction pay" — there are two callers, segment() and the peel. _group's _PH/_D merge is a related duplication but not a second copy of suffixy itself, so extraction does not consolidate it.

    How real is the input?

    Worth recording, because it bounds what this is worth. CLDR's sorting order is the Surname, Given shape. ja, ko and zh author 5, 6 and 6 sorting patterns and use a comma in none of them; en, ru and th use one in every pattern they author. The comma signals inversion, and CJK is already family-first, so there is nothing to invert.

    A comma in a CJK name is therefore a Western-system artifact rather than a native convention. That does not make the work pointless — VIAF, library authority files and HR systems all record CJK names in Surname, Given form regardless, and #312's 김, 민준씨 is exactly that shape. But 田中さん, V. is rare even among forms that are themselves imported, which argues for keeping this fix small rather than growing it into comma-handling work.

    Also noted, out of scope: CLDR keeps generation (Jr., III) and credentials (PhD, MD) as separate fields where this library has one suffix, and formats them differently — every en pattern reads {generation}, {credentials}. That bears on #296 and #291 and is worth a 2.2 investigation, but it is an API-surface change and does not belong in this fix.

  4. added a commit that references this issue on Aug 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions