Repository navigation
Glued honorific: "田中さん, V." does not peel where "田中さん, PhD" does #319
Description
Activity
Updated by #320, and measured
Three things have changed since this was written.
The gap this issue is built on has shrunk, and what's left is correct. The body says
V.,V,Ifall betweenis_suffix_lenientandis_suffix_strict. Still true — but #320 removed the other half. Measured across all 652 default vocabulary entries in every case variant, the two predicates now disagree on exactly seven forms:2. I I. V V. i. v.Before #320 the set also held
様. 殿. 氏. 군. 님. 씨. 양., and those were a bug. What remains is romanI/Vand a digit — precisely where the initial veto is doing its intended job, sinceSmith, I.really should read as an initial.So the framing shifts. The gap is not the defect and closing it is not the fix. The defect is that the peel re-derives suffix-ness token by token when the question it needs answered — is this run name text? — has already been decided by
segment(). That makes the "extractsuffixy" plan below the right one, for a slightly different reason than originally stated.The fix is measured, not predicted. Prototyped against current master:
input before after 田中さん, PhDfamily 田中, suffix さん unchanged 田中さん, V.family 田中さん, given V. family 田中, suffix さん, given V. 田中さん, Ph. D.family 田中さん, suffix Ph. D. family 田中, suffix さん, Ph. D.김, 민준씨family 김, given 민준, suffix 씨 unchanged Exactly one case row fails —
ja_honorific_glued_family_comma_suffixy_second_run, the row #318 added to record this as a known limit. The fix moves precisely the parked behavior and nothing else. Differential staysunexplained: 0.One correction to the body: "it is the third caller that makes the extraction pay" — there are two callers,
segment()and the peel._group's_PH/_Dmerge is a related duplication but not a second copy ofsuffixyitself, so extraction does not consolidate it.How real is the input?
Worth recording, because it bounds what this is worth. CLDR's
sortingorder is theSurname, Givenshape.ja,koandzhauthor 5, 6 and 6 sorting patterns and use a comma in none of them;en,ruandthuse one in every pattern they author. The comma signals inversion, and CJK is already family-first, so there is nothing to invert.A comma in a CJK name is therefore a Western-system artifact rather than a native convention. That does not make the work pointless — VIAF, library authority files and HR systems all record CJK names in
Surname, Givenform regardless, and #312's김, 민준씨is exactly that shape. But田中さん, V.is rare even among forms that are themselves imported, which argues for keeping this fix small rather than growing it into comma-handling work.Also noted, out of scope: CLDR keeps
generation(Jr., III) andcredentials(PhD, MD) as separate fields where this library has onesuffix, and formats them differently — everyenpattern reads{generation}, {credentials}. That bears on #296 and #291 and is worth a 2.2 investigation, but it is an API-surface change and does not belong in this fix.- added a commit that references this issue
on Aug 3, 2026 - added 10 commits that reference this issue
on Aug 22, 2026
Same credential, three spellings, two answers:
Not a regression — all three are byte-identical to 1.4.0 and to pre-#312 — but #312 fixed the first and left its siblings, so which behavior you get depends on whether the post-comma token passes an initial veto.
Why
#312 scoped the peel's site to the name-bearing runs:
segments[:2]underFAMILY_COMMA,segments[0]otherwise. That rests onsegments[1]being name text under a family comma. It usually is — but not always._segment.pypicksSUFFIX_COMMAonly whensuffixy(groups[1]) and len(groups[0]) > 1. When the post-comma part is entirely suffix-shaped and the pre-comma part is a single word, the structure falls through toFAMILY_COMMAwithsegments[1]holding post-nominals rather than given-name text._assignagrees with that reading — it runs_peel_leading_titlesover that run and routes it to title/suffix.So the peel walks into the one kind of run it was scoped away from, and the gap it was scoped against bites:
segment()admits a post-comma run onis_suffix_lenient_is_post_nominalasksis_suffix_strictV.,V,Iare lenient-True and strict-FalseAn initial-shaped token therefore becomes the peel site, ends in no listed tail, and the peel is silently abandoned. Adding a word before the comma (
Dr 田中さん, V.) flips the structure toSUFFIX_COMMAand it peels correctly.Why it was not fixed in #318
The clean condition is "include
segments[1]only when it is actually name text", which is whatsegment()'s ownsuffixyalready decides. Butsuffixyis a closure insidesegment(), and its comment warns it must stay in sync withgroup's_PH/_Dmerge — so a copy in_script_segmentwould be a third definition of one rule.The cheap substitute was measured and rejected:
all(is_suffix_lenient)fixesV.andJr.but notPh. D.(becausePh.alone is not lenient-suffix), leavingPhDandPh. D.inconsistent — trading one arbitrary split for another.田中さん, V.is pinned as-is by aparitycase row in #318 so the limit is recorded rather than latent.What the fix probably looks like
Extract
suffixyfromsegment()into_pipeline/_vocab.pyalongsideis_suffix_lenient/is_suffix_strict, sosegment,groupandscript_segmentshare one definition of "this run is post-nominals, not name". Then the peel's scope becomes: include the second run under a family comma unlesssuffixysays it is a suffix run.Care needed on the
Ph./D.merge —suffixytreats an adjacent pair as one unit, andgrouphas a matching merge that the comment says must not drift.Worth doing alongside anything else that wants the same predicate; it is the third caller that makes the extraction pay.
Found by whole-branch review on #318 (#312).