Repository navigation
Rai is parsed as a post-nominal suffix, consuming a common South Asian surname #342
Description
Activity
- added a commit that references this issue
on Aug 16, 2026 - added 2 commits that reference this issue
on Aug 22, 2026 Two things for whoever picks this up, both from the #291 close-out (2026-08-23).
1. This issue's collision is not a one-off — carry the measurement
raiis not an isolated bad entry. Measured across the whole set:- 575 of 579 alphabetic
SUFFIX_ACRONYMSentries leavefamilyempty in"John <word>". - Exactly 4 are ambiguous-gated:
ma,do,ed,jd. - Other un-gated entries that are borne as surnames:
ba(West African, Vietnamese),cha(Korean 차 — the corpus hasAhmad Jayadi, CHA),sa,se,om,mc.
That is consistent with this issue's own provenance note — the list arrived in
af5bdabas a bulk Wikipedia import and was never reviewed against surname collisions. Whichever of option A or B is chosen forrai, the choice is a precedent for the others, so it is worth stating the criterion rather than decidingraialone.decisions.md#vocabulary-collisions(C-i, with the positional qualifier from #360) is the criterion that already exists.The
ba/cha/sa/se/omclaims are model recall, not corpus-attested — per the #360 lesson, they need a human before they move.2. Please carry this
docs/design/decisions.mdamendment#291 closed working-as-designed, and the comma-suffix-arc Declined entry needs an amendment rather than a deletion. Suggested text for the entry beginning "Splitting the dead multi-word suffix entries into single-word entries":
Amended 2026-08-23 (#291 closed working-as-designed): the decline stands as measured — splitting into single words does cost
Smith, A.P.,John LeedandMary Nicet— but the inference drawn from it did not. "A multi-word entry is inert" was read as "a multi-word credential is unreachable", and that is false: the run predicateis_wholly_suffixreassembles adjacent suffix tokens, soparse("John Smith, MD PhD").suffixhas been'MD PhD'since 1.4.0, andpsm i/psm ii— two of the seven entries called dead — parse today from their component words. The third option neither the issue nor the spec considered is to leave the shipped vocabulary alone and let callers add the component words to aLexicon, which is what was chosen.Also declined with it:
SUFFIX_PHRASESand segment-level matching (unnecessary once the run predicate is measured), and amendment A6's glued-peel question (no phrase-matching unit, so_is_post_nominal's token-level test stays correct by construction).And the Excluded (MAIDEN_MARKERS) entry for Polish "z domu" should repoint from #291 to #434 — markers have no run predicate, so that question does not dissolve the way the suffix one did.
Per Derek: this amendment rides with whatever PR takes this issue, rather than landing as a separate docs PR.
- 575 of 579 alphabetic
- added 2 commits that reference this issue
on Aug 29, 2026 Shipped in 2.3.0 via #515.
raiandchaleftSUFFIX_ACRONYMS, andbajoinedSUFFIX_ACRONYMS_AMBIGUOUSbesidema,do,edandjd.HumanName("Aishwarya Rai")reads lastRai, which is 1.4.0's own answer — the 2.x releases read suffixRaiwith no last name at all.Lala Lajpat Raireads middleLajpat, lastRai;Kim Chareads lastCha;Anna Bareads lastBawith a suffix-or-name flag.The criterion, since your comment asked for one rather than a decision about
raialone. C-i (borne as a name in the position the claim acts on) is only the entry ticket. What decides among the three answers is a comparison of frequencies: how common the word is as a borne name in the trailing position against how common it is as a credential. Where the name reading dominates, the entry is removed — Rai and Cha are far more common as surnames than RAI ("RETA Authorized Instructor") and CHA (Certified Hotel Administrator) are as credentials — and a caller who needs the credential adds it back:Lexicon.default().add(suffix_acronyms={"cha"}). Where the two are roughly balanced, the word takes the ambiguous marking and the parse reports the fork: Ba is about as common a surname as BA is a credential, thema/Mashape. Where the credential dominates, the entry stays unambiguous. Length is a correlate, not the test. This supersedes thedecisions.md#vocabulary-collisionsline "rai — Rai IS a borne surname, so it earns the marking rather than moving", which now carries a dated supersession note.Of the other words your comment listed:
batook the marking.sastays where #296's audit put it — title and suffix dual, position decides.seandomstay unambiguous; OM is the Order of Merit and neither has surname evidence worth standing behind.mcis #454's word and closed by design there. All were treated as model recall rather than corpus-attested, per the #360 lesson, and none moved without a human.The accepted cost, pinned as case rows so the reversal is one measurement away.
John Smith RAIreads middleSmith, lastRAI, and the comma forms swap ends:John Smith, RAIreads firstRAI, lastJohn Smith, andAhmad Jayadi, CHAfirstCHA, lastAhmad Jayadi— once nothing after the comma is credential vocabulary, it is an ordinary family comma. The marking has a comma cost too:Smith, BAreads firstBA, asSmith, Edalready did. #516 asks whether shape and position could recover the credential reading without a wordlist.One curiosity:
John Smith R.A.I.still reads suffix, but by accident — S3 splits a token with interior periods into its chunks and the last chunkiis a Roman numeral in the suffix words, soX.Y.I.andR.A.V.read suffix whileR.A.X.andC.H.A.do not. Nothing is pinned on it.Your
decisions.mdamendment rode in verbatim on the comma-suffix-arc Declined entry. The Excluded (MAIDEN_MARKERS) "z domu" bullet you asked to repoint to #434 turned out to still call the marker pending while it has shipped since 2026-08-26, so it now records the exclusion as lifted. Rationale:decisions.md#suffix-acronym-collisions.- added 5 commits that reference this issue
on Oct 8, 2026
HumanName("Aishwarya Rai")returns an emptylastand puts the surname insuffix:'rai'is a bare entry inSUFFIX_ACRONYMS(nameparser/config/suffixes.py:747). It entered inaf5bdab("add post-nominal list from wikipedia. fix #93") as part of a bulk import, so it was never individually reviewed against surname collisions. Rai is a common surname across Hindi- and Bengali-speaking regions, and the bare form currently wins over the surname reading.Two candidate fixes — pick when this is picked up
Option A — remove
'rai'fromSUFFIX_ACRONYMS. Fixes both the two-token and three-token surname cases. Costs bare all-capsRAIas a post-nominal.Option B — add
'rai'toSUFFIX_ACRONYMS_AMBIGUOUS(keeping it inSUFFIX_ACRONYMS). This is the mechanism that already protectsJack Ma. Keeps the bare credential, but the period-gate only rescues the two-token case, soLala Lajpat Raistays misparsed — the same limitation'ma'has today withJack Wei Ma.Verified behavior at 2.1.0
Aishwarya Railast='',suffix='Rai'❌last='Rai'✅last='Rai'✅Lala Lajpat Railast='Lajpat',suffix='Rai'❌last='Rai'✅suffix='Rai'❌Aishwarya Rai Bachchanmiddle='Rai',last='Bachchan'✅John Smith R.A.I.suffix='R.A.I.'✅John Smith RAIsuffix='RAI'✅last='RAI'❌suffix='RAI'✅John Smith, RAIsuffix='RAI'✅last='John Smith'❌last='John Smith'❌So the trade is: A fixes one more surname case, B preserves one more credential case. Neither dominates. Note
John Smith, RAIregresses under both, which may be worth understanding before choosing.The dotted
R.A.I.form is unaffected either way, and not for the reason it looks like:period_joined_vocab(nameparser/_pipeline/_vocab.py:181) splits on periods and derives the suffix from the chunki— a suffix word (the Roman numeral, as in "John Smith I") — never fromrai.Workaround
Either behavior is reachable today without a release: