Skip to content

Rai is parsed as a post-nominal suffix, consuming a common South Asian surname #342

Description

@derek73

HumanName("Aishwarya Rai") returns an empty last and puts the surname in suffix:

>>> from nameparser import HumanName
>>> n = HumanName("Aishwarya Rai")
>>> n.last, n.suffix
('', 'Rai')

'rai' is a bare entry in SUFFIX_ACRONYMS (nameparser/config/suffixes.py:747). It entered in af5bdab ("add post-nominal list from wikipedia. fix #93") as part of a bulk import, so it was never individually reviewed against surname collisions. Rai is a common surname across Hindi- and Bengali-speaking regions, and the bare form currently wins over the surname reading.

Two candidate fixes — pick when this is picked up

Option A — remove 'rai' from SUFFIX_ACRONYMS. Fixes both the two-token and three-token surname cases. Costs bare all-caps RAI as a post-nominal.

Option B — add 'rai' to SUFFIX_ACRONYMS_AMBIGUOUS (keeping it in SUFFIX_ACRONYMS). This is the mechanism that already protects Jack Ma. Keeps the bare credential, but the period-gate only rescues the two-token case, so Lala Lajpat Rai stays misparsed — the same limitation 'ma' has today with Jack Wei Ma.

Verified behavior at 2.1.0

input today A (remove) B (ambiguous)
Aishwarya Rai last='', suffix='Rai' ❌ last='Rai' ✅ last='Rai' ✅
Lala Lajpat Rai last='Lajpat', suffix='Rai' ❌ last='Rai' ✅ suffix='Rai' ❌
Aishwarya Rai Bachchan middle='Rai', last='Bachchan' ✅ unchanged ✅ unchanged ✅
John Smith R.A.I. suffix='R.A.I.' ✅ unchanged ✅ unchanged ✅
John Smith RAI suffix='RAI' ✅ last='RAI' ❌ suffix='RAI' ✅
John Smith, RAI suffix='RAI' ✅ last='John Smith' ❌ last='John Smith' ❌

So the trade is: A fixes one more surname case, B preserves one more credential case. Neither dominates. Note John Smith, RAI regresses under both, which may be worth understanding before choosing.

The dotted R.A.I. form is unaffected either way, and not for the reason it looks like: period_joined_vocab (nameparser/_pipeline/_vocab.py:181) splits on periods and derives the suffix from the chunk i — a suffix word (the Roman numeral, as in "John Smith I") — never from rai.

Workaround

Either behavior is reachable today without a release:

from nameparser.config import CONSTANTS
CONSTANTS.suffix_acronyms.remove('rai')          # option A
CONSTANTS.suffix_acronyms_ambiguous.add('rai')   # option B

Activity

  1. added this to the v2.2 milestone on Aug 9, 2026
  2. derek73 commented on Aug 24, 2026

    @derek73
    OwnerAuthor

    Two things for whoever picks this up, both from the #291 close-out (2026-08-23).

    1. This issue's collision is not a one-off — carry the measurement

    rai is not an isolated bad entry. Measured across the whole set:

    • 575 of 579 alphabetic SUFFIX_ACRONYMS entries leave family empty in "John <word>".
    • Exactly 4 are ambiguous-gated: ma, do, ed, jd.
    • Other un-gated entries that are borne as surnames: ba (West African, Vietnamese), cha (Korean 차 — the corpus has Ahmad Jayadi, CHA), sa, se, om, mc.

    That is consistent with this issue's own provenance note — the list arrived in af5bdab as a bulk Wikipedia import and was never reviewed against surname collisions. Whichever of option A or B is chosen for rai, the choice is a precedent for the others, so it is worth stating the criterion rather than deciding rai alone. decisions.md#vocabulary-collisions (C-i, with the positional qualifier from #360) is the criterion that already exists.

    The ba/cha/sa/se/om claims are model recall, not corpus-attested — per the #360 lesson, they need a human before they move.

    2. Please carry this docs/design/decisions.md amendment

    #291 closed working-as-designed, and the comma-suffix-arc Declined entry needs an amendment rather than a deletion. Suggested text for the entry beginning "Splitting the dead multi-word suffix entries into single-word entries":

    Amended 2026-08-23 (#291 closed working-as-designed): the decline stands as measured — splitting into single words does cost Smith, A.P., John Leed and Mary Nicet — but the inference drawn from it did not. "A multi-word entry is inert" was read as "a multi-word credential is unreachable", and that is false: the run predicate is_wholly_suffix reassembles adjacent suffix tokens, so parse("John Smith, MD PhD").suffix has been 'MD PhD' since 1.4.0, and psm i/psm ii — two of the seven entries called dead — parse today from their component words. The third option neither the issue nor the spec considered is to leave the shipped vocabulary alone and let callers add the component words to a Lexicon, which is what was chosen.

    Also declined with it: SUFFIX_PHRASES and segment-level matching (unnecessary once the run predicate is measured), and amendment A6's glued-peel question (no phrase-matching unit, so _is_post_nominal's token-level test stays correct by construction).

    And the Excluded (MAIDEN_MARKERS) entry for Polish "z domu" should repoint from #291 to #434 — markers have no run predicate, so that question does not dissolve the way the suffix one did.

    Per Derek: this amendment rides with whatever PR takes this issue, rather than landing as a separate docs PR.

  3. modified the milestones: v2.2, v2.3 on Aug 30, 2026
  4. derek73 commented on Sep 8, 2026

    @derek73
    OwnerAuthor

    Shipped in 2.3.0 via #515. rai and cha left SUFFIX_ACRONYMS, and ba joined SUFFIX_ACRONYMS_AMBIGUOUS beside ma, do, ed and jd.

    HumanName("Aishwarya Rai") reads last Rai, which is 1.4.0's own answer — the 2.x releases read suffix Rai with no last name at all. Lala Lajpat Rai reads middle Lajpat, last Rai; Kim Cha reads last Cha; Anna Ba reads last Ba with a suffix-or-name flag.

    The criterion, since your comment asked for one rather than a decision about rai alone. C-i (borne as a name in the position the claim acts on) is only the entry ticket. What decides among the three answers is a comparison of frequencies: how common the word is as a borne name in the trailing position against how common it is as a credential. Where the name reading dominates, the entry is removed — Rai and Cha are far more common as surnames than RAI ("RETA Authorized Instructor") and CHA (Certified Hotel Administrator) are as credentials — and a caller who needs the credential adds it back: Lexicon.default().add(suffix_acronyms={"cha"}). Where the two are roughly balanced, the word takes the ambiguous marking and the parse reports the fork: Ba is about as common a surname as BA is a credential, the ma/Ma shape. Where the credential dominates, the entry stays unambiguous. Length is a correlate, not the test. This supersedes the decisions.md#vocabulary-collisions line "rai — Rai IS a borne surname, so it earns the marking rather than moving", which now carries a dated supersession note.

    Of the other words your comment listed: ba took the marking. sa stays where #296's audit put it — title and suffix dual, position decides. se and om stay unambiguous; OM is the Order of Merit and neither has surname evidence worth standing behind. mc is #454's word and closed by design there. All were treated as model recall rather than corpus-attested, per the #360 lesson, and none moved without a human.

    The accepted cost, pinned as case rows so the reversal is one measurement away. John Smith RAI reads middle Smith, last RAI, and the comma forms swap ends: John Smith, RAI reads first RAI, last John Smith, and Ahmad Jayadi, CHA first CHA, last Ahmad Jayadi — once nothing after the comma is credential vocabulary, it is an ordinary family comma. The marking has a comma cost too: Smith, BA reads first BA, as Smith, Ed already did. #516 asks whether shape and position could recover the credential reading without a wordlist.

    One curiosity: John Smith R.A.I. still reads suffix, but by accident — S3 splits a token with interior periods into its chunks and the last chunk i is a Roman numeral in the suffix words, so X.Y.I. and R.A.V. read suffix while R.A.X. and C.H.A. do not. Nothing is pinned on it.

    Your decisions.md amendment rode in verbatim on the comma-suffix-arc Declined entry. The Excluded (MAIDEN_MARKERS) "z domu" bullet you asked to repoint to #434 turned out to still call the marker pending while it has shipped since 2026-08-26, so it now records the exclusion as lifted. Rationale: decisions.md#suffix-acronym-collisions.

  5. self-assigned this
    on Sep 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions