Skip to content

Add Bengali honorifics to the default vocabulary #343

Description

@derek73

Every Bengali honorific is currently read as the given name, shifting the rest of the name one position right:

>>> from nameparser import parse
>>> n = parse("ড. মুহাম্মদ ইউনূস")        # Dr. Muhammad Yunus
>>> n.title, n.given, n.middle, n.family
('', 'ড.', 'মুহাম্মদ', 'ইউনূস')

This is new coverage, not a regression — no Bengali support has ever been claimed. docs/locales.rst:17 lists the default vocabulary as "Latin, Cyrillic, Greek, Arabic and Hebrew, plus Devanagari titles"; Bengali appears nowhere. A repo-wide sweep finds zero codepoints in U+0980–U+09FF across nameparser/, tests/ and docs/.

Why the Latin fallbacks can't rescue it

Unknown Latin honorifics usually survive on one of two heuristics. Neither can fire for Bengali or Devanagari, and the reason is structural rather than incidental:

  • _PERIOD_ABBREV (nameparser/_pipeline/_assign.py:49, ^[^\W\d_]{2,}\.$) promotes a leading abbreviation to a title — this is what makes Smt. and even Xyz. parse as titles. It requires a contiguous run of 2+ letter-category codepoints. Abugidas interleave combining marks (Mn/Mc) between letters, and a mark breaks the run. প্রফেসর. has five letters and still fails, because ্ at position 2 ends the run at length 1.
  • period_joined_vocab (nameparser/_pipeline/_vocab.py:181) needs an interior period (Lt.Gov.). Honorifics have a trailing one.

On top of that, single-letter abbreviations are actively claimed by something else: ড. matches is_initial (^(\w\.|[A-Z])$ — \w matches Bengali), so it is read as a given-name initial. That reading is correct for real Bengali initials (র. কে. নারায়ণ parses right today) — it just wins here because no vocabulary contests it.

So vocabulary is the only available route. Devanagari is in the same position and escapes only because #269 supplied it: डॉ. fails _PERIOD_ABBREV for exactly the same reason, and is rescued purely by the 'डॉ' entry.

Proposal

Two classes, matching the existing TITLES / FIRST_NAME_TITLES split. Both mirror the #269 Devanagari block (nameparser/config/titles.py:769) and its reasoning that native-script forms are safe where Latin transliterations are not.

Civil honorifics — TITLES only. A single following name reads as a family name, the way Dr. does.

entry gloss
ড Dr. (abbrev; matches ড. via edge-period normalization)
ডঃ Dr. (visarga spelling)
ডক্টর Doctor (full)
শ্রী Shri (Mr.)
শ্রীমতী Shrimati (Mrs.)
জনাব Janab (Mr.)
অধ্যাপক Professor
প্রফেসর Professor (borrowed)

Renunciate honorifics — TITLES and FIRST_NAME_TITLES. Renunciation abolishes the family name, so the single following name is a religious given name, not a surname. This is the Sister Mary / Pope Francis class, and family='' is the correct answer rather than a degraded one.

entry gloss
স্বামী Swami (Ramakrishna Order and general monastic)
শ্রীল Srila (Gaudiya Vaishnava, e.g. শ্রীল প্রভুপাদ)

স্বামী is the notable one: Vivekananda was Bengali, so স্বামী বিবেকানন্দ is the native spelling of the case that motivates this whole distinction.

No Latin twins, for the #269 reason: shri/sri collide with real given names (Sri Mulyani), the native-script forms cannot. Latin transliterations belong in an opt-in hi/bn pack — #345.

Deliberately excluded

Verified at 2.1.0

All cases fix, with no regression to Bengali initials or to Tagore:

input today with proposal
ড. মুহাম্মদ ইউনূস given=ড. ❌ title=ড., given=মুহাম্মদ, family=ইউনূস ✅
ডঃ মুহাম্মদ ইউনূস given=ডঃ ❌ title=ডঃ ✅
শ্রী অমর্ত্য সেন given=শ্রী ❌ title=শ্রী, given=অমর্ত্য, family=সেন ✅
শ্রীমতী মমতা ব্যানার্জী given=শ্রীমতী ❌ title=শ্রীমতী ✅
জনাব আবুল কালাম given=জনাব ❌ title=জনাব ✅
অধ্যাপক আনিসুজ্জামান given=অধ্যাপক ❌ title=অধ্যাপক, family=আনিসুজ্জামান ✅
স্বামী বিবেকানন্দ given=স্বামী, family=বিবেকানন্দ ❌ title=স্বামী, given=বিবেকানন্দ, family='' ✅
শ্রীল প্রভুপাদ given=শ্রীল ❌ title=শ্রীল, given=প্রভুপাদ ✅
শ্রী সেন given=শ্রী ❌ title=শ্রী, family=সেন ✅ (civil, contrast above)
র. কে. নারায়ণ given=র., middle=কে., family=নারায়ণ ✅ unchanged ✅
সত্যজিৎ রায় given=সত্যজিৎ, family=রায় ✅ unchanged ✅
রবীন্দ্রনাথ ঠাকুর given=রবীন্দ্রনাথ, family=ঠাকুর ✅ unchanged ✅

The শ্রী সেন / স্বামী বিবেকানন্দ pair is the point of the two-class split: same shape, opposite correct answers. Note the split only affects the title-plus-ONE-name case — স্বামী বিবেকানন্দ সরস্বতী is unaffected.

Open questions — decide before including

  • বেগম (Begum) — an honorific, but also appears as a name component in Bangladeshi usage, unlike শ্রী, which only occurs bound inside compounds (শ্রীকান্ত) and so cannot collide at token level. Same shape as the deferrals Provide constants in non-Latin scripts (Cyrillic, Greek, Arabic, Hebrew) #269 recorded for bare רב and בר.
  • মোঃ/মো. (Md.) — prefixes a large share of Bangladeshi male names, but reads more like a bound given-name element (Md. Abdul Karim) than a title, so BOUND_FIRST_NAMES may be the right home rather than TITLES. Worth its own analysis.

Implementation note

Three places enumerate which scripts the default vocabulary covers and will go stale otherwise:

  • docs/locales.rst:17 — the "five scripts" sentence
  • tests/v2/test_locales.py:1003 — the #269: non-Latin default vocabulary (Cyrillic, Greek, Arabic, Hebrew) section header
  • nameparser/config/titles.py:769 — the Devanagari block comment, whose "NO Latin twins on purpose" reasoning now covers Bengali too

Activity

  1. added this to the v2.2 milestone on Aug 22, 2026
  2. modified the milestones: v2.2, v2.3 on Aug 30, 2026
  3. derek73 commented on Sep 7, 2026

    @derek73
    OwnerAuthor

    Shipped in 2.3.0 via #513 — the first Bengali vocabulary in the default lexicon. Titles: ড, ডঃ, ডক্টর, ডাঃ, ডা, ডাক্তার, শ্রী, শ্রীমতী, জনাব, অধ্যাপক, প্রফেসর, বিচারপতি, মাওলানা, মুফতি, আলহাজ্ব, আলহাজ, মিঃ, মি, মিসেস, মোঃ, মো, মোসাঃ, মোসা, মোছাঃ, মোছা. Given-name titles: স্বামী, শ্রীল, গুরু, বাবা. Trailing: সাহেব, বাবু, মহারাজ in SUFFIX_WORDS, spaced only.

    The fork this issue verified at 2.1.0 holds on the shipped set: ড. মুহাম্মদ ইউনূস reads title ড., given মুহাম্মদ, family ইউনূস — the title vocabulary beats the is_initial reading — while র. কে. নারায়ণ, real initials with no entry behind them, is unchanged. Both are pinned. ডাঃ/ডা. (the Bangladeshi physician's abbreviation) ships beside ড./ডঃ (the PhD's).

    মোঃ/মো and the women's মোসাঃ/মোসা, মোছাঃ/মোছা went into TITLES rather than the bound-given-name set this issue suggested. Semantically they are name prefixes, but TITLES is the functional home: given must stay the name the person is addressed by (মোঃ আবদুল করিম is called আবদুল করিম) and the parser has no name-prefix field, while Arabic abdul is an inseparable half of one name where মোঃ is detachable. This mirrors Latin Md, whose 2015 origin in TITLES (the degree, #32) is now recorded. Every abbreviation carrying the visarga ships as two entries, the visarga form and the bare stem, because the lookup fold strips only edge periods and whitespace and never touches the mark (ঃ is a spacing combining mark), so মোঃ matches as written and মো. matches through the stem. Latin Mst waits for #345.

    বেগম is NOT shipped: it is a trailing name element far more often than a leading title. ঠাকুর stays excluded as this issue already recorded, pinned in the leading position where the exclusion is load-bearing (ঠাকুর রবীন্দ্রনাথ reads given ঠাকুর, no title), and কুমারী and শেখ join it. Deferred as trailing, then shipped with the suffix half: মহারাজ. শ্রীমৎ stays unverified.

    One documentation consequence: rules.md#H2's example of an unlisted abugida honorific used প্রফেসর, which this issue's vocabulary now lists, so the example moved to প্রকৌশলী. Sen (Engineer, deliberately not shipped). The clause itself is unchanged and still true.

    The line citation in this issue (test_locales.py:1003) was stale and was not edited; the #269: non-Latin default vocabulary section header now names Devanagari and Bengali. Rationale: the indic-honorifics entry of docs/design/decisions.md.

  4. self-assigned this
    on Sep 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions