Skip to content

Thai honorifics are glued to the given name (นายสมชาย → given "นายสมชาย") #317

Description

@derek73

Placeholder — filed so it isn't lost, not worked up.

Thai honorifics are conventionally written glued to the given name, and they currently end up inside it:

parse("นายสมชาย ใจดี")     # given 'นายสมชาย', family 'ใจดี'   ← นาย = Mr.
parse("นางสาวสุดา ใจดี")   # given 'นางสาวสุดา', family 'ใจดี' ← นางสาว = Miss

What already works

More than I expected, which is what makes this narrow:

parse("สมชาย ใจดี")   # given 'สมชาย', family 'ใจดี'   ← correct

Thai is given-name-first, so the positional default is already right — no script_orders entry needed. Names are spaced between given and family, so no segmentation is needed either. Thai is not in _SCRIPT_RANGES at all, and it may not need to be: #308's peel is licensed by its vocabulary rather than by the script it is written in.

So the whole gap is one mechanism: splitting a listed honorific off the front of a token.

Relationship to #308 and #312

#308 shipped the mirror of this — a listed honorific peeled off the end of a name token, routed to suffix. Thai needs the same surgery from the other side, routed to title (นาย is pre-nominal, not post-nominal), which means a separate vocabulary rather than a reuse of honorific_tails.

#312 is deciding how the existing peel is structured. Its answer matters here: if the peel becomes a sibling function inside script_segment, a head-peel joins as a third sibling; if it becomes its own stage, this belongs in that stage. Thai is the strongest evidence that a second inhabitant is coming, so it is worth weighing in #312 rather than after it.

Why the vetting may be easier here than for the Han equivalents

A leading peel is normally harder to vet than a trailing one, because the leading position is where surnames live — the argument that clears 양 in #308 is precisely that a surname LEADS, so a trailing-only gate never sees it. The Chinese familiar prefixes 老/小 (老王, 小王) fail badly on this: 小 is a common given-name character, and peeling it would cut 小明 — this repo's own worked example — in half.

Thai looks different. นาย, นาง, นางสาว are not name components in any position; they are closed-class address terms. So the "can this ever begin a name?" test may come back clean where the Han equivalents do not. That is an impression, not a vetted claim — it needs the same per-entry argument #308's vocabulary carries, from someone who reads Thai.

Open questions

  • The full entry set (นาย, นาง, นางสาว at minimum; ranks and academic titles are a separate question).
  • Whether Thai needs a _SCRIPT_RANGES entry at all, given the peel is vocabulary-licensed. Probably only if some other script-conditional behavior is wanted later.
  • Whether a head peel routes to title in every case, or whether some entries are better modelled as existing titles vocabulary once the token is split.
  • Royal and monastic titles are a much larger and more sensitive area — explicitly out of scope for this placeholder.

Raised while discussing #312; no one has asked for Thai support, so priority is unset.

Activity

  1. added a commit that references this issue on Aug 2, 2026
  2. derek73 commented on Aug 2, 2026

    @derek73
    OwnerAuthor

    Sourced answer to the initials question, if Thai ever gets a Script member

    Not the honorific gap this issue is about, but it lands on the same decision point, so recording it here.

    #320 added _policy._NO_INITIALS — the scripts whose characters cannot be an initial — plus a gate (test_every_script_is_classified_for_initials) that fails until any new Script member is classified. This issue notes that Thai "is not in _SCRIPT_RANGES at all, and it may not need to be", so the gate may never fire for Thai. But if a segmenter or an order rule ever earns Thai a member, the decision comes due, and it now has a source.

    Unicode CLDR personNames (common/main/th.xml) carries 20 locale-authored namePattern entries. None of them produces an initial. For comparison:

    locale locale-authored patterns that produce an initial
    ru 39 7 — {given-initial} {given2-initial} {surname}
    en 42 8
    zh 40 6, but initialPattern overridden to {0} — no period
    th 20 0
    ja 29 0
    ko 32 0

    So Thai belongs in _NO_INITIALS, alongside Han, hangul and kana.

    Two cautions on reading this data, both of which bit on the way to it:

    • initialPattern is the field you would reach for and it is the wrong one. root defaults it to {0}. and nearly every locale inherits (↑↑↑), so reading it naively reports that Thai, Japanese and Korean all abbreviate with a period — an artifact of nobody overriding a default, not a locale claim. What carries signal is whether the locale's own patterns use an initial.
    • Zero counts are strong evidence (a substantial authored pattern set that declines the feature); non-zero counts are weaker, because template derivation can supply them. chr shows six initial-using patterns identical in shape to en's, which looks template-derived rather than community-authored.

    CLDR is silent where its coverage tiers are thin — am has no locale-authored patterns, ii and vai no personNames block at all — so this method answers Thai but not the syllabaries generally.

  3. derek73 commented on Aug 3, 2026

    @derek73
    OwnerAuthor

    The second inhabitant arrived, and it is Japanese

    This issue asked whether the head-peel should be structured to accommodate a second inhabitant, calling Thai "the strongest evidence that a second inhabitant is coming." One has, from an unrelated direction.

    Japanese maiden markers need the same surgery. 旧姓 shipped in 2.1 as default vocabulary (#330), but it reaches only the spaced form. The form Japanese actually writes puts the marker inside brackets with a fullwidth colon, and that does not split:

    (旧姓:佐藤)  ->  '旧姓:佐藤'   ONE token
    

    Adding 旧姓: to MAIDEN_MARKERS was tried and does nothing — vocabulary matching is whole-token, and the token is the whole string. What it needs is exactly this issue's mechanism:

    นายสมชาย     ->  title นาย   + given สมชาย      (this issue)
    旧姓:佐藤    ->  marker 旧姓: + maiden 佐藤      (#329, moved here)
    

    Same operation, two vocabularies, two destination roles — title for Thai, consumed-into-maiden for Japanese.

    What that changes about the design

    A third destination, and one of them is consumption rather than assignment. This issue framed the head-peel as routing to title, needing "a separate vocabulary rather than a reuse of honorific_tails." Two consumers now, and the Japanese one does not route the peeled head anywhere — it drops it, the way _group's #274 rule drops a bare née. So the mechanism wants a peeled-head disposition (route to role R, or consume) rather than a fixed destination.

    The separator question is answered by making it part of the entry. 旧姓 alone would leave :佐藤 with the colon attached; 旧姓: leaves 佐藤 clean. Worth knowing when sizing the vocabulary format — entries may carry trailing punctuation, which the tail-peel's entries never do.

    The safety argument carries over unchanged. #308's rule is that an entry peels only where it can never END a name. The head-peel's mirror is: only where it can never BEGIN one. 旧姓: cannot. นาย/นาง/นางสาว cannot, which this issue already argued. The Chinese familiar prefixes 老/小 still cannot clear it, for the reason stated here — 小 is a common given-name character and 小明 would be cut in half.

    Status

    Not asking to widen this issue's scope — Thai remains the reason to build it. Recording the second inhabitant because it was predicted here, it arrived, and it constrains the design in one way the Thai case alone would not have surfaced: the peeled head is not always routed somewhere.

    #329 keeps the language-agnostic half it can fix without this mechanism (a tagged marker inside extracted maiden content going unconsumed).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions