Repository navigation
Thai honorifics are glued to the given name (นายสมชาย → given "นายสมชาย") #317
Description
Activity
- added a commit that references this issue
on Aug 2, 2026 Sourced answer to the initials question, if Thai ever gets a
ScriptmemberNot the honorific gap this issue is about, but it lands on the same decision point, so recording it here.
#320 added
_policy._NO_INITIALS— the scripts whose characters cannot be an initial — plus a gate (test_every_script_is_classified_for_initials) that fails until any newScriptmember is classified. This issue notes that Thai "is not in_SCRIPT_RANGESat all, and it may not need to be", so the gate may never fire for Thai. But if a segmenter or an order rule ever earns Thai a member, the decision comes due, and it now has a source.Unicode CLDR
personNames(common/main/th.xml) carries 20 locale-authorednamePatternentries. None of them produces an initial. For comparison:locale locale-authored patterns that produce an initial ru39 7 — {given-initial} {given2-initial} {surname}en42 8 zh40 6, but initialPatternoverridden to{0}— no periodth20 0 ja29 0 ko32 0 So Thai belongs in
_NO_INITIALS, alongside Han, hangul and kana.Two cautions on reading this data, both of which bit on the way to it:
initialPatternis the field you would reach for and it is the wrong one.rootdefaults it to{0}.and nearly every locale inherits (↑↑↑), so reading it naively reports that Thai, Japanese and Korean all abbreviate with a period — an artifact of nobody overriding a default, not a locale claim. What carries signal is whether the locale's own patterns use an initial.- Zero counts are strong evidence (a substantial authored pattern set that declines the feature); non-zero counts are weaker, because template derivation can supply them.
chrshows six initial-using patterns identical in shape toen's, which looks template-derived rather than community-authored.
CLDR is silent where its coverage tiers are thin —
amhas no locale-authored patterns,iiandvainopersonNamesblock at all — so this method answers Thai but not the syllabaries generally.- added a commit that references this issue
on Aug 2, 2026 The second inhabitant arrived, and it is Japanese
This issue asked whether the head-peel should be structured to accommodate a second inhabitant, calling Thai "the strongest evidence that a second inhabitant is coming." One has, from an unrelated direction.
Japanese maiden markers need the same surgery.
旧姓shipped in 2.1 as default vocabulary (#330), but it reaches only the spaced form. The form Japanese actually writes puts the marker inside brackets with a fullwidth colon, and that does not split:(旧姓:佐藤) -> '旧姓:佐藤' ONE tokenAdding
旧姓:toMAIDEN_MARKERSwas tried and does nothing — vocabulary matching is whole-token, and the token is the whole string. What it needs is exactly this issue's mechanism:นายสมชาย -> title นาย + given สมชาย (this issue) 旧姓:佐藤 -> marker 旧姓: + maiden 佐藤 (#329, moved here)Same operation, two vocabularies, two destination roles —
titlefor Thai, consumed-into-maidenfor Japanese.What that changes about the design
A third destination, and one of them is consumption rather than assignment. This issue framed the head-peel as routing to
title, needing "a separate vocabulary rather than a reuse ofhonorific_tails." Two consumers now, and the Japanese one does not route the peeled head anywhere — it drops it, the way_group's #274 rule drops a barenée. So the mechanism wants a peeled-head disposition (route to role R, or consume) rather than a fixed destination.The separator question is answered by making it part of the entry.
旧姓alone would leave:佐藤with the colon attached;旧姓:leaves佐藤clean. Worth knowing when sizing the vocabulary format — entries may carry trailing punctuation, which the tail-peel's entries never do.The safety argument carries over unchanged. #308's rule is that an entry peels only where it can never END a name. The head-peel's mirror is: only where it can never BEGIN one.
旧姓:cannot.นาย/นาง/นางสาวcannot, which this issue already argued. The Chinese familiar prefixes 老/小 still cannot clear it, for the reason stated here — 小 is a common given-name character and 小明 would be cut in half.Status
Not asking to widen this issue's scope — Thai remains the reason to build it. Recording the second inhabitant because it was predicted here, it arrived, and it constrains the design in one way the Thai case alone would not have surfaced: the peeled head is not always routed somewhere.
#329 keeps the language-agnostic half it can fix without this mechanism (a tagged marker inside extracted maiden content going unconsumed).
Placeholder — filed so it isn't lost, not worked up.
Thai honorifics are conventionally written glued to the given name, and they currently end up inside it:
What already works
More than I expected, which is what makes this narrow:
Thai is given-name-first, so the positional default is already right — no
script_ordersentry needed. Names are spaced between given and family, so no segmentation is needed either. Thai is not in_SCRIPT_RANGESat all, and it may not need to be: #308's peel is licensed by its vocabulary rather than by the script it is written in.So the whole gap is one mechanism: splitting a listed honorific off the front of a token.
Relationship to #308 and #312
#308 shipped the mirror of this — a listed honorific peeled off the end of a name token, routed to
suffix. Thai needs the same surgery from the other side, routed totitle(นาย is pre-nominal, not post-nominal), which means a separate vocabulary rather than a reuse ofhonorific_tails.#312 is deciding how the existing peel is structured. Its answer matters here: if the peel becomes a sibling function inside
script_segment, a head-peel joins as a third sibling; if it becomes its own stage, this belongs in that stage. Thai is the strongest evidence that a second inhabitant is coming, so it is worth weighing in #312 rather than after it.Why the vetting may be easier here than for the Han equivalents
A leading peel is normally harder to vet than a trailing one, because the leading position is where surnames live — the argument that clears 양 in #308 is precisely that a surname LEADS, so a trailing-only gate never sees it. The Chinese familiar prefixes 老/小 (老王, 小王) fail badly on this: 小 is a common given-name character, and peeling it would cut 小明 — this repo's own worked example — in half.
Thai looks different. นาย, นาง, นางสาว are not name components in any position; they are closed-class address terms. So the "can this ever begin a name?" test may come back clean where the Han equivalents do not. That is an impression, not a vetted claim — it needs the same per-entry argument #308's vocabulary carries, from someone who reads Thai.
Open questions
_SCRIPT_RANGESentry at all, given the peel is vocabulary-licensed. Probably only if some other script-conditional behavior is wanted later.titlein every case, or whether some entries are better modelled as existingtitlesvocabulary once the token is split.Raised while discussing #312; no one has asked for Thai support, so priority is unset.