Repository navigation
Transcribed foreign names in Chinese (威廉·莎士比亚) parse reversed or not at all #298
Description
Activity
Design amendment recorded during implementation (PR #305): the fix is codepoint-scoped, which reframes one of this issue's two examples.
The issue called the U+30FB form (
威廉・莎士比亚→ family威廉) "confidently wrong." Implementation surfaced that the same shape carries the opposite convention in Japanese typography:高橋・一郎is roster formatting, 姓・名, and 2.1's family-first reading of it (deliberately pinned during #272) is correct. A blanket "dot = transcription" rule would fix the Chinese example by misreading standard Japanese formatting.Resolution: each dot is read by its own typography's standard — U+00B7, the canonical 间隔号 of the GB punctuation standard, is the transcription marker (divides between classified-script characters, keeps source order, suppresses segmentation); U+30FB/U+FF65 remain pure separators with the script license reading their pieces (kanji pair → family-first roster; katakana → positional). This stays inside the project's doctrine: the orthography itself settles the convention, no language guessing.
Consequence, pinned as a chosen limitation beside the spaced form: a Chinese transcription typed with the Japanese dot (
威廉・莎士比亚) reads by the convention of the codepoint it was typed with. Only the Chinese dot rescues a transcription.- added a commit that references this issue
on Jul 30, 2026 - added 2 commits that reference this issue
on Aug 8, 2026 - added a commit that references this issue
on Aug 16, 2026 - added 7 commits that reference this issue
on Aug 22, 2026 - added a commit that references this issue
on Sep 22, 2026 - added 4 commits that reference this issue
on Oct 8, 2026
Chinese has no dedicated script for foreign names. Where Japanese writes transcriptions in katakana, Chinese writes them in ordinary Han characters — 威廉·莎士比亚 is William Shakespeare — and marks the boundary between name parts with the 间隔号, an interpunct written as U+00B7 (·), U+30FB (・), or U+FF65 (・) depending on the input method. Transcribed names keep their source order: 威廉 (William) is the given name. The interpunct is the only orthographic signal that distinguishes a transcription from a native Han name.
The 2.1 CJK defaults get these names wrong in two different ways, depending on which codepoint the dot happens to be:
script_orderslicenses them family-first, and the parts come out reversed. The katakana case is protected because pure katakana is deliberately excluded from the order license; a Han transcription has no such exemption.givenunsplit.Two halves to a fix, and the second is the important one:
Out of scope: the space-separated form.
parse("威廉 莎士比亚")also reads family-first today, but nothing distinguishes it by script from a native name like 毛 泽东, and family-first is right for the overwhelmingly more common native case. That limitation should be documented, not fixed.