Skip to content

Should Japanese 旧姓 (former surname) populate the maiden field? (山田(旧姓:佐藤)) #309

Description

@derek73

Japanese documents record a pre-marriage surname as 旧姓, conventionally parenthesized with a colon: 山田花子(旧姓:佐藤). Today the fullwidth parens are nickname delimiters (#273), so 旧姓:佐藤 presumably lands in nickname as literal text. The maiden-marker mechanism (#274: née, geb.) routes a following name to maiden, but it operates on tokens in the main stream — this convention puts the marker inside a delimited region, plus a colon.

Settling this needs: whether marker recognition should reach into extracted regions (a mechanism change), whether 旧姓 alone (spaced, unparenthesized: 佐藤 旧姓 山田?) occurs enough to matter, and what the colon does to tokenization. Filed as a question rather than a plan — the answer may be "document as unsupported."

Activity

  1. added this to the v2.1 milestone on Aug 1, 2026
  2. derek73 commented on Aug 3, 2026

    @derek73
    OwnerAuthor

    Punting in favour of #329, which covers this and generalizes it.

    Two findings changed the shape of the ask.

    The Japanese-specific part belongs in the default vocabulary, not the pack. 旧姓 satisfies the rule that already admitted the Cyrillic markers — "native-script entries cannot collide with Latin-script names, which is what makes them safe to enable by default." Matching is whole-token, _normalize("旧姓") is 旧姓, and neither character appears in any shipped surname, title, suffix, conjunction, particle or bound-given vocabulary. It sits alongside урожденная and geboren, not behind a Japanese-data declaration — locales.JA exists for things that need to know the data is Japanese, and a Han-script marker can only ever match Han text.

    Measured with it added: 山田花子 旧姓 佐藤 → family 山田花子, maiden 佐藤; 山田 花子 旧姓 佐藤 → given 花子, family 山田, maiden 佐藤.

    But the form this issue's title uses still does not work, and that half is language-agnostic:

    Parser(policy=Policy(maiden_delimiters={('(',')')}))
    山田(旧姓:佐藤)  ->  maiden '旧姓:佐藤'   # marker and colon ride along

    extract_delimited assigns bracketed content Role.MAIDEN wholesale before classify runs, so the marker inside is never tagged and _group's consuming rule never fires. née has the identical asymmetry — Jane Smith née Jones gives maiden Jones, Jane Smith (née Jones) gives maiden née Jones.

    So the vocabulary and the extraction fix are worth little apart, and are filed together as #329. Reopen if that lands and Japanese still needs something specific.

    Worth recording separately: patching maiden_delimiters in locales.JA is not an option either way. () is already a nickname delimiter, and _extract resolves the collision by exclusion — "a pair listed in maiden_delimiters is dropped from the effective nickname set" — so it would trade fullwidth-paren nicknames for maiden names wholesale. Measured: 山田 太郎(マイケル・ジャクソン) goes from nickname to maiden. That collision is a real ambiguity in the writing system; the only thing separating the two conventions is the marker inside the brackets.

  3. added a commit that references this issue on Aug 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions