Skip to content

Delimited maiden keeps its marker where the bare form drops it ((née Jones) vs née Jones) #329

Description

@derek73

Two halves of one gap. Neither is worth much without the other.

1. The same relationship gives two different maiden values

parse("Jane Smith née Jones")                     # maiden 'Jones'
parse("Jane Smith (née Jones)", maiden_delims)    # maiden 'née Jones'
parse("山田 花子(佐藤)", maiden_delims)            # maiden '佐藤'
parse("山田(旧姓:佐藤)", maiden_delims)           # maiden '旧姓:佐藤'

A caller comparing maiden across a dataset gets a spurious mismatch between two spellings of one person's name.

Why. Two paths, two treatments. Bare: classify tags the marker vocab:maiden-marker (_classify.py:59) and _group.py:324 consumes it, folding marker plus following piece into maiden (#274). Delimited: _extract assigns the bracketed content Role.MAIDEN wholesale at extract time, before classify runs — the marker inside is never tagged, so the consuming rule never fires.

Fix shape. Consume a leading maiden marker inside extracted maiden content, so both paths agree. Care on the anti-#100 span invariant: token spans must keep indexing the original string, so this is a role/span decision, not a text rewrite.

2. 旧姓 is not in the default maiden_markers

It belongs there by the rule that admitted the Cyrillic entries — native-script entries cannot collide with Latin-script names. Verified: whole-token matching, _normalize("旧姓") is 旧姓, and neither character appears in any shipped surname, title, suffix, conjunction, particle or bound-given vocabulary. Measured with it added, 山田花子 旧姓 佐藤 → family 山田花子, maiden 佐藤; 山田 花子 旧姓 佐藤 → given 花子, family 山田, maiden 佐藤.

Not the JA pack: that exists for things needing the Japanese-data declaration (segmentation, where a pure-Han string cannot say which language it is). A Han-script marker needs no declaration — it can only match Han text, so it sits alongside урожденная and geboren in config/maiden_markers.py.

Why together

The vocabulary alone reaches only the spaced form, which is not what Japanese usually writes. The extraction fix alone leaves 旧姓:佐藤 unrecognized as a marker even once markers are consumed. Together, 山田(旧姓:佐藤) with maiden_delimiters gives 佐藤.

Open questions

  • The separator. 旧姓:佐藤 carries a fullwidth colon. Is it part of the marker, part of the vocabulary, or stripped structurally? Decide rather than inherit.
  • Chinese and Korean. Chinese has 原姓 / 本姓; Korean women traditionally do not change surname, so the concept may not map. Wants the same per-entry vetting Provide constants in non-Latin scripts (Cyrillic, Greek, Arabic, Hebrew) #269's entries got, from someone who reads the languages.

Supersedes #309.

Activity

  1. added this to the v2.2 milestone on Aug 3, 2026
  2. changed the title [-]`maiden` keeps the marker when delimited and drops it when not (`née Jones` vs `Jones`), and `旧姓` isn't recognized[/-] [+]Delimited `maiden` keeps its marker where the bare form drops it (`née Jones` vs `Jones`)[/+] on Aug 3, 2026
  3. derek73 commented on Aug 3, 2026

    @derek73
    OwnerAuthor

    Half of this shipped in 2.1 — the scope narrows

    旧姓 is now in the default maiden_markers (#330, merged as 3af1e37), so the vocabulary half of this issue is done. The title has been narrowed accordingly.

    What shipped:

    parse("山田花子 旧姓 佐藤")    # family 山田花子, maiden 佐藤
    parse("山田 花子 旧姓 佐藤")   # given 花子, family 山田, maiden 佐藤

    What remains, and it is the language-agnostic half

    Delimited maiden content keeps its marker where the bare form consumes it:

    parse("Jane Smith née Jones")                     # maiden 'Jones'
    parse("Jane Smith (née Jones)", maiden_delims)    # maiden 'née Jones'
    parse("山田 花子(佐藤)", maiden_delims)            # maiden '佐藤'
    parse("山田(旧姓:佐藤)", maiden_delims)           # maiden '旧姓:佐藤'

    Why. _classify tags a bare marker vocab:maiden-marker (_classify.py:59) and _group.py:324 consumes it, folding marker plus the following piece into maiden (#274). _extract assigns bracketed content Role.MAIDEN wholesale at extract time, before classify runs — the marker inside is never tagged, so the consuming rule never fires.

    Fix shape. Consume a leading maiden marker inside extracted maiden content, so both paths agree. Care on the anti-#100 span invariant: token spans must keep indexing the original string, so this is a role/span decision, not a text rewrite.

    What #330 changes about the payoff

    Before it, fixing the extraction would have made 山田(旧姓:佐藤) yield 旧姓:佐藤 still — the marker would be consumable in principle but 旧姓 was not vocabulary. Now it is, so this fix alone takes that input to maiden 佐藤. The two halves were filed together because neither was worth much apart; the first landing is what makes this one worth doing on its own.

    Still open, unchanged

    • The separator. 旧姓:佐藤 carries a fullwidth colon after the marker. Part of the marker, part of the vocabulary, or stripped structurally? Decide rather than inherit.
    • Chinese and Korean. Chinese has 原姓 / 本姓; Korean women traditionally do not change surname, so the concept may not map. Wants the same per-entry vetting Provide constants in non-Latin scripts (Cyrillic, Greek, Arabic, Hebrew) #269's entries got, from someone who reads the languages. Independent of the extraction fix.

    Not a route to it

    Patching maiden_delimiters in locales.JA would not help and would cost: () is already a nickname delimiter and _extract resolves the collision by exclusion — "a pair listed in maiden_delimiters is dropped from the effective nickname set" — so it trades fullwidth-paren nicknames for maiden names wholesale. Measured: 山田 太郎(マイケル・ジャクソン) moves from nickname to maiden. The collision is a real ambiguity in the writing system; the only thing separating the two conventions is the marker inside the brackets, which is exactly what extraction has swallowed.

  4. changed the title [-]Delimited `maiden` keeps its marker where the bare form drops it (`née Jones` vs `Jones`)[/-] [+]Delimited `maiden` keeps its marker where the bare form drops it (`(née Jones)` vs `née Jones`)[/+] on Aug 3, 2026
  5. derek73 commented on Aug 3, 2026

    @derek73
    OwnerAuthor

    Scope narrowed again — the Japanese half is #317's mechanism, not this one

    Measuring the token structure corrected a claim in this issue and split it in two.

    Correction. This issue said the marker inside brackets "is never tagged, so the consuming rule never fires." That is false for the space-separated forms — classify tags it fine:

    (旧姓 佐藤)   ->  '旧姓' role=maiden tags=['vocab:maiden-marker']  +  '佐藤' role=maiden
    (née Jones)  ->  'née'  role=maiden tags=['vocab:maiden-marker']  +  'Jones' role=maiden
    

    What does not happen is the consuming — _group's #274 rule does not fold a tagged marker away for tokens extraction has already assigned Role.MAIDEN. That is what this issue is now scoped to, and it is tractable: no new mechanism, and the path is opt-in (Policy.maiden_delimiters defaults to frozenset(), four case rows, zero corpus entries).

    The Japanese form is a different problem. The fullwidth colon does not split the token:

    (旧姓:佐藤)  ->  '旧姓:佐藤'   ONE token, untagged
    

    So no amount of vocabulary reaches it. Adding 旧姓: to MAIDEN_MARKERS was tried and does nothing — matching is whole-token, and the token is 旧姓:佐藤.

    What it needs is a head-peel: split a listed vocabulary entry off the FRONT of a token. #308 shipped that for the tail (田中さん -> 田中 + さん); nothing does the head. That is exactly #317's ask, and #317 anticipated a second inhabitant:

    นายสมชาย     ->  title นาย   + given สมชาย      (#317)
    旧姓:佐藤    ->  marker 旧姓: + maiden 佐藤      (this)
    

    Same mechanism, two vocabularies, two destination roles. Moved there rather than duplicated here — see the comment on #317.

    This also answers the separator question that was open on this issue. With a head-peel, 旧姓 alone would leave :佐藤 with the colon attached, so 旧姓: is the right vocabulary entry — the separator belongs to the marker. It is only unanswerable while there is no mechanism to consume it.

    What remains here

    Consume a tagged maiden marker inside extracted maiden content, so the delimited path agrees with the bare one:

    parse("Jane Smith née Jones")                     # maiden 'Jones'
    parse("Jane Smith (née Jones)", maiden_delims)    # maiden 'née Jones'   <- this
    parse("山田 花子(旧姓 佐藤)", maiden_delims)        # maiden '旧姓 佐藤'    <- and this

    Care on the anti-#100 span invariant: token spans must keep indexing the original string, so this is a role/span decision rather than a text rewrite.

    Still open, unchanged: Chinese and Korean equivalents. Chinese has 原姓 / 本姓; Korean women traditionally do not change surname, so the concept may not map. Wants the same per-entry vetting #269's entries got, from someone who reads the languages — and it is independent of everything above.

  6. modified the milestones: v2.2, v2.1 on Aug 3, 2026
  7. derek73 commented on Aug 3, 2026

    @derek73
    OwnerAuthor

    Moved to v2.1.

    The case against was that this changes behavior released in 2.0. That holds much less weight than it looked: 2.0.0 shipped 2026-07-27, and Policy.maiden_delimiters is frozenset() by default — so the behavior being changed has been out for under a week, on a path a caller has to opt into. Almost nobody can be depending on née Jones landing in a field called maiden.

    Against that, this is the only open bug of the set that is Latin and reachable by any user of the feature. #322, #323 and #325 are narrower shapes.

    Scope is unchanged from the comment above: consume a tagged maiden marker inside extracted maiden content so the delimited path agrees with the bare one. The Japanese bracketed form still needs #317's head-peel and is not in scope here.

  8. added 2 commits that reference this issue on Aug 4, 2026
  9. added a commit that references this issue on Aug 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions