Repository navigation
Delimited maiden keeps its marker where the bare form drops it ((née Jones) vs née Jones) #329
Description
Activity
- changed the title
[-]`maiden` keeps the marker when delimited and drops it when not (`née Jones` vs `Jones`), and `旧姓` isn't recognized[/-][+]Delimited `maiden` keeps its marker where the bare form drops it (`née Jones` vs `Jones`)[/+]on Aug 3, 2026 Half of this shipped in 2.1 — the scope narrows
旧姓is now in the defaultmaiden_markers(#330, merged as3af1e37), so the vocabulary half of this issue is done. The title has been narrowed accordingly.What shipped:
parse("山田花子 旧姓 佐藤") # family 山田花子, maiden 佐藤 parse("山田 花子 旧姓 佐藤") # given 花子, family 山田, maiden 佐藤
What remains, and it is the language-agnostic half
Delimited maiden content keeps its marker where the bare form consumes it:
parse("Jane Smith née Jones") # maiden 'Jones' parse("Jane Smith (née Jones)", maiden_delims) # maiden 'née Jones' parse("山田 花子(佐藤)", maiden_delims) # maiden '佐藤' parse("山田(旧姓:佐藤)", maiden_delims) # maiden '旧姓:佐藤'
Why.
_classifytags a bare markervocab:maiden-marker(_classify.py:59) and_group.py:324consumes it, folding marker plus the following piece intomaiden(#274)._extractassigns bracketed contentRole.MAIDENwholesale at extract time, beforeclassifyruns — the marker inside is never tagged, so the consuming rule never fires.Fix shape. Consume a leading maiden marker inside extracted maiden content, so both paths agree. Care on the anti-#100 span invariant: token spans must keep indexing the original string, so this is a role/span decision, not a text rewrite.
What #330 changes about the payoff
Before it, fixing the extraction would have made
山田(旧姓:佐藤)yield旧姓:佐藤still — the marker would be consumable in principle but旧姓was not vocabulary. Now it is, so this fix alone takes that input to maiden佐藤. The two halves were filed together because neither was worth much apart; the first landing is what makes this one worth doing on its own.Still open, unchanged
- The separator.
旧姓:佐藤carries a fullwidth colon after the marker. Part of the marker, part of the vocabulary, or stripped structurally? Decide rather than inherit. - Chinese and Korean. Chinese has 原姓 / 本姓; Korean women traditionally do not change surname, so the concept may not map. Wants the same per-entry vetting Provide constants in non-Latin scripts (Cyrillic, Greek, Arabic, Hebrew) #269's entries got, from someone who reads the languages. Independent of the extraction fix.
Not a route to it
Patching
maiden_delimitersinlocales.JAwould not help and would cost:()is already a nickname delimiter and_extractresolves the collision by exclusion — "a pair listed in maiden_delimiters is dropped from the effective nickname set" — so it trades fullwidth-paren nicknames for maiden names wholesale. Measured:山田 太郎(マイケル・ジャクソン)moves from nickname to maiden. The collision is a real ambiguity in the writing system; the only thing separating the two conventions is the marker inside the brackets, which is exactly what extraction has swallowed.- The separator.
- changed the title
[-]Delimited `maiden` keeps its marker where the bare form drops it (`née Jones` vs `Jones`)[/-][+]Delimited `maiden` keeps its marker where the bare form drops it (`(née Jones)` vs `née Jones`)[/+]on Aug 3, 2026 Scope narrowed again — the Japanese half is #317's mechanism, not this one
Measuring the token structure corrected a claim in this issue and split it in two.
Correction. This issue said the marker inside brackets "is never tagged, so the consuming rule never fires." That is false for the space-separated forms —
classifytags it fine:(旧姓 佐藤) -> '旧姓' role=maiden tags=['vocab:maiden-marker'] + '佐藤' role=maiden (née Jones) -> 'née' role=maiden tags=['vocab:maiden-marker'] + 'Jones' role=maidenWhat does not happen is the consuming —
_group's #274 rule does not fold a tagged marker away for tokens extraction has already assignedRole.MAIDEN. That is what this issue is now scoped to, and it is tractable: no new mechanism, and the path is opt-in (Policy.maiden_delimitersdefaults tofrozenset(), four case rows, zero corpus entries).The Japanese form is a different problem. The fullwidth colon does not split the token:
(旧姓:佐藤) -> '旧姓:佐藤' ONE token, untaggedSo no amount of vocabulary reaches it. Adding
旧姓:toMAIDEN_MARKERSwas tried and does nothing — matching is whole-token, and the token is旧姓:佐藤.What it needs is a head-peel: split a listed vocabulary entry off the FRONT of a token. #308 shipped that for the tail (
田中さん-> 田中 + さん); nothing does the head. That is exactly #317's ask, and #317 anticipated a second inhabitant:นายสมชาย -> title นาย + given สมชาย (#317) 旧姓:佐藤 -> marker 旧姓: + maiden 佐藤 (this)Same mechanism, two vocabularies, two destination roles. Moved there rather than duplicated here — see the comment on #317.
This also answers the separator question that was open on this issue. With a head-peel,
旧姓alone would leave:佐藤with the colon attached, so旧姓:is the right vocabulary entry — the separator belongs to the marker. It is only unanswerable while there is no mechanism to consume it.What remains here
Consume a tagged maiden marker inside extracted maiden content, so the delimited path agrees with the bare one:
parse("Jane Smith née Jones") # maiden 'Jones' parse("Jane Smith (née Jones)", maiden_delims) # maiden 'née Jones' <- this parse("山田 花子(旧姓 佐藤)", maiden_delims) # maiden '旧姓 佐藤' <- and this
Care on the anti-#100 span invariant: token spans must keep indexing the original string, so this is a role/span decision rather than a text rewrite.
Still open, unchanged: Chinese and Korean equivalents. Chinese has 原姓 / 本姓; Korean women traditionally do not change surname, so the concept may not map. Wants the same per-entry vetting #269's entries got, from someone who reads the languages — and it is independent of everything above.
Moved to v2.1.
The case against was that this changes behavior released in 2.0. That holds much less weight than it looked: 2.0.0 shipped 2026-07-27, and
Policy.maiden_delimitersisfrozenset()by default — so the behavior being changed has been out for under a week, on a path a caller has to opt into. Almost nobody can be depending onnée Joneslanding in a field calledmaiden.Against that, this is the only open bug of the set that is Latin and reachable by any user of the feature. #322, #323 and #325 are narrower shapes.
Scope is unchanged from the comment above: consume a tagged maiden marker inside extracted maiden content so the delimited path agrees with the bare one. The Japanese bracketed form still needs #317's head-peel and is not in scope here.
- added a commit that references this issue
on Aug 3, 2026 - added a commit that references this issue
on Aug 16, 2026 - added 11 commits that reference this issue
on Aug 22, 2026 - added a commit that references this issue
on Aug 26, 2026 - added 4 commits that reference this issue
on Aug 29, 2026
Two halves of one gap. Neither is worth much without the other.
1. The same relationship gives two different
maidenvaluesA caller comparing
maidenacross a dataset gets a spurious mismatch between two spellings of one person's name.Why. Two paths, two treatments. Bare:
classifytags the markervocab:maiden-marker(_classify.py:59) and_group.py:324consumes it, folding marker plus following piece intomaiden(#274). Delimited:_extractassigns the bracketed contentRole.MAIDENwholesale at extract time, beforeclassifyruns — the marker inside is never tagged, so the consuming rule never fires.Fix shape. Consume a leading maiden marker inside extracted maiden content, so both paths agree. Care on the anti-#100 span invariant: token spans must keep indexing the original string, so this is a role/span decision, not a text rewrite.
2.
旧姓is not in the defaultmaiden_markersIt belongs there by the rule that admitted the Cyrillic entries — native-script entries cannot collide with Latin-script names. Verified: whole-token matching,
_normalize("旧姓")is旧姓, and neither character appears in any shipped surname, title, suffix, conjunction, particle or bound-given vocabulary. Measured with it added,山田花子 旧姓 佐藤→ family山田花子, maiden佐藤;山田 花子 旧姓 佐藤→ given花子, family山田, maiden佐藤.Not the JA pack: that exists for things needing the Japanese-data declaration (segmentation, where a pure-Han string cannot say which language it is). A Han-script marker needs no declaration — it can only match Han text, so it sits alongside
урожденнаяandgeboreninconfig/maiden_markers.py.Why together
The vocabulary alone reaches only the spaced form, which is not what Japanese usually writes. The extraction fix alone leaves
旧姓:佐藤unrecognized as a marker even once markers are consumed. Together,山田(旧姓:佐藤)withmaiden_delimitersgives佐藤.Open questions
旧姓:佐藤carries a fullwidth colon. Is it part of the marker, part of the vocabulary, or stripped structurally? Decide rather than inherit.Supersedes #309.