Skip to content

Parenthesized (née Jones) parses as a nickname where the bare née Jones gives maiden #335

Description

@derek73
>>> parse("Jane Smith née Jones").maiden
'Jones'
>>> parse("Jane Smith (née Jones)").maiden
''
>>> parse("Jane Smith (née Jones)").nickname
'née Jones'

The marker is right there in the clause, so the parse has enough information to know Jones is a birth surname — but ( ) is a nickname_delimiters pair by default and maiden_delimiters is empty, so the clause is claimed as a nickname before anything looks inside it.

Setting Policy(maiden_delimiters=frozenset({("(", ")")})) fixes this one shape, but it is a whole-parser switch for what is really a per-clause fact, and it reroutes every parenthesis away from nicknames.

Also affected

>>> parse('Jane "Janey" Smith (née Jones)').nickname
'Janey née Jones'          # two clauses merged into one nickname
>>> parse("Anna Müller (geb. Schmidt)").maiden
''
>>> parse("山田 花子(旧姓 佐藤)").maiden
''                          # bare form gives '佐藤'

The mechanism is already in the tree

_pipeline/_extract.py already lets clause content overrule the delimiter's verdict. _suffix_shaped is the precedent: "Andrew Perkins (MBA)" is not a nickname, because the content is suffix-shaped, so extract masks only the two delimiter spans and lets the inner content join the main token stream.

The same treatment gives the requested behavior for free — strip the parens, and #274's existing bare-marker consuming rule produces exactly what née Jones produces, rather than a parallel implementation that could drift from it.

The predicate is small, and extract already imports Lexicon and _normalize and holds state.lexicon:

def _maiden_marked(content: str, lexicon: Lexicon) -> bool:
    head = content.split()[:1]
    return bool(head) and _normalize(head[0]) in lexicon.maiden_markers

Prototyped against the shapes that matter (2026-08-06): née Jones, nee Jones, geb. Schmidt (normalization strips the period) and 旧姓 佐藤 all return True; MBA, Janey, Jones née and 旧姓:佐藤 all return False.

Out of scope

山田(旧姓:佐藤) stays broken. The fullwidth colon is no separator tokenize knows, so marker and name arrive as one token and there is nothing to drop — that is the head-peel #317 tracks. Worth doing the two together, so "delimited maiden markers work" is one story rather than two partial ones.

Why this is not a 2.1 change

It is a default-on parse change, and the work is in the blast radius rather than the code:

  • 山田 花子(旧姓 佐藤) is in corpus_cjk.jsonl, so the differential gate sees it and expected_since_2.0.0.toml needs the row reclassified out of the CJK order-flip rule.
  • Three case rows need rewriting: maiden_marker_delimited, maiden_marker_delimited_unaccented (both currently require Policy(maiden_delimiters=...) and would work by default) and maiden_marker_kyusei_delimited.
  • tools/differential/README.md's "What this gate does not cover" section uses this exact example as the thing the gate cannot see. That section stops being true.
  • It needs a release-log entry.

The deciding factor is the release story rather than the effort. Jane Smith (née Jones) is Latin-only and is in no corpus, so the gate would go on reporting 0 of N changed names are Latin-only while a real Latin-affecting change shipped — measured-true and substantively misleading. Fixing that means adding Latin corpus rows, which makes the count non-zero, which is a change to 2.1's central claim that only CJK-bearing names moved.

In 2.2 it lands clean: the new ledger picks it up as a matter of course, and the Latin diff is expected rather than an exception to a headline.

Related: #329 (the marker inside a configured maiden clause), #274 (the bare-marker consuming rule this would reuse), #317 (the glued/colon form left out of scope).

Activity

  1. added this to the v2.2 milestone on Aug 6, 2026
  2. derek73 commented on Aug 26, 2026

    @derek73
    OwnerAuthor

    Two corrections to this issue's body, both measured while implementing it.

    The proposed mechanism is the one to avoid. Masking the delimiter spans and reusing #274's bare-marker rule agrees with region extraction on every trailing clause and diverges on an interior one:

    input extract-as-region mask-only
    Jane Smith (née Jones) family Smith, maiden Jones same
    Jane (née Jones) Smith family Smith, maiden Jones family '', maiden Jones Smith

    M2's take runs to the end of the name, so masking discards the boundary the closing delimiter states. Extraction is also the smaller change, since extract already emits Role.MAIDEN regions and group already drops the marker inside one (#329) — the fix is a role swap at the existing emit site rather than a new masking branch.

    "Jane Smith (née Jones) is in no corpus" is stale. It is in two — corpus_rules.jsonl and corpus_issues.jsonl — since #414 harvested the rules doc's examples. That sentence was the argument for deferring past 2.1, so it is worth correcting even though the conclusion (land it in 2.2) held.

    Two things found on the way, neither of them this issue's:

    Fixed in #446.

  3. added a commit that references this issue on Aug 26, 2026
  4. added 7 commits that reference this issue on Aug 29, 2026
    3c7a14f
    d1f79ac
    7a611a0
    7355de6
    eba0b46
    b1d46f5
    a473f90
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions