Skip to content

Apply suffix_delimiter only at suffix-consumption sites - #206

Merged
derek73 merged 2 commits into
masterfrom
fix/suffix-delimiter-consumption-site
Jul 4, 2026
Merged

derek73 merged 2 commits into
masterfrom
fix/suffix-delimiter-consumption-site

Conversation

@derek73

@derek73 derek73 commented Jul 4, 2026

Copy link
Copy Markdown
Owner

Summary

  • suffix_delimiter previously split every post-comma segment unconditionally, before the parser decided which segment was actually a suffix group. This let it leak into first/middle-name segments in inverted format (e.g. "Doe, Mary - Kate, RN"), a case that was documented in Constants.suffix_delimiter's docstring as a known limitation.
  • Moved the delimiter split to the three places that actually consume a segment as suffixes: the suffix-comma vs. lastname-comma format detection (are_suffixes_after_comma), and the two suffix_list += parts[...] consumption points (suffix-comma branch's parts[1:], lastname-comma branch's parts[2:]).
  • The detection check now expands parts[1] on the delimiter to correctly recognize delimiter-joined suffixes (e.g. "RN - CRNA") while never expanding a segment that turns out to be a name.
  • Updated the docstring and replaced test_suffix_delimiter_inverted_format_known_limitation with test_suffix_delimiter_inverted_format_not_misparsed, asserting the corrected parse instead of just documenting the bug.

Test plan

  • pytest tests/test_suffixes.py -k suffix_delimiter -v — all 20 pass
  • Full suite: pytest -q — 1242 passed, 4 skipped, 22 xfailed, no regressions

derek73 added 2 commits July 4, 2026 02:36
Previously the delimiter split ran unconditionally on every post-comma
segment before the parser had decided which segment is actually a
suffix group. That let it leak into first/middle-name segments in
inverted format (e.g. "Doe, Mary - Kate, RN"), which was documented as
a known limitation.

Move the split to the three places that actually treat a segment as
suffixes: the suffix-comma vs. lastname-comma format detection, and
the two suffix_list += parts[...] consumption points. The detection
check now expands parts[1] to correctly recognize delimiter-joined
suffixes (e.g. "RN - CRNA") without ever expanding a segment that
turns out to be a name.
…ents

- Lock in hn.middle for the "Doe, Mary - Kate, RN" case so the stray
  hyphen artifact (a pre-existing tokenization quirk, reproducible
  without suffix_delimiter set) can't silently drift.
- Add coverage for the parts[1:] loop expanding more than one comma
  segment, for a multi-word token on one side of the delimiter in the
  detection check, and for delimiter no-op parity with the no-delimiter
  baseline when the format isn't detected as suffix-comma.
- Note in expand_suffix_delimiter's docstring that it's a no-op without
  a configured delimiter, and comment why detection needs the
  delimiter-expanded view of parts[1].
@derek73 derek73 self-assigned this Jul 4, 2026
@derek73 derek73 added this to the v1.3.0 milestone Jul 4, 2026
@derek73
derek73 merged commit b64a44f into master Jul 4, 2026
8 checks passed
@derek73
derek73 deleted the fix/suffix-delimiter-consumption-site branch July 4, 2026 09:46
derek73 added a commit that referenced this pull request Aug 24, 2026
The review round found the first draft shipped the inverse of the bug it
fixed. It let ANY piece open an entry, as the tail block always had --
safe there, because assign routes every tail piece to SUFFIX, which is
what `tail` means, and wrong off it, where a title piece routes to TITLE.

Two failures, one cause, neither visible to the gates that passed. The
`joined` tag is role-BLIND and the facade heals it for every role:

    "Smith, Rev. Dr."     title_list  ['Rev.','Dr.'] -> ['Rev. Dr.']
    "Smith Jr., Mr. Jr."  suffix      'Jr., Jr.'     -> 'Jr. Jr.'

The second glues a suffix backward across a comma the writer typed --
exactly what #429 exists to stop. The differential compares strings and
cannot see the first; the case table asserts the title STRING, which is
space-joined either way, and could not see it either.

Two joins that had been one, separated: WITHIN a piece the tag renders a
merged piece as one unit whatever role it holds; BETWEEN pieces it
continues an entry, and only a piece rendering into the same run may do
that. Sticky across a piece that is not in the entry, so an interleaved
title does not split its run ("Smith, MD Dr. PhD" -> 'MD PhD'); a
delimiter core still closes it.

Eight case rows and a facade test for the list views, which is the only
surface that shows the title collapse. Both regression guards verified
against a mutation copy -- they fail with the old condition restored.

Prose corrections, all measured by the reviewers:

- The round-trip claim was false AND backwards: str() of a fixed parse
  is a no-comma string, which re-parses with the comma back. master was
  the str-stable one. Struck from the release log and the case note.
- "one-word family comma" is not the condition -- there is no word-count
  gate, so "John Smith, Jr. III" moves too (1.4.0's reading), as does a
  title-led "Smith, Dr. MD PhD". Scope restated as it reads.
- The delimiter parity is #206 (021823e, "Apply suffix_delimiter only at
  suffix-consumption sites"), NOT #191, the German/Dutch vocabulary PR.
  Three code comments carried the error; corrected with it.
- The dormant-rule tell is #373's, and #426 the precedent for dropping a
  shadowed rule -- neither #424 entry mentions it.
- "boundary example" in the entry and both ledgers: the example FIRES,
  which is why the annotation came off.
- "filed rather than folded in" claimed an issue that does not exist.

C1 gains `_group.py` in `implemented:`, with the verbatim citation the
equality guard requires -- the whole-run half of the rule renders here.

Co-Authored-By: Claude Opus 5 <[email protected]>
derek73 added a commit that referenced this pull request Sep 6, 2026
rules.md#R1 has said it since 2026-08-23 -- a run of post-nominals
written with spaces renders with spaces, one written with commas
keeps them -- and every example under it was a comma path. The
NO-COMMA path violated it at 1.4.0, 2.0.0, 2.1.0 and 2.2.0 alike:
HumanName("John Smith MD PhD").suffix was 'MD, PhD', a comma nobody
typed. #429 fixed the family-comma path and recorded this one as
left alone.

The entry boundary is now read off the text instead of off segment
shape. A post-nominal continues the one before it iff they sit in
the same comma bucket -- comma_bucket, the function segment builds
segments with -- and nothing between them parts the run. group loses
the between-piece marking whole -- `tail` by index, #429's
segment_suffix_reading verdict, the per-piece gate and the
stickiness across an interleaved title -- and keeps core dropping
and the within-piece Ph. D. heal. The stickiness survives by
construction: 'Smith, MD Dr. PhD' has no comma between MD and PhD.

What parts a run is a name word -- GIVEN, MIDDLE or FAMILY -- or a
dropped delimiter core (#206), which is what 'Smith, MD - PhD -
FACS' puts between PhD and FACS under every policy, the half a
dropped-core test cannot reach. What does not part it is a token
whose role renders into a field other than the name and the suffix:
_RENDERS_ELSEWHERE beside the pass names all three of them, TITLE,
NICKNAME and MAIDEN. Nor does a dropped token that is NOT a core,
which is the maiden marker of a bare-marker clause -- dropped with
no role at all, so the role test cannot see it -- and that is why
the dropped arm reads the core set rather than treating every drop
as a boundary. The set is _vocab.delimiter_cores, group's own
derivation off Policy.extra_suffix_delimiters, imported rather than
repeated so the drop site and this pass cannot disagree. Measured
2026-09-06 against f5deae0: 'Smith, MD "Doc" PhD',
'Smith, MD (nee Jones) PhD' and 'Smith, MD nee Jones PhD' each
render 'MD PhD' at the parent and at HEAD, where a title-only
transparency would have regressed all three to 'MD, PhD'; and
'John Smith MD "Doc" PhD' rendered 'MD, PhD' at the parent -- the
bug this commit exists to fix, left unfixed by that draft.

Thirteen pre-existing corpus names move, all from a comma-joined
render to a space-joined one, none the other way, none with a
comma in the original between the two words. Nine cases.py rows
are re-pinned; 'MD PhD Jr., John' keeps its name and its point,
segment 0 being the family segment, and only its separator moves.
Eight assertions outside that table are re-pinned too: four in the
v1-parity suite (tests/test_suffixes.py x3 and
tests/test_capitalization.py) and four in the v2 API suite
(tests/v2/test_parser.py). v1 inserted a comma into a run the
writer had spaced, and that is a deliberate deviation recorded in
the release log by this bundle's records commit, not a parity
break. Four rules.md example lines outside R1 (P2's 'John Smith Mc
V', P5's 'abdul Smith Jr Ma', C1's 'Smith, John PhD I.', W3's
'田中さん 様.') carried the same comma and are corrected; all four are
corpus names, so the differential observed those edits
independently -- '田中さん 様.' sits in corpus_cjk_tolerated.jsonl and
carries its own ledger rule. C1's closing qualifier and W3's
tolerated: line were read and left verbatim: both describe a
reading, not a separator. Every *_list view other than suffix_list
is byte-identical over the corpus names, measured before and
after.

Round-tripping is stable on the shapes #429 fixed: str() of a fixed
parse is a no-comma string that used to re-parse with the comma
back.

Gate, with the new R1 example in the rules corpus: 360 / 254 / 166 /
28 intentional at 1.4.0 / 2.0.0 / 2.1.0 / 2.2.0, 0 unexplained, 0
radar unclassified. One name changes rule -- 'abdul Smith Jr V' at
2.0.0 and 2.1.0, whose diff grows past fix(#401)'s fields and lands
on a compound rule naming both mechanisms.

Refs #436, refs #437

Co-Authored-By: Claude Fable 5.1 <[email protected]>
derek73 added a commit that referenced this pull request Sep 27, 2026
…ew found stale

No parse changes. The full suite passes, all five differential gates
are unchanged, and frames per parse are identical to 4525770 on
CPython 3.11.16 for all ten reference names.

Reader dispatch:
- `_release_reads_off` takes
  `reader: Literal[TailReader.TRAILING, TailReader.GIVEN_SLOT]` and
  dispatches `if GIVEN_SLOT / elif TRAILING / else assert_never`.
- `_maiden_take` has three reader branches: the walk's first reading,
  the numeral view check, and the link release check. Each names
  TRAILING and GIVEN_SLOT explicitly and ends in `assert_never`.
- The title stop is gated `reader is not NONE`. `chained` is empty for
  NONE anyway, and the gate narrows the type.
- TailReader's promise, "a fourth member is a type error at every
  reader", now holds at every reader. mypy is clean, and
  test_an_unmapped_reader_is_a_loud_failure_rather_than_a_default
  passes.

`stop < trailing` is restored in the acronym fork. The floor is an
index, and the title chain's splice can leave the numeral stop in
front of it: in 'Jane Doe nee V Prof.' the fork sees stop 4 against
trailing 3 (instrumented in a copy of the tree). The paragraph that
said the condition guarded nothing is rewritten.

`tail_follows` is keyword-only and required on `_maiden_take`, and the
duplicated `reader is GIVEN_SLOT` test beside it is gone: group()
computes the flag from that reader. `_group_segment` keeps its
default, because its test callers do not pass it.

`_join_takes_the_member`:
- The docstring now names all three joins, each with its condition:
  - P2's chain, including a particle with a title behind it.
  - P6's attachment, for a released particle title after a family
    comma.
  - P5's bound-given join.
- The duplicated P5 rationale is gone.
- The `rules.md#P5:` citation now sits on one line, and
  test_doc_citations collects it.
- A note about a never-shipped intermediate reading is dropped.

Other code comments and docstrings:
- `_maiden_take`: all three stops, and a link the walk stops at, ask
  the release check; only the credential and title stops spare the
  first word.
- `_maiden_take`: the numeral view check reads only `Peel.numeral` for
  the NONE reader.
- `_maiden_take`: `peel_start` is the title-aware start for the other
  readers, and a `peel_start` past `trailing` is harmless because the
  walk never reaches `trailing`.
- `_maiden_take`: a frame comparison against a baseline that was never
  committed is dropped.
- The lone-core "keep in step" comments now say the #206 drop's copy
  adds `len(pieces) > 1`.
- `_run_neighbours` and the core-invariant test: "the #206 drop takes
  a LONE core out".
- `GIVEN_SLOT`'s comment names the H5 chain.
- `_pieces.py`: `trailing_start` names who stops at it. `trailing_titles`
  documents `floor`. `credential_at_the_given_slot` and `tail_reading`
  name their callers instead of counting them.

Tests and cases:
- Case a_title_behind_a_tail_bound_numeral_keeps_both: the title stop
  IS asked, and declines because 'V' stays a name word and the given
  part's chain stops at it.
- The clamp-era names are renamed and rewritten:
  - The case is now the_first_word_floor_holds_for_titles_too.
  - The unit test is now
    test_the_first_word_floor_holds_a_title_out_of_the_chain.
- Case notes name "the #538 commit d9d8049". A note describing a
  never-shipped state is dropped. test_group says "at 2.2.0 and 2.3.0".
- test_a_title_the_clause_gives_up_lands_in_title: the reader list is
  corrected. 'Jane Doe, PhD' is a tail segment and reaches NONE, and
  no head puts the clause before a family comma. The property holds at
  e0f1a2f over all 240 texts (checked).
- test_a_title_first_word_counts_as_a_word also asserts that both sides
  take a clause whose first word is the head. Its recorded control was
  re-measured 2026-09-26 with the floor removed entirely (`tail_reading`
  handed 1): all 30 pairs fail, on the head check alone, and 0 disagree
  on the suffix. The earlier "14 of 30" for a clamp variant is replaced.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
ethan42 pushed a commit to SavantEnvs/python-nameparser that referenced this pull request Oct 8, 2026
The link exception inside a maiden clause asks for a name word on
each side of the link. A delimiter core past the clause's first word
was an ordinary index to the neighbour walk and passed for that word,
so 'Smith, John, PhD née Puig Mr. - i Soler' under a ' - ' delimiter
kept 'i Soler' in the maiden name. The walk now steps over a core as
it steps over a connective, so the clause reads as the same text
written without the core. rules.md#M2's deviation becomes a statement
and two examples; decisions.md records the boundary reading declined.

The same stepping also applies to the `frozen` loop's own
`_run_neighbours` call, which reaches every credential tail and not
just a maiden clause: a connective of generational vocabulary beside
a declared delimiter core looks past it, and where a credential or
nothing stands beyond, it stays a lone suffix word and the core -- a
lone piece with nothing joined to it -- is dropped as derek73#206 drops any
lone core, so 'Smith, John, PhD - i Soler' reads suffix
'PhD - i Soler' -> 'PhD, i Soler'. Where a name word stands beyond
the core instead, the link joins across it and the core survives in
the suffix text ('Smith, John, Puig - i Soler' stays
'Puig - i Soler'); an ordinary connective's join still takes the
core as its neighbour regardless ('Smith, John, PhD - and MD' keeps
'PhD - and MD'). rules.md#P3 states the general separator-stepping
sentence this rests on; rules.md#R1 gains an example; both
cross-reference each other and M2. decisions.md's derek73#538 entry records
the measured population and the unrepaired cases.

Two new corpus names land in the ledgers. 'Smith, John, PhD née
Puig - i Soler' landed in _CORPUS_CLAIMS, in _NOT_A_VOCABULARY_COPY
in tests/v2/test_ledger_guards.py, and in the derek73#397 rule's alternation
in all five tools/differential/expected_since_*.toml ledgers, since
it reaches the DEFAULT facade there with no delimiter configured.
'Smith, John, PhD - i Soler' diffs at no baseline (the tree's own
default reading is unchanged from every baseline's), so it needed no
ledger alternation, only the corpus-wide growth in two 1.4.0
_CORPUS_CLAIMS comma rules (372 -> 373).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
ethan42 pushed a commit to SavantEnvs/python-nameparser that referenced this pull request Oct 8, 2026
…erek73#549)

A delimiter declared through extra_suffix_delimiters / suffix_delimiter
now separates a tail segment exactly as a comma typed in its place:
group cuts the segment at its cores before grouping, groups each part
as the comma twin's segment would be grouped, and drops the cores.
No join, maiden walk or link search can reach across a core, so the
derek73#538 stepping machinery goes: the cores parameter on four functions,
_maiden_take's seen index list, the skip arguments of peel_walk and
trailing_start, and the post-join derek73#206 drop block.

'Smith, John, PhD - and MD' reads suffix 'PhD, and MD' (1.4.0's
answer; 2.0-2.3 gave 'PhD - and MD'), and 'Smith, John, PhD née
Puig - i Soler' reads maiden 'Puig', suffix 'PhD, i Soler' again, as
2.3.0 did -- the boundary reading derek73#538 declined, taken now.

The default policy declares no delimiter; the differential gate exits
0 at all five baselines. One call per parse fewer.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant