Skip to content

Smith, Jr., Dr. Bart reports its last part as "consumed as suffix" when 'Dr.' is read as a title #629

Description

@derek73

A part past the second comma that the parser doesn't recognise gets a comma-structure report saying it was "consumed as suffix best-effort". Since #603 (rules.md#C2), a title word in that part reads as a title, so the report describes a reading the parse didn't make:

Input Fields Report says
John Smith, Jr., Dr. Bart title 'Dr.', suffix 'Jr., Bart' segment 'Dr. Bart' … consumed as suffix best-effort
Smith, Jr., Attorney Berg title 'Attorney', family 'Smith', suffix 'Jr., Berg' segment 'Attorney Berg' … consumed as suffix best-effort
,, Attorney Berg title 'Attorney', suffix 'Berg' the same

At 2.3.0 the wording was true, because the whole part became suffix. It went wrong this cycle and has not shipped. The report is made in segment (_pipeline/_segment.py, the COMMA_STRUCTURE emitter), before anything reads the part. Its referents then land in two fields.

Found by: #626's review sweeping every report's field phrase against where its tokens land.

Fix direction: word it from the part's final roles at assemble, as #626 does for named fields (for example "consumed as title and suffix best-effort"). Or reword it to say what is certain, that the part is beyond the recognised comma structures, without naming a field. Either way, add the phrase to _FIELD_CLAIMS in tests/v2/test_cases.py so the sweep checks it.

Open question: does a part holding a title still need the report at all? segment already excuses a part made only of titles and suffixes (titles_and_suffixes), so 'Dr. Bart' reports only because of 'Bart'.

Activity

  1. added a commit that references this issue on Oct 9, 2026
  2. derek73 commented on Oct 9, 2026

    @derek73
    OwnerAuthor

    Notes from #626's fix (PR #628), for whoever picks this up:

    • field_tail is the wrong tool here. Word a report's field from the word's final role (#626) #628 lets a report leave the field it names for assemble to fill in from the word's final role (PendingAmbiguity.field_tail). That works because each of those reports is about one piece whose tokens share one role. A part past the second comma doesn't: its words land in two fields. In John Smith, Jr., Freiherr von Richthofen, 'Freiherr' is a title and 'von Richthofen' are suffixes. So naming "the first token's field" would print "a title" for a part that is mostly suffix. The fix needs either a wording built from the part's distinct final roles ("consumed as title and suffix"), or a wording that names no field.
    • Add the phrasing to the guard. test_a_report_names_the_field_its_word_lands_in (tests/v2/test_cases.py) only checks wordings on its _FIELD_CLAIMS list, and "consumed as …" isn't on it. That's why the sweep passes today on the existing row John Smith, Jr., Freiherr von Richthofen. Whatever wording the fix chooses, add it there, then re-record test_the_field_sweep_sees_the_claims_it_checks.
    • Where it's tracked: decisions.md#A1 now carries Open: #629, and mechanisms.md's AMBIGUITY-AT-THE-DECISION-SITE entry notes this report as the exception.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions