A part past the second comma that the parser doesn't recognise gets a comma-structure report saying it was "consumed as suffix best-effort". Since #603 (rules.md#C2), a title word in that part reads as a title, so the report describes a reading the parse didn't make:
| Input |
Fields |
Report says |
John Smith, Jr., Dr. Bart |
title 'Dr.', suffix 'Jr., Bart' |
segment 'Dr. Bart' … consumed as suffix best-effort |
Smith, Jr., Attorney Berg |
title 'Attorney', family 'Smith', suffix 'Jr., Berg' |
segment 'Attorney Berg' … consumed as suffix best-effort |
,, Attorney Berg |
title 'Attorney', suffix 'Berg' |
the same |
At 2.3.0 the wording was true, because the whole part became suffix. It went wrong this cycle and has not shipped. The report is made in segment (_pipeline/_segment.py, the COMMA_STRUCTURE emitter), before anything reads the part. Its referents then land in two fields.
Found by: #626's review sweeping every report's field phrase against where its tokens land.
Fix direction: word it from the part's final roles at assemble, as #626 does for named fields (for example "consumed as title and suffix best-effort"). Or reword it to say what is certain, that the part is beyond the recognised comma structures, without naming a field. Either way, add the phrase to _FIELD_CLAIMS in tests/v2/test_cases.py so the sweep checks it.
Open question: does a part holding a title still need the report at all? segment already excuses a part made only of titles and suffixes (titles_and_suffixes), so 'Dr. Bart' reports only because of 'Bart'.
A part past the second comma that the parser doesn't recognise gets a
comma-structurereport saying it was "consumed as suffix best-effort". Since #603 (rules.md#C2), a title word in that part reads as a title, so the report describes a reading the parse didn't make:John Smith, Jr., Dr. BartSmith, Jr., Attorney Berg,, Attorney BergAt 2.3.0 the wording was true, because the whole part became suffix. It went wrong this cycle and has not shipped. The report is made in
segment(_pipeline/_segment.py, theCOMMA_STRUCTUREemitter), before anything reads the part. Its referents then land in two fields.Found by: #626's review sweeping every report's field phrase against where its tokens land.
Fix direction: word it from the part's final roles at assemble, as #626 does for named fields (for example "consumed as title and suffix best-effort"). Or reword it to say what is certain, that the part is beyond the recognised comma structures, without naming a field. Either way, add the phrase to
_FIELD_CLAIMSin tests/v2/test_cases.py so the sweep checks it.Open question: does a part holding a title still need the report at all?
segmentalready excuses a part made only of titles and suffixes (titles_and_suffixes), so 'Dr. Bart' reports only because of 'Bart'.