Skip to content

Should John Smith XYZ read XYZ as a credential when no wordlist has those letters? #516

Description

@derek73

parse("John Smith XYZ") reads middle Smith, family XYZ today, and so does parse("John Smith X.Y.Z."). Every post-nominal this library recognizes, it recognizes from a wordlist, so an unlisted credential is a family name — which is fine until the wordlist is wrong in the other direction, and #342 has just made it wrong for two words on purpose. Removing rai and cha in 2.3 cost John Smith RAI its credential reading, the accepted price of Aishwarya Rai keeping her surname. This issue asks whether shape and position could pay it back.

Two shapes, both position-scoped. (1) An ALL-CAPS acronym standing in the suffix position of a MIXED-CASE name: John Smith XYZ, John Smith, XYZ. The case contrast is the signal, so an all-caps name gives none — JOHN SMITH XYZ must keep reading family XYZ, and that carve-out is half the design. (2) A DOTTED acronym in that position regardless of case: John Smith X.Y.Z. and john smith x.y.z. should both read suffix x.y.z., because the periods carry the signal. Measured 2026-09-07, none of the four reads as a suffix today.

The dotted half is already half-implemented by accident, which is the strongest evidence here. The interior-period split (rules.md#S3) reads a trailing dotted token chunk by chunk, and a last chunk that is a Roman numeral (i, v, x) is suffix vocabulary, so John Smith X.Y.I. and John Smith R.A.V. DO read suffix while John Smith R.A.X. does not. Same shape, same position, different answer, decided by whether the last letter happens to be I, V or X. Nothing is pinned on that and it is not a feature.

Relation to #490, which proposed recognizing a credential by shape where it collides with existing vocabulary. This is #490's idea WITHOUT the collision part: the collision cases are the ones a wordlist already handles, and the gap is the letters no wordlist has. Recorded as a parking-lot bullet in decisions.md#suffix-acronym-collisions.

Relation to #383, which asks whether Jose E Maria Santos should join like Jose e Maria Santos does. That is the same signal read at a different position: a single letter is a conjunction only when the input is mixed case and the letter's case says so. Whatever criterion this issue settles on for "the input is mixed case and the word in question is upper case" is probably the criterion #383 wants too, and the two should share it rather than each growing one.

What needs deciding.

  • Whether an unlisted all-caps trailing token is a credential often enough to be worth the surnames it would eat (Vietnamese and Korean data is frequently upper-cased in whole records, which the mixed-case requirement is meant to exclude — is that enough?).
  • Whether the dotted form should be decided independently of the all-caps one, since the periods argue much harder.
  • Whether the roman-numeral fork should be unified with whatever lands or left alone as generational.
  • Whether either shape reports a suffix-or-name ambiguity.
  • Whether there should be a way to turn the new behavior off, and/or to force it to apply to all-caps or all-lower-case input, the way capitalized(force=True) overrides the mixed-case default for capitalization — and whether Should Jose E Maria Santos join like Jose e Maria Santos does? (the single-letter connective's Latin-capital veto was never adjudicated) #383's case-sensitive behavior should ride the same switch or a similar one.

Activity

  1. added this to the 2.4 milestone on Sep 9, 2026
  2. derek73 commented on Sep 14, 2026

    @derek73
    OwnerAuthor

    PR #527 answers this issue's switch question for the single-letter shape: there is no switch. The Lexicon.conjunctions_ambiguous subset is the knob (remove e to restore joining, add y for a Dutch-style reading), and the "written wholly in one case" fact is computed once per parse at classify time over the name's own words (a maiden clause and delimited content excluded). The all-caps acronym half of this issue would read that same fact at the suffix slot; the dotted half needs no case fact at all.

  3. added 3 commits that reference this issue on Sep 18, 2026
  4. added a commit that references this issue on Sep 19, 2026
  5. derek73 commented on Sep 19, 2026

    @derek73
    OwnerAuthor

    Shipped in 2.4 (PR #530), as two Policy switches with different defaults, which is the answer to the question this issue asked.

    unlisted_dotted_suffixes, default True: an unlisted token of two or more alphabetic period-separated chunks joins the ambiguous credential class by shape, and the words-to-spare rule reads it. John Smith X.Y.Z. gives suffix X.Y.Z., Jack X.Y.Z. keeps its family, John Smith, A.B. gives suffix A.B. and Smith, A.B. keeps the given-name reading from #29. All four report suffix-or-name. Case is irrelevant here because the periods are the signal. Whole-token vocabulary still wins (M.A., Ph.D.), a lone trailing period is not the shape (John Smith Xyz. stays a family name), digits are not chunks (John Smith 1.4), and leading runs such as J.R.R. Tolkien are untouched.

    unlisted_caps_suffixes, default False: the same for an all-caps word found in no wordlist, in a mixed-case name. It is off by default because French and Korean records write surnames in capitals. With it on, Jean Pierre DUPONT gives family Pierre, suffix DUPONT, and a swallowed family name is the worse failure. Two-word Jean DUPONT and Minjun KIM keep their family either way. With it on, John Smith RAI and Ahmad Jayadi, CHA read as credentials again, which is the payback for #342 that this issue proposed.

    One accident retired whether or not the first switch is on: a dotted word whose only vocabulary match was a single ASCII character was reading as a generational suffix. Jack X.Y.I. gives family X.Y.I. again, which is 1.4.0's reading, while Msc.Ed., JD.CPA, Lt.Gov. and J.씨 are untouched.

    Neither switch reaches the v1 Constants API. HumanName tracks the parser's defaults, so the dotted reading reaches it and cannot be turned off from there. Recorded at decisions.md#S2.

  6. added 7 commits that reference this issue on Sep 22, 2026
  7. added 2 commits that reference this issue on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions