Skip to content

Mc Donald and Ste Marie read the particle as the given name #360

Description

@derek73

Two abbreviation particles are in the may-be-given half of the particle vocabulary, so a name that opens with one is read as having a given name:

Mc Donald    given='Mc'   family='Donald'   ambiguity: particle-or-given
Ste Marie    given='Ste'  family='Marie'    ambiguity: particle-or-given

Both should give the whole name as the surname (family='Mc Donald', family='Ste Marie'), which is what membership in NON_GIVEN_NAME_PARTICLES produces. Neither Mc nor Ste is a standalone given name in any culture — they are contractions of Mac and Sainte.

Note this is the default name order, so it is a live misparse today, independent of #359.

Three words, three different answers

Raised as "mc, st and ste have no vowels so cannot be given names". Measured, the vowel heuristic is a good instinct but not the actual criterion:

  • mc — move it. Live misparse above. It is also in SUFFIX_ACRONYMS (MC, Master of Ceremonies), but that is trailing position and unaffected: John Smith MC still gives suffix='MC'.
  • ste — move it. Live misparse above, and it has a vowel — so the criterion that actually works is "abbreviation of a word that is never itself a name", not the vowel shape.
  • st — moving it changes nothing. st is in TITLES, and title handling consumes a leading match before particle logic runs:
    St Clair    title='St'  given=''  family='Clair'
    
    Mid-name is already correct (Jean St Clair -> family='St Clair'). Whether title='St' is the right reading for a leading St is a separate question worth its own issue if anyone cares.
  • mac — leave it. Same shape as mc, but Mac is a real given name, so it belongs in the may-be-given half. This is the case that shows the vowel test is doing real work.

Joined spellings are single tokens and unaffected either way (McDonald, StClair).

The wider question behind it

Only 9 of the 39 may-be-given members were ever individually justified in the NON_GIVEN_NAME_PARTICLES docstring: al, van, von, della, di, del, da, vander, abu. The other 27 are there by the docstring's stated conservative default — "When unsure, leave a word out: a missing member just means that name is not auto-fixed, whereas a wrong member misparses a real person."

So the split was never exhaustively curated, and mc/ste are not a one-off oversight so much as the first two cases anyone has looked at closely.

Same-shape candidates worth examining: aan, aen, heer, freiherr, freiherrin, te, tho, thoe, vel, vande.

Words in that list that genuinely ARE given names and must stay put: bar (Bar Refaeli — the docstring's own Hebrew note cites exactly this), le, do, bin, mac.

The evidence standard is the docstring's: a wrong never-given member misparses a real person, a missing one merely leaves a name un-auto-fixed. So each word needs a reason to move, not an absence of a reason to stay.

Scope note

This changes default-order parsing, which is a wider blast radius than #359 (which only affects name_order=FAMILY_FIRST). Worth its own commit and its own tools/differential run so the numbers attribute to a cause rather than arriving as one lump.

Split out of #359 to keep that PR scoped.

Activity

  1. self-assigned this
    on Aug 9, 2026
  2. derek73 commented on Aug 10, 2026

    @derek73
    OwnerAuthor

    Dependency worth knowing before moving words out of the ambiguous set.

    _group's PARTICLE_OR_GIVEN emitter — the branch that reports a particle chained onto the following piece — needs a leading piece that is both a title and an ambiguous particle. Since #367 that is the only shape reaching it, because an ordinary title is now transparent to the leading-particle exception and takes _assign's branch instead.

    That intersection is exactly three words today:

    titles & particles_ambiguous  ->  ['do', 'freiherr', 'st']
    

    Two of them (freiherr, st) are on this issue's candidate list. st is listed here as moot for the leading position because TITLES claims it first — true for the parse, but moving it still removes it from particles_ambiguous and so from this intersection.

    PR #370 adds tests/v2/test_parser.py::test_the_chained_emitter_is_still_reachable, which fails in two distinguishable ways:

    • the word this issue moves is gone from the set — the message prints what is left, e.g. need a lead from ['do', 'st'], and the five emitter tests just need a different lead
    • the set is empty — the emitter is unreachable, every test of it is measuring nothing, and the right move is to delete the emitter rather than repoint its tests

    So moving one or two of the three is cheap and self-explaining. Moving all three is a decision about whether _group's emitter should still exist, and the test will say so rather than leaving a green suite over dead code.

  3. added this to the v2.2 milestone on Aug 14, 2026
  4. derek73 commented on Aug 17, 2026

    @derek73
    OwnerAuthor

    The split has a third category: words that made neither list

    This issue's "wider question" covers the 39 may-be-given members that
    were never individually justified. Found while attempting #390: there is
    a category above that one — articles absent from the vocabulary
    entirely.

    de     never-given        dos    never-given
    del    ambiguous          do     ambiguous
    la     ambiguous          da     ambiguous
    los    ABSENT             das    ABSENT
    las    ABSENT             el     ABSENT
    

    The coverage is inconsistent within a single language: Portuguese
    dos is never-given while its plural partner das is missing; Spanish
    del (the de+el contraction) is present while the bare articles
    el, los, las are not.

    Why nobody noticed, measured

    Adding los/las/das/el to the never-given set is a no-op
    today
    — seven probe shapes, zero change:

    de los Santos · de las Casas · dos das Neves · de el Greco
    Maria de los Santos · Juan de las Casas · de los Santos Garcia
                                                        all unchanged
    

    P1's whole-remainder sweep folds to the end of the name regardless, so
    it produces the right answer without knowing los exists — right for a
    reason unrelated to why the answer is right.

    Why it is a prerequisite for #390, not an improvement

    Narrowing the fold to "the particle and the one word it attaches to"
    removes the thing covering the gap. #390's first attempt regressed:

    de los Santos   2.1.0: last='de los Santos'  →  given='Santos'  family='de los'
    de las Casas    2.1.0: last='de las Casas'   →  given='Casas'   family='de las'
    

    A first/last swap on a very common Hispanic shape — the worst field
    pair to get wrong. This has to land before the fold narrowing, and
    its verification cannot be the differential: 0 diffs is
    indistinguishable from no effect. The proof is unit tests plus the #390
    retry.

    C-i does not say "add them all"

    Applying decisions.md#vocabulary-collisions C-i per word, the naive
    fix would repeat #342's Rai mistake:

    word C-i verdict evidence
    los, las never-given Spanish plural articles, borne as no one's name
    das ambiguous, not never-given Anjali Das → family='Das', a common Bengali surname
    lo ambiguous Wei Lo → family='Lo', a Chinese surname
    el ambiguous Ahmed El Sayed, Arabic transliteration
    des, les, il, gli, os, as needs research plausible articles, unexamined
    a, o, i, d do not add single letters; i is already in suffix_words as a Roman numeral, the rest collide with initials
    y already handled in conjunctions (rules.md#P3)

    So this is the same C-i exercise the issue already asks for, over a
    different population — and three of the four articles that first looked
    like obvious never-given members are borne surnames.

    Scope

    Fits here as its own commit: the no-op measurement means it adds
    nothing to this issue's differential run, so mc/ste's numbers still
    attribute to mc/ste. Worth splitting out if the never-given moves
    and the new members should have separate issues — they are
    distinguishable jobs with different blast radii, and only the moves are
    observable today.

  5. derek73 commented on Aug 17, 2026

    @derek73
    OwnerAuthor

    Correction to the table above: C-i was underspecified, and das was judged wrong

    The C-i verdicts I posted are wrong for at least das, and the reason
    is a flaw in the criterion rather than in the individual call. Correcting
    it here because that comment is public and someone could act on it.

    What was wrong. C-i as first recorded asked whether a word is borne
    as an ordinary name in some tradition
    . Applied to this set's most
    load-bearing member it self-destructs: "De" is a borne Bengali and Odia
    surname
    — parse("Bimal De").family == "De" — so the criterion said
    mark de AMBIGUOUS, which would break de la Vega and every
    leading-particle reading. Nothing had tested the criterion against de,
    and because the error pointed at the safe side it would never have
    surfaced as a misparse.

    The missing qualifier is POSITION, and the existing set was already
    obeying it. P1 acts on the leading position, or a lone particle in the
    given role — so what matters is whether the naming use occupies that
    position:

    van   Vietnamese Văn in given position    collides    ambiguous  ✓
    bar   Bar Refaeli, given position         collides    ambiguous  ✓
    do    Đỗ leads a surname                  collides    ambiguous  ✓
    de    "De" is a TRAILING surname          no clash    never-given ✓
    

    So the test is "is it borne as a name in the position the vocabulary
    claim acts on"
    , not "does a bearer exist anywhere".
    decisions.md#vocabulary-collisions is corrected with a dated entry.

    Corrected verdicts:

    word old verdict corrected why
    das ambiguous never-given Das is a trailing surname; the rule acts leading. Measured: Anjali Das and Bimal Das are unchanged, and Maria das Neves GAINS its particle — family='Neves' today, family='das Neves' with it. The old reading would have declined a fix.
    lo ambiguous still ambiguous, but for the right reason Lo genuinely leads in romanized Chinese (Lo Wei), so it collides in the acting position.
    el ambiguous needs re-judging Ahmed El Sayed has El mid-name, not leading — where a particle chains forward and would produce family='El Sayed', better than today's family='Sayed'.
    los, las never-given unchanged articles, no naming use in any position.

    Das Anjali does change under never-given das (family-first shape,
    given='Das' → family='Das Anjali') — genuinely ambiguous, and the
    kind of thing C-ii's frequency judgment covers rather than C-i.

    How this was caught, since it is the argument for doing the records
    at all: writing a per-word evidence comment for de is what falsified
    the criterion. A verdict alone hides the reasoning; a record has to
    state why, which forces the collision question. One word into the
    audit.

  6. added a commit that references this issue on Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions