Skip to content

de Mesnil Jean reads Jean as part of the surname: the leading-particle fold takes the rest of the name #471

Description

@derek73

Under the default (given-first) order, a name opening with a never-given particle has its ENTIRE remainder read as the surname:

parse("de Mesnil Jean")            family 'de Mesnil Jean'    →  should be family 'de Mesnil',  given 'Jean'
parse("de la Cruz Juan Carlos")    family (whole string)      →  should be family 'de la Cruz', given 'Juan', middle 'Carlos'
parse("ibn Awf abdul Rahman")      family (whole string)      →  should be family 'ibn Awf',    given 'abdul Rahman'

Jean is not part of that surname. The fold is reaching past the thing it is entitled to.

The criterion

The fold takes the particle's own unit and ONE name word. Whether the name is surname-only is then decided by what is left over:

  • nothing left → surname-only is the only available reading, and the fold is right (de la Vega, Mc Donald, dos Santos, De Groot, Ste Marie, de los Santos)
  • something left → that something IS a name component, and the fold has no business taking it

The mechanism already exists: the family-first branch of rules.md#P1 has narrowed exactly this way since #395. Only the default order kept the wider reach, and nothing recorded a reason.

Measured

10 parses, 5 names, default order only, over the 1094-name corpus × 3 name_order values × middle_as_family off and on. Every surname-only name is untouched; the conjunction join (de la Vega y Santos Juan) and the bound-given pair (ibn Awf abdul Rahman) stay whole.

This is a v1 parity break, and that is why it is its own issue

Unlike the family-first work in #467 — where v1 has no name_order and so no answer to compare against — this changes the order every existing caller is on, and 1.4.0 has a definite answer that we currently match:

1.4.0   HumanName("de Mesnil Juan").last  ==  'de Mesnil Jean'

Prototyped, the differential gate goes red at the 1.4.0 baseline with 5 unexplained diffs (the five names above), and tests/test_particles.py::LastNamePrefixSplitTestCase::test_leading_non_first_name_prefix_with_middle_name_as_last fails outright. Landing it needs ledger rules at all three baselines, the v1 test repointed, and a release note saying default-order behaviour changed — none of which should be bundled into a family-first fix.

Scope note

The leftover must be laid out as the positions the declared order gives it, with the family slot already filled. The current code uses _name_positions(order, len(rest) + 1)[1:], which is correct only where the family comes first; under the default order the family is last, so the slice takes the wrong roles.

Activity

  1. self-assigned this
    on Aug 30, 2026
  2. added this to the v2.3 milestone on Aug 30, 2026
  3. changed the title [-]`de Mesnil Juan` reads `Juan` as part of the surname: the leading-particle fold takes the rest of the name[/-] [+]`de Mesnil Jean` reads `Jean` as part of the surname: the leading-particle fold takes the rest of the name[/+] on Sep 1, 2026
  4. added a commit that references this issue on Sep 8, 2026
  5. derek73 commented on Sep 8, 2026

    @derek73
    OwnerAuthor

    Closing as by design (#517). The greedy reach under the default order stays, and the reason is now written down as a principle rather than a limit.

    A name_order is a property of the data source, not of a string. A caller sets it to match how their records are written, and a declared family-first order outranks anything the parser could infer from the shape of one name (rules.md's Name order Background; it yields only to a vocabulary claim, O4, or a name's own script, W4). Under the default given-first order a string that opens with a never-given particle is a surname whose given name is simply absent, and surnames may be several words — Spanish double surnames above all — so the fold takes the rest of the name. de la Torre Vega is a surname-only string read right: family de la Torre Vega.

    de la Cruz Juan Carlos is a family-first listing read under the wrong order, and the remedy is the declaration or the comma. Policy(name_order=FAMILY_FIRST) already gives family de la Cruz, given Juan, middle Carlos (#395), and the family comma fixes the surname whatever order is declared: de la Torre Vega, Juan reads family de la Torre Vega, given Juan under both orders. de la Family Family Given and de la Family Given Given are undecidable under either order, which is what makes the comma the answer rather than a better heuristic.

    This supersedes only the REASON the 2026-08-17 P1 entry gave for the same outcome ("the parser cannot see language; write the comma"). That was right and weaker: the order is not a language judgement at all, it is a declaration about the data. The outcome, and the rest of that entry, stand. The #364 sentence "nothing ever argued for 'takes everything'" is answered in place.

    rules.md#P1 carries the Accepted clause with both faces as executable examples; de la Torre Vega joined the contract corpus and, at the 1.4.0 baseline only, the already-recorded initials-view class, so that gate's count is 368 with the claim recorded. Nothing in the parser, the tests' expectations, or the ledgers moved.

    The one follow-up this leaves open is whether an undecidable particle-first shape should report an ambiguity; that is an AmbiguityKind question and belongs with #449/#491/#348 if anywhere. Rationale: decisions.md#P1, the 2026-09-07 entry.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions