Skip to content

rules.md#R3's "a conjunction never initials" is unqualified: Juan de y has base y and initials J. #461

Description

@derek73

Reframed 2026-08-30. The original issue read this as initials() failing to honor rules.md#R3's "even then" clause. A fix on that reading was implemented and then backed out of #463 — it made initials() disagree with family_base about the same token, which is the defect the rest of that branch removes. The clause is the question, not the code.

R3 promises two things that do not agree

Initials take the first letter of each given, middle, and base family word; titles, suffixes, particles and nicknames contribute nothing — except the particles of a part whose every word is one, which are not acting as particles there (R2) and initial like any other name word. A CONJUNCTION never initials, so a base that is one contributes nothing even then.

The first sentence says initials follow the base. The last overrides it. Where they collide, the views split:

parse("Juan de y").family_base   # 'y'
parse("Juan de y").initials()    # 'J.'      -- the base word contributes nothing

The parse has already decided de is a working particle and y is the base — family_particles is 'de'. Only initials() disagrees.

The carve-out is right where the conjunction is doing its job

parse("Juan Velasquez y Garcia").family_base   # 'Velasquez y Garcia'
parse("Juan Velasquez y Garcia").initials()    # 'J. V. G.'   -- correct, y is joining

Here y links two surnames and must not initial. So the clause is not wrong in general; it is unqualified where it should be conditional.

The proposed rule

A conjunction with nothing to join is not acting as a conjunction, exactly as a particle with nothing to join is not acting as a particle. R2 already draws that line for particles and the machinery exists. Applied evenly:

today proposed
Juan Velasquez y Garcia J. V. G. J. V. G. unchanged — the y joins
Juan de y J. J. y. — the y joins nothing
Juan y J. J. y.

Every row then agrees with family_base, which is what R3's own first sentence promises.

The gap is wider than this, and the wider half is reachable by default

"A CONJUNCTION never initials" is unqualified, and a conjunction in the given group has always initialed. Measured over the 1094-name differential corpus, 25 names do it today:

'Duke of Edinburgh'     ->  'D. o. E.'
'Dean of Chemistry'     ->  'D. o. C.'
'John & Jane'           ->  'J. &. J.'

Unlike the family-side shape, this needs no custom vocabulary — so it can carry a rules.md example line and a real deviates: marker, which the family-side shape cannot.

Scope notes for whoever takes it

  • rules.md#R4 cross-references this clause. It grounds case repair's ungated conjunction conjunct in "the carve-out R3 states for initials". Settling R3 settles that sentence too, and R4 currently carries no trace of the question.
  • The all-particle shape cannot carry an example line. Every rules.md example parses with the default vocabulary, where particles ∩ conjunctions is empty — in the defaults and in every locale pack — so no input string reaches a conjunction inside an all-particle part. It is pinned in tests/v2/test_render.py instead, asserting today's output so that settling this fails the suite until the pin moves.
  • rules.md's preamble says grep deviates: is the deviation backlog, with three named exceptions. This gap fits none of them and can carry no marker on the family side, so R3 currently reads as fully implemented.
  • The readmission must not be widened by word count. Jon Dough and has base Dough and; Juan y Garcia has base y Garcia. A criterion phrased as "the base has a word the initials view skipped" condemns those and Velasquez y Garcia too. The mark is the criterion — it speaks about a whole part — not the word count.

Original report follows.


rules.md#R3 states a carve-out that initials() does not honor once a part is marked all-particle. Reproduction needs a Lexicon where a word is in both particles and conjunctions:

from nameparser import Lexicon, Parser
lex = Lexicon.default().add(particles={"y"})
p = Parser(lexicon=lex)

p.parse("Juan de y").initials()      # 'J. d. y.'
p.parse("Anh y Van").initials()      # 'A. y. V.'

Under the reframing above these outputs are correct — the part is all-particle, so its words are name words — and it is R3's clause that needs the condition.

Activity

  1. self-assigned this
    on Aug 29, 2026
  2. added this to the v2.2 milestone on Aug 30, 2026
  3. added a commit that references this issue on Aug 30, 2026
  4. changed the title [-]`initials()` initials a conjunction inside an all-particle part, where R3 says it never does[/-] [+]`rules.md#R3`'s "a conjunction never initials" is unqualified: `Juan de y` has base `y` and initials `J.`[/+] on Aug 30, 2026
  5. modified the milestones: v2.2, v2.3 on Aug 30, 2026
  6. modified the milestones: v2.3, 2.4 on Sep 9, 2026
  7. added 3 commits that reference this issue on Sep 20, 2026
  8. added a commit that references this issue on Sep 22, 2026
  9. derek73 commented on Sep 22, 2026

    @derek73
    OwnerAuthor

    Closed by #536.

    R3's clause is conditional now, and it is one rule for every group: a connective contributes nothing where it is joining, and a part holding nothing else is a part where it is joining nothing, so there it initials like any other name word. The whole part decides, never a word count.

    before now
    Juan Velasquez y Garcia J. V. G. unchanged, the y joins
    Jon Dough and J. D. unchanged, base Dough and
    Juan de y, Juan y J. J. y., agreeing with family_base
    Juan y Garcia J. G. J. y. G., the middle part holds nothing else
    John and Jane Smith J. a. J. S. J. J. S.
    Duke of Edinburgh, John & Jane D. o. E., J. &. J. D. E., J. J.

    The wider half this issue measured, a given-group connective initialing, was the code not following the rule rather than the rule being wrong, and it is the half that departs from 1.4.0 (J a J. S.). The lone-part half mostly restores 1.4.0: JUAN Y GARCIA and محمد و علي are back in parity and two ledger rules retired. 46 corpus names move, 18 gaining a letter and 28 losing one, identically on parse().initials() and HumanName.initials().

    The scope notes, in order. rules.md#R4 no longer borrows its grounding from this clause: a connective in a name part keeps its lowercase as R4's own rule, case repair did not move for any name whose roles did not move, and a connective the parse read as a generation repairs as the generation (John Smith i forced gives I). The all-particle custom-lexicon shape is unchanged and still pinned in tests/v2/test_render.py, its self-expiring pin rewritten to the settled rule. The readmission is keyed on a mark about the whole part, written by the same walk that writes R2's all-particle mark, so it cannot be widened by word count; revise recomputes both.

    The August bullet under R2 that said Juan y Garcia and Juan de y "may not move" is amended by a dated bullet rather than edited. Reasoning under R3 and R4 in docs/design/decisions.md.

  10. added 4 commits that reference this issue on Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions