Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
a965272
feat(differential): split the corpora into contract and radar tiers
derek73 Sep 1, 2026
10e48bc
feat(differential): corpus.jsonl carries the v1 test names it was scr…
derek73 Sep 1, 2026
c2dae22
feat(differential): shape inventory and an order-aware worker protocol
derek73 Sep 1, 2026
ee2ca00
feat(tests): case rows can declare the input shape they instantiate
derek73 Sep 1, 2026
a587393
style(tests): split one-line imports and annotate the fake workers
derek73 Sep 1, 2026
6edfd75
feat(differential): the shape-matrix contract corpus, generated from …
derek73 Sep 1, 2026
4d3cb03
test: the family-first/comma correspondence as a generated invariant
derek73 Sep 1, 2026
ae381b8
docs: document the two family-first input shapes and the comma corres…
derek73 Sep 1, 2026
6bfd87e
docs(differential): the corpora count claims catch up with the fifth …
derek73 Sep 1, 2026
30a5e83
docs(design): record the corpus-tier arc's decisions
derek73 Sep 1, 2026
d2aaf46
fix(differential): corpus object lines reject unknown and reserved keys
derek73 Sep 1, 2026
c897819
feat(differential): rules can scope to the default order, and order-b…
derek73 Sep 1, 2026
8fc8e0d
test(differential): the claims roster records a rule's order scope
derek73 Sep 1, 2026
e480889
test(differential): pin the dedup key, the scope sweep, the worker's …
derek73 Sep 1, 2026
0d893f6
docs: review-round corrections — recipes, the trio's causality, AGENT…
derek73 Sep 1, 2026
92b3b96
docs: the de Mesnil example's given name is Jean in the values too
derek73 Sep 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ Three committed contributor docs carry the parser's normative rules and their re

**Counting claims.** A bare count in prose is either an assertion or a liability, keyed by who observes its staleness: asserted counts (a test holds the number) fail CI at change time — the useful kind; dated snapshots ("51 sites at spec time") cannot go stale; standing present-tense prose counts are the forbidden class — promote to an assertion, add a date, or state the invariant and let a test count. After changing how many times something runs, sweep for counts, not for the thing's name.

**Release-log claims.** Quantified or universal behavior claims in release bullets must come from the differential gate's classified summary, be verified against rules.md examples, or -- for a view the gate cannot see -- carry a recompute recipe stored with the design entry the bullet cites; never write one from memory. The gate compares the seven role fields and `_ambiguities`, so `capitalized()`, `initials()` and any future render view are invisible to it (decisions.md#R4, #R3) and the first two sources cannot reach a claim about one: a gate run is byte-identical across the change, and an example line witnesses an output without counting anything. A recipe names the corpus files, the policy sweep, and -- the part that is easy to omit and fatal -- THE COMPARATOR, which must be something the shipped tree is not: #408's first recipe said to compare `initials()` against a folded-first partition, which is what `initials()` now IS, so it reproduced 0 where the bullet claimed 660 and was the only stated provenance for the number. Run the recipe as written before shipping the bullet. Cross-version numbers (a released wheel, the pre-change tree) are dated snapshots under Counting claims, since nothing in the repository re-runs them. Per-rule ledger toml comments asserting PARSER behavior cite rule IDs under the excerpt discipline; free prose is for ledger mechanics only (owned by tools/differential/README.md).
**Release-log claims.** Quantified or universal behavior claims in release bullets must come from the differential gate's classified summary, be verified against rules.md examples, or -- for a view the gate cannot see -- carry a recompute recipe stored with the design entry the bullet cites; never write one from memory. The classified summary covers the CONTRACT tier plus whatever radar diffs a rule classifies; a radar corpus's unmatched diffs are listed under UNCLASSIFIED (radar) and are not in it, so a claim quantified from the summary alone is silent about them. The gate compares the seven role fields and `_ambiguities`, so `capitalized()`, `initials()` and any future render view are invisible to it (decisions.md#R4, #R3) and the first two sources cannot reach a claim about one: a gate run is byte-identical across the change, and an example line witnesses an output without counting anything. A recipe names the corpus files, the policy sweep, and -- the part that is easy to omit and fatal -- THE COMPARATOR, which must be something the shipped tree is not: #408's first recipe said to compare `initials()` against a folded-first partition, which is what `initials()` now IS, so it reproduced 0 where the bullet claimed 660 and was the only stated provenance for the number. Run the recipe as written before shipping the bullet. Cross-version numbers (a released wheel, the pre-change tree) are dated snapshots under Counting claims, since nothing in the repository re-runs them. Per-rule ledger toml comments asserting PARSER behavior cite rule IDs under the excerpt discipline; free prose is for ledger mechanics only (owned by tools/differential/README.md).

**Working on docs/design/ has its own AGENTS.md.** `docs/design/AGENTS.md` carries the landing-a-design distillation checklist, the primary-source review rule, the dated-count convention, and the review axes. Claude Code loads it automatically when a session reads or edits anything under docs/design/; if your tool does not do nested discovery, read it yourself before touching those files or reviewing a change to them.

Expand Down Expand Up @@ -110,6 +110,11 @@ uv run sphinx-build -b html docs dist/docs
# code with the pipe's, so a failing run reads as a passing one. The
# classified summary it prints is the source for the release notes' behavior
# claims, including the count of changed names that are Latin-only.
# Exit 0 no longer means every diff is classified: since the tier split
# (#468) a radar corpus's unmatched diffs print under UNCLASSIFIED (radar)
# and cannot fail the run. Read that block. A radar diff worth a release
# note gets promoted (a cases.py row plus a shape tag) or classified with a
# rule BEFORE the log is drafted, not after.
# 2. Clear PRE_RELEASE in nameparser/_version.py — it carries 'dev' through the
# cycle (see step 9), so releasing is setting it to ''. VERSION should already
# be the version you are shipping; bump it here only if step 9 was skipped.
Expand Down
41 changes: 36 additions & 5 deletions docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -433,7 +433,7 @@ particle *ends* the name there is nothing ahead of it to join, and what
it belongs to is decided by what the writing says rather than by the
word. Two things say it, and both amount to someone stating that the
family name came first — a family comma, and a declared family-first
order — so a Dutch listing reads the same either way:
order:

.. doctest::

Expand All @@ -442,10 +442,41 @@ order — so a Dutch listing reads the same either way:
>>> family_first.parse("Jong Anke de").family # the order says so
'de Jong'

``FAMILY_FIRST`` is the only order this arises under, because it is the
only one that puts a trailing piece in the *middle*, where a particle
means nothing. ``FAMILY_FIRST_GIVEN_LAST`` puts it in the given slot,
where your own declaration says it is the given name, so it stays one:
That pair is not a coincidence but one shape written two ways: form 4
(``Title Family Given Middle Middle [Particle] [, Suffix]``) is form 2
(``Family [Suffix], Title Given (Nickname) Middle Middle[,] Suffix [,
Suffix]``) with the comma removed and the family folded inline —
titles included, so a title that form 2 writes after the comma leads
the name in form 4 instead. If your records are family-first without
commas, ``Policy(name_order=FAMILY_FIRST)`` reads them the way the
comma format is already read. Measured over the whole particle
vocabulary — every particle nameparser ships, crossed with three
families and three given names, 630 generated pairs in all — 603 of
630 agree (2026-08-30); the executable form of this correspondence is
``tests/v2/test_order_correspondence.py``.

Three limits keep that statement honest.

The correspondence covers one shape written two ways, not
comma-deletion in general: a name whose shape changes when the comma
is removed — a title or suffix crossing to a different position —
parses as the shape it becomes, not as a disagreeing reading of form
2. Where the trailing word is both particle and suffix vocabulary,
the two writings read it differently, and it is that ASYMMETRY rather
than a precedence that breaks the correspondence. The particle
attachment outranks the suffix reading on the comma side alone —
that is the scope the rule is stated in — so
``parse("Ménil, Christophe vd")`` reads family ``vd Ménil``, while
``family_first.parse("Ménil Christophe vd")`` reads family ``Ménil``
and suffix ``vd``. A listing ending in one of those three words
therefore does not correspond between the two writings — this is the
whole of the 27 disagreeing pairs, not scatter. And ``FAMILY_FIRST`` is the only order the correspondence
reaches at all, because it is the only one that puts a trailing piece
in the *middle*, where a particle means nothing; ``FAMILY_FIRST_GIVEN_LAST``
puts it in the given slot, where your own declaration already says
it is the given name, and no comma format writes the given name
last, so form 5 has no comma twin to correspond to in the first
place:

.. doctest::

Expand Down
Loading