diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 621a6b35..eb722c83 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -220,6 +220,8 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q TWO SECOND-ORDER MOVEMENTS, both recorded rather than repaired. The two one-case spellings part company, which is P3's own Accepted clause reaching a taken marker rather than a declined one: classify leaves a letter after a marker to the mixed-case rule, so `JANE DOE NEE PUIG I SOLER` never moved (the capital is an initial, and the initial veto kept the walk going all along) while `jane doe nee puig i soler` goes from family 'doe i soler' to maiden 'puig i soler'. And a report is GAINED where a longer clause ends on an ambiguous credential: `Jane Doe nee Puig i Ma` keeps maiden 'Puig i Ma' and says `suffix-or-name` where 0fbcaa0b read family 'Doe i Ma' in silence — the clause's own emitter, reaching a word the truncation had put out of its reach. A THIRD, found by the second review and the mirror of that one: a report is LOST where the clause takes a class member it used to end at. `Jane Doe nee MA i Soler` reads maiden 'MA i Soler' in silence where the parent 46651750 read maiden 'MA' with a `suffix-or-name` on it. That FOLLOWS from the link no longer ending the clause and is not a second decision: the report is the walk's own, raised on the LAST word the clause kept where a trailing rule was asked about it, and with the link joining, 'MA' has a name word behind it and is no longer that word — which is exactly what `Jane Doe nee MA Smith` has always done with the same acronym. The `y` twin is the oracle and it AGREES, in fields and in reports: `Jane Doe nee MA y Soler` reads maiden 'MA y Soler' in silence at the parent and here. Measured 2026-09-20 over the review grid, 24 names lose the report this way and all 24 agree with their `y` twin on every field and every report. - 2026-09-22 #397 follow-up — A DELIMITER CORE PAST THE CLAUSE'S FIRST WORD PASSES FOR THE NAME WORD BESIDE A LINK, RECORDED AS A DEVIATION RATHER THAN REPAIRED. The link exception above wants "a connective standing between two name words of the clause", and a separator the caller declared through `Policy.extra_suffix_delimiters` is structure, not a name word — so a link with one beside it joins nothing and should end the clause like any other suffix word. Between the marker and the clause's first word that already holds, the bound refusing the core before either piece test is asked. PAST that first word it does not: the core is an ordinary index to the run walk, which steps over connectives and nothing else, so it stands in for the name word on the link's left and the clause runs on past a title it would otherwise stop at. MEASURED 2026-09-22 under `extra_suffix_delimiters=(" - ",)`: `Smith, John, PhD née Puig Mr. - i Soler` reads maiden 'Puig Mr. i Soler' where its separator-less twin `Smith, John, PhD née Puig Mr. i Soler` stops at 'Puig Mr.'. THE POPULATION is the branch's own sweep, recorded with the code it describes (2026-09-21, corpus ∪ cases.py ∪ the property grids ∪ a 50,925-name generated set with cores, under thirteen core-bearing policies): the predicate is asked about a core in 51,072 of 900,023 calls, the answer differs from a core-skipping reading in 8,094 parses over 1,278 texts, and 1,824 of those move `maiden` on 288 texts — none of the 288 reachable at the default policy, `extra_suffix_delimiters` being empty there. NOT REPAIRED HERE: the fix threads the core set through three call sites into the run walk and moves the parent's reading as well, which makes it its own change rather than a rider on a review round. Open: #538. PINNED TWICE MEANWHILE. rules.md#M2 carries the shape as a `deviates: #538` example under an `extra_suffix_delimiters-dash` annotation — the first entry `tests/v2/rules_doc.py`'s registry has had for that field, named after the Policy field and carrying the delimiter in the suffix because the field's value is a set rather than a flag. And `tests/v2/pipeline/test_group.py::test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538` holds the pair at the piece level, named so nobody reads it as the contract. ONE COST OF THE DOC EXAMPLE, worth knowing before the repair lands: its string enters `corpus_rules.jsonl`, where the differential gate parses it with the DEFAULT facade — no delimiter declared, so the dash is an ordinary name word and #538's reading is off the path entirely. It moves there for the 2026-09-20 link fix instead, and is classified as that at all five baselines: added to the `fix(#397) a link inside a maiden clause stays in the birth name` alternation at the four 2.x ledgers (suffix 'PhD i Soler' → 'PhD', maiden 'Puig Mr. -' → 'Puig Mr. - i Soler', identical at each), and given its own rule at 1.4.0, where v1 had no maiden markers and read the whole suffix-comma tail as one suffix. +- 2026-09-26 #538 — A DELIMITER CORE IS STEPPED OVER BY THE LINK'S NEIGHBOUR SEARCH, AS A CONNECTIVE IS (Derek, 2026-09-26). The 2026-09-22 follow-up above recorded the deviation; this resolves it. `_run_neighbours` treats a lone core like a connective, so the word on a link's side is the one past the core and a maiden clause reads exactly as the same text written without the core: `Smith, John, PhD née Puig Mr. - i Soler` under `extra_suffix_delimiters=(" - ",)` now stops at maiden 'Puig Mr.' as its separator-less twin does. INVARIANT, pinned by `tests/v2/test_properties.py::test_a_delimiter_core_reads_as_if_it_were_not_written` over 81 generated texts, maiden field: 6 disagreed at e0f1a2fa, 0 after. Declined: THE BOUNDARY READING, which the 2026-09-22 entry's wording implied ("a link with one beside it joins nothing") — a core ending the neighbour search on its side. It fixes the reported shape equally, and it moves `Smith, John, PhD née Puig - i Soler` and `… Puig i - Soler` from maiden 'Puig i Soler' to maiden 'Puig', suffix 'PhD i Soler' — a birth-name link pushed into the credentials, where the skip reading leaves both as they read. Default-unreachable either way. THE SAME STEPPING APPLIES OUTSIDE A CLAUSE: the `frozen` loop's own `_run_neighbours` call also takes `cores`, so a connective beside a core looks past it; where a credential or nothing stands beyond, it stays a lone suffix word rather than joining, and the core it stood beside -- now a lone piece with nothing joined to it -- is dropped as #206 drops any lone core, the same way it already drops one between two ordinary post-nominals: `Smith, John, PhD - i Soler` 'PhD - i Soler' -> 'PhD, i Soler', `Smith, John, PhD - i - MD` 'PhD - i - MD' -> 'PhD, i, MD', `Smith, John, - i Puig` '- i Puig' -> 'i Puig', `Smith, John, Puig i -` 'Puig i -' -> 'Puig i'. MEASURED 2026-09-26: 191 of 8,097 generated core-bearing tail texts (the three heads `Smith, John,`, `Smith, John, Jr.,` and `John Smith,`, each followed by every 2-4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr.} containing a core) move `suffix` and no other field, none reachable at the default policy; the comparison is each text parsed under `extra_suffix_delimiters=(" - ",)` with `_run_neighbours` handed its cores and with it handed none. Not repaired here: a core the join MERGES into a joined piece -- interior (`… Puig Dr. i - y Soler`, still 'PhD i - y Soler') or at the edge a link joins across, a name word standing beyond it rather than a credential or nothing (`Smith, John, Puig - i Soler` 'Puig - i Soler', its separator-less twin 'Puig i Soler'; `Smith, John, Puig i - Soler` 'Puig i - Soler', measured 2026-09-26) -- is not dropped from the suffix text; the join is right, the surviving separator is the gap. Also in that family: an ORDINARY connective's join still takes a core as its neighbour, the stepping above being for generational vocabulary alone (rules.md#P3) -- `Smith, John, PhD - and MD` keeps suffix 'PhD - and MD' (measured 2026-09-26), pre-existing and unchanged here. Open: #549, which carries both the surviving core and this one. +- 2026-09-26 #535 — THE MAIDEN WALK READS THE TRAILING TITLE CHAIN, AND EVERY STOP ASKS ONE RELEASE QUESTION OF THE SPAN IT GIVES UP (Derek, 2026-09-26). This resolves the NOT TRANSPARENT HERE clause of the 2026-09-19 #533 entry above, which recorded `Jane Doe nee Smith MA Prof.` and `Jane Doe nee Smith Prof. MA` disagreeing and deferred the pair as one decision about what "trailing" means inside a clause. BOTH HALVES WERE TAKEN TOGETHER: a trailing title ends the clause (`Jane Doe nee Smith Prof.` reads title 'Prof.', maiden 'Smith', where 2.3.0 read maiden 'Smith Prof.'), and a credential or numeral standing in front of that title gets the stop it gets with the title absent (`… Smith MA Prof.` gives suffix 'MA', `… Smith V Prof.` suffix 'V'). Only the first half would have left the pair disagreeing in the other direction; H5's transparency is a statement about the peel and the chain read together to their fixed point, so the walk now reads the end of the clause the way assign reads the end of the name — `tail_reading`, the shared predicate mechanisms.md#ONE-PREDICATE-PER-QUESTION names — and over the forms the transparency property test below exercises, the two spellings give one answer. They split where something in or ahead of the clause stands to take the title — for example a particle chain (ahead of the clause or inside it), a bound-given join, a name left with no name word, or, after a family comma, a title word in front of a class member with a title behind it, which the given part's chain stops short of exactly as it does with no marker (`Doe, Jane nee Smith Rev. MA Prof.` reads maiden 'Smith Rev.', title 'Prof.', suffix 'MA', where `… Rev. Prof. MA` reads title 'Rev. Prof.', maiden 'Smith', suffix 'MA'; bare `Doe, Jane Rev. MA Prof.` reads middle 'Rev.') — and the Accepted pairs recorded here and in rules.md#M2 are examples of that, not a complete list. `tests/v2/test_properties.py::test_a_trailing_title_is_transparent_to_the_maiden_clause` holds that as an invariant over the two inputs, over heads and runs where the credential is given up and nothing ahead would take the title, with its negative control at e0f1a2fa in its docstring. ACCEPTED, THE TRANSPARENCY BOUNDARY: where the clause KEEPS the credential the spellings differ, because a clause is one contiguous run and a title inside the kept text cannot leave without the words behind it — `Doe nee Smith ba Prof.` reads title 'Prof.', maiden 'Smith ba', and `Doe nee Smith Prof. ba` maiden 'Smith Prof. ba'; `abdul nee Smith MA Prof.` reads title 'Prof.', family 'abdul', maiden 'Smith MA', and `abdul nee Smith Prof. MA` given 'abdul', maiden 'Smith Prof. MA' (2.3.0 kept every word in all four). The report differs with them: `Doe, Dr. nee Smith MA` and `… Prof. MA` report `suffix-or-name` on the kept MA, `… MA Prof.` reports nothing, the kept member no longer ENDING the clause, which is what M2's report is asked of. rules.md#M2 carries the first pair as an Accepted example. ACCEPTED, FURTHER SPLITS, measured 2026-09-26. Where something ahead of the clause would take the title, the spellings differ even with the credential given up in both, because the first-suffix-word stop ends the clause before the title check runs and the title check alone declines: `Jane van der Berg nee Smith PhD Prof.` reads title 'Prof.', maiden 'Smith', suffix 'PhD', and `… Prof. PhD` maiden 'Smith Prof.', suffix 'PhD' — H5's accepted particle-chain boundary inherited, the bare `Jane van der Berg Smith PhD Prof.` / `… Prof. PhD` splitting the same way — and `Berg, abdul nee Smith PhD Prof.` (bound-given join) and `Doe, Dr. nee Smith PhD Prof.` (no name word left) split likewise; the parent d9d80492 and 2.3.0 read all of these as the tree does, so only the claim is new. And a released particle with a title behind it is withdrawn: `Jane Doe nee Smith DO Prof.` reads title 'Prof.', maiden 'Smith DO', where `… Prof. DO` reads title 'Prof.', maiden 'Smith', suffix 'DO' (2.3.0 maiden 'Smith DO Prof.' and 'Smith Prof. DO'); `do` likewise. A particle INSIDE the clause does it from the other side, its chain able to take the credential behind it: `Jane Doe nee Smith do MA Prof.` reads title 'Prof.', maiden 'Smith do MA', and `… do Prof. MA` title 'Prof.', maiden 'Smith do', suffix 'MA' (2.3.0 kept every word in both). rules.md#M2 carries a pair of each shape as Accepted examples. ACCEPTED, H5'S REACH: `Jane Doe nee Smith King.` reads title 'King.', maiden 'Smith' (2.3.0 maiden 'Smith King.'), H5's accepted `Mary Jane King.` cost now reaching the last word of a birth name as it reaches the last word of a current one; rules.md#M2 carries it as an Accepted example. READER SCOPE IS UNCHANGED: the chain is consulted only where a trailing rule reads the clause's words (no comma, the part before a suffix comma, the given part after a family comma); before a family comma and in a tail segment the walk reads the peel alone and the clause keeps the title (`Doe nee Smith Prof., Jane` keeps maiden 'Smith Prof.'). THE FIRST-WORD FLOOR MOVED INTO THE CHAIN. `trailing_titles` and `tail_reading` take a `floor` (1 everywhere else), and the walk passes the position just past the marker's first word, so the chain never takes that word: `Jane Doe nee King.` keeps its maiden name (TITLES holds borne surnames and the marker announced one), and `Jane Doe nee Prof. Dr.` keeps 'Prof.' and gives up 'Dr.'. The first draft applied the floor as a clamp on the title stop after the chain had run, and that is measurably wrong rather than merely inelegant: a chain allowed to take the first word has already spliced it out of the count the re-peel reads, so `Jane Doe nee King. ba` kept 'ba' in the clause where `Jane Doe nee Smith ba` gives it up. `test_a_title_first_word_counts_as_a_word` is the invariant (a title-vocabulary first word against an ordinary one, the credential behind it read alike) and its docstring records the clamp's failure count. ONE RELEASE CHECK FOR THREE STOPS. The numeral, the credential and the title stop each ask `_release_reads_off` whether the name the take would leave reads what they give up as titles or suffixes with no join below the take absorbing it — rules.md#M2's invariant, "a word the clause gives up reads as a post-nominal or the clause keeps it", which the title stop now answers to as well. It is asked of the whole SPAN a stop gives up, not of the stop's own word: a first draft that checked the word alone broke the invariant on 513 of a 5,198-parse sweep (measured 2026-09-26 on that draft, which is not in the tree to recompute), because a stop gives up every word behind it — `DOE NEE SMITH PROF. MA` stopped at the title and handed the MA to the family, where the name left standing reads MA as a name; it keeps maiden 'SMITH PROF. MA' now. The question is asked per reader, the way that reader reads: the TRAILING reader is assign's own `tail_reading` over the view; the GIVEN_SLOT reader has its count settled by the comma, so it is the writing alone with a name word ahead. The name-word-ahead half is what keeps `Dr. nee Jones Smith Prof.` whole — the take would leave `Dr. Prof.`, whose Prof. would be the family name. THE JOINS. A released span holding a title with a particle ahead of it declines, because P2's chain runs on over a trailing title (rules.md#H5's Accepted `John van der Berg Prof.`): `Jane van der Berg nee Smith Prof.` keeps maiden 'Smith Prof.', and a released particle with a title behind it is withdrawn for the same reason (`Jane Doe nee Smith MA do Prof.` keeps maiden 'Smith MA do'). The bound-given (P5) half of `_join_takes_the_member` is asked for the GIVEN_SLOT reader alone: P5 is `BoundJoin.LENIENT` only after a family comma, and before one its STRICT reserve already refuses a join that would change a suffix reading, so asking there over-declined — `abdul nee Smith V` kept 'V' in the clause, where the scoped check lets it go to suffix as 2.3.0 did. The consequence is recorded by a case row rather than prevented: `abdul nee Smith Dr.` reads title 'Dr.', family 'abdul', maiden 'Smith', which is how `abdul Dr.` reads bare (2.3.0 read given 'abdul', maiden 'Smith Dr.'). THREE NUMERAL-STOP READINGS MOVE WITH IT, decided in rather than deferred, since each is the same invariant broken at the stop this change was already rewriting. The numeral stop never asked the join question: `Berg, abdul nee Smith V` read given 'abdul V' at 2.2.0 and 2.3.0 (given 'abdul nee', suffix 'V' at 2.0.0 and 2.1.0, #411's reserve differing between those pairs before the walk runs) and reads maiden 'Smith V' now. After a family comma the given slot reads a lone numeral as a suffix only where the given part is the LAST comma part — assign's own two-segment condition from #144 — and the walk now asks it too (`tail_follows`): `Doe, Jane nee Smith V, PhD` read middle 'V' at 2.2.0, 2.3.0 and e0f1a2fa, an M2 violation predating this change that its first draft had extended to `… V Prof., PhD`, and it reads maiden 'Smith V' now, as 2.0.0 read it. The withdrawal reaches through a title behind the numeral, so `Doe, Jane nee Smith Prof. V, PhD` keeps maiden 'Smith Prof. V' where `Doe, Jane nee Smith V Prof., PhD` gives up the title — the same asymmetry the bare given slot has with no marker in it (`Doe, Jane Prof. V, PhD` against `Doe, Jane V Prof., PhD`). And the numeral stop reads FROM the marker, unlike the other two, so a numeral straight after the marker is not held to the first-word floor, and it may decline the clause; examples, not a rule over every shape: `Dr. nee V` and `Doe, J. nee V` keep maiden 'V' as 2.3.0 did, though `Doe, J. V` reads suffix 'V', and `Doe nee V, Jane` and `Smith, John, PhD nee V` read no clause at all (family 'Doe nee V', suffix 'PhD nee V', as at 2.3.0) — `Jane Doe nee V Prof.` read maiden 'V Prof.' at 2.3.0 and reads family 'nee', suffix 'V', title 'Prof.' now, as `Jane Doe nee V` reads plus the title. After a family comma with another comma part behind the given one, the same #144 condition that keeps `Doe, Jane nee Smith V, PhD` whole now keeps a lone numeral in the clause as well: `Doe, Jane nee V, PhD` read given 'Jane', middle 'nee V', suffix 'PhD' at 2.3.0 and at the parent d9d80492 and reads maiden 'V', suffix 'PhD' now, as 2.0.0 and 2.1.0 read it, `Doe, Jane nee V, Jr.` likewise — the marker had been read as a name word there at 2.2.0 and 2.3.0, and a case row records it. A LINK THE WALK STOPS AT ASKS THE RELEASE QUESTION TOO. Reading the end of the clause through the title chain makes the link exception refuse a link it used to join — in `Jane Doe nee Smith i DO Prof.` the DO is the peel's once the title is chained — and a stop at a link gives up the words behind it. Unchecked, that put 'Doe i' in the middle name and 'DO Prof.' in the family (e0f1a2fa read maiden 'Smith i DO Prof.'), and `Berg, abdul nee Smith i V Prof., MD` read middle 'V'. So a link the exception refuses asks `_release_reads_off` of the run it would give up — only where the exception, bounded by the peel over the words as WRITTEN, would have joined it (the refusal is the title chain's), where a trailing rule reads the clause, and where the link is not the first word after the marker (a first-word stop declines the clause and gives nothing up); where that fails the clause keeps the link and walks on. Scoped that narrowly because a first version asked it of every refused link and kept links the name left standing reads off: `Doe, Jane nee Smith i V` kept maiden 'Smith i' where every release gives up suffix 'i V' (23,472 of a 411,936-parse link grid moved clean readings, measured 2026-09-26). With it, the given-part model reads the lenient numeral in both of assign's passes — the literal last piece first, then the last one standing once the chain has taken the titles behind it — and only where no comma part follows (`Doe, Jane nee Smith i V Prof.` gives up 'i V' and the title, as bare `Doe, Jane i V Prof.` reads them; `… i V Prof., PhD` keeps maiden 'Smith i V'). The same model moves the credential stop where a lenient numeral follows the credential: `Doe, Jane nee Smith MA V` now gives up 'MA V' (suffix 'MA V', maiden 'Smith') as bare `Doe, Jane MA V` reads it, where d9d80492 and 2.3.0 read maiden 'Smith MA', suffix 'V'; `… MA V, PhD` keeps maiden 'Smith MA V'. Over that link grid (links i, y, e, and, i y, y i, and i × 8 heads × bodies × `, MD` tail × 3 cases × 3 name orders), against the tree before the link check: 0 new violations, 6,195 fixed, and the clean readings that move now match the bare given part's (`DOE, JANE NEE SMITH MA I` gives up 'MA I', as `DOE, JANE MA I` reads suffix 'MA I'). `Jane Doe nee Smith i DO Prof.` now reads title 'Prof.', maiden 'Smith i DO' and reports the kept DO, and `Berg, abdul nee Smith i V Prof., MD` maiden 'Smith i V Prof.', suffix 'MD'. A RELEASED TITLE THAT IS ALSO A PARTICLE IS KEPT AFTER A FAMILY COMMA, because P6 attaches a particle trailing the given part to the family: released by the title stop, 'St.' in `Doe, Jane nee Smith St.` reads family 'St. Doe', so `_join_takes_the_member` declines it and the name reads maiden 'Smith St.', as e0f1a2fa and 2.3.0 read it, and so does `Doe, Jane nee Smith MA St.`. Titles only — a credential that is also a particle ('DO') is the given slot's own lean, which the credential stop has already asked (#533). A FOURTH MEASUREMENT, against e0f1a2fa rather than d9d80492: the M2 invariant over a fuzz of twelve heads ({`Jane Doe`, `Doe, Jane`, `Doe, Prof.`, `Jane Doe, PhD`, `Doe, J.`, `Berg, Jane van der`, `Jane van der Berg`, `Berg, abdul`, `abdul Berg`, `J. Doe`, `Prof. Jane Doe`, `Dr.`}), ` nee Smith ` and every one-to-three-word sequence over {MA, Ma, V, Prof., St., King., PhD, ba, do, DO, Jr., i, van, y} holding at least one of Prof., St., King., then nothing or `, MD`, as written, upper- and lower-cased (107,352 parses, measured 2026-09-26 on the narrowed tree; the intermediate tree above read 2,228): 588 texts that violated it at e0f1a2fa no longer do, and 12 newly do, all `Dr. nee Smith i MA|V` followed by `Prof.`, `King.` or `St.`, with or without `, MD`. Each is the title-carrying twin of `Dr. nee Smith i MA` / `Dr. nee Smith i V`, which read family 'i' at e0f1a2fa already: the bare-title head leaves no name word, the unguarded stop's class (#548), which the title now reaches because it leaves the clause. ACCEPTED: `Jane Doe (nee Smith Prof.)` reads title 'Prof.', maiden 'Smith' — bracket content ending in a period is suffix-shaped (S1), so the brackets are dropped and the clause is read bare, outside the reach of the delimiter precedence; rules.md#M2 carries it as an Accepted example. MEASURED 2026-09-26, the tree against its parent d9d80492, over three generated grids, and these are dated snapshots. Recipe: grid A is each head in {`Doe, Jane`, `Doe, J.`, `Doe, Prof.`, `Berg, abdul`, `Doe, Jane van`, `Jane Doe`, `J. Doe`, `Dr.`, `abdul`, `Jane van der Berg`, `John`}, then ` nee `, then every sequence of one to three words (repetition allowed) from {Smith, Jones, Prof., Dr., Sir, MA, DO, PhD, V, III, i, do, van, Ma, M.A., King., Rev., ba}, then either nothing or `, PhD`, each text as written, upper-cased and lower-cased, duplicates removed — 327,936 texts; grid B, aimed at the bound-given join, is the same construction over heads {`Dr. abdul`, `abdul rahman`, `Dr. abdul rahman`, `Berg, Dr. abdul`, `Berg, abdul rahman`, `abu`, `Berg, abu`, `Mr. abu`, `Dr.`, `Berg, Dr.`, `abdul`, `Berg, abdul`}, words {Smith, Prof., MA, V, do, van, PhD, Ma, ba, III, Jr., M.A.} and tails {nothing, `, PhD`, `, Jr., MD`}, plus these fourteen texts, as written only: `Jane Doe nee Ph. D. Prof.`, `Jane Doe nee Ph. D. Smith Prof.`, `Jane Doe z domu King. ba`, `Jane Doe z domu Smith Prof.`, `Jane Doe nee King. Prof. ba`, `Jane Doe nee King. Prof.`, `Doe, Jane nee Smith V,`, `Doe, Jane nee Smith V, ` (with the trailing space), `Doe, Jane, nee Smith V`, `Doe, Jane nee Smith V Prof., PhD`, `Doe, Jane nee Smith Prof. V, PhD`, `Doe, Jane nee Smith MA, PhD`, `Doe, Jane nee Smith V (Jr.)`, `Doe, Jane nee Smith V "Bo"` — 173,057 texts. The invariant tested is M2's: in a parse with a non-empty maiden field, every token written after the ` nee ` marker is roled maiden, title or suffix (so the two `z domu` texts are parsed and not checked). Grid C, a four-word probe of the title chain after a family comma, is each head in {`Doe, Jane`, `Jane Doe`, `Doe, J.`, `John Smith`, `Doe, Jane Mary`}, then ` nee `, then every sequence of four words (repetition allowed) from {Smith, Jones, Prof., Dr., King., MA, PhD, V, Jr., ba, do}, then either nothing or `, PhD`, as written only — 146,410 texts. Per grid, the counts are: texts that newly violate the invariant at the tree, texts that violated it at the parent and no longer do, texts that violate it at both. Grid A: 0, 2,012 and 17,854; grid C: 0, 1,368 and 15,534; grid B: 0, 1,222 and 11,149 (all re-measured 2026-09-26 on the tree with the link and particle-title checks below, as narrowed; an intermediate tree whose link check fired on every refused link read grid A 0, 3,772 and 16,094, its extra 1,760 fixes being links the as-written reading refuses too, which are #548's class), one of those last being the check itself counting the quoted nickname in `Doe, Jane nee Smith V "Bo"`. In grids A and B, every other remaining violation has the clause ending immediately before a word of the unambiguous suffix vocabulary written the way the plain suffix-word stop takes it (PhD, III, M.A., Jr., and a lower-case i or v; never a bare capital, which reads as an initial), which is where that stop ends a clause without asking a release question. DEFERRED: the walk's plain suffix-word stop — the one ending the clause at the first suffix word, as distinct from the trailing numeral, credential and title stops — asks no release question at all, and where the words behind it cannot read as post-nominals they land in a name part: on degenerate heads the stop word itself does (`Dr. nee Smith PhD Prof.` reads family 'PhD', `Doe nee Smith Jr. Prof., Jane` family 'Doe Prof.'), and with an ordinary head the words behind it do (`Doe, Jane nee Smith PhD Smith` reads middle 'Smith'). Unchanged here, and what "the clause keeps it" should mean where the kept words would follow a credential the clause itself ended at is its own question; rules.md#M2 carries `Doe nee Smith Jr. Prof., Jane` as an Accepted example meanwhile. The trailing numeral's stop has the same gap before a family comma, where it is made over the peel alone with no release question: `Doe nee Smith V, Jane` reads family 'Doe V', maiden 'Smith' (so did 2.2.0, 2.3.0 and the parent d9d80492; 2.0.0 and 2.1.0 read maiden 'Smith V'), and rules.md#M2 carries it beside the other. Open: #548. COST, measured 2026-09-26 on CPython 3.11.16 as profiler call events in one `Parser.parse` after a warm-up parse, parent → tree: `John Smith` 171 unchanged, as are `Jane Doe Prof.` and `John Smith MA`; `Jane Doe nee Smith` 244 → 246; `Jane Doe nee Smith MA` 377 → 388; `Jane Doe nee Smith Prof.` 276 → 372, the one shape that now runs the chain and a release check it never ran. Recompute: count `sys.setprofile` call events around the second of two `Parser().parse` calls, with the tree and d9d80492 each first on `sys.path`. ### N3 — the lone-word nickname rule @@ -389,6 +391,8 @@ The reconciled v1-style banks (`tests/test_*.py`) carried eight `@pytest.mark.xf - 2026-09-24 #478 — CROSS-REFERENCE, no change to this rule: case repair's hyphen clause is the one place a marked letter in a name written wholly in one case is read as the connective rather than an initial — `maria silva-e-sousa` repairs to `Maria Silva-e-Sousa` beside the spaced `Maria Silva E Sousa`, and `conjunctions_ambiguous` does not reach the hyphenated form. Decided there, with its cost (`J-E-P DUPONT` → `J-e-P Dupont`): decisions.md#R4's 2026-09-24 #478 bullet. +- 2026-09-26 #538 — CROSS-REFERENCE: rules.md#P3's separator sentence (a declared delimiter is read past as a connective is) is decided and measured in the 2026-09-26 #538 entry under `### M2`. + - Provenance: the single-letter-connective guard is v1's fix for Google Code issue 11 ("john e smith", 2013, commit 33676c9) — the "#11" citations that circulated pointed at a GitHub accident, not the real source. Recorded so the archaeology stays done. Excluded (Lexicon.conjunctions_ambiguous, the marked half of nameparser/config/conjunctions.py — an entry here reads as an INITIAL in a name written wholly in one case): @@ -547,6 +551,7 @@ Decided 2026-09-08 (was Open: [#316](https://github.com/derek73/python-nameparse - **Nothing else moved, and it was measured rather than argued.** Snapshot all seven fields, the ambiguity kinds and the recorded `order` for every name in `tools/differential/corpus*.jsonl` plus every string literal in `tests/v2/cases.py` (`ast.walk` over that file, which sweeps up the notes and the case ids too — harmless, they parse like anything else), on the branch tip and on the fix, and diff: 2682 strings before and 2681 after, 2674 of them on both trees, and NONE of those 2674 reads differently. The seven that differ are the case rows this round added and dropped, not parses that moved. The two readings above are inputs no corpus holds — as is every other member of the class the bullet below measures — and each takes a case row (`title_word_trailing_run_is_read_to_a_fixed_point`, `title_word_trailing_behind_a_bound_pair_at_the_peel_reserve`). The same round dropped the `Sir Jr` example from rules.md#S2 and its case row: it reads by the same branches as `Dr Jr`, the same roles and the same reported kind, the run being empty by then and no branch reading `vocab:given-title` — so the pair pinned one reading twice. Its leaving corpus_rules.jsonl moves two ledger rosters: the jr suffix-routing rule 7 → 6 corpus names in the 1.4.0 ledger, and the #489 peel-floor rule 4 → 3 in all four. +- 2026-09-26 #535 — CROSS-REFERENCE: the maiden half of rules.md#H5's "the chain reads PIECES" Accepted row now reads the title — a maiden clause is no longer a join that puts the trailing word out of the chain's reach, so `Mary Smith née Jones Prof.` reads title 'Prof.' — decided in the 2026-09-26 #535 entry under `### M2`. ### W1 — unspaced CJK division diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 2d728513..1c8637fa 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -55,7 +55,7 @@ Problem shape. "Which stage does X?" — asked before attributing behavior in pr ## ONE-PREDICATE-PER-QUESTION — one predicate answers it, and every other site calls that -Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from three sites that each needed the identical question answered — classify's own tag emission, this module's ambiguous_class_candidate, and _segment.py's multi-token run test — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold, because every one of the three callers' own FIRST conjunct is the caller-configured switch itself, `Policy.unlisted_caps_suffixes`, False by default, so the shared call is never reached at the default regardless of how many callers share it (decisions.md#S2)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain and M2's walk stop at — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. +Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from three sites that each needed the identical question answered — classify's own tag emission, this module's ambiguous_class_candidate, and _segment.py's multi-token run test — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold, because every one of the three callers' own FIRST conjunct is the caller-configured switch itself, `Policy.unlisted_caps_suffixes`, False by default, so the shared call is never reached at the default regardless of how many callers share it (decisions.md#S2)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. ## RENDER-HONORS-THE-PARSE — the parse decides it, the views honor it diff --git a/docs/design/rules.md b/docs/design/rules.md index e0e88686..4b3a0e71 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -320,11 +320,13 @@ H5. Rationale: a word abbreviated with a period at the END of a name report and not this rule's. Accepted: the chain reads PIECES, so a join that ran earlier puts the word out of reach. A particle chain (P2) has already - taken the trailing word into the family name, and a maiden - marker (M2) has already taken it into the maiden name; in - neither is a title word standing in the trailing slot at all. + taken the trailing word into the family name, and no title word + is standing in the trailing slot at all. A maiden clause is not + such a join: the clause's walk reads the end of the name through + this chain (M2), so a trailing title ends the clause and is a + title, where the name the take leaves reads it as a title (M2). "John van der Berg Prof." → family="van der Berg Prof." - "Mary Smith née Jones Prof." → maiden="Jones Prof." + "Mary Smith née Jones Prof." → title="Prof." Accepted: what the chain leaves is also what counts as a name word to spare (P5). A trailing title word is not one, so a bound given-name word behind one joins exactly as it joins with the @@ -534,6 +536,11 @@ P3. Rationale: connective words ("y", "of the") bind name words into nothing, and a word of that vocabulary ending a name, or standing before the credential a name ends with, is the generation it also spells. + A separator the caller declared is not a word of the name, and + the search reads past it as it reads past a connective: the word + on the connective's side is the first one beyond the separator, + so a connective of that vocabulary beside one joins exactly where + it would with the separator absent. Both questions this rule asks of a name — how many words it has, and whether it is written in one case — are asked of the name's OWN words: a maiden marker taken as one, and the words it takes @@ -635,7 +642,7 @@ P3. Rationale: connective words ("y", "of the") bind name words into same two words unjoined are two name words and H1 does not fire. P1's leading run is the second (#395, landed): its run takes the "Vega y Santos" join whole or stops before it. - history: decisions.md#P3 · interacts: H1, P1, M2, R3, R4, S2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_post_rules.py + history: decisions.md#P3 · interacts: H1, P1, M2, R1, R3, R4, S2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_post_rules.py P4. Rationale: a particle links forward from inside a name; at the very front there is no name yet to be inside. @@ -931,7 +938,7 @@ S1. Rationale: brackets set off more than nicknames — credentials as if written bare. "Andrew Perkins (MBA)" → suffix="MBA" "Andrew Perkins (Andy)" → nickname="Andy" · boundary - implemented: nameparser/_pipeline/_extract.py + interacts: M2 · implemented: nameparser/_pipeline/_extract.py S2. Rationale: generational suffixes and credentials are recognized by vocabulary; an acronym that is also an ordinary name is only @@ -1269,15 +1276,33 @@ M2. Rationale: a maiden marker announces that what follows it is the reading the name left standing reads the word as the credential, and never the first word after the marker — as the maiden name, and - the marker itself is dropped. + the marker itself is dropped. Where a trailing rule reads the + words, a trailing title ends it too: a period-marked word the + trailing title chain takes (H5), read together with the trailing + suffix run to where neither takes more, so a credential or a + numeral in front of the title stops the take exactly as it does + with the title absent. Neither the credential stop nor the title + stop takes the first word after the marker; the numeral stop is + not held to that, and a numeral there may decline the clause. One suffix word does not stop it. Where such a word is also a connective standing between two name words of the clause (P3), a link inside the birth name does not end it, and the words on both sides of the link are the maiden name. A link with the marker on one side of it, or with the trailing run on the other, is joining nothing there and ends the clause like any - other suffix word. - Those last two stops are each asked TWICE for one reason: the + other suffix word. Where a trailing rule reads the words, a link + that ends the clause only because a trailing title is read as + one — a link that, read over the words as written, would join — + gives up the words behind it only where the name left standing + reads them as post-nominals or titles; otherwise the clause keeps + the link and runs on. The link first after the marker is not + asked: stopping there declines the clause. + A separator the caller declared is structure rather than a name + word, and the link exception reads past it: the word on a link's + side is the one beyond the separator, so the clause reads as the + same clause written without it. + The trailing numeral, credential and title stops are each asked + TWICE for one reason: the count of words to spare includes the very words the marker removes, so a reading taken over the name as written can be wrong about the name the take would leave. WHICH rule does the @@ -1286,10 +1311,13 @@ M2. Rationale: a maiden marker announces that what follows it is the the part before a SUFFIX comma, which that rule reads the same way — that rule is the reader. After a family comma it is the reading the end of the given part takes, where the comma has - already settled the count and the writing decides alone. Before - a family comma, and in a part after a second one, no trailing - rule reads those words at all: the clause keeps them and says - nothing about them. + already settled the count and the writing decides alone. A lone + numeral reads as a suffix there only where no comma part follows + the given one, exactly as it does outside a marker's clause; with + one behind it the numeral stays name text and the clause keeps + it. Before a family comma, and in a part after a second one, no + trailing rule reads those words at all: the clause keeps them and + says nothing about them. That the credential stop spares the first word after the marker is a deliberate divergence from what certain suffix vocabulary gets in the same position, where the marker declines and stays @@ -1317,13 +1345,24 @@ M2. Rationale: a maiden marker announces that what follows it is the rule sees it, which would carry a word of the BIRTH name into the current one. In both the clause keeps the word, and reports it as it reports every member it keeps. + Where a trailing rule reads the words, the trailing numeral, the + trailing credential and the trailing title each give up a run + only where the name the take leaves + reads the WHOLE run, not only the word the stop is made at, as + titles or post-nominals; otherwise the clause keeps it. After a + family comma a released title that is also a particle is kept, + since the particle attachment (P6) would carry it into the family. Delimiters outrank every reading inside them. Where a recognized marker stands inside a delimited clause, the whole span is the maiden name whatever its last word is, and whether or not the pair is a configured maiden delimiter (M3): the writer drew the boundary, so no fork is called and nothing is reported. A word the writer left OUTSIDE the span is outside the clause and reads - as it would anywhere else. + as it would anywhere else. Bracketed content ending in a period + is the exception, because it never reaches this rule as a + delimited clause: a trailing period makes the content + suffix-shaped (M3 states this), so the brackets are dropped (S1) + and the clause is read as if written bare. A marker with nothing after it, or nothing before it, is just a word. A marker may be more than one word, and is then recognized only @@ -1345,6 +1384,7 @@ M2. Rationale: a maiden marker announces that what follows it is the "John née Jones Smith V" → maiden="Jones Smith" "John née Jones Smith V" → suffix="V" "Jane Smith née V" → suffix="V" + "Dr. nee V" → maiden="V" · boundary "J. née Jones Smith V" → maiden="Jones Smith V" · boundary "Jane née Jones J. V" → maiden="Jones J. V" · boundary "Jane Doe nee Smith MA" → maiden="Smith" @@ -1359,6 +1399,23 @@ M2. Rationale: a maiden marker announces that what follows it is the "Jane Doe nee Puig i Soler" → maiden="Puig i Soler" "Jane Doe nee Puig i" → maiden="Puig" · boundary "Jane Doe nee Puig i" → suffix="i" · boundary + "Jane Doe nee Smith i DO Prof." → maiden="Smith i DO" + "Doe, Jane nee Smith St." → maiden="Smith St." + "Smith, John, PhD née Puig Mr. - i Soler" extra_suffix_delimiters-dash → maiden="Puig Mr." + "Smith, John, PhD née Puig - i Soler" extra_suffix_delimiters-dash → maiden="Puig i Soler" + "Jane Doe nee Smith Prof." → maiden="Smith" + "Jane Doe nee Smith Prof." → title="Prof." + "Jane Doe nee Smith MA Prof." → suffix="MA" + "Jane Doe nee Smith Prof. MA" → maiden="Smith" + "Jane Doe nee Smith V Prof." → suffix="V" + "Jane Doe nee King." → maiden="King." · boundary + "Jane Doe nee Prof. Dr." → maiden="Prof." · boundary + "Dr. nee Jones Smith Prof." → maiden="Jones Smith Prof." · boundary + "Doe nee Smith Prof., Jane" → maiden="Smith Prof." · boundary + "Jane van der Berg nee Smith Prof." → maiden="Smith Prof." · boundary + "Berg, abdul nee Smith V" → maiden="Smith V" · boundary + "Jane Doe nee King. ba" → suffix="ba" · boundary + "Doe, Jane nee Smith V, PhD" → maiden="Smith V" "Jane Doe (nee Smith MA)" → maiden="Smith MA" "Jane Doe (nee Smith Ma)" → maiden="Smith Ma" "Jane Doe (nee Smith) MA" → suffix="MA" @@ -1392,24 +1449,57 @@ M2. Rationale: a maiden marker announces that what follows it is the a reading needs is taken over the name the take would leave rather than over the words as they stand. "John née Jones Smith Ma" → maiden="Jones Smith Ma" - Accepted: a trailing title is not transparent inside a clause, - and the two spellings disagree — the walk reads the trailing - credential run and not the title chain behind it, so a title - AFTER a member of that class hides it and a title before it - does not. - "Jane Doe nee Smith MA Prof." → maiden="Smith MA Prof." · boundary - "Jane Doe nee Smith Prof. MA" → maiden="Smith Prof." · boundary - Deviation: the link exception asks for a name word on each side, - and a separator the caller declared is structure rather than a - name word — so a link with one beside it is joining nothing and - ends the clause like any other suffix word. A declared separator - standing inside the clause, past its first word, is read as that - name word instead, and the clause runs on across a link it should - have ended at. The same clause written without the separator, - which leaves the title as the word on the link's left, does end - there. - "Smith, John, PhD née Puig Mr. - i Soler" extra_suffix_delimiters-dash → maiden="Puig Mr." deviates: #538 (today: maiden="Puig Mr. i Soler") - history: decisions.md#M2 · interacts: P2, P3, P5, P6, R2, M1, S2, H1, H5 · implemented: nameparser/_pipeline/_group.py + Accepted: bracket content ending in a period is suffix-shaped + (M3), so the brackets are dropped (S1) and the clause is read as + if written bare — the delimiter precedence above does not reach + it, and a trailing title inside gives itself up as the bare + clause's does. + "Jane Doe (nee Smith Prof.)" → title="Prof." + Accepted: the title is transparent only where the clause gives + the credential up. A clause is one run of words, so where it + KEEPS the credential a title written behind it can still leave, + while one written in front of it cannot leave without the words + behind it, and the two spellings differ. + "Doe nee Smith ba Prof." → maiden="Smith ba" + "Doe nee Smith Prof. ba" → maiden="Smith Prof. ba" + Accepted: where the clause gives the credential up in both + spellings, the title still leaves only where nothing ahead stands + to take it. A particle chain ahead runs on over a trailing title + (H5), and a bound given-name join or a name left with no name + word would absorb it too, so a title in front of the credential + stays in the clause while one behind it, which the first suffix + word has already cut off, leaves — the split H5 already accepts + for the same name written without a marker. + "Jane van der Berg nee Smith PhD Prof." → title="Prof." + "Jane van der Berg nee Smith Prof. PhD" → maiden="Smith Prof." + Accepted: a released particle with a title behind it is withdrawn, + because the particle chain would run on over the title (P2, H5); + so the particle written in front of the title stays in the clause, + while written behind it, it leaves with the title. A particle + standing INSIDE the clause, ahead of the credential, does the + same from the other side: its chain would take the credential, + so the clause keeps the credential, and a title can leave only + from behind it. + "Jane Doe nee Smith DO Prof." → maiden="Smith DO" + "Jane Doe nee Smith Prof. DO" → suffix="DO" + "Jane Doe nee Smith do MA Prof." → maiden="Smith do MA" + "Jane Doe nee Smith do Prof. MA" → suffix="MA" + Accepted: H5's reach into the ordinary surnames the title + vocabulary holds reaches the end of a clause as it reaches the end + of a name: a period written behind one ends the clause as a title + and takes the word out of the birth name. + "Jane Doe nee Smith King." → title="King." + Accepted: the stop at the first suffix word asks no question of + what it gives up, so a clause that ends there can hand a word + behind it to a name part, where the invariant stated above says + the clause keeps it. Before a family comma, where the statement + above says no trailing rule reads the words and the clause keeps + them, the trailing numeral's stop is still made, asked of the + peel alone with no such question, and a lone numeral written + there goes to the family. Open: #548. + "Doe nee Smith Jr. Prof., Jane" → family="Doe Prof." + "Doe nee Smith V, Jane" → family="Doe V" + history: decisions.md#M2 · interacts: P2, P3, P5, P6, R1, R2, M1, S1, S2, H1, H5 · implemented: nameparser/_pipeline/_group.py M3. Rationale: an enclosure says nothing about whether it means maiden, but a recognized marker word inside it does — the clause @@ -2014,6 +2104,7 @@ R1. Rationale: a field is a way of reading the parse, not a stored "Smith, MD PhD" → suffix="MD PhD" "John Smith MD PhD" → suffix="MD PhD" "John Smith, MD, Bart" → suffix="MD, Bart" + "Smith, John, PhD - i Soler" extra_suffix_delimiters-dash → suffix="PhD, i Soler" Accepted: a suffix value handed to revise() derives its entries the same way, from the value's own commas, so a name's rendered suffix revises back to itself wherever the value's words read as the @@ -2030,7 +2121,7 @@ R1. Rationale: a field is a way of reading the parse, not a stored instead. Stated without an example line because every line here names an input string, and this shape needs a field revised after the parse. - history: decisions.md#C1 · interacts: O3, P6, R3 · implemented: nameparser/_parser.py, nameparser/_pipeline/_post_rules.py, nameparser/_types.py + history: decisions.md#C1 · interacts: O3, P3, P6, R3, M2 · implemented: nameparser/_parser.py, nameparser/_pipeline/_post_rules.py, nameparser/_types.py R2. Rationale: callers need the surname with and without its particles — sorting wants "Vega", display wants "de la Vega". diff --git a/docs/release_log.rst b/docs/release_log.rst index 5aba1ab3..82795b8d 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,9 +22,13 @@ Release Log - **Fix a credential ending the given part of a family-comma listing being read as a middle name in silence.** ``HumanName("Doe, John MA")`` gives first ``John``, last ``Doe``, suffix ``MA``, where 2.0 through 2.3 gave middle ``MA`` -- and 1.4.0 gave the suffix, so this restores v1's reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone: ``Doe, John Ma`` keeps middle ``Ma``, written the way a name is written, and ``Doe, John Ed`` keeps middle ``Ed``. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields -- a declined name re-rendered without its comma, ``John Ma Doe``, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent -- ``Doe, John MA Smith`` gives middle ``MA Smith`` and reports nothing -- while a credential run or a trailing title is transparent to it: ``Doe, John MA PhD`` gives suffix ``MA PhD`` and ``Doe, John MA Prof.`` gives title ``Prof.`` with suffix ``MA``. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further: ``Doe, John Prof. MA`` now gives title ``Prof.`` where it gave middle ``Prof. MA``, the trailing-title chain reaching a word the credential used to hide; and ``Doe, John van MA`` gives last ``van Doe`` with suffix ``MA`` where it gave middle ``van MA``, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read: ``DOE, MARY JO MA``, ``doe, john ma``, ``田中, 太郎 MA`` and ``김, 민준 MA`` all give a suffix. The unlisted dotted spelling moves with them without the parity claim -- ``Doe, John X.Y.Z.`` gives suffix ``X.Y.Z.`` where 1.4.0 and 2.3.0 both gave a middle name -- to match the comma-less ``John Doe X.Y.Z.``. One word is carved out: ``do`` is the only member of this class that is also a surname particle, so capitals decide it and the particle reading keeps every other spelling. ``Doe, John DO`` gives suffix ``DO``, while ``Doe, John do``, ``Doe, John Do``, ``DOE, JOHN DO`` and ``doe, john do`` are unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, so ``SMITH, JOHN DO`` keeps last ``DO SMITH`` as ``NASCIMENTO, EDSON ARANTES DO`` does -- right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See the ``S2`` and ``P6`` entries of ``docs/design/decisions.md`` (closes #531) - - **Fix a maiden marker's clause swallowing a trailing credential in silence.** ``HumanName("Jane Doe nee Smith MA")`` gives maiden ``Smith`` with suffix ``MA``, where 2.0 through 2.3 gave maiden ``Smith MA`` and said nothing; 1.4.0 read the ``MA`` as a suffix too. ``Doe, Jane nee Smith MA`` moves with it, and so do the one-case spellings ``JANE DOE NEE SMITH MA`` and ``jane doe nee smith ma``. The words a marker takes now end where a trailing credential begins, which is what the marker's other two stops -- a suffix word, a trailing roman numeral -- have always done. Until this release it was the last trailing position in the library where a word of the ambiguous credential class was read without a report, and it was order-sensitive besides: ``Jane Doe nee Smith MA PhD`` gave maiden ``Smith MA`` while ``Jane Doe nee Smith PhD MA`` gave maiden ``Smith``, so whether the word was read at all depended on which side of the unambiguous credential the writer put it. Both now give maiden ``Smith``, with suffix ``MA PhD`` and ``PhD MA``. The writing still decides, exactly as it does for the same word ending a name with no clause: ``Jane Doe nee Smith Ma`` keeps maiden ``Smith Ma``, and ``Jane Doe nee Yo-Yo Ma`` keeps a two-word birth surname whole. The one member of this class that is also a surname particle keeps the carve-out it has outside a clause -- ``Doe, Jane nee Smith DO`` gives suffix ``DO`` while ``Doe, Jane nee Smith do`` and ``Doe, Jane nee Smith Do`` keep maiden ``Smith do`` and ``Smith Do``, and the comma-less ``Jane Doe nee Smith do`` gives suffix ``do`` as ``John Doe do`` does. Either reading is now reported, and there is no third: a word the clause gives up reads as a post-nominal, or the clause keeps it and says so. A name word behind the credential ends its reach and stays silent -- ``Jane Doe nee MA Smith`` gives maiden ``MA Smith`` and reports nothing -- and this stop never takes the first word after the marker, whatever its writing says: ``Jane Doe nee MA`` keeps maiden ``MA`` and reports, the marker having announced a name where there would otherwise be none, and ``Jane Doe nee MA PhD`` keeps it too. That differs on purpose from what a certain post-nominal gets there, ``Jane Smith nee PhD`` and ``Jane Smith nee V`` leaving the marker standing as an ordinary word as before. Where no trailing rule reads the clause's tail nothing is decided and the clause keeps every word: ``Smith nee Jones MA, Jane`` and ``Smith, John, Jr nee Jones MA`` both keep maiden ``Jones MA``, unchanged and with no ``suffix-or-name`` report. A trailing title is not transparent here and the two spellings disagree -- ``Jane Doe nee Smith MA Prof.`` is unchanged and silent while ``Jane Doe nee Smith Prof. MA`` gives maiden ``Smith Prof.`` with suffix ``MA`` -- which is recorded as a boundary rather than fixed. The clause also keeps a word it cannot promise a credential reading for, which is where three shapes that look like they should move do not. Where the part the word would land in holds no name of its own there is nothing to read it as a credential, so ``Doe, Dr. nee Smith MA`` and ``Jane Doe, Jr nee Smith MA`` both keep maiden ``Smith MA`` and report. Where a join would swallow it first the same applies, and it is the birth name that would lose the word: ``Berg, abdul nee Jones MA`` keeps maiden ``Jones MA`` rather than reading first ``abdul MA``, and ``Berg, Jane van der nee Smith DO`` keeps maiden ``Smith DO`` rather than letting the particle chain carry the ``DO`` into last ``van der DO Berg``. Each of those reads as 2.3.0 read it. Delimiters settle the question outright and always did: ``HumanName("Jane Doe (nee Smith MA)")`` keeps the whole span as the maiden name and reports nothing, the writer having drawn the boundary, while ``Jane Doe (nee Smith) MA`` gives suffix ``MA`` for the word left outside it. One name is a restoration rather than a change: ``John Smith nee Jones R.A.I.`` gives suffix ``R.A.I.`` again, as 2.3.0 read it, this unreleased cycle having moved it into the maiden name when the unlisted-dotted reading above took the word out of the certain-suffix class. See the ``M2`` and ``S2`` entries of ``docs/design/decisions.md`` (closes #533) + - **Fix a maiden marker's clause swallowing a trailing credential in silence.** ``HumanName("Jane Doe nee Smith MA")`` gives maiden ``Smith`` with suffix ``MA``, where 2.0 through 2.3 gave maiden ``Smith MA`` and said nothing; 1.4.0 read the ``MA`` as a suffix too. ``Doe, Jane nee Smith MA`` moves with it, and so do the one-case spellings ``JANE DOE NEE SMITH MA`` and ``jane doe nee smith ma``. The words a marker takes now end where a trailing credential begins, which is what the marker's other two stops -- a suffix word, a trailing roman numeral -- have always done. Until this release it was the last trailing position in the library where a word of the ambiguous credential class was read without a report, and it was order-sensitive besides: ``Jane Doe nee Smith MA PhD`` gave maiden ``Smith MA`` while ``Jane Doe nee Smith PhD MA`` gave maiden ``Smith``, so whether the word was read at all depended on which side of the unambiguous credential the writer put it. Both now give maiden ``Smith``, with suffix ``MA PhD`` and ``PhD MA``. The writing still decides, exactly as it does for the same word ending a name with no clause: ``Jane Doe nee Smith Ma`` keeps maiden ``Smith Ma``, and ``Jane Doe nee Yo-Yo Ma`` keeps a two-word birth surname whole. The one member of this class that is also a surname particle keeps the carve-out it has outside a clause -- ``Doe, Jane nee Smith DO`` gives suffix ``DO`` while ``Doe, Jane nee Smith do`` and ``Doe, Jane nee Smith Do`` keep maiden ``Smith do`` and ``Smith Do``, and the comma-less ``Jane Doe nee Smith do`` gives suffix ``do`` as ``John Doe do`` does. Either reading is now reported, and there is no third: a word the clause gives up reads as a post-nominal, or the clause keeps it and says so. A name word behind the credential ends its reach and stays silent -- ``Jane Doe nee MA Smith`` gives maiden ``MA Smith`` and reports nothing -- and this stop never takes the first word after the marker, whatever its writing says: ``Jane Doe nee MA`` keeps maiden ``MA`` and reports, the marker having announced a name where there would otherwise be none, and ``Jane Doe nee MA PhD`` keeps it too. That differs on purpose from what a certain post-nominal gets there, ``Jane Smith nee PhD`` and ``Jane Smith nee V`` leaving the marker standing as an ordinary word as before. Where no trailing rule reads the clause's tail nothing is decided and the clause keeps every word: ``Smith nee Jones MA, Jane`` and ``Smith, John, Jr nee Jones MA`` both keep maiden ``Jones MA``, unchanged and with no ``suffix-or-name`` report. A title written behind or in front of the credential usually does not change how the credential is read; where something in or ahead of the clause could take the title it can -- see the trailing-title bullet below. The clause also keeps a word it cannot promise a credential reading for, which is where three shapes that look like they should move do not. Where the part the word would land in holds no name of its own there is nothing to read it as a credential, so ``Doe, Dr. nee Smith MA`` and ``Jane Doe, Jr nee Smith MA`` both keep maiden ``Smith MA`` and report. Where a join would swallow it first the same applies, and it is the birth name that would lose the word: ``Berg, abdul nee Jones MA`` keeps maiden ``Jones MA`` rather than reading first ``abdul MA``, and ``Berg, Jane van der nee Smith DO`` keeps maiden ``Smith DO`` rather than letting the particle chain carry the ``DO`` into last ``van der DO Berg``. Each of those reads as 2.3.0 read it. Delimiters settle the question outright and always did: ``HumanName("Jane Doe (nee Smith MA)")`` keeps the whole span as the maiden name and reports nothing, the writer having drawn the boundary, while ``Jane Doe (nee Smith) MA`` gives suffix ``MA`` for the word left outside it. One name is a restoration rather than a change: ``John Smith nee Jones R.A.I.`` gives suffix ``R.A.I.`` again, as 2.3.0 read it, this unreleased cycle having moved it into the maiden name when the unlisted-dotted reading above took the word out of the certain-suffix class. See the ``M2`` and ``S2`` entries of ``docs/design/decisions.md`` (closes #533) - - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Doe, Jane nee Puig i Soler`` and ``Jane Doe née Kowalska i Nowak`` move with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. See the ``P3`` entry of ``docs/design/decisions.md`` (closes #397) + - **Fix a trailing title after a maiden marker being read as part of the maiden name.** ``HumanName("Jane Doe nee Smith Prof.")`` gives title ``Prof.`` with maiden ``Smith``, where 2.0 through 2.3 gave maiden ``Smith Prof.``, and ``Mary Smith née Jones Prof.`` moves the same way. A credential or roman numeral in front of the title is read as it is with the title absent, so the two spellings ``Jane Doe nee Smith MA Prof.`` and ``Jane Doe nee Smith Prof. MA`` now agree -- title ``Prof.``, maiden ``Smith``, suffix ``MA`` -- where 2.3.0 kept every word in the maiden name (``Smith MA Prof.`` and ``Smith Prof. MA``), and ``Jane Doe nee Smith V Prof.`` gives suffix ``V``. Where the clause keeps the word in front of the title they still differ, since a title cannot leave without the words behind it: ``Doe nee Smith ba Prof.`` gives title ``Prof.`` with maiden ``Smith ba``, while ``Doe nee Smith Prof. ba`` keeps maiden ``Smith Prof. ba`` as 2.3.0 kept all three words in both. They can also differ where something in or ahead of the clause could take the title -- a particle, a bound given name, or a name left with no name word: ``Jane Doe nee Smith do MA Prof.`` keeps maiden ``Smith do MA`` while ``Jane Doe nee Smith do Prof. MA`` gives maiden ``Smith do``, suffix ``MA``, both with title ``Prof.``. A period behind an ordinary surname that is also a title makes it one here as it does at the end of any name: ``Jane Doe nee Smith King.`` gives title ``King.`` with maiden ``Smith``, where 2.3.0 gave maiden ``Smith King.``. The given part after a family comma reads it the same way: ``Doe, Jane nee Smith MA Prof.`` gives title ``Prof.`` too. A title straight after the marker stays the maiden name, the marker having announced one: ``Jane Doe nee King.`` keeps maiden ``King.``, and ``Jane Doe nee Prof. Dr.`` keeps maiden ``Prof.`` and gives title ``Dr.``, where 2.3.0 gave maiden ``Prof. Dr.``. Where the title itself would end the clause, the clause keeps it wherever giving it up would put it in a name part, and each of these reads as 2.3.0 read it: before a family comma (``Doe nee Smith Prof., Jane`` keeps maiden ``Smith Prof.``), where no name word would be left in front of it (``Dr. nee Jones Smith Prof.``), and behind a particle whose chain would take it (``Jane van der Berg nee Smith Prof.``). Brackets around a clause ending in a period are dropped before any of this is read, as in 2.3.0, so ``Jane Doe (nee Smith Prof.)`` gives title ``Prof.`` too. One reading moves because of the name the title is now read against: ``abdul nee Smith Dr.`` gives title ``Dr.`` with last ``abdul``, as ``abdul Dr.`` does, where 2.3.0 gave first ``abdul`` with maiden ``Smith Dr.``. See the ``M2`` entry of ``docs/design/decisions.md`` (closes #535) + + - **Fix a roman numeral ending a maiden clause landing in the first or middle name.** ``HumanName("Berg, abdul nee Smith V")`` gives first ``abdul`` with maiden ``Smith V``, where 2.2 and 2.3 gave first ``abdul V`` -- a word of the birth name joined into the current given name -- and 2.0 and 2.1 gave first ``abdul nee``. ``Doe, Jane nee Smith V, PhD`` gives maiden ``Smith V``, as 2.0 and 2.1 did, where 2.2 and 2.3 gave middle ``V``: after a family comma a lone numeral ending the given part is a suffix only when no further comma part follows it, and the clause now asks that as the given part itself does. Without the credential tail the numeral still goes to the suffix (``Doe, Jane nee Smith V`` gives maiden ``Smith``, suffix ``V``). A numeral straight after the marker can leave the marker an ordinary word, as ``Jane Smith née V`` does, and where it does a title behind the numeral no longer changes that: ``Jane Doe nee V Prof.`` gives title ``Prof.``, suffix ``V``, middle ``Doe``, last ``nee``, where 2.3.0 gave maiden ``V Prof.``. Elsewhere the clause keeps it -- ``Dr. nee V`` keeps maiden ``V`` as before -- and ``Doe, Jane nee V, PhD`` now gives maiden ``V``, suffix ``PhD``, as 2.0 and 2.1 did, where 2.2 and 2.3 gave middle ``nee V`` -- the same further-comma-part condition as ``Doe, Jane nee Smith V, PhD`` above -- with ``Doe, Jane nee V, Jr.`` moving the same way. See the ``M2`` entry of ``docs/design/decisions.md`` (#535) + + - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Doe, Jane nee Puig i Soler`` and ``Jane Doe née Kowalska i Nowak`` move with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. A delimiter the caller declares through ``Policy(extra_suffix_delimiters=...)`` is read past the way a connective is, so a link beside one reads as it would with the delimiter absent: under ``(" - ",)``, ``Smith, John, PhD née Puig Mr. - i Soler`` keeps maiden ``Puig Mr.`` and ``Smith, John, PhD - i Soler`` keeps suffix ``PhD, i Soler``, both as 2.3.0 read them, where the link had first let the delimiter pass for a name word and given maiden ``Puig Mr. i Soler`` and suffix ``PhD - i Soler``. The default policy declares no such delimiter. See the ``P3`` and ``M2`` entries of ``docs/design/decisions.md`` (closes #397, closes #538) - **Fix a connective contributing no initial even where it is joining nothing.** ``parse("Juan de y").initials()`` gives ``J. y.``, where every release gave ``J.`` while ``family_base`` said ``y`` -- two views of one parse disagreeing about one token. A connective contributes nothing where it is JOINING, and initials like any other name word where its part holds nothing else for it to join. One rule for all three groups, so ``John and Jane Smith`` gives ``J. J. S.`` where 2.0 through 2.3 gave ``J. a. J. S.`` and 1.4.0 the run-together ``J a J. S.``, ``Duke of Edinburgh`` gives ``D. E.`` where 2.0 through 2.3 gave ``D. o. E.`` and 1.4.0 ``D o E.``, and ``John & Jane`` gives ``J. J.``. The question is asked of the whole part and never of a word count, so ``Jon Dough and`` has base ``Dough and`` and keeps ``J. D.``, and ``Juan Velasquez y Garcia`` keeps ``J. V. G.``. ``HumanName.initials()`` moves with the core -- over the differential corpora the two surfaces move on the same names and give the same values, reading one mark. Two names come back into 1.4.0 parity rather than away from it: ``JUAN Y GARCIA`` and ``محمد و علي`` both give the answer 1.4.0 gave. Parsing got cheaper by the same change -- the marks come off one pass instead of two, six fewer Python frames per name on 3.11. Two limits carried over from the 2.4 facade fix above: case repair still keeps such a connective lower-case, so ``initials()`` and ``capitalize()`` disagree about it on purpose, and a name restored from a pickle or a copy, or built from keyword fields, carries no tags and takes the older reading. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #461) diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index b66f6983..d7e81bec 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -41,13 +41,14 @@ import dataclasses from collections.abc import Iterable, Sequence, Set from enum import IntEnum -from typing import assert_never +from typing import Literal, assert_never from nameparser._lexicon import _run_addresses_by_given from nameparser._pipeline._pieces import ( credential_at_the_given_slot, is_leading_title, is_suffix_piece, is_title_piece, - leading_titles, peel_trailing, peel_walk, tail_reading, + is_trailing_title_word, + Peel, leading_titles, peel_trailing, peel_walk, tail_reading, trailing_start, trailing_start_past_titles, ) from nameparser._pipeline._state import ( @@ -80,9 +81,12 @@ class TailReader(IntEnum): """Which rule reads the words the maiden walk would leave standing - at the end of this segment -- the reader the acronym fork's second - check has to ask, since a stop is only right where that reader - takes the word (rules.md#M2, #533). + at the end of this segment -- the reader every stop's release + check (`_release_reads_off`) has to ask, since a stop is only right + where that reader takes what the stop gives up (rules.md#M2, #533, + #535). It also decides whether the walk reads the trailing title + chain at all: TRAILING and GIVEN_SLOT read the end of the clause + through `tail_reading`, NONE through the peel alone. NONE is a statement and not a default: before a family comma the words are the family the comma already named, and a tail segment @@ -90,9 +94,11 @@ class TailReader(IntEnum): there and the clause keeps what it has -- a stop would hand a word to `family` rather than to `suffix`. - A CLOSED set: `_maiden_take` dispatches on it exhaustively - (`typing.assert_never`), so a fourth member is a type error at - every reader until it is given a reading. group() is the one + A CLOSED set: `_maiden_take` dispatches on it exhaustively, ONCE + (`typing.assert_never`), and every later site there branches on + the narrowed reading that dispatch binds; `_release_reads_off` + dispatches on the two readers that reach it the same way. So a + fourth member is a type error until it is given a reading. group() is the one place (structure, segment index) is mapped onto it, and tests/v2/pipeline/test_group.py's `test_the_reader_is_pinned_to_the_structure_it_is_read_from` @@ -101,8 +107,8 @@ class TailReader(IntEnum): """ NONE = 0 # FAMILY_COMMA segment 0, and every tail segment - TRAILING = 1 # the S2 peel: NO_COMMA, SUFFIX_COMMA segment 0 - GIVEN_SLOT = 2 # #531's reading: FAMILY_COMMA segment 1 + TRAILING = 1 # S2 peel + H5 chain: NO_COMMA, SUFFIX_COMMA seg 0 + GIVEN_SLOT = 2 # #531's reading + H5 chain: FAMILY_COMMA seg 1 class BoundJoin(IntEnum): @@ -122,7 +128,10 @@ class BoundJoin(IntEnum): # rules.md#S2: "a trailing word of the suffix vocabulary reads as a # suffix" -- group does not decide that; it stops before whatever # trailing_start says the run is, so the chain and the maiden walk -# end where assign's peel begins (#424). +# end where assign's peel begins (#424). The maiden walk reads the +# title-aware `tail_reading` instead wherever a trailing rule reads +# the clause (#535), so there it ends where assign's peel and title +# chain together begin. # rules.md#P2: "a particle joins the words after it into one name # part, the join running until the next particle starts a group of # its own, a trailing suffix begins" -- and on to the maiden marker @@ -138,9 +147,11 @@ def _is_prefix_piece(piece: Sequence[int], ptags: Set[str], # rules.md#M2: "a recognized maiden marker standing after at least one -# name word takes the words after it" -- up to any suffix word, or the -# trailing numeral assign reads as the suffix, as the maiden name, the -# marker itself dropped (history: decisions.md#M2) +# name word takes the words after it" -- up to any suffix word, the +# trailing numeral or credential assign reads as the suffix, or (where +# a trailing rule reads the clause) a trailing title H5's chain takes, +# as the maiden name, the marker itself dropped (history: +# decisions.md#M2) # # A marker piece is a LONE marker -- M2's own "standing as a word of # its own". The consumer runs before every join but the Ph. D. merge @@ -250,7 +261,7 @@ def _a_name_word_ahead(view: Sequence[Sequence[int]], def _join_takes_the_member(view: Sequence[Sequence[int]], view_tags: Sequence[Set[str]], tokens: Sequence[WorkToken], - at: int) -> bool: + at: int, reader: TailReader) -> bool: """Whether a join BELOW this pass would absorb the member at `at`. A released word only reads as the credential while it is still the @@ -260,31 +271,209 @@ def _join_takes_the_member(view: Sequence[Sequence[int]], now from the other side ('Berg, Jane van der nee Smith DO' read family 'van der DO Berg'). - Two joins can reach it, and each is modelled by the shape it needs - rather than by running it: P2's chain, where the member is - particle vocabulary standing behind a particle piece or beside - another released one ('Jane Doe nee Smith DO DO' read family - 'DO DO'), and P5's bound-given join, where the member is the word - after the bound one ('Berg, abdul nee Jones MA' read given - 'abdul MA'). Both over-decline rather than predict: a shape that - only MIGHT join keeps its word in the maiden name, which is the + Three joins can reach it, and each is modelled by the shape it + needs rather than by running it: + + * P2's chain, where the member is particle vocabulary standing + behind a particle piece, beside another released one ('Jane Doe + nee Smith DO DO' read family 'DO DO'), or with a title behind + it, the chain running on over a trailing title (rules.md#H5's + Accepted 'John van der Berg Prof.'; 'Jane Doe nee Smith MA do + Prof.'). + * P6's attachment, after a family comma, where the member is a + TITLE that is also particle vocabulary: P6 attaches a particle + trailing the given part to the family ('Doe, Jane nee Smith St.' + would read family 'St. Doe'). Titles only -- a credential that + is also a particle ('DO') is the given slot's own lean, which + the credential stop has already asked (#533). + * P5's bound-given join, where the member is the word after the + bound one ('Berg, abdul nee Jones MA' read given 'abdul MA'). + Asked for the GIVEN_SLOT reader alone: that join is + `BoundJoin.LENIENT` only after a family comma, and before one + the reserve is `BoundJoin.STRICT`, which reads the same peel and + chain the release was just asked of and so never joins a span + that reading covered (rules.md#P5). + + All three over-decline rather than predict: a shape that only + MIGHT join keeps its word in the maiden name, which is the conservative direction M2's invariant asks for.""" member = tokens[view[at][0]] + # P6's half (see the docstring): a released title that is also a + # particle, after a family comma. + if (reader is TailReader.GIVEN_SLOT and "particle" in member.tags + and is_title_piece(view[at], view_tags[at], tokens)): + return True if "particle" in member.tags: if at and _is_prefix_piece(view[at - 1], view_tags[at - 1], tokens): return True + # a title behind it: the chain runs on over it (docstring) for q in range(at + 1, len(view)): if len(view[q]) == 1 and "particle" in tokens[view[q][0]].tags: return True - # P5 joins the first non-title piece to the one after it, so the + if is_title_piece(view[q], view_tags[q], tokens): + return True + # rules.md#P5: "a word the peel reads as a suffix unjoined must + # read so joined, or the join declines" -- why the STRICT reserve + # needs no model here and only the GIVEN_SLOT reader asks. P5 + # joins the first non-title piece to the one after it, so the # member is at risk exactly where it IS the one after it. The tag - # pair is tested before `leading_titles` is asked, which is what - # keeps the ordinary credential release from paying that frame. - return (at > 0 and len(view[at - 1]) == 1 + # pair is tested before `leading_titles` is asked, which keeps the + # ordinary credential release from paying that frame. + return (reader is TailReader.GIVEN_SLOT + and at > 0 and len(view[at - 1]) == 1 and "vocab:bound-given" in tokens[view[at - 1][0]].tags and leading_titles(view, view_tags, tokens) == at - 1) +# rules.md#M2: "a word the clause gives up reads as a post-nominal or +# the clause keeps it" -- asked once, here, for every stop the walk +# makes: the numeral, the credential and the title (#535). A stop +# gives up the whole SPAN behind the word it stops at, so the question +# is asked of the span and not of the word: 'DOE NEE SMITH PROF. MA' +# stopped at the title and handed the MA behind it to the family, +# where the name left standing ('DOE PROF. MA') reads MA as a name. +def _release_reads_off(view: Sequence[Sequence[int]], + view_tags: Sequence[Set[str]], + tokens: Sequence[WorkToken], + start: int, end: int, at: int, + reader: Literal[TailReader.TRAILING, + TailReader.GIVEN_SLOT], + one_case: bool | None, + reading: tuple[list[int], tuple[int, ...], Peel] + | None = None, + *, tail_follows: bool) -> bool: + """Whether the name the take would leave (`view`) reads every + piece in `start`..`end` as a title or a suffix, and no join below + this pass absorbs any piece from `at` on. + + `at` is the first released piece. `start` is where coverage is + asked from, which is `at` itself except where the caller has + already asked the first piece its own question (the numeral fork + and the given-slot credential test each read their word in a way + the span test does not, and pass `at + 1`). `end` is where it + stops: `len(view)`, except for the title stop when a numeral or + credential stop stands behind it -- that stop has already had its + own span asked, in the way ITS word reads: the span check does not + ask the given slot's lenient numeral itself -- the numeral fork has + ('Doe, Jane nee Smith Prof. V' gives the V up there) -- though the + GIVEN_SLOT title chain below counts it as the first pass's suffix, + so the title in front of it is asked only of itself. + + Each reader is asked the way it READS, because the release is only + right where that reader places the span: + + * TRAILING -- the no-comma path and the part before a suffix + comma -- is assign's own reading, `tail_reading` over the view, + the S2 peel and the H5 chain to their fixed point. A piece is + covered where that reading took it as a title or peeled it as a + suffix, or where group flagged it a credential outright. + * GIVEN_SLOT -- after a family comma -- has words to spare by + construction, so the question is the writing alone: a name word + must stand ahead (`_a_name_word_ahead`), and every piece of the + span must be a suffix piece, a class member #531's reading takes + as the credential, or a word H5's chain takes as a title. + + Then the joins. A title in the span with a particle ahead of it is + taken by P2's chain, which runs on over a trailing title + (rules.md#H5's Accepted 'John van der Berg Prof.'), and every piece + of the span is asked `_join_takes_the_member`. Both decline rather + than predict, which is the conservative direction M2 asks for. + + NONE never reaches this, and the annotation says so: no trailing + rule reads those words, so no stop is made for this to check. + + `reading` is the TRAILING reader's `tail_reading` of this same + view, where the caller has one in hand (the numeral fork reads it + first), so the view is read once rather than twice. + """ + if reader is TailReader.GIVEN_SLOT: + if not _a_name_word_ahead(view, view_tags, tokens, at): + return False + # The given part's title chain is read the way assign reads it + # after a family comma: over the pieces its FIRST suffix pass + # leaves, from the end. That pass takes a class member as the + # credential only where every piece behind it is taken too, so + # a member with a title behind it is still a name word there + # and the chain stops at it -- 'Doe, Jane Dr. MA Prof.' reads + # middle 'Dr.', and a clause that gave 'Dr. MA Prof.' up put + # 'Dr.' in the middle name (#535). `chain_ok[q]` says whether + # the chain, walking from the end, is still running at `q`. + # The one suffix these passes read which `is_suffix_piece` does + # not is the lenient numeral (#144): a suffix word the initial + # veto refuses, read as a suffix only where no comma part + # follows the given one (`tail_follows`, assign's own + # two-segment condition) and only where it is the given part's + # LAST word -- the literal last piece in the first pass, and + # the last one standing once the chain has taken the titles + # behind it in the second ('Doe, Jane i V Prof.' reads suffix + # 'i V', title 'Prof.', as 'Doe, Jane nee Smith i V Prof.' must + # be able to give them up). + last = len(view) - 1 + chain_ok = [False] * len(view) + members_ok = running = True + for q in range(last, -1, -1): + piece = view[q] + if (is_suffix_piece(piece, view_tags[q], tokens) + or (q == last and not tail_follows + and len(piece) == 1 + and "vocab:suffix" in tokens[piece[0]].tags)): + chain_ok[q] = running + continue + if (members_ok and len(piece) == 1 + and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags + and credential_at_the_given_slot(tokens[piece[0]], + one_case)): + chain_ok[q] = running + continue + members_ok = False + if not is_trailing_title_word(piece, view_tags[q], tokens): + running = False + chain_ok[q] = running + titles_behind = [False] * (len(view) + 1) + titles_behind[len(view)] = True + for q in range(last, -1, -1): + titles_behind[q] = (titles_behind[q + 1] and chain_ok[q] + and is_trailing_title_word( + view[q], view_tags[q], tokens)) + for q in range(start, end): + piece = view[q] + if is_suffix_piece(piece, view_tags[q], tokens): + continue + if (not tail_follows and titles_behind[q + 1] + and len(piece) == 1 + and "vocab:suffix" in tokens[piece[0]].tags): + continue + if is_trailing_title_word(piece, view_tags[q], tokens): + if chain_ok[q]: + continue + return False + if (len(piece) == 1 + and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags + and credential_at_the_given_slot(tokens[piece[0]], + one_case)): + continue + return False + elif reader is TailReader.TRAILING: + rest, chained, peeled = reading or tail_reading( + peel_walk(leading_titles(view, view_tags, tokens), view_tags), + view, view_tags, tokens, one_case) + covered = set(chained) + covered.update(rest[peeled.names:]) + if not all(q in covered or "suffix" in view_tags[q] + for q in range(start, end)): + return False + else: + assert_never(reader) + if (any(is_title_piece(view[q], view_tags[q], tokens) + for q in range(at, len(view))) + and any(_is_prefix_piece(view[q], view_tags[q], tokens) + for q in range(at))): + return False + return not any(_join_takes_the_member(view, view_tags, tokens, q, + reader) + for q in range(at, len(view))) + + # rules.md#M2: "a link inside the birth name does not end it" -- the # one shape the walk below steps over rather than stopping at. # rules.md#P3: "A connective that is also generational vocabulary @@ -300,7 +489,8 @@ def _link_joins_inside_the_clause(k: int, lo: int, hi: int, pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - beside: list[_Beside]) -> bool: + beside: list[_Beside], + cores: Set[str]) -> bool: """Whether the suffix piece at `k` is a connective PLACED TO JOIN between two name words of the clause `lo`..`hi`. @@ -330,7 +520,7 @@ def _link_joins_inside_the_clause(k: int, lo: int, hi: int, if not _is_conj_piece(pieces[k], ptags[k], tokens): return False if not beside: - beside.append(_run_neighbours(pieces, ptags, tokens)) + beside.append(_run_neighbours(pieces, ptags, tokens, cores)) return _between_name_words(k, lo, hi, pieces, ptags, tokens, beside[0]) @@ -341,6 +531,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], one_case: bool | None, reader: TailReader, ambiguities: list[PendingAmbiguity], + *, tail_follows: bool, ) -> MaidenIndices | None: """The PIECE indices the marker pass removes, split the way MaidenIndices declares them: the MARKER's pieces (one, or several @@ -372,8 +563,26 @@ def _maiden_take(pieces: Sequence[Sequence[int]], name is left standing. Nor is one given up to a reader that will not be there to read it, or to a JOIN that runs before the reader does: rules.md#M2's invariant is that a released word ends the - parse suffix-roled, and the two view checks below are what makes - the stop conservative enough to hold it (#533 review). + parse suffix-roled, and the shared release check + (`_release_reads_off`) is what makes the stop conservative enough + to hold it (#533 review). + + And up to a trailing TITLE since #535, where the reader has a + trailing rule: the walk reads the end of the name through H5's + chain as assign does, so the title ends the clause and a + credential or numeral in front of it gets the stops it gets with + the title absent. All three stops -- the numeral, the credential + and the title -- ask one question of the whole span they give up + (`_release_reads_off`), and so does a link the walk stops at. Only + the credential and title stops spare the first word after the + marker. The numeral stop does not: it reads FROM the marker by + design, so a numeral standing straight after it is not held to + the first-word floor and may decline the clause -- 'Jane Smith née + V' (rules.md#M2) stays a marker with nothing behind it and the + name has no maiden clause at all -- while 'Dr. nee V' keeps maiden + 'V'. Those are examples, not a rule over every shape: the numeral + fork and the reader decide it, and 'Doe, J. nee V' keeps maiden + 'V' though 'Doe, J. V' reads suffix 'V'. A tail segment's delimiter cores (`cores`, empty elsewhere) are structure, not words, and group() drops them after the pass -- @@ -386,6 +595,8 @@ def _maiden_take(pieces: Sequence[Sequence[int]], the segment as written, which is why this returns indices rather than a slice. """ + # the lone-core test; also in _run_neighbours and group()'s #206 + # drop, whose copy adds `len(pieces) > 1` -- keep in step seen = [k for k in range(len(pieces)) if not (len(pieces[k]) == 1 and tokens[pieces[k][0]].text in cores)] @@ -431,7 +642,42 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # under every one of the six policies. skip = frozenset(range(len(pieces))) - frozenset(seen) rest = peel_walk(seen[m], ptags, skip) - peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) + # #535: where a trailing rule reads these words, the walk reads the + # end of the name as that rule does -- the S2 peel and the H5 title + # chain to their fixed point (`tail_reading`), so a title behind + # the clause no longer hides the credential or numeral in front of + # it, and the title is itself a stop (below). `rest` comes back + # with the chained pieces SPLICED OUT, which is what keeps + # `rest[-1]` the numeral and `rest[peeled.names]` the first piece + # the peel took. Where no rule reads them (NONE) there is no chain + # to consult and the peel alone stands, as before. + # THE reader dispatch: exhaustive here, once, so a fourth + # `TailReader` member is a type error at this line; every later + # site branches on `reads`, which is the reader where a trailing + # rule reads these words and None where none does. + reads: Literal[TailReader.TRAILING, TailReader.GIVEN_SLOT] | None + # the walk as written, before any chain splices it: the link check + # below re-asks the link exception with its pre-#535 bound + written = rest + if reader is TailReader.NONE: + reads = None + chained: tuple[int, ...] = () + peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) + elif reader is TailReader.TRAILING or reader is TailReader.GIVEN_SLOT: + reads = reader + # the FIRST-WORD FLOOR, as the chain's own: `rest` opens with + # the marker run, and the word after it stays the maiden name + # whatever it is, so the chain may not take it -- taken, it + # left the count the re-peel reads, and 'Jane Doe nee King. ba' + # kept 'ba' in the clause where 'Jane Doe nee Smith ba' gives + # it up (#535 review). A marker run with nothing after it has + # no floor to set; the walk below declines it anyway. + first = seen[m + run] if m + run < len(seen) else None + floor = rest.index(first) + 1 if first in rest else 1 + rest, chained, peeled = tail_reading(rest, pieces, ptags, tokens, + one_case, floor) + else: + assert_never(reader) trailing = rest[-1] if peeled.numeral is not None else len(pieces) # The fork reads the piece before the numeral, and the take # REMOVES that piece: afterwards assign sees the piece before the @@ -448,17 +694,50 @@ def _maiden_take(pieces: Sequence[Sequence[int]], left = [i for i in seen if i < seen[m] or i >= trailing] view = [pieces[i] for i in left] view_tags = [ptags[i] for i in left] - # The same peel pair as above, read over the view -- and only - # `Peel.numeral` off it, because the bare-acronym fork COUNTS - # pieces and this view no longer holds the pieces it counted + # The same peel pair as above, read over the view -- and, for + # the NONE reader, only `Peel.numeral` off it (every other + # reader reads the view through `tail_reading` and asks the + # shared release check below), because the bare-acronym fork + # COUNTS pieces and this view no longer holds the pieces it + # counted # ('John née Jones Smith Ma' peeled over the pieces as written # reads the acronym as a credential with words to spare, and # once 'Jones Smith' has left it is the family of what # remains). The acronym fork builds a view of its own below. view_rest = peel_walk(leading_titles(view, view_tags, tokens), view_tags) - if peel_trailing(view_rest, view, view_tags, tokens, - one_case).numeral is None: + if reads is None: + kept = peel_trailing(view_rest, view, view_tags, tokens, + one_case).numeral is None + else: + # the view read through the chain too, and the span behind + # the numeral asked the shared release question -- which + # includes the joins, the half this fork never asked: + # 'Berg, abdul nee Smith V' handed the V to the bound-given + # join and read given 'abdul V' (2.2 and 2.3; 2.0 and 2.1 + # read given 'abdul nee', suffix 'V' -- #411's own reserve + # differs there too, before #535 ever runs) + # + # After a family comma the given slot reads a lone numeral + # as a suffix only where the given part is the LAST comma + # part (#144, and assign's own two-segment condition): with + # a credential tail behind it the numeral is a middle + # initial there, so a clause may not give it up -- + # 'Doe, Jane nee Smith V, PhD' read middle 'V' (#535 + # review). + view_reading = tail_reading(view_rest, view, view_tags, + tokens, one_case) + view_peel = view_reading[2] + at = left.index(trailing) + # `tail_follows` is only ever true for the GIVEN_SLOT + # reader: group() computes it from that reader. + kept = (view_peel.numeral is None + or tail_follows + or not _release_reads_off( + view, view_tags, tokens, at + 1, len(view), at, + reads, one_case, view_reading, + tail_follows=tail_follows)) + if kept: trailing = len(pieces) # #533: the ACRONYM fork, asked the way the numeral is -- the peel # over the pieces as they stand, then again over the name the take @@ -469,7 +748,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # peel, one view built once per take: O(pieces) for the take, not # per member, and no re-entrancy -- the predicate never calls the # walk that calls it. - if (reader is not TailReader.NONE and peeled.names < len(rest) + if (reads is not None and peeled.names < len(rest) and m + run + 1 < len(seen)): # THE FIRST-WORD FLOOR: the stop never takes the FIRST word # after the marker -- a class member standing alone there @@ -510,9 +789,9 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # reason `_assign.previous_kept` is: an inert branch is cheaper # than a question asked of the wrong shape, and the three # sibling sites (`_pieces.segment_suffix_reading`, the - # GIVEN_SLOT branch below, and the emitter at the end of this - # function) each pair a length test with a tag test the same - # way. + # GIVEN_SLOT branch of `_release_reads_off` above, and the + # emitter at the end of this function) each pair a length + # test with a tag test the same way. # # The tag is the CLASS the rule is stated in terms of, and it # is not redundant with the walk the way the length test is: @@ -528,14 +807,18 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # point, and a row for one name would have to be guessed for # the four interpreters only CI runs. # - # A THIRD condition stood here and is gone: `stop < trailing` - # guarded nothing, `stop` being the larger of a walked piece - # and the floor and both bounded by `trailing`, so at worst - # `stop == trailing` and the assignment below sets `trailing` - # to what it already is. Measured over 1,760,904 parses: - # `stop > trailing` never once, `stop == trailing` 15,696 - # times, and the tag test declined every one of those. - if (len(head) == 1 + # And `stop < trailing`, DEFENSIVE: a stop at or past the + # numeral stop would move the clause's end the wrong way. The + # walked piece is bounded by `trailing`, but the FLOOR is an + # index, and the title chain's splice can leave the numeral + # stop in front of it ('Jane Doe nee V Prof.' reaches `stop > + # trailing` above, though its head is no class member and the + # tag test below declines it anyway). Measured 2026-09-26: no + # input on the maiden grids has a class-member head with `stop + # >= trailing`, and dropping the guard changes no reading there -- + # kept because a caller's vocabulary can list a title as a + # class member. + if (stop < trailing and len(head) == 1 and AMBIGUOUS_ACRONYM_TAG in tokens[head[0]].tags): left = [i for i in seen if i < seen[m] or i >= stop] view = [pieces[i] for i in left] @@ -548,7 +831,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # member itself reads as the family name ('JOHN NEE JONES # SMITH MA PHD' left 'JOHN MA PHD', whose MA is the family). at = left.index(stop) - if reader is TailReader.GIVEN_SLOT: + if reads is TailReader.GIVEN_SLOT: # after a family comma the words to spare are there by # construction, so the reader is #531's -- the member's # own reading, asked through the one predicate that @@ -560,33 +843,51 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # rather than handing one of them to the current # name's middle -- which is what the clause-less # 'Doe, Jane MA do' does with them, middle 'MA' and - # family 'do Doe'). And a slot the take would - # leave nobody to read is no slot: `_a_name_word_ahead` - # is that half, asked first because it is the cheaper - # question and because with no name word ahead the - # answer below is about a name that would not exist. - takes = _a_name_word_ahead(view, view_tags, tokens, at) and all( - is_suffix_piece(view[q], view_tags[q], tokens) - or (len(view[q]) == 1 - and AMBIGUOUS_ACRONYM_TAG in tokens[view[q][0]].tags - and credential_at_the_given_slot( - tokens[view[q][0]], one_case)) - for q in range(at, len(view))) - elif reader is TailReader.TRAILING: - takes = trailing_start( - leading_titles(view, view_tags, tokens), - view, view_tags, tokens, - one_case=one_case) <= at + # family 'do Doe'). The member is asked here, the span + # behind it and the name word ahead of it by the + # shared check below. + takes = credential_at_the_given_slot(tokens[head[0]], + one_case) + start = at + 1 else: - assert_never(reader) + # TRAILING: the peel over the view IS the member's + # question here, so the shared check asks it from the + # member itself: covered means the reading of the name + # left standing took this piece as the credential, so + # one `tail_reading` answers both questions. + takes = True + start = at # rules.md#M2: "a word the clause gives up reads as a - # post-nominal or the clause keeps it". Both readers above - # ask what a TRAILING rule makes of the member, and a join - # below this pass runs first and can take the word out of - # that rule's reach entirely, so the release is withdrawn - # where one would (#533 review). - if takes and not _join_takes_the_member( - view, view_tags, tokens, at): + # post-nominal or the clause keeps it" -- of the whole span + # the stop gives up, and with the joins below this pass + # asked too (#533 review, #535). + if takes and _release_reads_off(view, view_tags, tokens, + start, len(view), at, reads, + one_case, + tail_follows=tail_follows): + trailing = stop + # rules.md#M2 (#535): a trailing title the H5 chain takes ends the + # clause too -- asked as the other two stops are, over the name the + # take would leave (`_release_reads_off`), because the chain needs + # a name word to stand behind and 'Dr. nee Jones Smith Prof.' + # leaves 'Dr. Prof.', whose Prof. would be the family name. The + # FIRST-WORD FLOOR is the chain's own (`floor` above): a title + # standing straight after the marker stays the maiden name, the + # marker having announced one ('Jane Doe nee King.'), and where + # titles follow it only the ones behind it go ('Jane Doe nee Prof. + # Dr.' reads maiden 'Prof.', title 'Dr.'). + if chained and reads is not None: + stop = min(chained) + if stop < trailing: + left = [i for i in seen if i < seen[m] or i >= stop] + view = [pieces[i] for i in left] + view_tags = [ptags[i] for i in left] + at = left.index(stop) + end = (left.index(trailing) if trailing < len(pieces) + else len(view)) + if _release_reads_off(view, view_tags, tokens, at, end, at, + reads, one_case, + tail_follows=tail_follows): trailing = stop # The walk starts past the WHOLE marker: a phrase's second word is # the marker, not the first word it takes. With nothing behind the @@ -613,39 +914,37 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # ' - ', which is the pair # test_a_core_between_the_marker_and_the_first_word_is_below_lo # holds (all four readings measured 2026-09-22). - # NOT theoretical and not a whole claim about cores, both settled - # by measurement 2026-09-21 over corpus u cases.py u the property - # grids u a 50,925-name generated set with cores, under thirteen - # core-bearing policies: 25,536 of 596,392 maiden takes had a core - # standing exactly there, so the bound is load-bearing -- and PAST - # `lo` a core is no longer below it, is an ordinary index to - # `_run_neighbours` (which steps over connectives and nothing - # else), and DOES pass for the name word on a link's side. That - # reading is pinned as it stands rather than repaired here - # (test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538): - # `_between_name_words` is asked about a core in 51,072 of 900,023 - # calls, the answer differs from a core-skipping reading in 8,094 - # parses over 1,278 texts, and 1,824 of those move `maiden` on 288 - # texts -- none of them reachable at the default policy, which is - # why rules.md#M2 states it with a policy annotation beside the - # marker. The repair is `cores` threaded through - # three call sites into `_run_neighbours`, which is its own change - # (#538, and rules.md#M2 carries it as a `deviates:` example). - # `peel_start` is where assign's trailing run begins over - # the pieces as WRITTEN, so the generation or credential a clause - # ends with is never the name word on a link's right ('... nee Puig - # i III', '... i MA', whose MA carries no `vocab:suffix` tag for - # the piece test to refuse it by). + # NOT theoretical, settled by measurement 2026-09-21 over corpus u + # cases.py u the property grids u a 50,925-name generated set with + # cores, under thirteen core-bearing policies: 25,536 of 596,392 + # maiden takes had a core standing exactly there, so the bound is + # load-bearing. PAST `lo` the bound no longer reaches a core, and + # `_run_neighbours` steps over it as it steps over a connective + # (#538): a core is structure, so the word on a link's side is the + # one past it, and the clause reads as the same text written + # without the core ('Smith, John, PhD née Puig Mr. - i Soler' ends + # at the link after the title, as '... Puig Mr. i Soler' does -- + # the title is the word on the link's left and refuses). + # `peel_start` is where assign's trailing run begins -- over the + # pieces as written for a NONE reader (`trailing_start`'s whole + # answer, read off the peel pair above rather than re-running it), + # and for any other reader `tail_reading`'s title-aware answer + # (`rest[peeled.names]`, `rest` having been read through the H5 + # chain, so a chained title is not in it). Either way the + # generation or credential a clause ends with is never the name + # word on a link's right ('... nee Puig i III', '... i MA', whose + # MA carries no `vocab:suffix` tag for the piece test to refuse it + # by). # - # `peel_start` is `trailing_start`'s whole answer, read off the - # peel pair above rather than re-running it, which is the reading - # that function's own docstring sends this caller here for. It is - # never past the walk's own stop, `trailing`: the numeral fork's - # `trailing` is the walk's LAST piece and the acronym fork's `stop` - # is a max over this one, so the two never disagree about where the - # clause ends, only about what the exception may reach across. - # Measured 2026-09-20 with a probe here over the whole suite -- - # 93,408 reaches of this site, `peel_start > trailing` 0 of them. + # It can stand PAST `trailing`: the title stop sets `trailing` to + # a piece the chain read past ('Jane Doe nee Smith Prof.' reaches + # here with `peel_start` 5 and `trailing` 4). That is harmless: the + # walk below never reaches `trailing`, so the link exception is + # never asked about a piece at or past it, and a bound past the + # walk's own end cannot let a link join across anything the walk + # visits. A dated snapshot from before #535, measured 2026-09-20 + # with a probe here over the whole suite: 93,408 reaches of this + # site, `peel_start > trailing` 0 of them. lo = seen[m + run] peel_start = (rest[peeled.names] if peeled.names < len(rest) else len(pieces)) @@ -653,13 +952,52 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # The link exception's memo cell, filled inside the predicate on # the first connective it is asked about (see its docstring). beside: list[_Beside] = [] - while (j < len(seen) and seen[j] < trailing - and (not is_suffix_piece(pieces[seen[j]], ptags[seen[j]], - tokens) - or _link_joins_inside_the_clause(seen[j], lo, peel_start, + while j < len(seen) and seen[j] < trailing: + k = seen[j] + if (not is_suffix_piece(pieces[k], ptags[k], tokens) + or _link_joins_inside_the_clause(k, lo, peel_start, pieces, ptags, tokens, - beside))): - j += 1 + beside, cores)): + j += 1 + continue + # rules.md#M2: "a word the clause gives up reads as a + # post-nominal or the clause keeps it" -- asked here of a LINK + # the exception refused only BECAUSE the title chain moved its + # bound, and of no other stop. The exception reads the end of + # the name through the chain where a trailing rule reads the + # clause, so a link can now stop the walk where, bounded by the + # peel over the words as written, it joined ('Jane Doe nee + # Smith i DO Prof.': the DO is the peel's once the title is + # chained), and that new stop gives up the words behind it + # too. Where the name left standing would not read that run as + # post-nominals and titles, the clause keeps the link as it + # kept it before -- otherwise 'Doe i' became a middle name + # (#535). A link the as-written bound refuses too stops as it + # always did ('Doe, Jane nee Smith i V' gives suffix 'i V'), + # and a suffix word that is no link is #548's question. So is + # a link standing FIRST after the marker: stopping there + # declines the clause outright, which gives nothing up. The + # as-written peel is read only here, on a refused link. + if (j > m + run and reads is not None + and _is_conj_piece(pieces[k], ptags[k], tokens)): + as_written = peel_trailing(written, pieces, ptags, tokens, + one_case) + written_start = (written[as_written.names] + if as_written.names < len(written) + else len(pieces)) + if _link_joins_inside_the_clause(k, lo, written_start, + pieces, ptags, tokens, + beside, cores): + left = [i for i in seen if i < seen[m] or i >= k] + view = [pieces[i] for i in left] + view_tags = [ptags[i] for i in left] + at = left.index(k) + if not _release_reads_off(view, view_tags, tokens, at, + len(view), at, reads, one_case, + tail_follows=tail_follows): + j += 1 + continue + break # j == m + run means nothing followed the marker but a suffix, so # the pass declines and the marker stays ordinary words # (rules.md#M2). @@ -698,7 +1036,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # the 905,796-parse oracle at the same frame counts. last = pieces[seen[j - 1]] word = tokens[last[0]] - if (reader is not TailReader.NONE + if (reads is not None and not word.tags.isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS)): ambiguities.append(PendingAmbiguity( AmbiguityKind.SUFFIX_OR_NAME, @@ -720,8 +1058,17 @@ def _maiden_take(pieces: Sequence[Sequence[int]], def _run_neighbours(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], - tokens: Sequence[WorkToken]) -> _Beside: - """The nearest non-connective piece on each side of every index. + tokens: Sequence[WorkToken], + cores: Set[str]) -> _Beside: + """The nearest piece that is neither a connective nor a delimiter + core, on each side of every index. + + A tail segment's delimiter CORE (`cores`, empty off a tail segment + and at the default policy) is stepped over exactly as a connective + is: it is structure the caller declared, the #206 drop takes a LONE + core out of the output, and a link read with it present must read as the + same text read without it (rules.md#M2, #538). Asked inline, beside + `_is_conj_piece`, so an empty set costs no frame. EVERY MEMBER OF ONE RUN HAS THE SAME ANSWER, which is the whole of the fix: `_between_name_words` used to walk the run itself, so a @@ -780,7 +1127,11 @@ def _run_neighbours(pieces: Sequence[Sequence[int]], prev = -1 for k in range(n): left[k] = prev - is_conj = _is_conj_piece(pieces[k], ptags[k], tokens) + # the lone-core test; also in _maiden_take and group()'s #206 + # drop, whose copy adds `len(pieces) > 1` -- keep in step + is_conj = (_is_conj_piece(pieces[k], ptags[k], tokens) + or (len(pieces[k]) == 1 + and tokens[pieces[k][0]].text in cores)) conj[k] = is_conj if not is_conj: prev = k @@ -838,6 +1189,8 @@ def _is_rootname(piece: Sequence[int], ptags: Set[str], # is the CALLER's half: `_group_segment`'s `frozen` set is where a # connective this refuses is placed as the generation, and # `_link_joins_inside_the_clause` is the maiden walk's. +# rules.md#P3: "the search reads past it as it reads past a +# connective" (#538) -- `cores` in `beside`, below. # `Sequence[Sequence[int]]` rather than `Sequence[Piece]`, widened # when the maiden walk became a second caller: this reads a piece and # never edits one, and `_maiden_take` holds its pieces at the wider @@ -876,9 +1229,11 @@ def _between_name_words(k: int, lo: int, hi: int, one ('Carod i y Rovira'), so the word this rule is about is the first one past the run -- and where the run runs out ('Juan i e') there is no name word on that side at all. `beside` is where that - stepping already happened: `_run_neighbours` walked every run once - for the whole segment, so this reads an index rather than walking - to it. A SENTINEL OUT OF RANGE is how "the run ran out" arrives -- + stepping already happened, a lone delimiter core stepped over the + same way too (#538, rules.md#P3's separator sentence): + `_run_neighbours` walked every run once for the whole segment, so + this reads an index rather than walking to it. A SENTINEL OUT OF + RANGE is how "the run ran out" arrives -- -1 on the left, `len(pieces)` on the right -- and each bound test below turns it into False, exactly as the walk did when it ran here and stopped at the same place. `lo` is never negative and @@ -930,6 +1285,7 @@ def _group_segment(seg: tuple[int, ...], additional: int, one_case: bool | None, reader: TailReader, maiden_ambiguities: list[PendingAmbiguity], + tail_follows: bool = False, ) -> tuple[list[Piece], list[set[str]], MaidenTake | None]: pieces: list[Piece] = [[i] for i in seg] ptags: list[set[str]] = [set() for _ in seg] @@ -1085,7 +1441,8 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), # returns what it took, and group() records the drop and the roles. taken: MaidenTake | None = None take = _maiden_take(pieces, ptags, tokens, cores, one_case, - reader, maiden_ambiguities) + reader, maiden_ambiguities, + tail_follows=tail_follows) if take is not None: marker_ks, maiden_ks = take taken = ([i for k in marker_ks for i in pieces[k]], @@ -1184,8 +1541,11 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), # of connectives has the same nearest name word on # each side, and asking per member walked the run once # per member -- quadratic in its length, 3.8x per - # doubling measured at `b9ed1429`. - beside = _run_neighbours(pieces, ptags, tokens) + # doubling measured at `b9ed1429`. `cores` is passed + # deliberately here too: a link beside a declared + # delimiter reads as it would with the delimiter absent + # (#538). + beside = _run_neighbours(pieces, ptags, tokens, cores) if not _between_name_words(k, lo, hi, pieces, ptags, tokens, beside): frozen.add(piece[0]) @@ -1678,7 +2038,9 @@ def group(state: ParseState) -> ParseState: opens_the_name=(seg_idx == 0 and not family_comma), one_case=state.one_case, reader=reader, - maiden_ambiguities=ambiguities) + maiden_ambiguities=ambiguities, + tail_follows=(reader is TailReader.GIVEN_SLOT + and len(state.segments) > 2)) # the marker is dropped and the maiden name's tokens become # MAIDEN (#274); which pieces those are was settled in # _group_segment, before the joins @@ -1721,6 +2083,10 @@ def group(state: ParseState) -> ParseState: if seg_cores: kept: list[int] = [] for k in range(len(pieces)): + # the lone-core test; also in _maiden_take and + # _run_neighbours -- keep in step. This copy alone adds + # `len(pieces) > 1`: a segment that is nothing but its + # core keeps it is_core = (len(pieces[k]) == 1 and tokens[pieces[k][0]].text in seg_cores and len(pieces) > 1) diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index 1c9c2403..6af347a9 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -410,7 +410,8 @@ def trailing_start(start: int, pieces: Sequence[Sequence[int]], """Where assign's trailing suffix run begins, read over the pieces as they stand from `start`: the index of the first piece the S2 peel takes, or len(pieces) when it takes none (#424). What P2's - chain and M2's walk stop before -- each had asked "is this a + chain stops before, and what M2's walk stops before where no + trailing rule reads the clause -- each had asked "is this a suffix?" with the suffix-piece test, which vetoes a bare 'V' as an initial (the #401 question), and so took a trailing numeral, or a bare acronym with words to spare, into the family or the @@ -421,7 +422,12 @@ def trailing_start(start: int, pieces: Sequence[Sequence[int]], leave, where the acronym fork's piece COUNT no longer describes the name -- calls the `peel_walk` + `peel_trailing` pair this wraps and reads the half it wants (#533). A `numeral_only` flag - lived here for that one caller and cost it a frame.""" + lived here for that one caller and cost it a frame. The maiden + walk asks that pair only for the NONE reader; every other reader + reads the same pieces through `tail_reading`, which runs this + pair's peel and H5's title chain to a fixed point. So this + function's answer is the whole answer only where no trailing + title chain also reads the pieces.""" rest = peel_walk(start, ptags, skip) peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) return rest[peeled.names] if peeled.names < len(rest) else len(pieces) @@ -446,9 +452,12 @@ def trailing_start(start: int, pieces: Sequence[Sequence[int]], # and the frame argument holds transitively because both of them ask # inline -- `AMBIGUOUS_ACRONYM_TAG in tok.tags` after a # `len(piece) == 1` at assign's given-part trailing slot, and the -# same pair inside the `all(...)` of `_group.py`'s `_maiden_take` -# view check. So no non-member piece reaches this function down that -# route either. +# same pair in `_group.py`'s `_release_reads_off` (its GIVEN_SLOT +# branch) and in the acronym fork of `_maiden_take` (#535 review +# folded the view check's own `all(...)` into the shared release +# check, but the pre-check pair travelled with it rather than moving +# into this function). So no non-member piece reaches this function +# down that route either. def listed_lean(token: WorkToken, one_case: bool | None) -> Lean | None: """`ambiguous_lean` for a LISTED bare-ambiguous token, or None if the token is not tagged a listed member, is admitted by SHAPE @@ -471,13 +480,15 @@ def credential_at_the_given_slot(token: WorkToken, particle vocabulary reads as the credential on a POSITIVE lean alone, P6's attachment keeping every other spelling. - Two callers since #533 -- assign's walk over the given part, and - the maiden walk's second check over the name the take would leave - (rules.md#M2) -- so the reading is a function rather than a - condition written twice - (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It is a - text-and-tags question, which is what puts it in this module - rather than beside either caller. + Called from assign's walk over the given part, from + `_release_reads_off`'s GIVEN_SLOT branch (the shared release check + every maiden-walk stop asks, rules.md#M2), and from the acronym + fork of `_maiden_take`, asked there of the member itself rather + than the span around it. One function rather than a condition + written at each + (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It is a text-and-tags + question, which is what puts it in this module rather than beside + any caller. The membership half of that contract is CHECKED rather than trusted, because getting it wrong is silent: all three of @@ -597,7 +608,8 @@ def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]], # WORD is while disagreeing, deliberately, about what a title SHAPE is. def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], - tokens: Sequence[WorkToken]) -> int: + tokens: Sequence[WorkToken], + floor: int = 1) -> int: """How many pieces of `rest` the trailing title chain LEAVES standing: `rest[:kept]` are the name pieces and `rest[kept:]` the period-marked title words the chain took, in piece order. Counted @@ -607,9 +619,13 @@ def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], the segment's pieces that the segment's own suffix reading does not claim, and in `tail_reading` the leftovers of whichever peel is current. - Floor: one name piece stands, so a name is never all title -- and - an empty `rest` returns 0, which is what leaves assign's - bare-suffix carve-out reached exactly as before. + Floor: `floor` leading positions of `rest` are never taken -- + 1 by default, so one name piece stands and a name is never all + title; the maiden walk passes the position just past the marker's + first word, so the chain never takes that word (rules.md#M2's + first-word floor, `tail_reading` passes it through). An empty + `rest` returns 0, which is what leaves assign's bare-suffix + carve-out reached exactly as before. ONE WORD per piece, the same gate the leading peel's give-back uses: a joined unit is not the shape this reads, and the tokens of @@ -631,7 +647,7 @@ def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], stops (decisions.md#parse-cost). """ k = len(rest) - while k > 1: + while k > floor: idx = rest[k - 1] piece = pieces[idx] # no #323 veto on the shape here, unlike is_leading_title's: @@ -647,6 +663,24 @@ def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], return k +# rules.md#H5: "successive single words that wear the abbreviation +# shape and are title vocabulary chain into the title from the end" +# -- one word's half of that, for a caller asking it of a piece the +# chain did not walk to: the maiden walk's release check, after a +# family comma (#535). `trailing_titles` asks the same three tests +# INLINE, in the same order, for the frame budget its docstring +# states; keep the two in step. +def is_trailing_title_word(piece: Sequence[int], ptags: Set[str], + tokens: Sequence[WorkToken]) -> bool: + """Whether one piece is a word H5's trailing chain takes: a lone + token wearing the abbreviation shape that the title vocabulary + lists. Position is the caller's question -- this answers only + what the word is.""" + return (len(piece) == 1 + and _PERIOD_ABBREV.match(tokens[piece[0]].text) is not None + and is_title_piece(piece, ptags, tokens)) + + # rules.md#H5: "the title is TRANSPARENT to the suffix reading: where # two or more name words stand, what stands once the chain is taken # reads exactly as it would read written without the title, plus the @@ -655,6 +689,7 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], one_case: bool | None, + floor: int = 1, ) -> tuple[list[int], tuple[int, ...], Peel]: """The S2 peel and the H5 chain read together to a FIXED POINT: peel, chain, splice the chained pieces out, peel again over what @@ -681,23 +716,33 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], acronym and re-exposed the first title, reading family 'Prof.' with suffix 'MA' where 'John Prof. MA' reads family 'MA'. - One function for two readers -- assign's placement and group's - bound-given reserve (P5), which must count the name words assign - will leave. Deriving that agreement twice is what left the two + One function for assign's placement, group's bound-given reserve + (P5), which must count the name words assign will leave, and the + maiden walk and its release check (rules.md#M2), which must end + the clause where assign's reading of the name will begin. Deriving that agreement twice is what left the two disagreeing at S2's bare-ambiguous reserve: 'abdul rahman MA' declined the join and 'abdul rahman MA Prof.' took it (mechanisms.md#ONE-PREDICATE-PER-QUESTION). `rest` is a peel_walk list and is not mutated -- the splice rebinds this local -- so a caller's own reference still names the - walk it built. Both callers read the one returned here instead, + walk it built. Every caller reads the one returned here instead, which is the one the final peel partitions. + + `floor` is `trailing_titles`' own: how many leading positions of + `rest` the chain may never take. 1 everywhere but the maiden walk, + whose `rest` opens with the marker and whose FIRST word after it + stays the maiden name whatever it is (rules.md#M2) -- a chain that + took that word spliced it out of the count the re-peel reads, and + 'Jane Doe nee King. ba' lost the suffix 'Jane Doe nee Smith ba' + keeps. The splice only ever removes positions at or + past the floor, so the floor names the same pieces every pass. """ titled: list[int] = [] while True: peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) kept = trailing_titles(rest[:peeled.names], pieces, ptags, - tokens) + tokens, floor) if kept == peeled.names: return rest, tuple(titled), peeled # the chain's pieces reach this list back to front, so each diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 7a900499..16cc3657 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -3713,34 +3713,311 @@ def _check_cjk_shape_purity(self) -> None: shape=1), Case("the_peel_never_reaches_a_title_behind_the_member", "Jane Doe nee Smith MA Prof.", - {"given": "Jane", "family": "Doe", "maiden": "Smith MA Prof."}, - classification="fix(#274)", - ambiguities=(), - notes="the H5 BOUNDARY, recorded rather than fixed: the " - "maiden walk has always read peel_trailing alone, and " - "the trailing-title chain lives in tail_reading, which " - "trailing_start does not run -- so the peel breaks at " - "'Prof.' and never reaches the member behind it. This " - "change inherits that boundary rather than creating " - "it, and a follow-up carries the question of whether " - "the walk should move onto tail_reading (both forks, " - "numeral included). Silent, because nothing was asked", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "MA", "maiden": "Smith"}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="the H5 boundary #533 recorded, now closed: the walk " + "reads the end of the name through the trailing-title " + "chain as assign does (tail_reading), so 'Prof.' is a " + "stop and no longer hides the member in front of it, " + "which the credential stop then gives up as it does in " + "'Jane Doe nee Smith MA'. The report is assign's peel " + "of 'MA'. 2.3.0 read maiden 'Smith MA Prof.'", shape=1), Case("a_title_in_front_of_the_member_is_the_other_spelling", "Jane Doe nee Smith Prof. MA", - {"given": "Jane", "family": "Doe", "suffix": "MA", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "MA", "maiden": "Smith"}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="the pair is the finding, and it now AGREES: with the " + "title read as the chain reads it, 'Prof. MA' and " + "'MA Prof.' give one answer, as they do at the given " + "part's trailing slot (#531's 'Doe, John MA Prof.' / " + "'Doe, John Prof. MA') and in 'Jane Doe MA Prof.'. " + "2.3.0 read maiden 'Smith Prof. MA', no suffix", + shape=1), + Case("a_trailing_title_ends_the_maiden_clause", + "Jane Doe nee Smith Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "maiden": "Smith"}, + classification="fix(#535)", + notes="rules.md#M2's title stop: a period-marked word the " + "trailing title chain takes (H5) ends the clause, as " + "'Jane Doe Prof.' reads the title. 2.3.0 read maiden " + "'Smith Prof.'", + shape=1), + Case("a_title_behind_the_numeral_hides_it_no_longer", + "Jane Doe nee Smith V Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "V", "maiden": "Smith"}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="#424's numeral stop through the title: reads as " + "'Jane Doe nee Smith Prof. V' and as 'Jane Doe nee " + "Smith V' plus the title. 2.3.0 read maiden 'Smith V " + "Prof.'", + shape=1), + Case("a_declined_member_before_a_title_stays_and_reports", + "Jane Doe nee Smith Ma Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "maiden": "Smith Ma"}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="the writing declines 'Ma' as it does in 'Jane Doe nee " + "Smith Ma', and the clause now ENDS on it, so M2's " + "report of the member it keeps fires. 2.3.0 read maiden " + "'Smith Ma Prof.' and reported nothing", + shape=1), + Case("a_trailing_title_after_a_family_comma_clause", + "Doe, Jane nee Smith MA Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "MA", "maiden": "Smith"}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="the given part's own slot reads the chain too " + "('Doe, Jane MA Prof.' gives title and suffix). 2.3.0 " + "read maiden 'Smith MA Prof.'", + shape=2), + Case("the_first_word_after_the_marker_stays_even_as_a_title", + "Jane Doe nee King.", + {"given": "Jane", "family": "Doe", "maiden": "King."}, + notes="rules.md#M2's first-word floor covers the title stop: " + "the marker announced a name, and 'King.' is a borne " + "surname the title vocabulary also lists. Unchanged " + "from 2.3.0"), + Case("the_first_word_floor_holds_for_titles_too", + "Jane Doe nee Prof. Dr.", + {"title": "Dr.", "given": "Jane", "family": "Doe", + "maiden": "Prof."}, + classification="fix(#535)", + notes="the floor is the title chain's own: the chain may not " + "take the first word after the marker, so it takes 'Dr.' " + "and stops at 'Prof.', which stays the maiden name, as " + "'Doe, J. nee MA ba' keeps 'MA' and gives up 'ba'. 2.3.0 " + "read maiden 'Prof. Dr.'"), + Case("a_title_stop_that_leaves_no_name_word_is_no_stop", + "Dr. nee Jones Smith Prof.", + {"title": "Dr.", "maiden": "Jones Smith Prof."}, + notes="the view check: the take would leave 'Dr. Prof.', " + "where H5's chain has no name word to stand behind and " + "'Prof.' would be the family name, so the clause keeps " + "it. Unchanged from 2.3.0"), + Case("no_title_stop_before_a_family_comma", + "Doe nee Smith Prof., Jane", + {"given": "Jane", "family": "Doe", "maiden": "Smith Prof."}, + notes="before a family comma no trailing rule reads these " + "words ('Doe Prof., Jane' keeps family 'Doe Prof.'), so " + "there is no chain to consult and the clause keeps the " + "title. Unchanged from 2.3.0"), + Case("a_particle_ahead_would_chain_the_released_title", + "Jane van der Berg nee Smith Prof.", + {"given": "Jane", "family": "van der Berg", "maiden": "Smith Prof."}, + notes="P2's chain runs on over a trailing title (rules.md#H5's " + "Accepted 'John van der Berg Prof.'), so releasing the " + "title would put it in the family; the clause keeps it. " + "Unchanged from 2.3.0"), + Case("a_numeral_the_bound_join_would_take_stays_maiden", + "Berg, abdul nee Smith V", + {"given": "abdul", "family": "Berg", "maiden": "Smith V"}, + classification="fix(#535)", + notes="the numeral stop now asks M2's join question too: " + "released, the V was taken by the bound-given join " + "after the comma (P5) and read given 'abdul V' at " + "2.3.0 -- a word of the birth name in the current one"), + Case("a_period_final_bracket_reads_as_the_bare_clause", + "Jane Doe (nee Smith Prof.)", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "maiden": "Smith"}, + classification="fix(#535)", + notes="rules.md#M2 Accepted: bracket content ending in a " + "period is not extracted, so the brackets drop and the " + "clause is read bare -- and the bare clause now gives " + "the title up. 2.3.0 read maiden 'Smith Prof.' for the " + "same reason the bare form did"), + Case("the_floor_keeps_the_first_word_out_of_the_chains_count", + "Jane Doe nee King. ba", + {"given": "Jane", "family": "Doe", "suffix": "ba", + "maiden": "King."}, classification="fix(#533)", ambiguities=("suffix-or-name",), - notes="the boundary's other side, and the pair is the " - "finding: the two spellings DISAGREE here, where at " - "the given part's own trailing slot they agree " - "(#531's 'Doe, John MA Prof.' and 'Doe, John Prof. " - "MA' land on one answer). The peel reaches 'MA' " - "because nothing stands behind it, so the clause stops " - "and 'Prof.' stays maiden text. 1.4.0 read family " - "'Prof.', suffix 'MA'", - shape=1), + notes="2.3.0 read maiden 'King. ba' as it read 'Smith ba' -- " + "no credential stop before #533; the floor keeps the " + "first word in the chain-and-peel count, so the " + "reading #533 gave stands"), + Case("a_given_slot_numeral_with_a_credential_tail_stays", + "Doe, Jane nee Smith V, PhD", + {"given": "Jane", "family": "Doe", "suffix": "PhD", + "maiden": "Smith V"}, + classification="fix(#535)", + notes="rules.md#M2's given-slot reader declines the numeral " + "where a third comma part follows (#144's own " + "condition, asked of the maiden clause too): the given " + "part is not the LAST comma part, so the release is " + "withdrawn and the clause keeps 'V'. 2.3.0 read middle " + "'V', maiden 'Smith', suffix 'PhD' -- an M2 violation " + "predating #535"), + Case("a_lone_numeral_before_a_credential_tail_stays_maiden", + "Doe, Jane nee V, PhD", + {"given": "Jane", "family": "Doe", "suffix": "PhD", + "maiden": "V"}, + classification="fix(#535)", + notes="the numeral stop reads FROM the marker and is not held " + "to the first-word floor; here the given slot does not " + "read a lone numeral as a suffix with a third comma part " + "behind it (#144's condition, asked of the clause since " + "#535), so the clause keeps 'V', as 2.0.0 and 2.1.0 read " + "it. 2.2.0, 2.3.0 and the #538 commit d9d80492 read " + "middle 'nee V', the marker a name word; 'Doe, Jane nee " + "V, Jr.' moves the same way"), + Case("the_given_title_chain_stops_at_a_member_with_a_title_behind", + "Doe, Jane nee Smith Rev. MA Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "MA", "maiden": "Smith Rev."}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="after a family comma the given part's title chain runs " + "from the end over what its first suffix pass leaves, " + "and that pass leaves 'MA' a name word while a title " + "stands behind it -- 'Doe, Jane Dr. MA Prof.' reads " + "middle 'Dr.' -- so the chain stops at 'MA' and 'Rev.' " + "stays in the clause; only 'MA' and 'Prof.' leave. 2.3.0 " + "and the #538 commit d9d80492 read maiden 'Smith Rev. " + "MA Prof.'"), + Case("a_lenient_numeral_leaves_with_the_title_in_front", + "Doe, Jane nee Smith Prof. V", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "V", "maiden": "Smith"}, + classification="fix(#535)", + notes="the lenient trailing numeral (#144) is a suffix of the " + "given part's first pass once the numeral stop has asked " + "it, so the title in front of it leaves the clause too, " + "as 'Doe, Jane Prof. V' reads title 'Prof.', suffix 'V'. " + "2.3.0 read maiden 'Smith Prof.', suffix 'V'"), + Case("a_title_behind_that_numeral_still_stays", + "Doe, Jane nee Smith V Prof., PhD", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "PhD", "maiden": "Smith V"}, + classification="fix(#535)", + notes="the same third-comma-part decline reaches through the " + "title chain: the numeral stays maiden text and the " + "title behind it still leaves the clause. 2.3.0 read " + "maiden 'Smith V Prof.', suffix 'PhD', with no title at " + "all"), + Case("the_bound_given_join_is_the_given_slots_alone", + "abdul nee Smith V", + {"given": "abdul", "suffix": "V", "maiden": "Smith"}, + ambiguities=("suffix-or-name",), + notes="P5's bound-given join is LENIENT only after a family " + "comma; before one the STRICT reserve declines a join " + "that would change a suffix reading, so the release " + "here is not withdrawn and the clause gives 'V' up as " + "'Berg, abdul nee Smith V' does not. Unchanged from " + "2.3.0"), + Case("a_title_released_from_the_bound_given_pair", + "abdul nee Smith Dr.", + {"title": "Dr.", "family": "abdul", "maiden": "Smith"}, + classification="fix(#535)", + notes="the same STRICT-reserve reading for a title behind the " + "bound word: 'Dr.' leaves the clause and 'abdul' reads " + "as the family the title stands in front of, as " + "'abdul Dr.' reads it bare. 2.3.0 read given 'abdul', " + "maiden 'Smith Dr.'"), + Case("a_numeral_straight_after_the_marker_declines_through_a_title", + "Jane Doe nee V Prof.", + {"title": "Prof.", "given": "Jane", "middle": "Doe", + "family": "nee", "suffix": "V"}, + classification="fix(#535)", + ambiguities=("suffix-or-name",), + notes="the numeral stop reads FROM the marker by design, " + "unlike the credential and title stops, so a numeral " + "standing straight after it is not held to the " + "first-word floor, and here it declines the clause -- " + "as 'Jane " + "Smith née V' does bare (rules.md#M2) -- and the title " + "behind the numeral then reads the declined name as " + "'Jane Doe nee V' plus the title would. 2.3.0 read " + "maiden 'V Prof.'"), + Case("a_title_behind_a_tail_bound_numeral_keeps_both", + "Doe, Jane nee Smith Prof. V, PhD", + {"given": "Jane", "family": "Doe", "suffix": "PhD", + "maiden": "Smith Prof. V"}, + classification="fix(#535)", + notes="NOT the transparent twin of 'Doe, Jane nee Smith V " + "Prof., PhD' (title 'Prof.', maiden 'Smith V'): the " + "third comma part's credential tail withdraws the " + "numeral's release, so 'V' stays a name word, and the " + "title stop is then asked and declines -- the given " + "part's chain runs from the end and stops at 'V', so " + "it never reaches 'Prof.', and the title stays maiden " + "text in front of the numeral. The " + "same asymmetry is the bare given slot's own, with no " + "marker in it -- 'Doe, Jane Prof. V, PhD' reads middle " + "'Prof. V' (no title) where 'Doe, Jane V Prof., PhD' " + "reads title 'Prof.', middle 'V'. 2.3.0 read middle " + "'V', maiden 'Smith Prof.'"), + Case("a_link_the_walk_stops_at_gives_up_only_what_reads_off", + "Jane Doe nee Smith i DO Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "maiden": "Smith i DO"}, + classification="fix(#397/#535)", + ambiguities=("suffix-or-name",), + notes="with the title chained, the DO is the trailing peel's, " + "so the link exception refuses 'i' and the walk would " + "stop there -- and a stop at a link gives up the words " + "behind it. The name left standing ('Jane Doe i DO " + "Prof.') does not read 'i DO' as post-nominals, so the " + "clause keeps the link and the DO (reporting the kept " + "credential), and only the title leaves. 2.3.0 read " + "middle 'Doe i', family 'DO Prof.', maiden 'Smith' ('i' " + "was a plain suffix word there); the #538 commit " + "d9d80492 read maiden 'Smith i DO Prof.'"), + Case("a_link_whose_run_would_land_in_a_name_part_stays", + "Berg, abdul nee Smith i V Prof., MD", + {"given": "abdul", "family": "Berg", "suffix": "MD", + "maiden": "Smith i V Prof."}, + classification="fix(#397)", + notes="the given-slot twin of the row above: giving up " + "'i V Prof.' would leave the V a middle name behind the " + "bound pair, so the clause keeps the whole run. 2.3.0 " + "read middle 'V', title 'Prof.', suffix 'i, MD', maiden " + "'Smith'; the #538 commit d9d80492 read as this does"), + Case("a_released_particle_title_is_kept_after_a_family_comma", + "Doe, Jane nee Smith St.", + {"given": "Jane", "family": "Doe", "maiden": "Smith St."}, + classification="fix(#274)", + notes="'St.' is a title and a particle. Released after a " + "family comma, P6 would attach it to the family ('Doe, " + "Jane St.' reads family 'St. Doe'), a word of the birth " + "name carried into the current one, so the clause keeps " + "it. Unchanged from 2.3.0"), + Case("a_released_particle_title_behind_a_credential_is_kept", + "Doe, Jane nee Smith MA St.", + {"given": "Jane", "family": "Doe", "maiden": "Smith MA St."}, + classification="fix(#274)", + notes="the same guard with a credential in front of the " + "title: releasing 'MA St.' would hand 'St.' to the " + "family, so the clause keeps both. Unchanged from 2.3.0"), + Case("a_split_credential_behind_the_member_counts_as_released", + "Jane Doe nee Smith MA Ph. D.", + {"given": "Jane", "family": "Doe", "suffix": "MA Ph. D.", + "maiden": "Smith"}, + classification="fix(#533)", + ambiguities=("suffix-or-name",), + notes="the TRAILING reader's release check counts a piece " + "group already flagged a credential ('Ph. D.', merged " + "from two tokens) as read off, so 'MA Ph. D.' leaves " + "whole. 2.3.0 read maiden 'Smith MA', suffix 'Ph. D.'"), + Case("a_split_credential_behind_the_member_after_a_family_comma", + "Doe, Jane nee Smith MA Ph. D.", + {"given": "Jane", "family": "Doe", "suffix": "MA Ph. D.", + "maiden": "Smith"}, + classification="fix(#533)", + ambiguities=("suffix-or-name",), + notes="the given-slot twin of the row above. 2.3.0 read " + "maiden 'Smith MA', suffix 'Ph. D.'"), Case("a_connective_behind_the_member_stops_the_peel", "Jane Doe nee Smith MA y", {"given": "Jane", "family": "Doe", "maiden": "Smith MA y"}, @@ -7620,16 +7897,15 @@ def _check_cjk_shape_purity(self) -> None: "2026-09-09)"), Case("title_word_trailing_after_a_maiden_take", "Mary Smith née Jones Prof.", - {"given": "Mary", "family": "Smith", "maiden": "Jones Prof."}, - classification="fix(#274)", - notes="negative control for the trailing walk, and rules.md#" - "H5's M2 boundary: M2's take runs to the name's end " - "and the title is inside what it takes, so no title " - "word is in trailing position at all. Unchanged by " - "#316/#489 -- master reads the same (measured " - "2026-09-09). 1.4.0 had no maiden support and read " - "first 'Mary' / middle 'Smith née Jones' / last " - "'Prof.'"), + {"title": "Prof.", "given": "Mary", "family": "Smith", + "maiden": "Jones"}, + classification="fix(#535)", + notes="rules.md#H5's M2 boundary, removed by #535: the maiden " + "walk now reads the trailing title chain, so the title " + "ends the clause and reads as 'Mary Smith Prof.' would " + "read it. 2.3.0 read maiden 'Jones Prof.'; 1.4.0 had no " + "maiden support and read first 'Mary' / middle 'Smith " + "née Jones' / last 'Prof.'"), Case("title_word_trailing_ahead_of_a_maiden_marker", "Mary Jones Prof. née Smith", {"title": "Prof.", "given": "Mary", "family": "Jones", diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index 1077db49..e01fc463 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -11,7 +11,8 @@ from nameparser._pipeline._classify import classify from nameparser._pipeline._extract import extract_delimited, _maiden_marked from nameparser._pipeline._group import ( - TailReader, _group_segment, group, marker_run_length, + TailReader, _group_segment, _release_reads_off, group, + marker_run_length, ) from nameparser._pipeline._script_segment import script_segment from nameparser._pipeline._segment import segment @@ -1142,6 +1143,19 @@ def test_an_unmapped_reader_is_a_loud_failure_rather_than_a_default( maiden_ambiguities=[]) +def test_the_release_check_fails_loudly_on_an_unmapped_reader( +) -> None: + """`_release_reads_off` dispatches on the two readers that reach it + and ends with `assert_never`, as `_maiden_take`'s one dispatch + does. Unreachable at runtime by construction, so reached here with + a value outside the enum -- the loudness pinned, and the line kept + from being an uncovered statement.""" + with pytest.raises(AssertionError): + _release_reads_off([[0]], [set()], [], 0, 1, 0, + cast(TailReader, 99), None, # type: ignore[arg-type] + tail_follows=False) + + def test_the_reader_is_pinned_to_the_structure_it_is_read_from( monkeypatch: pytest.MonkeyPatch, ) -> None: @@ -1810,42 +1824,62 @@ def test_a_core_between_the_marker_and_the_first_word_is_below_lo( assert _maiden_texts(plain) == ["-", "i", "Jones"] -def test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538( +def test_a_core_beside_a_link_is_stepped_over_like_a_connective( ) -> None: - """A KNOWN-WRONG reading, pinned so the repair has to move it. - - rules.md#M2 gives the link exception a name word on each side, - and a delimiter core is structure rather than a name word -- so - the clause below should end where its separator-less twin ends. - It does not. Update this test when #538 lands: the assertion - beneath the first parse is the deviation, not the contract, and - rules.md#M2's `deviates: #538` example is its other half. + """rules.md#M2 gives the link exception a name word on each side, + and a delimiter core is structure the caller declared rather than + a name word -- so the word on a link's side is the one past the + core, and the clause reads as the same text written without it + (#538). + + Past the clause's first word the core used to be an ordinary index + to the neighbour walk, which stepped over connectives and nothing + else, so it stood in for the name word on the link's left and the + clause ran on past the link it otherwise ends at. The population is + decisions.md's 2026-09-22 #397 follow-up entry: none of it + reachable at the default policy, `extra_suffix_delimiters` being + empty there. """ - # WHAT IS NOT TRUE OF A CORE PAST `lo`, pinned as it reads rather - # than as it ought to: inside the clause a core is an ordinary - # index to `_run_neighbours`, which steps over CONNECTIVES and - # nothing else, so it stands as the name word on the link's left - # and the clause runs on past a title it would otherwise stop at. - # `_between_name_words` is asked about a core on one side or the - # other in 51,072 of 900,023 calls over the population above, and - # the answer differs from a core-skipping reading in 8,094 parses - # (1,278 texts); 1,824 of those move the `maiden` field, on 288 - # texts. None of the 288 is reachable at the default policy, - # `extra_suffix_delimiters` being empty there -- so the one of - # them rules.md#M2 now carries as a `deviates: #538` example (this - # row's first text) enters corpus_rules.jsonl as a name the gate - # parses with the DEFAULT facade, where it moves for the link fix - # and not for this. Reported, not fixed: the repair is `cores` - # threaded through three call sites into `_run_neighbours`, not a - # one-liner (#538). out = _grouped("Smith, John, PhD née Puig Mr. - i Soler", policy=_DASH, lexicon=_LINK_LEX) - assert _maiden_texts(out) == ["Puig", "Mr.", "i", "Soler"] + assert _maiden_texts(out) == ["Puig", "Mr."] # the same clause with the core taken out of it: the title IS the # word on the link's left and refuses, so the clause ends there. without = _grouped("Smith, John, PhD née Puig Mr. i Soler", policy=_DASH, lexicon=_LINK_LEX) assert _maiden_texts(without) == ["Puig", "Mr."] + # and between two NAME words the core is stepped over too, so the + # link joins exactly as it joins written without the core -- the + # skip reading, not a boundary one, which would have ended the + # clause at 'Puig' and pushed the link into the credentials. + between = _grouped("Smith, John, PhD née Puig - i Soler", + policy=_DASH, lexicon=_LINK_LEX) + assert _maiden_texts(between) == ["Puig", "i", "Soler"] + + +def test_a_core_beside_a_link_in_a_credential_tail_is_dropped() -> None: + """The same stepping applies outside a maiden clause, in the + `frozen` loop's own `_run_neighbours` call. In BOTH texts below, + what stands beyond the core is a credential or nothing -- never a + name word -- so the link's neighbour search, stepping past the + core, finds no name word there either and stays a lone suffix + word rather than joining. The core is then a lone piece with + nothing joined to it, which the #206 drop removes exactly as it + removes any lone core, and the entries it stood between separate + the way they already do in 'PhD - MD' -> 'PhD, MD' (#538). Where a + name word stands beyond the core instead, the link joins across it + and the core survives in the suffix text -- not this test's shape; + rules.md#P3's separator sentence states the join, decisions.md's + #538 entry the surviving core. + + RECORDED NEGATIVE CONTROL: at e0f1a2fa, before the frozen loop + stepped over a core, these read suffix 'PhD - i Soler' and + '- i Puig' -- the core kept inside the piece the link joined. + """ + dash = Parser(policy=Policy(extra_suffix_delimiters=frozenset({" - "}))) + assert str(dash.parse("Smith, John, PhD - i Soler").suffix) == \ + "PhD, i Soler" + assert str(dash.parse("Smith, John, - i Puig").suffix) == "i Puig" def test_a_marker_with_nothing_after_it_declines_before_the_bound( @@ -1967,3 +2001,157 @@ def test_a_frozen_link_is_still_absorbed_by_a_neighbours_join() -> None: assert out.family == "Puig y i" assert out.middle == "Carod Rovira" assert out.suffix == "" + + +def test_the_title_stop_needs_a_name_word_left_standing() -> None: + """rules.md#M2 (#535): the title stop is asked over the name the + take would leave. 'Dr. nee Jones Smith Prof.' would leave 'Dr. + Prof.', where H5's chain has no name word to stand behind and the + title would read as the family name -- so the clause keeps it.""" + out = _grouped("Dr. nee Jones Smith Prof.", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["Jones", "Smith", "Prof."] + # the control: with a name word ahead of the marker the same clause + # gives the title up + ok = _grouped("Jane Doe nee Jones Smith Prof.", lexicon=Lexicon.default()) + assert _maiden_texts(ok) == ["Jones", "Smith"] + + +def test_a_released_title_behind_a_particle_is_withdrawn() -> None: + """P2's chain runs on over a trailing title (rules.md#H5 Accepted), + so a title the clause gives up with a particle ahead of it in the + remaining name would join the family; the release is withdrawn.""" + out = _grouped("Jane van der Berg nee Smith Prof.", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["Smith", "Prof."] + out = _grouped("Jane Doe nee Smith MA do Prof.", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["Smith", "MA", "do"] + + +def test_the_numeral_stop_asks_the_join_question() -> None: + """The numeral fork never asked whether a join below the take + would absorb the word it releases: after a family comma the + bound-given join took the V ('Berg, abdul nee Smith V' read given + 'abdul V' at 2.2.0 and 2.3.0). It asks now, through the shared + check.""" + out = _grouped("Berg, abdul nee Smith V", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["Smith", "V"] + # the control: with no bound word the numeral is released as before + ok = _grouped("Berg, Jane nee Smith V", lexicon=Lexicon.default()) + assert _maiden_texts(ok) == ["Smith"] + + +def test_the_first_word_floor_holds_a_title_out_of_the_chain() -> None: + """A title straight after the marker stays the maiden name: the + floor is the title chain's own, so the chain never takes that word + and stops at it, taking only the titles behind it.""" + assert _maiden_texts(_grouped("Jane Doe nee King.", lexicon=Lexicon.default())) == ["King."] + assert _maiden_texts(_grouped("Jane Doe nee Prof. Dr.", lexicon=Lexicon.default())) == ["Prof."] + + +def test_the_floor_keeps_the_first_word_in_the_chains_count() -> None: + """rules.md#M2 (#535 review): the floor lives in the chain now, + not in a clamp read after the fact -- the chain may not take the + first word after the marker, so the re-peel's count still holds + it and a trailing suffix behind a chained title reads as + 'Jane Doe nee Smith ba' reads it.""" + out = _grouped("Jane Doe nee King. ba", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["King."] + # the control: an unlisted first word gives the same count + ok = _grouped("Jane Doe nee Smith ba", lexicon=Lexicon.default()) + assert _maiden_texts(ok) == ["Smith"] + + +def test_a_given_slot_numeral_with_a_credential_tail_stays() -> None: + """rules.md#M2/#144 (#535 review): after a family comma the given + slot reads a lone numeral as a suffix only where the given part is + the LAST comma part -- a third comma part behind it withdraws the + release, so 'Doe, Jane nee Smith V, PhD' keeps the V where + 'Doe, Jane nee Smith V' (no tail) gives it up.""" + out = _grouped("Doe, Jane nee Smith V, PhD", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["Smith", "V"] + ok = _grouped("Doe, Jane nee Smith V", lexicon=Lexicon.default()) + assert _maiden_texts(ok) == ["Smith"] + + +def test_the_bound_given_half_of_the_join_model_is_the_given_slots_alone() -> None: + """rules.md#P5 (#535 review): the LENIENT bound-given join only + applies after a family comma; before one the STRICT reserve + declines a join that would change a suffix reading, so a numeral + the peel already reads as a suffix stays given up -- 'abdul nee + Smith V' (no comma) keeps the release where 'Berg, abdul nee + Smith V' (after a comma) withdraws it.""" + out = _grouped("abdul nee Smith V", lexicon=Lexicon.default()) + assert _maiden_texts(out) == ["Smith"] + ok = _grouped("Berg, abdul nee Smith V", lexicon=Lexicon.default()) + assert _maiden_texts(ok) == ["Smith", "V"] + + +def test_the_given_title_chain_stops_at_a_member_with_a_title_behind() -> None: + """rules.md#M2 after a family comma (#535): the given part's title + chain is read from the end over what its first suffix pass leaves, + and that pass takes a class member only where every piece behind + it is taken too -- so 'MA' with 'Prof.' behind it is a name word + there and the chain stops at it ('Doe, Jane Dr. MA Prof.' reads + middle 'Dr.'). A clause that released 'Rev.' as a title would put + it in the middle name; it keeps it instead, and only the title + behind the member and the member itself leave. And the lenient + trailing numeral (#144) counts as that pass's suffix once the + numeral stop has asked it, so 'Prof. V' leaves the clause whole, + as 'Doe, Jane Prof. V' reads.""" + rev = _grouped("Doe, Jane nee Smith Rev. MA Prof.", + lexicon=Lexicon.default()) + assert _maiden_texts(rev) == ["Smith", "Rev."] + num = _grouped("Doe, Jane nee Smith Prof. V", lexicon=Lexicon.default()) + assert _maiden_texts(num) == ["Smith"] + + +def test_a_link_the_walk_stops_at_gives_up_only_a_run_that_reads_off( +) -> None: + """rules.md#M2 (#535): with the title chained, the link exception + can refuse a link it used to join, and a stop at a link gives up + the words behind it. 'i DO Prof.' left standing behind 'Jane Doe' + does not read as post-nominals, so the clause keeps the link and + the DO ('Doe i' was the middle name without the check); 'i MA + Prof.' does, so there the link and the credential leave.""" + kept = _grouped("Jane Doe nee Smith i DO Prof.", + lexicon=Lexicon.default()) + assert _maiden_texts(kept) == ["Smith", "i", "DO"] + given_up = _grouped("Jane Doe nee Smith i MA Prof.", + lexicon=Lexicon.default()) + assert _maiden_texts(given_up) == ["Smith"] + + +def test_only_a_link_the_title_chain_refused_asks_the_release_question( +) -> None: + """The link check is asked only where reading the title chain made + the link exception refuse a link it joined over the words as + written. A link refused either way stops as it always did: 'Doe, + Jane nee Smith i V' gives up 'i V' (a check that fired here kept a + dangling 'i' in the birth name). And the given part's lenient + numeral counts as a suffix where only chained titles stand behind + it, as bare 'Doe, Jane i V Prof.' reads suffix 'i V' -- but not + where another comma part follows, where it is a middle initial.""" + plain = _grouped("Doe, Jane nee Smith i V", lexicon=Lexicon.default()) + assert _maiden_texts(plain) == ["Smith"] + titled = _grouped("Doe, Jane nee Smith i V Prof.", + lexicon=Lexicon.default()) + assert _maiden_texts(titled) == ["Smith"] + tail = _grouped("Doe, Jane nee Smith i V Prof., PhD", + lexicon=Lexicon.default()) + assert _maiden_texts(tail) == ["Smith", "i", "V"] + + +def test_a_released_particle_title_is_kept_after_a_family_comma() -> None: + """rules.md#M2 with P6 (#535): after a family comma a released + title that is also a particle would be attached to the family by + P6 ('Doe, Jane St.' reads family 'St. Doe'), so the clause keeps + it. A title that is no particle still leaves, and so does the + credential DO, which the given slot's own lean reads (#533); with + no comma P6 does not run and 'St.' leaves as a title.""" + kept = _grouped("Doe, Jane nee Smith St.", lexicon=Lexicon.default()) + assert _maiden_texts(kept) == ["Smith", "St."] + title = _grouped("Doe, Jane nee Smith Prof.", lexicon=Lexicon.default()) + assert _maiden_texts(title) == ["Smith"] + credential = _grouped("Doe, Jane nee Smith DO", lexicon=Lexicon.default()) + assert _maiden_texts(credential) == ["Smith"] + no_comma = _grouped("Jane Doe nee Smith St.", lexicon=Lexicon.default()) + assert _maiden_texts(no_comma) == ["Smith"] diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index d34e6001..6bfdefeb 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1490,6 +1490,57 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: "fix(#424/#445) accepted: the maiden walk keeps a bare acronym": ("John née Jones Smith MA", "John nee Jones Smith MA PHD", "JOHN NEE JONES SMITH MA"), + # #535's title-stop rule is a literal list over the SLOT, for the + # same reason fix(#533)'s rules give: a superstring with a name + # word standing ahead of the marker's release, or behind the + # title, must not match. Used at 2.0.0-2.3.0, where the rule is + # #535-only for the names it still carries. + "fix(#535) a trailing title ends the maiden clause": + ("Dr. Jane Doe nee Smith Prof.", "Jane Doe nee Smith Prof. Dr."), + # 1.4.0's copy: every name here differs from that baseline for + # #274's reason too (no maiden reading at all), so the issue is + # relabeled -- same probes, same reason. + "fix(#274/#535) a trailing title ends the maiden clause": + ("Dr. Jane Doe nee Smith Prof.", "Jane Doe nee Smith Prof. Dr."), + # The lone-marker shape's own probe, at 1.4.0: a name word standing + # ahead of the marker or behind the title must not match either -- + # 'Dr. nee Jones Smith Prof.' is the boundary where only a TITLE + # precedes the marker, so no name word is left standing there. + "fix(#274) a trailing title after the marker with nothing else in the name": + ("Jane Dr. nee Jones Smith Prof.", "Dr. nee Jones Smith Prof. Jr."), + # The 1.4.0-only Dr.-postnominal rule's own probe: a name word + # standing ahead of the marker or behind the title must not match. + "fix(#274/#296/#535) a trailing title after a dropped postnominal Dr., with no maiden reading either": + ("Dr. Jane Doe nee Prof. Dr.", "Jane Doe nee Prof. Dr. Jr."), + # The numeral rule's own probe: a name word ahead of the released + # numeral must not match either. Used at 2.2.0/2.3.0, where the + # rule is #535-only. + "fix(#535) the numeral stop asks the join question": + ("Berg, Jane nee Smith V", "Dr. Berg, abdul nee Smith V"), + # This baseline's copy, shared with 1.4.0 (both read 'Berg, abdul + # nee Smith V' identically -- #274 moves nothing here, so 1.4.0 + # carries the same two-issue label rather than a third): the + # bound-given reserve (#411) already decides the shape before + # #535 exists, so the issue is relabeled -- same probes, same + # reason. + "fix(#411/#535) the numeral stop asks the join question": + ("Berg, Jane nee Smith V", "Dr. Berg, abdul nee Smith V"), + # The #399 sibling's own probe, at 2.0.0/2.1.0: a name word + # standing ahead of the marker or behind the title must not match. + "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words": + ("Dr. Jane van der Berg nee Smith Prof.", + "Jane van der Berg nee Smith Prof. Jr."), + # The dropped-Dr.-postnominal rule's own probe, at 2.0.0/2.1.0. + "fix(#296/#535) a trailing title after a dropped postnominal Dr.": + ("Dr. Jane Doe nee Prof. Dr.", "Jane Doe nee Prof. Dr. Jr."), + # The title-in-front-of-the-credential rule's own probe, at + # 2.0.0-2.3.0. + "fix(#533/#535) a title in front of the credential the clause gives up": + ("Dr. Jane Doe nee Smith Prof. MA", "Jane Doe nee Smith Prof. MA Jr."), + # The given-slot-numeral-with-a-tail rule's own probe, at 2.2.0 + # and 2.3.0. + "fix(#535) a given-slot numeral with a credential tail stays": + ("Dr. Doe, Jane nee Smith V, PhD", "Doe, Jane nee Smith V, PhD Jr."), } @@ -2188,6 +2239,13 @@ class _LatinCopy(NamedTuple): "Jane Doe nee MA PhD"}), frozenset({"Jane Doe \\(nee Smith MA\\)", "Jane Doe \\(nee Smith Ma\\)", "Jane Doe \\(nee Smith\\) MA"}), + # rules.md#M2's two live examples of the delimiter-core fix (#538), + # one alternative per corpus name. A list of names, not a copy of + # any wordlist: what selects them is the doc's own choice of + # examples for a policy this corpus does not configure, so there + # is no vocabulary here for the alternation to drift from. + frozenset({"Smith, John, PhD née Puig Mr\\. - i Soler", + "Smith, John, PhD née Puig - i Soler"}), # fix(#400)'s two openings: start-of-name or just after a family # comma. `abd` joins forward on the given side wherever that side # begins, and the alternation is over ANCHORS, not over words -- @@ -2200,6 +2258,66 @@ class _LatinCopy(NamedTuple): # _vocab.D are fixed regexes, so there is no wordlist here for the # alternation to drift from. frozenset({" John Smith", " Van Johnson", ", Jr\\."}), + # #535's title-stop rule, one alternative per corpus name -- a + # list of names, not a copy of any wordlist: what selects them is + # the SHAPE the release check turns on (a name word left standing + # for the title chain to end behind, no join below the marker + # pass absorbing the released span), and no vocabulary decides + # that. Three sets, one per baseline's remaining membership after + # the review that split out every name whose diff from that + # baseline has an EARLIER cause too (each such name got its own + # joint-labelled rule instead, verified against the parent tree + # d9d80492): 1.4.0 is additionally relabelled fix(#274/#535), + # since v1 has no maiden reading at all and so differs from every + # remaining name for #274's reason as well as #535's, and loses + # 'Jane Doe nee Prof. Dr.' too (#274/#296, v1's missing maiden + # reading plus dr still being suffix vocabulary there); 2.0.0/2.1.0 + # lose 'Jane van der Berg nee Smith Prof.' (#399, a two-trailing- + # word marker clause that baseline's own particle-chain rule + # cannot yet reach), 'Jane Doe nee Prof. Dr.' (#296 alone there, + # dr leaving the postnominal vocabulary) and 'Jane Doe nee Smith + # Prof. MA' (#533, the credential the clause gives up); 2.2.0/2.3.0 + # lose only the last of those three, #296 and #399 both predating + # those baselines. 2.3.0 alone also carries 'Doe nee Smith ba + # Prof.', a rules.md#M2 Accepted example: at every earlier + # baseline its diff has #342's cause too ('ba' was a plain suffix + # word there) and it sits in a joint-labelled rule instead. + frozenset({"Doe, Jane nee Smith MA Prof\\.", + "Jane Doe \\x28nee Smith Prof\\.\\x29", + "Jane Doe nee Smith King\\.", + "Jane Doe nee Smith MA Prof\\.", + "Jane Doe nee Smith Ma Prof\\.", + "Jane Doe nee Smith Prof\\.", + "Jane Doe nee Smith Prof\\. MA", + "Jane Doe nee Smith V Prof\\.", + "Mary Smith née Jones Prof\\."}), + frozenset({"Doe, Jane nee Smith MA Prof\\.", + "Jane Doe \\x28nee Smith Prof\\.\\x29", + "Jane Doe nee Smith King\\.", + "Jane Doe nee Smith MA Prof\\.", + "Jane Doe nee Smith Ma Prof\\.", + "Jane Doe nee Smith Prof\\.", + "Jane Doe nee Smith V Prof\\.", + "Mary Smith née Jones Prof\\."}), + frozenset({"Doe, Jane nee Smith MA Prof\\.", + "Jane Doe \\x28nee Smith Prof\\.\\x29", + "Jane Doe nee Prof\\. Dr\\.", + "Jane Doe nee Smith King\\.", + "Jane Doe nee Smith MA Prof\\.", + "Jane Doe nee Smith Ma Prof\\.", + "Jane Doe nee Smith Prof\\.", + "Jane Doe nee Smith V Prof\\.", + "Mary Smith née Jones Prof\\."}), + frozenset({"Doe nee Smith ba Prof\\.", + "Doe, Jane nee Smith MA Prof\\.", + "Jane Doe \\x28nee Smith Prof\\.\\x29", + "Jane Doe nee Prof\\. Dr\\.", + "Jane Doe nee Smith King\\.", + "Jane Doe nee Smith MA Prof\\.", + "Jane Doe nee Smith Ma Prof\\.", + "Jane Doe nee Smith Prof\\.", + "Jane Doe nee Smith V Prof\\.", + "Mary Smith née Jones Prof\\."}), # #436/#437's Latin movers, one corpus name per alternative -- a # list of names, not a copy of any wordlist, so there is no # vocabulary for it to drift from. The rule's subject is the @@ -2637,39 +2755,73 @@ class _LatinCopy(NamedTuple): # at 2.2.0, and 'John née Jones Smith MA' needs a compound rule of # its own below 2.2.0 where #445's move is still in the diff. # - # The movers, at 2.2.0 and 2.3.0 (seventeen): + # The movers, at 2.2.0, sixteen since #535 dropped 'Jane Doe nee + # Smith Prof. MA' -- the title chain now ends the clause before + # this rule's credential reading is asked, so the name moved to + # its own fix(#535) rule: frozenset({r"Doe, J\. nee MA ba", r"Doe, Jane Q\. nee Smith MA", "Doe, Jane nee Smith DO", "Doe, Jane nee Smith MA", "Doe, Jane nee Smith ma", "JANE DOE NEE SMITH MA", "JANE DOE NEE YO-YO MA", r"Jane Doe geb\. Smith MA", "Jane Doe nee Smith MA", "Jane Doe nee Smith MA JD", "Jane Doe nee Smith MA PhD", "Jane Doe nee Smith Ma JD", - r"Jane Doe nee Smith Prof\. MA", "Jane Doe nee Smith V MA", + "Jane Doe nee Smith V MA", "Jane Doe nee Smith do", "John née Jones Smith MA", "Maria Kowalska z domu Nowak MA", "jane doe nee smith ma"}), - # and at 2.0.0 and 2.1.0 (fourteen). + # and at 2.3.0, seventeen for the same reason plus one: 'Jane Doe + # nee King. ba' joins here too (the review's own floor boundary, + # this baseline never having split its credential at all). + frozenset({r"Doe, J\. nee MA ba", r"Doe, Jane Q\. nee Smith MA", + "Doe, Jane nee Smith DO", "Doe, Jane nee Smith MA", + "Doe, Jane nee Smith ma", "JANE DOE NEE SMITH MA", + "JANE DOE NEE YO-YO MA", r"Jane Doe geb\. Smith MA", + r"Jane Doe nee King\. ba", + "Jane Doe nee Smith MA", "Jane Doe nee Smith MA JD", + "Jane Doe nee Smith MA PhD", "Jane Doe nee Smith Ma JD", + "Jane Doe nee Smith V MA", + "Jane Doe nee Smith do", + "John née Jones Smith MA", "Maria Kowalska z domu Nowak MA", + "jane doe nee smith ma"}), + # and at 2.0.0 and 2.1.0, thirteen for the same reason. frozenset({r"Doe, Jane Q\. nee Smith MA", "Doe, Jane nee Smith DO", "Doe, Jane nee Smith MA", "Doe, Jane nee Smith ma", "JANE DOE NEE SMITH MA", "JANE DOE NEE YO-YO MA", r"Jane Doe geb\. Smith MA", "Jane Doe nee Smith MA", "Jane Doe nee Smith MA JD", "Jane Doe nee Smith MA PhD", "Jane Doe nee Smith Ma JD", - r"Jane Doe nee Smith Prof\. MA", "Jane Doe nee Smith V MA", + "Jane Doe nee Smith V MA", "Jane Doe nee Smith do", "jane doe nee smith ma"}), - # The names the clause KEEPS and now reports, at 2.2.0 and 2.3.0 - # (eight), + # The names the clause KEEPS and now reports, at 2.3.0 (eight), + # plus 'Doe nee Smith Prof. ba' (2026-09-26), rules.md#M2's + # Accepted example of a title in front of a kept credential, + # which #535 moves nothing on, + frozenset({"Berg, Jane van der nee Smith DO", "Berg, abdul nee Jones MA", + r"Doe nee Smith Prof\. ba", + r"Doe, Dr\. nee Smith MA", "Doe, Jane nee Smith Do", + "Doe, Jane nee Smith MA do", "Doe, Jane nee Smith Ma", + "Doe, Jane nee Smith do", "JOHN NEE JONES SMITH MA PHD", + "Jane Doe nee MA", "Jane Doe nee MA PhD", + "Jane Doe nee Smith DO DO", "Jane Doe nee Smith Ma", + "Jane Doe nee Yo-Yo Ma", "John née Jones Smith Ma"}), + # and at 2.2.0 (nine): 'Jane Doe nee King. ba' joins here (the + # review's own floor boundary; unlike at 2.3.0, this baseline + # already split the credential, so only the report is new). frozenset({"Berg, Jane van der nee Smith DO", "Berg, abdul nee Jones MA", r"Doe, Dr\. nee Smith MA", "Doe, Jane nee Smith Do", "Doe, Jane nee Smith MA do", "Doe, Jane nee Smith Ma", "Doe, Jane nee Smith do", "JOHN NEE JONES SMITH MA PHD", + r"Jane Doe nee King\. ba", "Jane Doe nee MA", "Jane Doe nee MA PhD", "Jane Doe nee Smith DO DO", "Jane Doe nee Smith Ma", "Jane Doe nee Yo-Yo Ma", "John née Jones Smith Ma"}), - # and at 2.0.0 and 2.1.0 (five). - frozenset({r"Doe, J\. nee MA ba", "Doe, Jane nee Smith Ma", - "Jane Doe nee MA", "Jane Doe nee Smith Ma", + # and at 2.0.0 and 2.1.0 (nine): the review added 'Jane Doe nee + # King. ba' beside 'Doe, J. nee MA ba', the same FLOOR boundary. + frozenset({r"Doe, Dr\. nee Smith MA", r"Doe, J\. nee MA ba", + "Doe, Jane nee Smith Ma", r"Jane Doe nee King\. ba", + "Jane Doe nee MA", "Jane Doe nee MA PhD", + "Jane Doe nee Smith DO DO", "Jane Doe nee Smith Ma", "Jane Doe nee Yo-Yo Ma"}), # The two names where the released member lands somewhere other # than the trailing peel. One set, shared by the 1.4.0 rule and @@ -2728,15 +2880,16 @@ class _LatinCopy(NamedTuple): "abdul Smith Jr V"}), # #397's maiden-clause rule, one corpus name per alternative -- a # list of names, not a copy of any wordlist. What selects the - # three is the PLACEMENT of the link inside a clause, which no + # first two is the PLACEMENT of the link inside a clause, which no # vocabulary decides; a member spelled as the shape (a bare letter # after a maiden marker) would reach every clause name in the # corpora and pre-excuse the readings the walk must refuse. The - # third member is rules.md#M2's `deviates: #538` example, which - # the corpus parses at the DEFAULT policy and so for this rule's - # own sentence rather than for #538's. + # last two are rules.md#M2's two live examples of the #538 fix, + # which the corpus parses at the DEFAULT policy and so for this + # rule's own sentence rather than for #538's. frozenset({"Doe, Jane nee Puig i Soler", "Jane Doe nee Puig i Soler", - "Smith, John, PhD née Puig Mr\\. - i Soler"}), + "Smith, John, PhD née Puig Mr\\. - i Soler", + "Smith, John, PhD née Puig - i Soler"}), # #397's join and its one-case report, one corpus name per # alternative -- lists of names, not copies of any wordlist. The # join's subject is a SHAPE the vocabulary participates in at one @@ -3328,8 +3481,21 @@ def _claim(rule: dict) -> _Claim: # 'Smith, John, PhD née Puig Mr. - i Soler'. Reach again -- # it carries a marker -- and verified name by name; no role # joined the list. + # 2026-09-26, #538: 78 -> 79, rules.md#M2's second live + # example, 'Smith, John, PhD née Puig - i Soler', added to the + # corpus. Reach again -- it carries a marker -- and verified + # name by name; no role joined the list. + # 2026-09-26, #535: 79 -> 89, the corpus growth from the + # title-stop rebuild ('Jane Doe nee Smith Prof.' and its + # siblings) -- every one carries a marker and this regex + # reaches it; no role joined the list. + # 2026-09-26, #535 review: 89 -> 91, the two rules.md#M2 + # examples the floor/tail-follows fix added ('Jane Doe nee + # King. ba', 'Doe, Jane nee Smith V, PhD'). Reach again -- + # both carry a marker -- and verified name by name; no role + # joined the list. "fix(#274) maiden markers consumed": - _Claim(78, ('family', 'maiden', 'middle'), '08e08f62622f', None), + _Claim(105, ('family', 'maiden', 'middle'), 'c48af99f8184', None), # 2026-09-19, #533: 5 -> 6, the same one new corpus name # '田中 太郎 旧姓 佐藤 MA' as the CJK rule above. "fix(cjk-maiden-marker) maiden marker consumed, compounding with the CJK order flip": @@ -3443,8 +3609,24 @@ def _claim(rule: dict) -> _Claim: # 2026-09-25, #540: 369 -> 371, 'John Smith, PhD MEng' and # 'john smith, phd meng', the credential-run comma rows' # shape landing here too. + # 2026-09-26, #538: 371 -> 372, 'Smith, John, PhD née Puig - i + # Soler', rules.md#M2's second live example, its comma landing + # in this rule's reach too. Reach again, verified name by + # name. + # 2026-09-26, #538 (frozen-loop half): 372 -> 373, 'Smith, + # John, PhD - i Soler', rules.md#R1's new example, its comma + # landing in this rule's reach too. Reach again, verified name + # by name. + # 2026-09-26, #535: 373 -> 376, three comma-fronted names the + # title-stop rebuild added ('Berg, abdul nee Smith V' and two + # of its siblings), each a lone post-comma piece this rule + # already reaches. Reach again, verified name by name. + # 2026-09-26, #535 review: 376 -> 377, 'Doe, Jane nee Smith V, + # PhD', rules.md#M2's third-comma-part example, its comma + # landing in this rule's reach too. Reach again, verified name + # by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(371, ('given', 'suffix', 'title'), 'af9e870177b7', None), + _Claim(380, ('given', 'suffix', 'title'), '658cb8403f4b', None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3461,8 +3643,10 @@ def _claim(rule: dict) -> _Claim: # the two 'Dr.' names now, both having moved to the #316 # rule at the end of the ledger, and the comment there # records the handover. + # 2026-09-26, #535: 13 -> 14, 'Jane Doe nee Prof. Dr.' joined + # the corpus and ends in " Dr." like the rest this rule reaches. "fix(#296) dr is not postnominal vocabulary, so a trailing Dr. is a name word": - _Claim(13, ('family', 'suffix'), "fb9c68f36d0b", None), + _Claim(14, ('family', 'suffix'), "2e229a1fed16", None), "fix(#296) a credential-only comma string reads a name and its postnominal": _Claim(2, ('family', 'given', 'suffix', 'title'), "3f983ff71dee", None), # 2026-09-18: 18 -> 20. Two new corpus names, 'Smith, MA' and @@ -3516,8 +3700,17 @@ def _claim(rule: dict) -> _Claim: # 2026-09-25, #540: 369 -> 371, 'John Smith, PhD MEng' and # 'john smith, phd meng', the credential-run comma rows' # shape landing here too. + # 2026-09-26, #538: 371 -> 372, the same one new comma name as + # the rule above and for the same reason. + # 2026-09-26, #538 (frozen-loop half): 372 -> 373, the same + # one new comma name as the rule above and for the same + # reason. + # 2026-09-26, #535: 373 -> 376, the same three comma-fronted + # names as the rule above and for the same reason. + # 2026-09-26, #535 review: 376 -> 377, the same one new comma + # name as the rule above and for the same reason. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(371, ('family', 'given'), 'af9e870177b7', None), + _Claim(380, ('family', 'given'), '658cb8403f4b', None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -3602,8 +3795,11 @@ def _claim(rule: dict) -> _Claim: # lenient post-comma join takes the released member before # assign can read it, so the clause form agrees with the # bare 'Berg, abdul MA'. Reach, not explanation. + # 2026-09-26, #535: 2 -> 3, 'Berg, abdul nee Smith V' joining + # the corpus -- the same bound-given/maiden shape. Reach, not + # explanation. "fix(#411) the bound-given reserve stops counting words the maiden name takes": - _Claim(2, ('given', 'maiden', 'middle'), "b2500b6f6dfc", None), + _Claim(3, ('given', 'maiden', 'middle'), "9305df1c05c5", None), "fix(#400/#274) bound-given join and maiden consumption in one name": _Claim(1, ('family', 'given', 'maiden', 'middle'), "6bed6d349342", None), "fix(#411/S2) a declining bound-given join leaves the suffix reading after a family comma": @@ -3777,8 +3973,10 @@ def _claim(rule: dict) -> _Claim: # with a bound-given word, 'Berg, abdul MA' and 'Berg, abdul # nee Jones MA' -- the P5 pair this change added to record # that the clause form now agrees with the bare one. + # 2026-09-26, #535: 43 -> 44, 'Berg, abdul nee Smith V' -- + # another opening bound-given word. Reach again. "fix(initials-per-word) a bound-given run initials each word (facade, since 2.0.0)": - _Claim(43, ('_initials',), "2a1c728285b8", ('DEFAULT',)), + _Claim(44, ('_initials',), "4d6f497bebf7", ('DEFAULT',)), # 2026-09-18: 109 -> 110. One new corpus name, # 'john van der berg ma' -- rules.md#P2's one-case contrast, # and a particle chain like every other member. @@ -3786,8 +3984,10 @@ def _claim(rule: dict) -> _Claim: # as the connective-run rule above, 'Carod y de Rovira i', # whose 'de Rovira' is a particle chain; two rules reach one # name and neither widened. Verified name by name. + # 2026-09-26, #535: 112 -> 113, 'Jane van der Berg nee Smith + # Prof.' -- another particle chain, 'van der Berg'. Reach again. "fix(initials-per-word) a particle chain inside a name part initials each word (facade, since 2.0.0)": - _Claim(112, ('_initials',), 'b3b3b696a56e', ('DEFAULT',)), + _Claim(115, ('_initials',), '5f9056683f2f', ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": @@ -4007,8 +4207,109 @@ def _claim(rule: dict) -> _Claim: # 2026-09-22, #397 follow-up: new rule, one literal name -- # rules.md#M2's `deviates: #538` example, which this baseline # reads as one long suffix. + # 2026-09-26, #538: 1 -> 2, rules.md#M2's second live example, + # 'Smith, John, PhD née Puig - i Soler', joining the same + # alternation for the same reason. "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all": - _Claim(1, ('maiden', 'suffix'), 'e20491ebfe62', None), + _Claim(2, ('maiden', 'suffix'), '168b8bdef5db', None), + # New rule (#274): one corpus name, 'Dr. nee Jones Smith + # Prof.' -- only a title precedes the marker, so no name word + # is left standing ahead of it. Its diff from this baseline is + # #274's alone (v1 has no maiden reading), verified against + # the parent tree d9d80492: before #535 ran, it already read + # title 'Dr.', maiden 'Jones Smith Prof.', identically to + # HEAD, so #535 moves nothing here. `given` moves because + # 'nee' is v1's given text; the sibling '#274 maiden markers + # consumed' rule above does not declare it, which is why this + # name needs its own rule rather than joining that one's + # alternation. + "fix(#274) a trailing title after the marker with nothing else in the name": + _Claim(1, ('family', 'given', 'maiden', 'middle'), + 'b6132e37d7a8', None), + # New rule (#274/#296/#535): one corpus name, 'Jane Doe nee + # Prof. Dr.' -- moved out of the nine-name list below (review): + # this baseline still has 'dr' in the suffix vocabulary + # (#296 had not run yet), so it reads suffix 'Dr.' rather than + # keeping the word in the maiden clause; #274 (no maiden + # reading at all) is why it reads middle 'Doe nee' and last + # 'Prof.' where the tree reads family 'Doe' and maiden 'Prof.'. + # Verified against the parent tree + # d9d80492: before #535, it already reads maiden 'Prof. Dr.' + # (dr having left suffix vocabulary by #296), differing from + # this baseline on {family, maiden, middle, suffix}; #535 then + # moves the word on into `title`. + "fix(#274/#296/#535) a trailing title after a dropped postnominal Dr., with no maiden reading either": + _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), + '193350645868', None), + # New rule (#274/#535): eight corpus names -- the title-stop + # literal list, minus 'Dr. nee Jones Smith Prof.' and 'Jane + # Doe nee Prof. Dr.' (each its own rule above). Relabelled + # from a bare '#535' rule (review): this baseline has no + # maiden/title reading at all, so every name here differs from + # it for #274's reason too, verified against the parent tree + # d9d80492 -- before #535 ran, seven of these eight already + # differed from this baseline on exactly {family, maiden, + # middle}, and the eighth (the family-comma one, 'Doe, Jane + # nee Smith MA Prof.') on exactly {maiden, middle, suffix} + # instead, `family` there already matching. #535 additionally + # moves `title` (and `suffix` for the credential-bearing ones + # among the seven; 1.4.0 is compared on the facade, where + # `_ambiguities` does not exist). + "fix(#274/#535) a trailing title ends the maiden clause": + _Claim(9, ('family', 'maiden', 'middle', 'suffix', 'title'), + '6be225d75666', None), + # New rule (#411/#535): one corpus name, 'Berg, abdul nee + # Smith V' -- the numeral-join literal, relabelled from a bare + # '#535' rule (review): this baseline and 2.0.0 read the name + # IDENTICALLY (given 'abdul nee', middle 'Smith', family + # 'Berg', suffix 'V'), so #274 (no maiden reading at all) + # moves nothing here and the label is the same two-issue one + # 2.0.0/2.1.0 carry, not a third. Two causes move it at this + # baseline: #411, the bound-given reserve, which the parent + # tree d9d80492 shows already deciding `given`/`maiden` before + # #535 ran; and #535, which moves the released V from given + # back into maiden. + "fix(#411/#535) the numeral stop asks the join question": + _Claim(1, ('given', 'maiden', 'middle', 'suffix'), + 'c8d3b254b804', None), + # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe + # nee Smith King.' to the title-stop rule, and the + # two rules below hold the Accepted 'ba' pair, one name each. + "fix(#274/#342/#445/#535) a title behind a credential the clause keeps still leaves it": + _Claim(1, ('family', 'given', 'maiden', 'middle', 'title'), + 'b8aa7d51a16c', None), + "fix(#274/#342/#445/#533) a title in front of a credential the clause keeps stays in it": + _Claim(1, ('family', 'given', 'maiden', 'middle', 'suffix'), + 'cb346cb8419a', None), + # 2026-09-26: one-name rules for rules.md#M2's second-round + # Accepted examples (the particle pair, the particle-chain + # title pair, and 'Dr. nee V' at 1.4.0). + "fix(#274/#535) a particle in front of a trailing title stays in the clause": + _Claim(1, ('family', 'maiden', 'middle', 'title'), + 'a1484fe026d1', None), + "fix(#274/#535) a particle behind a trailing title leaves the clause with it": + _Claim(1, ('family', 'maiden', 'middle', 'title'), + '97e10b3d2aa8', None), + "fix(#274/#316) a trailing title behind a maiden clause's first suffix word": + _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), + 'd58b847181aa', None), + "fix(#274) a numeral straight after the marker stays maiden where nothing reads it as a suffix": + _Claim(1, ('family', 'given', 'maiden'), + '4ff67af4187f', None), + # 2026-09-26: one-name rules for rules.md#M2's + # Accepted examples (the particle-inside-the-clause pair, and + # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). + "fix(#274/#535) a particle inside the clause keeps the credential in front of a trailing title": + _Claim(1, ('family', 'maiden', 'middle', 'title'), + '647659c5eafc', None), + "fix(#274/#535) a particle inside the clause lets a credential behind the title go": + _Claim(1, ('family', 'maiden', 'middle', 'title'), + '8db3f5842115', None), + # 2026-09-26: the one-name rule for rules.md#M2's example + # of a link the clause stops at and keeps. + "fix(#274/#397/#535) a link the clause stops at keeps a run that would not read off": + _Claim(1, ('family', 'maiden', 'middle', 'title'), + '1b74e094fed9', None), }, "expected_since_2.0.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -4207,8 +4508,11 @@ def _claim(rule: dict) -> _Claim: # lenient post-comma join takes the released member before # assign can read it, so the clause form agrees with the # bare 'Berg, abdul MA'. Reach, not explanation. + # 2026-09-26, #535: 2 -> 3, 'Berg, abdul nee Smith V' joining + # the corpus -- the same bound-given/maiden shape. Reach, not + # explanation. "fix(#411) the bound-given reserve stops counting words the maiden name takes": - _Claim(2, ('given', 'maiden', 'middle'), "b2500b6f6dfc", None), + _Claim(3, ('given', 'maiden', 'middle'), "9305df1c05c5", None), "fix(#412) a connective join no longer absorbs the maiden marker beside it": _Claim(2, ('family', 'maiden'), "51c0eb36b5c5", None), "fix(#418) the connective carve-out counts the name the maiden clause leaves behind": @@ -4249,8 +4553,10 @@ def _claim(rule: dict) -> _Claim: # `middle` left the ROLES in the same edit, and the reach # grew with the rules corpus; the 1.4.0 roster above carries # both, and says the same at the other two baselines. + # 2026-09-26, #535: 13 -> 14, 'Jane Doe nee Prof. Dr.' joined + # the corpus and ends in " Dr." like the rest this rule reaches. "fix(#296) dr is not postnominal vocabulary, so a trailing Dr. is a name word": - _Claim(13, ('family', 'suffix'), "fb9c68f36d0b", None), + _Claim(14, ('family', 'suffix'), "2e229a1fed16", None), "fix(#296) a credential-only comma string reads a name and its postnominal": _Claim(2, ('suffix', 'title'), "3f983ff71dee", None), # 2026-09-18: 18 -> 20. Two new corpus names, 'Smith, MA' and @@ -4486,14 +4792,22 @@ def _claim(rule: dict) -> _Claim: # the member lands in one, leaves the other and is reported, # so a widening taking any of the three alone would change # the roles here before it reached the gate. + # 2026-09-26, #535: 15 -> 14, 'Jane Doe nee Smith Prof. MA' + # LEAVING this rule's alternation -- the title chain now ends + # the clause before the credential is asked, so the name has + # its own fix(#535) rule instead. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(15, ('_ambiguities', 'maiden', 'suffix'), '1b2218cb6296', None), - # The declining half: eight corpus names at 2.2.0 and 2.3.0, - # five at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role + _Claim(14, ('_ambiguities', 'maiden', 'suffix'), 'e5cd5fb89c8b', None), + # The declining half: fourteen corpus names at 2.2.0, eight + # at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause # gave the member up. + # 2026-09-26: 8 -> 9, 'Jane Doe nee King. ba' -- this + # baseline already splits suffix 'ba' from maiden 'King.', so + # the whole diff is #533's; #535 moves no role here either, + # only the report. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(8, ('_ambiguities',), '329a6472b344', None), + _Claim(9, ('_ambiguities',), 'f999caeb63dc', None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -4558,9 +4872,102 @@ def _claim(rule: dict) -> _Claim: # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name # rules.md#M2's `deviates: #538` example adds, which this # rule's own alternation now names. Verified name by name. + # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, + # 'Smith, John, PhD née Puig - i Soler', joining the same + # alternation for the same reason. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(3, ('family', 'maiden', 'middle', 'suffix'), - '4de0e7570bd6', ('DEFAULT',)), + _Claim(4, ('family', 'maiden', 'middle', 'suffix'), + '4556aa8f93e5', ('DEFAULT',)), + # New rule (#399): one corpus name, 'Jane van der Berg nee + # Smith Prof.' -- a sibling of the general #399 rule above, + # for the two-trailing-word shape its own anchor + # (`\\sn[ée]e\\s+\\S+$`) cannot reach: per rules.md#P2, "a + # maiden marker takes the words after it (M2), or the name + # ends", and the chain here swallows both trailing words. Its + # diff from this baseline is #399's alone, verified against + # the parent tree d9d80492: before #535 ran, it already read + # family 'van der Berg', maiden 'Smith Prof.', identically to + # HEAD. + "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words": + _Claim(1, ('family', 'maiden'), '6b2d1fa71195', None), + # New rule (#296/#535): one corpus name, 'Jane Doe nee Prof. + # Dr.'. This baseline still has 'dr' in the suffix vocabulary + # (#296 had not run yet), so it reads suffix 'Dr.'. The parent + # tree d9d80492, before #535, already reads maiden 'Prof. + # Dr.' (dr having left the suffix vocabulary by #296); #535 + # then moves that word on into `title`, returning maiden to + # 'Prof.' -- the same value this baseline reads there, so + # `maiden` does not move net baseline-to-HEAD. + "fix(#296/#535) a trailing title after a dropped postnominal Dr.": + _Claim(1, ('suffix', 'title'), '193350645868', None), + # New rule (#533/#535): one corpus name, 'Jane Doe nee Smith + # Prof. MA'. #533 already gives 'MA' up to the credential + # reading and reports the fork (`suffix`, `maiden`, + # `_ambiguities`) before #535 exists, verified against the + # parent tree d9d80492; #535 then reads the title chain + # through the credential, so the title standing in front of + # 'MA' moves `title` and `maiden` further. + "fix(#533/#535) a title in front of the credential the clause gives up": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + 'cb85a2cbf9f2', None), + # New rule (#535): seven corpus names -- the title-stop + # literal list, minus the three names above whose diff from + # this baseline has an earlier cause too. `title`, `maiden` + # and `suffix` move for the credential-bearing ones, and + # `_ambiguities` for the declined-member one. + "fix(#535) a trailing title ends the maiden clause": + _Claim(8, ('_ambiguities', 'maiden', 'suffix', 'title'), + '3216299d78a6', None), + # New rule (#411/#535): one corpus name, 'Berg, abdul nee + # Smith V', relabelled from a bare '#535' rule (review): #411 + # (the bound-given reserve) already decides `given`/`maiden` + # at this baseline before #535 exists, verified against the + # parent tree d9d80492; #535 then moves the released V from + # given back into maiden. + "fix(#411/#535) the numeral stop asks the join question": + _Claim(1, ('given', 'maiden', 'middle', 'suffix'), + 'c8d3b254b804', None), + # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe + # nee Smith King.' to the title-stop rule, and the + # two rules below hold the Accepted 'ba' pair, one name each. + "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it": + _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'middle', + 'title'), 'b8aa7d51a16c', None), + "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it": + _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), + 'cb346cb8419a', None), + # 2026-09-26: one-name rules for rules.md#M2's second-round + # Accepted examples (the particle pair, the particle-chain + # title pair, and 'Dr. nee V' at 1.4.0). + "fix(#535) a particle in front of a trailing title stays in the clause": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + 'a1484fe026d1', None), + "fix(#533/#535) a particle behind a trailing title leaves the clause with it": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '97e10b3d2aa8', None), + "fix(#316/#399) a trailing title behind a maiden clause's first suffix word": + _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), + 'd58b847181aa', None), + "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps": + _Claim(1, ('family', 'maiden'), + '58cf2330fed6', None), + # 2026-09-26: one-name rules for rules.md#M2's + # Accepted examples (the particle-inside-the-clause pair, and + # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). + "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + '647659c5eafc', None), + "fix(#533/#535) a particle inside the clause lets a credential behind the title go": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '8db3f5842115', None), + "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family": + _Claim(1, ('family', 'maiden'), + '339279b78f2f', None), + # 2026-09-26: the one-name rule for rules.md#M2's example + # of a link the clause stops at and keeps. + "fix(#397/#535) a link the clause stops at keeps a run that would not read off": + _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), + '1b74e094fed9', None), }, # The 2.3 cycle's first rule, and a facade-only render fix: every # role is identical, so `_initials` alone. Reach and digest as in @@ -4822,14 +5229,30 @@ def _claim(rule: dict) -> _Claim: # the member lands in one, leaves the other and is reported, # so a widening taking any of the three alone would change # the roles here before it reached the gate. + # 2026-09-26, #535: 18 -> 17, 'Jane Doe nee Smith Prof. MA' + # LEAVING this rule's alternation -- the title chain now ends + # the clause before the credential is asked, so the name has + # its own fix(#535) rule instead. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(18, ('_ambiguities', 'maiden', 'suffix'), '3f7d5fd3cc4b', None), - # The declining half: eight corpus names at 2.2.0 and 2.3.0, - # five at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role + _Claim(17, ('_ambiguities', 'maiden', 'suffix'), '6bc9b2ac9772', None), + # The declining half: fourteen corpus names at this baseline, + # eight at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause # gave the member up. + # 2026-09-26: 14 -> 15, 'Jane Doe nee King. ba' -- this + # baseline already splits suffix 'ba' from maiden 'King.', so + # the whole diff is #533's; #535 moves no role here either, + # only the report. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(14, ('_ambiguities',), '24fe6a424eda', None), + _Claim(15, ('_ambiguities',), '410b3a9f7f62', None), + # New rule (#535): one corpus name, 'Doe, Jane nee Smith V, + # PhD' -- after a family comma the given slot reads a lone + # numeral as a suffix only where the given part is the LAST + # comma part (#144), asked of the maiden clause too since this + # review; a third comma part behind it withdraws the release, + # moving `middle` and `maiden`. + "fix(#535) a given-slot numeral with a credential tail stays": + _Claim(1, ('maiden', 'middle'), 'c668f28295a7', None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -4877,9 +5300,71 @@ def _claim(rule: dict) -> _Claim: # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name # rules.md#M2's `deviates: #538` example adds, which this # rule's own alternation now names. Verified name by name. + # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, + # 'Smith, John, PhD née Puig - i Soler', joining the same + # alternation for the same reason. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(3, ('family', 'maiden', 'middle', 'suffix'), - '4de0e7570bd6', ('DEFAULT',)), + _Claim(4, ('family', 'maiden', 'middle', 'suffix'), + '4556aa8f93e5', ('DEFAULT',)), + # New rule (#533/#535): one corpus name, 'Jane Doe nee Smith + # Prof. MA'. #533 already gives 'MA' up to the credential + # reading and reports the fork (`suffix`, `maiden`, + # `_ambiguities`) before #535 exists, verified against the + # parent tree d9d80492 (#296 and #399, the other two-cause + # names at 2.0.0/2.1.0, both predate this baseline and so move + # nothing extra here); #535 then reads the title chain through + # the credential, so the title standing in front of 'MA' + # moves `title` and `maiden` further. + "fix(#533/#535) a title in front of the credential the clause gives up": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + 'cb85a2cbf9f2', None), + # New rule (#535): eight corpus names -- the title-stop + # literal list, minus the one name above whose diff from this + # baseline has an earlier cause too. `title`, `maiden`, + # `suffix` and `_ambiguities` move between them. + "fix(#535) a trailing title ends the maiden clause": + _Claim(9, ('_ambiguities', 'maiden', 'suffix', 'title'), + '0b7fa8ce0b7b', None), + # New rule (#535): one corpus name, 'Berg, abdul nee Smith V', + # the numeral-join literal. #535 is the sole cause here (the + # parent tree d9d80492 already reads it identically to this + # baseline before #535 ran). + "fix(#535) the numeral stop asks the join question": + _Claim(1, ('given', 'maiden'), 'c8d3b254b804', None), + # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe + # nee Smith King.' to the title-stop rule, and the + # two rules below hold the Accepted 'ba' pair, one name each. + "fix(#342/#535) a title behind a credential the clause keeps still leaves it": + _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'middle', + 'title'), 'b8aa7d51a16c', None), + "fix(#342/#533) a title in front of a credential the clause keeps stays in it": + _Claim(1, ('_ambiguities', 'maiden', 'suffix'), 'cb346cb8419a', None), + # 2026-09-26: one-name rules for rules.md#M2's second-round + # Accepted examples (the particle pair, the particle-chain + # title pair, and 'Dr. nee V' at 1.4.0). + "fix(#535) a particle in front of a trailing title stays in the clause": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + 'a1484fe026d1', None), + "fix(#533/#535) a particle behind a trailing title leaves the clause with it": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '97e10b3d2aa8', None), + "fix(#316) a trailing title behind a maiden clause's first suffix word": + _Claim(1, ('family', 'middle', 'suffix', 'title'), + 'd58b847181aa', None), + # 2026-09-26: one-name rules for rules.md#M2's + # Accepted examples (the particle-inside-the-clause pair, and + # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). + "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + '647659c5eafc', None), + "fix(#533/#535) a particle inside the clause lets a credential behind the title go": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '8db3f5842115', None), + # 2026-09-26: the one-name rule for rules.md#M2's example + # of a link the clause stops at and keeps. + "fix(#397/#535) a link the clause stops at keeps a run that would not read off": + _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), + '1b74e094fed9', None), }, "expected_since_2.1.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -5054,8 +5539,11 @@ def _claim(rule: dict) -> _Claim: # lenient post-comma join takes the released member before # assign can read it, so the clause form agrees with the # bare 'Berg, abdul MA'. Reach, not explanation. + # 2026-09-26, #535: 2 -> 3, 'Berg, abdul nee Smith V' joining + # the corpus -- the same bound-given/maiden shape. Reach, not + # explanation. "fix(#411) the bound-given reserve stops counting words the maiden name takes": - _Claim(2, ('given', 'maiden', 'middle'), "b2500b6f6dfc", None), + _Claim(3, ('given', 'maiden', 'middle'), "9305df1c05c5", None), "fix(#412) a connective join no longer absorbs the maiden marker beside it": _Claim(2, ('family', 'maiden'), "51c0eb36b5c5", None), "fix(#418) the connective carve-out counts the name the maiden clause leaves behind": @@ -5096,8 +5584,10 @@ def _claim(rule: dict) -> _Claim: # `middle` left the ROLES in the same edit, and the reach # grew with the rules corpus; the 1.4.0 roster above carries # both, and says the same at the other two baselines. + # 2026-09-26, #535: 13 -> 14, 'Jane Doe nee Prof. Dr.' joined + # the corpus and ends in " Dr." like the rest this rule reaches. "fix(#296) dr is not postnominal vocabulary, so a trailing Dr. is a name word": - _Claim(13, ('family', 'suffix'), "fb9c68f36d0b", None), + _Claim(14, ('family', 'suffix'), "2e229a1fed16", None), "fix(#296) a credential-only comma string reads a name and its postnominal": _Claim(2, ('suffix', 'title'), "3f983ff71dee", None), # 2026-09-18: 18 -> 20. Two new corpus names, 'Smith, MA' and @@ -5318,14 +5808,22 @@ def _claim(rule: dict) -> _Claim: # the member lands in one, leaves the other and is reported, # so a widening taking any of the three alone would change # the roles here before it reached the gate. + # 2026-09-26, #535: 15 -> 14, 'Jane Doe nee Smith Prof. MA' + # LEAVING this rule's alternation -- the title chain now ends + # the clause before the credential is asked, so the name has + # its own fix(#535) rule instead. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(15, ('_ambiguities', 'maiden', 'suffix'), '1b2218cb6296', None), - # The declining half: eight corpus names at 2.2.0 and 2.3.0, - # five at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role + _Claim(14, ('_ambiguities', 'maiden', 'suffix'), 'e5cd5fb89c8b', None), + # The declining half: fourteen corpus names at 2.2.0, eight + # at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause # gave the member up. + # 2026-09-26: 8 -> 9, 'Jane Doe nee King. ba' -- this + # baseline already splits suffix 'ba' from maiden 'King.', so + # the whole diff is #533's; #535 moves no role here either, + # only the report. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(8, ('_ambiguities',), '329a6472b344', None), + _Claim(9, ('_ambiguities',), 'f999caeb63dc', None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -5396,9 +5894,102 @@ def _claim(rule: dict) -> _Claim: # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name # rules.md#M2's `deviates: #538` example adds, which this # rule's own alternation now names. Verified name by name. + # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, + # 'Smith, John, PhD née Puig - i Soler', joining the same + # alternation for the same reason. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(3, ('family', 'maiden', 'middle', 'suffix'), - '4de0e7570bd6', ('DEFAULT',)), + _Claim(4, ('family', 'maiden', 'middle', 'suffix'), + '4556aa8f93e5', ('DEFAULT',)), + # New rule (#399): one corpus name, 'Jane van der Berg nee + # Smith Prof.' -- a sibling of the general #399 rule above, + # for the two-trailing-word shape its own anchor + # (`\\sn[ée]e\\s+\\S+$`) cannot reach: per rules.md#P2, "a + # maiden marker takes the words after it (M2), or the name + # ends", and the chain here swallows both trailing words. Its + # diff from this baseline is #399's alone, verified against + # the parent tree d9d80492: before #535 ran, it already read + # family 'van der Berg', maiden 'Smith Prof.', identically to + # HEAD. + "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words": + _Claim(1, ('family', 'maiden'), '6b2d1fa71195', None), + # New rule (#296/#535): one corpus name, 'Jane Doe nee Prof. + # Dr.'. This baseline still has 'dr' in the suffix vocabulary + # (#296 had not run yet), so it reads suffix 'Dr.'. The parent + # tree d9d80492, before #535, already reads maiden 'Prof. + # Dr.' (dr having left the suffix vocabulary by #296); #535 + # then moves that word on into `title`, returning maiden to + # 'Prof.' -- the same value this baseline reads there, so + # `maiden` does not move net baseline-to-HEAD. + "fix(#296/#535) a trailing title after a dropped postnominal Dr.": + _Claim(1, ('suffix', 'title'), '193350645868', None), + # New rule (#533/#535): one corpus name, 'Jane Doe nee Smith + # Prof. MA'. #533 already gives 'MA' up to the credential + # reading and reports the fork (`suffix`, `maiden`, + # `_ambiguities`) before #535 exists, verified against the + # parent tree d9d80492; #535 then reads the title chain + # through the credential, so the title standing in front of + # 'MA' moves `title` and `maiden` further. + "fix(#533/#535) a title in front of the credential the clause gives up": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + 'cb85a2cbf9f2', None), + # New rule (#535): seven corpus names -- the title-stop + # literal list, minus the three names above whose diff from + # this baseline has an earlier cause too. `title`, `maiden` + # and `suffix` move for the credential-bearing ones, and + # `_ambiguities` for the declined-member one. + "fix(#535) a trailing title ends the maiden clause": + _Claim(8, ('_ambiguities', 'maiden', 'suffix', 'title'), + '3216299d78a6', None), + # New rule (#411/#535): one corpus name, 'Berg, abdul nee + # Smith V', relabelled from a bare '#535' rule (review): #411 + # (the bound-given reserve) already decides `given`/`maiden` + # at this baseline before #535 exists, verified against the + # parent tree d9d80492; #535 then moves the released V from + # given back into maiden. + "fix(#411/#535) the numeral stop asks the join question": + _Claim(1, ('given', 'maiden', 'middle', 'suffix'), + 'c8d3b254b804', None), + # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe + # nee Smith King.' to the title-stop rule, and the + # two rules below hold the Accepted 'ba' pair, one name each. + "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it": + _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'middle', + 'title'), 'b8aa7d51a16c', None), + "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it": + _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), + 'cb346cb8419a', None), + # 2026-09-26: one-name rules for rules.md#M2's second-round + # Accepted examples (the particle pair, the particle-chain + # title pair, and 'Dr. nee V' at 1.4.0). + "fix(#535) a particle in front of a trailing title stays in the clause": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + 'a1484fe026d1', None), + "fix(#533/#535) a particle behind a trailing title leaves the clause with it": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '97e10b3d2aa8', None), + "fix(#316/#399) a trailing title behind a maiden clause's first suffix word": + _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), + 'd58b847181aa', None), + "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps": + _Claim(1, ('family', 'maiden'), + '58cf2330fed6', None), + # 2026-09-26: one-name rules for rules.md#M2's + # Accepted examples (the particle-inside-the-clause pair, and + # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). + "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + '647659c5eafc', None), + "fix(#533/#535) a particle inside the clause lets a credential behind the title go": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '8db3f5842115', None), + "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family": + _Claim(1, ('family', 'maiden'), + '339279b78f2f', None), + # 2026-09-26: the one-name rule for rules.md#M2's example + # of a link the clause stops at and keeps. + "fix(#397/#535) a link the clause stops at keeps a run that would not read off": + _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), + '1b74e094fed9', None), }, "expected_since_2.3.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -5525,14 +6116,30 @@ def _claim(rule: dict) -> _Claim: # the member lands in one, leaves the other and is reported, # so a widening taking any of the three alone would change # the roles here before it reached the gate. + # 2026-09-26, #535: 18 -> 17, 'Jane Doe nee Smith Prof. MA' + # LEAVING this rule's alternation -- the title chain now ends + # the clause before the credential is asked, so the name has + # its own fix(#535) rule instead. + # 2026-09-26: 17 -> 18, 'Jane Doe nee King. ba' -- unlike at + # 2.2.0, this baseline never splits the credential at all, so + # it joins the movers here instead of the declining half + # below; the whole diff is #533's and #535 moves nothing. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(18, ('_ambiguities', 'maiden', 'suffix'), '3f7d5fd3cc4b', None), - # The declining half: eight corpus names at 2.2.0 and 2.3.0, - # five at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role + _Claim(18, ('_ambiguities', 'maiden', 'suffix'), '537d3586cd79', None), + # The declining half: fourteen corpus names at this baseline, + # eight at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause # gave the member up. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(14, ('_ambiguities',), '24fe6a424eda', None), + _Claim(15, ('_ambiguities',), 'e30f2d27ccde', None), + # New rule (#535): one corpus name, 'Doe, Jane nee Smith V, + # PhD' -- after a family comma the given slot reads a lone + # numeral as a suffix only where the given part is the LAST + # comma part (#144), asked of the maiden clause too since this + # review; a third comma part behind it withdraws the release, + # moving `middle` and `maiden`. + "fix(#535) a given-slot numeral with a credential tail stays": + _Claim(1, ('maiden', 'middle'), 'c668f28295a7', None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -5580,9 +6187,60 @@ def _claim(rule: dict) -> _Claim: # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name # rules.md#M2's `deviates: #538` example adds, which this # rule's own alternation now names. Verified name by name. + # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, + # 'Smith, John, PhD née Puig - i Soler', joining the same + # alternation for the same reason. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(3, ('family', 'maiden', 'middle', 'suffix'), - '4de0e7570bd6', ('DEFAULT',)), + _Claim(4, ('family', 'maiden', 'middle', 'suffix'), + '4556aa8f93e5', ('DEFAULT',)), + # New rule (#533/#535): one corpus name, 'Jane Doe nee Smith + # Prof. MA'. #533 already gives 'MA' up to the credential + # reading and reports the fork (`suffix`, `maiden`, + # `_ambiguities`) before #535 exists, verified against the + # parent tree d9d80492 (#296 and #399, the other two-cause + # names at 2.0.0/2.1.0, both predate this baseline and so move + # nothing extra here); #535 then reads the title chain through + # the credential, so the title standing in front of 'MA' + # moves `title` and `maiden` further. + "fix(#533/#535) a title in front of the credential the clause gives up": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + 'cb85a2cbf9f2', None), + # New rule (#535): eight corpus names -- the title-stop + # literal list, minus the one name above whose diff from this + # baseline has an earlier cause too. `title`, `maiden`, + # `suffix` and `_ambiguities` move between them. + "fix(#535) a trailing title ends the maiden clause": + _Claim(10, ('_ambiguities', 'maiden', 'suffix', 'title'), + '67b72e8f082f', None), + # New rule (#535): one corpus name, 'Berg, abdul nee Smith V', + # the numeral-join literal. #535 is the sole cause here (the + # parent tree d9d80492 already reads it identically to this + # baseline before #535 ran). + "fix(#535) the numeral stop asks the join question": + _Claim(1, ('given', 'maiden'), 'c8d3b254b804', None), + # 2026-09-26: one-name rules for rules.md#M2's second-round + # Accepted examples (the particle pair, the particle-chain + # title pair, and 'Dr. nee V' at 1.4.0). + "fix(#535) a particle in front of a trailing title stays in the clause": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + 'a1484fe026d1', None), + "fix(#533/#535) a particle behind a trailing title leaves the clause with it": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '97e10b3d2aa8', None), + # 2026-09-26: one-name rules for rules.md#M2's + # Accepted examples (the particle-inside-the-clause pair, and + # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). + "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": + _Claim(1, ('_ambiguities', 'maiden', 'title'), + '647659c5eafc', None), + "fix(#533/#535) a particle inside the clause lets a credential behind the title go": + _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), + '8db3f5842115', None), + # 2026-09-26: the one-name rule for rules.md#M2's example + # of a link the clause stops at and keeps. + "fix(#397/#535) a link the clause stops at keeps a run that would not read off": + _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), + '1b74e094fed9', None), }, } @@ -7061,7 +7719,12 @@ class _Excluded(NamedTuple): # none of their diff either: the fix(#335/#533) rule carries # it at every baseline that has one, and `absorbed_by` stays # empty, which is the half of this record that matters. - _Excluded(62, "7fdfb86d426e", ()), + # 62 -> 63 for #535's delimited title example, 'Jane Doe (nee + # Smith Prof.)', whose parentheses match this shape too. Same + # non-effect: the fix(#535) title-stop rule carries the name's + # diff at every baseline that has one, and `absorbed_by` stays + # empty. + _Excluded(63, "a2bd839d58dd", ()), } diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index b40d142d..a31a51b8 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -497,6 +497,144 @@ def test_a_word_the_clause_gives_up_lands_in_suffix() -> None: f"name; the rest cannot exercise M2 at all") +def test_a_trailing_title_is_transparent_to_the_maiden_clause() -> None: + """rules.md#H5's transparency, asked where #535 found it broken: + a maiden clause ending ' Prof.' and one ending 'Prof. ' + must read alike in every field but the order of the title words, + wherever a trailing rule reads the clause -- no comma, the part + before a suffix comma, the given part after a family comma -- and + the run behind the title reads as post-nominals, i.e. the clause + gives the credential up, and nothing ahead stands to take the + title (no particle chain, no bound-given join, no name left + empty). The invariant is claimed over the forms below and no + wider. Where the clause KEEPS the credential the spellings differ, + a clause being one contiguous run: 'Doe nee Smith ba Prof.' gives + the title up, 'Doe nee Smith Prof. ba' cannot; and where something + in or ahead of the clause would take the title they differ too + ('Jane van der Berg nee Smith PhD Prof.' against '... Prof. PhD', + 'Jane Doe nee Smith do MA Prof.' against '... do Prof. MA') -- + rules.md#M2's Accepted pairs, which are examples rather than a + complete list. No head or run below has a particle, a bound given + name or an empty name in or ahead of the clause. + + An invariant over two INPUTS (docs/design/AGENTS.md axis 11), so it + consults no rule statement. RECORDED NEGATIVE CONTROL: at e0f1a2fa, + before the walk read the chain, these 36 pairs disagreed on 72 + fields ('Jane Doe nee Smith MA Prof.' read maiden 'Smith MA Prof.' + against 'Smith Prof.' and suffix 'MA'). + """ + forms = ("Jane Doe nee Smith {}", "Doe, Jane nee Smith {}", + "JANE DOE NEE SMITH {}", "jane doe nee smith {}", + "Prof. Jane Doe nee Smith {}", "Jane Doe nee Smith {}, PhD") + runs = ("MA", "V", "PhD", "Jr.", "MA JD", "M.A.") + failures = [] + for form in forms: + for tail in runs: + title = "Prof." + if form == form.upper(): + tail, title = tail.upper(), title.upper() + elif form == form.lower(): + tail, title = tail.lower(), title.lower() + a = parse(form.format(f"{tail} {title}")) + b = parse(form.format(f"{title} {tail}")) + for field in ("given", "middle", "family", "suffix", "maiden"): + if str(getattr(a, field)) != str(getattr(b, field)): + failures.append( + f"{form!r} {tail!r} {field}: " + f"{str(getattr(a, field))!r} vs " + f"{str(getattr(b, field))!r}") + if sorted(str(a.title).split()) != sorted(str(b.title).split()): + failures.append(f"{form!r} {tail!r} title") + assert not failures, "\n".join(failures) + + +def test_a_title_the_clause_gives_up_lands_in_title() -> None: + """rules.md#M2's invariant for the title stop (#535): a + period-marked title word written after the marker ends the parse + in the maiden name or in the title field, never in a name part. + + Over heads reaching all three readers and the guards: no comma + (TRAILING), the given part after a family comma (GIVEN_SLOT), a + tail segment behind a suffix comma ('Jane Doe, PhD', NONE), a bare + title head (the view check), a particle head (the chain), a + bound-given head (P5). No head puts the clause before a family + comma, the other place NONE reads. It holds at e0f1a2fa too, + where the clause kept every title, so its control is a mutation: + RECORDED NEGATIVE CONTROL, re-measured 2026-09-26 after the link + and particle-title bodies joined: the title stop's release check + replaced by `if True:` fails 83 tokens across 78 of these 240 + texts (54 across 49 of the first 160). 240 parses, about 0.03s on + CPython 3.11 (measured the same day). + """ + heads = ("Jane Doe", "Doe, Jane", "Doe, Prof.", "Dr.", "J.", + "Jane van der Berg", "Berg, abdul", "Jane Doe, PhD") + bodies = ("Smith Prof.", "Smith MA Prof.", "Smith Prof. MA", + "Smith V Prof.", "Smith Ma Prof.", "Smith Prof. Dr.", + "Prof.", "Prof. Dr.", "Smith King.", "Smith MA do Prof.", + "Smith i DO Prof.", "Smith i MA Prof.", "Smith St.", + "Smith MA St.", "Smith V St.") + failures = [] + for head in heads: + for body in bodies: + for text in (f"{head} nee {body}", + f"{head} nee {body}".upper()): + name = parse(text) + marker = text.lower().index(" nee ") + 5 + for tok in name.tokens: + if (tok.span is not None and tok.span.start >= marker + and "vocab:title" in tok.tags + and tok.text.endswith(".") + and tok.role not in (Role.MAIDEN, Role.TITLE)): + failures.append( + f"{text!r}: {tok.text!r} -> {tok.role.value}") + assert not failures, "\n".join(failures) + + +def test_a_title_first_word_counts_as_a_word() -> None: + """rules.md#M2's first-word floor, over two INPUTS: the floor is + the title chain's own (`trailing_titles` takes it), so whether the + first word after the marker IS title vocabulary ('King.') or not + ('Smith') must not change what a trailing credential behind it + reads as -- the floor holds either word out of the chain's count, + so the credential is read the same way in both. Both sides must + also take a clause whose first word is that head, so two declined + clauses cannot pass by agreeing about nothing. + + An invariant over two INPUTS (docs/design/AGENTS.md axis 11), so + it consults no rule statement. RECORDED NEGATIVE CONTROL, measured + 2026-09-26 on a copy of this tree with the floor removed entirely + (`tail_reading` handed 1 in `_maiden_take`): every one of the 30 + pairs fails, on the head check alone -- the chain takes 'King.', + the clause declines, and the suffixes still agree (0 of 30 + disagree), which is why the head check is here. This replaces a + figure recorded for a clamp-based variant, which a later attempt + could not reproduce from its description. + """ + forms = ("Jane Doe nee {} {}", "Doe, Jane nee {} {}") + runs = ("ba", "MA", "V", "PhD", "MA JD") + heads = ("King.", "Smith") + failures = [] + for form in forms: + for tail in runs: + texts = [form.format(h, tail) for h in heads] + for variant in (texts, [t.upper() for t in texts], + [t.lower() for t in texts]): + a, b = (parse(t) for t in variant) + if str(a.suffix) != str(b.suffix): + failures.append( + f"{variant[0]!r} vs {variant[1]!r}: " + f"{str(a.suffix)!r} vs {str(b.suffix)!r}") + # both sides take a clause, and its first word is the + # head -- otherwise agreeing suffixes could be two + # declined clauses agreeing about nothing + for name, text in ((a, variant[0]), (b, variant[1])): + head = text.split(" nee ", 1)[-1].split(" NEE ", 1)[-1] + if str(name.maiden).split()[:1] != head.split()[:1]: + failures.append( + f"{text!r}: maiden {str(name.maiden)!r}") + assert not failures, "\n".join(failures) + + def test_no_two_ambiguities_name_the_same_token_span() -> None: """The maiden walk and assign's trailing peel both report at this class, and group's particle-chain emitter stands beside them. None @@ -2580,3 +2718,46 @@ def test_the_case_only_walk_can_fail( failures = _case_only_violations() assert failures, "the case-only walk cannot see a substitution" assert any("'Ph.D.' -> 'PhD'" in line for line in failures), failures + + +def test_a_delimiter_core_reads_as_if_it_were_not_written() -> None: + """#538: under a core-bearing policy, a maiden clause reads exactly + as the same text written without the core. The core is structure + the caller declared, the #206 drop takes a LONE core out of the + output, and + rules.md#M2's link exception asks for a NAME word on each side of + the link -- so the word on a link's side is the one past the core. + + Asked of the `maiden` field only, because that is the field #538 + is about: a core that lands inside a connective run the join + merges ('Puig Dr. i - y Soler') stays in the SUFFIX text under this + policy and the default alike, which is a separate gap. + + RECORDED NEGATIVE CONTROL: at e0f1a2fa, before the core was stepped + over, 6 of these 81 texts disagreed ('Puig Mr. - i Soler' read + maiden 'Puig Mr. i Soler' against 'Puig Mr.', and the same for + 'Puig Dr. - i y Soler', under each of the three heads). + """ + dash = Parser(policy=Policy(extra_suffix_delimiters=frozenset({" - "}))) + heads = ("Smith, John, PhD née", "Smith, John, MD née", + "Doe, Jane, PhD nee") + bodies = ("Puig Mr. i Soler", "Puig i Soler", "Puig Dr. i y Soler", + "Puig i i Soler", "Carod i Rovira Mr.", "Puig Mr. i Dr. Soler", + "Jones Smith i Soler", "Puig y Soler", "Puig Jr. i Soler") + failures = [] + total = 0 + for head in heads: + for body in bodies: + parts = body.split() + # a core in every gap PAST the first word: one between the + # marker and that word is below the clause's bound already + # (test_a_core_between_the_marker_and_the_first_word_is_below_lo) + for gap in range(1, len(parts)): + written = " ".join(parts[:gap] + ["-"] + parts[gap:]) + total += 1 + got = str(dash.parse(f"{head} {written}").maiden) + want = str(dash.parse(f"{head} {body}").maiden) + if got != want: + failures.append(f"{head} {written!r}: {got!r} != {want!r}") + assert total == 81 + assert not failures, "\n".join(failures) diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index ed4b3072..d4ab6602 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -30,6 +30,7 @@ "Berg, abd née Jones" "Berg, abdul V" "Berg, abdul nee Jones MA" +"Berg, abdul nee Smith V" "Berg, abdul van" "Berg, abdul vd" "Carod i" @@ -37,7 +38,14 @@ "Carod y de Rovira i" "Davis Royce, Ed" "Del Toro" +"Doe nee Smith Jr. Prof., Jane" +"Doe nee Smith Prof. ba" +"Doe nee Smith Prof., Jane" +"Doe nee Smith V, Jane" +"Doe nee Smith ba Prof." "Doe, Dr. nee Smith MA" +"Doe, Jane nee Smith St." +"Doe, Jane nee Smith V, PhD" "Doe, Jane, and Jr." "Doe, John DO" "Doe, John DO Ed" @@ -67,6 +75,8 @@ "Dr. abdul salam" "Dr. juan garcia" "Dr. med. univ. Margit Popp, MSc" +"Dr. nee Jones Smith Prof." +"Dr. nee V" "Duke of Edinburgh" "Esq. Smith" "Freiherr von Berg MA" @@ -102,17 +112,29 @@ "Jane (née Jones) Smith" "Jane Doe (nee Smith MA)" "Jane Doe (nee Smith Ma)" +"Jane Doe (nee Smith Prof.)" "Jane Doe (nee Smith) MA" +"Jane Doe nee King." +"Jane Doe nee King. ba" "Jane Doe nee MA" "Jane Doe nee MA Smith" +"Jane Doe nee Prof. Dr." "Jane Doe nee Puig i" "Jane Doe nee Puig i Soler" "Jane Doe nee Smith DO DO" +"Jane Doe nee Smith DO Prof." +"Jane Doe nee Smith King." "Jane Doe nee Smith MA" "Jane Doe nee Smith MA Prof." "Jane Doe nee Smith Ma" +"Jane Doe nee Smith Prof." +"Jane Doe nee Smith Prof. DO" "Jane Doe nee Smith Prof. MA" +"Jane Doe nee Smith V Prof." "Jane Doe nee Smith X.Y.Z." +"Jane Doe nee Smith do MA Prof." +"Jane Doe nee Smith do Prof. MA" +"Jane Doe nee Smith i DO Prof." "Jane Smith (Nee)" "Jane Smith (Nee) (Jones)" "Jane Smith (née Jones)" @@ -127,6 +149,9 @@ "Jane née Jones J. V" "Jane née Jones Smith" "Jane née Jr y Jones" +"Jane van der Berg nee Smith PhD Prof." +"Jane van der Berg nee Smith Prof." +"Jane van der Berg nee Smith Prof. PhD" "Jane van der Berg née" "Jane van der Berg née Jones" "Jane van der Berg née PhD" @@ -288,6 +313,8 @@ "Smith, John Prof." "Smith, John V" "Smith, John V." +"Smith, John, PhD - i Soler" +"Smith, John, PhD née Puig - i Soler" "Smith, John, PhD née Puig Mr. - i Soler" "Smith, John, and" "Smith, Jr." diff --git a/tools/differential/corpus_shapes.jsonl b/tools/differential/corpus_shapes.jsonl index c8f65f2e..7dcb23ed 100644 --- a/tools/differential/corpus_shapes.jsonl +++ b/tools/differential/corpus_shapes.jsonl @@ -38,9 +38,12 @@ {"name": "Jane Doe nee Smith MA y", "shape": 1} {"name": "Jane Doe nee Smith Ma", "shape": 1} {"name": "Jane Doe nee Smith Ma JD", "shape": 1} +{"name": "Jane Doe nee Smith Ma Prof.", "shape": 1} {"name": "Jane Doe nee Smith PhD MA", "shape": 1} +{"name": "Jane Doe nee Smith Prof.", "shape": 1} {"name": "Jane Doe nee Smith Prof. MA", "shape": 1} {"name": "Jane Doe nee Smith V", "shape": 1} +{"name": "Jane Doe nee Smith V Prof.", "shape": 1} {"name": "Jane Doe nee Smith XYZ", "shape": 1} {"name": "Jane Doe nee Smith do", "shape": 1} {"name": "Jane Doe nee Yo-Yo Ma", "shape": 1} @@ -142,6 +145,7 @@ {"name": "Doe, Jane nee Smith DO", "shape": 2} {"name": "Doe, Jane nee Smith Do", "shape": 2} {"name": "Doe, Jane nee Smith MA", "shape": 2} +{"name": "Doe, Jane nee Smith MA Prof.", "shape": 2} {"name": "Doe, Jane nee Smith MA do", "shape": 2} {"name": "Doe, Jane nee Smith Ma", "shape": 2} {"name": "Doe, Jane nee Smith do", "shape": 2} diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 4f6459b4..ed55acb5 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4269,6 +4269,195 @@ issue = "fix(#533) accepted: an unlisted dotted acronym ending a maiden clause l name_regex = "^Jane Doe nee Smith X\\.Y\\.Z\\.$" fields = ["family", "maiden", "middle", "suffix"] +[[change]] +issue = "fix(#274) a trailing title after the marker with nothing else in the name" +# 'Dr. nee Jones Smith Prof.' moves the same fields at this baseline +# whether or not #535 exists: v1 has no maiden reading at all, so +# 'nee' stays given text and 'Jones Smith Prof.' stays split across +# middle/family -- the whole diff is #274's (maiden markers +# consumed), and #535 moves NOTHING here (verified: the parent tree +# d9d80492, before #535, already reads title 'Dr.', maiden 'Jones +# Smith Prof.', identically to HEAD). Not folded into the '#274 +# maiden markers consumed' rule above because that rule's `fields` +# lack `given`, which only this name's shape moves. Only a TITLE +# precedes the marker here, so no name word is left standing ahead of +# it, and it is this +# absence, not the marker's own reading, that this rule's SLOT is +# about. Literal-anchored for the reason fix(#533)'s rules give: that +# SLOT is not a vocabulary a regex could describe more broadly. +name_regex = "^Dr\\. nee Jones Smith Prof\\.$" +fields = ["family", "given", "maiden", "middle"] + +[[change]] +issue = "fix(#274/#296/#535) a trailing title after a dropped postnominal Dr., with no maiden reading either" +# 'Jane Doe nee Prof. Dr.' moves at this baseline for two reasons v1 +# and 2.0.0/2.1.0 both still carry: v1 has no maiden reading at all +# (#274, so it reads middle 'Doe nee' and last 'Prof.' where the +# tree reads family 'Doe' and maiden 'Prof.') AND 'dr' is still suffix +# vocabulary there (#296 had not yet run) -- rules.md#S2: "A trailing +# word of the suffix vocabulary reads as a suffix" -- so this +# baseline reads suffix 'Dr.' where the parent tree d9d80492, before +# #535, already reads maiden 'Prof. Dr.' (dr having left the suffix +# vocabulary by #296, so the clause keeps it as a name word). #535 +# then moves that word on into `title`, returning maiden to 'Prof.'. +# `given` does NOT move (both read 'Jane'). Literal-anchored for the +# reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Prof\\. Dr\\.$" +fields = ["family", "maiden", "middle", "suffix", "title"] + +[[change]] +issue = "fix(#274/#535) a trailing title ends the maiden clause" +# v1 has no maiden reading at all, so every one of these names' diff +# from this baseline is #274's (maiden markers consumed) as well as +# #535's (the title stop) -- verified against the parent tree +# d9d80492: before #535, seven of these eight already differed from +# 1.4.0 on exactly {family, maiden, middle}, and the eighth (the +# family-comma one, 'Doe, Jane nee Smith MA Prof.') on exactly +# {maiden, middle, suffix} instead -- `family` there already matches, +# the comma having fixed it. #535 additionally moves `title` (and +# `suffix` for the credential-bearing ones among the +# seven). rules.md#M2: "Where a trailing rule reads the words, a +# trailing title ends it too" -- the walk reads the end of the name +# through the trailing title chain -- rules.md#H5: "successive +# single words that wear the abbreviation shape and are title +# vocabulary chain into the title from the end" -- so the title +# leaves the clause and the credential or numeral in front of it gets +# the stop it gets with the title absent; the pair 'Smith MA Prof.' / +# 'Smith Prof. MA' now agrees. Literal-anchored for the reason +# fix(#533)'s rules give: the subject is a SLOT, and a regex would +# claim the names whose view check keeps the title as readily as the +# ones that release it. +name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +fields = ["middle", "family", "title", "suffix", "maiden"] + +[[change]] +issue = "fix(#274/#342/#445/#535) a title behind a credential the clause keeps still leaves it" +# rules.md#M2's Accepted transparency boundary, the half that gives +# the title up: the clause keeps 'ba' (one name word left, none to +# spare) and the title behind it still ends the clause, so title +# 'Prof.', family 'Doe', maiden 'Smith ba'. Four causes together, +# checked against the parent tree d9d80492 (family 'Doe', maiden +# 'Smith ba Prof.', no report): before #535, v1 has no maiden reading (#274), #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family; +# #535 then reads the title off the end of the clause. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith ba Prof\\.$" +fields = ["title", "given", "middle", "family", "maiden"] + +[[change]] +issue = "fix(#274/#342/#445/#533) a title in front of a credential the clause keeps stays in it" +# rules.md#M2's Accepted transparency boundary, the half that keeps +# the title: a title in FRONT of the kept 'ba' cannot leave without +# it, so the clause keeps 'Smith Prof. ba' and reports the kept +# credential. The parent tree d9d80492 reads it exactly as HEAD +# does, so #535 moves nothing here: v1 has no maiden reading (#274), #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family, and #533 reports +# the credential the clause keeps. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith Prof\\. ba$" +fields = ["given", "middle", "family", "suffix", "maiden"] + +[[change]] +issue = "fix(#274/#535) a particle in front of a trailing title stays in the clause" +# rules.md#M2's Accepted pair on a released particle: the span 'DO +# Prof.' holds a particle with a title behind it, and P2's chain would +# run on over the title, so the particle's release is withdrawn and +# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' +# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.', so the +# move is #535's together with #274 (v1 had no maiden reading at +# all). Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith DO Prof\\.$" +fields = ["title", "middle", "family", "maiden"] + +[[change]] +issue = "fix(#274/#535) a particle behind a trailing title leaves the clause with it" +# The other half of that pair: written behind the title the particle +# is released with nothing behind it, so the title and 'DO' both leave +# the clause. v1 had no maiden reading (#274); #535 reads the title +# off the clause, as for 'Jane Doe nee Smith Prof. MA', which this +# ledger also labels #274/#535. Literal-anchored for the SLOT reason +# fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. DO$" +fields = ["title", "middle", "family", "maiden"] + +[[change]] +issue = "fix(#274/#316) a trailing title behind a maiden clause's first suffix word" +# rules.md#M2's Accepted pair on a title something ahead would take: +# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' +# stands behind the clause and reads as a trailing title. The parent +# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; +# v1 had no maiden reading (#274) and read no trailing title (#316). +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" +fields = ["title", "middle", "family", "suffix", "maiden"] + +[[change]] +issue = "fix(#274) a numeral straight after the marker stays maiden where nothing reads it as a suffix" +# rules.md#M2's boundary: a numeral straight after the marker declines +# the clause only where the name left standing reads it as a suffix, +# and 'Dr.' alone reads nothing, so the clause keeps 'V'. Every 2.x +# baseline reads it so; v1 had no maiden reading (#274) and read given +# 'nee', family 'V'. Literal-anchored: one corpus name, the doc's own +# example. +name_regex = "^Dr\\. nee V$" +fields = ["given", "family", "maiden"] + +[[change]] +issue = "fix(#274/#535) a particle inside the clause keeps the credential in front of a trailing title" +# rules.md#M2's Accepted pair on a particle INSIDE the clause: the +# particle's chain would take the credential behind it, so the clause +# keeps 'Smith do MA' and only the title behind the credential leaves. +# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no +# report, and v1 had no maiden reading at all (#274), so the move is +# #535's together with #274. Literal-anchored: one corpus name, the +# doc's own example. +name_regex = "^Jane Doe nee Smith do MA Prof\\.$" +fields = ["title", "middle", "family", "maiden"] + +[[change]] +issue = "fix(#274/#535) a particle inside the clause lets a credential behind the title go" +# The other half of that pair: with the title in front of 'MA', the +# credential is released and the title leaves with it, so maiden +# 'Smith do', suffix 'MA', title 'Prof.'. v1 had no maiden reading +# (#274); #535 reads the title off the clause, as for 'Jane Doe nee +# Smith Prof. MA', which this ledger also labels #274/#535. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do Prof\\. MA$" +fields = ["title", "middle", "family", "maiden"] + +[[change]] +issue = "fix(#274/#397/#535) a link the clause stops at keeps a run that would not read off" +# rules.md#M2: a link that ends the clause gives up the words behind +# it only where the name left standing reads them as post-nominals or +# titles. With the title chained the link exception refuses 'i', but +# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps +# the link and the DO (reporting the kept credential) and only 'Prof.' +# leaves. Checked against the parent tree d9d80492, which read maiden +# 'Smith i DO Prof.': v1 had no maiden reading (#274) and no Catalan +# link, 'i' being a plain suffix word, and #397 joins the link inside +# the clause; #535 then reads the title off. Literal-anchored: one +# corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith i DO Prof\\.$" +fields = ["title", "middle", "family", "maiden"] + +[[change]] +issue = "fix(#411/#535) the numeral stop asks the join question" +# 'Berg, abdul nee Smith V' reads IDENTICALLY at this baseline and at +# 2.0.0 -- given 'abdul nee', middle 'Smith', family 'Berg', suffix +# 'V' -- so #274 (v1 has no maiden reading) moves nothing here; this +# name simply never exercises a maiden reading before #411 lands. +# Two reasons together move it from this baseline to HEAD: the +# bound-given reserve (#411) decides how much of the tail the given +# join takes before #535 ever asks anything -- rules.md#P5: "The +# marker and the words it will take are not among the words to +# spare: they leave the name, so counting them asks the question +# about a name that will not exist" -- and the parent tree d9d80492, +# before #535, already reads given 'abdul V', maiden 'Smith', +# differing from this baseline; #535 is what moves the released V +# from given back into maiden (`given` 'abdul V' -> 'abdul', `maiden` +# 'Smith' -> 'Smith V'). Literal-anchored for the same SLOT reason +# fix(#533)'s rules give. +name_regex = "^(?:Berg, abdul nee Smith V)$" +fields = ["given", "middle", "suffix", "maiden"] + [[change]] issue = "fix(#434/#533) a marker PHRASE takes the maiden name, and its clause ends at the credential" # 'Maria Kowalska z domu Nowak MA'. fix(#274)'s regex is the @@ -4290,7 +4479,8 @@ fields = ["family", "maiden", "middle"] # #397 review: A LINK INSIDE A MAIDEN CLAUSE STAYS IN THE BIRTH NAME # (2026-09-20; a third rule added 2026-09-22 with rules.md#M2's -# `deviates: #538` example). Separate rules rather than one, because +# `deviates: #538` example -- a live example since #538, 2026-09-26). +# Separate rules rather than one, because # at this baseline OTHER changes ride along in the same diffs and one # rule has to explain the whole of each: the marker leaving the name # (#274), R1's space-joined rendering of a suffix run v1 wrote with a @@ -4343,11 +4533,11 @@ fields = ["family", "maiden", "middle", "suffix"] [[change]] issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all" # 'Smith, John, PhD née Puig Mr. - i Soler', which arrives from -# rules.md#M2 rather than from a report: it is the doc's -# `deviates: #538` example, and the doc reaches that deviation with a -# configured ' - ' this corpus does not apply. What the gate sees is -# the DEFAULT facade reading, where the dash is an ordinary name word -# like any other. +# rules.md#M2 rather than from a report: it is one of rules.md#M2's +# live examples of the #538 fix (under a configured ' - ' it reads +# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# so what the gate sees is the DEFAULT facade reading, where the dash +# is an ordinary name word like any other. # # v1 had no maiden markers at all, so everything behind the suffix # comma stayed in the suffix: suffix 'PhD née Puig Mr. - i Soler' @@ -4361,10 +4551,18 @@ issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the su # about a FAMILY comma, where v1 left a middle name behind and the # field list says so. This tail leaves v1 a suffix and nothing else. # -# Literal, one name, for the reason those rules give: the shape would +# Literal, two names, for the reason those rules give: the shape would # be "a maiden marker and a bare letter", which reaches every clause -# name in the corpora. -name_regex = "^Smith, John, PhD née Puig Mr\\. - i Soler$" +# name in the corpora. 'Smith, John, PhD née Puig - i Soler' joined +# this alternation with #538 (2026-09-26): it is rules.md#M2's second +# live example, and it reaches the DEFAULT facade the same way its +# neighbour does -- no delimiter configured here either, so the dash +# is an ordinary name word. Measured at this baseline: suffix +# 'PhD née Puig - i Soler' -> 'PhD', maiden '' -> 'Puig - i Soler' -- +# the same shape as the neighbour's own reading with 'Mr.' dropped +# from it; the tree's reading (suffix 'PhD', maiden 'Puig - i Soler') +# is unchanged by the fix, and only the corpus name is new. +name_regex = "^(?:Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" fields = ["maiden", "suffix"] [[change]] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 5888e41d..11f376da 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2818,7 +2818,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # Literal-anchored for the reason fix(#531)'s rules give: the subject # is a SLOT, so a regex for it would claim the names whose WRITING # declines the word as readily as the ones it takes. -name_regex = "^(?:Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" +name_regex = "^(?:Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -2846,16 +2846,21 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # Doe nee Smith Ma', 'Jane Doe nee Yo-Yo Ma', 'Doe, Jane nee Smith # Ma'); a member that is the ONLY word after the marker stays the # maiden name whatever its writing says ('Jane Doe nee MA'); and -# 'Doe, J. nee MA ba' is the CLAMP, whose roles this baseline already -# reads the way the tree does -- maiden 'MA', suffix 'ba' -- so from -# here the clamp is visible only as the report it raises. +# 'Doe, J. nee MA ba' and 'Jane Doe nee King. ba' are the FLOOR's own +# pair, whose roles this baseline already reads the way the tree does +# -- maiden 'MA'/'King.', suffix 'ba' -- so from here the floor is +# visible only as the report it raises. The floor keeps a chained +# first word in the re-peel's count, which the parent's walk (it had +# no chain at all) never removed it from in the first place; so #535 +# moves nothing here either, and the second name's whole diff is +# #533's, exactly as the first's. # # Reporting a DECLINED fork is #530's stated rule -- the report # tracks the fork consulted, not the lean -- and it is why this is a # rule of its own rather than a widening of the one above: `fields` # is `_ambiguities` alone, so it cannot absorb a role diff on any of -# the five. -name_regex = "^(?:Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane nee Smith Ma|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" +# the six. +name_regex = "^(?:Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane nee Smith Ma|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee King\\. ba|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" fields = ["_ambiguities"] [[change]] @@ -2951,6 +2956,215 @@ issue = "fix(#533) restores the suffix reading a 2.4 retag had moved into the ma name_regex = "^John Smith nee Jones R\\.A\\.I\\.$" fields = ["_ambiguities"] +[[change]] +issue = "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words" +# 'Jane van der Berg nee Smith Prof.' moves {family, maiden} at this +# baseline whether or not #535 exists -- verified against the parent +# tree d9d80492, which already reads family 'van der Berg', maiden +# 'Smith Prof.' identically to HEAD, before #535 touched anything. +# It is the #399 rule's own shape: rules.md#P2: "a maiden marker +# takes the words after it (M2), or the name ends" -- but the chain +# here swallows the marker and BOTH trailing words before the maiden +# consumer could see any of them, and the general rule's anchor takes +# exactly one word after the marker (`\\sn[ée]e\\s+\\S+$`), so this +# name's two-word clause needs its own literal rather than widening +# that regex's reach past what its own review measured. #535 moves +# nothing here. +name_regex = "^Jane van der Berg nee Smith Prof\\.$" +fields = ["family", "maiden"] + +[[change]] +issue = "fix(#296/#535) a trailing title after a dropped postnominal Dr." +# 'Jane Doe nee Prof. Dr.' moves {suffix, title} at this baseline for +# two reasons together: this baseline still has 'dr' in the suffix +# vocabulary (#296 had not run yet), so it reads suffix 'Dr.' -- +# rules.md#S2: "A trailing word of the suffix vocabulary reads as a +# suffix" -- where the parent tree d9d80492, before #535, already +# reads maiden 'Prof. Dr.' (dr having left the suffix vocabulary by +# #296, so the clause keeps it as a name word). #535 then moves that +# word on into `title`, returning maiden to 'Prof.' -- the SAME value +# this baseline reads there, which is why `maiden` itself does not +# move net baseline-to-HEAD even though it moves in between. +# Literal-anchored for the reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Prof\\. Dr\\.$" +fields = ["suffix", "title"] + +[[change]] +issue = "fix(#533/#535) a title in front of the credential the clause gives up" +# 'Jane Doe nee Smith Prof. MA' moves at this baseline for two +# reasons together: #533 already gives 'MA' up to the credential +# reading and reports the fork (`suffix='MA'`, `maiden='Smith +# Prof.'`, SUFFIX_OR_NAME -- rules.md#S2's ambiguity flag: "either +# reading carries the ambiguity flag") before #535 exists -- verified +# against the parent tree d9d80492, which already differs from this +# baseline exactly that way. #535 then reads the trailing title chain +# through the credential, so the title standing in front of 'MA' +# leaves the maiden name too (`title='Prof.'`, `maiden='Smith'`). +# Literal-anchored for the reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. MA$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +issue = "fix(#535) a trailing title ends the maiden clause" +# rules.md#M2: "Where a trailing rule reads the words, a trailing +# title ends it too" -- the walk reads the end of the name through +# the trailing title chain -- rules.md#H5: "successive single words +# that wear the abbreviation shape and are title vocabulary chain +# into the title from the end" -- so the title leaves the clause and +# the credential or numeral in front of it gets the stop it gets with +# the title absent; the pair 'Smith MA Prof.' / 'Smith Prof. MA' now +# agrees. Literal-anchored for the reason fix(#533)'s rules give: the +# subject is a SLOT, and a regex would claim the names whose view +# check keeps the title as readily as the ones that release it. Three +# names moved to their own rules above ('Jane van der Berg nee Smith +# Prof.', 'Jane Doe nee Prof. Dr.', 'Jane Doe nee Smith Prof. MA'), +# each because the parent tree d9d80492 already differed from this +# baseline before #535 touched it; the seven names here are ones +# where #535 is the sole cause, verified the same way. +name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +issue = "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it" +# rules.md#M2's Accepted transparency boundary, the half that gives +# the title up: the clause keeps 'ba' (one name word left, none to +# spare) and the title behind it still ends the clause, so title +# 'Prof.', family 'Doe', maiden 'Smith ba'. Three causes together, +# checked against the parent tree d9d80492 (family 'Doe', maiden +# 'Smith ba Prof.', no report): before #535, #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family; +# #535 then reads the title off the end of the clause. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith ba Prof\\.$" +fields = ["title", "given", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it" +# rules.md#M2's Accepted transparency boundary, the half that keeps +# the title: a title in FRONT of the kept 'ba' cannot leave without +# it, so the clause keeps 'Smith Prof. ba' and reports the kept +# credential. The parent tree d9d80492 reads it exactly as HEAD +# does, so #535 moves nothing here: #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family, and #533 reports +# the credential the clause keeps. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith Prof\\. ba$" +fields = ["given", "family", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#535) a particle in front of a trailing title stays in the clause" +# rules.md#M2's Accepted pair on a released particle: the span 'DO +# Prof.' holds a particle with a title behind it, and P2's chain would +# run on over the title, so the particle's release is withdrawn and +# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' +# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this +# baseline does, so #535 is the cause alone. Literal-anchored for the +# SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith DO Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" +# The other half of that pair: written behind the title the particle +# is released with nothing behind it, so the title and 'DO' both leave +# the clause. Two causes together, checked against the parent tree +# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already +# gives 'DO' up and reports the fork; #535 then reads the title in +# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. DO$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#316/#399) a trailing title behind a maiden clause's first suffix word" +# rules.md#M2's Accepted pair on a title something ahead would take: +# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' +# stands behind the clause and reads as a trailing title. The parent +# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; +# two causes, #399 (this baseline's particle chain swallowed the +# marker, middle 'van der Berg nee Smith PhD') and #316 (no trailing +# title reading here, family 'Prof.'). Literal-anchored: one corpus +# name, the doc's own example. +name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" +fields = ["title", "middle", "family", "suffix", "maiden"] + +[[change]] +issue = "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps" +# The other half of that pair: the title in front of 'PhD' stays in +# the clause, the particle chain ahead being able to take it +# (rules.md#H5's 'John van der Berg Prof.'). The parent tree d9d80492 +# reads it as HEAD does, so #535 moves nothing; #399 is the cause here +# -- this baseline's particle chain swallowed the marker and read +# family 'van der Berg nee Smith Prof.'. Literal-anchored: one corpus +# name, the doc's own example. +name_regex = "^Jane van der Berg nee Smith Prof\\. PhD$" +fields = ["family", "maiden"] + +[[change]] +issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" +# rules.md#M2's Accepted pair on a particle INSIDE the clause: the +# particle's chain would take the credential behind it, so the clause +# keeps 'Smith do MA' and only the title behind the credential leaves. +# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no +# report, as this baseline does, so #535 is the whole cause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do MA Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" +# The other half of that pair: with the title in front of 'MA', the +# credential is released and the title leaves with it, so maiden +# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked +# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix +# 'MA', reported): #533 already gives 'MA' up and reports the fork; +# #535 then reads the title in front of it off the clause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do Prof\\. MA$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family" +# rules.md#M2's Accepted row under #548: before a family comma the +# trailing numeral's stop is made over the peel alone, with no release +# question, so a lone numeral behind the clause goes to the family +# ('Doe V'). This baseline read maiden 'Smith V'; 2.2.0, 2.3.0 and the +# parent tree d9d80492 all read family 'Doe V', so the move predates +# #535 and is #424's -- the walk stopping before the trailing numeral +# -- recorded as accepted while #548 is open. Literal-anchored: one +# corpus name, the doc's own example. +name_regex = "^Doe nee Smith V, Jane$" +fields = ["family", "maiden"] + +[[change]] +issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" +# rules.md#M2: a link that ends the clause gives up the words behind +# it only where the name left standing reads them as post-nominals or +# titles. With the title chained the link exception refuses 'i', but +# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps +# the link and the DO (reporting the kept credential) and only 'Prof.' +# leaves. Checked against the parent tree d9d80492, which read maiden +# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word +# that ended the clause, and #397 joins the link inside the clause; +# #535 then reads the title off. Literal-anchored: one corpus name, +# the doc's own example. +name_regex = "^Jane Doe nee Smith i DO Prof\\.$" +fields = ["title", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#411/#535) the numeral stop asks the join question" +# 'Berg, abdul nee Smith V' moves at this baseline for two reasons +# together: #411 (the bound-given reserve) already decides how much +# of the tail the given join takes before #535 exists -- rules.md#P5: +# "The marker and the words it will take are not among the words to +# spare: they leave the name, so counting them asks the question +# about a name that will not exist" -- verified against the parent +# tree d9d80492, which already reads given 'abdul V', maiden 'Smith', +# differing from this baseline's given 'abdul nee', middle 'Smith', +# suffix 'V'. #535 then moves the released V from given back into +# maiden (given 'abdul V' -> 'abdul', maiden 'Smith' -> 'Smith V'). +# Literal-anchored for the same SLOT reason fix(#533)'s rules give. +name_regex = "^(?:Berg, abdul nee Smith V)$" +fields = ["given", "middle", "suffix", "maiden"] + [[change]] issue = "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family" # 'John née Jones Smith MA', whose diff has TWO causes from here and @@ -3235,25 +3449,34 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # twin is the property invariant in tests/v2/test_properties.py. # # The THIRD name arrives from rules.md#M2 rather than from a report: -# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's -# `deviates: #538` example, and the doc reaches the deviation with a -# configured ' - ' that this corpus does not apply. Parsed with the -# DEFAULT facade, as the gate parses it, the dash is an ordinary name -# word, so the link has one on each side and the clause keeps -# 'i Soler' where this baseline left it in the suffix: suffix +# 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's +# live examples of the #538 fix (under a configured ' - ' it reads +# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# so the gate parses it with the DEFAULT facade, where the dash is an +# ordinary name word: the link has one on each side and the clause +# keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). # That is this rule's own sentence and no part of #538, whose reading # needs the delimiter declared; {maiden, suffix} is a subset of the # union above. # -# Literal-anchored to the three. The shape is "a one-letter connective +# Literal-anchored to the four. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" +# +# 'Smith, John, PhD née Puig - i Soler' joined the alternation with +# #538 (2026-09-26): it is rules.md#M2's second live example, and it +# reaches the DEFAULT facade the same way its neighbour does -- no +# delimiter configured here either, so the dash is an ordinary name +# word. Measured at this baseline: suffix 'PhD i Soler' -> 'PhD', +# maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the +# neighbour's own reading with 'Mr.' dropped from it, unchanged by +# the fix; only the corpus name is new. +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 8f7acdaf..1c95431e 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2705,7 +2705,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # Literal-anchored for the reason fix(#531)'s rules give: the subject # is a SLOT, so a regex for it would claim the names whose WRITING # declines the word as readily as the ones it takes. -name_regex = "^(?:Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" +name_regex = "^(?:Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -2733,16 +2733,21 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # Doe nee Smith Ma', 'Jane Doe nee Yo-Yo Ma', 'Doe, Jane nee Smith # Ma'); a member that is the ONLY word after the marker stays the # maiden name whatever its writing says ('Jane Doe nee MA'); and -# 'Doe, J. nee MA ba' is the CLAMP, whose roles this baseline already -# reads the way the tree does -- maiden 'MA', suffix 'ba' -- so from -# here the clamp is visible only as the report it raises. +# 'Doe, J. nee MA ba' and 'Jane Doe nee King. ba' are the FLOOR's own +# pair, whose roles this baseline already reads the way the tree does +# -- maiden 'MA'/'King.', suffix 'ba' -- so from here the floor is +# visible only as the report it raises. The floor keeps a chained +# first word in the re-peel's count, which the parent's walk (it had +# no chain at all) never removed it from in the first place; so #535 +# moves nothing here either, and the second name's whole diff is +# #533's, exactly as the first's. # # Reporting a DECLINED fork is #530's stated rule -- the report # tracks the fork consulted, not the lean -- and it is why this is a # rule of its own rather than a widening of the one above: `fields` # is `_ambiguities` alone, so it cannot absorb a role diff on any of -# the five. -name_regex = "^(?:Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane nee Smith Ma|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" +# the six. +name_regex = "^(?:Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane nee Smith Ma|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee King\\. ba|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" fields = ["_ambiguities"] [[change]] @@ -2959,6 +2964,215 @@ fields = ["_ambiguities", "maiden", "suffix"] # criterion and its measured blast radius. # --------------------------------------------------------------- +[[change]] +issue = "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words" +# 'Jane van der Berg nee Smith Prof.' moves {family, maiden} at this +# baseline whether or not #535 exists -- verified against the parent +# tree d9d80492, which already reads family 'van der Berg', maiden +# 'Smith Prof.' identically to HEAD, before #535 touched anything. +# It is the #399 rule's own shape: rules.md#P2: "a maiden marker +# takes the words after it (M2), or the name ends" -- but the chain +# here swallows the marker and BOTH trailing words before the maiden +# consumer could see any of them, and the general rule's anchor takes +# exactly one word after the marker (`\\sn[ée]e\\s+\\S+$`), so this +# name's two-word clause needs its own literal rather than widening +# that regex's reach past what its own review measured. #535 moves +# nothing here. +name_regex = "^Jane van der Berg nee Smith Prof\\.$" +fields = ["family", "maiden"] + +[[change]] +issue = "fix(#296/#535) a trailing title after a dropped postnominal Dr." +# 'Jane Doe nee Prof. Dr.' moves {suffix, title} at this baseline for +# two reasons together: this baseline still has 'dr' in the suffix +# vocabulary (#296 had not run yet), so it reads suffix 'Dr.' -- +# rules.md#S2: "A trailing word of the suffix vocabulary reads as a +# suffix" -- where the parent tree d9d80492, before #535, already +# reads maiden 'Prof. Dr.' (dr having left the suffix vocabulary by +# #296, so the clause keeps it as a name word). #535 then moves that +# word on into `title`, returning maiden to 'Prof.' -- the SAME value +# this baseline reads there, which is why `maiden` itself does not +# move net baseline-to-HEAD even though it moves in between. +# Literal-anchored for the reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Prof\\. Dr\\.$" +fields = ["suffix", "title"] + +[[change]] +issue = "fix(#533/#535) a title in front of the credential the clause gives up" +# 'Jane Doe nee Smith Prof. MA' moves at this baseline for two +# reasons together: #533 already gives 'MA' up to the credential +# reading and reports the fork (`suffix='MA'`, `maiden='Smith +# Prof.'`, SUFFIX_OR_NAME -- rules.md#S2's ambiguity flag: "either +# reading carries the ambiguity flag") before #535 exists -- verified +# against the parent tree d9d80492, which already differs from this +# baseline exactly that way. #535 then reads the trailing title chain +# through the credential, so the title standing in front of 'MA' +# leaves the maiden name too (`title='Prof.'`, `maiden='Smith'`). +# Literal-anchored for the reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. MA$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +issue = "fix(#535) a trailing title ends the maiden clause" +# rules.md#M2: "Where a trailing rule reads the words, a trailing +# title ends it too" -- the walk reads the end of the name through +# the trailing title chain -- rules.md#H5: "successive single words +# that wear the abbreviation shape and are title vocabulary chain +# into the title from the end" -- so the title leaves the clause and +# the credential or numeral in front of it gets the stop it gets with +# the title absent; the pair 'Smith MA Prof.' / 'Smith Prof. MA' now +# agrees. Literal-anchored for the reason fix(#533)'s rules give: the +# subject is a SLOT, and a regex would claim the names whose view +# check keeps the title as readily as the ones that release it. Three +# names moved to their own rules above ('Jane van der Berg nee Smith +# Prof.', 'Jane Doe nee Prof. Dr.', 'Jane Doe nee Smith Prof. MA'), +# each because the parent tree d9d80492 already differed from this +# baseline before #535 touched it; the seven names here are ones +# where #535 is the sole cause, verified the same way. +name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +issue = "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it" +# rules.md#M2's Accepted transparency boundary, the half that gives +# the title up: the clause keeps 'ba' (one name word left, none to +# spare) and the title behind it still ends the clause, so title +# 'Prof.', family 'Doe', maiden 'Smith ba'. Three causes together, +# checked against the parent tree d9d80492 (family 'Doe', maiden +# 'Smith ba Prof.', no report): before #535, #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family; +# #535 then reads the title off the end of the clause. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith ba Prof\\.$" +fields = ["title", "given", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it" +# rules.md#M2's Accepted transparency boundary, the half that keeps +# the title: a title in FRONT of the kept 'ba' cannot leave without +# it, so the clause keeps 'Smith Prof. ba' and reports the kept +# credential. The parent tree d9d80492 reads it exactly as HEAD +# does, so #535 moves nothing here: #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family, and #533 reports +# the credential the clause keeps. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith Prof\\. ba$" +fields = ["given", "family", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#535) a particle in front of a trailing title stays in the clause" +# rules.md#M2's Accepted pair on a released particle: the span 'DO +# Prof.' holds a particle with a title behind it, and P2's chain would +# run on over the title, so the particle's release is withdrawn and +# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' +# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this +# baseline does, so #535 is the cause alone. Literal-anchored for the +# SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith DO Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" +# The other half of that pair: written behind the title the particle +# is released with nothing behind it, so the title and 'DO' both leave +# the clause. Two causes together, checked against the parent tree +# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already +# gives 'DO' up and reports the fork; #535 then reads the title in +# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. DO$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#316/#399) a trailing title behind a maiden clause's first suffix word" +# rules.md#M2's Accepted pair on a title something ahead would take: +# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' +# stands behind the clause and reads as a trailing title. The parent +# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; +# two causes, #399 (this baseline's particle chain swallowed the +# marker, middle 'van der Berg nee Smith PhD') and #316 (no trailing +# title reading here, family 'Prof.'). Literal-anchored: one corpus +# name, the doc's own example. +name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" +fields = ["title", "middle", "family", "suffix", "maiden"] + +[[change]] +issue = "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps" +# The other half of that pair: the title in front of 'PhD' stays in +# the clause, the particle chain ahead being able to take it +# (rules.md#H5's 'John van der Berg Prof.'). The parent tree d9d80492 +# reads it as HEAD does, so #535 moves nothing; #399 is the cause here +# -- this baseline's particle chain swallowed the marker and read +# family 'van der Berg nee Smith Prof.'. Literal-anchored: one corpus +# name, the doc's own example. +name_regex = "^Jane van der Berg nee Smith Prof\\. PhD$" +fields = ["family", "maiden"] + +[[change]] +issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" +# rules.md#M2's Accepted pair on a particle INSIDE the clause: the +# particle's chain would take the credential behind it, so the clause +# keeps 'Smith do MA' and only the title behind the credential leaves. +# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no +# report, as this baseline does, so #535 is the whole cause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do MA Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" +# The other half of that pair: with the title in front of 'MA', the +# credential is released and the title leaves with it, so maiden +# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked +# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix +# 'MA', reported): #533 already gives 'MA' up and reports the fork; +# #535 then reads the title in front of it off the clause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do Prof\\. MA$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family" +# rules.md#M2's Accepted row under #548: before a family comma the +# trailing numeral's stop is made over the peel alone, with no release +# question, so a lone numeral behind the clause goes to the family +# ('Doe V'). This baseline read maiden 'Smith V'; 2.2.0, 2.3.0 and the +# parent tree d9d80492 all read family 'Doe V', so the move predates +# #535 and is #424's -- the walk stopping before the trailing numeral +# -- recorded as accepted while #548 is open. Literal-anchored: one +# corpus name, the doc's own example. +name_regex = "^Doe nee Smith V, Jane$" +fields = ["family", "maiden"] + +[[change]] +issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" +# rules.md#M2: a link that ends the clause gives up the words behind +# it only where the name left standing reads them as post-nominals or +# titles. With the title chained the link exception refuses 'i', but +# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps +# the link and the DO (reporting the kept credential) and only 'Prof.' +# leaves. Checked against the parent tree d9d80492, which read maiden +# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word +# that ended the clause, and #397 joins the link inside the clause; +# #535 then reads the title off. Literal-anchored: one corpus name, +# the doc's own example. +name_regex = "^Jane Doe nee Smith i DO Prof\\.$" +fields = ["title", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#411/#535) the numeral stop asks the join question" +# 'Berg, abdul nee Smith V' moves at this baseline for two reasons +# together: #411 (the bound-given reserve) already decides how much +# of the tail the given join takes before #535 exists -- rules.md#P5: +# "The marker and the words it will take are not among the words to +# spare: they leave the name, so counting them asks the question +# about a name that will not exist" -- verified against the parent +# tree d9d80492, which already reads given 'abdul V', maiden 'Smith', +# differing from this baseline's given 'abdul nee', middle 'Smith', +# suffix 'V'. #535 then moves the released V from given back into +# maiden (given 'abdul V' -> 'abdul', maiden 'Smith' -> 'Smith V'). +# Literal-anchored for the same SLOT reason fix(#533)'s rules give. +name_regex = "^(?:Berg, abdul nee Smith V)$" +fields = ["given", "middle", "suffix", "maiden"] + [[change]] issue = "fix(#397) the Catalan/Polish link joins two surnames" # 'i' is connective vocabulary since #397, and a connective counts as @@ -3146,25 +3360,34 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # twin is the property invariant in tests/v2/test_properties.py. # # The THIRD name arrives from rules.md#M2 rather than from a report: -# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's -# `deviates: #538` example, and the doc reaches the deviation with a -# configured ' - ' that this corpus does not apply. Parsed with the -# DEFAULT facade, as the gate parses it, the dash is an ordinary name -# word, so the link has one on each side and the clause keeps -# 'i Soler' where this baseline left it in the suffix: suffix +# 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's +# live examples of the #538 fix (under a configured ' - ' it reads +# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# so the gate parses it with the DEFAULT facade, where the dash is an +# ordinary name word: the link has one on each side and the clause +# keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). # That is this rule's own sentence and no part of #538, whose reading # needs the delimiter declared; {maiden, suffix} is a subset of the # union above. # -# Literal-anchored to the three. The shape is "a one-letter connective +# Literal-anchored to the four. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" +# +# 'Smith, John, PhD née Puig - i Soler' joined the alternation with +# #538 (2026-09-26): it is rules.md#M2's second live example, and it +# reaches the DEFAULT facade the same way its neighbour does -- no +# delimiter configured here either, so the dash is an ordinary name +# word. Measured at this baseline: suffix 'PhD i Soler' -> 'PhD', +# maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the +# neighbour's own reading with 'Mr.' dropped from it, unchanged by +# the fix; only the corpus name is new. +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 54cecc2f..30901436 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1282,7 +1282,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # would claim the names whose WRITING declines the word -- which keep # their maiden reading and have the rule below -- as readily as the # ones it takes. _MUST_NOT_MATCH carries both directions. -name_regex = "^(?:Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" +name_regex = "^(?:Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -1315,17 +1315,23 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # so the clause keeps it and the clause's own emitter is what raises # the fork ('JOHN NEE JONES SMITH MA PHD'); and P6's attachment keeps # every `do` spelling the capitals do not take ('Doe, Jane nee Smith -# do', 'Doe, Jane nee Smith MA do'). +# do', 'Doe, Jane nee Smith MA do'); and 'Jane Doe nee King. ba' is +# the FLOOR's own boundary, whose roles this baseline already reads +# the way the tree does -- maiden 'King.', suffix 'ba' -- so from +# here it is visible only as the report it raises. The floor keeps a +# chained first word in the re-peel's count, which the parent's walk +# (it had no chain at all) never removed it from in the first place; +# so #535 moves nothing here either, and the whole diff is #533's. # # Reporting a DECLINED fork is #530's stated rule -- the report # tracks the fork consulted, not the lean -- and it is why this is a # rule of its own rather than a widening of the one above: `fields` # is `_ambiguities` alone, so it cannot absorb a role diff on any of -# the eight. +# the nine. # # Literal-anchored: the class is the slot's declining half, and a # regex for it would claim the seventeen movers above. -name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|JOHN NEE JONES SMITH MA PHD|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" +name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|JOHN NEE JONES SMITH MA PHD|Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" fields = ["_ambiguities"] [[change]] @@ -1410,6 +1416,171 @@ fields = ["_ambiguities", "maiden", "suffix"] # criterion and its measured blast radius. # --------------------------------------------------------------- +[[change]] +issue = "fix(#533/#535) a title in front of the credential the clause gives up" +# 'Jane Doe nee Smith Prof. MA' moves at this baseline for two +# reasons together: #533 already gives 'MA' up to the credential +# reading and reports the fork (`suffix='MA'`, `maiden='Smith +# Prof.'`, SUFFIX_OR_NAME -- rules.md#S2's ambiguity flag: "either +# reading carries the ambiguity flag") before #535 exists -- verified +# against the parent tree d9d80492, which already differs from this +# baseline exactly that way. #535 then reads the trailing title chain +# through the credential, so the title standing in front of 'MA' +# leaves the maiden name too (`title='Prof.'`, `maiden='Smith'`). +# Literal-anchored for the reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. MA$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +issue = "fix(#535) a trailing title ends the maiden clause" +# rules.md#M2: "Where a trailing rule reads the words, a trailing +# title ends it too" -- the walk reads the end of the name through +# the trailing title chain -- rules.md#H5: "successive single words +# that wear the abbreviation shape and are title vocabulary chain +# into the title from the end" -- so the title leaves the clause and +# the credential or numeral in front of it gets the stop it gets with +# the title absent; the pair 'Smith MA Prof.' / 'Smith Prof. MA' now +# agrees. Literal-anchored for the reason fix(#533)'s rules give: the +# subject is a SLOT, and a regex would claim the names whose view +# check keeps the title as readily as the ones that release it. One +# name moved to its own rule above ('Jane Doe nee Smith Prof. MA'), +# because the parent tree d9d80492 already differed from this +# baseline before #535 touched it; the eight names here are ones +# where #535 is the sole cause, verified the same way. ('Jane Doe nee +# Prof. Dr.' and 'Jane van der Berg nee Smith Prof.' need no such +# split at THIS baseline -- #296 and #399 both predate it, so this +# baseline already reads them as HEAD does apart from #535's own move, +# unlike at 2.0.0/2.1.0.) +name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Prof\\. Dr\\.|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +fields = ["title", "maiden", "suffix", "_ambiguities"] + +[[change]] +issue = "fix(#342/#535) a title behind a credential the clause keeps still leaves it" +# rules.md#M2's Accepted transparency boundary, the half that gives +# the title up: the clause keeps 'ba' (one name word left, none to +# spare) and the title behind it still ends the clause, so title +# 'Prof.', family 'Doe', maiden 'Smith ba'. Two causes together, +# checked against the parent tree d9d80492 (family 'Doe', maiden +# 'Smith ba Prof.', no report): before #535, #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word; +# #535 then reads the title off the end of the clause. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith ba Prof\\.$" +fields = ["title", "given", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#342/#533) a title in front of a credential the clause keeps stays in it" +# rules.md#M2's Accepted transparency boundary, the half that keeps +# the title: a title in FRONT of the kept 'ba' cannot leave without +# it, so the clause keeps 'Smith Prof. ba' and reports the kept +# credential. The parent tree d9d80492 reads it exactly as HEAD +# does, so #535 moves nothing here: #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #533 reports +# the credential the clause keeps. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Doe nee Smith Prof\\. ba$" +fields = ["suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#535) a particle in front of a trailing title stays in the clause" +# rules.md#M2's Accepted pair on a released particle: the span 'DO +# Prof.' holds a particle with a title behind it, and P2's chain would +# run on over the title, so the particle's release is withdrawn and +# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' +# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this +# baseline does, so #535 is the cause alone. Literal-anchored for the +# SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith DO Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" +# The other half of that pair: written behind the title the particle +# is released with nothing behind it, so the title and 'DO' both leave +# the clause. Two causes together, checked against the parent tree +# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already +# gives 'DO' up and reports the fork; #535 then reads the title in +# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. DO$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#316) a trailing title behind a maiden clause's first suffix word" +# rules.md#M2's Accepted pair on a title something ahead would take: +# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' +# stands behind the clause and reads as a trailing title. The parent +# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; +# #316 is the whole cause at this baseline, which read middle 'van der +# Berg PhD', family 'Prof.'. Literal-anchored: one corpus name, the +# doc's own example. +name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" +fields = ["title", "middle", "family", "suffix"] + +[[change]] +issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" +# rules.md#M2's Accepted pair on a particle INSIDE the clause: the +# particle's chain would take the credential behind it, so the clause +# keeps 'Smith do MA' and only the title behind the credential leaves. +# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no +# report, as this baseline does, so #535 is the whole cause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do MA Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" +# The other half of that pair: with the title in front of 'MA', the +# credential is released and the title leaves with it, so maiden +# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked +# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix +# 'MA', reported): #533 already gives 'MA' up and reports the fork; +# #535 then reads the title in front of it off the clause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do Prof\\. MA$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" +# rules.md#M2: a link that ends the clause gives up the words behind +# it only where the name left standing reads them as post-nominals or +# titles. With the title chained the link exception refuses 'i', but +# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps +# the link and the DO (reporting the kept credential) and only 'Prof.' +# leaves. Checked against the parent tree d9d80492, which read maiden +# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word +# that ended the clause, and #397 joins the link inside the clause; +# #535 then reads the title off. Literal-anchored: one corpus name, +# the doc's own example. +name_regex = "^Jane Doe nee Smith i DO Prof\\.$" +fields = ["title", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#535) the numeral stop asks the join question" +# rules.md#M2 (#535): the numeral stop now asks the join question +# too -- released, the numeral was taken by the bound-given join +# after a family comma and read as part of the given name +# ('Berg, abdul nee Smith V' read given 'abdul V' at this +# baseline and at the other 2.2/2.3 baseline, before #535 ever ran; +# 2.0/2.1 read given 'abdul nee', suffix 'V' instead -- their own +# copy of this rule is fix(#411/#535) for that reason). +# Literal-anchored for the same SLOT reason fix(#533)'s rules give. +name_regex = "^(?:Berg, abdul nee Smith V)$" +fields = ["given", "maiden"] + +[[change]] +issue = "fix(#535) a given-slot numeral with a credential tail stays" +# 'Doe, Jane nee Smith V, PhD' moves {middle, maiden} at this +# baseline: after a family comma the given slot reads a lone numeral +# as a suffix only where the given part is the LAST comma part +# (#144), asked of the maiden clause too since this review -- a third +# comma part behind it withdraws the release, so the clause keeps 'V' +# where it would otherwise give it up. Verified against the parent +# tree d9d80492, which already reads middle 'V', maiden 'Smith' +# exactly as this baseline does -- an M2 violation #535's own arc had +# carried rather than caused, now fixed. Literal-anchored for the +# same SLOT reason fix(#533)'s rules give. +name_regex = "^Doe, Jane nee Smith V, PhD$" +fields = ["middle", "maiden"] + [[change]] issue = "fix(#397) the Catalan/Polish link joins two surnames" # 'i' is connective vocabulary since #397, and a connective counts as @@ -1601,25 +1772,34 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # twin is the property invariant in tests/v2/test_properties.py. # # The THIRD name arrives from rules.md#M2 rather than from a report: -# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's -# `deviates: #538` example, and the doc reaches the deviation with a -# configured ' - ' that this corpus does not apply. Parsed with the -# DEFAULT facade, as the gate parses it, the dash is an ordinary name -# word, so the link has one on each side and the clause keeps -# 'i Soler' where this baseline left it in the suffix: suffix +# 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's +# live examples of the #538 fix (under a configured ' - ' it reads +# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# so the gate parses it with the DEFAULT facade, where the dash is an +# ordinary name word: the link has one on each side and the clause +# keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). # That is this rule's own sentence and no part of #538, whose reading # needs the delimiter declared; {maiden, suffix} is a subset of the # union above. # -# Literal-anchored to the three. The shape is "a one-letter connective +# Literal-anchored to the four. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" +# +# 'Smith, John, PhD née Puig - i Soler' joined the alternation with +# #538 (2026-09-26): it is rules.md#M2's second live example, and it +# reaches the DEFAULT facade the same way its neighbour does -- no +# delimiter configured here either, so the dash is an ordinary name +# word. Measured at this baseline: suffix 'PhD i Soler' -> 'PhD', +# maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the +# neighbour's own reading with 'Mr.' dropped from it, unchanged by +# the fix; only the corpus name is new. +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 1eb3ec9f..318204a5 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -578,17 +578,20 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # `do` collision, where the capitals decide ('Doe, Jane nee Smith DO' # beside 'Jane Doe nee Smith do', whose comma-less spelling P6 never # reaches); the clause that leaves ONE name word, which rules.md#M4 -# makes the family ('John née Jones Smith MA'); and the CLAMP, where +# makes the family ('John née Jones Smith MA'); and the FLOOR, where # the member is the only word the marker would otherwise leave and # only what stands BEHIND it is given up ('Doe, J. nee MA ba' keeps -# maiden 'MA' and reads suffix 'ba'). +# maiden 'MA' and reads suffix 'ba'; 'Jane Doe nee King. ba' is the +# same shape with a title-vocabulary first word -- this baseline's +# own reading of it never split the credential at all, so the whole +# diff is #533's and #535 moves nothing here). # # Literal-anchored for the reason fix(#531)'s rules give, and it is # the same reason here: the subject is a SLOT, so a regex for it # would claim the names whose WRITING declines the word -- which keep # their maiden reading and have the rule below -- as readily as the # ones it takes. _MUST_NOT_MATCH carries both directions. -name_regex = "^(?:Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" +name_regex = "^(?:Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee King\\. ba|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -629,9 +632,14 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # is `_ambiguities` alone, so it cannot absorb a role diff on any of # the eight. # +# 'Doe nee Smith Prof. ba' (a rules.md#M2 Accepted example) is +# the same declining half: one name word is left, so the clause +# keeps 'ba' and reports it, and the title in front of it cannot +# leave without it -- the parent tree d9d80492 reads it as HEAD +# does, so #535 moves nothing. # Literal-anchored: the class is the slot's declining half, and a # regex for it would claim the seventeen movers above. -name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|JOHN NEE JONES SMITH MA PHD|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" +name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe nee Smith Prof\\. ba|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|JOHN NEE JONES SMITH MA PHD|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" fields = ["_ambiguities"] [[change]] @@ -716,6 +724,138 @@ fields = ["_ambiguities", "maiden", "suffix"] # criterion and its measured blast radius. # --------------------------------------------------------------- +[[change]] +issue = "fix(#533/#535) a title in front of the credential the clause gives up" +# 'Jane Doe nee Smith Prof. MA' moves at this baseline for two +# reasons together: #533 already gives 'MA' up to the credential +# reading and reports the fork (`suffix='MA'`, `maiden='Smith +# Prof.'`, SUFFIX_OR_NAME -- rules.md#S2's ambiguity flag: "either +# reading carries the ambiguity flag") before #535 exists -- verified +# against the parent tree d9d80492, which already differs from this +# baseline exactly that way. #535 then reads the trailing title chain +# through the credential, so the title standing in front of 'MA' +# leaves the maiden name too (`title='Prof.'`, `maiden='Smith'`). +# Literal-anchored for the reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. MA$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +issue = "fix(#535) a trailing title ends the maiden clause" +# rules.md#M2: "Where a trailing rule reads the words, a trailing +# title ends it too" -- the walk reads the end of the name through +# the trailing title chain -- rules.md#H5: "successive single words +# that wear the abbreviation shape and are title vocabulary chain +# into the title from the end" -- so the title leaves the clause and +# the credential or numeral in front of it gets the stop it gets with +# the title absent; the pair 'Smith MA Prof.' / 'Smith Prof. MA' now +# agrees. Literal-anchored for the reason fix(#533)'s rules give: the +# subject is a SLOT, and a regex would claim the names whose view +# check keeps the title as readily as the ones that release it. One +# name moved to its own rule above ('Jane Doe nee Smith Prof. MA'), +# because the parent tree d9d80492 already differed from this +# baseline before #535 touched it; the ten names here are ones +# where #535 is the sole cause, verified the same way. Two of them +# are rules.md#M2 Accepted examples: 'Jane Doe nee Smith King.', +# H5's ordinary-surname reach at the end of a clause, and 'Doe nee +# Smith ba Prof.', the title leaving from behind a credential the +# clause keeps (the parent tree reads maiden 'Smith ba Prof.'). ('Jane Doe nee +# Prof. Dr.' and 'Jane van der Berg nee Smith Prof.' need no such +# split at THIS baseline -- #296 and #399 both predate it, so this +# baseline already reads them as HEAD does apart from #535's own move, +# unlike at 2.0.0/2.1.0.) +name_regex = "^(?:Doe nee Smith ba Prof\\.|Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Prof\\. Dr\\.|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +fields = ["title", "maiden", "suffix", "_ambiguities"] + +[[change]] +issue = "fix(#535) a particle in front of a trailing title stays in the clause" +# rules.md#M2's Accepted pair on a released particle: the span 'DO +# Prof.' holds a particle with a title behind it, and P2's chain would +# run on over the title, so the particle's release is withdrawn and +# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' +# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this +# baseline does, so #535 is the cause alone. Literal-anchored for the +# SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith DO Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" +# The other half of that pair: written behind the title the particle +# is released with nothing behind it, so the title and 'DO' both leave +# the clause. Two causes together, checked against the parent tree +# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already +# gives 'DO' up and reports the fork; #535 then reads the title in +# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. +# Literal-anchored for the SLOT reason fix(#533)'s rules give. +name_regex = "^Jane Doe nee Smith Prof\\. DO$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" +# rules.md#M2's Accepted pair on a particle INSIDE the clause: the +# particle's chain would take the credential behind it, so the clause +# keeps 'Smith do MA' and only the title behind the credential leaves. +# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no +# report, as this baseline does, so #535 is the whole cause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do MA Prof\\.$" +fields = ["title", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" +# The other half of that pair: with the title in front of 'MA', the +# credential is released and the title leaves with it, so maiden +# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked +# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix +# 'MA', reported): #533 already gives 'MA' up and reports the fork; +# #535 then reads the title in front of it off the clause. +# Literal-anchored: one corpus name, the doc's own example. +name_regex = "^Jane Doe nee Smith do Prof\\. MA$" +fields = ["title", "suffix", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" +# rules.md#M2: a link that ends the clause gives up the words behind +# it only where the name left standing reads them as post-nominals or +# titles. With the title chained the link exception refuses 'i', but +# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps +# the link and the DO (reporting the kept credential) and only 'Prof.' +# leaves. Checked against the parent tree d9d80492, which read maiden +# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word +# that ended the clause, and #397 joins the link inside the clause; +# #535 then reads the title off. Literal-anchored: one corpus name, +# the doc's own example. +name_regex = "^Jane Doe nee Smith i DO Prof\\.$" +fields = ["title", "middle", "family", "maiden", "_ambiguities"] + +[[change]] +issue = "fix(#535) the numeral stop asks the join question" +# rules.md#M2 (#535): the numeral stop now asks the join question +# too -- released, the numeral was taken by the bound-given join +# after a family comma and read as part of the given name +# ('Berg, abdul nee Smith V' read given 'abdul V' at this +# baseline and at the other 2.2/2.3 baseline, before #535 ever ran; +# 2.0/2.1 read given 'abdul nee', suffix 'V' instead -- their own +# copy of this rule is fix(#411/#535) for that reason). +# Literal-anchored for the same SLOT reason fix(#533)'s rules give. +name_regex = "^(?:Berg, abdul nee Smith V)$" +fields = ["given", "maiden"] + +[[change]] +issue = "fix(#535) a given-slot numeral with a credential tail stays" +# 'Doe, Jane nee Smith V, PhD' moves {middle, maiden} at this +# baseline: after a family comma the given slot reads a lone numeral +# as a suffix only where the given part is the LAST comma part +# (#144), asked of the maiden clause too since this review -- a third +# comma part behind it withdraws the release, so the clause keeps 'V' +# where it would otherwise give it up. Verified against the parent +# tree d9d80492, which already reads middle 'V', maiden 'Smith' +# exactly as this baseline does -- an M2 violation #535's own arc had +# carried rather than caused, now fixed. Literal-anchored for the +# same SLOT reason fix(#533)'s rules give. +name_regex = "^Doe, Jane nee Smith V, PhD$" +fields = ["middle", "maiden"] + [[change]] issue = "fix(#397) the Catalan/Polish link joins two surnames" # 'i' is connective vocabulary since #397, and a connective counts as @@ -918,25 +1058,34 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # twin is the property invariant in tests/v2/test_properties.py. # # The THIRD name arrives from rules.md#M2 rather than from a report: -# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's -# `deviates: #538` example, and the doc reaches the deviation with a -# configured ' - ' that this corpus does not apply. Parsed with the -# DEFAULT facade, as the gate parses it, the dash is an ordinary name -# word, so the link has one on each side and the clause keeps -# 'i Soler' where this baseline left it in the suffix: suffix +# 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's +# live examples of the #538 fix (under a configured ' - ' it reads +# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# so the gate parses it with the DEFAULT facade, where the dash is an +# ordinary name word: the link has one on each side and the clause +# keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). # That is this rule's own sentence and no part of #538, whose reading # needs the delimiter declared; {maiden, suffix} is a subset of the # union above. # -# Literal-anchored to the three. The shape is "a one-letter connective +# Literal-anchored to the four. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" +# +# 'Smith, John, PhD née Puig - i Soler' joined the alternation with +# #538 (2026-09-26): it is rules.md#M2's second live example, and it +# reaches the DEFAULT facade the same way its neighbour does -- no +# delimiter configured here either, so the dash is an ordinary name +# word. Measured at this baseline: suffix 'PhD i Soler' -> 'PhD', +# maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the +# neighbour's own reading with 'Mr.' dropped from it, unchanged by +# the fix; only the corpus name is new. +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"]