Skip to content

Recover PDF heading structure for the bill ledger, measured by agreement with the XML twin - #734

Open
mattzamora wants to merge 9 commits into
civictechdc:developfrom
mattzamora:feat/524-pdf-section-convergence
Open

mattzamora wants to merge 9 commits into
civictechdc:developfrom
mattzamora:feat/524-pdf-section-convergence

Conversation

@mattzamora

@mattzamora mattzamora commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Related issue

Closes #732
Closes #524

Advances:

Related: #551 (epic: PDF heading detection rests on layered geometry rules), #198 (font name as a
margin discriminator), #557 (a display-only heading change can alter the structure tree untested),
#552 (epic: levels and addresses derived from display labels), #706 (XML and PDF summaries report
different change categories for the same bill).

What does this change?

The PDF reader now recovers a bill's headings much closer to how the bill's XML labels them. It
reads each detected heading together with the lines around it, instead of one printed line at a
time, and corrects five misreadings:

  • A long account name wrapped onto a second line was read as two headings:
    FAMILY HOUSING OPERATION AND MAINTENANCE, / ARMY on 118-hr-4366 made the first half an
    agency, and each later family-housing account was nested under it.
  • Two separate headings printed one above the other were sometimes read as one.
  • A capitalized line inside quoted text of another law, or one that finishes a sentence, was
    read as a heading of the bill.
  • A department printed partway through a title was missed.
  • An agency heading was applied to accounts after it that are not its own.

Where the print does not show the evidence, the original reading is kept.

Accuracy is checked two ways, both with the bill's XML as the reference:

  • Headings. The PDF's account headings are matched against the XML's. Across the nine bills
    of the account-precision check, the lowest recall / precision goes from 0.744 / 0.750 to
    0.967 / 0.967, with seven of the nine at 1.000 / 1.000. On 118-hr-4820, PDF account headings
    with no match in the XML go from 17 to 0.
  • Where dollar amounts sit. Each amount's chain of headings in the PDF is compared with its
    chain in the XML (the ledger, table below). The same chain: 62.3% → 87.6% on the committed
    corpus, and 39.5% → 74.2% on four bills never used to design the rules.

One earlier assumption this changes. ADR 0012 and docs/bill-structure.md recorded that an
agency heading followed by prose looks identical to an account by size, case and position, and
accepted that as a gap. The letters do differ: GPO sets an agency in title-case small caps (each
word's first letter larger) and an account in even small caps. This PR reads that pattern and
uses it to decide one heading versus two and to stop an agency at an account that cannot be its
own. It does not re-label such an agency, which is still reported as an account, so the gap is
narrowed, not closed. Both documents are updated.

Where the change lives, for reviewers reading the code:

  • parsers/pdf_heading_passes.py (new): converge_headings, ordered fail-closed passes that
    re-read the detected headings against the whole reading order, after detection and before
    divisions are assigned. Order and rationale: ADR 0022 (Proposed).
  • parsers/pdf_text.py: one new signal from the existing glyph walk, no new PDFium calls.
    LineGeom.initial_caps is the letters' case pattern (GPO sets an agency in title case in small
    caps, an account in even small caps; a wrapped name's two lines print alike), and
    size_min/size_max the letter size range (capitals at body size are the department style).
    Prose lines are not judged.
  • parsers/pdf_anchors.py: Anchor.caps records the case pattern; _breadcrumb_core lets a
    department end the agency above it, and _agency_reaches stops a carried-over agency at an
    account that prints like it cannot be its child.
  • parsers/pdf_heading_exceptions.py (new): two named wordings decide a line break where
    format cannot (SALARIES AND EXPENSES on its own line; the fifteen executive departments). It
    is the one heading module on the ADR 0018 vocabulary-gate allowlist, and ADR 0018 is rewritten
    to admit it and nothing else.
  • tests/test_pdf_ledger_location.py + tests/ledger_location.py (new): the ledger gate. The
    ledger is every dollar amount in a bill with the breadcrumb it sits under; each amount in
    the XML twin's ledger is aligned with the PDF's and graded T0 (same location) … T4 (different
    heading), pinned exactly per version. python -m tests.ledger_location prints the totals for
    whichever parser is importable, which is how the "before" column below was produced.

Evidence

Ledger, measured by the same scorer on develop (c636448) and on this branch:

corpus before corpus after holdout before holdout after
same location (T0) 62.3% 87.6% 39.5% 74.2%
OK: same, label-only, or true but shallower (T0–T2) 72.1% 92.7% 51.2% 82.9%
wrong parent (T3) 3,648 944 1,064 376
different heading (T4) 99 41 45 12
  • Corpus: the 33 committed dual-format versions whose XML carries appropriations headings
    (13,422 amounts), as python -m tests.ledger_location prints them.
  • Holdout: four FY2024 House bills fetched after the rules were frozen and never used to design
    them: 118-hr-4368, 118-hr-4394, 118-hr-4665, 118-hr-4821 (10 versions, 2,272 amounts). Not
    committed, so they stay unseen for the follow-up that broadens the corpus.
  • Reconciliation (119-hr-1, 117-hr-5376, 117-hr-1319, 115-hr-1; 18 versions, 10,140 amounts):
    different heading 84 → 4.
  • No version got worse on either measure: run on develop, the report finds 0 of the 42
    pinned versions better than this branch; 0 of 10 holdout and 0 of 17 fetched reconciliation
    versions either.

Why OK stops near 93% rather than 99%. The answer key is DeltaTrack's existing XML reader,
which leaves a heading with no text of its own out of the breadcrumb (Food and Drug Administration under Department of Health and Human Services). The PDF now keeps it, and is
graded "wrong parent" where it matches the page: of the 944 wrong-parent amounts, up to 892 have
a PDF parent the XML file tags as a heading and the reader drops. This PR leaves the reader alone
and says so in ADR 0022; changing it is a separate decision (#733, the XML reader drops a heading that carries no text of its own).

#524's verification, criterion by criterion:

  1. Never fabricate hierarchy. NATIONAL OCEANIC AND ATMOSPHERIC ADMINISTRATION on
    114-hr-2578 (fetched; not in the corpus) stays an agency over its own five accounts in all four
    versions, before and after. Corpus-wide, no version's not-tolerated count rose.
  2. Join where the print carries the evidence. 118-hr-4820: all fourteen wrapped names listed on
    A wrapped account heading in a PDF is read as an agency plus an account #524 are read as one account each, including the RAILROAD REHABILITATION AND IMPROVEMENT FINANCING PROGRAM case A wrapped account heading in a PDF is read as an agency plus an account #524 had set aside. Unmatched PDF account headings 17 → 0.
  3. Decline where it is absent. An unknown case pattern keeps the detectors' reading
    (TestFailsClosed).

The account-precision floors #524 names are re-derived from this parser revision:
TestCorpusAccountPrecision's lowest recall/precision goes 0.744 / 0.750 → 0.967 / 0.967 (seven
of nine bills at 1.000 / 1.000), and the floors go 0.70 → 0.95.

#648 (advanced, not closed). Renumbered cards on the 17 accepted corpus pairs: 156 → 143.
Fixed: GINIA.— → WEST VIRGINIA.— and about a dozen other wrap-only renames. Kept, as #648
requires: the four genuine SEC. renumberings in 115-hr-5895 and the genuine heading edit
(SECURITY → SECRUITY). Not fixed: …NAVY AND MARINE CORPS still reads as AND MARINE CORPS on
the Senate strike-all print, and formatters/canonical._pdf_move, the second site #648 names, is
untouched.

#500 (an account printed directly under a section catchline is dropped) is not affected: its
strict expected failure still fails.

Pins regenerated, each with its cause

  • tests/data/pdf_canonical_baseline.json: change counts mostly fall as wrapped names become one
    block (117-hr-4502 v1→v2 1500→1456); 118-hr-8752 v1→v2 37→38 (next bullet).
  • tests/test_pipeline_parity.py + ADR 0014 table: 118-hr-8752 PDF 37→38. The engrossed version
    adds a SPENDING REDUCTION ACCOUNT heading the PDF now reads as its own block, and SEC.
    552–567 file under it as they do in the XML (which also keeps GENERAL PROVISIONS above it).
    117-hr-4502 gap to the XML +85→+41, band unchanged.
  • examples/hr8752_pdf_diff.html and web/webapp/sample/example.html: the same 8752 change,
    and nothing else (the breadcrumb headings of SEC. 552–567, plus the added heading block).
  • tests/data/round1_pairing_sentinel.json: only parser_revision re-stamped (ADR 0019); no
    pairing digest moved.
  • tests/data/pdf/anchors_golden/: 118-s-4795, seven wrapped account names are now one account
    each and one fragment is now its full name; the frozen .pre-agency-anchors baseline is
    re-anchored for exactly those eight. major_vocab.json: additions only (mid-title
    departments).
  • tests/test_pdf_round1_revocation.py: split population 224+6 → 225+6 (three pairs' block
    boundaries moved; per-pair deltas in the docstring).
  • tests/test_pdf_anchor_golden.py: account-precision floors 0.70 → 0.95 (above).
  • tests/test_pdf_size_detection.py: synthetic paragraphs end with a period before a heading, as
    GPO prints them; a line following unfinished prose is never a heading. What each test asserts
    is unchanged.

Known residuals (ADR 0022, Consequences)

  • Two same-styled headings stacked with nothing between them are read as one
    (ATOMIC ENERGY DEFENSE ACTIVITIES over NATIONAL NUCLEAR SECURITY ADMINISTRATION, Two stacked department headings merge into one when the upper line nearly fills the column #501): the
    print carries no format signal there.
  • A department name the major detector itself glues (OVERSEAS CONTINGENCY OPERATIONS DEPARTMENT OF DEFENSE) is left as detected.
  • A wrapped account name opening with the word TITLE stays split (115-hr-5895).
  • Reconciliation prints carry no subtitle/part level to read, so their amounts are at best true but
    shallower.

Docs

ADR 0022 (new, Proposed) + index row; ADRs 0018, 0012 and 0014 rewritten in place;
docs/bill-structure.md (the ledger, the heading forms read across lines, the case-pattern
finding); docs/source-signal-inventory.md (the new signal); TESTING.md (accuracy check 7 and
the ledger pin, with the before/after commands).

For reviewers

  • Where to look hardest: _segment and _agency_reaches. Each rule has a test that fails when
    the rule is removed (28 mutations, all caught).
  • Named exceptions. They are the one place wording decides structure. Measured against the rest
    of this change, the SALARIES AND EXPENSES rule moves 15 amounts (9 to "wrong parent" because
    the heading it separates is one the XML reader drops, 6 to "shallower"); the departments rule
    moves none and corrects heading names. To drop them: delete the module, its allowlist line and
    the named_split call, then regenerate the pins.
  • ADR 0012 wording. Besides the case-pattern note, "body-size" → "heading-band" for the
    prose-leading agency line, per the ADR's own spike and bill-structure.md.
  • Overlaps with open work. Draft fix(pdf): rejoin every printed word break, deciding the hyphen from the document (#650) #682 (rejoin every printed word break) edits pdf_text._attach_geometry and regenerates the same
    three pins (canonical baseline, sentinel, revocation). Whichever lands second regenerates them.
    If Widen the skip ceiling to the whole suite, fix five CWD-dependent defects, add a cwd-independence CI job #721 (skip ceiling for the whole suite) lands first, the ledger gate's update-mode skip needs
    the same allowlist entry as test_pdf_canonical_baseline's.
  • Cost. On the largest numbered print (117-hr-4502 v2, 1,040 pages), extraction plus anchors
    takes about 0.8 s longer (≈18.4 s → ≈19.2 s, mean of three runs; the machine's run-to-run spread
    is about 1 s).
  • Pre-existing, not changed here: ADR 0014's table lists 115-hr-5895 XML = 300; develop
    measures 302.

How to test

uv run ruff check . && uv run ruff format --check .
uv run pytest -m "not slow and not browser"
uv run pytest -m browser --run-browser
uv run pytest -m slow --deselect tests/test_govinfo_corpus_parity.py
uv run pytest tests/test_pdf_heading_passes.py -v          # one test per rule

# The ledger, before and after, from the same scorer
git archive origin/develop src | tar -x -C /tmp/before
PYTHONPATH=/tmp/before/src uv run python -m tests.ledger_location
uv run python -m tests.ledger_location

To see it on a page: 118-hr-4366/1_reported-in-house.pdf, p. 10 l. 13–14, FAMILY HOUSING OPERATION AND MAINTENANCE, / ARMY. On develop the drill-down shows … › FAMILY HOUSING OPERATION AND MAINTENANCE, › ARMY, and each later family-housing account inherits the fragment;
on this branch it is one account under DEPARTMENT OF DEFENSE.

Ran locally on Windows at the head of this branch: ruff check and format clean; fast 2,083
passed; browser 41 passed; slow (every CI slow gate) 1,755 passed, 0 failed. The fast tier's 16
failures are this machine's, not the change's: missing executable bits, CRLF line endings in the
committed renders, and CI-report checks. The same 16 fail on a clean develop checkout, and each
of the first two commits, applied alone on top of develop, gives the same fast result and a
green slow tier.

Checklist

  • Linked the issue above (Closes #...)
  • Ran the CI gates locally and they pass (see What CI checks)
  • New or changed behavior has tests
  • For a bug fix: the test fails without the fix, and I ran it both ways to check
  • Disclosed AI assistance below, if any

AI assistance

Claude Code (Claude Opus 5.5) wrote the code, tests and docs under the contributor's direction. The
design came out of an investigation reviewed case by case against the printed pages and the XML.

🤖 Generated with Claude Code

Ordered, fail-closed passes re-read detected headings across the reading order. pdf_text records each line's letter case pattern; pdf_anchors uses it to bound an agency's reach. Pins re-derived for the new parser revision.
Grades every amount T0-T4 against the XML twin and pins the tiers per version. python -m tests.ledger_location prints before/after totals.
ADR 0022 (Proposed). ADR 0018 admits named exceptions; ADR 0012 notes the case pattern. bill-structure.md, signal inventory and TESTING.md updated.

@willhea willhea left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @mattzamora! I love this line of work. I reviewed with GPT Sol 6.1 and arrived at the following. I think these changes will get this PR good to go. I'm reluctant to amend the ADR for a list of keywords. I would rather err on the side of preserving the existing reading when the printed evidence is ambiguous, even if that leaves some wrapped headings split. Recovering fewer headings is an acceptable tradeoff if it avoids inferring hierarchy from keyword matches.

Reviewed at e5c5b6e. I’d like to retain the typography approach, with three changes before merging:

  1. Keep ADR 0018 intact. Please remove the named wording exceptions and their vocabulary-gate exemption. I’m concerned that these encode answers to known examples rather than establish that the printed evidence distinguishes the structures. Where format is insufficient, preserve the existing reading. We should accept some missed joins rather than invent hierarchy.

  2. Fix the sentence filter deleting real headings. In 118-s-4928/1_reported-in-senate.pdf, page 117, the preceding paragraph ends with (Public Law 111–241) without a period in the actual print. The new filter removes OFFICE OF INSPECTOR GENERAL and SALARIES AND EXPENSES, filing $274 million under the preceding PAYMENT TO THE POSTAL SERVICE FUND account. Missing terminal punctuation alone should not override clear heading evidence. Please add a small behavioral regression test for this case while retaining the control that rejects heading-shaped text embedded in prose.

  3. Correct the ledger scorer before regenerating the results. tier() can award T0 when the account and immediate parent match despite wrong ancestors above them. Changing three Tax Court amounts to a fabricated division, title, and department left every tier count unchanged. T2 also uses sets, so reordered ancestors can be accepted as shallower. T0 should require identical normalized paths; T2 should require an ordered ancestry with only omissions. Wrong or reordered ancestors must be rejected.

Please demonstrate that the regression tests fail under the corresponding faulty behavior, then regenerate the affected measurements and baselines after these changes.

I also ran three additional pairs—119-H.R.3944, 119-H.R.1968, and 118-H.R.815—as PDF→PDF and XML→XML. The core smoke checks passed, and disabling the wording exceptions produced byte-identical PDF canonical outputs to this PR on all three pairs. That supports removing the exceptions, although the smoke test does not establish heading or matching accuracy.

@mattzamora

Copy link
Copy Markdown
Contributor Author

Thanks @willhea, this was a careful review, and all three points held up when I checked them. They are addressed in five commits on top of e5c5b6e, ending at e3d3926:

Commit What it does
8621d8a Point 1: drop the named exceptions, restore ADR 0018
18a1be3 Point 2, plus the bugs it uncovered
6fa6cbd One follow-on level rule (narrows ADR 0012; see below)
7680e3f Point 3: the whole-path scorer, plus a per-amount check
e3d3926 Regenerated baselines and measurements, and the open issues

I agree with the principle you set: where the print is ambiguous, keep the existing reading and accept fewer recovered headings rather than infer hierarchy from wording. Every rule added below reads the print or grammar, never appropriations vocabulary. The one ADR amended, 0012, is amended on a fact about the text, not a keyword list; details under "Also new".

1. ADR 0018 kept intact

The named wording exceptions are gone: the module, and its vocabulary-gate exemption. ADR 0018 and tests/test_structure_vocabulary_gate.py are byte-identical to develop again.

TestNoWordingDecidesALineBreak pins the behaviour you asked for: two stacked lines printed alike stay one heading, whatever they say. A familiar sub-account or department name no longer splits them.

What it costs, measured on 118 PDFs (53 committed, 65 fetched, 44 bills):

  • The sub-account rule: its removal changes 4 headings in 8 versions (19 of 60,508 amounts). Stacked names such as OFFICE OF TERRORISM AND FINANCIAL INTELLIGENCE over its own SALARIES AND EXPENSES are glued again.
  • The department rule: it fired 54 times and changed no output.

ADR 0022 now records named exceptions as a rejected alternative with these numbers, and the glued stacks as a known residual.

This agrees with your three extra pairs. As you say, byte-identical canonical output shows the exceptions were inert there, not that headings are right; the ledger measurement below is the accuracy check. I haven't rerun those three pairs on the final branch.

2. The sentence filter no longer deletes real headings

A line is now exempt from the unfinished-sentence filter when it is set apart:

  • it is centred;
  • and it is followed by an indented paragraph, another heading, or nothing.

That is a fact about the print, so missing terminal punctuation alone can't override it.

  • On 118 PDFs: 11 wrongly deleted headings restored, 0 fake headings let through, 0 real headings lost.
  • 118-s-4928 p.117: OFFICE OF INSPECTOR GENERAL and SALARIES AND EXPENSES are back, and the $274M is filed under them rather than under PAYMENT TO THE POSTAL SERVICE FUND.
  • The rest are Energy and Water's opening headings (118-hr-4394, 3 versions), which the filter had also deleted.

Tests:

  • TestHeadingsAfterUnpunctuatedProse checks the real p.117 page: the headings survive and the $274M sits under them.
  • Two synthetic tests cover the shape: an unpunctuated paragraph end, and an interrupted enacting clause.
  • The control that rejects heading-shaped text inside prose is kept, and is proved below.

What fixing this uncovered. Checking the restored headings exposed a bug under the case pattern the passes rely on. pdf_text._initial_caps split words at space glyphs, but the glyph walk drops PDFium's generated spaces. So some lines were judged by their first letter alone: 140 misread lines in the corpus, for example and Efficiency on 118-s-4928 p.75.

  • Words are now split at letter gaps, and all-capital acronyms don't vote.
  • The same gap rule fixes _first_word_right, used by the line-fullness veto: 4 merged NATO headings split, 0 broken.

With the case pattern read correctly, four heading rules no longer held. Each change below was checked against the raw XML <header> tags on the 118 PDFs:

Rule Change Measured on 118 PDFs
Repeat-words veto Removed. A name repeating its own words is no evidence of two headings. 9 changes, all wrong
Line-fullness veto Kept; a line ending in a small word continues. 7 of 7 toward the XML
Level of a split piece A piece printed like the account below it takes the account's level. 202 amounts better, 0 worse
Title-name continuation The rest of a TITLE n—NAME line is not a heading. 3 fragments removed

One of these needs a word on your keyword concern: the line-fullness veto's small-word override. A line ending in OF, AND, THE and similar (…GRANTS FOR CONSTRUCTION OF) always continues onto the next line.

  • It uses the small-word list the PR already had, TITLE_CASE_SMALL_WORDS: function words, not appropriations vocabulary.
  • At e5c5b6e the rule covered only AND / OR; it now covers the whole list.
  • The vocabulary gate scans the module with no exemption.

If you'd rather keep it to AND / OR, the cost is those 7 headings.

3. The scorer compares whole paths

tier() now compares every level:

  • T0 needs identical normalized paths.
  • T2 needs the XML's ancestors in order, with only omissions.
  • A wrong, extra or reordered ancestor is T3.

Your experiment, rerun on 117-hr-4502 v2: three Tax Court amounts moved under a made-up division, title and department.

  • Old tier(): every count unchanged.
  • New tier(): those 3 amounts move from T0 to T3.

tests/test_ledger_location_scorer.py pins the cases on hand-built paths: made-up, wrong, reordered, missing and extra ancestors, a label variant, and a different account.

The stricter scorer changes the picture in two ways:

  • Most regrades are a convention. Measured when the scorer was changed, T0 on the branch fell by 2,592 amounts. 2,375 of them are a title's name being displaced by the major heading below it in the PDF breadcrumb (true but shallower, now T2).
  • It exposes real errors the old scorer hid: 142 glued stacks and 16 amounts under a repeated level.

Regression tests fail under the faulty behaviour

Each check is a pytest plugin that puts the old behaviour back. Every one fails exactly the test written for it, and nothing else; the unmutated run passes all 101 tests.

Faulty behaviour restored Fails
The named exceptions (the current tests run on the e5c5b6e tree) both no-wording tests
No set-apart exemption (the filter as shipped) 2 synthetic tests + the real p.117 test
Any centred line exempt (a looser rule) the embedded-in-prose control
Old tier() 5 path rows: made-up, wrong, reordered, missing and extra ancestors
A totals-only comparison the hidden-regression test
Old case pattern / acronyms voting / old first-word gap their TestInitialCaps / TestFirstWordRight tests
Repeat veto / AND-OR-only grammar / old piece level / old title-name rule one test each
The 6fa6cbd rule off / without its $ check one test each

Regenerated measurements and baselines

Regenerated with the final parser and scorer:

  • the ledger baseline;
  • the PDF canonical baselines: 9 pairs, each count moving by 0–5 changes, mostly fewer false adds and moves;
  • the round-1 pairing sentinel (the parser revision moved);
  • the round-1 revocation pin: 225 → 217 accepted, 6 declined unchanged. The two pairs that moved are in its docstring.

ADR 0022's results, develop → this branch, both graded by the new scorer:

corpus (33 versions) holdout (4 FY2024 House bills, 10 versions)
same location (T0) 58.3% → 78.7% 29.2% → 53.8%
OK (T0–T2) 66.7% → 91.8% 45.1% → 83.0%
wrong parent (T3) 4,373 → 1,055 1,202 → 375
misfiled (T4) 99 → 40 45 → 12

Beyond the table:

  • Versions without appropriations headings (32 fetched): misfiled 368 → 4.
  • Totals: no version's totals got worse, committed or fetched.
  • Per amount: 6,413 better and 105 worse. 103 of the worse are ADMINISTRATIVE PROVISIONS—<agency> headings. The PDF now reads them in full, but the XML reader drops them. develop had kept only the fragment ADMINISTRATION, which T1 passed as a variant.
  • Account-name precision/recall: 7 of 9 bills still at 1.000; the lowest is 0.967.

Also new in these commits

  • 6fa6cbd narrows ADR 0012's boundary 1. A title-case heading is an agency, not an account, when both hold:

    • no dollar amount appears in its own text;
    • the next heading is account-style.

    Examples: BUREAU OF RECLAMATION, GREAT LAKES ST. LAWRENCE SEAWAY DEVELOPMENT CORPORATION.

    • Result: 183 amounts toward the XML, 0 away.
    • Title case alone was measured and rejected: 265 amounts the wrong way against 8.
    • The $ is read as a fact about the text, the way a period is, not as appropriations wording. ADR 0012 says so, and the gap it accepts still stands for an agency whose own text carries money.
  • scripts/heading_precision.py: recall now credits a PDF agency name that the XML tags appropriations-intermediate. Without it, the change above counts BUREAU OF RECLAMATION as a missed account (115-hr-5895 recall 0.982 → 0.945) though the XML sets it on the agency level. Precision is unchanged, accounts only, and all nine bills score as before.

  • A per-amount check (python -m tests.ledger_location --save before.json, then --against before.json). It follows each amount from one parser to another and lists every amount that got worse, because totals hide a regression whenever another amount improves. The tier totals remain the formal result that the pin holds; the check is a working tool, not a gate. It grades both runs with the checked-out scorer, so a scorer change can't pass for a parser change.

  • Docs:

    • TESTING.md's ledger section: tier table, the two readings, what the answer key gets wrong.
    • ADR 0022: tier definitions, rejected alternatives, results, residuals, open issues.
    • One pointer sentence in CONTRIBUTING.md's review guidance.

Open issues

ADR 0022 records these:

Local runs on Windows:

  • Slow suite: 1,755 passed, 0 failed.
  • Fast suite: 16 failures, the same known Windows-environment ones as on a clean develop.
  • Ruff: clean.

CI needs fork workflows approved.

🤖 Generated with Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants