Recover PDF heading structure for the bill ledger, measured by agreement with the XML twin - #734
mattzamora wants to merge 9 commits into
Conversation
Ordered, fail-closed passes re-read detected headings across the reading order. pdf_text records each line's letter case pattern; pdf_anchors uses it to bound an agency's reach. Pins re-derived for the new parser revision.
Grades every amount T0-T4 against the XML twin and pins the tiers per version. python -m tests.ledger_location prints before/after totals.
ADR 0022 (Proposed). ADR 0018 admits named exceptions; ADR 0012 notes the case pattern. bill-structure.md, signal inventory and TESTING.md updated.
willhea
left a comment
There was a problem hiding this comment.
Hey @mattzamora! I love this line of work. I reviewed with GPT Sol 6.1 and arrived at the following. I think these changes will get this PR good to go. I'm reluctant to amend the ADR for a list of keywords. I would rather err on the side of preserving the existing reading when the printed evidence is ambiguous, even if that leaves some wrapped headings split. Recovering fewer headings is an acceptable tradeoff if it avoids inferring hierarchy from keyword matches.
Reviewed at e5c5b6e. I’d like to retain the typography approach, with three changes before merging:
-
Keep ADR 0018 intact. Please remove the named wording exceptions and their vocabulary-gate exemption. I’m concerned that these encode answers to known examples rather than establish that the printed evidence distinguishes the structures. Where format is insufficient, preserve the existing reading. We should accept some missed joins rather than invent hierarchy.
-
Fix the sentence filter deleting real headings. In
118-s-4928/1_reported-in-senate.pdf, page 117, the preceding paragraph ends with(Public Law 111–241)without a period in the actual print. The new filter removesOFFICE OF INSPECTOR GENERALandSALARIES AND EXPENSES, filing $274 million under the precedingPAYMENT TO THE POSTAL SERVICE FUNDaccount. Missing terminal punctuation alone should not override clear heading evidence. Please add a small behavioral regression test for this case while retaining the control that rejects heading-shaped text embedded in prose. -
Correct the ledger scorer before regenerating the results.
tier()can award T0 when the account and immediate parent match despite wrong ancestors above them. Changing three Tax Court amounts to a fabricated division, title, and department left every tier count unchanged. T2 also uses sets, so reordered ancestors can be accepted as shallower. T0 should require identical normalized paths; T2 should require an ordered ancestry with only omissions. Wrong or reordered ancestors must be rejected.
Please demonstrate that the regression tests fail under the corresponding faulty behavior, then regenerate the affected measurements and baselines after these changes.
I also ran three additional pairs—119-H.R.3944, 119-H.R.1968, and 118-H.R.815—as PDF→PDF and XML→XML. The core smoke checks passed, and disabling the wording exceptions produced byte-identical PDF canonical outputs to this PR on all three pairs. That supports removing the exceptions, although the smoke test does not establish heading or matching accuracy.
|
Thanks @willhea, this was a careful review, and all three points held up when I checked them. They are addressed in five commits on top of
I agree with the principle you set: where the print is ambiguous, keep the existing reading and accept fewer recovered headings rather than infer hierarchy from wording. Every rule added below reads the print or grammar, never appropriations vocabulary. The one ADR amended, 0012, is amended on a fact about the text, not a keyword list; details under "Also new". 1. ADR 0018 kept intactThe named wording exceptions are gone: the module, and its vocabulary-gate exemption. ADR 0018 and
What it costs, measured on 118 PDFs (53 committed, 65 fetched, 44 bills):
ADR 0022 now records named exceptions as a rejected alternative with these numbers, and the glued stacks as a known residual. This agrees with your three extra pairs. As you say, byte-identical canonical output shows the exceptions were inert there, not that headings are right; the ledger measurement below is the accuracy check. I haven't rerun those three pairs on the final branch. 2. The sentence filter no longer deletes real headingsA line is now exempt from the unfinished-sentence filter when it is set apart:
That is a fact about the print, so missing terminal punctuation alone can't override it.
Tests:
What fixing this uncovered. Checking the restored headings exposed a bug under the case pattern the passes rely on.
With the case pattern read correctly, four heading rules no longer held. Each change below was checked against the raw XML
One of these needs a word on your keyword concern: the line-fullness veto's small-word override. A line ending in
If you'd rather keep it to 3. The scorer compares whole paths
Your experiment, rerun on 117-hr-4502 v2: three Tax Court amounts moved under a made-up division, title and department.
The stricter scorer changes the picture in two ways:
Regression tests fail under the faulty behaviourEach check is a pytest plugin that puts the old behaviour back. Every one fails exactly the test written for it, and nothing else; the unmutated run passes all 101 tests.
Regenerated measurements and baselinesRegenerated with the final parser and scorer:
ADR 0022's results, develop → this branch, both graded by the new scorer:
Beyond the table:
Also new in these commits
Open issuesADR 0022 records these:
Local runs on Windows:
CI needs fork workflows approved. 🤖 Generated with Claude Code |
Related issue
Closes #732
Closes #524
Advances:
Related: #551 (epic: PDF heading detection rests on layered geometry rules), #198 (font name as a
margin discriminator), #557 (a display-only heading change can alter the structure tree untested),
#552 (epic: levels and addresses derived from display labels), #706 (XML and PDF summaries report
different change categories for the same bill).
What does this change?
The PDF reader now recovers a bill's headings much closer to how the bill's XML labels them. It
reads each detected heading together with the lines around it, instead of one printed line at a
time, and corrects five misreadings:
FAMILY HOUSING OPERATION AND MAINTENANCE,/ARMYon 118-hr-4366 made the first half anagency, and each later family-housing account was nested under it.
read as a heading of the bill.
Where the print does not show the evidence, the original reading is kept.
Accuracy is checked two ways, both with the bill's XML as the reference:
of the account-precision check, the lowest recall / precision goes from 0.744 / 0.750 to
0.967 / 0.967, with seven of the nine at 1.000 / 1.000. On 118-hr-4820, PDF account headings
with no match in the XML go from 17 to 0.
chain in the XML (the ledger, table below). The same chain: 62.3% → 87.6% on the committed
corpus, and 39.5% → 74.2% on four bills never used to design the rules.
One earlier assumption this changes. ADR 0012 and
docs/bill-structure.mdrecorded that anagency heading followed by prose looks identical to an account by size, case and position, and
accepted that as a gap. The letters do differ: GPO sets an agency in title-case small caps (each
word's first letter larger) and an account in even small caps. This PR reads that pattern and
uses it to decide one heading versus two and to stop an agency at an account that cannot be its
own. It does not re-label such an agency, which is still reported as an account, so the gap is
narrowed, not closed. Both documents are updated.
Where the change lives, for reviewers reading the code:
parsers/pdf_heading_passes.py(new):converge_headings, ordered fail-closed passes thatre-read the detected headings against the whole reading order, after detection and before
divisions are assigned. Order and rationale: ADR 0022 (Proposed).
parsers/pdf_text.py: one new signal from the existing glyph walk, no new PDFium calls.LineGeom.initial_capsis the letters' case pattern (GPO sets an agency in title case in smallcaps, an account in even small caps; a wrapped name's two lines print alike), and
size_min/size_maxthe letter size range (capitals at body size are the department style).Prose lines are not judged.
parsers/pdf_anchors.py:Anchor.capsrecords the case pattern;_breadcrumb_corelets adepartment end the agency above it, and
_agency_reachesstops a carried-over agency at anaccount that prints like it cannot be its child.
parsers/pdf_heading_exceptions.py(new): two named wordings decide a line break whereformat cannot (
SALARIES AND EXPENSESon its own line; the fifteen executive departments). Itis the one heading module on the ADR 0018 vocabulary-gate allowlist, and ADR 0018 is rewritten
to admit it and nothing else.
tests/test_pdf_ledger_location.py+tests/ledger_location.py(new): the ledger gate. Theledger is every dollar amount in a bill with the breadcrumb it sits under; each amount in
the XML twin's ledger is aligned with the PDF's and graded T0 (same location) … T4 (different
heading), pinned exactly per version.
python -m tests.ledger_locationprints the totals forwhichever parser is importable, which is how the "before" column below was produced.
Evidence
Ledger, measured by the same scorer on
develop(c636448) and on this branch:(13,422 amounts), as
python -m tests.ledger_locationprints them.them: 118-hr-4368, 118-hr-4394, 118-hr-4665, 118-hr-4821 (10 versions, 2,272 amounts). Not
committed, so they stay unseen for the follow-up that broadens the corpus.
different heading 84 → 4.
develop, the report finds 0 of the 42pinned versions better than this branch; 0 of 10 holdout and 0 of 17 fetched reconciliation
versions either.
Why OK stops near 93% rather than 99%. The answer key is DeltaTrack's existing XML reader,
which leaves a heading with no text of its own out of the breadcrumb (
Food and Drug AdministrationunderDepartment of Health and Human Services). The PDF now keeps it, and isgraded "wrong parent" where it matches the page: of the 944 wrong-parent amounts, up to 892 have
a PDF parent the XML file tags as a heading and the reader drops. This PR leaves the reader alone
and says so in ADR 0022; changing it is a separate decision (#733, the XML reader drops a heading that carries no text of its own).
#524's verification, criterion by criterion:
NATIONAL OCEANIC AND ATMOSPHERIC ADMINISTRATIONon114-hr-2578 (fetched; not in the corpus) stays an agency over its own five accounts in all four
versions, before and after. Corpus-wide, no version's not-tolerated count rose.
A wrapped account heading in a PDF is read as an agency plus an account #524 are read as one account each, including the
RAILROAD REHABILITATION AND IMPROVEMENT FINANCING PROGRAMcase A wrapped account heading in a PDF is read as an agency plus an account #524 had set aside. Unmatched PDF account headings 17 → 0.(
TestFailsClosed).The account-precision floors #524 names are re-derived from this parser revision:
TestCorpusAccountPrecision's lowest recall/precision goes 0.744 / 0.750 → 0.967 / 0.967 (sevenof nine bills at 1.000 / 1.000), and the floors go 0.70 → 0.95.
#648 (advanced, not closed). Renumbered cards on the 17 accepted corpus pairs: 156 → 143.
Fixed:
GINIA.— → WEST VIRGINIA.—and about a dozen other wrap-only renames. Kept, as #648requires: the four genuine
SEC.renumberings in 115-hr-5895 and the genuine heading edit(
SECURITY → SECRUITY). Not fixed:…NAVY AND MARINE CORPSstill reads asAND MARINE CORPSonthe Senate strike-all print, and
formatters/canonical._pdf_move, the second site #648 names, isuntouched.
#500 (an account printed directly under a section catchline is dropped) is not affected: its
strict expected failure still fails.
Pins regenerated, each with its cause
tests/data/pdf_canonical_baseline.json: change counts mostly fall as wrapped names become oneblock (117-hr-4502 v1→v2 1500→1456); 118-hr-8752 v1→v2 37→38 (next bullet).
tests/test_pipeline_parity.py+ ADR 0014 table: 118-hr-8752 PDF 37→38. The engrossed versionadds a
SPENDING REDUCTION ACCOUNTheading the PDF now reads as its own block, and SEC.552–567 file under it as they do in the XML (which also keeps
GENERAL PROVISIONSabove it).117-hr-4502 gap to the XML +85→+41, band unchanged.
examples/hr8752_pdf_diff.htmlandweb/webapp/sample/example.html: the same 8752 change,and nothing else (the breadcrumb headings of SEC. 552–567, plus the added heading block).
tests/data/round1_pairing_sentinel.json: onlyparser_revisionre-stamped (ADR 0019); nopairing digest moved.
tests/data/pdf/anchors_golden/: 118-s-4795, seven wrapped account names are now one accounteach and one fragment is now its full name; the frozen
.pre-agency-anchorsbaseline isre-anchored for exactly those eight.
major_vocab.json: additions only (mid-titledepartments).
tests/test_pdf_round1_revocation.py: split population 224+6 → 225+6 (three pairs' blockboundaries moved; per-pair deltas in the docstring).
tests/test_pdf_anchor_golden.py: account-precision floors 0.70 → 0.95 (above).tests/test_pdf_size_detection.py: synthetic paragraphs end with a period before a heading, asGPO prints them; a line following unfinished prose is never a heading. What each test asserts
is unchanged.
Known residuals (ADR 0022, Consequences)
(
ATOMIC ENERGY DEFENSE ACTIVITIESoverNATIONAL NUCLEAR SECURITY ADMINISTRATION, Two stacked department headings merge into one when the upper line nearly fills the column #501): theprint carries no format signal there.
OVERSEAS CONTINGENCY OPERATIONS DEPARTMENT OF DEFENSE) is left as detected.TITLEstays split (115-hr-5895).shallower.
Docs
ADR 0022 (new, Proposed) + index row; ADRs 0018, 0012 and 0014 rewritten in place;
docs/bill-structure.md(the ledger, the heading forms read across lines, the case-patternfinding);
docs/source-signal-inventory.md(the new signal);TESTING.md(accuracy check 7 andthe ledger pin, with the before/after commands).
For reviewers
_segmentand_agency_reaches. Each rule has a test that fails whenthe rule is removed (28 mutations, all caught).
of this change, the
SALARIES AND EXPENSESrule moves 15 amounts (9 to "wrong parent" becausethe heading it separates is one the XML reader drops, 6 to "shallower"); the departments rule
moves none and corrects heading names. To drop them: delete the module, its allowlist line and
the
named_splitcall, then regenerate the pins.prose-leading agency line, per the ADR's own spike and
bill-structure.md.pdf_text._attach_geometryand regenerates the samethree pins (canonical baseline, sentinel, revocation). Whichever lands second regenerates them.
If Widen the skip ceiling to the whole suite, fix five CWD-dependent defects, add a cwd-independence CI job #721 (skip ceiling for the whole suite) lands first, the ledger gate's update-mode skip needs
the same allowlist entry as
test_pdf_canonical_baseline's.takes about 0.8 s longer (≈18.4 s → ≈19.2 s, mean of three runs; the machine's run-to-run spread
is about 1 s).
measures 302.
How to test
To see it on a page:
118-hr-4366/1_reported-in-house.pdf, p. 10 l. 13–14,FAMILY HOUSING OPERATION AND MAINTENANCE,/ARMY. Ondevelopthe drill-down shows… › FAMILY HOUSING OPERATION AND MAINTENANCE, › ARMY, and each later family-housing account inherits the fragment;on this branch it is one account under
DEPARTMENT OF DEFENSE.Ran locally on Windows at the head of this branch: ruff check and format clean; fast 2,083
passed; browser 41 passed; slow (every CI slow gate) 1,755 passed, 0 failed. The fast tier's 16
failures are this machine's, not the change's: missing executable bits, CRLF line endings in the
committed renders, and CI-report checks. The same 16 fail on a clean
developcheckout, and eachof the first two commits, applied alone on top of
develop, gives the same fast result and agreen slow tier.
Checklist
Closes #...)AI assistance
Claude Code (Claude Opus 5.5) wrote the code, tests and docs under the contributor's direction. The
design came out of an investigation reviewed case by case against the printed pages and the XML.
🤖 Generated with Claude Code