Show each version's dollar amounts, typed, in three report views - #736
mattzamora wants to merge 11 commits into
Conversation
Ordered, fail-closed passes re-read detected headings across the reading order. pdf_text records each line's letter case pattern; pdf_anchors uses it to bound an agency's reach. Pins re-derived for the new parser revision.
Grades every amount T0-T4 against the XML twin and pins the tiers per version. python -m tests.ledger_location prints before/after totals.
ADR 0022 (Proposed). ADR 0018 admits named exceptions; ADR 0012 notes the case pattern. bill-structure.md, signal inventory and TESTING.md updated.
willhea
left a comment
There was a problem hiding this comment.
Hey @mattzamora! Thank you for the work on this. There’s useful progress here in making the financial extraction inspectable and connecting it to the report.
I’d like to close this PR and continue the methodology work in docs/research/financial-semantics/. We deliberately removed the financial displays in #681 / #671, and I don’t think we have sufficient evidence to bring them back yet. This is a disagreement about readiness and scope, so I’m declining the current approach rather than requesting incremental changes to the UI.
I reviewed these examples at 624e77ed and attempted to disprove them against the source bill text and the generated reports. Three remain concrete classification or amount-selection errors; one demonstrates a limitation in what the comparison includes.
To reproduce, check out this PR and run these commands from the repository root:
uv sync --locked
uv run python diff_bill.py compare \
tests/corpus/118-hr-4366/4_engrossed-amendment-senate.xml \
tests/corpus/118-hr-4366/5_engrossed-amendment-house.xml \
--format html -o hr4366-v4-v5.html
uv run python diff_bill.py compare \
tests/corpus/118-hr-8752/1_reported-in-house.xml \
tests/corpus/118-hr-8752/2_engrossed-in-house.xml \
--format html -o hr8752-v1-v2.html1. An appropriation is displayed and calculated as a rescission.
Open hr4366-v4-v5.html, select Financials – Version A, and find MEDICAL SERVICES. The opening $71 billion amount is labeled “rescission.”
The source provides $71 billion for medical services and separately rescinds $4,933,113,000 later in the paragraph. Version B changes that later rescission to $3,034,205,000. The classifier applies the later rescission language to the opening appropriation, and the comparison treats the $71 billion as negative money given out.
2. The comparison excludes a funding change already present in the ledger.
In the same report, find COMPENSATION AND PENSIONS in the financial comparison, then inspect its clauses in both version views.
The comparison shows a $10,416,509,000 increase from the opening appropriation. A separate advance appropriation changes from $181,390,281,000 to $182,310,515,000—a $920,234,000 increase—which is present in the ledger but excluded from the comparison.
Using only the first clause is an intentional implementation choice. However, the broader “Net change in money given out” label does not communicate that limitation. These amounts also concern different fiscal years, so this example should not be read as a proposed corrected bill-wide total.
3. A spending cap is selected as the account’s appropriation.
Open hr8752-v1-v2.html, select Financials – Version A, and find Coast Guard → OPERATIONS AND SUPPORT.
The displayed appropriation is $31 million. The source describes boat-related spending as “not to exceed a total of $31,000,000”; the account’s appropriation is $10,554,261,000.
The row does carry a review warning, but the $31 million still enters the comparison calculation. Both versions select that same figure, so this example demonstrates an incorrect selected amount, not an incorrect nonzero delta.
4. Allocations are incorrectly marked as amendments to existing law.
In that report, export the Version A financial CSV and inspect §211(a). The six allocations—including $600 million for physical barriers—have in_amended_law=true.
The source section allocates funds made available by this bill; it does not amend another law. Its six allocations sum exactly to the stated $1,390,338,000 parent amount. “As follows:” appears to trigger the amended-law classification. This is an incorrect label and export value; it is not evidence that these allocations should be added again to a funding total.
I’d like the next step to be resolving these cases in the research folder, with independently justified expected results. Matching the notebook’s outputs is useful regression evidence, but it does not establish that those outputs correctly represent the legislation.
Once we’re confident in the methodology, I’d like to resolve generating and validating the financial data separately from displaying it. The tables, totals, warnings, exports, and comparison behavior involve substantial design choices that deserve their own review.
Thank you again for the contribution. I’d welcome continuing this work through the research first.
…dc#115) Schema 3.1 adds an optional financial field: for each version on its own, every dollar amount with its chain of headings, its clause, a type read from the wording by the research notebook's rules (classifier 1.0, moved into the package unchanged) and flags. A change still carries no money. PDF sections are read so they give the notebook's rows: gutter and running header off, page-seam hyphens rejoined, paragraphs split where the XML has nodes, quote marks removed for the rules only. The report gains Inferred Financial Comparison, Financials - Version A and Version B, each with a CSV export and a not-authoritative alert. Find also searches collapsed rows; each tab keeps its scroll position; jump targets clear the sticky bar at its measured height; the change navigator and changes export show only on Changes and Full bill. The PDF heading keeps the whole bill title and an upload label is the file name as given. The civictechdc#671 no-money checks now look where civictechdc#671 drew the line, the Changes view. Canonical baselines re-pinned: the digests move by the new field only.
ADR 0023 (Proposed): the per-side ledger, reading PDF sections for the notebook's values, why the classifier is versioned and how it improves, the three views, and the open question on money in the API's JSON. ADR 0006 amended in place; schema doc at 3.1; TESTING.md gains check 8 and the financial rows pin; research README records that the classifier moved into the product.
624e77e to
f08e6cc
Compare
Related issue
Refs #115 (a first cut of its sub-amount decomposition; see "What is not done")
Refs #195 (every dollar amount of a version, by section)
Depends on #734. This branch is built on #734's branch, so its first four commits are
#734's; review from the fifth. I will rebase onto
developonce #734 merges.Related: #735 (a PDF's last section carries the back matter; found building this), #724 (the research notebook's summary table; it touches
classify_bill.py's outputcolumns, not its rules, so the parity checked here is unaffected), #147 (financial semantics),
#175 (consuming the leveled tree), #521 (headings that own no text), #518 (unnumbered
sections).
What does this change?
The report gets three new tabs after Changes and Full bill, showing the money in the two
versions being compared:
version, with each dollar amount's type (appropriation, rescission, cap, earmark, …), a
category key with counts, sorting, and a row that opens onto the section's text highlighted
clause by clause. "Export Inferred Financials" downloads that version's whole ledger as CSV.
Version A's amount beside Version B's, and the difference in money given out. A row opens onto
that change's word diff and links to its section in Version A and Version B. "Export
Comparison" downloads it as CSV.
The types come from the research notebook's rules (
docs/research/financial-semantics/), movedinto the package unchanged as classifier version 1.0. What is new is reading them from the PDF:
the product mostly compares PDFs, so the same bill read from its PDF has to give the notebook's
rows. Every view carries an information alert: the types are not authoritative, users must
verify values against the bill text, the classifications will change as the application
improves, some sections and amounts may be missed, and feedback goes to #congressional-tech on
the Civic Tech DC Slack.
The data is a new optional
financialfield in the canonical document (schema 3.1): for eachversion separately, every dollar amount with its chain of headings, its clause, its type and
flags. A change still carries no money (ADR 0006); the ledger is per side and unpaired. Design
and reasoning: ADR 0023 (Proposed).
How it is checked
Against the research notebook, on H.R. 4366:
The three PDF rows missing on v4 are sections the PDF reader does not start (lettered numbers:
SEC. 109A.,SEC. 119A.,SEC. 119B.); the two rows that absorb them are flagged on the page("may hold more than one section"). Every
$figure in the text is shown: 141 of 141 (v1), 206of 206 (v2), on both pipelines.
The classifier is versioned so it can improve without changing silently. The parity rows
are frozen under the version that produced them (
tests/data/financial_rows/). A rule changethat moves any row fails the test until
CLASSIFIERis bumped, and regenerating refuses towrite different rows under the same version. Reports record and show the version. There are no
releases to pin, so an earlier version is the commit before its changelog line. How to improve
it, step by step, and the known candidates (the Senate's earmark wording is typed "unknown";
lettered sections; "For purposes of" openers): ADR 0023 and TESTING.md.
Each rule was broken on purpose and a test went red (quote-mark removal, hyphen rejoin, heading
split, mid-sentence guard, subsection split, quotation guard, running-header skip, amended law,
and the version gate).
Changes to existing report behaviour
These came with the new tabs and touch the report's existing code (
diff_html.py); each istested in a browser:
(
← n / N →) show only on Changes and Full bill, where they apply.also searches rows that are collapsed: it fills them in and opens the one with the match.
Ctrl+F cannot see collapsed text; this can.
the last one left the page.
one-row bar but not the two-row bar these tabs need; it now follows the bar's measured
height. (Before, a Full bill contents link could land its section under the bar.)
their 940px reading width.
the file name exactly as given, extension included (
web/app.py): the app cannot inferanything from someone's file naming.
"…"; the XML report never cut it).
Decisions deliberately not made here
financial? The report's own "Download diff.json"strips it, so what a staffer uploads to an AI assistant is unchanged. The API
(
output=json) and the CLIs' canonical JSON do include it, and a machine reading only thatfile sees the classifier's types without the page's alert. Options: keep it (it is typed per
clause, flagged and versioned), strip it there too, or add a machine-readable caveat to the
field. ADR 0023, Consequences.
and Full bill restyles their layout, and their change cards are prose, which reads poorly at
200+ characters a line.
What is not done
across versions and a per-account change narrative are not built. The comparison deliberately
does not pair clauses across versions (it links to each version's section instead): a
clause-by-clause comparison was mocked up from real data, and pairing went wrong exactly where
the classifier types one clause differently in the two versions.
shallower (no reliable bill-type signal to show it on).
Found while building this, filed separately: a PDF's last section carries everything printed
after it (short title, attestation, back cover), in 27 of 27 line-numbered corpus versions,
which also stops a renumbered last section being recognized as moved (#735).
For reviewers
financial._prose_blocks(how a PDF section's text is read beforethe rules see it) and
financial_views.comparison_rows(the comparison uses the differ's ownpairing by span overlap; no second matcher).
financial.pyandformatters/financial_views.pyare added, the twomodules that read or render money wording.
diff_html.pystays free of it: the views' CSSand script live in
financial_views.py, and the generic hooks it adds (find:prepare,find:reveal,data-find-reveal,body[data-active-view]) carry no financial vocabulary.--info/--info-foregroundfor thealert, and
--sticky-bar-height(default 64px, set by script).duplicated (every range points into
full_text, and the page reads the text from there).test_pdf_compare.py::test_compare_pdfs_html_returns_standalone_reportand
test_format_html.py::test_format_html_real_bill_renders_no_financial_presentationasserted that "Financial Summary" appears nowhere in a report. Version A / B now carry that
heading (the research notebook's). The tests keep every Remove the financial table and paired-amount claims until amounts can be typed to accounts #671 marker check
(
financial-table,financial-callout,data-financial), assert the Changes view carries nomoney presentation at all, and pin the heading to exactly the two financial views.
tests/data/canonical_baseline.json(27 pairs) andtests/data/pdf_canonical_baseline.json(17 compared, 6 declined unchanged): only digests andbyte counts move; every change count and summary is identical. Rebuilding all 50 pinned pairs
with this branch, removing
financialand settingschema_versionback to3.0, reproducesevery stored digest exactly, so the diff itself did not move; the new field is the whole
change. New pin:
tests/data/financial_rows/(above). The committed examples and the servedsample are re-rendered for the new views.
tab with no address of its own, where history entries do not behave the same across
browsers.
How to test
To see it: run the web app (
uvicorn web.app:app --port 8077), open/compare.html, uploadtests/corpus/118-hr-4366/1_reported-in-house.pdfand2_engrossed-in-house.pdf, and open thethree new tabs. For a large case use
4_engrossed-amendment-senate.pdf→5_engrossed-amendment-house.pdf; the XML pair gives the same ledger rows.Ran locally on Windows at the head of this branch, each CI step on its own: ruff check and
format clean; browser 48 passed; corpus gates 1,390 passed (26 declared skips); remaining slow
431 passed (7 declared skips); external validation 35 passed. Not run locally: the packaging
gate (
tests/test_engine_installs.py, which builds the wheel and installs it into a throwawayenvironment); it is left to CI here. The fast tier's 15 failures are
this machine's, not the change's: the same tests fail identically on an untouched
developcheckout (no executable bit on Windows, backslash paths, the CI-report checks). The
committed-examples check fails on Windows only because the renderer writes the OS line ending;
a fresh render here equals the committed LF bytes once CR is removed, for every example. Both
commits were checked the same way on a clean worktree.
Checklist
test_engine_installs.py), left to CIAI assistance
Claude Code (Claude Opus 5.5) wrote the code, tests and docs under the contributor's direction.
The views follow the research notebook's report and a comparison layout chosen from mockups;
each change to existing behaviour was reviewed in the running app.
🤖 Generated with Claude Code