Skip to content

Show each version's dollar amounts, typed, in three report views - #736

Draft
mattzamora wants to merge 11 commits into
civictechdc:developfrom
mattzamora:feat/115-financial-views
Draft

mattzamora wants to merge 11 commits into
civictechdc:developfrom
mattzamora:feat/115-financial-views

Conversation

@mattzamora

Copy link
Copy Markdown
Contributor

Related issue

Refs #115 (a first cut of its sub-amount decomposition; see "What is not done")
Refs #195 (every dollar amount of a version, by section)

Depends on #734. This branch is built on #734's branch, so its first four commits are
#734's; review from the fifth. I will rebase onto develop once #734 merges.

Related: #735 (a PDF's last section carries the back matter; found building this), #724 (the research notebook's summary table; it touches classify_bill.py's output
columns, not its rules, so the parity checked here is unaffected), #147 (financial semantics),
#175 (consuming the leveled tree), #521 (headings that own no text), #518 (unnumbered
sections).

What does this change?

The report gets three new tabs after Changes and Full bill, showing the money in the two
versions being compared:

  • Financials – Version A and Financials – Version B: every money-bearing section of one
    version, with each dollar amount's type (appropriation, rescission, cap, earmark, …), a
    category key with counts, sorting, and a row that opens onto the section's text highlighted
    clause by clause. "Export Inferred Financials" downloads that version's whole ledger as CSV.
  • Inferred Financial Comparison: the changes from the Changes tab whose sections hold money,
    Version A's amount beside Version B's, and the difference in money given out. A row opens onto
    that change's word diff and links to its section in Version A and Version B. "Export
    Comparison" downloads it as CSV.

The types come from the research notebook's rules (docs/research/financial-semantics/), moved
into the package unchanged as classifier version 1.0. What is new is reading them from the PDF:
the product mostly compares PDFs, so the same bill read from its PDF has to give the notebook's
rows. Every view carries an information alert: the types are not authoritative, users must
verify values against the bill text, the classifications will change as the application
improves, some sections and amounts may be missed, and feedback goes to #congressional-tech on
the Civic Tech DC Slack.

The data is a new optional financial field in the canonical document (schema 3.1): for each
version separately, every dollar amount with its chain of headings, its clause, its type and
flags. A change still carries no money (ADR 0006); the ledger is per side and unpaired. Design
and reasoning: ADR 0023 (Proposed).

How it is checked

Against the research notebook, on H.R. 4366:

clause rows matching the notebook appropriation total
As reported (v1), from XML 108 of 108 equal
As reported (v1), from PDF 108 of 108 equal
Senate amendment (v4), from XML 622 of 622 equal
Senate amendment (v4), from PDF 619 of 622 equal

The three PDF rows missing on v4 are sections the PDF reader does not start (lettered numbers:
SEC. 109A., SEC. 119A., SEC. 119B.); the two rows that absorb them are flagged on the page
("may hold more than one section"). Every $ figure in the text is shown: 141 of 141 (v1), 206
of 206 (v2), on both pipelines.

The classifier is versioned so it can improve without changing silently. The parity rows
are frozen under the version that produced them (tests/data/financial_rows/). A rule change
that moves any row fails the test until CLASSIFIER is bumped, and regenerating refuses to
write different rows under the same version. Reports record and show the version. There are no
releases to pin, so an earlier version is the commit before its changelog line. How to improve
it, step by step, and the known candidates (the Senate's earmark wording is typed "unknown";
lettered sections; "For purposes of" openers): ADR 0023 and TESTING.md.

Each rule was broken on purpose and a test went red (quote-mark removal, hyphen rejoin, heading
split, mid-sentence guard, subsection split, quotation guard, running-header skip, amended law,
and the version gate).

Changes to existing report behaviour

These came with the new tabs and touch the report's existing code (diff_html.py); each is
tested in a browser:

  • "Export and share" is renamed "Export and share changes", and it and the change navigator
    (← n / N →) show only on Changes and Full bill, where they apply.
  • The find box is larger, with a magnifying-glass icon ("Search this view…"), and it now
    also searches rows that are collapsed: it fills them in and opens the one with the match.
    Ctrl+F cannot see collapsed text; this can.
  • Each tab keeps its own scroll position. Previously a tab came back at whatever position
    the last one left the page.
  • Jumps land below the sticky bar. Jump targets had a fixed 64px clearance, enough for the
    one-row bar but not the two-row bar these tabs need; it now follows the bar's measured
    height. (Before, a Full bill contents link could land its section under the bar.)
  • The three financial tabs use the full window on a wide screen; Changes and Full bill keep
    their 940px reading width.
  • The header lists "Before:" and "After:" on their own lines; for an upload, the label is
    the file name exactly as given, extension included (web/app.py): the app cannot infer
    anything from someone's file naming.
  • The PDF report's heading keeps the whole bill title (it was cut at 140 characters with
    "…"; the XML report never cut it).

Decisions deliberately not made here

  1. Should the API's JSON output carry financial? The report's own "Download diff.json"
    strips it, so what a staffer uploads to an AI assistant is unchanged. The API
    (output=json) and the CLIs' canonical JSON do include it, and a machine reading only that
    file sees the classifier's types without the page's alert. Options: keep it (it is typed per
    clause, flagged and versioned), strip it there too, or add a machine-readable caveat to the
    field. ADR 0023, Consequences.
  2. Full window width for every tab? Only the financial tabs widen here. Widening Changes
    and Full bill restyles their layout, and their change cards are prose, which reads poorly at
    200+ characters a line.

What is not done

Found while building this, filed separately: a PDF's last section carries everything printed
after it (short title, attestation, back cover), in 27 of 27 line-numbered corpus versions,
which also stops a renumbered last section being recognized as moved (#735).

For reviewers

  • Where to look hardest: financial._prose_blocks (how a PDF section's text is read before
    the rules see it) and financial_views.comparison_rows (the comparison uses the differ's own
    pairing by span overlap; no second matcher).
  • ADR 0018 allowlist: financial.py and formatters/financial_views.py are added, the two
    modules that read or render money wording. diff_html.py stays free of it: the views' CSS
    and script live in financial_views.py, and the generic hooks it adds (find:prepare,
    find:reveal, data-find-reveal, body[data-active-view]) carry no financial vocabulary.
  • Palette: seven type colour pairs (the notebook's), --info / --info-foreground for the
    alert, and --sticky-bar-height (default 64px, set by script).
  • Size: the ledger adds about 94 KB to the H.R. 4366 v1→v2 report; the bill text is not
    duplicated (every range points into full_text, and the page reads the text from there).
  • Two Remove the financial table and paired-amount claims until amounts can be typed to accounts #671 tests now check where Remove the financial table and paired-amount claims until amounts can be typed to accounts #671 drew its line. test_pdf_compare.py::test_compare_pdfs_html_returns_standalone_report
    and test_format_html.py::test_format_html_real_bill_renders_no_financial_presentation
    asserted that "Financial Summary" appears nowhere in a report. Version A / B now carry that
    heading (the research notebook's). The tests keep every Remove the financial table and paired-amount claims until amounts can be typed to accounts #671 marker check
    (financial-table, financial-callout, data-financial), assert the Changes view carries no
    money presentation at all, and pin the heading to exactly the two financial views.
  • Pins regenerated, with the evidence. tests/data/canonical_baseline.json (27 pairs) and
    tests/data/pdf_canonical_baseline.json (17 compared, 6 declined unchanged): only digests and
    byte counts move; every change count and summary is identical. Rebuilding all 50 pinned pairs
    with this branch, removing financial and setting schema_version back to 3.0, reproduces
    every stored digest exactly, so the diff itself did not move; the new field is the whole
    change. New pin: tests/data/financial_rows/ (above). The committed examples and the served
    sample are re-rendered for the new views.
  • Tabs add no browser history (Back leaves the report): reports are often written into a
    tab with no address of its own, where history entries do not behave the same across
    browsers.

How to test

uv run ruff check . && uv run ruff format --check .
uv run pytest -m "not slow and not browser"
uv run pytest -m browser --run-browser
uv run pytest -m slow tests/test_financial_corpus.py -v     # PDF/XML against the frozen rows

To see it: run the web app (uvicorn web.app:app --port 8077), open /compare.html, upload
tests/corpus/118-hr-4366/1_reported-in-house.pdf and 2_engrossed-in-house.pdf, and open the
three new tabs. For a large case use 4_engrossed-amendment-senate.pdf →
5_engrossed-amendment-house.pdf; the XML pair gives the same ledger rows.

Ran locally on Windows at the head of this branch, each CI step on its own: ruff check and
format clean; browser 48 passed; corpus gates 1,390 passed (26 declared skips); remaining slow
431 passed (7 declared skips); external validation 35 passed. Not run locally: the packaging
gate (tests/test_engine_installs.py, which builds the wheel and installs it into a throwaway
environment); it is left to CI here. The fast tier's 15 failures are
this machine's, not the change's: the same tests fail identically on an untouched develop
checkout (no executable bit on Windows, backslash paths, the CI-report checks). The
committed-examples check fails on Windows only because the renderer writes the OS line ending;
a fresh render here equals the committed LF bytes once CR is removed, for every example. Both
commits were checked the same way on a clean worktree.

Checklist

  • Linked the issue above
  • Ran the CI gates locally and they pass (see What CI checks), except the packaging gate (test_engine_installs.py), left to CI
  • New or changed behavior has tests
  • For a bug fix: the test fails without the fix, and I ran it both ways to check
  • Disclosed AI assistance below, if any

AI assistance

Claude Code (Claude Opus 5.5) wrote the code, tests and docs under the contributor's direction.
The views follow the research notebook's report and a comparison layout chosen from mockups;
each change to existing behaviour was reviewed in the running app.

🤖 Generated with Claude Code

Ordered, fail-closed passes re-read detected headings across the reading order. pdf_text records each line's letter case pattern; pdf_anchors uses it to bound an agency's reach. Pins re-derived for the new parser revision.
Grades every amount T0-T4 against the XML twin and pins the tiers per version. python -m tests.ledger_location prints before/after totals.
ADR 0022 (Proposed). ADR 0018 admits named exceptions; ADR 0012 notes the case pattern. bill-structure.md, signal inventory and TESTING.md updated.

@willhea willhea left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @mattzamora! Thank you for the work on this. There’s useful progress here in making the financial extraction inspectable and connecting it to the report.

I’d like to close this PR and continue the methodology work in docs/research/financial-semantics/. We deliberately removed the financial displays in #681 / #671, and I don’t think we have sufficient evidence to bring them back yet. This is a disagreement about readiness and scope, so I’m declining the current approach rather than requesting incremental changes to the UI.

I reviewed these examples at 624e77ed and attempted to disprove them against the source bill text and the generated reports. Three remain concrete classification or amount-selection errors; one demonstrates a limitation in what the comparison includes.

To reproduce, check out this PR and run these commands from the repository root:

uv sync --locked

uv run python diff_bill.py compare \
  tests/corpus/118-hr-4366/4_engrossed-amendment-senate.xml \
  tests/corpus/118-hr-4366/5_engrossed-amendment-house.xml \
  --format html -o hr4366-v4-v5.html

uv run python diff_bill.py compare \
  tests/corpus/118-hr-8752/1_reported-in-house.xml \
  tests/corpus/118-hr-8752/2_engrossed-in-house.xml \
  --format html -o hr8752-v1-v2.html

1. An appropriation is displayed and calculated as a rescission.

Open hr4366-v4-v5.html, select Financials – Version A, and find MEDICAL SERVICES. The opening $71 billion amount is labeled “rescission.”

The source provides $71 billion for medical services and separately rescinds $4,933,113,000 later in the paragraph. Version B changes that later rescission to $3,034,205,000. The classifier applies the later rescission language to the opening appropriation, and the comparison treats the $71 billion as negative money given out.

2. The comparison excludes a funding change already present in the ledger.

In the same report, find COMPENSATION AND PENSIONS in the financial comparison, then inspect its clauses in both version views.

The comparison shows a $10,416,509,000 increase from the opening appropriation. A separate advance appropriation changes from $181,390,281,000 to $182,310,515,000—a $920,234,000 increase—which is present in the ledger but excluded from the comparison.

Using only the first clause is an intentional implementation choice. However, the broader “Net change in money given out” label does not communicate that limitation. These amounts also concern different fiscal years, so this example should not be read as a proposed corrected bill-wide total.

3. A spending cap is selected as the account’s appropriation.

Open hr8752-v1-v2.html, select Financials – Version A, and find Coast Guard → OPERATIONS AND SUPPORT.

The displayed appropriation is $31 million. The source describes boat-related spending as “not to exceed a total of $31,000,000”; the account’s appropriation is $10,554,261,000.

The row does carry a review warning, but the $31 million still enters the comparison calculation. Both versions select that same figure, so this example demonstrates an incorrect selected amount, not an incorrect nonzero delta.

4. Allocations are incorrectly marked as amendments to existing law.

In that report, export the Version A financial CSV and inspect §211(a). The six allocations—including $600 million for physical barriers—have in_amended_law=true.

The source section allocates funds made available by this bill; it does not amend another law. Its six allocations sum exactly to the stated $1,390,338,000 parent amount. “As follows:” appears to trigger the amended-law classification. This is an incorrect label and export value; it is not evidence that these allocations should be added again to a funding total.

I’d like the next step to be resolving these cases in the research folder, with independently justified expected results. Matching the notebook’s outputs is useful regression evidence, but it does not establish that those outputs correctly represent the legislation.

Once we’re confident in the methodology, I’d like to resolve generating and validating the financial data separately from displaying it. The tables, totals, warnings, exports, and comparison behavior involve substantial design choices that deserve their own review.

Thank you again for the contribution. I’d welcome continuing this work through the research first.

…dc#115)

Schema 3.1 adds an optional financial field: for each version on its own, every dollar amount with its chain of headings, its clause, a type read from the wording by the research notebook's rules (classifier 1.0, moved into the package unchanged) and flags. A change still carries no money. PDF sections are read so they give the notebook's rows: gutter and running header off, page-seam hyphens rejoined, paragraphs split where the XML has nodes, quote marks removed for the rules only.

The report gains Inferred Financial Comparison, Financials - Version A and Version B, each with a CSV export and a not-authoritative alert. Find also searches collapsed rows; each tab keeps its scroll position; jump targets clear the sticky bar at its measured height; the change navigator and changes export show only on Changes and Full bill. The PDF heading keeps the whole bill title and an upload label is the file name as given. The civictechdc#671 no-money checks now look where civictechdc#671 drew the line, the Changes view. Canonical baselines re-pinned: the digests move by the new field only.
ADR 0023 (Proposed): the per-side ledger, reading PDF sections for the notebook's values, why the classifier is versioned and how it improves, the three views, and the open question on money in the API's JSON. ADR 0006 amended in place; schema doc at 3.1; TESTING.md gains check 8 and the financial rows pin; research README records that the classifier moved into the product.
@mattzamora
mattzamora force-pushed the feat/115-financial-views branch from 624e77e to f08e6cc Compare September 30, 2026 18:57

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants