From d5cdf330072b8c71952d81845bbd23fdea3afdea Mon Sep 17 00:00:00 2001 From: Will Hea <1123070+willhea@users.noreply.github.com> Date: Sun, 4 Oct 2026 23:54:46 -0400 Subject: [PATCH] docs(adr): propose 0025, the bar for bringing dollar figures back Money has been out of the report since #681 until the team is "confident" in how it is added back, and nothing said what confident means. ADR 0025 (Proposed) defines it: every figure typed on its own, matched on meaning against official sources, zero tolerance for errors that change a total, uncertainty carried per figure in the data, and money returning in stages ordered by claim strength. ADR 0001 is rewritten to say its money table is the goal and is gated by 0025. The financial-semantics research README gains a research question and exit criteria. Refs #147 Co-Authored-By: Claude Opus 5.5 --- docs/decisions/0001-structured-money-diff.md | 32 ++- .../0025-financial-confidence-criteria.md | 251 ++++++++++++++++++ docs/decisions/README.md | 1 + docs/research/financial-semantics/README.md | 36 +++ 4 files changed, 310 insertions(+), 10 deletions(-) create mode 100644 docs/decisions/0025-financial-confidence-criteria.md diff --git a/docs/decisions/0001-structured-money-diff.md b/docs/decisions/0001-structured-money-diff.md index d7393454..9b21f6e8 100644 --- a/docs/decisions/0001-structured-money-diff.md +++ b/docs/decisions/0001-structured-money-diff.md @@ -44,10 +44,18 @@ table. Parse each bill version into the bill's own section structure and diff that, which answers question 1 for any bill type: added, removed, modified, and moved sections, -with formatting noise suppressed. For appropriations bills, parse a structured -model of accounts and amounts on top and diff it too, answering question 2: an -account-level table of paired old → new amounts with a breadcrumb path to each -account, plus flags for floor-amendment annotations. +with formatting noise suppressed. For appropriations bills, a structured model of +accounts and amounts sits on top and is diffed too, answering question 2: an +account-level table of old → new amounts with a path to each account. + +Question 2 is the goal, not yet the product. The paired-amount table was removed in +#681 (#671), because a paragraph's figures (appropriations, sub-allocations, +ceilings, rescissions) could not be told apart, and a figure with the wrong meaning +makes every total built on it wrong. The export now carries each node's dollar +figures per version and unpaired, which states what a version contains without +claiming what changed ([0006](0006-canonical-diff-contract.md)). The money model +returns in stages, each gated on every figure being typed and its type agreeing with +an official source ([0025](0025-financial-confidence-criteria.md), proposed). This splits the system into two layers, one per question: @@ -56,19 +64,23 @@ This splits the system into two layers, one per question: reconstruct the section structure, then diff sections. Most parser effort lives here. It is unavoidable for PDF-only inputs (draft and pre-introduction bills have no XML upstream), where the noise is severe. -2. **Money model (question 2)**: map normalized text to accounts and paired - amounts. A smaller slice of code, and the appropriations-specific layer. +2. **Money model (question 2)**: attach each dollar figure to its account, type + what it means, and pair amounts across versions. The appropriations-specific + layer, and the one whose correctness is hardest to establish. ## Consequences -- The tool answers both questions that generic differs cannot: a clean - section-level diff for any bill, and the account-level money table for - appropriations. This is the core capability and the reason the project exists. +- The tool can answer both questions that generic differs cannot: a clean + section-level diff for any bill, which ships today, and the account-level money + table for appropriations, which is gated as above. The second is the reason the + project exists, so the bar for shipping it is set by evidence rather than by + how much of it is built. - Extraction noise that a generic PDF redline would surface is the extraction layer's job list, not a defect of the approach. Suppressing it is required work, especially for PDF-only bills. - Fidelity effort should go where it is **load-bearing for correctness**, since a - misread or misattributed amount is a money error and the worst failure mode. The + misread, mistyped or misattributed amount is a money error and the worst failure + mode. The next priority is **verifiability**: link each amount to its exact source location so a reader can confirm it. Cosmetic reproduction of the source document's typography and layout is lower value, since the reader already has diff --git a/docs/decisions/0025-financial-confidence-criteria.md b/docs/decisions/0025-financial-confidence-criteria.md new file mode 100644 index 00000000..de82ea1f --- /dev/null +++ b/docs/decisions/0025-financial-confidence-criteria.md @@ -0,0 +1,251 @@ +# 25. Bring dollar figures back only when each one is typed and its type agrees with an official source + +- Status: Proposed +- Date: 2026-10-04 + +## Context + +DeltaTrack exists to answer two questions about a pair of bill versions: what text +changed, and what money changed ([0001](0001-structured-money-diff.md)). Today it answers +only the first. #681 (#671) removed the financial table and every paired amount from the +report and the export, because the pipeline could not say what a dollar figure *is*. The +rule since then has been that money stays out until the team is confident in how it is +added back. Nobody has written down what "confident" means, so there is no way to tell +when a proposal clears the bar. Without that, every attempt is judged case by case. The +research folder (`docs/research/financial-semantics/`) has no research question or exit +criteria either. + +**The meaning of a figure is the hard part, not finding it.** An appropriations paragraph +carries several kinds of number: the appropriation itself, sub-allocations carved out of +it, ceilings ("not to exceed $X"), rescissions of earlier money, advance appropriations +for a later fiscal year, and figures inside text the bill writes into another law. They +look identical on the page. A figure can appear in the source, the diff and the view with +the same value and still be wrong, because it was given the wrong meaning. Adding it to a +total then produces a wrong number. + +**The errors found so far are errors of meaning.** Review of #736 at `624e77ed` found +four, each traced to the bill text: + +1. **H.R. 4366, Medical Services (v4 to v5).** The opening $71 billion appropriation is + typed as a rescission, because "is hereby rescinded" appears later in the same + paragraph, applied to a separate $4.9 billion. The comparison counts the $71 billion + as negative money. +2. **H.R. 4366, Compensation and Pensions.** An advance appropriation rises by $920 + million, from $181.39 billion to $182.31 billion. It is in the ledger but left out of + the comparison, which uses only the first clause. It also funds a different fiscal + year from the opening figure, so the two cannot simply be added. +3. **H.R. 8752, Coast Guard Operations and Support (v1 to v2).** A $31 million + boat-related ceiling is selected as the account's appropriation instead of + $10,554,261,000. +4. **H.R. 8752, §211(a).** Six allocations that sum to their stated $1,390,338,000 parent + are flagged as amendments to existing law, because "as follows:" triggers that flag. + +Review of #724 found a fifth: its Title I rollup of H.R. 4366 came to $15.56 billion +against the committee report's $17.474 billion. The $1.91 billion gap is exactly the +figures after the first in six sections that each appropriate several named amounts. +The same rollup merged three different `SALARIES AND EXPENSES` accounts into one row, +because it grouped them by their leaf heading. + +These are four kinds of failure: the wrong sign (1), the wrong figure picked (3), a figure +missing (2 and #724), and a wrong label (4 and the merged accounts). Case 1 and the #724 +shortfall come straight from one design choice in the research classifier +(`classify_bill.py`). It assigns a type to a clause rather than to a figure, gives the +first clause the type of the whole paragraph, and keeps one figure per clause, so a +clause with nine appropriations has nowhere to record the other eight. Case 3 starts as a +pattern miss ("not to exceed a total of" escapes the ceiling rule), and the same choice +turns it into a wrong total: keeping one figure means the wrong pick displaces the right +one. Tuning patterns cannot fix the design, because the data has no slot for "this +figure is X." + +**The existing independent check would pass these errors.** [0009](0009-validation-ground-truth.md) +validates extraction against Senate committee reports. It counts an account as recalled +when the report's figure appears anywhere in the extraction under the right agency, and +it leaves out rescission rows. In case 3 the $10.55 billion figure is present, so the +account passes even though the $31 million cap was picked as the appropriation. That +check measures whether a figure was found. It says nothing about the meaning assigned +to it. + +**Official sources do not cover every version.** A Senate committee report covers the +Senate-reported version of a bill. House reports print their account tables as images. +A diff usually compares other versions: engrossed text, the chambers' amendments, the +enrolled bill. For many versions there is no independent source at all. + +**Two existing decisions constrain the answer.** [0006](0006-canonical-diff-contract.md) +removed paired amounts rather than caveating them. The export is read by machines (the +report invites a staffer to upload `diff.json` to an AI assistant), and a caveat in the +report or the documentation does not travel with the file. A "not authoritative" banner +over typed figures, as #736 proposes, runs into the same problem. +[0018](0018-text-triggers-are-financial-only.md) lets appropriations wording interpret a +figure but never decide which account it belongs to, so a figure filed under the wrong +account is a structure defect, not a typing one. + +## Decision + +Money returns to DeltaTrack in stages. Each stage ships when its criteria below pass as a +CI gate and the maintainer signs off. Being "confident" means exactly that: a passing +gate against official sources, with every disagreement explained. + +### Every figure gets its own type + +The unit of confidence is the single dollar figure, not the clause or the paragraph. Each +figure in a money-bearing provision is typed on two levels. + +Its **effect** is what totals depend on, and it is required in every stage: + +| Effect | Research labels (`classify_bill.py`) | +|---|---| +| Adds money | `appropriation`, including advance appropriations | +| Removes money | `rescission` | +| Neither | `sub_allocation`, `earmark`, `availability` (parts of a parent figure); `cap`; `transfer`; `authorization`; `fee`; `restriction`; `directive`; and figures inside amended law | +| Unresolved | `unknown` | + +Its **role** is the finer label in the right-hand column. A role is useful for the change +narrative #115 asks for, but only the effect is required to ship a stage. The research +vocabulary is reused rather than replaced. What changes is that the label attaches to a +figure, so a paragraph can hold an appropriation and a rescission at once. + +Two further properties are part of a figure's type, because a total is wrong without +them: + +- **The fiscal year it funds.** An advance appropriation and a current-year appropriation + in the same account are not summed together. +- **The account it belongs to**, identified by its full position in the bill's + structure, not its leaf heading. Per 0018 this comes from the tree, never from the + wording. + +### The evidence is official sources, matched on meaning + +The reference is an official source written independently of the bill text: + +- the Senate committee report's account tables, the same documents 0009 uses; +- the explanatory statement for an enacted bill, which states account amounts for the + enrolled version. It is a candidate; its availability as parseable text has to be + confirmed per bill. + +Matching is on meaning, which is stricter than 0009's recall rule: + +- the figure DeltaTrack types as an account's appropriation equals the source's + appropriation for that account; +- each rescission row in the source matches a figure typed as a rescission, and each + limitation row matches a figure typed as a ceiling; +- for totals, a title's typed appropriations reconcile to the source's `Total, title` + row. + +Hand-tagged examples, such as the five cases above, become acceptance cases: regression +tests with expected results justified from the bill text and, where one exists, the +official source. They are not evidence of accuracy. Hand tagging is often wrong, and +expected values written by someone reading the same bill as the classifier share its +blind spots (0009). + +**Versions with no official source are covered by extrapolation, and the limits are +stated.** The engine is deterministic ([0008](0008-deterministic-engine.md)), so a +provision whose text is identical to one in a validated version gets the same types. +Where a provision's text changed, confidence rests on how the classifier performs across +the validation set rather than on direct evidence for that version. + +### The validation set covers every kind of bill that appropriates + +- **Every vehicle kind:** regular, omnibus, supplemental, continuing resolution, + reconciliation, and an authorizing law with an appropriations division. These are the + `vehicle` values #739 proposes for the corpus manifest. +- **At least two versions of each bill**, because a diff compares versions. +- **An authorization bill as a negative control:** no figure that is authorized rather + than appropriated may be typed as adding money. +- **A holdout year** that the rules were not tuned on, as 0009 requires. +- **PDF against XML:** where the PDF and XML trees and provision text agree, the figures + get identical types. Where they disagree, no agreement is required. That gap belongs + to the structural work, not this record. + +### The bar: every disagreement explained, and none is ours + +As in 0009, the threshold is not a percentage. Every disagreement with the official +source is hand-traced, and none may be a DeltaTrack error. A disagreement is acceptable +only when the source and the bill legitimately state the figure differently (an +indefinite "such sums" account, a total the bill itemizes, a typo in the report). A +figure left `unresolved` where the source states its type counts as a disagreement and +as our error, so marking everything unresolved cannot pass. A percentage target invites +tuning toward the number and hides which errors matter. + +**Zero tolerance applies to any error that changes a number a reader would add up:** + +- the wrong effect: an appropriation typed as a rescission, or the reverse; +- a figure that adds no money counted as if it did, such as a ceiling, a sub-allocation + or an authorization; +- money filed under the wrong account; +- a missing figure that is not flagged. + +Label errors that change no total, such as a wrong role or a wrong amended-law flag, +still count against the bar above. Each must be traced and fixed, but one does not block +a stage on its own. + +### Uncertainty lives in the data, per figure + +A figure the classifier cannot type is marked `unresolved` in the document, not omitted +and not guessed. The document states its own coverage: how many figures each version +holds and how many are typed. A total is shown only when no figure beneath it is +unresolved. A consumer holding only `diff.json` can therefore see what is known and what +is not, which is the property 0006 requires. A banner can still tell a reader to check +the bill text, but it does not count as handling uncertainty. + +### Money returns in order of how strong the claim is + +1. **Each figure's type, per version, unpaired.** This states what a version contains. + Gate: the matching rules, the validation set and the zero-tolerance rule above. +2. **Totals**, rolled up from account to title. Gate: title totals reconcile to the + official source. +3. **Paired changes between versions.** 0006 already calls pairing a separate claim, and + case 2 is a pairing failure, not a typing one. Its criteria are not set here. They + need their own decision once stage 1 has shipped. + +Each stage ships to the export and the report together. Shipping by surface (export +first, then the report) was considered and rejected: the export is the surface a caveat +cannot follow, so it is not the safer place to start. + +### Sign-off + +The maintainer approves these criteria, and later approves each new hand-traced +disagreement before it joins the explained list. Whether a stage passes is shown by the +gate, not by judgement. + +## Alternatives + +- **A percentage threshold**, for example 95% of figures typed correctly. Rejected: it + allows a wrong-sign error inside the 5%, and it rewards tuning toward the sample. 0009 + holds extraction to the same every-miss-explained bar this record adopts. +- **Reuse the 0009 recall check unchanged.** Rejected: it passes case 3, because it asks + whether a figure is present, not what it was taken to mean. +- **A hand-labelled set as the primary reference.** Rejected as the measure of accuracy. + Hand tagging is often inaccurate, and expected values written from the same bill text + share the classifier's blind spots. Kept for acceptance cases. +- **CBO cost estimates as a reference.** Rejected: they report budget authority at the + subcommittee or title level after scorekeeping adjustments, so most disagreements would + be accounting conventions rather than typing errors. +- **Show typed figures now, behind a "not authoritative" banner (#736).** Rejected for + the reasons 0006 removed paired amounts: the banner does not travel with the export, + and no caveat makes a wrong-sign figure safe to add up. +- **Return everything at once.** Rejected: pairing has failure modes of its own, and + waiting on them would hold back per-version types that may be ready sooner. + +## Consequences + +- There is a definite answer to "is it ready?" A proposal to show money names the stage + it targets and shows that stage's gate passing. +- **The research classifier needs a per-figure data model** before any stage can pass. + That is a design change, not a pattern change. Its labels carry over. +- **The validation harness needs new work.** Matching on meaning against committee-report + tables means parsing the rescission and limitation rows that the recall check + currently excludes, and checking which figure was typed, not just whether it was + found. The committee-report answer keys also need to be guarded against drift (#293), + since this record makes them load-bearing for money as well as extraction. +- **#736's per-version ledger matches stage 1 in shape**, but it types clauses and leans + on a banner. To ship it would need per-figure types and a passing stage 1 gate. Its ADR + 0023 would need to defer to this record. +- **Coverage depends on #739** for the `vehicle` field and the added bill kinds. +- **A known limitation remains.** The provisions a diff surfaces are the ones whose text + changed, and in a version with no official source those are exactly the figures with + no direct evidence. Extrapolation from the validation set is the best available, and + the report should not present those figures as checked against a source. +- **Money stays out of the product longer** than a banner-first approach would allow. + That is the intended cost. +- [0001](0001-structured-money-diff.md) is updated to say that its money table is the + goal and that it is gated by this record. diff --git a/docs/decisions/README.md b/docs/decisions/README.md index e2975bda..3ff18ae6 100644 --- a/docs/decisions/README.md +++ b/docs/decisions/README.md @@ -103,3 +103,4 @@ presents itself as the architecture in force. | [0019](0019-observation-identity.md) | Accepted | Identify a parsed observation by its source, its parser revision and its ordinal; never by its text | | [0020](0020-matching-stages.md) | Accepted | Separate retrieval, correspondence evidence, correspondence assignment and change classification | | [0021](0021-naming-authority-and-boundaries.md) | Accepted | Name things in the vocabulary an outside reader already speaks, scoped to the boundary being named | +| [0025](0025-financial-confidence-criteria.md) | Proposed | Bring dollar figures back only when each one is typed and its type agrees with an official source | diff --git a/docs/research/financial-semantics/README.md b/docs/research/financial-semantics/README.md index b0cbe22c..ed223bf1 100644 --- a/docs/research/financial-semantics/README.md +++ b/docs/research/financial-semantics/README.md @@ -2,6 +2,42 @@ Financial classifier and analysis tools for DeltaTrack bill data. +## Research question + +Can every dollar figure in a bill version be given the right meaning (what kind of +money it is, which account it belongs to, and which fiscal year it funds), and can +that meaning be checked against an official source written independently of the +classifier? + +Finding a figure is not the question. Dollar amounts were removed from the report in +#681 (#671) because figures were found but given the wrong meaning, and a wrong +meaning makes every total built on it wrong. The bar for bringing them back is +[ADR 0025](../../decisions/0025-financial-confidence-criteria.md) (proposed); this +folder is where the work to meet it happens. Tracking: epic #147. + +## Exit criteria + +The research is done, and its classifier is ready to move into the product, when the +first stage of ADR 0025 passes: + +1. **Every figure is typed on its own**, not by clause or paragraph, with its effect + (adds money, removes money, neither, or unresolved), its role (the labels in + `classify_bill.py`), its fiscal year and its account. +2. **The types are matched against official sources on meaning**: the figure typed as + an account's appropriation equals the committee report's appropriation, and + rescission and limitation rows match figures of that type. +3. **The validation set covers every kind of bill that appropriates**, with at least + two versions each, an authorization bill as a negative control and a holdout year. +4. **No zero-tolerance error remains**: no wrong sign, no non-money figure counted as + money, no money under the wrong account, no unflagged missing figure. +5. **Every disagreement with the source is hand-traced**, and none is a classifier + error. +6. **The known review cases pass as acceptance tests**: the four from #736 and the + Title I shortfall from #724, each with expected results justified from the bill + text. + +ADR 0025 holds the reasoning behind each criterion. + ## Files | File | Purpose |