diff --git a/docs/decisions/0001-structured-money-diff.md b/docs/decisions/0001-structured-money-diff.md index d7393454..9b21f6e8 100644 --- a/docs/decisions/0001-structured-money-diff.md +++ b/docs/decisions/0001-structured-money-diff.md @@ -44,10 +44,18 @@ table. Parse each bill version into the bill's own section structure and diff that, which answers question 1 for any bill type: added, removed, modified, and moved sections, -with formatting noise suppressed. For appropriations bills, parse a structured -model of accounts and amounts on top and diff it too, answering question 2: an -account-level table of paired old → new amounts with a breadcrumb path to each -account, plus flags for floor-amendment annotations. +with formatting noise suppressed. For appropriations bills, a structured model of +accounts and amounts sits on top and is diffed too, answering question 2: an +account-level table of old → new amounts with a path to each account. + +Question 2 is the goal, not yet the product. The paired-amount table was removed in +#681 (#671), because a paragraph's figures (appropriations, sub-allocations, +ceilings, rescissions) could not be told apart, and a figure with the wrong meaning +makes every total built on it wrong. The export now carries each node's dollar +figures per version and unpaired, which states what a version contains without +claiming what changed ([0006](0006-canonical-diff-contract.md)). The money model +returns in stages, each gated on every figure being typed and its type agreeing with +an official source ([0025](0025-financial-confidence-criteria.md), proposed). This splits the system into two layers, one per question: @@ -56,19 +64,23 @@ This splits the system into two layers, one per question: reconstruct the section structure, then diff sections. Most parser effort lives here. It is unavoidable for PDF-only inputs (draft and pre-introduction bills have no XML upstream), where the noise is severe. -2. **Money model (question 2)**: map normalized text to accounts and paired - amounts. A smaller slice of code, and the appropriations-specific layer. +2. **Money model (question 2)**: attach each dollar figure to its account, type + what it means, and pair amounts across versions. The appropriations-specific + layer, and the one whose correctness is hardest to establish. ## Consequences -- The tool answers both questions that generic differs cannot: a clean - section-level diff for any bill, and the account-level money table for - appropriations. This is the core capability and the reason the project exists. +- The tool can answer both questions that generic differs cannot: a clean + section-level diff for any bill, which ships today, and the account-level money + table for appropriations, which is gated as above. The second is the reason the + project exists, so the bar for shipping it is set by evidence rather than by + how much of it is built. - Extraction noise that a generic PDF redline would surface is the extraction layer's job list, not a defect of the approach. Suppressing it is required work, especially for PDF-only bills. - Fidelity effort should go where it is **load-bearing for correctness**, since a - misread or misattributed amount is a money error and the worst failure mode. The + misread, mistyped or misattributed amount is a money error and the worst failure + mode. The next priority is **verifiability**: link each amount to its exact source location so a reader can confirm it. Cosmetic reproduction of the source document's typography and layout is lower value, since the reader already has diff --git a/docs/decisions/0025-financial-confidence-criteria.md b/docs/decisions/0025-financial-confidence-criteria.md new file mode 100644 index 00000000..de82ea1f --- /dev/null +++ b/docs/decisions/0025-financial-confidence-criteria.md @@ -0,0 +1,251 @@ +# 25. Bring dollar figures back only when each one is typed and its type agrees with an official source + +- Status: Proposed +- Date: 2026-10-04 + +## Context + +DeltaTrack exists to answer two questions about a pair of bill versions: what text +changed, and what money changed ([0001](0001-structured-money-diff.md)). Today it answers +only the first. #681 (#671) removed the financial table and every paired amount from the +report and the export, because the pipeline could not say what a dollar figure *is*. The +rule since then has been that money stays out until the team is confident in how it is +added back. Nobody has written down what "confident" means, so there is no way to tell +when a proposal clears the bar. Without that, every attempt is judged case by case. The +research folder (`docs/research/financial-semantics/`) has no research question or exit +criteria either. + +**The meaning of a figure is the hard part, not finding it.** An appropriations paragraph +carries several kinds of number: the appropriation itself, sub-allocations carved out of +it, ceilings ("not to exceed $X"), rescissions of earlier money, advance appropriations +for a later fiscal year, and figures inside text the bill writes into another law. They +look identical on the page. A figure can appear in the source, the diff and the view with +the same value and still be wrong, because it was given the wrong meaning. Adding it to a +total then produces a wrong number. + +**The errors found so far are errors of meaning.** Review of #736 at `624e77ed` found +four, each traced to the bill text: + +1. **H.R. 4366, Medical Services (v4 to v5).** The opening $71 billion appropriation is + typed as a rescission, because "is hereby rescinded" appears later in the same + paragraph, applied to a separate $4.9 billion. The comparison counts the $71 billion + as negative money. +2. **H.R. 4366, Compensation and Pensions.** An advance appropriation rises by $920 + million, from $181.39 billion to $182.31 billion. It is in the ledger but left out of + the comparison, which uses only the first clause. It also funds a different fiscal + year from the opening figure, so the two cannot simply be added. +3. **H.R. 8752, Coast Guard Operations and Support (v1 to v2).** A $31 million + boat-related ceiling is selected as the account's appropriation instead of + $10,554,261,000. +4. **H.R. 8752, §211(a).** Six allocations that sum to their stated $1,390,338,000 parent + are flagged as amendments to existing law, because "as follows:" triggers that flag. + +Review of #724 found a fifth: its Title I rollup of H.R. 4366 came to $15.56 billion +against the committee report's $17.474 billion. The $1.91 billion gap is exactly the +figures after the first in six sections that each appropriate several named amounts. +The same rollup merged three different `SALARIES AND EXPENSES` accounts into one row, +because it grouped them by their leaf heading. + +These are four kinds of failure: the wrong sign (1), the wrong figure picked (3), a figure +missing (2 and #724), and a wrong label (4 and the merged accounts). Case 1 and the #724 +shortfall come straight from one design choice in the research classifier +(`classify_bill.py`). It assigns a type to a clause rather than to a figure, gives the +first clause the type of the whole paragraph, and keeps one figure per clause, so a +clause with nine appropriations has nowhere to record the other eight. Case 3 starts as a +pattern miss ("not to exceed a total of" escapes the ceiling rule), and the same choice +turns it into a wrong total: keeping one figure means the wrong pick displaces the right +one. Tuning patterns cannot fix the design, because the data has no slot for "this +figure is X." + +**The existing independent check would pass these errors.** [0009](0009-validation-ground-truth.md) +validates extraction against Senate committee reports. It counts an account as recalled +when the report's figure appears anywhere in the extraction under the right agency, and +it leaves out rescission rows. In case 3 the $10.55 billion figure is present, so the +account passes even though the $31 million cap was picked as the appropriation. That +check measures whether a figure was found. It says nothing about the meaning assigned +to it. + +**Official sources do not cover every version.** A Senate committee report covers the +Senate-reported version of a bill. House reports print their account tables as images. +A diff usually compares other versions: engrossed text, the chambers' amendments, the +enrolled bill. For many versions there is no independent source at all. + +**Two existing decisions constrain the answer.** [0006](0006-canonical-diff-contract.md) +removed paired amounts rather than caveating them. The export is read by machines (the +report invites a staffer to upload `diff.json` to an AI assistant), and a caveat in the +report or the documentation does not travel with the file. A "not authoritative" banner +over typed figures, as #736 proposes, runs into the same problem. +[0018](0018-text-triggers-are-financial-only.md) lets appropriations wording interpret a +figure but never decide which account it belongs to, so a figure filed under the wrong +account is a structure defect, not a typing one. + +## Decision + +Money returns to DeltaTrack in stages. Each stage ships when its criteria below pass as a +CI gate and the maintainer signs off. Being "confident" means exactly that: a passing +gate against official sources, with every disagreement explained. + +### Every figure gets its own type + +The unit of confidence is the single dollar figure, not the clause or the paragraph. Each +figure in a money-bearing provision is typed on two levels. + +Its **effect** is what totals depend on, and it is required in every stage: + +| Effect | Research labels (`classify_bill.py`) | +|---|---| +| Adds money | `appropriation`, including advance appropriations | +| Removes money | `rescission` | +| Neither | `sub_allocation`, `earmark`, `availability` (parts of a parent figure); `cap`; `transfer`; `authorization`; `fee`; `restriction`; `directive`; and figures inside amended law | +| Unresolved | `unknown` | + +Its **role** is the finer label in the right-hand column. A role is useful for the change +narrative #115 asks for, but only the effect is required to ship a stage. The research +vocabulary is reused rather than replaced. What changes is that the label attaches to a +figure, so a paragraph can hold an appropriation and a rescission at once. + +Two further properties are part of a figure's type, because a total is wrong without +them: + +- **The fiscal year it funds.** An advance appropriation and a current-year appropriation + in the same account are not summed together. +- **The account it belongs to**, identified by its full position in the bill's + structure, not its leaf heading. Per 0018 this comes from the tree, never from the + wording. + +### The evidence is official sources, matched on meaning + +The reference is an official source written independently of the bill text: + +- the Senate committee report's account tables, the same documents 0009 uses; +- the explanatory statement for an enacted bill, which states account amounts for the + enrolled version. It is a candidate; its availability as parseable text has to be + confirmed per bill. + +Matching is on meaning, which is stricter than 0009's recall rule: + +- the figure DeltaTrack types as an account's appropriation equals the source's + appropriation for that account; +- each rescission row in the source matches a figure typed as a rescission, and each + limitation row matches a figure typed as a ceiling; +- for totals, a title's typed appropriations reconcile to the source's `Total, title` + row. + +Hand-tagged examples, such as the five cases above, become acceptance cases: regression +tests with expected results justified from the bill text and, where one exists, the +official source. They are not evidence of accuracy. Hand tagging is often wrong, and +expected values written by someone reading the same bill as the classifier share its +blind spots (0009). + +**Versions with no official source are covered by extrapolation, and the limits are +stated.** The engine is deterministic ([0008](0008-deterministic-engine.md)), so a +provision whose text is identical to one in a validated version gets the same types. +Where a provision's text changed, confidence rests on how the classifier performs across +the validation set rather than on direct evidence for that version. + +### The validation set covers every kind of bill that appropriates + +- **Every vehicle kind:** regular, omnibus, supplemental, continuing resolution, + reconciliation, and an authorizing law with an appropriations division. These are the + `vehicle` values #739 proposes for the corpus manifest. +- **At least two versions of each bill**, because a diff compares versions. +- **An authorization bill as a negative control:** no figure that is authorized rather + than appropriated may be typed as adding money. +- **A holdout year** that the rules were not tuned on, as 0009 requires. +- **PDF against XML:** where the PDF and XML trees and provision text agree, the figures + get identical types. Where they disagree, no agreement is required. That gap belongs + to the structural work, not this record. + +### The bar: every disagreement explained, and none is ours + +As in 0009, the threshold is not a percentage. Every disagreement with the official +source is hand-traced, and none may be a DeltaTrack error. A disagreement is acceptable +only when the source and the bill legitimately state the figure differently (an +indefinite "such sums" account, a total the bill itemizes, a typo in the report). A +figure left `unresolved` where the source states its type counts as a disagreement and +as our error, so marking everything unresolved cannot pass. A percentage target invites +tuning toward the number and hides which errors matter. + +**Zero tolerance applies to any error that changes a number a reader would add up:** + +- the wrong effect: an appropriation typed as a rescission, or the reverse; +- a figure that adds no money counted as if it did, such as a ceiling, a sub-allocation + or an authorization; +- money filed under the wrong account; +- a missing figure that is not flagged. + +Label errors that change no total, such as a wrong role or a wrong amended-law flag, +still count against the bar above. Each must be traced and fixed, but one does not block +a stage on its own. + +### Uncertainty lives in the data, per figure + +A figure the classifier cannot type is marked `unresolved` in the document, not omitted +and not guessed. The document states its own coverage: how many figures each version +holds and how many are typed. A total is shown only when no figure beneath it is +unresolved. A consumer holding only `diff.json` can therefore see what is known and what +is not, which is the property 0006 requires. A banner can still tell a reader to check +the bill text, but it does not count as handling uncertainty. + +### Money returns in order of how strong the claim is + +1. **Each figure's type, per version, unpaired.** This states what a version contains. + Gate: the matching rules, the validation set and the zero-tolerance rule above. +2. **Totals**, rolled up from account to title. Gate: title totals reconcile to the + official source. +3. **Paired changes between versions.** 0006 already calls pairing a separate claim, and + case 2 is a pairing failure, not a typing one. Its criteria are not set here. They + need their own decision once stage 1 has shipped. + +Each stage ships to the export and the report together. Shipping by surface (export +first, then the report) was considered and rejected: the export is the surface a caveat +cannot follow, so it is not the safer place to start. + +### Sign-off + +The maintainer approves these criteria, and later approves each new hand-traced +disagreement before it joins the explained list. Whether a stage passes is shown by the +gate, not by judgement. + +## Alternatives + +- **A percentage threshold**, for example 95% of figures typed correctly. Rejected: it + allows a wrong-sign error inside the 5%, and it rewards tuning toward the sample. 0009 + holds extraction to the same every-miss-explained bar this record adopts. +- **Reuse the 0009 recall check unchanged.** Rejected: it passes case 3, because it asks + whether a figure is present, not what it was taken to mean. +- **A hand-labelled set as the primary reference.** Rejected as the measure of accuracy. + Hand tagging is often inaccurate, and expected values written from the same bill text + share the classifier's blind spots. Kept for acceptance cases. +- **CBO cost estimates as a reference.** Rejected: they report budget authority at the + subcommittee or title level after scorekeeping adjustments, so most disagreements would + be accounting conventions rather than typing errors. +- **Show typed figures now, behind a "not authoritative" banner (#736).** Rejected for + the reasons 0006 removed paired amounts: the banner does not travel with the export, + and no caveat makes a wrong-sign figure safe to add up. +- **Return everything at once.** Rejected: pairing has failure modes of its own, and + waiting on them would hold back per-version types that may be ready sooner. + +## Consequences + +- There is a definite answer to "is it ready?" A proposal to show money names the stage + it targets and shows that stage's gate passing. +- **The research classifier needs a per-figure data model** before any stage can pass. + That is a design change, not a pattern change. Its labels carry over. +- **The validation harness needs new work.** Matching on meaning against committee-report + tables means parsing the rescission and limitation rows that the recall check + currently excludes, and checking which figure was typed, not just whether it was + found. The committee-report answer keys also need to be guarded against drift (#293), + since this record makes them load-bearing for money as well as extraction. +- **#736's per-version ledger matches stage 1 in shape**, but it types clauses and leans + on a banner. To ship it would need per-figure types and a passing stage 1 gate. Its ADR + 0023 would need to defer to this record. +- **Coverage depends on #739** for the `vehicle` field and the added bill kinds. +- **A known limitation remains.** The provisions a diff surfaces are the ones whose text + changed, and in a version with no official source those are exactly the figures with + no direct evidence. Extrapolation from the validation set is the best available, and + the report should not present those figures as checked against a source. +- **Money stays out of the product longer** than a banner-first approach would allow. + That is the intended cost. +- [0001](0001-structured-money-diff.md) is updated to say that its money table is the + goal and that it is gated by this record. diff --git a/docs/decisions/README.md b/docs/decisions/README.md index e2975bda..3ff18ae6 100644 --- a/docs/decisions/README.md +++ b/docs/decisions/README.md @@ -103,3 +103,4 @@ presents itself as the architecture in force. | [0019](0019-observation-identity.md) | Accepted | Identify a parsed observation by its source, its parser revision and its ordinal; never by its text | | [0020](0020-matching-stages.md) | Accepted | Separate retrieval, correspondence evidence, correspondence assignment and change classification | | [0021](0021-naming-authority-and-boundaries.md) | Accepted | Name things in the vocabulary an outside reader already speaks, scoped to the boundary being named | +| [0025](0025-financial-confidence-criteria.md) | Proposed | Bring dollar figures back only when each one is typed and its type agrees with an official source | diff --git a/docs/research/financial-semantics/README.md b/docs/research/financial-semantics/README.md index b0cbe22c..ed223bf1 100644 --- a/docs/research/financial-semantics/README.md +++ b/docs/research/financial-semantics/README.md @@ -2,6 +2,42 @@ Financial classifier and analysis tools for DeltaTrack bill data. +## Research question + +Can every dollar figure in a bill version be given the right meaning (what kind of +money it is, which account it belongs to, and which fiscal year it funds), and can +that meaning be checked against an official source written independently of the +classifier? + +Finding a figure is not the question. Dollar amounts were removed from the report in +#681 (#671) because figures were found but given the wrong meaning, and a wrong +meaning makes every total built on it wrong. The bar for bringing them back is +[ADR 0025](../../decisions/0025-financial-confidence-criteria.md) (proposed); this +folder is where the work to meet it happens. Tracking: epic #147. + +## Exit criteria + +The research is done, and its classifier is ready to move into the product, when the +first stage of ADR 0025 passes: + +1. **Every figure is typed on its own**, not by clause or paragraph, with its effect + (adds money, removes money, neither, or unresolved), its role (the labels in + `classify_bill.py`), its fiscal year and its account. +2. **The types are matched against official sources on meaning**: the figure typed as + an account's appropriation equals the committee report's appropriation, and + rescission and limitation rows match figures of that type. +3. **The validation set covers every kind of bill that appropriates**, with at least + two versions each, an authorization bill as a negative control and a holdout year. +4. **No zero-tolerance error remains**: no wrong sign, no non-money figure counted as + money, no money under the wrong account, no unflagged missing figure. +5. **Every disagreement with the source is hand-traced**, and none is a classifier + error. +6. **The known review cases pass as acceptance tests**: the four from #736 and the + Title I shortfall from #724, each with expected results justified from the bill + text. + +ADR 0025 holds the reasoning behind each criterion. + ## Files | File | Purpose |