Develop to Prod: Increase Diff Granularity - #603
Merged
Merged
Conversation
…L ones The updater skipped versions without XML, so a PDF-only version got no committee_report entry at all. Report metadata describes the legislative version; it does not depend on whether DeltaTrack happens to hold that version's XML, PDF, or both, and #295 is about making the corpus self-describing at (bill, version, chamber). One corpus version was affected, and it is the example the issue uses to establish version/chamber pairing: 114-hr-2029 / 4_reported-in-senate, formats = ["pdf"]. It is a Senate-reported print, so the Senate committee report explains it -- the same S. Rept. 114-57 the Senate amendment that follows already carried. All 57 versions now carry an entry: 35 paired with a report, 22 recording citation = "none" with the reason none applies. Two gates, because absence and emptiness fail differently. One asserts every manifested version has an entry at all, since a missing key is indistinguishable from "no report applies" to a reader. The other asserts each entry says something: real sources, or exactly one sentinel carrying a reason. The first also asserts the corpus still contains a non-XML version, so it cannot quietly stop testing the format independence it exists for. Regenerating offline exposed a second defect, introduced by the previous commit rather than this one. Offline mode recovers a bill's report sources by reading back the pairings still recorded, so a bill whose every pairing the lineage rule dropped reads as a bill that never had reports. 118-hr-2882 keeps H. Rept. 118-364 on no version, and an offline run duly rewrote its three accurate reasons ("predates the committee report", "no senate committee report explains this stage", "authored in response to the other chamber's amendment") into "bill has no committee reports", which is false about a bill that has one and reads as authoritative. Offline mode now fills only versions that lack an entry when it can recover nothing, and says so; it never overwrites a reason computed when the sources were known. --refresh remains the authoritative answer and produced the committed manifest. An offline run is now idempotent against it, verified. The three reasons are pinned by a test that fails on the generic wording. The first attempt to prove the completeness gate could fail did not: restoring the XML-only filter and re-running the updater left the existing entry in place, so the condition never occurred and the gate passed for the wrong reason. Removing the entry from a manifested version -- the scenario the gate is actually for, a version added without a pairing -- turns it red naming the version and its formats. Degrading a reason to the generic wording turns the pin red. Lineage, conference, granule, error-page and orphan-fixture behavior are unchanged. Full suite: 2594 passed, 31 skipped (unchanged), 15 xfailed (rc=0), ruff clean. Refs #295 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add Content-Security-Policy header (#282)
) The #500 xfail test only asserted the account (OPERATIONS AND SUPPORT) survived catchline suppression, but the bug also drops the agency (MANAGEMENT DIRECTORATE). Update the test to assert both using inline filtering (not _by_kind) so a partial fix (restoring only the account) remains XFAIL.
Review findings on #531, all four addressed. The shell-version check asserted an UPPER BOUND on the fragment count, which zero satisfies. A fragment extractor that found nothing would have left both of its assertions true (0 <= 2, and nothing missing out of nothing), so the five shell cases would have gone green while asserting nothing -- the same fail-open channel the check was written to close, reintroduced one line below the comment explaining it. The count is now asserted as equal, which is safe here because the fixtures are committed and immutable: the number is a fact about a specific file, not a baseline that drifts. Confirmed by re-running the fault injection: with fragment extraction returning nothing, the module now fails all 24 cases where it previously failed 19 and passed the 5 shells. The calibration script judged every version against the global floor, so the one version on an approved degraded floor printed "BELOW FLOOR" on a case the suite passes. A warning that fires on a known-good state trains the reader to discount the warning that matters. It now reports each version against the floor the test actually holds it to, keeps the degraded version out of the healthy-population figure, says so, and exits non-zero only when something genuinely needs attention. TESTING.md said the check confirms every sentence of the official text appears in the PDF. It compares body-prose passages of eight words or more, deduplicated, excluding the table of contents and quoted blocks, by containment rather than position. Each of those narrowings is now stated with the reason it exists, and the limits section says the check confirms a passage is present somewhere in the version, not that it appears in the right place. The comment above _CORPUS_EXPANDING_MODULES still described "these two modules" after the dict grew to three. Its case-count figures belong to the two original modules and are left attributed to them rather than extended to a third by guesswork. Refs #7 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
test: assert both account and agency survive catchline suppression (#530)
The Financial Summary lists one row per amount entry, so on a real appropriations bill it fills the first several screens and the change cards start below it. It now renders inside a <details> that is closed by default, with the entry count in the summary so its size stays visible while collapsed. revealCard already walks ancestor <details>, so find-in-page hits and #change-N jumps still open it. Relocating the table to its own tab stays open as a later option; this is the smaller fix that makes the report readable now. Committed example reports regenerated and the served sample re-copied from the PDF example, per the freshness gate in test_committed_examples. Closes #2 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The caret sat on the <summary>, so it inherited body text and rendered at the sidebar navigation's 11px next to a 24px serif heading. At that size it read as decoration and the heading did not look like a control. Moving it onto the <h2> lets it be sized in `em` against the heading, so it tracks the heading's scale rather than being pinned to a number. A hover background and caret color change make the target legible as a toggle. Refs #2 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidebar navigation, change card groups, and the full-bill table of contents all pinned their caret to 10px. Beside 13px sidebar text that reads as a control; beside a 24px heading it reads as decoration, which is why the Financial Summary heading did not look clickable. Each caret is now sized in `em` against its own label, so the affordance stays proportionate wherever it appears. The browser gate asserts the caret-to-label ratio is the same across all three sizes rather than pinning pixels, which would re-encode the numbers the rule exists to remove. A per-caret band alone was too weak: a caret pinned to 10px still sits inside any sane band next to 13px text, so the spread between carets is what the assertion rests on. Confirmed by re-pinning the sidebar caret, which the spread check catches and the band check did not. Refs #2 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The caret was declared four times, once per collapsible, so keeping them in proportion was a convention nothing enforced, and the browser test asserting the caret-to-label ratio existed only to catch that drift. A test is the wrong tool for a rule CSS can hold structurally. Each collapsible now opts in with class="disclosure" on whichever element carries the label (the <summary>, or a heading inside it), and a single rule sizes the caret in `em` and flips the glyph on [open]. The four duplicated declarations are gone, so the drift the test guarded is no longer expressible, and the test is removed with them. The rendered output is unchanged: measured caret sizes are still 11.05px/13px, 13.6px/16px and 20.4px/24px. The markup assertions that pinned a class-less <summary>/<h2> are updated, and two summary-text regexes now allow attributes. Also: the first draft of the new CSS comment used the phrase "Financial Summary", which broke two tests asserting that phrase is absent from a report with no amounts. The stylesheet ships inside every report, so comment prose is part of the output. The comment now says so. Refs #2 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The report tab is opened on the Compare click, because a pop-up opened later is blocked (#41), so it necessarily sits empty until the server finishes rendering. On a large bill that was tens of seconds of a blank tab, indistinguishable from a crash, inviting the reader to close it or re-click a render about to finish. The tab now carries a self-contained "Diff in progress" placeholder with an animated spinner, overwritten by the report when it arrives. Self-contained because a tab written via document.write has no origin to resolve relative URLs from. The opener link is severed on the final write rather than the first. The browser test holds `fetch` mid-flight (the pattern the upload tests already use) and asserts both the placeholder and its replacement by the report. Closes #490 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`CORPUS_SWEEP=1` failed on four `_KNOWN_DUPLICATE_COUNTS` entries that had drifted below their true counts. The entries name bills present only in a local `bills/` tree, so no CI run and no clean checkout ever evaluated them and nothing could hold them current, while the parser reaching further (#465, #146, #188) raised the real counts underneath them. Drop the six non-manifested keys rather than repinning them. Repinning restores the mode until the next parser change and leaves the same gap open; and half of those keys were unreachable even under the sweep, because `sweep_bill_dirs` yields one directory per bill id with the committed copy winning, so a download-only version of a partly-committed bill is shadowed. Measured on a fully fetched machine, only the two 116-hr-133 entries were reachable at all. A sweep-only file is now reported rather than asserted: it is still parsed, so a crash or empty tree still fails, and `-rs` prints the measured count. This is the treatment `test_corpus_tree_properties` already gives its PDF-layout registry for the same reason — a registry calibrated on the committed corpus cannot judge an uncalibrated superset. `test_known_duplicate_counts_names_manifest_fixtures` keeps the class closed. It keys on the new `manifest_xml_ids()` rather than the collected file list, which widens with the sweep and would let a sweep-only key look live on a fetched machine — the same fail-open it exists to close. Also correct TESTING.md's "strict superset" claim, which the by-bill widening contradicts, and record that these dicts are `<=` ceilings: a parser change that reduces collisions leaves them silently loose, as #474 did to several committed values days after #482 measured them. Verified with the four bills fetched and the sweep on, checking collected cases rather than the result alone: sweep goes 3 failures to 1, and the one that remains is a genuine parser finding on 116-hr-133 (one section unreached), not a stale constant. Default slow suite 692 passed. The new guard was confirmed able to fail by injecting a bogus key. Closes #496 Refs #482, #478, #126, #474 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A bill written without TITLE divisions routes its sections through walk_body_sections, which had no appropriations-* branch. _extract_section_text therefore absorbed the whole account hierarchy into the section's own text and no account node was ever created: in 118-hr-9468 two agencies and their accounts, including Compensation and Pensions at $2,285,513,000 and Readjustment Benefits at $596,969,000, collapsed into one 382-character entry with a blank name and, because the section carries no <enum>, an empty match_path. The money still appeared in a node's text, so the amount-conservation gates passed; what was lost was which account it belonged to. Carve the appropriations children out of the section's own text and walk them, the same handling the title path already has. The walk itself is extracted to _walk_section_appro_children and shared, so the account naming rules (#474's split-account pending header) and the address rules now apply on both paths. The accounts address off their own agency names rather than off the section: _build_paths already omits an empty title from both paths, so the section's empty match_path anchors nothing and need not. Only the walk is shared, not the section's own node. The two callers build that node's display_path by different conventions, and unifying them would have re-cased 24,662 of the 25,191 plain body-level sections in the corpus to fix 7. Refs #474, #465, #147
…ure (#485) Commit 118-hr-9468 (introduced and enrolled) under tests/corpus/ and manifest it, so CI can see this shape at all: it is the corpus's only bill written without TITLE divisions, and neither bill exhibiting the defect was a fixture. TestUntitledBillAppropriations asserts the account NAMES and ADDRESSES rather than the amounts. Both amounts landed inside the collapsed node's text while the defect was present, so every amount-conservation gate passed on this bill; a money-watching test would have gone green on the broken build. All five cases were run against the pre-fix walker and fail there, each for its own reason. Two corpus gates needed a decision once the fixture enrolled them: test_every_section_reaches_a_node keyed on the section's own id. A section whose every child is an account now emits no node, so its content is represented by the accounts instead. Measured: 2 of the 1,467 sections with appropriations children across the committed fixtures take that branch, both in this bill, which is why the title path's identical behaviour had never presented it to this gate. Requiring a node here would mean emitting the empty, address-less placeholder #485 exists to remove, so the gate now accepts representation by an account, and test_appropriations_section_relaxation_still_fails_closed builds the case the corpus cannot supply to prove the relaxed branch still goes red on a real collapse. test_no_duplicate_match_paths: this bill has two enum-less body-level sections (the enacting lead-in and the short title) that both address to the empty tuple. That is a separate, pre-existing defect, not appropriations-related; the fix REDUCED it from 2 duplicates to 1. Recorded in _KNOWN_DUPLICATE_COUNTS, which is a ceiling rather than a pin, so a later fix tightens it without a test edit. Refs #474
The new fixture trips test_every_dollar_amount_appears_in_a_node's shell-bill threshold without being a shell: it is a complete enacted supplemental that appropriates to two accounts, so it carries exactly two amounts and falls under the gate's <3 cutoff. The undeclared-skip ceiling failed the CI corpus-gate step, which is that ceiling doing its job. Recorded with the distinction stated, because unlike every other entry in this list it is not a coverage gap: both amounts are asserted by name against the named account each belongs to in TestUntitledBillAppropriations, which is a stronger claim than the skipped gate makes. Caught by running the workflow's own module selection rather than the whole suite. The ceiling is scoped to that selection, so a full-suite run stays green on it -- the local check that matches CI is 'pytest -m slow <the corpus-gate module list>'.
…their own Review on #504: the stand-in accepted any surviving appropriations child as proof the section arrived, which is broader than the invariant it documents. A section can hold both prose and accounts, and there the parser does emit a node for the section. A regression dropping that node while keeping an account would have passed green with the section-level prose gone, and invisibly: the accounts and their dollar figures survive either way, so the money gates see nothing. Decide it with the parser's own carve and extractor rather than by tag, so the gate's condition IS walk_body_sections' emit condition and the two cannot drift. Adds the mixed prose+accounts negative case, and a synthetic division > section > appropriations-* case. #485 recorded the division shape as untested because no local bill has it, so the fix's reach into that caller rested on the two paths sharing one function rather than on a check.
…488) The gate read 3 of the 13 committed PDF/XML pairs. Two of the six fabricated anchors found while reviewing #473 sat on documents it never opened, so no tolerance could have caught them. Derive the pair list by walking tests/corpus/, and record per-document expectations in EXPECTED so a new pair runs immediately and test_expectations_cover_every_committed_pair fails until it is written down. Two documents needed a decision rather than a straight add: - 115-hr-5895/5_enrolled-bill.pdf: the anchor pipeline declines enrolled prints (no GPO margin line numbers, #141), which reads as 36 of 36 missed. Asserted positively as a declined layout instead of excluded, so a document that starts emitting anchors fails too. - 113-hr-3547/4_engrossed-amendment-senate.pdf: genuinely zero catchline-bearing subsections, which the MIN_CATCHLINES anti-vacuity floor could not tell from an empty extraction. The floor is replaced by an exact per-document pin of the oracle's count, which distinguishes the two by construction and also catches the denominator shrinking on the large fixtures. Refs #488, #473, #141 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q58Luk1QqHoGPaCzFGfb1L
#488) test_xml_subsection_nodes imports the fixture list from test_pdf_subsection_recall, so deriving it there widened this module from 3 documents to 13 and broke its bill-keyed tables. Decoupling would have preserved a 3-document scope that was never chosen here either: the gap was inherited through the shared import, not decided. Widening measures clean on this branch. Emission equals the oracle exactly and the quoted-block leak count is zero on all 13, so nothing needed a tolerance. Also replaces the two >= floors with exact per-document pins in EXPECTED, which now carries the subsection denominator alongside the catchline one. The floors had to sit under the smallest document (MIN_ALL_PAIRS 900 against 952 actual on 119-hr-1, >= 3 catchlines against 934), and the convergence floor could not admit 113-hr-3547, which genuinely has none. CLEAN is now derived from the recorded PDF residue rather than named: a document converges when the pipeline handles its layout and it carries no known false positive. 11 of 13 qualify. Before: 12 failures, 24 skipped (one an empty-parametrization skip). After: 2226 passed, 23 skipped, 15 xfailed. Refs #488, #188 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q58Luk1QqHoGPaCzFGfb1L
Rebasing onto develop grew the committed PDF/XML pair set from 13 to 24. The derived list picked all of them up on the rebase, and test_expectations_cover_every_committed_pair failed until each was measured and recorded, which is the mechanism this change exists to provide: a new pair is read on the run it lands, and cannot sit unmeasured. All 11 new documents measure clean (0 missed, 0 false positives, emission equal to the oracle, 0 quoted-block leaks). 117-hr-4432 is a better example than 113-hr-3547 of why the MIN_CATCHLINES floor had to go: it carries 74 subsections and not one bears a catchline, so "no subsection of this shape exists" is plainly not "the extractor returned nothing". The docstring now names it. Also drops the per-document table from the module docstring. EXPECTED is that table, and a prose copy would be a second list to keep in step. Refs #488 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q58Luk1QqHoGPaCzFGfb1L
0.4 and 0.6 decide whether a pair of bill sections is "the same text, edited or moved" or "two unrelated changes". They were written in five and four places, kept in step only by comments asserting they matched. A partial recalibration would not have failed: the XML and PDF pipelines would have disagreed about what counts as a move, each self-consistent, every test green. #368 and #170 are both open and both would move these values. New src/deltatrack/similarity.py owns SIMILARITY_THRESHOLD, MOVE_THRESHOLD and the three helpers that travel with them. diff_bill, diff_pdf, formatters/_text and the tests import from it. A module, not a constants file: it is cohesive around one concept, and formatters/_text importing the cutoff from diff_bill would point the rendering layer at the differ. Deliberately not a general constants.py, argued in #492: the other numeric constants have one consumer each and carry measurement-derived justifications that centralising would separate them from. The formatters/_text site was a bare default argument whose only caller never passes it, so nothing named the value that decides whether a reader sees an inline word-diff or two stacked paragraphs. Confirmed live before rewiring: a pair scoring 0.429 renders inline at 0.4 and stacked at 0.6. Names are public (text_similarity, not _text_similarity). A module whose purpose is cross-module import should not publish underscore names; diff_pdf was already importing three private ones from diff_bill. Two copies knowingly left: scripts/p2_catalog_survey.py and p3_prototypes.py are frozen replicas of what a past study measured (they reimplement the ratio function too), so wiring them live would change what a recorded result means. Noted in the module docstring. Refactor with no behaviour change; full suite green (2168 passed, 23 skipped, 15 xfailed). Closes #492 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q58Luk1QqHoGPaCzFGfb1L
) Account anchors came from one of two paths: glyph-size bands, or a fallback that walked back from a `For necessary expenses of` line to the nearest uppercase heading. The fallback made an appropriations-specific English phrase load-bearing for STRUCTURE, which #114 rules out: text triggers may interpret dollar amounts, never name accounts or build hierarchy. Measured over the 68-PDF corpus before removal, the fallback fired on 32 files and produced two account anchors in total, both on 119-hr-1 and both wrong (a wrapped section-heading fragment, `EXISTING ``FREE FILE'' PROGRAM AND ANY ``DIRECT`, labelled `account`). The other 30 were 11 enrolled/public-law files the product already declines as unnumbered layouts (#141) and 19 introduced/engrossed shells of 3-14 pages holding no accounts. So the tradeoff the issue flagged -- size-fail bills losing account breadcrumbs -- costs no correct structure. When bands are not derivable the structure now degrades to TITLE/SEC., the universal legislative tokens, instead of guessing an appropriations-specific one. Tests: the three account-block cases in test_diff_pdf (#56 card range, rename, duplicate-heading pairing) were reaching the retired path because their synthetic pages carried no glyph sizes -- a shape no published bill presents. They now supply sizes and exercise the production size detector, so they assert the same intent against the path real bills take. The legacy dedup case is dropped: its invariant is held corpus-wide by test_no_value_equal_duplicate_anchors, at a better altitude than a synthetic pair. The anchor goldens are unchanged, confirming every committed fixture PDF already took the size path. Refs: #114
… gate (#114) ADR 0018 locks the rule the #114 work implements: appropriations-genre English (`For necessary expenses of`, `RESCISSION`, `INCLUDING TRANSFER OF FUNDS`) may interpret dollar amounts, never name accounts or build hierarchy. Structure comes from format — glyph size, position, and the universal `TITLE N` / `SEC. N` / enumerator tokens, which are explicitly carved out. When the format signal is absent the parser degrades to TITLE/SEC. rather than substituting a guess: a shallower breadcrumb that is true beats a deeper one that is invented, because a consumer can see the first is shallow and cannot see the second is wrong. The committee-report parser keeps its qualifier vocabulary — it parses the comparative-statement table that is our independent ground truth (ADR 0009), a different document and not bill structure. The gate scans the structural modules' string literals (via ast, so prose about the rule in comments and docstrings does not trip it) for that vocabulary. It is an absence assertion, so it is built to prove it can fire: one test feeds it the exact retired `_FOR_NECESSARY_EXPENSES` line and fails if that is NOT flagged. Without that, a detector that had quietly stopped matching would be indistinguishable from permanent compliance. Also corrects two in-code ADR references that pointed at 0012 (heading levels) before the number was checked; 0018 is the free slot. Refs: #114
…igger notes The ADR index in AGENTS.md is generated and gated by test_adr_index.py, which caught the missing entry. Also drops the now-unused SizeBands import and rewrites the test_pdf_size_detection module docstring, which still told readers to keep body text clear of "For necessary expenses" so cases could not fall through to a fallback that no longer exists. Refs: #114
… the real bill Review of PR #512 found three gaps. All three were reproduced before being fixed. 1. The ADR described a degrade the parser does not have. It said that without the format signal "the parser emits TITLE/SEC. and stops", but `_anchors_from_page` runs unconditionally and deliberately keeps emitting enumerator-derived run-in subsections, which #114 carves out as universal grammar. Verified: on a size-less page the parser still emits `(a) In general` and `(b) Deadline`; on 119-hr-1 it emits 11 titles, 355 sections and 936 subsections. The code was right and the new contract was wrong, so the contract is corrected here -- in the ADR, the extract_anchors/breadcrumb docstrings, and the fallback test comments. What degrades is the appropriations-specific interior levels (account/agency/major/ grouping), not everything below the title. 2. The enforcement gate failed open on ordinary reintroductions. It matched contiguous substrings over a fixed three-file list, so `r"^For\s+necessary\s+ expenses\s+of"` passed, and so did moving the trigger into a new helper module. Both are now covered: literals are normalised to letters before matching (regex escapes dropped first, so the `s` in `\s` cannot pose as a letter of the phrase), and the scanned surface is everything under src/deltatrack minus an allowlist naming the modules permitted this vocabulary. A new module is therefore guarded by default. Added tests prove the gate fires on escaped/spaced/bracketed variants and on a trigger planted in a freshly-created module, and that the allowlist names only modules that exist. 3. The motivating failure had no real-PDF regression. The goldens pin three fixtures, all of which take the size path, so none could witness the 119-hr-1 behaviour the removal was justified by. `119-hr-1/1_reported-in-house.pdf` IS committed and does take the degraded path (bands None at coverage 1.0), so it now carries a direct gate: no account/agency/major/grouping anchors, universal structure intact, and the specific false fragment recovered instead as the whole subsection catchline it belongs to. The precondition (bands is None) is asserted too, so the test cannot silently stop guarding if #508 moves this bill onto the size path. Also corrects the ADR's own arithmetic: the fallback produced two anchors, one on each of two versions of the same bill, which is ONE distinct piece of structure -- not "two wrong anchors" in context and "one incorrect anchor" in consequences. And records that the census used the parser's real condition (bands AND coverage >= 0.85), since low-coverage documents with valid bands also used the fallback. Refs: #114
The cross-format gates only run where both formats of a version are committed, so the 32 versions carrying XML alone were invisible to them. This commits 27 PDFs and takes per-version format parity from 24 of 57 versions to 51. Every added PDF was measured against the dollar-amount cross-check before being committed: all 27 recall their XML amounts from PDF text with nothing missing, so none needed a new baseline. Four bills that had NO amount cross-check at all -- 117-hr-2471, 116-hr-1865, 115-hr-1625, 115-hr-244 -- now have one; their enrolled print is the only PDF each has. The twelve Senate reported prints were the best value per byte and had been deferred only because nobody had fetched them: 4.0 MB for 2,524 amounts and 3,237 anchors. 118-s-4795's PDF is byte-identical to the committed CJS fixture, so git stores one blob under both paths rather than a second copy. Declarations this needed, each recording a property of a document rather than excusing a defect: - The nine enrolled prints contribute no structure -- enrolled bills carry no GPO margin line numbers, so extract_anchors returns nothing (#141). They are declared in the two zero-anchor registries, which still ASSERT on them (documented layout, unnumbered classification, intact text layer) rather than skipping. - The zero-anchor class's "text layer intact" check moves from absolute floors (>10,000 chars, >=10 section enumerators) to its own XML as the oracle. The floors encoded "omnibus-sized" and could not admit a genuinely short enrolled print: 118-hr-9468 is 3 pages, and its PDF text EXCEEDS its XML (5,222 vs 4,594 chars). The relative form is stricter where it matters -- a multi-thousand-page omnibus extracting 10,001 chars passed the old floor and fails this one. - Two meta-tests keyed on real single-format stages as their negative examples. Parity removed those examples, so they now run against a stub manifest and test the resolution logic rather than the corpus's shape. Both modules already warned in their own docstrings that this would happen; the fault-injection check (collapsing format resolution) reddens both. Five engrossed-amendment PDFs are WITHHELD, and that is the more useful half of the result. They exposed three defects in the subsection detector -- 9 catchline-bearing subsections reaching no anchor, 8 quoted-block subsections leaking as anchors, and an XML-side divergence that looks like the #11 amendment-doc class. The subsection gates assert recall and precision absolutely by design, so committing these would mean weakening a gate rather than recording a fact. Filed as #519 with the exact misses; each is a ready-made regression fixture the moment it lands, the same posture the manifest already records for 114-hr-2029 v4 (#434) and 115-hr-244 v5 (#330). Closes #126 Refs #488, #519, #434, #330 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The parity comments read as "every manifested version now carries both formats", which the same change makes false: five engrossed-amendment PDFs are deliberately withheld for #519 and 114-hr-2029 v4 stays PDF-only while the multi-<legis-body> section drop is open. A comment that overstates coverage is the kind a later session reasons from, so the six exceptions are now named wherever the claim appears. - TESTING.md: 51 of 57 versions dual-format; six deliberate exceptions. - corpus_manifest.toml: the amount cross-check covers every DUAL-FORMAT version — the five XML-only ones have no PDF to check against. - test_corpus_manifest.py: drop "no xml-only counterpart since #126" and "since #126 every manifested version carries a pdf". The stub manifest's real justification is that the six remaining single-format versions are each single-format only until their withholding issue closes, so keying on one would reset the same trap with a later fuse. For the pair test, no real pair can reach the negative branch at all: the ids come from adjacent_pdf_pairs(), which enumerates PDFs on disk. - conftest.py: 113-hr-3547 v1 gained its XML in this PR, so the pdf-only example is now 114-hr-2029 v4. - Two zero-anchor registries: parity phrasing corrected in place. Test behavior, the subsection gates, and the #519 withholding strategy are untouched: the diff contains no executable lines. Refs #126, #519 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Updated tests/corpus_manifest.toml: 114-hr-2029/4_reported-in-senate now has both PDF and XML formats (was XML withheld while #434 was open) - Updated tests/test_corpus_manifest.py: Fixed test expectations for format parity (114-hr-2029/4_reported-in-senate now has both formats) - Updated tests/test_corpus_properties.py: Added duplicate match_path baseline for new XML fixture - Updated tests/test_corpus_tree_properties.py: Added find_bill_body import (was changed in PR branch) - Updated tests/test_pdf_subsection_recall.py: Added EXPECTED entry for new PDF/XML pair
1. Restored test_corpus_properties.py from current develop: - Restored test_no_top_level_legis_body_is_silently_discarded() (#434 regression test) - Restored exact equality baselines for duplicate match_paths (#513) - Restored _assert_duplicate_baseline() and bidirectional test - Kept #520-specific addition: 114-hr-2029/4_reported-in-senate.xml baseline (119) 2. Updated format-parity narrative to 52/57 dual-format (was 51/57): - 114-hr-2029/4 now has both XML and PDF (since #434 landed) - No PDF-only versions remain; 5 XML-only are the #519 withheld engrossed amendments - Updated: corpus_manifest.toml, TESTING.md, conftest.py, test_corpus_manifest.py - Updated: test_corpus_tree_properties.py, test_pdf_division_recall.py comments 3. Lint clean: ruff check and ruff format --check pass 4. All tests pass: - Fast suite: 1669 passed - Slow suite: 1103 passed - Relevant modules: 554 passed
…-refresh The rebase conflict in corpus_manifest.toml was resolved by taking the develop version (with #126 format parity changes) and re-running scripts/update_manifest_with_reports.py --refresh to regenerate all committee_report pairings on top.
#126 format parity Resolved corpus_manifest.toml conflict by taking develop base and re-running scripts/update_manifest_with_reports.py --refresh to regenerate all committee_report pairings. test_ci_workflow.py kept at HEAD version.
Remove tomlkit from the engine's install dependencies
…arget Path(bills_dir) / target silently discards bills_dir whenever target is absolute, so an explicit --bills-dir combined with an absolute slug/target used to resolve against the wrong corpus with nothing in the output to say so. Reject the combination at the CLI boundary instead. The bare absolute-directory listing #426 added is unaffected, since it never names --bills-dir. Closes #454 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
COMPARE_RATE_LIMIT_PER_MINUTE arrived with PR #390 but never reached the operator-facing cap list, so the one limit an operator can trip without touching the app was the one not written down. It now sits alongside the upload-size and concurrency caps. The keying rule carries a constraint worth recording before it bites: _rate_limit_key reads the last entry of the last X-Forwarded-For header, which is the real client only while Apache is the outermost proxy. Put a CDN in front and that entry becomes the CDN edge address, collapsing every user behind it into one shared bucket. The key has to change as part of adding the CDN, not after the 429s start. The replacement key carries its own condition, so it is stated in the same paragraph rather than left to a private runbook: receiving a CDN client-IP header is not on its own a reason to trust it. The origin has to accept it solely from the CDN or overwrite untrusted copies, because a key a direct caller can set for itself removes per-IP limiting entirely rather than over-applying it. Also records that the limit is a slowapi default limit on ASGI middleware rather than a per-route decorator, so a future API route inherits the budget unless it opts out, and adds rate limiting to the request-flow map, which listed the app's other guards but not this one. Closes #394 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, and gate sensitivity Three probes for the #581 prerequisite question, covering the gaps the #586 probes leave open. All read-only; none touches production code. probe_node_identity.py answers where an ADR 0019 ordinal can still be recovered. match_nodes returns the same objects its source BillTree.nodes holds -- no copy, no reconstruction, and the recovery is a bijection over all 27 corpus pairs. That result is worthless without controls, so it carries three: a value-equal copy substituted for the real node (identity reddens on 27/27, value equality answers silently on 27/27 with an address it did not earn), value collisions within one tree (none today, which is why the copy control and not the collision count is the argument), and output position used as the ordinal (wrong on 38,104 of 46,942 addresses, 81.2%). probe_round2_migration.py measures the overlap probe_splits.py states it does not. 46.0% of the 496 selected moves involve a removal or addition created by classification's similarity split, and 145 have both sides so created -- so moving the second retrieval round ahead of classification is not a reordering while that split remains a classification-time decision. It also runs the legacy (ri, ai) sort key and the ADR 0019 ordinal key through one greedy loop: they select different correspondence on 3 of the 16 selecting pairs, 20 links apart. The architectural address and the legacy ordering key are two different values. probe_canonical_sensitivity.py proves the byte-identity gate can fire on a correspondence change rather than assuming it. Four passes, so a drifted duplicate cannot be mistaken for the injected fault: production, the duplicated loop under the production key, the duplicated loop under the ordinal key, production restored. The three reddened pairs carry identical change counts, identical byte counts and identical summaries -- only the digest moves, which is the case a count-based check would pass straight through. Refs #581
#591 moved the similarity revocation out of classification and into apply_similarity_revocation, which runs ahead of the classification loop. Every figure these probes report is unchanged by that (verified: identical output at 0f07dc4 and at the pre-#591 develop 97f91ba), but two of them now describe the architecture inaccurately, and one of those inaccuracies inverts a conclusion. probe_round2_migration.py Terminology. "Created by classification" was true only of the pre-#591 loop. The removal and the addition are now produced by the revocation stage, before classification runs. Renamed throughout to "revoked" / "revocation-produced" / "left unmatched by the similarity rule". The conclusion this changes, which matters more than the wording. The figure used to be evidence that a second retrieval round CANNOT move ahead of classification: 46.0% of selected moves involved an entry that did not exist until classification had run. Those two observations are now separated before classification, so a round-2 pass placed after the revocation stage does see them. The same 228/496 now sizes an ordering constraint INSIDE assignment -- round 2 must run after the revocation -- rather than a dependency on classification output. What still blocks a literal relocation is representational: reconcile_moves consumes NodeDiff, which is classification output. Every field it reads is derivable from the BillNode pair, so that is a mechanical change owing its own evidence, not an obstacle. Recorded in the docstring so the next slice starts from the current constraint. A fail-open closed. split_element_ids transcribed the revocation condition as `old_norm != new_norm and text_similarity(...) < SIMILARITY_THRESHOLD`, while production gates on the emptiness of a word-level diff. The two agree on all 15,034 path-matched corpus pairings (measured, 327 == 327, zero disagreements), so no reported number was wrong -- but nothing asserted that agreement, and a divergence would have silently misclassified the population. The probe now asks production's own pairing_survives_similarity_rule, and cross-checks that verdict against what apply_similarity_revocation structurally did: one revocation must add exactly one tuple, or the run stops. Proven able to fire by replacing the stage with a passthrough (fires on the first pair carrying a revocation, before any figure prints). probe_node_identity.py The pre-classification sequence is two stages now, not one, and the second rebuilds the tuple list -- which is exactly where a copy would be introduced unnoticed. Measuring only match_nodes would leave a future ObservationRef wired at a seam one stage short of classification. Identity survives both stages: 0 foreign objects, 0 double-claimed, 0 omitted, and the revocation stage adds exactly 327 tuples. Control C now runs at both stages, and the two figures are ASSERTED equal rather than reported side by side. They are equal by construction -- the revocation stage replaces a pairing in place, so per-side observation order cannot move -- so printing them as two independent agreeing measurements would overstate the evidence. As an assertion it is a tripwire for the one mutation #591's docstring names as dangerous: appending the replacements elsewhere moves canonical output while leaving every change count untouched. Proven able to fire (81.2% -> 89.7% wrong addresses under an appends-at-the-end revocation stage). probe_canonical_sensitivity.py is unchanged. Its mutation point is reconcile_moves, which #591 did not touch, and its duplicated loop is still verbatim against current source. MEASUREMENTS, all reproduced on current develop: identity 27/27 pairs, 0 copies, bijection intact, both stages value equality answers with an unearned address on 27/27 (control A) position as ordinal 38,104 wrong of 46,942 (81.2%), equal at both stages revocation population 327 pairings; 228 of 496 selected moves (46.0%) touch one ordinal vs (ri, ai) 3 of 16 selecting pairs change the selected set, 20 links canonical sensitivity the same 3 pairs redden; counts, byte counts and summaries identical, only the digest moves Every fail-closed guard in all three probes was fault-injected and confirmed to fire: the element_id bridge (repeated and empty), the duplicate-loop agreement check, the predicate/stage cross-check, the canonical VACUOUS guard, the canonical equivalence guard, and the new in-place tripwire. Refs #581
The previous commit's docstring states that what now blocks relocating round 2 is representational rather than populational: the removals and additions reconcile_moves consumes are already present as unmatched observations before classification runs, and only their NodeDiff representation is classification output. That was reasoning, not evidence, and it is the load-bearing sentence for the next slice -- if part of the population genuinely appeared only after classification, the relocation would be a behaviour change rather than a rewrite. Section E measures it. Every (old, None) out of the revocation stage must correspond to a "removed" record reaching reconcile_moves, and every (None, new) to an "added" one. On the committed corpus: 1207 == 1207 and 16321 == 16321, no pair mismatched. Compared PER PAIR rather than in aggregate, because two pairs whose errors cancelled would agree on the totals and disagree on every document. Proven able to fire against exactly that: a doctored pair with one unmatched observation too many and another with one too few, whose totals match exactly, is still refused. Refs #581
The list said 'three places' over four bullets, and section E made it five. A hardcoded count of the thing directly beneath it is drift by construction, so the count is gone rather than corrected, and section E is named in the summary that lists what the probe measures. Refs #581
The docstring credited the pre-#591 measurement to develop 97f91ba, which is where #591 loaded the pre-refactor module for its own comparison, not where these probes ran. They ran on this branch at c3b6387, based on af83c28. The two bases differ by documentation only, so the numbers are unaffected -- but a provenance line that names the wrong commit is the kind of claim a later session would build on without re-deriving. Refs #581
…y groups test_the_engine_install_does_not_drag_in_the_web_channel only checked three hardcoded package names (fastapi, uvicorn, httpx), all from the web group, despite its own failure message describing a broader contract: no delivery-channel OR tooling dependency should leak into [project.dependencies]. #533 found a dev-group package (tomlkit) declared there too, and the hardcoded probe had no way to catch it. The generalized version parses the web/fetch/dev groups straight out of pyproject.toml and resolves each declared package to its actual import name(s) via importlib.metadata.packages_distributions() on the running interpreter (dev/web/fetch are default groups, so whatever they declare is already installed here) -- so a package added to any group is covered automatically, with no second list to keep in sync. Verified the new probe actually catches what the old one couldn't: re-adding pyyaml (a dev-group package) to [project.dependencies] fails it; reverting passes it clean. Closes #593
The runbook says main has no merge queue and that pip-audit (production deps) is a required check on main. Both are GitHub branch-protection settings, not derivable from repository files, so no repository consistency test can go red when they change. Say so explicitly next to the claims rather than letting a reader mistake the green suite for coverage of them.
docs/release.md step 4 tells a maintainer to watch the post-merge security run on main as one of the three runs that test the promotion merge commit. ci.yml's push branches were already pinned for both branches; security.yml's push trigger was not, so removing main from it would silently retire the run the runbook watches.
) docs/release.md drifted into seven factual errors before it merged (#544). Its load-bearing claims about machinery were all checkable from committed files yet none was. Four contracts are now pinned, each with a demonstrated negative control: - The Pages job publishes committed files and must not render. Executable run steps are parsed from the workflow YAML (comments cannot count) and watched for the fetch/diff commands #42 removed and for the current renderer. - The publishing actions the runbook names by role still appear as uses: entries. The version pin is deliberately not part of the contract. - Every test the runbook cites by name resolves to a real test, read from the AST of the referenced module (bare citations resolve against whichever module defines them). - The security workflow still fires on pushes to main (test_ci_workflow.py).
docs(web-compare): add the per-IP rate-limit cap and its CDN caveat
The upper bound of the member-count calibration was 200,000 -- but that is exactly the member count issue #306 used to demonstrate inode exhaustion. extract_archive() only refuses on member_count > MAX_MEMBER_COUNT, so a ceiling raised to exactly 200,000 would still pass the calibration test while re-enabling that exact attack: the archive would be accepted, not refused. Pin the ceiling to the current production value (100,000) instead, which is the value the code's own comment justifies as a deliberate 2x safety margin under the #306 threshold ("refusing the 200,000-member archive... by a factor of two"). A bound anywhere below 200,000 technically still refuses that one archive, but only after letting the margin erode to nothing; pinning to the documented value keeps the margin the comment claims, not just technical non-equality. Also split each ceiling's combined test into two: one asserting the numeric calibration band, one proving the real (unpatched) ceiling still extracts an ordinary small archive. The previous single test per ceiling read as if it extracted something derived from the real corpus figures (162 MiB / 10,564 members); it never did, and did not need to -- the two pieces of evidence are now named for what they actually check. Verified against negative controls: MAX_MEMBER_COUNT = 200_000 now fails the calibration test (previously the whole point of this fix), as do 50_000_000 (#447's original fault injection) and MAX_UNCOMPRESSED_BYTES = 3_900_000_000. Restoring the real constants returns the full archive test suite to green.
Every figure these probes report was already correct. What was missing is that a future run reporting a WRONG figure would still have exited zero. Five gaps, each closed and each proven closed by injecting the fault it claims to catch. probe_node_identity.py -- validation, not reporting The identity and bijection results were accumulated into counters and printed. A regression would have printed a non-zero count and exited zero, which is a finding only if someone reads it. validate_identity() now raises IdentityFailure unless foreign objects, double claims and omissions are all zero, and BOTH production stages go through it on every corpus pair. Injecting a revocation stage that emits dataclass copies aborts the run on the first pair. Control A tested the wrong thing. It built replace(original) and asserted its id was absent from the index -- a restatement of the expression under test, not a test of the validator. The copy is now substituted into an otherwise-real tuple sequence and pushed through the SAME validator, which must reject it (27/27), while a value-keyed map still accepts it with an address it did not earn (27/27). Control B stays observational: zero collisions today is not the argument. Scope stated: object identity is a run-local mechanism for recovering the complete parser-sequence ordinal. It is not ADR 0019 Observation identity, which is (source_sha256, parser_revision, node_ordinal). probe_round2_migration.py -- structure, order, and a pinned baseline The revocation cross-check compared CARDINALITY: extra output tuples against predicate revocations. Revoking the same number of different pairings would have passed it while every population figure described the wrong observations. revoked_pairings() now walks match_nodes output to revocation output by object identity, requiring each input tuple to appear either unchanged or as its own two halves adjacent and in place, and refusing substitution, reordering, wrong partner, omission, duplication, trailing output and unexpected shape. That population is then cross-checked against the predicate as sets of object identities. Injected a stage revoking an equal-sized DISJOINT set: 44 revoked by the stage, 44 by the predicate, zero in common -- refused. Section E compared counts per pair. Equal counts do not establish that the same observations become the removed/added lists, still less that their ORDER is preserved -- and legacy (ri, ai) is a position in exactly those lists, so order is the part worth having. It now compares ordered element-id sequences and names the index of the first divergence. A same-count, same-SET reorder is refused. The headline figures are pinned. They are historical behaviour, not ADR policy: nothing here is a target. But a Phase-1 slice claims to preserve matching behaviour, so a figure that moves is a regression or an intentional change owing evidence, and both deserve a failing run rather than a new number printed where a reader has to remember the old one. probe_canonical_sensitivity.py -- the prediction is now enforced The docstring predicted exactly three reddened pairs; the code required only that pass 3 be non-empty. A fault that reddened all 27 would have read as success, which is what a broken harness produces. Pass 3 must now equal the three keys exactly, imported from probe_round2_migration so the two probes cannot drift into disagreeing about which pairs matter. Injected a wrong expectation naming an unrelated pair, and a strict subset: both refused, each naming what was missing and what was unexpected. Its element_id -> ordinal bridge built a dict inline, which would silently collapse two observations onto one address. It now goes through the same fail-closed ordinal_bridge, refused on both a duplicated and an empty id. MEASUREMENTS, unchanged on the rebased head -- which is the point, since the corrections were to the guards and not to the arithmetic: identity 27/27 both stages, validated rather than counted copied-node control rejected 27/27 by the validator; value equality accepted 27/27 position-as-ordinal 38,104 wrong of 46,942 (81.2%), equal at both stages revocations 327, structurally walked and identity-cross-checked overlap 228 of 496 selected moves (46.0%); 145 both sides ordinal vs (ri, ai) 3 of 16 selecting pairs; 20 links; 0 order-only section E 1207 and 16321 observations, ordered sequences identical canonical exactly the predicted 3 pairs redden, no more, no fewer Refs #581
Review on #597 found two problems with the generalized gate: - `assert len(forbidden) > 3` was a false-green risk: `web` alone already clears that floor, so a code regression back to checking only `web` (the exact scope of the hardcoded probe this replaced) would still pass. Fixed by deriving the set of groups to check from pyproject.toml's `[dependency-groups]` table directly (no hardcoded group-name tuple), and asserting each group individually resolves at least one forbidden import name -- so a group that silently lost its coverage (emptied, or dropped from the loop) fails loudly and names the group, instead of hiding behind a large sibling group's count. - The failure message claimed "[project.dependencies] has re-acquired" a dependency, which is stronger than what importability alone establishes: the package could also have arrived as a legitimate transitive dependency of a real runtime dependency. Reworded to state both possibilities and fail closed, forcing a conscious packaging-boundary decision either way. Verified: reproduced the PyYAML fault injection against the corrected assertion (still fails, now names the owning group in the message); added a negative control emptying the `fetch` group's package list, which fails with a message naming `fetch` specifically -- proving per-group coverage is enforced, not just a large union.
The two *_at_default_extracts_a_normal_archive tests added alongside the calibration assertions duplicated coverage that already existed: test_creates_the_destination_and_writes_members already extracts a small archive under both real, unpatched ceilings, and test_an_archive_at_the_ceiling_is_still_extracted already proves the byte ceiling isn't wired to reject everything. Issue #447 is about the untested calibration of the production constants specifically, not about extraction succeeding under them, so the calibration assertions are the only new coverage this needs.
evidence(#581): probes for ordinal recoverability, round-2 migration cost, and canonical gate sensitivity
…solute-path fix: hard error when --bills-dir conflicts with an absolute compare target
…bration test: two-sided calibration for archive extraction ceilings (#447)
…trol through one validator The negative control re-parsed the push trigger and asserted the inverse of the live guard's membership check, so it proved only that parsing could detect a mainless fixture, not that the same verdict the live guard depends on would reject it. Both tests now call _security_push_failures, so a rewrite that accepts a mainless trigger (or stops reading push) flips both red.
Review on #597 found a remaining false-green: the gate ran `import {name}` and treated ModuleNotFoundError as evidence of absence. But a forbidden package can be genuinely installed and still raise ModuleNotFoundError on import, if importing it fails partway through on one of ITS OWN missing dependencies (something the engine-only install never pulled in). Demonstrated concretely: a package installed into a throwaway venv whose __init__.py does `import this_module_absolutely_does_not_exist_xyz` gives `import broken_thing` a ModuleNotFoundError (rc=1) even though the distribution is on disk -- the old assertions would have called it absent. Replaced import-based detection with `importlib.metadata.distribution(name)` run inside the engine-only venv, which asks what ADR 0016 actually protects (is the distribution installed) independent of whether importing it succeeds. importlib.metadata's own lookup already normalizes PEP 503 name variants (case, "-"/"_"/"."), so the distribution-to-import-name mapping this used to need (packages_distributions(), _canonicalize()) is gone entirely -- `_group_distribution_names()` is now a pure parser over pyproject.toml with no dependency on the current interpreter's installed packages. Verified: reproduced the PyYAML fault injection against the new metadata check (fails, names 'pyyaml' and the owning 'dev' group); reran the fetch=[] structural negative control (fails closed, names 'fetch'); both reverted cleanly. Full local suite, ruff check, and ruff format all pass.
…ed verdict functions The render and publishing-action-role negative controls recomputed their verdict inline, so they proved parsing could detect a bad fixture but not that the same verdict the live guard depends on would reject it. Both contracts now share one validator each (_rendering_offenders, _missing_publishing_roles): the live guard asserts it is empty for update-examples.yml and the negative control asserts it fires on a deliberately bad workflow, so a rewrite that stops matching goes red in both places.
…ll-gate fix: generalize the engine-install gate to cover any tooling/web dependency
…tency Pin docs/release.md against the workflows and tests it describes (#547)
This was referenced Oct 4, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This pulls in a newer bill drop down navigator, and section controls. Behind the scenes there are new docs on standard terms, language, and style guide standards the bills and our system are starting to use.