Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
31 changes: 0 additions & 31 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -104,37 +104,6 @@ docs/research/provision-matching/probes/form_*.html
docs/research/provision-matching/probes/labels/
docs/research/provision-matching/probes/merged_labels.json

# The PDF bake-off's P2 holdout corpus: 88 govinfo documents, 16.4 MB, downloaded by
# probes/fetch_holdout.py. Same reasoning as /bills and bills_corpus above — bill source
# material is fetched, not vendored. What makes this safe here specifically is that
# results/holdout_membership.json (committed, and covered by the spike's frozen
# PRESERVED-MANIFEST.txt) records the govinfo package id, sha256 and byte count of every
# one of the 88 files, and the fetcher verifies each download against it. A re-issued or
# withdrawn package therefore fails loudly instead of being scored as the historical input.
# No trailing slash, per the #319 symlink reasoning above.
docs/research/pdf-backend-bakeoff/holdout

# External-validity probe evidence (A46). Running any xNN probe rewrites its own evidence
# file, so after the A46 cleanup these reappear as untracked artifacts and a later `git add -A`
# would silently re-commit the working material that was retired. Ignored rather than left
# loose. They are regenerated by running the probe; git history holds the removed versions.
#
# The negations are the artifacts something actually READS or that a byte-frozen document
# cites — dropping one un-tracks a live gate input, so only change this list together with
# the consumer that justifies it. A new committed artifact needs a new negation here.
docs/research/pdf-backend-bakeoff/validation/external-validity/results/x*.json
# G2 reads this one; G6 defects ORACLE_INTEGRATION_NOT_VERIFIED without the other.
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x2_contract_assertions.json
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x26_control_oracle.json
# Cited as MEASURED by PRE-REGISTRATION.md, which is byte-frozen and can never be repointed.
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x00_design_pilot.json
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x02_oracle_reference_defects.json
# Named by retained METHODOLOGY_SURFACE files (cross_engine_control, run_hybrid, run_extended,
# reconstruct_extended_corrected), which may not be edited for a cosmetic reason.
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x09_skeleton_cross_engine.json
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x11_provenance_chain.json
!docs/research/pdf-backend-bakeoff/validation/external-validity/results/x13_x_arm.json

# Build artifacts. The project became buildable in #398, so `uv build` now produces
# these where it never used to. `build/` matters beyond tidiness: some backends copy
# the package's own .py files into it, and a directory holding Python that is neither
Expand Down
8 changes: 8 additions & 0 deletions docs/decisions/0002-pdfium-single-engine.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,14 @@ shipped in PRs #38 and #40.
point of extraction quality for this tool.
- One engine instead of two means one set of text quirks to understand and one
cleaning path to maintain, at the cost of that path being PDFium-specific.
- PDFium is the engine for the anchoring and chrome behaviour above, not for a
measured accuracy lead, and no backend has one. Six backends read through one
per-glyph contract, on published GPO bills of effectively one typesetting class,
showed no winner: pdfminer.six agreed best with the XML on text and headings, while
PDFium compiled to WebAssembly reproduced current output exactly and ran about 8×
faster in the browser. No design for plugging a second backend into the extractor
was selected; a hybrid of engines and an extended per-glyph contract were both
prototyped and neither was validated ([evidence](https://github.com/civictechdc/DeltaTrack/blob/4171e32d93e869725a86ae72eab70fa355a9919f/docs/research/pdf-backend-bakeoff/CLOSEOUT.md)).
- The engine-vs-engine parity check could not survive pdfplumber's removal, so the
regression guard is now a golden snapshot: five curated pages, each exercising
one cleaner path (soft-hyphen reconstruction, VerDate-glue, watermark-glue,
Expand Down
10 changes: 10 additions & 0 deletions docs/decisions/0003-pdfjs-client-side-viability.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,16 @@ server-side engine choice in ADR 2.
- Two engines across two channels (PDFium server-side, PDF.js client-side) means
two extraction paths that must be kept in agreement; divergence on edge cases is
a maintenance cost if both channels ship.
- Porting to TypeScript is not the only browser route. The Python engine itself runs
in a page under Pyodide: the XML comparison produces byte-identical canonical JSON
and HTML there, checked by `scripts/pyodide_parity.py`. Its one obstacle is the
module-scope `pypdfium2` import in `parsers/pdf_text.py`, which the XML comparison
reaches through module-level imports and which the check stubs out until
[#751](https://github.com/civictechdc/DeltaTrack/issues/751) makes it lazy. For PDFs
on that route, PDFium compiled to WebAssembly (`@embedpdf/pdfium`) exposes the
per-glyph data the extractor uses and reproduced current output on the published
bills tested, which would keep one PDF engine across channels rather than two. The
PDF path has not yet run end to end in a browser ([evidence](https://github.com/civictechdc/DeltaTrack/blob/4171e32d93e869725a86ae72eab70fa355a9919f/docs/research/pdf-backend-bakeoff/CLOSEOUT.md)).
- **Open risk:** the spike covered only published GPO bills, which have clean text
layers. Draft and pre-introduction PDFs (watermarked, possibly image-only) were
not tested and are the documents where extraction is hardest and most
Expand Down
11 changes: 11 additions & 0 deletions docs/decisions/0011-local-only-processing.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,3 +99,14 @@ Alternatives:
- Telemetry, crash reporting, or "send us the file that failed" diagnostics that would
carry bill content off-device are foreclosed by this rule. Diagnostics must be local
or content-free.
- **A browser channel cannot guarantee zero egress with Content Security Policy
alone.** Tested against script deliberately trying to exfiltrate, a strict policy
(`connect-src 'none'`, and no `'unsafe-inline'` in `script-src`, which Speculation
Rules prefetching otherwise gets through) blocked every subresource mechanism tried.
Two mechanisms are outside CSP entirely: `window.open` opens a new browsing context
that can carry content in its URL and leaves the page in place, and WebRTC reaches a
STUN server under any page-level policy, a covert signal rather than a content
channel. Closing them needs a browser- or device-level control such as enterprise
policy. A browser channel's no-egress claim must name that dependency and be
verified at the network layer against a control shown to observe egress
([evidence](https://github.com/civictechdc/DeltaTrack/blob/4171e32d93e869725a86ae72eab70fa355a9919f/docs/research/pdf-backend-bakeoff/CLOSEOUT.md)).
79 changes: 0 additions & 79 deletions docs/research/pdf-backend-bakeoff/CLOSEOUT.md

This file was deleted.

Loading
Loading