You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Two different scanned PDFs compare as "no changes" instead of being declined #550
If you compare two scanned PDFs, DeltaTrack tells you nothing changed, even when the two documents share no text at all. It can't read either file, but it doesn't say so.
Some background, for anyone new to this:
The two pipelines. DeltaTrack compares bill versions either from their XML (structured text Congress publishes) or from their PDF (the printed page). The PDF pipeline exists for drafts that circulate only as PDF (ADR 0010). It reads the PDF's text layer, the selectable text stored behind the page image.
Scanned (image-only) PDFs. A scanned or photocopied page is just a picture. It has no text layer, so there is nothing for DeltaTrack to read.
Declining. When DeltaTrack knows it can't give a trustworthy answer, it declines: it refuses, with a message written for the user, instead of guessing. In code that is UnsupportedLayoutError; the web app returns it as an HTTP 422 carrying the message. Today the only decline is for PDFs printed without margin line numbers, such as enrolled bills, added by PDF anchor extraction yields 0 anchors on enrolled bills (no margin line numbers) #141 (closed).
How it shows up
Checked on develop (d3935ee), 2026-10-05. Pages 3–22 and 23–42 of tests/corpus/118-hr-4366/1_reported-in-house.pdf were rasterized into two 20-page image-only PDFs:
The command line agrees.diff_pdf.py writes a normal report that says "No changes found between these versions."
The control works. The same two page ranges with their text intact extract 370 lines and compare as 89 changes.
Page count decides the outcome. Scans of 1 to 49 pages are answered "no changes". Scans of 50 pages or more are declined, but with the enrolled-bill message ("no printed line numbers … Use the XML version"). That is wrong for a scan, since a scanned draft has no XML version.
One scanned side isn't caught either. The committed one-page image-only fixture tests/data/p3/imageonly.pdf, compared against a text PDF, is answered as 1 removed and 1 added.
Why it matters
The web app's compare page has PDF selected by default (web/webapp/compare.html:39). A first-time visitor who uploads two scanned drafts gets "no changes", which looks the same as two versions that really are identical. That is a confident wrong answer, the outcome the #141 decline exists to prevent.
Cause
The decline check only judges documents with at least 50 extracted lines (_MIN_LINES_FOR_GUARD, src/deltatrack/compare/pdf.py:94, applied in _is_unnumbered_layout at line 106). That floor is deliberate: short real prints have too few margin numbers to judge. But an image-only page still yields one empty line, so a 20-page scan looks like a 20-line document and is exempt. Nothing checks whether any text was extracted at all.
Done when
Comparing two image-only PDFs of any length, including 1 page and 20 pages, is declined.
Comparing an image-only PDF with a text PDF is declined.
The decline message describes a scan, not missing line numbers. Proposed wording: "This PDF appears to contain images rather than extractable text and cannot currently be compared."
The web app returns that message as a 422.
No committed corpus PDF is newly declined. tests/test_pdf_canonical_baseline.py stays unchanged (23 pairs, 6 declined), and the sparse real print tests/corpus/113-hr-3547/2_engrossed-in-house.pdf (4 pages, 22 non-empty lines, 697 characters) is still compared.
Scope
In scope: detecting a PDF with no usable text layer and declining it, in compare/pdf.py.
Code:src/deltatrack/compare/pdf.py, lines 94–112 (the guard) and 149–150 (where the decline is raised).
Signal: on the fixtures above, every image-only page extracted zero characters. "No non-empty text on the page" is the simplest candidate. Check it against the 113-hr-3547 print above before relying on a looser per-page floor.
Tests to extend:tests/test_pdf_compare.py, next to test_short_document_is_not_declined (line 465) and the web 422 test that follows it.
Fixtures: build them inside the test rather than committing binaries. Rendering corpus pages with pypdfium2 and saving them with Pillow as a PDF takes about a second for 2×20 pages. Pillow is available through the reportlab dev dependency, and tests/test_pdf_watermark_recall.py shows the pytest.importorskip pattern. Don't depend on tests/data/p3/imageonly.pdf: Remove the closed PDF backend study #783 (open PR that removes the closed PDF backend study) deletes it.
Existing reproduction script: on develop it is docs/research/pdf-backend-bakeoff/probes/make_scan_fixture.py, and Remove the closed PDF backend study #783 moves it to scripts/. It needs macOS (qlmanage, sips). On Linux, the in-test approach above does the same job.
Unverified
Only synthetic scans were tested, not a real scanned draft. A real scan may carry a partial OCR text layer that falls between readable and unreadable.
The web app path was not run. The outcome follows from web/app.py:287–292, which declines only when UnsupportedLayoutError is raised.
Whether the XML pipeline has an equivalent gap for its own inputs.
The two 20-page fixtures cited above are regenerable from a committed corpus bill, so the reproduction does not depend on any scratch directory:
uv run --with pypdf python scripts/make_scan_fixture.py --out /tmp/scans
(Moved from the PDF bake-off probes to scripts/ in #783; pypdf is not a project dependency, hence --with.)
It rasterizes two disjoint 20-page windows of tests/corpus/118-hr-4366/1_reported-in-house.pdf, then re-runs the check and reports whether the defect still reproduces. Current output on spike/pdf-backend-bakeoff (production code there is identical to develop):
What's wrong
If you compare two scanned PDFs, DeltaTrack tells you nothing changed, even when the two documents share no text at all. It can't read either file, but it doesn't say so.
Some background, for anyone new to this:
UnsupportedLayoutError; the web app returns it as an HTTP 422 carrying the message. Today the only decline is for PDFs printed without margin line numbers, such as enrolled bills, added by PDF anchor extraction yields 0 anchors on enrolled bills (no margin line numbers) #141 (closed).How it shows up
Checked on
develop(d3935ee), 2026-10-05. Pages 3–22 and 23–42 oftests/corpus/118-hr-4366/1_reported-in-house.pdfwere rasterized into two 20-page image-only PDFs:diff_pdf.pywrites a normal report that says "No changes found between these versions."tests/data/p3/imageonly.pdf, compared against a text PDF, is answered as 1 removed and 1 added.Why it matters
The web app's compare page has PDF selected by default (
web/webapp/compare.html:39). A first-time visitor who uploads two scanned drafts gets "no changes", which looks the same as two versions that really are identical. That is a confident wrong answer, the outcome the #141 decline exists to prevent.Cause
The decline check only judges documents with at least 50 extracted lines (
_MIN_LINES_FOR_GUARD,src/deltatrack/compare/pdf.py:94, applied in_is_unnumbered_layoutat line 106). That floor is deliberate: short real prints have too few margin numbers to judge. But an image-only page still yields one empty line, so a 20-page scan looks like a 20-line document and is exempt. Nothing checks whether any text was extracted at all.Done when
tests/test_pdf_canonical_baseline.pystays unchanged (23 pairs, 6 declined), and the sparse real printtests/corpus/113-hr-3547/2_engrossed-in-house.pdf(4 pages, 22 non-empty lines, 697 characters) is still compared.Scope
compare/pdf.py.diff_pdf.pyas a traceback; that is A document the tool deliberately declines reaches the command line as a traceback, not as the explanation it prepared #695 (open: a declined document reaches the CLI as a traceback, not as the explanation).Where to start
src/deltatrack/compare/pdf.py, lines 94–112 (the guard) and 149–150 (where the decline is raised).tests/test_pdf_compare.py, next totest_short_document_is_not_declined(line 465) and the web 422 test that follows it.pypdfium2and saving them with Pillow as a PDF takes about a second for 2×20 pages. Pillow is available through thereportlabdev dependency, andtests/test_pdf_watermark_recall.pyshows thepytest.importorskippattern. Don't depend ontests/data/p3/imageonly.pdf: Remove the closed PDF backend study #783 (open PR that removes the closed PDF backend study) deletes it.developit isdocs/research/pdf-backend-bakeoff/probes/make_scan_fixture.py, and Remove the closed PDF backend study #783 moves it toscripts/. It needs macOS (qlmanage,sips). On Linux, the in-test approach above does the same job.Unverified
web/app.py:287–292, which declines only whenUnsupportedLayoutErroris raised.