Skip to content

Two different scanned PDFs compare as "no changes" instead of being declined #550

Description

@willhea

What's wrong

If you compare two scanned PDFs, DeltaTrack tells you nothing changed, even when the two documents share no text at all. It can't read either file, but it doesn't say so.

Some background, for anyone new to this:

  • The two pipelines. DeltaTrack compares bill versions either from their XML (structured text Congress publishes) or from their PDF (the printed page). The PDF pipeline exists for drafts that circulate only as PDF (ADR 0010). It reads the PDF's text layer, the selectable text stored behind the page image.
  • Scanned (image-only) PDFs. A scanned or photocopied page is just a picture. It has no text layer, so there is nothing for DeltaTrack to read.
  • Declining. When DeltaTrack knows it can't give a trustworthy answer, it declines: it refuses, with a message written for the user, instead of guessing. In code that is UnsupportedLayoutError; the web app returns it as an HTTP 422 carrying the message. Today the only decline is for PDFs printed without margin line numbers, such as enrolled bills, added by PDF anchor extraction yields 0 anchors on enrolled bills (no margin line numbers) #141 (closed).

How it shows up

Checked on develop (d3935ee), 2026-10-05. Pages 3–22 and 23–42 of tests/corpus/118-hr-4366/1_reported-in-house.pdf were rasterized into two 20-page image-only PDFs:

scan20.pdf   pages=20 extracted_lines=20 chars=0 guard_declines=False
scan20b.pdf  pages=20 extracted_lines=20 chars=0 guard_declines=False
compare_pdfs(scan20, scan20b) -> changes=0 summary={}
  • The command line agrees. diff_pdf.py writes a normal report that says "No changes found between these versions."
  • The control works. The same two page ranges with their text intact extract 370 lines and compare as 89 changes.
  • Page count decides the outcome. Scans of 1 to 49 pages are answered "no changes". Scans of 50 pages or more are declined, but with the enrolled-bill message ("no printed line numbers … Use the XML version"). That is wrong for a scan, since a scanned draft has no XML version.
  • One scanned side isn't caught either. The committed one-page image-only fixture tests/data/p3/imageonly.pdf, compared against a text PDF, is answered as 1 removed and 1 added.

Why it matters

The web app's compare page has PDF selected by default (web/webapp/compare.html:39). A first-time visitor who uploads two scanned drafts gets "no changes", which looks the same as two versions that really are identical. That is a confident wrong answer, the outcome the #141 decline exists to prevent.

Cause

The decline check only judges documents with at least 50 extracted lines (_MIN_LINES_FOR_GUARD, src/deltatrack/compare/pdf.py:94, applied in _is_unnumbered_layout at line 106). That floor is deliberate: short real prints have too few margin numbers to judge. But an image-only page still yields one empty line, so a 20-page scan looks like a 20-line document and is exempt. Nothing checks whether any text was extracted at all.

Done when

  • Comparing two image-only PDFs of any length, including 1 page and 20 pages, is declined.
  • Comparing an image-only PDF with a text PDF is declined.
  • The decline message describes a scan, not missing line numbers. Proposed wording: "This PDF appears to contain images rather than extractable text and cannot currently be compared."
  • The web app returns that message as a 422.
  • No committed corpus PDF is newly declined. tests/test_pdf_canonical_baseline.py stays unchanged (23 pairs, 6 declined), and the sparse real print tests/corpus/113-hr-3547/2_engrossed-in-house.pdf (4 pages, 22 non-empty lines, 697 characters) is still compared.

Scope

Where to start

  • Code: src/deltatrack/compare/pdf.py, lines 94–112 (the guard) and 149–150 (where the decline is raised).
  • Signal: on the fixtures above, every image-only page extracted zero characters. "No non-empty text on the page" is the simplest candidate. Check it against the 113-hr-3547 print above before relying on a looser per-page floor.
  • Tests to extend: tests/test_pdf_compare.py, next to test_short_document_is_not_declined (line 465) and the web 422 test that follows it.
  • Fixtures: build them inside the test rather than committing binaries. Rendering corpus pages with pypdfium2 and saving them with Pillow as a PDF takes about a second for 2×20 pages. Pillow is available through the reportlab dev dependency, and tests/test_pdf_watermark_recall.py shows the pytest.importorskip pattern. Don't depend on tests/data/p3/imageonly.pdf: Remove the closed PDF backend study #783 (open PR that removes the closed PDF backend study) deletes it.
  • Existing reproduction script: on develop it is docs/research/pdf-backend-bakeoff/probes/make_scan_fixture.py, and Remove the closed PDF backend study #783 moves it to scripts/. It needs macOS (qlmanage, sips). On Linux, the in-test approach above does the same job.

Unverified

  • Only synthetic scans were tested, not a real scanned draft. A real scan may carry a partial OCR text layer that falls between readable and unreadable.
  • The web app path was not run. The outcome follows from web/app.py:287–292, which declines only when UnsupportedLayoutError is raised.
  • Whether the XML pipeline has an equivalent gap for its own inputs.

Activity

  1. willhea commented on Aug 6, 2026

    @willhea
    CollaboratorAuthor

    The two 20-page fixtures cited above are regenerable from a committed corpus bill, so the reproduction does not depend on any scratch directory:

    uv run --with pypdf python scripts/make_scan_fixture.py --out /tmp/scans
    

    (Moved from the PDF bake-off probes to scripts/ in #783; pypdf is not a project dependency, hence --with.)

    It rasterizes two disjoint 20-page windows of tests/corpus/118-hr-4366/1_reported-in-house.pdf, then re-runs the check and reports whether the defect still reproduces. Current output on spike/pdf-backend-bakeoff (production code there is identical to develop):

    scan20.pdf   pages=20 extracted_lines= 20 guard_declines=False
    scan20b.pdf  pages=20 extracted_lines= 20 guard_declines=False
    
    compare_pdfs(two disjoint 20-page scans) -> changes=0, summary={}
    (guard exemption threshold: fewer than 50 extracted lines)
    
    DEFECT REPRODUCED
    

    macOS only as written: it uses qlmanage and sips to rasterize. On Linux, substitute pdftoppm plus img2pdf.

  2. added a commit that references this issue on Aug 14, 2026
  3. added 2 commits that reference this issue on Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions