Skip to content

The XML comparison cannot load in a browser because it imports a PDF library it never uses #751

Description

@willhea

What needs doing

Loading DeltaTrack's XML comparison also loads the PDF parser and a PDF library it never uses. That stops the XML comparison loading in a web page without a workaround, which only matters if DeltaTrack ends up running inside the browser. It also ties the XML parser to the PDF parser's private code: an edit to the PDF reader changes the XML parser's recorded revision. Make the PDF library load only when a PDF is actually opened, and give the rules both parsers share a home that belongs to neither.

Some background, for anyone new to this:

  • Pyodide. Python compiled to run inside a web page. One delivery option in Open decision: how DeltaTrack reaches staffers (delivery channel), waiting on Hill IT, workflow and packaging signals #112 (open: how DeltaTrack reaches staffers) is running the real engine in the user's browser, so bill text never leaves their machine.
  • pypdfium2. The PDF library the PDF pipeline uses. It has native code and is stubbed out for the browser runs.
  • ADR 0003 records that the XML comparison gives byte-identical output under Pyodide, checked by scripts/pyodide_parity.py, and names this import as the one obstacle.
  • The shared run-in rule. Both parsers recognise a run-in subsection heading ("(a) Catchline.—") the same way, deliberately, so both pipelines derive the same label. Today the XML parser gets that rule by importing private names from the PDF parser.
  • Parser revision. ADR 0019 records which parser code produced a stored judgment, as a hash over the XML parser and every module it imports (tests/round1_identity.parser_revision). Its docstring calls the set "deliberately over-broad".

How it shows up (checked 2026-10-06 on develop, 224af4f)

  • Modules that load the PDF parser and pypdfium2: importing deltatrack.compare.xml, deltatrack.formatters.canonical, deltatrack.bill_tree, deltatrack.diff_bill or deltatrack.structure_tree loads parsers.pdf_anchors, parsers.pdf_text and pypdfium2. With pypdfium2 made unimportable, import deltatrack.compare.xml fails.
  • What The report renderer loads the PDF differ and parsers, so the viewer and the diff can't be worked on apart #801 already fixed: deltatrack.formatters.diff_html, the report renderer, no longer loads any of them since The report renderer loads the PDF differ and parsers, so the viewer and the diff can't be worked on apart #801 (Load the report viewer apart from the diff engine #802). tests/test_import_direction.py holds that.
  • Browser parity check: scripts/pyodide_parity.py --node-dir <dir> passed on 2026-10-05: 3 fixtures × 2 outputs identical between CPython 3.12 and Pyodide 314.0.7. It passes only because it substitutes a stub that raises if PDFium is called.
  • The private import: bill_tree.py:8 imports _RUNIN_QUOTED_LINE and _match_runin_subsection, two private names, from parsers.pdf_anchors.
  • The revision: because bill_tree imports pdf_anchors, and pdf_anchors imports pdf_text, the XML parser revision hashes both PDF modules. The 2026-10-04 audit found that a comment-only edit to pdf_text.py moves it and fails all 27 round-1 provenance checks.

Cause

  • The library: src/deltatrack/parsers/pdf_text.py:28-29 imports pypdfium2 at module scope. Every other use of pypdfium2 in that file is inside a function body.

  • The parser: the XML path reaches pdf_text.py through parsers/pdf_anchors.py along several routes:

    • bill_tree.py:8 (the private run-in rule);
    • structure_tree.py:36 and formatters/canonical.py:23 (Anchor, anchor_positions, breadcrumb_for);
    • diff_pdf.py:58 (Anchor, extract_anchors).

    The PDF-free pieces live in the PDF parser module, so importing them loads the rest.

Done when

  • pdf_text.py imports pypdfium2 inside the functions that use it, not at module scope.
  • The run-in subsection rule both parsers share lives in a module that belongs to neither parser, under public names. bill_tree.py imports nothing from parsers.pdf_anchors.
  • The PDF-free anchor pieces the XML path needs (Anchor, anchor_positions, breadcrumb_for) are importable without loading pdf_text. Either move them, or keep pdf_anchors free of a module-level pdf_text import.
  • A fast test imports the XML comparison, the XML parser and the canonical producers with pypdfium2 blocked (for example sys.modules["pypdfium2"] = None in a subprocess). It goes red when the module-level import is put back. tests/test_import_direction.py is the place.
  • The XML parser revision no longer hashes the PDF parser modules. Stored provenance is re-derived on purpose in the same PR, saying so.
  • PDF output is unchanged: tests/test_pdf_canonical_baseline.py passes without regenerating the baseline. XML output is unchanged too: tests/test_canonical_baseline.py.
  • scripts/pyodide_parity.py still passes, including its --mutate negative control. The stub can then go, or stay as a tripwire.

Where to start

  • A trial run: on a scratch copy of develop (2026-10-05), moving the two pypdfium2 imports into extract_print_pages and _page_glyph_sizes was enough. With pypdfium2 blocked, the XML pair 118-hr-8752 1→2 compared normally.
  • The shared rule: _RUNIN_QUOTED_LINE and _match_runin_subsection in src/deltatrack/parsers/pdf_anchors.py. bill_tree._subsection_label explains why both parsers use one rule. tests/test_xml_subsection_nodes.py checks the XML labels against PDF recall.
  • The revision: parser_revision and _revision_over in tests/round1_identity.py.
  • The browser check: run npm install pyodide in a directory outside the checkout, then uv run python scripts/pyodide_parity.py --node-dir <dir>.

Unverified

  • The trial above didn't run the PDF pipeline or the test suite.
  • Whether pypdfium2 now has a Pyodide build. The study found none in August 2026.

History

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions