Skip to content

Close out the PDF external-validity spike as VOID - #728

Closed
willhea wants to merge 1 commit into
developfrom
spike-external-validity-closeout
Closed

willhea wants to merge 1 commit into
developfrom
spike-external-validity-closeout

Conversation

@willhea

@willhea willhea commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

Closes out the PDF backend bake-off's external-validity run. The run is void under its own pre-registered controls, and this lands the verdict plus the artifacts it is derived from, so the finding stops living only on an unpushable local branch.

Why void

§5.6's N-B control failed, and its stated consequence is "run void" — the adjudicator is unreliable independently of any architecture.

control AI route human route combined status
N-A (injected defects transcribed, not repaired) 7/8 3/8 10/16 FAIL
N-B (XML-corroborated headings agree) 8/8 6/8 14/16 FAIL
N-C (no heading where none printed) 4/4 4/4 8/8 PASS

Every N-B failure and five of six N-A failures are human-route. The AI's single N-A failure is the one fixture whose injected defect is not legible in the render (#727), where both routes failed identically. Setting it aside, the AI route is 8/8, 8/8, 4/4.

Both human failure modes are systematic rather than noisy:

  • spacing defects normalised away — DELETE 3/3 but WELD 0/3 and SPLIT 0/2. OPERATIONSAND SUPPORT transcribed OPERATIONS AND SUPPORT.
  • familiar spelling substituted — RESCISSION read as RECISSION, which is both N-B failures. The AI route is 4/4 on the same word.

Weld and split are exactly the defect classes the bake-off exists to detect, so the ground truth silently repaired the thing being measured. That makes the failure disqualifying rather than a tolerance question.

Why the research questions have no answer

RQ1 — its D-frame design required every one of 13,992 census regions human-adjudicated. 45 human answers exist in total. The human arm was never executed at census scale, so RQ1 was foreclosed before the controls ran.

RQ2 — figures were computed but M2 and M3 are void via the R1 text gate, H and X are identical on every pooled C metric, and the run is void via N-B.

Not a sampling problem

§4.5 adequacy returned GENERALISABLE (2583 occurrences, 7 filled strata), all 17 documents scored, extraction clean across 4190 pages. Adjudication failed, not sampling — which is why re-running with different documents would not help, and why re-running the human arm is not recommended.

What is worth keeping

  1. An AI image-adjudicator outperformed human ground truth on this study's own controls. The study was designed assuming the opposite.
  2. Human adjudicators read for meaning — they normalise spacing and correct spelling without noticing. Any protocol using human transcription as ground truth for character-level defects needs a control that catches this.
  3. A compositional heading definition needs an explicit page-furniture rule. Read strictly, this one makes page numbers headings; nobody read it that way, and the routes diverged on bill designators as a result (register entry A62.4).

What's in the diff

  • RESULTS-EXTERNAL-VALIDITY.md — the closeout, at the spike root rather than under validation/external-validity/ because F9's PROTECTED_SUFFIXES covers .md under the study directory.
  • results/metrics.json — the scored output. Every figure above is derivable from it.
  • results/DEVIATIONS.md — advances A56 → A62.4. Append-only: develop's copy is a byte-exact prefix of this one, verified before overwriting.

Deliberately excluded: oracle_adjudicated.json and ORACLE-PROVENANCE.json. They are answer-key material (mode 0600 in the working tree) and publishing them would contaminate any successor run.

The apparatus, the 307 MB answer key and the 15,437 rendered stimuli remain on the local branch pdf-study-continuation-execution, which cannot be pushed because it tracks oracle_key.json.

Follow-ups, not in this PR

🤖 Generated with Claude Code

https://claude.ai/code/session_01XmeDh5BVXMBkwZ7zCX1oca

The external-validity run fails its own pre-registered N-B control, whose §5.6
consequence is "run void". This lands the closeout document plus the two
artifacts every figure in it is derived from.

WHAT VOIDED IT. Controls require both adjudication routes with no tolerance.
N-A 10/16 FAIL, N-B 14/16 FAIL, N-C 8/8 PASS. Every N-B failure and five of six
N-A failures are human-route. Both human failure modes are systematic: spacing
defects normalised away (WELD 0/3, SPLIT 0/2) and RESCISSION read as RECISSION.
Those are the defect classes the bake-off exists to detect, so the ground truth
repaired the thing being measured. Excluding the one fixture whose injected
defect is not legible in the render (#727), the AI route is 8/8, 8/8, 4/4.

WHY THE RQs HAVE NO ANSWER. RQ1's D-frame design required every one of 13992
census regions human-adjudicated; 45 human answers exist. RQ2's figures are
recorded but M2 and M3 are void via the R1 text gate, H and X are identical on
every pooled C metric, and the run is void via N-B.

NOT THE POPULATION. §4.5 adequacy returned GENERALISABLE (2583 occurrences, 7
strata), all 17 documents scored, extraction clean across 4190 pages.
Adjudication failed, not sampling, so a different document sample would not
help.

WHAT IS INCLUDED AND WHY. metrics.json carries the control verdicts, R1 pair
facts, M1-M5 and the adequacy verdict. DEVIATIONS.md advances from A56 to A62.4
and is append-only against develop (develop's copy is a byte-exact prefix).
oracle_adjudicated.json and ORACLE-PROVENANCE.json are deliberately NOT included:
they are answer-key material and publishing them would contaminate any successor
run.

The apparatus, the 307 MB key and the rendered stimuli stay on the local branch
pdf-study-continuation-execution, which cannot be pushed because it tracks
oracle_key.json.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XmeDh5BVXMBkwZ7zCX1oca
@willhea

willhea commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by #740, a one-page closeout built from this PR's base that publishes no scored outputs, register entries, blind identifiers or individual judgments. Closing this one unmerged.

@willhea willhea closed this Sep 30, 2026
@willhea
willhea deleted the spike-external-validity-closeout branch October 5, 2026 02:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant