Conversation
The external-validity run fails its own pre-registered N-B control, whose §5.6 consequence is "run void". This lands the closeout document plus the two artifacts every figure in it is derived from. WHAT VOIDED IT. Controls require both adjudication routes with no tolerance. N-A 10/16 FAIL, N-B 14/16 FAIL, N-C 8/8 PASS. Every N-B failure and five of six N-A failures are human-route. Both human failure modes are systematic: spacing defects normalised away (WELD 0/3, SPLIT 0/2) and RESCISSION read as RECISSION. Those are the defect classes the bake-off exists to detect, so the ground truth repaired the thing being measured. Excluding the one fixture whose injected defect is not legible in the render (#727), the AI route is 8/8, 8/8, 4/4. WHY THE RQs HAVE NO ANSWER. RQ1's D-frame design required every one of 13992 census regions human-adjudicated; 45 human answers exist. RQ2's figures are recorded but M2 and M3 are void via the R1 text gate, H and X are identical on every pooled C metric, and the run is void via N-B. NOT THE POPULATION. §4.5 adequacy returned GENERALISABLE (2583 occurrences, 7 strata), all 17 documents scored, extraction clean across 4190 pages. Adjudication failed, not sampling, so a different document sample would not help. WHAT IS INCLUDED AND WHY. metrics.json carries the control verdicts, R1 pair facts, M1-M5 and the adequacy verdict. DEVIATIONS.md advances from A56 to A62.4 and is append-only against develop (develop's copy is a byte-exact prefix). oracle_adjudicated.json and ORACLE-PROVENANCE.json are deliberately NOT included: they are answer-key material and publishing them would contaminate any successor run. The apparatus, the 307 MB key and the rendered stimuli stay on the local branch pdf-study-continuation-execution, which cannot be pushed because it tracks oracle_key.json. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XmeDh5BVXMBkwZ7zCX1oca
This was referenced Sep 22, 2026
Collaborator
Author
|
Superseded by #740, a one-page closeout built from this PR's base that publishes no scored outputs, register entries, blind identifiers or individual judgments. Closing this one unmerged. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes out the PDF backend bake-off's external-validity run. The run is void under its own pre-registered controls, and this lands the verdict plus the artifacts it is derived from, so the finding stops living only on an unpushable local branch.
Why void
§5.6's N-B control failed, and its stated consequence is "run void" — the adjudicator is unreliable independently of any architecture.
Every N-B failure and five of six N-A failures are human-route. The AI's single N-A failure is the one fixture whose injected defect is not legible in the render (#727), where both routes failed identically. Setting it aside, the AI route is 8/8, 8/8, 4/4.
Both human failure modes are systematic rather than noisy:
OPERATIONSAND SUPPORTtranscribedOPERATIONS AND SUPPORT.RESCISSIONread asRECISSION, which is both N-B failures. The AI route is 4/4 on the same word.Weld and split are exactly the defect classes the bake-off exists to detect, so the ground truth silently repaired the thing being measured. That makes the failure disqualifying rather than a tolerance question.
Why the research questions have no answer
RQ1 — its D-frame design required every one of 13,992 census regions human-adjudicated. 45 human answers exist in total. The human arm was never executed at census scale, so RQ1 was foreclosed before the controls ran.
RQ2 — figures were computed but M2 and M3 are void via the R1 text gate, H and X are identical on every pooled C metric, and the run is void via N-B.
Not a sampling problem
§4.5 adequacy returned GENERALISABLE (2583 occurrences, 7 filled strata), all 17 documents scored, extraction clean across 4190 pages. Adjudication failed, not sampling — which is why re-running with different documents would not help, and why re-running the human arm is not recommended.
What is worth keeping
What's in the diff
RESULTS-EXTERNAL-VALIDITY.md— the closeout, at the spike root rather than undervalidation/external-validity/because F9'sPROTECTED_SUFFIXEScovers.mdunder the study directory.results/metrics.json— the scored output. Every figure above is derivable from it.results/DEVIATIONS.md— advances A56 → A62.4. Append-only: develop's copy is a byte-exact prefix of this one, verified before overwriting.Deliberately excluded:
oracle_adjudicated.jsonandORACLE-PROVENANCE.json. They are answer-key material (mode 0600 in the working tree) and publishing them would contaminate any successor run.The apparatus, the 307 MB answer key and the 15,437 rendered stimuli remain on the local branch
pdf-study-continuation-execution, which cannot be pushed because it tracksoracle_key.json.Follow-ups, not in this PR
NOT_EVALUABLE_NO_R1_PAIRSwhilerule3_gatesaccepts onlyNOT_EVALUABLE, so a zero-evidence R1 falls through to PASS. Fail-open; fix before any successor run.score_canonical()exists on the scoring side; there is no equivalent on the decision side. Not built here on purpose: the void verdict is already machine-computed inmetrics.jsoncontrol_verdicts, so a decision module would re-emit a foregone conclusion. To be filed as an issue.🤖 Generated with Claude Code
https://claude.ai/code/session_01XmeDh5BVXMBkwZ7zCX1oca