Skip to content

The PDF study's conclusion cannot be derived from its evidence, only asserted by hand #729

Description

@willhea

What's wrong

The PDF backend bake-off's external-validity study can compute its measurements by a reproducible path that reads committed evidence and writes a committed artifact, but it cannot produce its conclusion the same way. There is no operation anywhere that loads the study's committed results and writes a decision. The decision logic exists and is tested, but it has only ever been run on values typed into a test probe by hand.

The study's decision module (probes/decide_architecture.py) implements the frozen Rule 0 / Rule 1 / Rule 3 machinery and exposes decide(inputs: DecisionInputs) at line 698. DecisionInputs (line 248) is the bundle of facts that machinery consumes. Every place in the repository that constructs one is a control probe:

$ grep -rln 'DecisionInputs(' docs/research/pdf-backend-bakeoff/validation/external-validity/probes/
.../probes/x28_decide_architecture.py
.../probes/x31_dframe_budget_routes.py

Both are verification probes, and x28 says so in its own self-description:

"artifacts_created": "NONE of frames.json, oracle_key.json, oracle_blind.json,
 oracle_adjudicated.json, s1_control.json, cross_engine_control.json, metrics.json,
 scores.json, EXECUTION-START.json",
"fixture_caveat": "every payload is real `score_metrics.score(...)` output; Rule 3 gate
 STATUSES are overwritten to reach later rules and are fixtures, never evidence",

results/scores.json is named in the apparatus as an artifact that is expected to exist. It appears in x28's FORBIDDEN_ARTIFACTS list (line 1360), which enumerates the outputs a control probe must never create, alongside metrics.json and the oracle files. Nothing on any path writes it.

How it surfaced

Closing out the external-validity run (#728) required stating what the study concluded. There was no artifact to read the conclusion from, and no command that would produce one.

Checked on develop at c636448b:

$ ls .../external-validity/results/scores.json
ls: results/scores.json: No such file or directory

$ grep -rn 'scores\.json' .../external-validity/probes/
x28_decide_architecture.py:1360:    "results/scores.json",          # FORBIDDEN_ARTIFACTS
x28_decide_architecture.py:1418:        "scores.json, EXECUTION-START.json",   # artifacts_created: NONE
x21_build_oracle.py:1620:        "metrics.json, scores.json, EXECUTION-START.json",

Every reference is a probe declaring it does not write the file.

Why it matters

The scoring side of this study had the same shape and it was treated as a defect worth removing. Deviation-register entry A57 replaced the scoring path with score_canonical(), which "loads the committed inputs, verifies the oracle's origin, scores and writes as ONE operation", specifically so that a caller could not pass in a payload and have a metric computed from facts of their own choosing. The register records the rejected alternative: a payload argument "made 'verify this oracle, write a metric computed from that one' expressible".

The decision side still has exactly that shape, and worse, because it has no canonical path at all rather than a canonical path with an injection point. The practical consequences:

  • A conclusion cannot be derived, only asserted. Anyone wanting the study's architecture verdict has to hand-assemble DecisionInputs and run decide(), which reproduces precisely the pattern A57 removed. The closeout in Close out the PDF external-validity spike as VOID #728 states the verdict in prose for this reason.
  • Rule 0 / Rule 1 / Rule 3 have never executed against real committed evidence. They have executed against fixtures that are real scorer output with gate statuses overwritten to reach later branches. That is correct for a control probe and is not evidence about the live path.
  • There is already a known supplied input on this side. x28 records an open ruling that Rule 1's M4 condition "has NO producer -- the scorer emits per-arm M4 counts only. The verdict is SUPPLIED and its absence REFUSES". A canonical operation would have to resolve or explicitly refuse on that.

Not urgent: the run this would have concluded is void under its own controls (#728), so nothing is currently blocked on it. It matters before any successor run, and it matters for anyone reading the closeout who expects the conclusion to be machine-derived in the way the measurements are.

What to do

Unresolved; the options differ in cost and in what they promise.

Option What it gives Cost
Build decide_canonical(), mirroring score_canonical(): load committed metrics.json and the run's provenance, assemble DecisionInputs internally with no caller-supplied payload, write results/scores.json Symmetry with the scoring side, a derivable conclusion Must resolve the M4-has-no-producer ruling, or make it refuse explicitly. Touches the result-bearing surface, so it needs a deviation entry and a new authorization before it can run
Document the asymmetry and leave it Zero cost now Anyone resuming the study hits the same wall, and the apparatus keeps naming an artifact it never produces
Remove results/scores.json from the apparatus's artifact lists Stops the apparatus promising something it does not do Loses the record that a decision artifact was ever intended

The first option is the one that matches the study's own standard, but it is not free and should not be started without deciding the M4 question first.

Verification

"The probes still pass" is not proof of a fix here, because they pass now, with no canonical decision path existing at all. A fix should be shown to:

  • produce results/scores.json from committed inputs alone, with no payload parameter on the public entry point;
  • refuse, rather than proceed, when a required input has no producer (the M4 case above);
  • be caught going RED by a control that supplies a deliberately wrong committed input, in the way A57's scoring controls were demonstrated to fail before being trusted.

Unverified

I did not check whether HARNESS-PLAN.md section 6 or PRE-REGISTRATION section 7.2 specify a required output path or filename for the decision artifact. If either does, that constrains the first option and should be read before designing it.

I also did not check BillTrax or any other consumer for a dependency on results/scores.json existing.

Refs #728 (closes out the external-validity run as void; states the verdict in prose for the reason above)

🤖 Generated with Claude Code

https://claude.ai/code/session_01XmeDh5BVXMBkwZ7zCX1oca

Activity

  1. added theissue type on Sep 22, 2026
  2. willhea commented on Oct 4, 2026

    @willhea
    CollaboratorAuthor

    Closing as not planned. #740 retired the PDF external-validity study as inconclusive, with no architecture recommendation (docs/research/pdf-backend-bakeoff/CLOSEOUT.md), and the closeout records this as a known limit. Worth revisiting only if a successor study starts.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions