You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
The PDF study's conclusion cannot be derived from its evidence, only asserted by hand #729
The PDF backend bake-off's external-validity study can compute its measurements by a reproducible path that reads committed evidence and writes a committed artifact, but it cannot produce its conclusion the same way. There is no operation anywhere that loads the study's committed results and writes a decision. The decision logic exists and is tested, but it has only ever been run on values typed into a test probe by hand.
The study's decision module (probes/decide_architecture.py) implements the frozen Rule 0 / Rule 1 / Rule 3 machinery and exposes decide(inputs: DecisionInputs) at line 698. DecisionInputs (line 248) is the bundle of facts that machinery consumes. Every place in the repository that constructs one is a control probe:
Both are verification probes, and x28 says so in its own self-description:
"artifacts_created": "NONE of frames.json, oracle_key.json, oracle_blind.json,
oracle_adjudicated.json, s1_control.json, cross_engine_control.json, metrics.json,
scores.json, EXECUTION-START.json",
"fixture_caveat": "every payload is real `score_metrics.score(...)` output; Rule 3 gate
STATUSES are overwritten to reach later rules and are fixtures, never evidence",
results/scores.json is named in the apparatus as an artifact that is expected to exist. It appears in x28's FORBIDDEN_ARTIFACTS list (line 1360), which enumerates the outputs a control probe must never create, alongside metrics.json and the oracle files. Nothing on any path writes it.
How it surfaced
Closing out the external-validity run (#728) required stating what the study concluded. There was no artifact to read the conclusion from, and no command that would produce one.
Checked on develop at c636448b:
$ ls .../external-validity/results/scores.json
ls: results/scores.json: No such file or directory
$ grep -rn 'scores\.json' .../external-validity/probes/
x28_decide_architecture.py:1360: "results/scores.json", # FORBIDDEN_ARTIFACTS
x28_decide_architecture.py:1418: "scores.json, EXECUTION-START.json", # artifacts_created: NONE
x21_build_oracle.py:1620: "metrics.json, scores.json, EXECUTION-START.json",
Every reference is a probe declaring it does not write the file.
Why it matters
The scoring side of this study had the same shape and it was treated as a defect worth removing. Deviation-register entry A57 replaced the scoring path with score_canonical(), which "loads the committed inputs, verifies the oracle's origin, scores and writes as ONE operation", specifically so that a caller could not pass in a payload and have a metric computed from facts of their own choosing. The register records the rejected alternative: a payload argument "made 'verify this oracle, write a metric computed from that one' expressible".
The decision side still has exactly that shape, and worse, because it has no canonical path at all rather than a canonical path with an injection point. The practical consequences:
A conclusion cannot be derived, only asserted. Anyone wanting the study's architecture verdict has to hand-assemble DecisionInputs and run decide(), which reproduces precisely the pattern A57 removed. The closeout in Close out the PDF external-validity spike as VOID #728 states the verdict in prose for this reason.
Rule 0 / Rule 1 / Rule 3 have never executed against real committed evidence. They have executed against fixtures that are real scorer output with gate statuses overwritten to reach later branches. That is correct for a control probe and is not evidence about the live path.
There is already a known supplied input on this side.x28 records an open ruling that Rule 1's M4 condition "has NO producer -- the scorer emits per-arm M4 counts only. The verdict is SUPPLIED and its absence REFUSES". A canonical operation would have to resolve or explicitly refuse on that.
Not urgent: the run this would have concluded is void under its own controls (#728), so nothing is currently blocked on it. It matters before any successor run, and it matters for anyone reading the closeout who expects the conclusion to be machine-derived in the way the measurements are.
What to do
Unresolved; the options differ in cost and in what they promise.
Option
What it gives
Cost
Build decide_canonical(), mirroring score_canonical(): load committed metrics.json and the run's provenance, assemble DecisionInputs internally with no caller-supplied payload, write results/scores.json
Symmetry with the scoring side, a derivable conclusion
Must resolve the M4-has-no-producer ruling, or make it refuse explicitly. Touches the result-bearing surface, so it needs a deviation entry and a new authorization before it can run
Document the asymmetry and leave it
Zero cost now
Anyone resuming the study hits the same wall, and the apparatus keeps naming an artifact it never produces
Remove results/scores.json from the apparatus's artifact lists
Stops the apparatus promising something it does not do
Loses the record that a decision artifact was ever intended
The first option is the one that matches the study's own standard, but it is not free and should not be started without deciding the M4 question first.
Verification
"The probes still pass" is not proof of a fix here, because they pass now, with no canonical decision path existing at all. A fix should be shown to:
produce results/scores.json from committed inputs alone, with no payload parameter on the public entry point;
refuse, rather than proceed, when a required input has no producer (the M4 case above);
be caught going RED by a control that supplies a deliberately wrong committed input, in the way A57's scoring controls were demonstrated to fail before being trusted.
Unverified
I did not check whether HARNESS-PLAN.md section 6 or PRE-REGISTRATION section 7.2 specify a required output path or filename for the decision artifact. If either does, that constrains the first option and should be read before designing it.
I also did not check BillTrax or any other consumer for a dependency on results/scores.json existing.
Refs #728 (closes out the external-validity run as void; states the verdict in prose for the reason above)
Closing as not planned. #740 retired the PDF external-validity study as inconclusive, with no architecture recommendation (docs/research/pdf-backend-bakeoff/CLOSEOUT.md), and the closeout records this as a known limit. Worth revisiting only if a successor study starts.
What's wrong
The PDF backend bake-off's external-validity study can compute its measurements by a reproducible path that reads committed evidence and writes a committed artifact, but it cannot produce its conclusion the same way. There is no operation anywhere that loads the study's committed results and writes a decision. The decision logic exists and is tested, but it has only ever been run on values typed into a test probe by hand.
The study's decision module (
probes/decide_architecture.py) implements the frozen Rule 0 / Rule 1 / Rule 3 machinery and exposesdecide(inputs: DecisionInputs)at line 698.DecisionInputs(line 248) is the bundle of facts that machinery consumes. Every place in the repository that constructs one is a control probe:Both are verification probes, and
x28says so in its own self-description:results/scores.jsonis named in the apparatus as an artifact that is expected to exist. It appears inx28'sFORBIDDEN_ARTIFACTSlist (line 1360), which enumerates the outputs a control probe must never create, alongsidemetrics.jsonand the oracle files. Nothing on any path writes it.How it surfaced
Closing out the external-validity run (#728) required stating what the study concluded. There was no artifact to read the conclusion from, and no command that would produce one.
Checked on
developatc636448b:Every reference is a probe declaring it does not write the file.
Why it matters
The scoring side of this study had the same shape and it was treated as a defect worth removing. Deviation-register entry A57 replaced the scoring path with
score_canonical(), which "loads the committed inputs, verifies the oracle's origin, scores and writes as ONE operation", specifically so that a caller could not pass in a payload and have a metric computed from facts of their own choosing. The register records the rejected alternative: a payload argument "made 'verify this oracle, write a metric computed from that one' expressible".The decision side still has exactly that shape, and worse, because it has no canonical path at all rather than a canonical path with an injection point. The practical consequences:
DecisionInputsand rundecide(), which reproduces precisely the pattern A57 removed. The closeout in Close out the PDF external-validity spike as VOID #728 states the verdict in prose for this reason.x28records an open ruling that Rule 1's M4 condition "has NO producer -- the scorer emits per-arm M4 counts only. The verdict is SUPPLIED and its absence REFUSES". A canonical operation would have to resolve or explicitly refuse on that.Not urgent: the run this would have concluded is void under its own controls (#728), so nothing is currently blocked on it. It matters before any successor run, and it matters for anyone reading the closeout who expects the conclusion to be machine-derived in the way the measurements are.
What to do
Unresolved; the options differ in cost and in what they promise.
decide_canonical(), mirroringscore_canonical(): load committedmetrics.jsonand the run's provenance, assembleDecisionInputsinternally with no caller-supplied payload, writeresults/scores.jsonresults/scores.jsonfrom the apparatus's artifact listsThe first option is the one that matches the study's own standard, but it is not free and should not be started without deciding the M4 question first.
Verification
"The probes still pass" is not proof of a fix here, because they pass now, with no canonical decision path existing at all. A fix should be shown to:
results/scores.jsonfrom committed inputs alone, with no payload parameter on the public entry point;Unverified
I did not check whether
HARNESS-PLAN.mdsection 6 or PRE-REGISTRATION section 7.2 specify a required output path or filename for the decision artifact. If either does, that constrains the first option and should be read before designing it.I also did not check BillTrax or any other consumer for a dependency on
results/scores.jsonexisting.Refs #728 (closes out the external-validity run as void; states the verdict in prose for the reason above)
🤖 Generated with Claude Code
https://claude.ai/code/session_01XmeDh5BVXMBkwZ7zCX1oca