This was written agentically; verify its assertions:
Slice 9 of epic #152. Design: docs/superpowers/specs/2026-08-13-token-benchmark-design.md, sections "Staged plan" stage 3, "Sampling, not a grid".
Why
Only if stage 2 justifies it: fill the remaining cells (2 arms × 3 effort tiers) to produce the deliverable — acceptance score vs blended USD, two curves. The matrix is sampled, never exhaustive: default rounds are the drift control plus current main per tier; a new model enters only after the frozen arm establishes that model's own baseline.
What
- Remaining cells at the stage-2-set budget, with drift controls per round and
A,B,B,A alternation.
bench/plot.py: the two curves as a pure function of bench/results/ (no hand-carried numbers).
- Epic close-out: verdicts for Q1 (already met) and Q2 written against the pre-registered rules, with the longitudinal machinery (ledger, fixture versioning, drift arm, per-merge Tier 0) left running as standing infrastructure.
Acceptance
🤖 Co-authored by Claude Fable 5.
This was written agentically; verify its assertions:
Slice 9 of epic #152. Design:
docs/superpowers/specs/2026-08-13-token-benchmark-design.md, sections "Staged plan" stage 3, "Sampling, not a grid".Why
Only if stage 2 justifies it: fill the remaining cells (2 arms × 3 effort tiers) to produce the deliverable — acceptance score vs blended USD, two curves. The matrix is sampled, never exhaustive: default rounds are the drift control plus current
mainper tier; a new model enters only after the frozen arm establishes that model's own baseline.What
A,B,B,Aalternation.bench/plot.py: the two curves as a pure function ofbench/results/(no hand-carried numbers).Acceptance
fixture_versionpoints never share a line🤖 Co-authored by Claude Fable 5.