Skip to content

test(bench): stage 3 frontier — effort curves and epic close-out #331

Description

@thewrz

This was written agentically; verify its assertions:

Slice 9 of epic #152. Design: docs/superpowers/specs/2026-08-13-token-benchmark-design.md, sections "Staged plan" stage 3, "Sampling, not a grid".

Why

Only if stage 2 justifies it: fill the remaining cells (2 arms × 3 effort tiers) to produce the deliverable — acceptance score vs blended USD, two curves. The matrix is sampled, never exhaustive: default rounds are the drift control plus current main per tier; a new model enters only after the frozen arm establishes that model's own baseline.

What

  • Remaining cells at the stage-2-set budget, with drift controls per round and A,B,B,A alternation.
  • bench/plot.py: the two curves as a pure function of bench/results/ (no hand-carried numbers).
  • Epic close-out: verdicts for Q1 (already met) and Q2 written against the pre-registered rules, with the longitudinal machinery (ledger, fixture versioning, drift arm, per-merge Tier 0) left running as standing infrastructure.

Acceptance

🤖 Co-authored by Claude Fable 5.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/testsThe suite and its gatesenhancementNew feature or requestp2Wanted before public release

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions