Skip to content

Bring the plan's status, counts and costs back to what is true - #69

Merged
leggetter merged 1 commit into
mainfrom
plan-status-refresh
Aug 29, 2026
Merged

Bring the plan's status, counts and costs back to what is true#69
leggetter merged 1 commit into
mainfrom
plan-status-refresh

Conversation

@leggetter

Copy link
Copy Markdown
Collaborator

An audit of .plans/delivery-plan.md and .plans/evals-page-brief.md against the repository, the releases and the schedule. Six claims had gone stale in ways that would mislead a reader taking the file as current.

Claim Was Is
Status date 18 August 28 August
The scoreboard "open as a pull request on the website" live, publishing from a release; three releases cut; publishing held under #66
Scenarios eighteen (fifteen benchmark) twenty-two (nineteen benchmark)
Weak-model skills delta −3 −2 — both −3 readings are from the credit-outage day and no run since reproduced them
Stage balance two each against eleven build thirteen build, two investigate, four resolve
Weekly run twelve scenarios, docs-only/+MCP arms, ~$50 four experiments weekly and six monthly over nineteen scenarios; per-pair cost not re-measured since the suite grew
Concurrency ninety pairs a hundred and fourteen

The cost section is the one worth reading. Rather than pick a number, it now states what the schedule runs and that the two costings in the file disagree — $1.44 a pair from the ten-scenario re-baseline, ~$0.80 from #24's three-attempt matrix — with the reason (composition: the expensive scenarios run the agent's code and are a smaller fraction of nineteen than of ten). A costing nobody can reproduce is worse than none.

Open questions 1 and 3 were written while the org key gated Phase 3. Phase 3 shipped without it, so it is now the difference between a matrix that finishes in an hour and one that takes most of a day.

The page brief's counts move with the plan's (the grid is 6 x 19, 114 cells), and it gains a second warning: the website team is iterating on /evals, so neither the app in this repository nor the published page is a stable base to design against right now.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK

Six claims in the published plan had gone stale, all in ways that would mislead
someone reading it as current.

- Status was dated 18 August and said the scoreboard was "open as a pull
  request on the website". It is live, publishes from a release rather than
  from whatever ran last, and three releases have been cut. Publishing is held
  as of 27 August, so the status now says so and points at #66 for the
  condition rather than restating it.
- Eighteen scenarios and fifteen benchmark: it is twenty-two and nineteen.
- The weak model's skills delta was recorded as -3. Both -3 readings are from
  13 August, the credit-outage day, and the eight runs since read -2. The sign
  replicated and the magnitude did not, which is worth saying in the same
  breath as the number.
- Investigate and resolve were described as two scenarios each against eleven
  for build. Thirteen build, two investigate, four resolve.
- The weekly cost section still assumed twelve scenarios and docs-only/+MCP
  arms, neither of which exists. Replaced with what the schedule actually runs
  — four experiments weekly, six monthly, nineteen scenarios, one attempt —
  and an explicit statement that per-pair cost has not been re-measured since
  the suite grew, along with the two figures in this file that disagree about
  it. A costing nobody can reproduce is worse than none.
- Ninety pairs is a hundred and fourteen, which is the whole argument for the
  org key.

Open questions 1 and 3 were written when the org key gated Phase 3. Phase 3
shipped without it, so it is now the difference between an hour and most of a
day rather than a blocker.

The page brief's counts move with the plan's, and it gains the fact that the
website team is iterating on /evals — so neither the app in this repository nor
the published page is a stable base to design against right now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
@leggetter
leggetter merged commit 2e324c6 into main Aug 29, 2026
2 checks passed
@leggetter
leggetter deleted the plan-status-refresh branch August 29, 2026 10:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant