Bring the plan's status, counts and costs back to what is true - #69
Merged
Conversation
Six claims in the published plan had gone stale, all in ways that would mislead someone reading it as current. - Status was dated 18 August and said the scoreboard was "open as a pull request on the website". It is live, publishes from a release rather than from whatever ran last, and three releases have been cut. Publishing is held as of 27 August, so the status now says so and points at #66 for the condition rather than restating it. - Eighteen scenarios and fifteen benchmark: it is twenty-two and nineteen. - The weak model's skills delta was recorded as -3. Both -3 readings are from 13 August, the credit-outage day, and the eight runs since read -2. The sign replicated and the magnitude did not, which is worth saying in the same breath as the number. - Investigate and resolve were described as two scenarios each against eleven for build. Thirteen build, two investigate, four resolve. - The weekly cost section still assumed twelve scenarios and docs-only/+MCP arms, neither of which exists. Replaced with what the schedule actually runs — four experiments weekly, six monthly, nineteen scenarios, one attempt — and an explicit statement that per-pair cost has not been re-measured since the suite grew, along with the two figures in this file that disagree about it. A costing nobody can reproduce is worse than none. - Ninety pairs is a hundred and fourteen, which is the whole argument for the org key. Open questions 1 and 3 were written when the org key gated Phase 3. Phase 3 shipped without it, so it is now the difference between an hour and most of a day rather than a blocker. The page brief's counts move with the plan's, and it gains the fact that the website team is iterating on /evals — so neither the app in this repository nor the published page is a stable base to design against right now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An audit of
.plans/delivery-plan.mdand.plans/evals-page-brief.mdagainst the repository, the releases and the schedule. Six claims had gone stale in ways that would mislead a reader taking the file as current.The cost section is the one worth reading. Rather than pick a number, it now states what the schedule runs and that the two costings in the file disagree — $1.44 a pair from the ten-scenario re-baseline, ~$0.80 from #24's three-attempt matrix — with the reason (composition: the expensive scenarios run the agent's code and are a smaller fraction of nineteen than of ten). A costing nobody can reproduce is worse than none.
Open questions 1 and 3 were written while the org key gated Phase 3. Phase 3 shipped without it, so it is now the difference between a matrix that finishes in an hour and one that takes most of a day.
The page brief's counts move with the plan's (the grid is 6 x 19, 114 cells), and it gains a second warning: the website team is iterating on
/evals, so neither the app in this repository nor the published page is a stable base to design against right now.🤖 Generated with Claude Code
https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK