Skip to content

Correct the weak-model skills delta to −2 - #59

Merged
leggetter merged 1 commit into
mainfrom
correct-skills-figure
Aug 26, 2026
Merged

Correct the weak-model skills delta to −2#59
leggetter merged 1 commit into
mainfrom
correct-skills-figure

Conversation

@leggetter

Copy link
Copy Markdown
Collaborator

AGENTS.md carried −3 for the weak model's skills delta since 13 August. It is −2.

Both −3 readings are from 13 August and nothing since has reproduced them:

run delta
08-13 ×2 −3
08-14 0
08-17 ×3, 08-18, 08-19, 08-24, 08-25 −2

13 August is the credit-outage day AGENTS.md documents under costs — thirty-seven Codex jobs failed and twenty-two earlier runs were scored with the agent never making a tool call. Whether these two snapshots contain affected rows needs the job start times against the outage window, which is the method that file prescribes and which was never applied to this figure. Flagged in the text rather than asserted.

The sign is replicated and unaffected. Only the magnitude was wrong.

What else changed

  • The recomputation is inline, so the next person derives it from results/runs/ instead of quoting the paragraph. That is how it went stale.
  • A warning against a conflation I made in conversation: Loop 2's corrected Outpost measurement is +2 for skills, and it does not contradict this. Five Outpost scenarios across three models is not fifteen-plus scenarios on the weak model. Stating the scenario set and the model with any skills delta is now a rule.
  • Scenario counts corrected: twenty-two scenarios, nineteen benchmark, nine of nineteen discriminate — was eighteen/fifteen and eight of fifteen.

Related, not in this PR

#2 has the full recomputation. The composition of the delta has changed completely: transform-001-reshape-payload is worse with skills in all ten runs and is the only durable one; the other three originals dissolved; and verification-002 inverted to a skills win holding for eight consecutive runs with no skill change to explain it (the submodule pin has not moved since 12 August).

#24 updated with the same correction.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Nt2Zgjw7STjrnFXYKRRVAA

AGENTS.md carried -3 for eleven days. Both -3 readings are from 13
August and no run since has reproduced them: eight of the nine later runs
read -2, one read 0. 13 August is the credit-outage day this same file
documents under costs, where jobs were scored without the agent acting —
and the figure was never re-derived after that cause was identified,
which is exactly the mistake the "identify a bad population by its cause"
trap was written about.

Adds the recomputation inline so the next person derives it rather than
quoting the paragraph, and a warning against reading Loop 2's Outpost-only
+2 as a contradiction: five scenarios across three models is not fifteen
scenarios on one, and the two were conflated once into a claim that the
sign had flipped, which no run supports.

Scenario counts corrected while here: twenty-two scenarios, nineteen
benchmark, and nine of nineteen discriminate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nt2Zgjw7STjrnFXYKRRVAA
@leggetter
leggetter merged commit 7356e80 into main Aug 26, 2026
2 checks passed
@leggetter
leggetter deleted the correct-skills-figure branch August 26, 2026 11:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant