Record what the do-not-ask instruction bought - #63
Merged
Conversation
Four cells had tripped ASKED_AND_STOPPED. Re-run under the new base prompt, none of them does: three now pass, and the fourth made 68 tool calls, stopped cleanly and failed 3/5 on the handler-side signature checks — a capability failure on its merits rather than an agent waiting for a person. About $3. Recorded with the caveat that matters more than the number. The detector going 4->0 is the direct result, because that is what the instruction targets. The three flips to passing are consistent with it and are not evidence of it: failing cells, re-run once, no control arm. Reading them as three scenarios bought would be Loop 2's mistake in miniature. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nt2Zgjw7STjrnFXYKRRVAA
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up measurement for #62. Step 3 of the plan: re-run the four cells that had tripped
ASKED_AND_STOPPEDand check whether the instruction changes anything.Result
The detector goes 4 → 0.
resolve-002×claude-code-sonnet-5outpost-003×claude-code-sonnet-5-no-skillstransform-001×codex-gpt-5.4-miniverification-002×codex-gpt-5.4-miniThe remaining failure is now a real one. It created the ElevenLabs source, passed the genuine and forged signature checks, and failed the two handler-side Hookdeck signature checks. Its report ends by offering further work rather than blocking on a question, which is why the detector does not fire. That is a capability failure on its merits.
Total cost about $3.
The caveat, which matters more than the number
Read this as one measurement, not four.
The detector going 4 → 0 is the direct result, because suppressing that behaviour is exactly what the instruction targets. The three flips to passing are consistent with it and are not evidence of it — these were failing cells, re-run once, with no control arm, and a failing cell that is re-run can flip on its own.
Claiming the instruction bought three scenarios would be Loop 2's mistake in miniature: a striking number from a single uncontrolled pass. AGENTS.md now says so in as many words.
Not a loop
This changed our instrument, not the product, the docs or the skills, so it does not go in
LOOPS.mdand it belongs in a release'sBenchmarksection rather thanShipped. Noting it explicitly because the shape is loop-like enough to be filed wrongly.Publishing is still held
EVALS_PUBLISH=falsestays until the milestone-2 run lands a snapshot measured end-to-end under the new prompt. Monday's cron and the 1 September full matrix would otherwise publish a mix of treatments.🤖 Generated with Claude Code
https://claude.ai/code/session_01Nt2Zgjw7STjrnFXYKRRVAA