The harness is everything wrapped around the model call. A plain call-tools-until-done loop works for a short request and fails in specific, predictable ways on a long one. Four mechanisms address those failures.
A model that has lost the thread calls the same tool with the same arguments, reads the same failure, and calls it again. Left alone it burns the entire turn budget doing that, and the run ends with nothing to show and no explanation.
Every tool call is fingerprinted by name and arguments, normalised so that a re-serialised identical call still counts as the same one. After three, the model is told plainly:
You have called read_file with the same arguments several times and it is not getting you anywhere. Do not call it again. Either try a different approach, or say what is blocking you.
That nudge fires once, not on every subsequent call — otherwise the history fills with the same message. After six, the run stops and says why.
agent:
repeat_limit: 3 # identical calls tolerated before the nudgeDifferent arguments are not repetition. Reading twenty files in a row is fine.
You notice a run going the wrong way. Waiting for it to finish wastes the work; interrupting it throws away the work that was good.
/steer <instruction> queues a note and delivers it after the current batch of
tools completes — the first moment the model can act on it without discarding
anything already underway.
/steer stop looking at the tests, the bug is in the parser
It reports back when nothing is running, so the note is never silently dropped.
Models describe work they did not do. Not usually, and not deliberately, but often enough that a long autonomous run needs a check.
With agent.verify_replies on, a finished answer goes to a second, cheap model
along with the list of tools that actually ran. The critic is asked one thing:
was the request carried out? Not whether it was done well, and not whether more
could be done — only whether what was asked for actually happened.
When it finds something missing, the turn does not end. The gap is fed back:
Before finishing: the second file was never written. Do that now, or explain why it cannot be done.
This is bounded by agent.verify_max, because a critic that is never satisfied
would loop forever. Two passes is the default.
agent:
verify_replies: true
verify_max: 2It is off by default because it costs an extra call per turn. On a long unattended run — a cron job, a standing goal — it earns that cost back the first time it catches a silent omission.
Sub-agents are not verified: whoever delegated the work checks it.
A goal outlives a turn.
/goal get the test suite green, then open a pull request
From then on the goal is part of every system prompt, and when the agent thinks it is finished, a judge model decides whether the goal is actually met. If not, the judge names one concrete next step and the loop continues with it.
/goal status where it stands, and how many iterations it has taken
/goal pause hold it without losing it
/goal resume pick it back up
/goal clear drop it
The judge is told that a plan is not completion and an intention is not doing — but also not to invent extra work. If the goal is met, it says so.
Iterations are capped:
agent:
goal_max_iterations: 10At the cap the goal pauses rather than stopping, so /goal resume continues from
where it got to. If no judge model is available, the loop ends rather than
running forever.
/learn turns what just happened into a skill.
The session transcript goes to a model with one instruction: write the procedure someone would want next time, with the specific commands, paths, and gotchas that actually came up — and skip everything particular to this one conversation. If nothing general was learned, it says so and writes nothing.
What comes back is saved into the skill library, front matter and all, and is offered the next time something similar comes up.
/learn distil the whole session
/learn the deploy sequence focus on one part of it
This is the loop that makes the agent better at your work specifically: solve something once, keep the procedure, do it faster next time.
A long autonomous run with all four:
/goal migrate the config loader to the new schema- The agent works. The repetition guard catches it re-reading the same file and redirects it.
- You notice it editing the wrong package:
/steer the loader is in internal/config, not internal/settings. - It says it is done. Verification notices the tests were never run and sends it back.
- It runs them, fixes two failures, and says it is done again.
- The judge agrees the goal is met and the run ends.
/learnkeeps the migration procedure for the next one.
One model's blind spot is often another's strength. /panel asks several the
same question independently and synthesises one answer.
model:
panel:
- anthropic/claude-sonnet-4.5
- openai/gpt-5
- google/gemini-3-pro/panel is it safe to run this migration on a live database?
The value is in the disagreement. Where independent answers diverge is usually where the question was ambiguous or the problem is genuinely hard — exactly what a single sample hides.
The synthesiser is told not to average: where the answers differ it says so and follows the one best supported by its own reasoning, rather than producing something none of them said. A model that fails to answer is named, so a quiet failure is not mistaken for a smaller panel.
It costs one call per model plus one to synthesise, so it is a command rather than a mode.
Every write and edit copies what was there first, keyed by session.
/rollback list what changed
/rollback all put all of it back
/rollback <path> put one file back
The first copy of a file wins, so rolling back returns it to how it was before the session started rather than before the most recent edit — undoing one step of five is rarely what "undo this" means. A file that did not exist before is deleted, because that is what restoring it means.
Bare /rollback only lists. Undoing work takes a second word.
| Setting | Default | What it controls |
|---|---|---|
agent.repeat_limit |
3 | Identical calls before the nudge |
agent.verify_replies |
off | Check answers against the request |
agent.verify_max |
2 | How many times a turn can be sent back |
agent.goal_max_iterations |
10 | Iterations before a goal pauses |
agent.max_turns |
200 | Hard ceiling on model calls per run |
tools.max_tool_calls_per_turn |
32 | Tool budget before the model is told to wrap up |
The two hard limits are the backstop. Everything above them is about failing usefully rather than failing quietly.