Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,9 @@ summary: Timeline of guardrail helper changes mirrored from Sweetistics and rela
# Changelog

## 2026-09-29: GPT-6.1 Sol routing
- Correct the remaining Autoreview prose summary to GPT-6.1 Sol xhigh, matching its executable and documented defaults.
- Move hardest tasks and Reflect judgment, divergent, and synthesis roles to GPT-6.1 Sol high; align their setup and fallback defaults.
- Replace Luna discovery, investigation, Swarm, comment-audit, and live-verification routes with GPT-6.1 Sol low; align setup templates, skill defaults, and the plan validator.
- Upgrade GPT-6 Sol implementation and synthesis roles to GPT-6.1 Sol high; replace Astra judgment, prose, hardest-task, and reflection roles with GPT-6 Sol high.
- Use GPT-6.1 Sol high, Opus 5.5 high, and Fable 5.1 high for Arena runners; keep Sol 6.1 high and Opus high in its cross-judge pool.
- Use GPT-6.1 Sol xhigh for Architect, Interrogate, and Codex Autoreview; keep Architect’s three candidates and one Sol reviewer alongside Opus xhigh in Interrogate.
Expand Down
2 changes: 1 addition & 1 deletion skills/autoreview/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -311,7 +311,7 @@ CLI flags and environment variables override these defaults. Pi does not get a b

Claude also supports `--fallback-model a,b` for availability-based fallback chains ([model-config](https://code.claude.com/docs/en/model-config)). Current Claude docs note that auth, billing, rate-limit, request-size, and transport errors do not trigger fallback, and the changelog documents interactive-session support in `v2.1.166`.

Autoreview defaults to Astra at `medium` reasoning. Explicit model requests remain supported, but Codex review never switches models automatically.
Autoreview defaults to `gpt-6.1-sol` at `xhigh` reasoning. Explicit model requests remain supported, but Codex review never switches models automatically.

Examples matching current `main` behavior:

Expand Down
2 changes: 1 addition & 1 deletion skills/how/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ When in doubt, take the simple path.

Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Spawn all explorers concurrently:

- `model`: your configured `how explorer` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6-luna` at `high` reasoning)
- `model`: your configured `how explorer` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `low` reasoning)

Each explorer gets the prompt in `references/explorer-prompt.md` with its angle filled in. Then go to Step 3.

Expand Down
2 changes: 1 addition & 1 deletion skills/no-comments/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ Use the caller's files or diff. Otherwise use the current diff against the base

## Steps

1. Spawn a Codex collaboration agent with the Comment Sicko instructions from `references/comment-sicko.md`, using your configured `Comment Sicko` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6-luna` at `high` reasoning). Pass the scope. Do not restate its rules.
1. Spawn a Codex collaboration agent with the Comment Sicko instructions from `references/comment-sicko.md`, using your configured `Comment Sicko` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `low` reasoning). Pass the scope. Do not restate its rules.
2. Inspect its report and diff. Reject application-code edits, scope escapes, exception-protected deletions, misstated `MUST KILL` reasons, and flags that treat kept intentional code as guilty. Reshape flags on our-code surprises stay actionable. Do not restore those comments. A keep survives only with proof it is about something we cannot change. Audit missed scoped lint and TypeScript suppressions. Correctness or safety suppressions stay actionable `MUST KILL`s. Restore deletions only with exact exceptions and scoped proof. Before accepting thin `IMPORTANT` or `do not remove` kills or keeps, run `$how` or `$why` on their symbol. If a kill is ambiguous, do not restore. If a keep is refuted or still ambiguous, delete it. Revert and rerun one rejected report with the failure named. Reject a second, report it open, and fail `$no-comments`.
3. Fix trivial accepted flags directly by deleting a dead path, dropping a parameter, or using the real API. If any fix needs a shape, run `$architect` once for the accepted set and surrounding code. Stop at the sketch. Architect shapes. Step 4 implements.
4. Implement the smallest root-cause fix in scope. Remove every named workaround. If the root cause is out of scope, land the smallest in-scope fix and report the rest open. The **principle-fix-root-causes** and **principle-redesign-from-first-principles** skills guide intent only. Neither authorizes widening the fence nor fixing instances outside it. Never bolt on symptom guards.
Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,7 @@ Read the leaf skill in full for any principle you apply. Each entry names when i

**Use the `poteto-agent` instructions in `references/poteto-agent.md` for any subagent you spawn inside a playbook step** (code-writing delegates, ad-hoc helpers). `$poteto-mode` and `poteto-agent` route through the same wrapper. Routed workflow skills (`how`, `why`, `interrogate`, `reflect`, `swarm`) set their own subagent instructions for diverse-model review. Respect what the skill prescribes, don't override to `poteto-agent`.

**Defaults for every spawn.** `fork_turns: "none"` (Codex forks the whole parent context by default), concurrent when independent, file pointers not inlined context, explicit model per role (configurable via `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md`. Defaults `gpt-6.1-sol` at `high` reasoning for code, `gpt-6-sol` at `high` reasoning for prose and judgment). Code delegates tier by difficulty. The hardest changes (cross-cutting design, gnarly concurrency, subtle algorithms) go to your strongest judgment model (`gpt-6-sol` at `high` reasoning), whether the task needs judgment on vague intent or is a precisely specified sequence of steps to execute to the letter. Trivial mechanical edits go to your fast code model. Per-role lines in `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` override these defaults and the model choices in the routed skills (`how`, `why`, `arena`, `swarm`, `architect`, `interrogate`, `reflect`). A role with no row keeps its default.
**Defaults for every spawn.** `fork_turns: "none"` (Codex forks the whole parent context by default), concurrent when independent, file pointers not inlined context, explicit model per role (configurable via `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md`. Defaults `gpt-6.1-sol` at `high` reasoning for code, `gpt-6-sol` at `high` reasoning for prose and judgment). Code delegates tier by difficulty. The hardest changes (cross-cutting design, gnarly concurrency, subtle algorithms) go to your strongest judgment model (`gpt-6.1-sol` at `high` reasoning), whether the task needs judgment on vague intent or is a precisely specified sequence of steps to execute to the letter. Trivial mechanical edits go to your fast code model. Per-role lines in `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` override these defaults and the model choices in the routed skills (`how`, `why`, `arena`, `swarm`, `architect`, `interrogate`, `reflect`). A role with no row keeps its default.

You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal.

Expand Down
4 changes: 2 additions & 2 deletions skills/poteto-mode/playbooks/multi-phase-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
6. Run `node ${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/scripts/check-plan.mjs <plan.md>` and fix every line it prints (the **encode-lessons-in-structure** principle skill).
7. Hand back. Post the plan path and the script's output, then stop. Execution starts on the operator's explicit go, under the execution playbook the plan names.

**Verification.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked (the **prove-it-works** principle skill). That sentence is the verification rule. Every verification block opens with it. The live block is mandatory. Ten lanes on `gpt-6-luna` at the PR head drive the real surface through its control skill, per the **swarm** skill. Each lane is one box with a concrete scenario, the screenshot it saves, and its pass predicate. One lane is the **Regression lane against trunk.** It runs the same load-bearing scenario on trunk and head. If trunk does not have the feature, the lane records that fact and gates the behavior the diff adds plus the end state the user waits for instead of inventing a trunk result. The perf gate is dual-sided. Trunk and head must both produce the named metric. If trunk lacks the feature, also isolate the work the diff adds and set an absolute budget for that work plus the end-to-end state the user waits for. Do not claim a ratio between unlike scenarios. The perf block names the metric, the interleaved probe, the trunk baseline measured first, and the rule with the number that fails. A PR that changes an interaction is review-gated. The operator reviews it in chat with screenshots and a video before merge. A PR that changes no interaction writes `**Review gate.** None. <PR id> is not review-gated.` and no boxes under it.
**Verification.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked (the **prove-it-works** principle skill). That sentence is the verification rule. Every verification block opens with it. The live block is mandatory. Ten lanes on `gpt-6.1-sol` with `low` reasoning at the PR head drive the real surface through its control skill, per the **swarm** skill. Each lane is one box with a concrete scenario, the screenshot it saves, and its pass predicate. One lane is the **Regression lane against trunk.** It runs the same load-bearing scenario on trunk and head. If trunk does not have the feature, the lane records that fact and gates the behavior the diff adds plus the end state the user waits for instead of inventing a trunk result. The perf gate is dual-sided. Trunk and head must both produce the named metric. If trunk lacks the feature, also isolate the work the diff adds and set an absolute budget for that work plus the end-to-end state the user waits for. Do not claim a ratio between unlike scenarios. The perf block names the metric, the interleaved probe, the trunk baseline measured first, and the rule with the number that fails. A PR that changes an interaction is review-gated. The operator reviews it in chat with screenshots and a video before merge. A PR that changes no interaction writes `**Review gate.** None. <PR id> is not review-gated.` and no boxes under it.

**Control skill.** Pick it by surface. Browser, Electron, and web UIs use `control-ui`. CLIs and TUIs use `control-cli`. Native mobile uses whatever simulator-driving skill the repo has. A PR that touches two surfaces gets lanes on both. A surface with no control skill is a risk in Appendix C, and its live block still names how each lane drives it.

Expand Down Expand Up @@ -98,7 +98,7 @@ Each live lane runs in its own isolated worktree at the PR head. Drive through `

- [ ] <Test file and the case it gains.> Run `<command>`.

**Verify, live.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked. Ten lanes on `gpt-6-luna` at the PR head, per the boot recipe.
**Verify, live.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked. Ten lanes on `gpt-6.1-sol` with `low` reasoning at the PR head, per the boot recipe.

- [ ] Lane 1. Regression lane against trunk. Run <the same load-bearing scenario> at trunk and head. If trunk lacks the feature, record that and gate <the behavior the diff adds plus the end state the user waits for>. Save `<slug>.png`. Pass when <predicate>.
- [ ] Lane 2. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
Expand Down
12 changes: 6 additions & 6 deletions skills/poteto-mode/references/models.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,16 +17,16 @@ One line per role. Delete a line to fall back to the skill default.
- perf-issue: `gpt-6.1-sol` at `high`
- hillclimb: `gpt-6.1-sol` at `high`
- judgment and prose: `gpt-6-sol` at `high`
- hardest tasks: `gpt-6-sol` at `high`
- how explorer: `gpt-6-luna` at `high`
- hardest tasks: `gpt-6.1-sol` at `high`
- how explorer: `gpt-6.1-sol` at `low`
- how explainer: `gpt-6.1-sol` at `high`
- why investigators: `gpt-6-luna` at `high`
- why investigators: `gpt-6.1-sol` at `low`
- why synthesizer: `gpt-6.1-sol` at `high`
- reflect tooling: `gpt-6.1-sol` at `high`
- reflect judgment, divergent, synthesizer: `gpt-6-sol` at `high`
- reflect judgment, divergent, synthesizer: `gpt-6.1-sol` at `high`
- arena runners: `gpt-6.1-sol` at `high`, `claude-opus-5-5` at `high`, `claude-fable-5-1` at `high`
- arena cross-judge pool: `gpt-6.1-sol` at `high`, `claude-opus-5-5` at `high`
- swarm workers: `gpt-6-luna` at `high`
- swarm workers: `gpt-6.1-sol` at `low`
- architect runners: `gpt-6.1-sol` at `xhigh`, `claude-opus-5-5` at `high`, `claude-fable-5-1` at `high`
- interrogate reviewers: `gpt-6.1-sol` at `xhigh`, `claude-opus-5-5` at `xhigh`
- Comment Sicko: `gpt-6-luna` at `high`
- Comment Sicko: `gpt-6.1-sol` at `low`
2 changes: 1 addition & 1 deletion skills/poteto-mode/scripts/check-plan.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import process from "node:process";

const RULE =
"Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked.";
const LANES = "Ten lanes on `gpt-6-luna` at the PR head";
const LANES = "Ten lanes on `gpt-6.1-sol` with `low` reasoning at the PR head";
const SUB_BLOCKS = [
"Depends on.",
"Files.",
Expand Down
6 changes: 3 additions & 3 deletions skills/reflect/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,15 +31,15 @@ One message, three agent calls, explicit `model:` and `reasoning_effort` on each

| Lens | `model` | Prompt template |
|---|---|---|
| Judgment | your configured `reflect judgment` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6-sol` at `high` reasoning) | `references/judgment-reviewer.md` |
| Judgment | your configured `reflect judgment` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `high` reasoning) | `references/judgment-reviewer.md` |
| Tooling | your configured `reflect tooling` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `high` reasoning) | `references/tooling-reviewer.md` |
| Divergent | your configured `reflect judgment` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6-sol` at `high` reasoning) | `references/divergent-reviewer.md` |
| Divergent | your configured `reflect judgment` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `high` reasoning) | `references/divergent-reviewer.md` |

Pass each template verbatim, substituting the transcript path or digest where marked. Reviewers return findings in the agent response body.

### 3. Synthesize

One agent call, using your configured `reflect judgment` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6-sol` at `high` reasoning). The synthesizer's quality check includes spot-verifying citations, which can require MCP access. Readonly strips MCPs. Use `references/synthesizer.md` verbatim, with each reviewer's full output inlined where marked. The synthesizer returns a structured Accepted / Rejected / Backlog list.
One agent call, using your configured `reflect judgment` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `high` reasoning). The synthesizer's quality check includes spot-verifying citations, which can require MCP access. Readonly strips MCPs. Use `references/synthesizer.md` verbatim, with each reviewer's full output inlined where marked. The synthesizer returns a structured Accepted / Rejected / Backlog list.

### 4. Structural enforcement check

Expand Down
12 changes: 6 additions & 6 deletions skills/setup-pstack/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,19 +60,19 @@ One line per role. Delete a line to fall back to the skill default.
- perf-issue: `gpt-6.1-sol` at `high`
- hillclimb: `gpt-6.1-sol` at `high`
- judgment and prose: `gpt-6-sol` at `high`
- hardest tasks: `gpt-6-sol` at `high`
- how explorer: `gpt-6-luna` at `high`
- hardest tasks: `gpt-6.1-sol` at `high`
- how explorer: `gpt-6.1-sol` at `low`
- how explainer: `gpt-6.1-sol` at `high`
- why investigators: `gpt-6-luna` at `high`
- why investigators: `gpt-6.1-sol` at `low`
- why synthesizer: `gpt-6.1-sol` at `high`
- reflect tooling: `gpt-6.1-sol` at `high`
- reflect judgment, divergent, synthesizer: `gpt-6-sol` at `high`
- reflect judgment, divergent, synthesizer: `gpt-6.1-sol` at `high`
- arena runners: `gpt-6.1-sol` at `high`, `claude-opus-5-5` at `high`, `claude-fable-5-1` at `high`
- arena cross-judge pool: `gpt-6.1-sol` at `high`, `claude-opus-5-5` at `high`
- swarm workers: `gpt-6-luna` at `high`
- swarm workers: `gpt-6.1-sol` at `low`
- architect runners: `gpt-6.1-sol` at `xhigh`, `claude-opus-5-5` at `high`, `claude-fable-5-1` at `high`
- interrogate reviewers: `gpt-6.1-sol` at `xhigh`, `claude-opus-5-5` at `xhigh`
- Comment Sicko: `gpt-6-luna` at `high`
- Comment Sicko: `gpt-6.1-sol` at `low`
```

### 6. Confirm
Expand Down
2 changes: 1 addition & 1 deletion skills/swarm/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ Call `update_plan` with one entry per phase before launching anything.
1. State the done predicate and the artifact or report the swarm must return.
2. Choose the shape. Partition into slices, race N workers on identical briefs, or mix both. For a race or mixed shape, declare `first pass`, `rank all`, or `best-of` before spawning.
3. Set N from the user or derive it from the shape. N is total workers, not the Codex concurrency limit.
4. Pick the worker model from `swarm workers` in `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` when present. Otherwise use `gpt-6-luna` at `high` reasoning. For a model race, name each arm's model up front.
4. Pick the worker model from `swarm workers` in `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` when present. Otherwise use `gpt-6.1-sol` at `low` reasoning. For a model race, name each arm's model up front.
5. Give each worker its own writable output when it writes.

## Phase B: Fan out
Expand Down
2 changes: 1 addition & 1 deletion skills/why/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ Aim for a complete **coverage map**, not a minimal one. Document the null, don't
Launch all matching investigators concurrently so they run concurrently. Don't ask one agent to cover multiple MCPs.

Subagent config (each):
- `model`: your configured `why investigators` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6-luna` at `high` reasoning)
- `model`: your configured `why investigators` model from `${CODEX_HOME:-$HOME/.codex}/skills/poteto-mode/references/models.md` (default `gpt-6.1-sol` at `low` reasoning)
- Codex has no read-only switch for collaboration agents. Instruct every agent not to edit files, change repository state, or perform external writes.

Each investigator gets:
Expand Down
Loading