Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,9 @@ summary: Timeline of guardrail helper changes mirrored from Sweetistics and rela

# Changelog

## 2026-09-06: Astra Reasoning Defaults
- Defaulted routine Astra work in PStack skills and Poteto playbooks to medium reasoning; retained high for reviews and reasoning-heavy tasks, including Autoreview, with fast mode still forbidden.

## 2026-09-04: Autoreview Astra Default
- Changed the autoreview skill and helper default to `gpt-6-astra` at high reasoning, using Sol as the access-only fallback and preserving explicit model selection; skill instructions require standard mode.

Expand Down
6 changes: 3 additions & 3 deletions skills/arena/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,12 +24,12 @@ The N candidates will receive the same prompt, so the prompt is the contract. Ge

1. State the artifact each candidate is producing.
2. Derive the rubric. State what success looks like for *this* task, then turn it into 3-6 concrete gradeable criteria. Concrete: `Adds a --dry-run flag that skips writes`. Vague: `code is correct`. The rubric is the picker's tool in Phase D; candidates only see the task.
3. Pick the runners. For judgment-sensitive work, cycle through `gpt-6-astra`, `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna`, each at `high` reasoning; repeat the pool when `N > 4`. For generation-bound work, use `gpt-5.6-luna` at `high` reasoning for every runner. Use standard mode; never use fast mode or substitute another model or reasoning level. Spawn more when the arena covers multiple design directions.
4. Assign output paths. Each candidate writes to its own location (an authorized Codex worktree where possible, otherwise `/tmp/arena-<slug>/candidate-<n>/`). N candidates writing to the same path is shared mutable state and fails the the **separate-before-serializing-shared-state** principle skill test.
3. Pick the runners. For judgment-sensitive work, cycle through `gpt-6-astra`, `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna`, using Astra at `medium` reasoning by default and the other models at `high`; repeat the pool when `N > 4`. Use Astra `high` for reviews or reasoning-heavy candidates involving architectural or algorithmic tradeoffs, ambiguous root causes, or cross-system constraints. For generation-bound work, use `gpt-5.6-luna` at `high` reasoning for every runner. Use standard mode; never use fast mode or substitute another model or reasoning level. Spawn more when the arena covers multiple design directions.
4. Assign output paths. Each candidate writes to its own location (an authorized Codex worktree where possible, otherwise `/tmp/arena-<slug>/candidate-<n>/`). N candidates writing to the same path is shared mutable state and fails the **separate-before-serializing-shared-state** principle skill test.

## Phase B: Fan out

Spawn all N Codex collaboration agents in one concurrent batch when capacity allows and otherwise in waves, each with its Phase A model, `reasoning_effort: "high"`, `fork_turns: "none"`, the task, the path to the shared grounding, its own output path, and instructions to produce both the artifact and a short rationale.
Spawn all N Codex collaboration agents in one concurrent batch when capacity allows and otherwise in waves, each with its Phase A model and explicit `reasoning_effort`, `fork_turns: "none"`, the task, the path to the shared grounding, its own output path, and instructions to produce both the artifact and a short rationale.

The rationale is mandatory. Without it, the parent cannot tell whether a candidate's structure is principled or accidental, which makes Phase E grafting unreliable. Each rationale names the alternatives the candidate considered and what it rejected.

Expand Down
2 changes: 1 addition & 1 deletion skills/figure-it-out/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Present the framing and tradeoffs before committing to a long run. Reversible wo
Decompose into atomic, independently-landable units. Sequence riskiest-unknown-first so option value stays high. Scaffold and verification come before features (the **foundational-thinking** principle skill).

- Build the verification harness before the work, with the baseline captured from the pre-change state, so the check reads as "old value vs new value".
- For one-way-door design decisions, run the **architect** skill (it runs **arena**) with isolated, opinionated candidates and a separately prompted judge. Follow **arena**'s model pool and judge contract, all at `high` reasoning and never fast mode. Skip it for mechanical work whose shape is already concrete. A second arena over a settled design is over-engineering (the **laziness-protocol** principle skill).
- For one-way-door design decisions, run the **architect** skill (it runs **arena**) with isolated, opinionated candidates and a separately prompted judge. Follow **arena**'s model and effort contracts, using Astra `high` for these consequential design decisions. Never use fast mode. Skip it for mechanical work whose shape is already concrete. A second arena over a settled design is over-engineering (the **laziness-protocol** principle skill).
- Decide what fans out. Parallelize only across genuine seams, and give each worker its own worktree or branch (the **separate-before-serializing-shared-state** principle skill). Don't over-fan.
- Write the designed phase list down. That list is what the human reviews.

Expand Down
8 changes: 4 additions & 4 deletions skills/how/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,11 +14,11 @@ Two modes:

## Codex Agent Contracts

Every agent uses `reasoning_effort: "high"` and `fork_turns: "none"` in standard mode.
Every agent uses `fork_turns: "none"` in standard mode. Set `reasoning_effort` explicitly by role below.

- Explorers use `model: "gpt-5.6-luna"`.
- Explainers and synthesizers use `model: "gpt-6-astra"`.
- Architectural critics use `gpt-6-astra`, `gpt-5.6-sol`, and `gpt-5.6-terra`, one critic per model.
- Explorers use `model: "gpt-5.6-luna"` at `high` reasoning.
- Explainers and synthesizers use `model: "gpt-6-astra"` at `medium` reasoning. Use `high` when explaining complex cross-system behavior or reconciling conflicting evidence requires substantial reasoning.
- Architectural critics use `gpt-6-astra`, `gpt-5.6-sol`, and `gpt-5.6-terra`, one critic per model, all at `high` reasoning because these are reviews.

Never use fast mode. If a required configuration is unavailable, stop and report the blocker rather than substituting another model or reasoning level.

Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ Read the leaf skill in full for any principle you apply. Each entry names when i

**Use the `poteto-agent` instructions in `references/poteto-agent.md` for any subagent you spawn inside a playbook step** (code-writing delegates, ad-hoc helpers). `$poteto-mode` and `poteto-agent` route through the same wrapper. Routed workflow skills (`how`, `why`, `interrogate`, `reflect`, `swarm`) set their own subagent instructions; respect what the skill prescribes, don't override them with `poteto-agent`.

**Model routing.** Routed workflow skills own their model contracts. For subagents spawned directly inside a playbook step, use `model: "gpt-6-astra"` for implementation, synthesis, and judgment; `model: "gpt-5.6-luna"` for read-only discovery and procedural verification; and `model: "gpt-5.6-sol"` for an independent review of Astra work. If Sol performed the work being reviewed, use `model: "gpt-5.6-terra"` for that review. Comment Sicko follows the **no-comments** skill and uses Luna. Every spawn uses `reasoning_effort: "high"`, `fork_turns: "none"`, and standard mode. Never use fast mode. Never substitute another model or reasoning level. If the required configuration is unavailable, stop and report the blocker. Run concurrently when the work is independent, with file pointers rather than inlined context. Codex collaboration agents share the checkout, so concurrent writers need disjoint file ownership; serialize shared writes unless the user explicitly authorized isolated worktree tasks.
**Model routing.** Routed workflow skills own their model contracts. For subagents spawned directly inside a playbook step, use `model: "gpt-6-astra"` for implementation, synthesis, and judgment; `model: "gpt-5.6-luna"` for read-only discovery and procedural verification; and `model: "gpt-5.6-sol"` for an independent review of Astra work. If Sol performed the work being reviewed, use `model: "gpt-5.6-terra"` for that review. Comment Sicko follows the **no-comments** skill and uses Luna. For Astra, use `reasoning_effort: "medium"` by default for implementation, explanation, and synthesis. Use `high` for reviews and reasoning-heavy work: architectural or algorithmic tradeoffs, ambiguous root-cause analysis, concurrency or performance diagnosis, or decisions spanning several systems. Task size or duration alone does not justify high. Assign the effort explicitly in each spawn. Sol, Terra, and Luna retain `high`. Every spawn uses `fork_turns: "none"` and standard mode. Never use fast mode. Never substitute another model or reasoning level. If the required configuration is unavailable, stop and report the blocker. Run concurrently when the work is independent, with file pointers rather than inlined context. Codex collaboration agents share the checkout, so concurrent writers need disjoint file ownership; serialize shared writes unless the user explicitly authorized isolated worktree tasks.

You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt in an independent fresh context. Agreement is high-signal.

Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/playbooks/bug-fix.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspender

1. Reproduce it yourself on the matching surface via the verification skill or tool (Non-negotiables). Don't hand the repro to the user. A debug or instrumentation protocol that says to ask the user does not override this; you drive the instrumented runtime. Ask the user only with a stated, specific reason the verification surface cannot reach the target, and only after driving it as far as it goes. Won't reproduce directly, force it: synthesize the trigger, tighten conditions, or instrument until it fires. A bug you can't reproduce, you can't prove fixed.
2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Codex Goal mode (`/goal`). Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out; a design grounded on a plausible-but-unconfirmed cause can be unanimously wrong while the real cause sits one subsystem over.
3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation to a subagent using `gpt-6-astra` at high reasoning with a specific scope; review the diff.
3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation to a subagent using `gpt-6-astra` at medium reasoning, or high for reasoning-heavy work under **Model routing**, with a specific scope; review the diff.
4. Verify on the same surface; the original repro now passes. "Inconclusive" or wrong-surface is not a pass; flag it. Unit tests show branch behavior, not bug absence.
5. When commits are authorized, stage them so the failing repro lands before the fix in git history; the diff tells the story. Otherwise keep the failing repro and fix as local changes in that order. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path; skip it when the test would be expensive, integration-heavy, or unclear.
This is the canonical **sequence-verifiable-units** principle skill, the failing test first and the fix on top.
Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/playbooks/eval.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ Evals test how a change affects agent behavior before promoting it: a new skill
1. **Frame.** State what variant is under test and what behavior counts as success. Write the rubric (3-6 concrete criteria) for the judge only. Hold it back from candidates.
2. **Set up sanitized environments.** Per-candidate working dir with the variant in place. Plant any context an organic task would have: a project skeleton, the skills the candidate would naturally read.
3. **Author one organic prompt.** What a user would type. No leakage of what's being measured.
4. **Spawn N parallel candidates** in independent fresh contexts per the **arena** skill's Phase B. Use the Phase A model pool declared by **arena**, at high reasoning. Each works in its own sanitized dir; same prompt to each.
4. **Spawn N parallel candidates** in independent fresh contexts per the **arena** skill's Phase B. Use the Phase A model pool and reasoning efforts declared by **arena**. Each works in its own sanitized dir; same prompt to each.
5. **Spawn one blinded judge** in another independent fresh context per the **arena** skill's Phase C. The judge uses `gpt-6-astra` at high reasoning and sees outputs by sanitized label and the rubric, never a run label.
6. **Verify the chain from transcripts, not self-report.** Read each candidate's exact current-thread rollout under `${CODEX_HOME:-$HOME/.codex}/sessions`, resolved from `CODEX_THREAD_ID`. Do not read unrelated sessions. Look at which files each candidate actually opened. Citing a principle is not reading its leaf skill, and reading it is not applying it. Grade chain-following from the files it really read plus the shape of the code, never from the candidate's own claims.
7. **Read every candidate output yourself** end to end. Compare to the judge's verdict. Disagreement means a model is biased or the rubric is ambiguous. Synthesize.
Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/playbooks/feature.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
- **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize.
- **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants.
- **Smallest safe decomposition.** If one worker is best, name why.
4. Delegate code-writing to a subagent using `gpt-6-astra` at high reasoning with a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain** — a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic — and success criteria); review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). You can spawn a subagent even though you are one; "the app is small" and "a subagent cannot spawn one" are both wrong. A subagent forbidden to spawn satisfies this by owning the diff directly with the same review separation; no "standing by" reply that waits on a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally when commits are authorized.
4. Delegate code-writing to a subagent using `gpt-6-astra` at medium reasoning, or high for reasoning-heavy work under **Model routing**, with a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain** — a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic — and success criteria); review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). You can spawn a subagent even though you are one; "the app is small" and "a subagent cannot spawn one" are both wrong. A subagent forbidden to spawn satisfies this by owning the diff directly with the same review separation; no "standing by" reply that waits on a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally when commits are authorized.
5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass; flag it.
6. Rebase into small, ordered commits; stack follow-ups.
Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next.
Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/playbooks/hillclimb.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Core discipline: one change, one measurement, keep or revert. Never stack untest
3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. This is the run's memory. Read it before each attempt so the search accumulates instead of circling. Keep it out of the tree (gitignored) so it survives reverts.
4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something".
5. Loop, one hypothesis per iteration:
- Hand the change to a subagent using `gpt-6-astra` at high reasoning with a tight scope; supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel subagents, each with disjoint file ownership; use isolated worktree tasks only when the user authorized them (the **separate-before-serializing-shared-state** principle skill).
- Hand the change to a subagent using `gpt-6-astra` at medium reasoning, or high for reasoning-heavy work under **Model routing**, with a tight scope; supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel subagents, each with disjoint file ownership; use isolated worktree tasks only when the user authorized them (the **separate-before-serializing-shared-state** principle skill).
- Measure before and after with the frozen harness, and run the regression gate.
- Accept only when the metric moves past noise and the gate stays green. Otherwise revert the change in full; a tweak that "might help" does not ride along.
- When commits are authorized, make one commit per accepted fix, staging only the files you changed (`git add <files>`, never `-A`). Log the row either way, kept or reverted.
Expand Down
2 changes: 1 addition & 1 deletion skills/poteto-mode/playbooks/multi-phase-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ Tests alone are not sufficient verification. A PR is verified only when its unit

### Verdict and merge, for every PR

- [ ] At the merge-ready head SHA, run the swarm per `${CODEX_HOME:-$HOME/.codex}/skills/swarm/SKILL.md`. The ten live lanes from the PR's **Verify, live** block use `gpt-5.6-luna` at high reasoning. One gates lane, the perf lane from its **Verify, perf** block, and one audit lane that reads the diff and the receipts and distrusts the PR body use `gpt-6-astra` at high reasoning.
- [ ] At the merge-ready head SHA, run the swarm per `${CODEX_HOME:-$HOME/.codex}/skills/swarm/SKILL.md`. The ten live lanes from the PR's **Verify, live** block use `gpt-5.6-luna` at high reasoning. The gates lane and the perf lane from its **Verify, perf** block use `gpt-6-astra` at medium reasoning; use high if they require reasoning-heavy diagnosis under **Model routing**. One audit lane that reads the diff and the receipts and distrusts the PR body uses `gpt-6-astra` at high reasoning.
- [ ] Clean only when every lane is `PASS`. Findings go back to the owner. A new head gets a fresh swarm and a fresh verdict.
- [ ] <The merge or append rule from the execution playbook, with the patch-id rule from `playbooks/shipping.md`.>

Expand Down
Loading
Loading