diff --git a/CHANGES.md b/CHANGES.md index 9e791df2..59b3613a 100644 --- a/CHANGES.md +++ b/CHANGES.md @@ -1,5 +1,13 @@ # CHANGES — applied substitutions +## Unreleased: project model sheets and GPT-6.1 Sol + +- `setup-pstack` asks on every run whether to configure the global sheet or a project sheet. A project sheet lives at `/.claude/pstack-models.md` (Claude Code) or `/.codex/pstack-models.md` (Codex), starts from the global assignments, and is listed in `.git/info/exclude` so it never reaches a commit. It needs no CLAUDE.md include and no AGENTS.md block. +- `provider-dispatch.md` gains a "Sheet scope" section: the parent reads the project path once before dispatch, an existing project sheet replaces the global sheet whole, and a project without one uses the global sheet. The two are never merged per role. +- Add the `sol-6.1` stock family, `codex:gpt-6.1-sol@high`, which Codex CLI 0.160.0 lists as its latest coding model. It takes the first-run `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb` roles. `sol-6` stays selectable; existing sheets keep their rows until setup reassigns them. +- `architect runners` gets its own first-run default, `codex:gpt-6-astra@high, claude:fable@max`, from a new "Default architect panel" line. The other three panel roles keep the four-lane default. +- The matrix test and the static invariant check cover the new row, the architect line, and the scope contract. + ## Unreleased: pstack-flex becomes its own distribution - The marketplace is now `pstack-flex` (was `open-pstack`) in both the Claude Code and Codex marketplace files. Install with `pstack@pstack-flex`. The plugin keeps the name `pstack`, so skill names such as `pstack:poteto-mode` are unchanged. Manifests, package names, and docs point at `thisguymartin/pstack-flex`; attribution to open-pstack, pstack-claude, and Cursor pstack stays in README and NOTICE. diff --git a/README.md b/README.md index 04e937ac..0bafdd48 100644 --- a/README.md +++ b/README.md @@ -39,7 +39,8 @@ Every pstack role (who writes code, who explores, who sits on a review panel) ma | `fable` | `claude:fable@max` | Claude Code login | judgment, prose, explanation, hardest tasks, panels | | `opus` | `claude:opus@max` | Claude Code login | panels | | `astra` | `codex:gpt-6-astra@high` | Codex (ChatGPT) login | panels | -| `sol-6` | `codex:gpt-6-sol@high` | Codex (ChatGPT) login | feature, refactoring, bug-fix, perf-issue, hillclimb | +| `sol-6.1` | `codex:gpt-6.1-sol@high` | Codex (ChatGPT) login | feature, refactoring, bug-fix, perf-issue, hillclimb | +| `sol-6` | `codex:gpt-6-sol@high` | Codex (ChatGPT) login | none; selectable | | `luna` | `codex:gpt-6-luna@high` | Codex (ChatGPT) login | how explorer, swarm workers | | `sol` | `codex:gpt-5.6-sol@max` | Codex (ChatGPT) login | none; selectable | | `grok` | `grok:grok-4.7@xhigh` | Grok CLI login | panels | @@ -48,7 +49,7 @@ Every pstack role (who writes code, who explores, who sits on a review panel) ma | `minimax` | `minimax:MiniMax-M3@high` | `MINIMAX_API_KEY` | none; selectable | | `minimax-preview` | `minimax:MiniMax-M3.1-Flash-Preview@high` | `MINIMAX_API_KEY` (Token Plan) | none; selectable | -The default review panel is `claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.7@xhigh, claude:opus@max`: four lanes across three providers. Any family can take any role. Panels must span at least two providers, and two models from one provider count as one, because the adversarial signal comes from model diversity. +The default review panel is `claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.7@xhigh, claude:opus@max`: four lanes across three providers. Architect sketches default to `codex:gpt-6-astra@high, claude:fable@max`. Any family can take any role. Panels must span at least two providers, and two models from one provider count as one, because the adversarial signal comes from model diversity. The DeepSeek and MiniMax lanes run the stock `claude` binary against the lab's Anthropic-compatible endpoint with that lab's key, in an isolated config directory, with inherited Anthropic routing stripped. A lane refuses to start if it finds a claude.ai login in that directory, so a subscription credential can never reach a third-party endpoint. Their receipts keep real token usage but set `costUsd` to null (Claude Code prices at Anthropic rates); the price table is in [docs/LANES.md](docs/LANES.md). Anthropic does not support pointing Claude Code at non-Anthropic endpoints; use synthetic data for gateway testing and keep keys in your local environment. @@ -120,7 +121,7 @@ Use pstack:setup-pstack to configure pstack. Setup is assignment-first. It shows the role map, asks which roles to change, asks one effort per assigned family, probes only those families with a real one-turn run, and writes nothing until every probe passes and you confirm. A fresh run proposes the defaults in the table above. An existing sheet keeps its assignments until you change a named role. -The sheet lives at `~/.claude/pstack-models.md` (Claude Code) or `~/.codex/pstack-models.md` (Codex). It is global, not per project. Change it by rerunning setup rather than editing it by hand, so every choice is probed before it is saved. +Setup first asks which scope to configure. The global sheet lives at `~/.claude/pstack-models.md` (Claude Code) or `~/.codex/pstack-models.md` (Codex). A project sheet lives at `.claude/pstack-models.md` or `.codex/pstack-models.md` in the repository root. It replaces the global sheet for that project, stays out of git through `.git/info/exclude`, and starts from your global assignments. A project without one uses the global sheet; delete the project sheet to go back. Change either sheet by rerunning setup rather than editing it by hand, so every choice is probed before it is saved. A model sheet from an earlier release keeps its panel. To take the new defaults, delete those role lines and run setup again; setup fills missing roles from the defaults. A `grok:grok-4.6` entry keeps running until the next setup run asks you to replace it. diff --git a/UPSTREAM-FLEX.md b/UPSTREAM-FLEX.md index b13e5a8c..a1153803 100644 --- a/UPSTREAM-FLEX.md +++ b/UPSTREAM-FLEX.md @@ -32,9 +32,9 @@ All flex changes are additive and live in port-owned files so upstream merges st - `plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts` and `flex-providers.test.ts` (new) - Gateway-provider hooks in `runner/{types,commands,run,parse-output,cli}.ts` and their tests -- The three GPT-6 rows in the stock model matrix, the "Default panel" section, the "Flex model matrix" section, and the route-table columns in `references/provider-dispatch.md` +- The four GPT-6 rows in the stock model matrix, the "Default panel", "Default architect panel", and "Sheet scope" sections, the "Flex model matrix" section, and the route-table columns in `references/provider-dispatch.md` - The first-run sheet in `skills/setup-pstack/SKILL.md` and the default descriptors named in `arena`, `architect`, `interrogate`, `how`, `swarm`, and the `feature`, `refactoring`, `bug-fix`, `perf-issue`, and `hillclimb` playbooks -- The assignment-first restructure of `skills/setup-pstack/SKILL.md` +- The assignment-first restructure of `skills/setup-pstack/SKILL.md`, and its project-or-global scope question, project sheet paths, and `.git/info/exclude` write - `docs/LANES.md`, this file, the README fork section, and the NOTICE/LICENSE/CHANGES additions - `plugins/pstack/hooks/session-start-context.md`, which the fork rewrote from open-pstack's auto-fire mandate into an opt-in gate, and the docs lines that describe it - The `intake` and `diff-behavior` skills: `plugins/pstack/skills/intake/` and `plugins/pstack/skills/diff-behavior/`. Upstream has no equivalent, so they never conflict. diff --git a/UPSTREAM.md b/UPSTREAM.md index c9c5704f..d6c051f9 100644 --- a/UPSTREAM.md +++ b/UPSTREAM.md @@ -18,7 +18,7 @@ The table above is the current Cursor sync point. Open Pstack 1.5.0 imports this - Commits `799151d` and `6fecddb` add and relocate `make-bot-ui`. It depends on Cursor routines, webhook events, and UI primitives that Claude Code and Codex do not share. - Four `disable-model-invocation: true` lines from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`. Poteto-mode invokes those skills by name, and the flag blocks that route on Claude Code. -- The default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` from `23a56e2`, `889ec4b`, and `70b2dc8` are not applied. Those frequent code-writing roles stay on a Codex model (`codex:gpt-6-sol@high` in pstack-flex). +- The default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` from `23a56e2`, `889ec4b`, and `70b2dc8` are not applied. Those frequent code-writing roles stay on a Codex model (`codex:gpt-6.1-sol@high` in pstack-flex). - `5bf2b15`'s setup budget question, its `# budget` line, and its step down to a lower detected effort are not applied. Setup already asks one requested effort per assigned family, and the step-down would silently lower a requested effort. - `12d587d`'s rule that reruns a rejected configured entry on its family default or the closest valid slug is not applied. An unavailable model stays a named dropout per `provider-dispatch.md`. - The expected-runtime column in `70b2dc8`'s `children.tsv` and its expected-runtime stuck test are not applied. A lane is stuck only on affirmative failure evidence. diff --git a/docs/LANES.md b/docs/LANES.md index 36156629..c1528dff 100644 --- a/docs/LANES.md +++ b/docs/LANES.md @@ -8,22 +8,23 @@ Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against | Kind | Lanes | Auth | Billing | Route | | --- | --- | --- | --- | --- | -| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna`, `codex:gpt-5.6-sol`, `grok:grok-4.7` | each CLI's own login | that CLI's plan | native or external per the route table | +| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-6-astra`, `codex:gpt-6.1-sol`, `codex:gpt-6-sol`, `codex:gpt-6-luna`, `codex:gpt-5.6-sol`, `grok:grok-4.7` | each CLI's own login | that CLI's plan | native or external per the route table | | Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner | A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide). ## GPT-6 Codex families -Three stock Codex families carry the first-run defaults. They use the same ChatGPT login as `codex:gpt-5.6-sol`: +Three of the four stock GPT-6 Codex families carry the first-run defaults. They use the same ChatGPT login as `codex:gpt-5.6-sol`: | Family | Descriptor at default requested effort | First-run roles | Codex's own description | | --- | --- | --- | --- | -| astra | `codex:gpt-6-astra@high` | every panel (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) | Frontier tier for the most demanding work | -| sol-6 | `codex:gpt-6-sol@high` | `feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb` | Coding and everyday workhorse | +| astra | `codex:gpt-6-astra@high` | every panel (`arena runners`, `arena cross-judge pool`, `interrogate reviewers`) and `architect runners`, where it pairs with Fable | Frontier tier for the most demanding work | +| sol-6.1 | `codex:gpt-6.1-sol@high` | `feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb` | Latest workhorse model for coding and everyday work | +| sol-6 | `codex:gpt-6-sol@high` | none; selectable | Previous generation workhorse model | | luna | `codex:gpt-6-luna@high` | `how explorer`, `swarm workers` | Fast, low-cost tier for easier tasks | -A fresh `/setup-pstack` run proposes these. An existing sheet keeps its assignments until you change a named role in setup; `codex:gpt-5.6-sol` remains a selectable family for that. The `sol-6` family is separate from the `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus GPT-6 Sol does not satisfy the two-provider rule; the default panel spans Claude, Codex, and Grok. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27. +A fresh `/setup-pstack` run proposes these. An existing sheet keeps its assignments until you change a named role in setup; `codex:gpt-5.6-sol` remains a selectable family for that. The `sol-6.1`, `sol-6`, and `sol` families are separate, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus GPT-6 Sol does not satisfy the two-provider rule; the default panel spans Claude, Codex, and Grok. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra, GPT-6.1 Sol, and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.160.0 model list, checked 2026-10-04. ## Multiple models per provider diff --git a/docs/USAGE.md b/docs/USAGE.md index 069b85b9..5a691fa5 100644 --- a/docs/USAGE.md +++ b/docs/USAGE.md @@ -83,11 +83,11 @@ Daily flow: `pstack-keys -> claude -> /pstack:poteto-mode`. Alternatives, the th Setup is assignment-first: pick which roles run on which families, answer one effort question per **assigned** family, and only assigned families get probed. Unassigned families are skipped, not errors. Every probe is a real one-turn run — a failed probe writes nothing. Three configurations that make sense: -**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults. GPT-6 Sol writes code, Luna explores and verifies, and the panel spans three providers: +**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults. GPT-6.1 Sol writes code, Luna explores and verifies, and the panel spans three providers: ```text -feature, refactoring: codex:gpt-6-sol@high -bug-fix: codex:gpt-6-sol@high +feature, refactoring: codex:gpt-6.1-sol@high +bug-fix: codex:gpt-6.1-sol@high how explorer: codex:gpt-6-luna@high swarm workers: codex:gpt-6-luna@high arena runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.7@xhigh, claude:opus@max @@ -244,7 +244,7 @@ pstack's side is the lane journal. While `~/.pstack-flex/lanes/` exists, `pstack ## The GPT-6 Codex models -`astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), and `luna` (`codex:gpt-6-luna@high`) are stock families and the first-run defaults: Astra on every panel, GPT-6 Sol on the solo code-writing roles, Luna on exploration and swarm work. They need only your Codex login. Each gets its own effort question and live probe. See [GPT-6 Codex families](LANES.md#gpt-6-codex-families). +`astra` (`codex:gpt-6-astra@high`), `sol-6.1` (`codex:gpt-6.1-sol@high`), and `luna` (`codex:gpt-6-luna@high`) are stock families and the first-run defaults: Astra on every panel and, with Fable, on architect sketches; GPT-6.1 Sol on the solo code-writing roles; Luna on exploration and swarm work. They need only your Codex login. Each gets its own effort question and live probe. See [GPT-6 Codex families](LANES.md#gpt-6-codex-families). A sheet written before this release keeps its assignments. To move a role, run `/setup-pstack` and name it; every role you do not change keeps its descriptor, and `codex:gpt-5.6-sol` stays selectable. For example, this row keeps GPT-5.6 Sol on bug fixes while the rest of the sheet takes the new defaults: diff --git a/docs/reference.md b/docs/reference.md index 7483874d..b4abc367 100644 --- a/docs/reference.md +++ b/docs/reference.md @@ -86,7 +86,7 @@ The Codex build shares one `skills/` tree with the Claude Code build. Nothing is - **Tool and built-in mapping.** Claude tool names and built-in skills resolve through [`codex-tools.md`](../plugins/pstack/skills/poteto-mode/references/codex-tools.md). Model execution resolves separately through [`provider-dispatch.md`](../plugins/pstack/skills/poteto-mode/references/provider-dispatch.md), so Codex can keep Sol native while invoking Claude and Grok externally. - **Subagents.** The `Agent` tool maps to Codex `spawn_agent` / `wait_agent`, enabled by `multi_agent = true`. Parallel fan-out is multiple `spawn_agent` calls in one turn. If the native Codex lane is unavailable, record that lane as a dropout; external Claude and Grok lanes still run, and no provider is silently substituted. There is no `poteto-agent` subagent type on Codex; route ad-hoc subagents by dispatching a `spawn_agent` told to read `poteto-mode` first. - **Opt-in.** Codex runs the plugin's `hooks/` SessionStart hook and adds the opt-in gate to each session as a developer message (observed on Codex 0.157.1). Codex records trust for the hook in `~/.codex/config.toml` under `hooks.state`. Enter `pstack:poteto-mode` by name, or add a standing instruction to `~/.codex/AGENTS.md` if you want every non-trivial task routed into it. After a plugin update, run `codex plugin marketplace upgrade pstack-flex` so the installed copy carries the current gate. -- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per assigned family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-6 Astra high, Grok 4.7 xhigh, and Opus max. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. The GPT-6 Astra, Sol, and Luna Codex families are stock: GPT-6 Sol high carries `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high carries `how explorer` and `swarm workers`; Astra high sits on every panel. GPT-5.6 Sol remains a selectable family. In Codex, every Codex family uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; the Codex families and Grok use the runner. Children never detect the parent or reroute themselves. The solo code roles stay on a Codex model instead of upstream's Fable default because it costs less for these frequent delegated code roles. +- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per assigned family (`low`, `medium`, `high`, `xhigh`, `max`). It asks first whether to configure the global sheet or a private project sheet, which replaces the global one for that repository. The first-run panel is Fable max, GPT-6 Astra high, Grok 4.7 xhigh, and Opus max; architect sketches use GPT-6 Astra high and Fable max. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. The GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and Luna Codex families are stock: GPT-6.1 Sol high carries `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high carries `how explorer` and `swarm workers`; Astra high sits on every panel. GPT-6 Sol and GPT-5.6 Sol remain selectable families. In Codex, every Codex family uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; the Codex families and Grok use the runner. Children never detect the parent or reroute themselves. The solo code roles stay on a Codex model instead of upstream's Fable default because it costs less for these frequent delegated code roles. Verified upstream in fresh installed open-pstack Claude Code and Codex sessions (pstack-flex's own live-test record is in [LIVE-GATE.md](LIVE-GATE.md)): the user-facing skills are discovered and namespaced under `pstack`; both parents fan out the configured panel through the documented native/external route table, retain long-running handles without a default timeout, and cross-judge only after every candidate is terminal. The `principle-*` leaves remain available for `poteto-mode` to read by path. Claude honors their `user-invocable: false` metadata; Codex 0.149.0 does not ([open-pstack #8](https://github.com/ericlitman/open-pstack/issues/8)). @@ -201,7 +201,7 @@ The port is editorial, not mechanical. Anywhere upstream pstack assumed Cursor-s | Cursor's `/goal` (standing objective across turns) | The program objective written into the run's standing orders and restated in the todolist | | The Cursor agent store (path in the system prompt) | `~/.claude/orchestrate//`, which survives the session restarts a multi-day program expects | | Model rule `~/.cursor/rules/pstack-models.mdc` | Override sheet `~/.claude/pstack-models.md`, included from `CLAUDE.md` | -| Multi-model panels (arena, architect, interrogate) | Provider dispatch owns the default panel: `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.7@xhigh`, `claude:opus@max`. Same-provider lanes stay native; external lanes use the bundled runner. | +| Multi-model panels (arena, architect, interrogate) | Provider dispatch owns the default panel: `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.7@xhigh`, `claude:opus@max`, and the architect default: `codex:gpt-6-astra@high`, `claude:fable@max`. Same-provider lanes stay native; external lanes use the bundled runner. | ### Cross-vendor dispatch @@ -219,7 +219,7 @@ The earlier port collapsed panels to Claude-only models. The bundled runner rest - **`automations/benny/`** (upstream `0452e08`, the only pstack change between `e46364b` and v0.10.0) — a dormant Slack issue-triage and reproduce-and-fix automation pack built on Cursor's event-triggered automations. It registers no slash skills even upstream, so excluding it changes nothing about the ported plugin's behavior. Porting it would require Cursor's event-trigger runtime, Slack, and tracker plumbing that Open Pstack does not provide. - **`docs/guide/`** (upstream `02c03a9`, `0b7ef5b`, `424829e`) — the ten-chapter usage tutorial and its six screenshots (2.3 MB). It teaches pstack through Cursor's UI, sticky mode, and cloud agents, so a faithful port would be a rewrite rather than a sync, and none of it ships as skill content. Read it upstream at [cursor/plugins/pstack/docs/guide](https://github.com/cursor/plugins/tree/main/pstack/docs/guide); the concepts map through the substitution table above. - **`make-bot-ui`** (upstream `799151d`, relocated by `6fecddb`) uses Cursor routines, webhook events, hosted bot state, and Cursor UI primitives that have no shared Claude Code and Codex mapping. A provider-specific rewrite would be a separate feature, not an upstream sync. -- **Solo code defaults** (upstream `23a56e2` moved them to Fable; `889ec4b` and `70b2dc8` move them to Grok 4.7). pstack-flex keeps `bug-fix`, `perf-issue`, and `hillclimb` on a Codex model (`codex:gpt-6-sol@high`). +- **Solo code defaults** (upstream `23a56e2` moved them to Fable; `889ec4b` and `70b2dc8` move them to Grok 4.7). pstack-flex keeps `bug-fix`, `perf-issue`, and `hillclimb` on a Codex model (`codex:gpt-6.1-sol@high`). - **Sticky mode** (upstream `#144`) — Cursor-only `mode`/`icon`/`color`/`reminder` frontmatter with no Claude Code equivalent. The port's 0.9.5 SessionStart hook was the analog. pstack-flex turned that hook into an opt-in gate, so start `poteto-mode` by name. - **`is_background: true` on `poteto-agent`** (upstream `99559f2`) — Cursor names this key differently. Claude-native frontier definitions use `background: true`; ad-hoc `poteto-agent` calls remain background dispatches at the call site. - **`cursor-team-kit` beyond the seven imported skills** — the rest either duplicate Claude Code built-ins (`verify-this` → the `verify` skill and built-in verification discipline; `check-compiler-errors` → LSP diagnostics; `control-cli`/`control-ui` → `run`/`verify`, already the substitution targets) or overlap skills this port ships (`loop-on-ci`, `review-and-ship`, `weekly-review` vs `babysit`, `fix-ci`, `make-pr-easy-to-review`, `what-did-i-get-done`). `pr-review-canvas` is Cursor-UI-specific. diff --git a/plugins/pstack/skills/architect/SKILL.md b/plugins/pstack/skills/architect/SKILL.md index 35d199c3..39ec4009 100644 --- a/plugins/pstack/skills/architect/SKILL.md +++ b/plugins/pstack/skills/architect/SKILL.md @@ -31,7 +31,7 @@ Skip Phase A only when the work is genuinely greenfield with no surrounding syst Run the **arena** skill with the design-sketch task and the Phase A grounding artifacts. Pass `references/runner-prompt.md` as each runner's prompt. Each candidate produces a design package shaped per `references/rationale-template.md`. -Take the runners from `architect runners` in the current harness's pstack model sheet, in place of Arena's `arena runners`. If the sheet or that line is missing, use `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.7@xhigh`, `claude:opus@max`. +Take the runners from `architect runners` in the current harness's pstack model sheet, in place of Arena's `arena runners`. If the sheet or that line is missing, use `codex:gpt-6-astra@high`, `claude:fable@max`. Design it twice. Require at least two structurally distinct candidates before synthesis, even when the first looks sufficient. This is the **exhaust-the-design-space** principle skill made concrete. Whole-shape alternatives, not point fixes inside one shape. diff --git a/plugins/pstack/skills/poteto-mode/SKILL.md b/plugins/pstack/skills/poteto-mode/SKILL.md index 680f7215..8cd59fd4 100644 --- a/plugins/pstack/skills/poteto-mode/SKILL.md +++ b/plugins/pstack/skills/poteto-mode/SKILL.md @@ -89,7 +89,7 @@ Read the leaf skill in full for any principle you apply. Each entry names when i For `inherit-parent`, `auto`, or an unconfigured native ad-hoc helper, prefer `poteto-agent`. `/poteto-mode` and `poteto-agent` route through the same wrapper. A provider-qualified role instead follows provider dispatch: Claude's shipped frontier agent definitions select the model alias and requested effort, Codex passes both to `spawn_agent`, and external providers run through the deterministic launcher. Routed workflow skills set the task and access mode. Do not override their choices. -**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. First-run defaults use GPT-6 Sol high for feature/refactoring, bug fixes, performance work, and hillclimbing; GPT-6 Luna high for exploration and swarm work; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the Fable, GPT-6 Astra, Grok 4.7, Opus panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. Each code playbook's configured model comes from its line (`feature, refactoring`, `bug-fix`, `perf-issue`, or `hillclimb`), and the hardest changes read `hardest tasks`. Prose and judgment read `judgment and prose`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel. +**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. First-run defaults use GPT-6.1 Sol high for feature/refactoring, bug fixes, performance work, and hillclimbing; GPT-6 Luna high for exploration and swarm work; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; the Fable, GPT-6 Astra, Grok 4.7, Opus panel for model-diverse judgment; and GPT-6 Astra plus Fable for architect sketches. A project sheet replaces the global sheet per provider dispatch's sheet scope. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. Each code playbook's configured model comes from its line (`feature, refactoring`, `bug-fix`, `perf-issue`, or `hillclimb`), and the hardest changes read `hardest tasks`. Prose and judgment read `judgment and prose`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel. You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md index 60326923..8cb11dec 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md @@ -6,7 +6,7 @@ Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspender 1. Reproduce it yourself on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs) (Non-negotiables), even when a debug or instrumentation protocol says to ask the user to reproduce. Ask the user only with a stated, specific reason the control surface cannot reach the target, and only after driving it as far as it goes. If it won't reproduce directly, synthesize the trigger, tighten conditions, or instrument until it fires. 2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Claude Code's `loop` skill. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out. -3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope. +3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-6.1-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope. 4. Verify on the same surface. The original repro now passes. "Inconclusive" or wrong-surface is not a pass. Flag it. Unit tests show branch behavior, not bug absence. 5. Stage the commits so the failing repro lands before the fix in git history. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path. Skip it when the test would be expensive, integration-heavy, or unclear. This is the canonical **sequence-verifiable-units** principle skill, the failing test first and the fix on top. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/feature.md b/plugins/pstack/skills/poteto-mode/playbooks/feature.md index 42b1b390..ad94cdf0 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/feature.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/feature.md @@ -9,7 +9,7 @@ - **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize. - **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants. - **Smallest safe decomposition.** If one worker is best, name why. -4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. +4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `codex:gpt-6.1-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. 5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it. 6. Rebase into small, ordered commits. Stack follow-ups. Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md index bdde9697..92930d1a 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md @@ -9,7 +9,7 @@ Core discipline: one change, one measurement, keep or revert. Never stack untest 3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. Read it before each attempt. Keep it out of the tree (gitignored). 4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something". 5. Loop, one hypothesis per iteration: - - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-6-sol@high`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill). + - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-6.1-sol@high`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill). - Measure before and after with the frozen harness, and run the regression gate. - Accept only when the metric moves past noise and the gate stays green. Otherwise revert the change in full. A tweak that "might help" is not kept. - One commit per accepted fix, staging only the files you changed (`git add `, never `-A`). Log the row either way, kept or reverted. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md index 0f755d3e..0f8fb96b 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md @@ -13,7 +13,7 @@ - **Redundancy.** The wait hangs on one slow instance or attempt. Duplicate the work (replicas, hedged requests, speculative execution) and take the fastest result. The trace has to show the wait dominates and the system has headroom. - **Lazy evaluation.** Cost lands on results that are never used or not needed yet (eager init on the boot path, rendering offscreen items). Defer the work until first use. - **Scheduling.** The work must happen, but not during the interactive moment. Move it to where nobody is waiting: idle callbacks, a background warmup after boot, precompute before the user arrives, cleanup after the frame commits. The win is perceived latency, so measure the interactive path, not total work done. -3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-6-sol@high`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace. +3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-6.1-sol@high`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace. Apply the **sequence-verifiable-units** principle skill, verifying each attempt before trying the next. 4. Parse and compare the artifacts (JSON to sqlite, diff). "Inconclusive" or wrong-surface is not a pass. Flag it. 5. Cite the measurement in the PR. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md index 7cfb980b..b792bdc4 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md @@ -8,7 +8,7 @@ If the cleanup reveals a missing feature or a real bug, split it out and ship th 2. Name the structure the code is missing per **principle-model-the-domain**. Boring code stays when the shape is already clear and local. The reshape must delete branches or invalid states, not add indirection. 3. Name the target shape. State what the module layout, types, and call graph should be if built today (**principle-foundational-thinking**, **principle-redesign-from-first-principles**). If the target crosses a function boundary, run the **architect** skill for parallel design exploration of the shape before the move. 4. Subtract before you add. Delete dead code, collapse one-caller wrappers, drop redundant validators, and remove orphan references before introducing the new shape (**principle-subtract-before-you-add**). The smallest change that reaches the target shape ships (**principle-laziness-protocol**). A speculative cleanup that "might help" gets reverted. -5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). +5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `codex:gpt-6.1-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). 6. Prove behavior is unchanged on the real artifact, not "it compiles" (**principle-prove-it-works**). For larger reshapes, run an equivalence check: a script that diffs old-vs-new outputs, a recorded baseline replayed against the new code, or a smoke run on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs). 7. Confirm the change is worth keeping. The success measure is reduced reader load (**principle-minimize-reader-load**). If the diff does not lower reader load somewhere, revert it. 8. Rebase into small ordered commits. A subtraction commit, then the reshape, then any follow-on cleanup. Shape them with the **sequence-verifiable-units** principle skill, so each behavior-preserving slice stays green before the next. Run **Opening a PR**. diff --git a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md index af81c2ba..1392e029 100644 --- a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md +++ b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md @@ -15,6 +15,7 @@ pstack model choices are provider-qualified descriptors: | grok | grok-4.7-xhigh-fast | grok | grok-4.7 | xhigh | low medium high xhigh max | - | | opus | opus | claude | opus | max | low medium high xhigh max | opus | | astra | - | codex | gpt-6-astra | high | low medium high xhigh max | - | +| sol-6.1 | - | codex | gpt-6.1-sol | high | low medium high xhigh max | - | | sol-6 | - | codex | gpt-6-sol | high | low medium high xhigh max | - | | luna | - | codex | gpt-6-luna | high | low medium high xhigh max | - | @@ -22,15 +23,38 @@ The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. `fable` and `opus` are Claude Code's rolling aliases. Claude resolves each alias to the latest available family revision. A runner receipt keeps the requested alias in `model` and the concrete provider-reported revision in `reportedModel`; verification accepts only a numeric `claude-fable-*` or `claude-opus-*` revision from the matching family. -The `astra`, `sol-6`, and `luna` rows are the GPT-6 Codex families (pstack-flex addition). These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent. Each is its own family with its own requested effort and probe; `sol-6` is independent of `sol`, so an existing GPT-5.6 Sol assignment stays unchanged until setup reassigns the role. All Codex families count as one provider for panel diversity. +The `astra`, `sol-6.1`, `sol-6`, and `luna` rows are the GPT-6 Codex families (pstack-flex addition). These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent. Each is its own family with its own requested effort and probe; `sol-6.1`, `sol-6`, and `sol` are independent of each other, so an existing GPT-6 Sol or GPT-5.6 Sol assignment stays unchanged until setup reassigns the role. All Codex families count as one provider for panel diversity. ## Default panel -The first-run panel roles (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) use these four lanes, one per entry, at each family's default effort: +The first-run panel roles (`arena runners`, `arena cross-judge pool`, `interrogate reviewers`) use these four lanes, one per entry, at each family's default effort: `claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.7@xhigh, claude:opus@max` -This line is the single source for the panel default. `setup-pstack`'s first-run sheet and the `arena`, `architect`, and `interrogate` skills copy it verbatim; the static invariant check fails when they drift. Solo code-writing roles (`feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb`) default to the `sol-6` row; exploration and swarm roles default to the `luna` row. +This line is the single source for the panel default. `setup-pstack`'s first-run sheet and the `arena` and `interrogate` skills copy it verbatim; the static invariant check fails when they drift. Solo code-writing roles (`feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb`) default to the `sol-6.1` row; exploration and swarm roles default to the `luna` row. + +## Default architect panel + +The first-run `architect runners` role uses two lanes, one per entry, at each family's default effort: + +`codex:gpt-6-astra@high, claude:fable@max` + +This line is the single source for the architect default. `setup-pstack`'s first-run sheet and the `architect` skill copy it verbatim; the static invariant check fails when they drift. + +## Sheet scope + +pstack-flex addition. A model sheet is either global or project-scoped. + +| Parent | Global sheet | Project sheet | +|---|---|---| +| Claude Code | `~/.claude/pstack-models.md` | `/.claude/pstack-models.md` | +| Codex | `~/.codex/pstack-models.md` | `/.codex/pstack-models.md` | + +The project root is the top level of the repository's primary checkout, so every worktree of one repository shares one project sheet. Read it as the parent directory of `git rev-parse --path-format=absolute --git-common-dir`. Outside a git repository there is no project sheet. + +Before the first configured role launches in a run, the parent reads its project sheet path once. If the file exists, it is the model sheet for the whole run and replaces the global sheet, including a global sheet already loaded into context. If the file does not exist, the global sheet applies. Never merge the two role by role: every sheet carries every documented role, so one sheet always answers. Say which sheet is in use when reporting a panel. + +A project sheet is private to the machine. `setup-pstack` writes it, lists it in `.git/info/exclude`, and never commits it. Deleting the file returns the project to the global sheet. ## Flex model matrix diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts index 6e849198..48e258a2 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts @@ -30,15 +30,14 @@ const MATRIX_HEADER = [ "Claude-native agent stem", ] as const; -const FAMILY_ORDER = ["fable", "sol", "grok", "opus", "astra", "sol-6", "luna"] as const; -const GPT6_FAMILIES = ["astra", "sol-6", "luna"] as const; +const FAMILY_ORDER = ["fable", "sol", "grok", "opus", "astra", "sol-6.1", "sol-6", "luna"] as const; +const GPT6_FAMILIES = ["astra", "sol-6.1", "sol-6", "luna"] as const; const PROVIDERS = ["claude", "codex", "grok"] as const; const DESCRIPTOR_RE = /(claude|codex|grok):[a-z0-9.-]+@(low|medium|high|xhigh|max)/g; const PANEL_ROLES = [ "arena runners", "arena cross-judge pool", - "architect runners", "interrogate reviewers", ] as const; const SHEET_ROLES = [ @@ -187,11 +186,14 @@ function defaultDescriptor(row: MatrixRow): string { return `${row.provider}:${row.model}@${row.defaultEffort}`; } -function parseDefaultPanel(markdown: string): string[] { +function parseDefaultPanel( + markdown: string, + heading = "## Default panel" +): string[] { const lines = markdown.split(/\r?\n/); - const start = lines.findIndex((line) => line.trim() === "## Default panel"); + const start = lines.findIndex((line) => line.trim() === heading); if (start < 0) { - throw new Error("missing ## Default panel"); + throw new Error(`missing ${heading}`); } for (let i = start + 1; i < lines.length; i++) { if (lines[i].startsWith("## ")) { @@ -201,7 +203,7 @@ function parseDefaultPanel(markdown: string): string[] { return lines[i].match(DESCRIPTOR_RE) ?? []; } } - throw new Error("## Default panel has no descriptor line"); + throw new Error(`${heading} has no descriptor line`); } function parseFrontmatter(text: string): { @@ -241,6 +243,10 @@ describe("model matrix", () => { const rows = parseModelMatrix(dispatch); const setup = readFileSync(SETUP_PATH, "utf8"); const panel = parseDefaultPanel(dispatch); + const architectPanel = parseDefaultPanel( + dispatch, + "## Default architect panel" + ); it("owns the effort universe and first-run defaults", () => { expect([...EFFORTS]).toEqual(["low", "medium", "high", "xhigh", "max"]); @@ -261,6 +267,7 @@ describe("model matrix", () => { ["grok", "xhigh"], ["opus", "max"], ["astra", "high"], + ["sol-6.1", "high"], ["sol-6", "high"], ["luna", "high"], ]); @@ -324,6 +331,7 @@ describe("model matrix", () => { ); expect(gpt6Rows.map((row) => [row.family, row.model])).toEqual([ ["astra", "gpt-6-astra"], + ["sol-6.1", "gpt-6.1-sol"], ["sol-6", "gpt-6-sol"], ["luna", "gpt-6-luna"], ]); @@ -334,13 +342,16 @@ describe("model matrix", () => { expect(row.defaultEffort).toBe("high"); expect(row.selectableEfforts).toEqual([...EFFORTS]); expect(row.claudeNativeAgentStem).toBeNull(); - expect(sheet).toContain(defaultDescriptor(row)); + // sol-6 stays selectable; sol-6.1 took its first-run roles. + if (row.family !== "sol-6") { + expect(sheet).toContain(defaultDescriptor(row)); + } } expect(new Set(rows.map((row) => row.family)).size).toBe(rows.length); expect(new Set(rows.map((row) => `${row.provider}:${row.model}`)).size) .toBe(rows.length); - // Solo code roles ride the sol-6 row; exploration and swarm ride luna. - const sol6 = defaultDescriptor(rows.find((row) => row.family === "sol-6")!); + // Solo code roles ride the sol-6.1 row; exploration and swarm ride luna. + const sol6 = defaultDescriptor(rows.find((row) => row.family === "sol-6.1")!); const luna = defaultDescriptor(rows.find((row) => row.family === "luna")!); for (const role of ["feature, refactoring", "bug-fix", "perf-issue", "hillclimb"]) { expect(sheet).toContain(`${role}: ${sol6}\n`); @@ -351,12 +362,15 @@ describe("model matrix", () => { expect(setup).toContain("Its model matrices (stock and flex)"); expect(setup).toContain("Read the model matrices, stock and flex."); expect(setup).toContain("any stock or flex matrix family"); - expect(setup).toContain("Offer every stock family, including Astra, GPT-6 Sol, and Luna, when changing `architect runners`"); + expect(setup).toContain("Offer every stock family, including Astra, GPT-6.1 Sol, GPT-6 Sol, and Luna, when changing `architect runners`"); expect(setup).toContain("Read each model, proposed effort, and selectable efforts from its row."); expect(setup).toContain("outside the stock and flex matrix families"); expect(setup).toContain( "| Astra | Astra matrix row + selected effort | external runner | native `spawn_agent` |" ); + expect(setup).toContain( + "| GPT-6.1 Sol | sol-6.1 matrix row + selected effort | external runner | native `spawn_agent` |" + ); expect(setup).toContain( "| GPT-6 Sol | sol-6 matrix row + selected effort | external runner | native `spawn_agent` |" ); @@ -386,6 +400,45 @@ describe("model matrix", () => { expect(providers.size).toBeGreaterThanOrEqual(2); }); + it("owns the architect default: Astra and Fable at matrix default efforts", () => { + expect(architectPanel).toEqual([ + "codex:gpt-6-astra@high", + "claude:fable@max", + ]); + const byDescriptor = new Set(rows.map(defaultDescriptor)); + for (const descriptor of architectPanel) { + expect(byDescriptor.has(descriptor)).toBe(true); + } + expect(firstRunSheet(setup)).toContain( + `architect runners: ${architectPanel.join(", ")}\n` + ); + }); + + it("scopes a sheet to the project or globally, and always asks which", () => { + const scopeStart = dispatch.indexOf("## Sheet scope"); + const scopeEnd = dispatch.indexOf("## Flex model matrix"); + expect(scopeStart).toBeGreaterThan(-1); + expect(scopeEnd).toBeGreaterThan(scopeStart); + const scope = dispatch.slice(scopeStart, scopeEnd); + for (const path of [ + "`~/.claude/pstack-models.md`", + "`/.claude/pstack-models.md`", + "`~/.codex/pstack-models.md`", + "`/.codex/pstack-models.md`", + ]) { + expect(scope).toContain(path); + expect(setup).toContain(path); + } + expect(scope).toContain("replaces the global sheet"); + expect(scope).toContain("If the file does not exist, the global sheet applies."); + expect(scope).toContain("Never merge the two role by role"); + expect(setup).toContain("### 1. Establish the parent and scope"); + expect(setup).toContain("Ask every run; never infer the scope"); + expect(setup).toContain("load the global sheet instead as the starting assignments"); + expect(setup).toContain("`.git/info/exclude`"); + expect(setup).toContain("Never add the sheet to a tracked `.gitignore`, stage it, or commit it."); + }); + it("passes each GPT-6 family's selected model and effort to the existing runner", () => { for (const row of rows.filter((row) => (GPT6_FAMILIES as readonly string[]).includes(row.family))) { for (const effort of row.selectableEfforts) { diff --git a/plugins/pstack/skills/setup-pstack/SKILL.md b/plugins/pstack/skills/setup-pstack/SKILL.md index f7ff08d4..63c5f322 100644 --- a/plugins/pstack/skills/setup-pstack/SKILL.md +++ b/plugins/pstack/skills/setup-pstack/SKILL.md @@ -1,19 +1,19 @@ --- name: setup-pstack -description: Configure pstack's provider-qualified models, per-family requested effort, and parent-owned routes per role. Verifies native and external Claude, Codex, Grok, DeepSeek, and MiniMax lanes before writing the override sheet. Use for /setup-pstack, "configure pstack models", or changing pstack's model choices. +description: Configure pstack's provider-qualified models, per-family requested effort, and parent-owned routes per role, in the global sheet or a private project sheet. Verifies native and external Claude, Codex, Grok, DeepSeek, and MiniMax lanes before writing the override sheet. Use for /setup-pstack, "configure pstack models", "set up pstack models for this project", or changing pstack's model choices. --- # Setup pstack -Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrices (stock and flex), descriptor grammar, and route table are the contract. Choose one requested effort per assigned matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback. +Configure one portable model sheet for the current parent harness, in the scope the operator picks: global, or private to one project. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrices (stock and flex), descriptor grammar, and route table are the contract. Choose one requested effort per assigned matrix family. Do not add a runtime resolver or a weaker-model fallback. The only configuration files are the global sheet and the optional project sheet that the Sheet scope section of provider dispatch defines. -Claude Code writes `~/.claude/pstack-models.md` and loads it from `~/.claude/CLAUDE.md` with: +In global scope, Claude Code writes `~/.claude/pstack-models.md` and loads it from `~/.claude/CLAUDE.md` with: ```text @~/.claude/pstack-models.md ``` -Codex writes `~/.codex/pstack-models.md`. Codex has no `@` include, so mirror the sheet's exact bytes inside one bounded block in `~/.codex/AGENTS.md` and retain the sheet as the editable source of truth: +In global scope, Codex writes `~/.codex/pstack-models.md`. Codex has no `@` include, so mirror the sheet's exact bytes inside one bounded block in `~/.codex/AGENTS.md` and retain the sheet as the editable source of truth: ```text @@ -21,21 +21,30 @@ Codex writes `~/.codex/pstack-models.md`. Codex has no `@` include, so mirror th ``` +In project scope, the sheet is `/.claude/pstack-models.md` on Claude Code and `/.codex/pstack-models.md` on Codex. It needs no include and no mirrored block: the parent reads the project path before dispatch, and an existing project sheet replaces the global sheet for that project. + ## Steps -### 1. Establish the parent +### 1. Establish the parent and scope Use the harness and tool surface running this skill: Claude Code or Codex. Environment markers may corroborate that top-level answer, but do not launch a child and ask it to detect where it came from. Record the parent because the same descriptor takes a different route in each harness. +Then ask which scope this run configures. Ask every run; never infer the scope from the working directory or from which sheets exist. Before asking, resolve the project root per the Sheet scope section and state both sheet paths and whether each file exists. + +- **Project**: this repository only. The sheet replaces the global sheet for work in this project and stays private to this machine. +- **Global**: every project that has no project sheet. + +Outside a git repository, say that only global scope is available and continue with it. A project run never reads or writes the parent integration files or the other scope's sheet. The scoped sheet is the file for the chosen scope. + ### 2. Load current state -Read the current parent-specific sheet when it exists. Before matrix validation, normalize only the rolling-alias predecessors that earlier pstack releases generated. A provider-qualified Claude model is migratable when its model component starts with `claude-fable-` or `claude-opus-` and the remaining revision contains only digits and hyphens. Replace that component in memory with `fable` or `opus`, preserving the provider, effort, role, and lane order. Record each original and normalized descriptor for the confirmation in step 7. This migration is valid loaded state and does not require a separate operator choice. +Read the scoped sheet when it exists. In project scope with no project sheet, load the global sheet instead as the starting assignments and say so; if neither exists, this is a first run. Before matrix validation, normalize only the rolling-alias predecessors that earlier pstack releases generated. A provider-qualified Claude model is migratable when its model component starts with `claude-fable-` or `claude-opus-` and the remaining revision contains only digits and hyphens. Replace that component in memory with `fable` or `opus`, preserving the provider, effort, role, and lane order. Record each original and normalized descriptor for the confirmation in step 7. This migration is valid loaded state and does not require a separate operator choice. -Treat the normalized values as current role-to-family assignments. Overlay those rows on the complete first-run role map in step 7. Materialize any missing documented role row from that map on the next successful write. A duplicate role row is inconsistent state; report it and resolve it before probing. A row whose role is not in the step 7 role map, such as `how critics`, is from a retired role. Drop it and list it at confirmation. A bare host-native slug from an older sheet is also invalid because it does not say which provider owns it. A versioned Claude model outside the two migration families remains inconsistent state. If the sheet is missing, use the complete first-run role map and the model matrix's Default effort cells. +Treat the normalized values as current role-to-family assignments. Overlay those rows on the complete first-run role map in step 7. Materialize any missing documented role row from that map on the next successful write. A duplicate role row is inconsistent state; report it and resolve it before probing. A row whose role is not in the step 7 role map, such as `how critics`, is from a retired role. Drop it and list it at confirmation. A bare host-native slug from an older sheet is also invalid because it does not say which provider owns it. A versioned Claude model outside the two migration families remains inconsistent state. If no sheet was loaded, use the complete first-run role map and the model matrix's Default effort cells. Then ask whether to keep these role-to-family assignments or change named roles. Keeping them is the default. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments. A changed role may use any stock or flex matrix family, `inherit-parent`, or `auto`. -Offer every stock family, including Astra, GPT-6 Sol, and Luna, when changing `architect runners` or another configurable role. Read each model, proposed effort, and selectable efforts from its row. The Codex families are separate families even though they share the Codex provider; changing one family's effort does not change another's. GPT-6 Sol uses the `sol-6` family; the `sol` family keeps GPT-5.6 Sol for sheets that still assign it. +List every stock and flex family by name when the operator changes a role, so Claude, Codex, Grok, DeepSeek, and MiniMax are all on offer and any role can move to any of them. Offer every stock family, including Astra, GPT-6.1 Sol, GPT-6 Sol, and Luna, when changing `architect runners` or another configurable role. Read each model, proposed effort, and selectable efforts from its row. The Codex families are separate families even though they share the Codex provider; changing one family's effort does not change another's. GPT-6.1 Sol uses the `sol-6.1` family and GPT-6 Sol the `sol-6` family; the `sol` family keeps GPT-5.6 Sol for sheets that still assign it. ### 3. Parse per-family efforts @@ -53,13 +62,14 @@ Ask one effort question for each assigned family. Name each model, its current o ### 5. Probe the requested pairs -Probe only the selected `provider:model@effort` pair of each assigned family. Run one probe per family in the role map, even when two families share a provider. Do not enumerate or offer older models as substitutes. A failed probe writes nothing: report the failing pair and provider, stop, and keep the active sheet plus parent integration bytes unchanged. A failed model demands explicit repair or role reassignment before saving. A failed first run creates neither artifact. +Probe only the selected `provider:model@effort` pair of each assigned family. Run one probe per family in the role map, even when two families share a provider. Do not enumerate or offer older models as substitutes. A failed probe writes nothing: report the failing pair and provider, stop, and keep the scoped sheet plus parent integration bytes unchanged. A failed model demands explicit repair or role reassignment before saving. A failed first run creates neither artifact. | Family | Pair source | Claude parent route | Codex parent route | Availability proof | |---|---|---|---|---| | Fable | Fable matrix row + selected effort | native Agent `pstack-fable-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe | | Sol | Sol matrix row + selected effort | `codex exec` | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | | Astra | Astra matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| GPT-6.1 Sol | sol-6.1 matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | | GPT-6 Sol | sol-6 matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | | Luna | Luna matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | | Grok | Grok matrix row + selected effort | Grok CLI | Grok CLI | `grok models` must list the requested model; one-turn probe | @@ -78,6 +88,7 @@ Receipts and native transcripts prove the requested effort and the route. They d Build the new sheet in memory. Do not write it yet. - First run: start from the complete role assignments in step 7, with the step 2 role changes applied. +- New project sheet seeded from the global sheet: treat it as a rerun of the loaded global rows. The global sheet itself is not rewritten. - Rerun: start from the normalized complete role map from step 2, with the step 2 role changes applied, preserving each loaded row's lane order and family (or alias) per lane. Require every documented role to remain present and non-empty, `architect runners` to keep at least two entries, and the final role map to contain at least one assigned matrix family. There is no requirement to assign every matrix family. @@ -90,7 +101,7 @@ Rewrite every matrix-family descriptor to `provider:model@` and `` pair in `~/.codex/AGENTS.md`. Replace that whole bounded block on a rerun. Insert one block at the end on first run. If either marker is missing, duplicated, or reversed, stop and report inconsistent state instead of guessing a boundary. Snapshot every target's current bytes. Write the sheet and parent integration only after every requested pair passes and the operator confirms. Read both targets back and compare them with the in-memory render. If either write or readback fails, restore every snapshot and report the failure. An unchanged rerun must produce byte-identical sheet and integration content after normalization. -Do not copy the model sheet between harnesses without rerunning the parent-specific probes; route availability can differ even on the same host. +Do not copy the model sheet between harnesses or between projects without rerunning the parent-specific probes; route availability can differ even on the same host. ### 9. Behavioral smoke Before declaring setup complete, run one small read-only mixed panel from this parent: every distinct chosen descriptor, distinct output/receipt paths, and an independent cross-judge when at least two providers are assigned. Launch Claude-native agents and every external process in the background with retained handles, then drain them. Verify the native transcript entries and every external receipt. A structural config check or unit test is not a substitute. -Report the sheet path, parent route table, requested-effort probe results, smoke results, and external elapsed/token/cost receipts. Re-running this skill re-probes and updates the same sheet. Do not claim the provider exposed hidden applied-effort observability. +Report the scope, the sheet path, parent route table, requested-effort probe results, smoke results, and external elapsed/token/cost receipts. For a project sheet, add that deleting the file returns the project to the global sheet. Re-running this skill in the same scope re-probes and updates the same sheet. Do not claim the provider exposed hidden applied-effort observability. diff --git a/tests/skill-collision-repro.sh b/tests/skill-collision-repro.sh index b9c237d7..b7599069 100755 --- a/tests/skill-collision-repro.sh +++ b/tests/skill-collision-repro.sh @@ -73,32 +73,38 @@ else fi # Static invariant (CHANGES maintenance note): provider-dispatch's "## Default panel" -# line is the default panel. The setup-pstack panel rows and the arena, architect, -# and interrogate defaults copy it verbatim. +# line is the default panel. The setup-pstack panel rows and the arena and +# interrogate defaults copy it verbatim. "## Default architect panel" plays the +# same part for architect runners. setup="$repo/plugins/pstack/skills/setup-pstack/SKILL.md" dispatch="$repo/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md" quad_of() { { grep -oE '(claude|codex|grok):[a-z0-9.-]+@(low|medium|high|xhigh|max)' || true; } | tr '\n' ' ' | sed 's/ $//'; } -canon_panel="$(awk ' - $0 == "## Default panel" { in_panel = 1; next } +panel_under() { awk -v heading="$1" ' + $0 == heading { in_panel = 1; next } in_panel && /^## / { exit } in_panel && /^`/ { print; exit } -' "$dispatch" | quad_of)" +' "$dispatch" | quad_of; } +canon_panel="$(panel_under "## Default panel")" +canon_architect="$(panel_under "## Default architect panel")" panel_bad="" [ -n "$canon_panel" ] || panel_bad="could not read the canonical panel from $dispatch"$'\n' +[ -n "$canon_architect" ] || panel_bad="${panel_bad}could not read the canonical architect panel from $dispatch"$'\n' # Anchor on the panel's last slug rather than a hard-coded one, so a model swap in # setup-pstack cannot leave this check hunting for a slug nobody ships any more. -anchor="${canon_panel##* }" -# arena and architect each state the panel on one line; interrogate lists it +# arena and architect each state their panel on one line; interrogate lists it # as one slug per row of its Reviewer A/B/C/D table (upstream #167). for name in arena architect; do skill="$repo/plugins/pstack/skills/$name/SKILL.md" + want="$canon_panel" + [ "$name" = architect ] && want="$canon_architect" + anchor="${want##* }" n="$(grep -Fc "$anchor" "$skill" || true)" if [ "$n" != "1" ]; then panel_bad="$panel_bad$skill: expected exactly 1 default-panel line, found $n"$'\n' continue fi got="$(grep -F "$anchor" "$skill" | quad_of)" - [ "$got" = "$canon_panel" ] || panel_bad="$panel_bad$skill: [$got] != [$canon_panel]"$'\n' + [ "$got" = "$want" ] || panel_bad="$panel_bad$skill: [$got] != [$want]"$'\n' done interrogate="$repo/plugins/pstack/skills/interrogate/SKILL.md" got="$(grep -E '^\| Reviewer [A-Z] \|' "$interrogate" | quad_of)" @@ -106,13 +112,15 @@ got="$(grep -E '^\| Reviewer [A-Z] \|' "$interrogate" | quad_of)" while IFS= read -r line; do got="$(printf '%s\n' "$line" | quad_of)" [ "$got" = "$canon_panel" ] || panel_bad="$panel_bad$setup role row: [$got] != [$canon_panel]"$'\n' -done < <(grep -E '^(arena runners|arena cross-judge pool|architect runners|interrogate reviewers):' "$setup") +done < <(grep -E '^(arena runners|arena cross-judge pool|interrogate reviewers):' "$setup") +got="$(grep -E '^architect runners:' "$setup" | quad_of)" +[ "$got" = "$canon_architect" ] || panel_bad="$panel_bad$setup architect runners: [$got] != [$canon_architect]"$'\n' if [ -n "$panel_bad" ]; then note "FAIL: the default model panel is not identical across provider dispatch, the panel skills, and setup-pstack:" note "$panel_bad" fail=1 else - note "ok: default model panel identical across provider dispatch + 3 panel skills + setup-pstack ($canon_panel)" + note "ok: default model panel identical across provider dispatch + 3 panel skills + setup-pstack ($canon_panel; architect $canon_architect)" fi plugin="$repo/plugins/pstack" @@ -334,14 +342,14 @@ else fi sol_descriptor="$(awk -F '|' ' - $2 ~ /^[[:space:]]*sol-6[[:space:]]*$/ { + $2 ~ /^[[:space:]]*sol-6\.1[[:space:]]*$/ { for (i = 4; i <= 6; i++) gsub(/^[[:space:]]+|[[:space:]]+$/, "", $i) print $4 ":" $5 "@" $6 } ' "$dispatch")" solo_code_bad="" if [ -z "$sol_descriptor" ]; then - solo_code_bad="could not read the sol-6 row from $dispatch"$'\n' + solo_code_bad="could not read the sol-6.1 row from $dispatch"$'\n' fi for role in bug-fix perf-issue hillclimb; do setup_descriptor="$(sed -n "s/^${role}: //p" "$setup")" @@ -355,11 +363,11 @@ for role in bug-fix perf-issue hillclimb; do fi done if [ -n "$solo_code_bad" ]; then - note "FAIL: solo code roles must use the sol-6 row:" + note "FAIL: solo code roles must use the sol-6.1 row:" note "$solo_code_bad" fail=1 else - note "ok: solo code roles stay on the sol-6 row ($sol_descriptor)" + note "ok: solo code roles stay on the sol-6.1 row ($sol_descriptor)" fi codex_manifest="$plugin/.codex-plugin/plugin.json"