diff --git a/CHANGES.md b/CHANGES.md index 92906368..11330ddf 100644 --- a/CHANGES.md +++ b/CHANGES.md @@ -1,7 +1,31 @@ # CHANGES — applied substitutions +## Unreleased: GPT-6 Codex families become stock + +- Move `astra`, `sol-6`, and `luna` from the additional matrix into the stock model matrix. The first-run sheet now assigns GPT-6 Sol high to `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high to `how explorer` and `swarm workers`; and Astra high in place of GPT-5.6 Sol on every panel. `codex:gpt-5.6-sol` stays selectable; existing sheets are untouched until a role is reassigned in setup. +- `provider-dispatch.md` gains a "Default panel" section that is the single source for the four panel lanes. `setup-pstack`, `arena`, `architect`, and `interrogate` copy it; the static invariant reads that line instead of deriving the panel from every matrix row, and the solo-code invariant checks the `sol-6` row. +- The matrix test asserts seven stock rows, the GPT-6 descriptors in the first-run sheet, and the panel contract. `UPSTREAM-FLEX.md` records the new permanent conflict surface on upstream syncs. Tracked in [pstack-flex #17](https://github.com/thisguymartin/pstack-flex/issues/17). + +## Unreleased: multiple gateway model choices + +- Add DeepSeek V4 Pro and MiniMax M3.1 Flash Preview as independent setup families alongside existing Flash and M3 choices. Keep provider-owned routing and existing descriptors. +- Validate unique model families and provider/model pairs instead of requiring one row per provider; cover both parent routes and substituted-model rejection. +- Document preview Token Plan access, model-specific thinking semantics, and required installed live validation. Tracked in [pstack-flex #5](https://github.com/thisguymartin/pstack-flex/issues/5). + This port applies the Cursor → Claude Code substitutions in skill bodies. Earlier drafts left them flagged; this revision resolves them. A later pass added a Codex build that shares the same skills; see [Codex port](#codex-port) below. +## pstack-flex (unreleased) — gateway lanes and optional families + +Fork of open-pstack v1.4.1. Additive changes, all in port-owned files: + +- Runner: new gateway providers `deepseek` and `minimax` (`runner/flex-providers.ts`). Each spawns the stock `claude` binary with the exact claude argv, plus injected environment: the lab's Anthropic-compatible endpoint, `ANTHROPIC_AUTH_TOKEN` from `DEEPSEEK_API_KEY`/`MINIMAX_API_KEY`, model pins, and an isolated `CLAUDE_CONFIG_DIR` (`~/.pstack-flex/`). Inherited `ANTHROPIC_*` values are deleted before injection so a parent's credentials or endpoint never bleed into a gateway child. +- OAuth-leak guard: a gateway lane refuses to start (in-process, `unauthenticated` receipt, exit 77) when its API key variable is missing or when its config dir carries a claude.ai OAuth credentials file, so a claude.ai login can never be pointed at a third-party endpoint. +- Gateway preflight is `claude --version`; the one-shot invocation is the real auth test. Gateway receipts force `costUsd` to null (the CLI prices at Anthropic rates) and match served models case-insensitively, falling back to `modelEvidence: "pinned-argv"` like Codex. +- `provider-dispatch.md`: new additive "Flex model matrix" section, extended route table, gateway preflight semantics, and the panel-diversity rule (arena runners and interrogate reviewers span at least two providers unless the operator explicitly confirms otherwise). The stock model matrix is byte-unchanged. +- Optional GPT-6 families: the additional model matrix declares `codex:gpt-6-astra`, `codex:gpt-6-sol`, and `codex:gpt-6-luna` with default effort `high` and selectable `low`, `medium`, `high`, `xhigh`, and `max`. Setup can assign and probe each for `architect runners` or another configurable role. Codex parents use native `spawn_agent`; Claude Code parents use the external Codex runner. GPT-6 Sol has its own `sol-6` family. Stock models and first-run role assignments stay unchanged. +- `setup-pstack`: role assignments are selected first, and only assigned families get effort questions and probes; there is no requirement to assign every matrix family (mirrors upstream PR #73 / issue #72). The first-run sheet, its stock quad, and the fail-closed write rules are unchanged. +- Tests: the model-matrix contract gains a flex-matrix section check cross-validated against the runner's gateway specs; runner, commands, parse-output, and CLI tests cover env injection, the guard, cost nulling, and case-insensitive verification. Nothing in the suite performs network I/O. + ## 1.4.1 syncs to Cursor pstack 0.15.1 Open Pstack 1.4.1 tracks Cursor pstack 0.15.1 at `f8abeddd1862dc73704e3d719dd73df0d51b8c71`. Poteto-mode now requires each claim to include its evidence or a measured, inferred, or guess label in the same sentence. Agents also run any check they can run themselves instead of handing that check to the user. No playbook, model, runtime, or dependency changed. diff --git a/LICENSE b/LICENSE index 6b540023..643dd5d0 100644 --- a/LICENSE +++ b/LICENSE @@ -1,6 +1,7 @@ MIT License Copyright (c) 2026 Lauren Tan +Copyright (c) 2026 Martin Patino (pstack-flex modifications) Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal diff --git a/NOTICE.md b/NOTICE.md index 0ec8a23e..a1be00dc 100644 --- a/NOTICE.md +++ b/NOTICE.md @@ -2,6 +2,10 @@ This plugin is a port of upstream MIT-licensed work. All upstream copyright notices and license terms are preserved. The open-pstack history begins from `michael-denyer/pstack-claude` through proven import commit `053ed78732e3b71826933170eafe7f7782dda844`. +## pstack-flex provenance + +This repository, **pstack-flex** (Martin Patino), is a fork of [ericlitman/open-pstack](https://github.com/ericlitman/open-pstack) at v1.4.1 (`de67e6b40511814171e5e4c8ad7af3b79f07c9ee`), which ports [Lauren Tan's pstack](https://github.com/cursor/plugins/tree/main/pstack) (Cursor) to Claude Code and Codex. Provenance chain: pstack-flex <- ericlitman/open-pstack <- cursor/plugins/pstack. All licenses remain MIT; every upstream license and notice file is preserved. The flex gateway providers, optional families, docs, and tests are (c) 2026 Martin Patino, MIT, and are inventoried in [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md). + ## Upstream sources | Component | Upstream | Copyright | License | License file | diff --git a/README.md b/README.md index 64901115..8f61368c 100644 --- a/README.md +++ b/README.md @@ -1,16 +1,16 @@ -# open-pstack +# pstack-flex -[![CI](https://github.com/ericlitman/open-pstack/actions/workflows/ci.yml/badge.svg)](https://github.com/ericlitman/open-pstack/actions/workflows/ci.yml) -[![Latest release](https://img.shields.io/github/v/release/ericlitman/open-pstack)](https://github.com/ericlitman/open-pstack/releases/latest) -[![MIT license](https://img.shields.io/github/license/ericlitman/open-pstack)](LICENSE) +[![CI](https://github.com/thisguymartin/pstack-flex/actions/workflows/ci.yml/badge.svg)](https://github.com/thisguymartin/pstack-flex/actions/workflows/ci.yml) +[![Fork of open-pstack v1.4.1](https://img.shields.io/badge/fork%20of-open--pstack%20v1.4.1-blue)](https://github.com/ericlitman/open-pstack/releases/tag/v1.4.1) +[![MIT license](https://img.shields.io/github/license/thisguymartin/pstack-flex)](LICENSE) -**Open Pstack brings [Lauren Tan (@poteto)](https://x.com/poteto)'s [pstack](https://github.com/cursor/plugins/tree/main/pstack) to Claude Code and Codex.** Its job is to stay as close to her original work as possible while translating the parts that depend on Cursor. +**pstack-flex runs [Lauren Tan (@poteto)](https://x.com/poteto)'s [pstack](https://github.com/cursor/plugins/tree/main/pstack) in Claude Code and Codex on the models you actually have.** It is a fork of [ericlitman/open-pstack](https://github.com/ericlitman/open-pstack), which translates pstack's Cursor-specific parts for Claude Code and Codex. open-pstack assumes four frontier subscriptions. This fork keeps its skills and workflows and changes one thing: which models setup accepts and how they are reached. Lauren built pstack from the skills she uses to ship code at Cursor. In a [55-minute interview with Denis Labelle](https://x.com/DenisLabelle/status/2091337807939706928), she says that she shipped 1,000 pull requests in one month after steadily improving how her agents work and verify their results. > If you want to go fast, go deep first. -Open Pstack is an unofficial community project that makes pstack work in Claude Code and Codex. If Cursor is your main coding environment, use [Lauren's original pstack](https://github.com/cursor/plugins/tree/main/pstack). If Claude Code or Codex is your main coding environment, use this repository. +If Cursor is your main coding environment, use [Lauren's original pstack](https://github.com/cursor/plugins/tree/main/pstack). If you hold all four subscriptions and want the closest translation, use [open-pstack](https://github.com/ericlitman/open-pstack). If you want to pick your own models, pay per token where it makes sense, or run with no subscription at all, use this repository. ## What pstack does @@ -23,23 +23,63 @@ The normal entry point is `poteto-mode`. You give it a task in plain language. I - compares designs when the choice matters; - favors small, simple changes over extra machinery; - asks several models to challenge important decisions when useful; -- runs the code and checks real behavior instead of stopping at “the tests pass”; and +- runs the code and checks real behavior instead of stopping at "the tests pass"; and - carries the work through review, continuous integration (CI), and a ready-to-merge pull request when asked. ![How pstack routes a task through focused skills, real-app proof, and a review-ready pull request](assets/pstack-workflow.png) pstack does not ask you to trust an agent on day one. It helps the agent leave evidence you can inspect. Start with supervised work. Let it run more work in parallel only after its checks have earned that trust in your own repositories. +## The models + +Every pstack role (who writes code, who explores, who sits on a review panel) maps to one family. A family is one `(provider, model)` pair with its own requested effort and its own live probe in setup. These are the families pstack-flex ships: + +| Family | Descriptor at default effort | Needs | First-run role | +| --- | --- | --- | --- | +| `fable` | `claude:fable@max` | Claude Code login | judgment, prose, explanation, hardest tasks, panels | +| `opus` | `claude:opus@xhigh` | Claude Code login | panels | +| `astra` | `codex:gpt-6-astra@high` | Codex (ChatGPT) login | panels | +| `sol-6` | `codex:gpt-6-sol@high` | Codex (ChatGPT) login | feature, refactoring, bug-fix, perf-issue, hillclimb | +| `luna` | `codex:gpt-6-luna@high` | Codex (ChatGPT) login | how explorer, swarm workers | +| `sol` | `codex:gpt-5.6-sol@max` | Codex (ChatGPT) login | none; selectable | +| `grok` | `grok:grok-4.6@xhigh` | Grok CLI login | panels | +| `deepseek` | `deepseek:deepseek-flash@high` | `DEEPSEEK_API_KEY` | none; selectable | +| `deepseek-pro` | `deepseek:deepseek-v4-pro@high` | `DEEPSEEK_API_KEY` | none; selectable | +| `minimax` | `minimax:MiniMax-M3@high` | `MINIMAX_API_KEY` | none; selectable | +| `minimax-preview` | `minimax:MiniMax-M3.1-Flash-Preview@high` | `MINIMAX_API_KEY` (Token Plan) | none; selectable | + +The default review panel is `claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh`: four lanes across three providers. Any family can take any role. Panels must span at least two providers, and two models from one provider count as one, because the adversarial signal comes from model diversity. + +The DeepSeek and MiniMax lanes run the stock `claude` binary against the lab's Anthropic-compatible endpoint with that lab's key, in an isolated config directory, with inherited Anthropic routing stripped. A lane refuses to start if it finds a claude.ai login in that directory, so a subscription credential can never reach a third-party endpoint. Their receipts keep real token usage but set `costUsd` to null (Claude Code prices at Anthropic rates); the price table is in [docs/LANES.md](docs/LANES.md). Anthropic does not support pointing Claude Code at non-Anthropic endpoints; use synthetic data for gateway testing and keep keys in your local environment. + +### How a role becomes a lane + +```mermaid +flowchart LR + S["pstack-models.md
role -> provider:model@effort"] --> P["Parent harness
(Claude Code or Codex)"] + P -->|"parent's own provider"| N["Native subagent
Agent / spawn_agent"] + P -->|"any other provider"| R["pstack-runner
one process per lane"] + R --> C1["codex CLI"] + R --> C2["grok CLI"] + R --> C3["claude CLI + env
DeepSeek or MiniMax endpoint"] + N --> O["Output + receipt
model, effort, tokens, status"] + C1 --> O + C2 --> O + C3 --> O +``` + +The parent resolves every route once, before fan-out. Children never detect the harness or pick a model. A lane that cannot start drops out with a named receipt; nothing substitutes a weaker model or invents a timeout. + ## Install -You need a current Claude Code or Codex installation. For the full four-model review, install and sign in to the Claude Code, Codex, and Grok command-line tools. [Bun](https://bun.sh) runs the small local tool that starts models outside the app you are using. You can still use the core workflows with fewer models. +You need a current Claude Code or Codex installation and [Bun](https://bun.sh) for the lane runner. Sign in only to the CLIs whose plans you have (Claude Code, Codex, Grok), and export `DEEPSEEK_API_KEY` or `MINIMAX_API_KEY` in the shell that starts your session for the gateway lanes. Any subset works, down to a zero-subscription setup on two keys. ### Claude Code Run these commands inside Claude Code: ```text -/plugin marketplace add ericlitman/open-pstack +/plugin marketplace add thisguymartin/pstack-flex /plugin install pstack@open-pstack /reload-plugins ``` @@ -49,7 +89,7 @@ Run these commands inside Claude Code: Run these commands in your shell: ```shell -codex plugin marketplace add ericlitman/open-pstack --ref main +codex plugin marketplace add thisguymartin/pstack-flex --ref main codex plugin add pstack@open-pstack ``` @@ -64,8 +104,6 @@ Start a new Codex task after installation so it can discover the new skills and ## Get started -Lauren's original setup has two steps. Open Pstack keeps the same flow. - ### 1. Set up the models In Claude Code, run: @@ -80,9 +118,9 @@ In Codex, ask: Use pstack:setup-pstack to configure pstack. ``` -Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-5.6 Sol, Grok 4.6, and Opus. +Setup is assignment-first. It shows the role map, asks which roles to change, asks one effort per assigned family, probes only those families with a real one-turn run, and writes nothing until every probe passes and you confirm. A fresh run proposes the defaults in the table above. An existing sheet keeps its assignments until you change a named role. -An older model sheet starts using the rolling aliases in memory as soon as this release is installed. Run setup once after updating to persist that migration. It replaces versioned Fable and Opus entries while preserving every role assignment and effort selection. +The sheet lives at `~/.claude/pstack-models.md` (Claude Code) or `~/.codex/pstack-models.md` (Codex). It is global, not per project. Change it by rerunning setup rather than editing it by hand, so every choice is probed before it is saved. ### 2. Use poteto-mode @@ -102,7 +140,7 @@ Use pstack:poteto-mode. Add saved filters to search. Keep the design simple, ver For that feature, poteto-mode should first understand how search works today. It should decide how the data should be represented before writing code, implement the smallest complete version, run the feature the way a user would, review the result, and prepare the pull request. -That is the main workflow. The other skills are there when poteto-mode needs them or when you want to call one directly. +That is the main workflow. The other skills are there when poteto-mode needs them or when you want to call one directly. **[docs/USAGE.md](docs/USAGE.md)** is the longer walkthrough: three setup configurations (full frontier, hybrid saver, zero-subscription), copy-paste examples for the daily skills, how to read receipts, and troubleshooting. ## Useful skills @@ -120,11 +158,9 @@ That is the main workflow. The other skills are there when poteto-mode needs the Plugin skills include `pstack:` in their name. In Claude Code, invoke a native skill such as `/pstack:architect`. In Codex, ask for the skill, such as `Use pstack:architect for this design.` See the [technical reference](docs/reference.md) for the full list. -## Models and token use - -Some pstack workflows use one model. Skills such as `architect`, `arena`, and `interrogate` can run several models in parallel. Each model run uses the subscription and token allowance of its own command-line tool. +## Cost -`setup-pstack` lets you choose the models, one requested effort per model family, and how many run in parallel. A model from the app you are using runs inside that app. Other models run through their own command-line tools. Open Pstack does not quietly replace a failed model with a weaker one. +Some workflows use one model. `architect`, `arena`, and `interrogate` run several in parallel. Subscription lanes spend that CLI's plan; gateway lanes bill per token on the lab's account. The cost playbook in [docs/USAGE.md](docs/USAGE.md#cost-playbook) shows where the gateway lanes pay off: high-volume code-writing roles on DeepSeek Flash, long-context reading on MiniMax M3, and one frontier lane plus two gateway lanes for a three-provider panel at a fraction of three subscriptions. Keep `judgment and prose` and `hardest tasks` on your strongest lane; they are the last roles to economize. pstack-flex never replaces a failed model with a cheaper one; a lane that fails is reported, not swapped. ## Claude Code and Codex @@ -133,38 +169,27 @@ Both apps read the same pstack skills. Only the way they start those skills and | | Claude Code | Codex | | --- | --- | --- | | Start poteto-mode | Claude loads a small startup instruction that can route non-trivial work into it. You can also run `/pstack:poteto-mode` yourself. | Ask for `pstack:poteto-mode` by name. Codex does not load the Claude startup instruction. | -| Runs inside the app | Claude models stay inside Claude Code. | The Sol model stays inside Codex. | -| Other models | Codex and Grok run through their signed-in command-line tools. | Claude and Grok run through their signed-in command-line tools. | +| Runs inside the app | Claude models stay inside Claude Code. | The Codex families stay inside Codex. | +| Other models | The Codex families and Grok run through their signed-in command-line tools. | Claude and Grok run through their signed-in command-line tools. | +| Gateway models | DeepSeek and MiniMax always run through the external runner with an isolated config directory, never as a native agent. | Same. | | Skills and workflows | Shared with Codex. | Shared with Claude Code. | -Grok can take part in a multi-model review. You cannot use Grok as the main app running pstack. +Grok, DeepSeek, and MiniMax can take part in a multi-model review. You cannot use any of them as the main app running pstack. -## Learn from the original +## Upstream Lauren's [pstack guide](https://github.com/cursor/plugins/tree/main/pstack/docs/guide) walks through a real task, verification, and longer unattended runs. It uses Cursor's interface, but the ideas are the same. Use the translated skill invocations above in Claude Code or Codex. -This repository also keeps: - -- [the original README](README-UPSTREAM.md), unchanged; -- [the technical reference](docs/reference.md) for every skill, dependency, and Claude Code or Codex detail; -- [the upstream sync record](UPSTREAM.md) and update process; -- [the change record](CHANGES.md) for every adaptation; and -- [the attribution record](NOTICE.md) for pstack and the imported Cursor Team Kit skills. - -## Staying close to Lauren's pstack - -Open Pstack 1.4.1 tracks pstack 0.15.1 at Cursor commit [`f8abeddd1862dc73704e3d719dd73df0d51b8c71`](https://github.com/cursor/plugins/commit/f8abeddd1862dc73704e3d719dd73df0d51b8c71). - -The two projects have separate version numbers. The pstack version identifies Lauren's upstream content. The Open Pstack version identifies the Claude Code and Codex package built from it. +This repository tracks two upstreams. [UPSTREAM.md](UPSTREAM.md) records the Cursor pstack commit open-pstack imported (0.15.1 at [`f8abedd`](https://github.com/cursor/plugins/commit/f8abeddd1862dc73704e3d719dd73df0d51b8c71)) and how new pstack releases are brought over. [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md) records the open-pstack fork point (v1.4.1), which files this fork owns, and the merge procedure. The fork keeps every upstream skill body as-is except the default model descriptors; its own changes are the model matrix, the first-run sheet, the gateway providers in the runner, setup's assignment-first flow, and the docs. -In this repository, “upstream” means Lauren's original pstack. Open Pstack does not promise instant updates. It records the exact version it follows, reviews new changes in order, and changes only what Claude Code and Codex require. New pstack behavior belongs in Lauren's project first whenever possible. +Also kept here: [the original README](README-UPSTREAM.md), unchanged; [the technical reference](docs/reference.md) for every skill and harness detail; [the change record](CHANGES.md); and [the attribution record](NOTICE.md). ## Contributing -Fixes for Claude Code or Codex and help bringing over new pstack releases are welcome. Search [GitHub Issues](https://github.com/ericlitman/open-pstack/issues) before opening a new issue. For larger behavior changes, explain why the change belongs in Open Pstack instead of Lauren's original project. +Fixes for Claude Code or Codex, new lanes, and help bringing over new pstack releases are welcome. Search this repository's [GitHub Issues](https://github.com/thisguymartin/pstack-flex/issues) before opening a new one. For changes to upstream-derived content, explain why the change belongs here instead of in open-pstack or Lauren's original project. -Read [UPSTREAM.md](UPSTREAM.md) before changing content brought over from Lauren's pstack. Pull requests must keep one shared skill tree for Claude Code and Codex and pass the repository's tests, type checks, plugin validation, and static checks. +Read [UPSTREAM.md](UPSTREAM.md) and [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md) before changing content brought over from either upstream. Pull requests must keep one shared skill tree for Claude Code and Codex and pass the repository's tests, type checks, plugin validation, and static checks. Nothing merges until the exact candidate is installed and the changed behavior passes a live test from the real user surface in every affected harness; the [pull request template](.github/pull_request_template.md) records that evidence, and a PR without it stays a draft. Adding a gateway provider has its own checklist in [docs/LANES.md](docs/LANES.md#adding-a-gateway-provider). ## License -MIT. pstack was created by Lauren Tan. Open Pstack builds on Michael Denyer's [pstack-claude](https://github.com/michael-denyer/pstack-claude) port and includes attributed MIT-licensed work from Cursor Team Kit and Superpowers. See [NOTICE.md](NOTICE.md) and the preserved license files for details. +MIT. pstack was created by Lauren Tan. open-pstack builds on Michael Denyer's [pstack-claude](https://github.com/michael-denyer/pstack-claude) port and includes attributed MIT-licensed work from Cursor Team Kit and Superpowers. pstack-flex is a fork of open-pstack. See [NOTICE.md](NOTICE.md) and the preserved license files for details. diff --git a/UPSTREAM-FLEX.md b/UPSTREAM-FLEX.md new file mode 100644 index 00000000..9d4b234d --- /dev/null +++ b/UPSTREAM-FLEX.md @@ -0,0 +1,47 @@ +# Flex fork synchronization + +pstack-flex layers on top of open-pstack's own upstream tracking. Two sync relationships exist: + +1. `cursor/plugins/pstack` -> `ericlitman/open-pstack` — documented in [UPSTREAM.md](UPSTREAM.md), unchanged by this fork. +2. `ericlitman/open-pstack` -> `thisguymartin/pstack-flex` — this document. + +## Fork point + +| Source | Value | +| --- | --- | +| Repository | `https://github.com/ericlitman/open-pstack.git` | +| Tag | `v1.4.1` | +| Commit | `de67e6b40511814171e5e4c8ad7af3b79f07c9ee` | +| Tracks Cursor pstack | `0.15.1` (`f8abedd`) | + +The fork keeps full upstream history. The `upstream` remote points at ericlitman/open-pstack. + +## What the fork owns + +All flex changes are additive and live in port-owned files so upstream merges stay cheap: + +- `plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts` and `flex-providers.test.ts` (new) +- Gateway-provider hooks in `runner/{types,commands,run,parse-output,cli}.ts` and their tests +- The three GPT-6 rows in the stock model matrix, the "Default panel" section, the "Flex model matrix" section, and the route-table columns in `references/provider-dispatch.md` +- The first-run sheet in `skills/setup-pstack/SKILL.md` and the default descriptors named in `arena`, `architect`, `interrogate`, `how`, `swarm`, and the `feature`, `refactoring`, `bug-fix`, `perf-issue`, and `hillclimb` playbooks +- The assignment-first restructure of `skills/setup-pstack/SKILL.md` +- `docs/LANES.md`, this file, the README fork section, and the NOTICE/LICENSE/CHANGES additions + +Every upstream skill body is byte-unchanged except for the default-descriptor mentions listed above. Since [#17](https://github.com/thisguymartin/pstack-flex/issues/17), the stock matrix carries three fork-owned GPT-6 rows and the first-run sheet uses them, so those two surfaces conflict on every upstream sync and are resolved by hand: keep the fork's rows and defaults, take upstream's wording for everything else. + +## Merge procedure + +```shell +git fetch upstream +git switch -c merge-rehearsal +git merge --no-ff --no-commit upstream/main +# inspect, resolve, run the full local gate, then merge for real or abort +``` + +Expected conflict surface on future upstream releases: + +- `plugins/pstack/skills/setup-pstack/SKILL.md` — upstream issue #88 (1.5.0, syncing Cursor pstack 0.15.5) folds upstream PR #73, which moves setup to the same assignment-first, probe-only-assigned shape this fork already uses. Resolve toward upstream's wording wherever it covers the same rule; keep the flex families and the diversity rule. +- `plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts` — upstream 1.5.0 changes the stock panel to three lanes. Take upstream's stock assertions verbatim; the flex-matrix describe block is fork-owned and should survive as-is. +- `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md` — the four upstream rows and their prose are upstream's; the GPT-6 rows, the "Default panel" section, and the flex section are fork-owned. Upstream's own panel changes land in the "Default panel" line only if the fork wants them. + +After every merge: run the full local gate (`bun install --frozen-lockfile`, `bun run test`, `bun run typecheck`, manifest JSON parse, `PSTACK_STATIC_ONLY=1 bash tests/skill-collision-repro.sh`), then record the installed version, action, and observed result for each affected harness in the pull request before tagging. diff --git a/UPSTREAM.md b/UPSTREAM.md index 6d725c48..71b5b08d 100644 --- a/UPSTREAM.md +++ b/UPSTREAM.md @@ -18,7 +18,7 @@ The table above is the current Cursor sync point. Open Pstack 1.4.1 imports this - Commits `799151d` and `6fecddb` add and relocate `make-bot-ui`. It depends on Cursor routines, webhook events, and UI primitives that Claude Code and Codex do not share. - Four `disable-model-invocation: true` lines from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`. Poteto-mode invokes those skills by name, and the flag blocks that route on Claude Code. -- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on `codex:gpt-5.6-sol@max` for cost. +- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on a Codex model (`codex:gpt-6-sol@high` in pstack-flex) for cost. - The Claude manifest does not take the logo field from `efa2a53` because Claude Code has no schema for it. The shared asset is exposed through the Codex manifest instead. ## Check for changes diff --git a/docs/LANES.md b/docs/LANES.md new file mode 100644 index 00000000..9cd07efb --- /dev/null +++ b/docs/LANES.md @@ -0,0 +1,149 @@ +# Lanes: models, providers, and cost control + +pstack-flex's reason to exist: you choose which models run and what they cost. This document covers the lane concepts, the gateway environment reference, prices, the zero-subscription walkthrough, and the safety rules. + +Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against each provider's own docs before relying on a number. + +## Lane kinds + +| Kind | Lanes | Auth | Billing | Route | +| --- | --- | --- | --- | --- | +| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna`, `codex:gpt-5.6-sol`, `grok:grok-4.6` | each CLI's own login | that CLI's plan | native or external per the route table | +| Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner | + +A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide). + +## GPT-6 Codex families + +Three stock Codex families carry the first-run defaults. They use the same ChatGPT login as `codex:gpt-5.6-sol`: + +| Family | Descriptor at default requested effort | First-run roles | Codex's own description | +| --- | --- | --- | --- | +| astra | `codex:gpt-6-astra@high` | every panel (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) | Frontier tier for the most demanding work | +| sol-6 | `codex:gpt-6-sol@high` | `feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb` | Coding and everyday workhorse | +| luna | `codex:gpt-6-luna@high` | `how explorer`, `swarm workers` | Fast, low-cost tier for easier tasks | + +A fresh `/setup-pstack` run proposes these. An existing sheet keeps its assignments until you change a named role in setup; `codex:gpt-5.6-sol` remains a selectable family for that. The `sol-6` family is separate from the `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus GPT-6 Sol does not satisfy the two-provider rule; the default panel spans Claude, Codex, and Grok. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27. + +## Multiple models per provider + +The flex matrix now includes four independently assignable model families: + +| Family | Descriptor at default requested effort | Selection guidance | +| --- | --- | --- | +| deepseek | `deepseek:deepseek-flash@high` | Existing everyday option | +| deepseek-pro | `deepseek:deepseek-v4-pro@high` | Candidate for difficult debugging, architecture, and review | +| minimax | `minimax:MiniMax-M3@high` | Existing MiniMax option | +| minimax-preview | `minimax:MiniMax-M3.1-Flash-Preview@high` | Preview coding option with tunable thinking | + +These are choices, not automatic replacements or a performance ranking. Existing sheets keep their assignments. In `/setup-pstack`, assign named roles to the desired model family; efforts and probes are independent per model, even for models sharing a key. Two models from one provider count as one provider for panel diversity. No runtime routing change or new configuration file is needed. + +As of 2026-09-27, [MiniMax's model guide](https://platform.minimax.io/docs/guides/models-intro) restricts M3.1 Flash Preview to Token Plan and MiniMax Code. For gateway access, supply the eligible Token Plan key as `MINIMAX_API_KEY`; the live probe must confirm entitlement. It is not a zero-subscription option. A working M3 call does not establish preview access. + +[MiniMax's Anthropic API](https://platform.minimax.io/docs/api-reference/text-anthropic-api) documents always-on thinking for the preview and `output_config.effort` from `low` to `max`. Higher effort increases thinking latency; the matrix proposes `high`, while the API defaults to `max` when omitted. M3 defaults to thinking off at the API and needs adaptive thinking to enable it. Its requested effort flag is not evidence of the preview's depth controls. Verify the installed CLI forwards the intended parameters; receipts prove requested effort, not hidden applied depth. [DeepSeek documents V4 Pro through its Anthropic endpoint](https://api-docs.deepseek.com/guides/anthropic_api). + +Before recommending a fastest or strongest default, compare the same synthetic coding tasks for correctness, completion time, tool-call reliability, token usage, and actual provider billing. Preview pricing and plan limits must be checked against the active plan rather than inferred from M3 rates. + +## Gateway environment reference + +Set by you: + +| Variable | Required | Meaning | +| --- | --- | --- | +| `DEEPSEEK_API_KEY` / `MINIMAX_API_KEY` | yes, per lane | the lab's API key; the lane refuses to start without it | +| `DEEPSEEK_BASE_URL` / `MINIMAX_BASE_URL` | no | endpoint override; defaults are in the flex model matrix | +| `PSTACK_FLEX_DEEPSEEK_CONFIG_DIR` / `PSTACK_FLEX_MINIMAX_CONFIG_DIR` | no | config-dir override; default `~/.pstack-flex/` | +| `DEEPSEEK_MAX_CONTEXT_TOKENS` / `MINIMAX_MAX_CONTEXT_TOKENS` | no | context-cap override for the claude CLI | + +Injected by the runner at spawn time (never written to disk, never in receipts): `ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN`, the model pins (`ANTHROPIC_MODEL`, the opus/sonnet/haiku alias defaults, `CLAUDE_CODE_SUBAGENT_MODEL`), `CLAUDE_CODE_ATTRIBUTION_HEADER=0`, `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1`, and `CLAUDE_CONFIG_DIR`. The runner first removes inherited `ANTHROPIC_*` values and Claude Code cloud-provider flags from the parent session. + +## Storing keys + +Keys reach a lane through the environment only; the runner never writes them to disk, receipts, or sheets. So key hygiene is entirely about how your shell gets them. Do not put raw keys in dotfiles or committed `.env` files. + +Recommended: your OS keychain, loaded on demand. + +- **macOS** (built in, encrypted at rest, unlocks with login): + + ```zsh + # once per key — prompts for the value, nothing lands in shell history + security add-generic-password -a "$USER" -s pstack-deepseek -w + security add-generic-password -a "$USER" -s pstack-minimax -w + + # in .zshrc: a function, not an export — keys enter env only when called + pstack-keys() { + export DEEPSEEK_API_KEY=$(security find-generic-password -a "$USER" -s pstack-deepseek -w) + export MINIMAX_API_KEY=$(security find-generic-password -a "$USER" -s pstack-minimax -w) + } + ``` + +- **Linux**: `pass` (GPG-encrypted, git-syncable) or `secret-tool` (libsecret) with the same load-on-demand function shape. +- **1Password CLI**: `op run --env-file=.env.tpl -- claude` injects the keys at process start with biometric unlock and exports nothing into the shell permanently. +- **direnv**: fine for per-project scoping (gateway lanes are per-project opt-in anyway), but a raw `.envrc` is plaintext — have it call the keychain instead of holding the key. + +Honest threat model: encryption at rest protects against dotfile repos, backups, and file theft. Once a key is in process env, any process running as your user can read it — the same exposure your CLI OAuth credential files already have. Keychain storage plus two ops controls is the right amount: **set spend caps on the DeepSeek and MiniMax dashboards** (the real blast-radius limiter) and rotate keys if a machine is ever compromised. + +## Prices (verified 2026-09-25 — re-check before budgeting) + +| Lane | Price per million tokens | Notes | +| --- | --- | --- | +| DeepSeek V4.1-Flash (`deepseek-flash`) | $0.30 in / $1.20 out peak; $0.15 / $0.60 off-peak; cache hits near-free | Off-peak windows: 01:00-04:00 and 06:00-10:00 UTC on weekdays. The discount is automatic on DeepSeek's side; pstack-flex surfaces the window but never delays your work to hit it. MIT open weights. | +| DeepSeek V4-Pro | $1.32 / $3.96 peak; half off-peak | Stronger model for hard lanes; assign it per role if wanted. | +| MiniMax M3 (`MiniMax-M3`) | $0.30 / $1.20 at up to 512K input; higher above | 1M context. Custom community model license (irrelevant for API use). | +| Claude / Codex / Grok subscription lanes | plan-dependent | Billed by each provider's plan, not per token here. | + +Gateway receipts always report `costUsd: null`: the claude CLI computes `total_cost_usd` at Anthropic list prices, which would be fiction for third-party traffic. Token usage in receipts is real — multiply it by the table above. + +## Zero-subscription walkthrough + +Goal: run poteto-mode and its panels with no Claude, ChatGPT, or Grok plan — only two API keys. The Claude Code binary is a free download; a subscription is only needed to reach Anthropic's servers. + +1. Install the claude CLI, Bun, and this plugin as usual. Do not run `claude login` anywhere in this setup. +2. Export `DEEPSEEK_API_KEY` and `MINIMAX_API_KEY`. +3. Make the parent session itself a DeepSeek session — same mechanism as a lane, applied to your interactive shell: + + ```shell + export ANTHROPIC_BASE_URL="" + export ANTHROPIC_AUTH_TOKEN="$DEEPSEEK_API_KEY" + export ANTHROPIC_MODEL="deepseek-flash" + export CLAUDE_CODE_SUBAGENT_MODEL="deepseek-flash" + export CLAUDE_CONFIG_DIR="$HOME/.pstack-flex/parent-deepseek" + claude + ``` + + Native `claude:*` lanes spawned by this parent inherit the endpoint, so the fable/opus role slots ride DeepSeek too. +4. Run `/setup-pstack`. Assign roles across the `deepseek` and `minimax` families (a `budget-duo` style panel), skip the unassigned stock families, and let the probes confirm both endpoints. +5. Panels keep real diversity: DeepSeek and MiniMax are two distinct providers, which satisfies the two-provider panel rule without any override. + +Quality note: this trades peak capability for cost control. The hardest-task role on a frontier subscription lane is a config choice you can add later without touching anything else. + +## Safety and policy + +- **Unsupported, not prohibited.** Anthropic's docs state that routing Claude Code to non-Claude models through gateways is not supported. No terms clause or enforcement against pointing the unmodified binary at a third-party endpoint was found (2026-09-25), but a CLI update can break compatibility without notice. Pin the claude CLI version on machines that depend on gateway lanes and bump it deliberately. +- **Never a claude.ai login on a gateway path.** Do not run `claude login` or `claude setup-token` inside any `~/.pstack-flex/` config dir. The runner enforces this: a gateway lane refuses to start when its config dir carries an OAuth credentials file. Caveat: on macOS the CLI may store credentials in the Keychain where the file check cannot see them — the rule above is the real defense; the check is a backstop. +- **Privacy: gateway lanes are opt-in per project.** Do not send client or customer code to third-party providers by default. Keep sensitive repositories on subscription lanes, and enable gateway lanes deliberately, per project. +- **No runner fallback.** A failed gateway lane is a named dropout receipt. A reported model mismatch fails the lane. If the endpoint reports no model, the receipt says `modelVerified: false` and `modelEvidence: "pinned-argv"`; this cannot prove which model the gateway served. Confirm supported model slugs during the live probe. + +## Optional lanes + +- **OpenRouter (off by default).** OpenRouter has no Anthropic-format endpoint, so a lane needs a local translator that serves `/v1/messages` — musistudio/claude-code-router or a version-pinned LiteLLM — with `ANTHROPIC_BASE_URL` pointed at it. That is one extra long-running local process, which is why it is documented rather than shipped. Expect roughly a 5.5% credit fee on top of provider list prices. If you build it, model it as another gateway provider in `flex-providers.ts`. +- **Local via Ollama (planned).** Ollama serves an Anthropic-compatible API since v0.14, so a `local` gateway provider pointed at it is the natural next lane: full compute control, zero per-token cost, your hardware. Not wired in yet. + +## Adding a gateway provider + +Any lab that serves an Anthropic-compatible `/v1/messages` endpoint can become a gateway lane. The runner, parser, and preflight branch on `isGatewayProvider`, so no `switch` needs a new case. + +1. Add the provider name to `GATEWAY_PROVIDERS` in `plugins/pstack/skills/poteto-mode/scripts/runner/types.ts`. +2. Add its row to `GATEWAY_SPECS` in `runner/flex-providers.ts`: API key variable, base URL default, override variables, and context-window default. Typecheck fails until this row exists. +3. Add its row to the "Flex model matrix" in `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md`. `model-matrix.test.ts` fails until the key variable and base URL match the spec. +4. Add its probe row to the table in `plugins/pstack/skills/setup-pstack/SKILL.md`, its variables to the gateway environment reference above, and its prices to the price table. +5. Run the live validation checklist below for the new lane before merging. + +## Live validation checklist (before merge or rollout, real keys, never in CI) + +- V1: one DeepSeek probe through the runner (`--provider deepseek --model deepseek-flash --effort high`, read-only). Expect a `complete` receipt with `costUsd: null`; record the `reportedModel` string and confirm the base-URL default against DeepSeek's current guide; confirm `--effort` is accepted end-to-end. +- V2: same for MiniMax (`MiniMax-M3`); record the served-model casing. +- New-model gate: install the exact candidate and run `/setup-pstack` from both real Claude Code and Codex surfaces. Select Flash plus Pro and M3 plus Preview, verify independent efforts and probes, then run a read-only mixed panel. Record installed version/commit, surface, action, requested model/effort, served model, and observed result. Verify a failed preview entitlement probe leaves the sheet unchanged and does not select M3. A fake CLI regression test is not this gate. +- V3: run `claude auth status --json` inside a fresh flex config dir with `ANTHROPIC_AUTH_TOKEN` set and record the output here. On macOS, confirm whether `claude login` under an explicit `CLAUDE_CONFIG_DIR` writes `.credentials.json` or the Keychain. +- V4: the zero-subscription walkthrough above, end to end, on a machine with no stored provider logins. +- V5: OAuth guard live: `claude login` inside a scratch flex config dir, run a lane, confirm the refusal receipt, then delete that login. diff --git a/docs/USAGE.md b/docs/USAGE.md new file mode 100644 index 00000000..5a49334a --- /dev/null +++ b/docs/USAGE.md @@ -0,0 +1,251 @@ +# Using pstack-flex + +The walkthrough: what this plugin is, how work flows through it, how to set it up on the models you actually have, and copy-paste examples for the skills you will use daily. Lane mechanics and pricing live in [LANES.md](LANES.md); the fork's delta over upstream is in [UPSTREAM-FLEX.md](../UPSTREAM-FLEX.md). + +## What this is + +pstack is a plugin of engineering skills, playbooks, and small local tools for coding agents — not a model, not a service. You hand `poteto-mode` a task; it matches the task to a playbook, works the steps, and leaves evidence (diffs, runs, receipts) you can inspect instead of asking for trust. Its sharpest edge is multi-model adversarial review: several different model families challenge important work, because the adversarial signal comes from model diversity, not assigned personas. + +pstack-flex adds one thing on top: **you choose the models and the compute**. Any subset of families works, and two open labs — DeepSeek and MiniMax — are first-class lanes on plain API keys, down to a zero-subscription setup. + +If you also use my [thisguyskills](https://github.com/thisguymartin/skills) collection: that repo decides **what** to build (shaping, spec, Linear, handoff) and its handoff ends with "Use `pstack:poteto-mode`" — which is exactly where this repo picks up. + +## The big picture + +```mermaid +flowchart TD + T([Your task]) --> P["/pstack:poteto-mode"] + P --> PB[Playbook match
feature, bug-fix, refactoring, perf, ...] + PB --> S[Skills fire per step
how, tdd, interrogate, arena, ...] + S --> F{Lane fan-out} + F --> N1["claude:fable / claude:opus
native Agent (Claude sub)"] + F --> N2["codex:gpt-6-astra / gpt-6-sol / gpt-6-luna
native or codex CLI (ChatGPT sub)"] + F --> N3["grok:grok-4.6
grok CLI (Grok sub)"] + F --> G1["deepseek:deepseek-flash
runner + env -> DeepSeek API (key)"] + F --> G2["minimax:MiniMax-M3
runner + env -> MiniMax API (key)"] + N1 --> R[Outputs + receipts] + N2 --> R + N3 --> R + G1 --> R + G2 --> R + R --> V[Verification: run it, judge it,
cross-model consensus] + V --> PR([Review-ready PR]) +``` + +Every lane is a real agent process with tools and file access. The parent harness (your Claude Code or Codex session) resolves the route once; children never pick their own models. + +## Install this fork + +The marketplace keeps upstream's name (`open-pstack`), so only the source changes. + +Claude Code: + +```text +/plugin marketplace add thisguymartin/pstack-flex +/plugin install pstack@open-pstack +/reload-plugins +``` + +Codex: + +```shell +codex plugin marketplace add thisguymartin/pstack-flex --ref main +codex plugin add pstack@open-pstack +``` + +Plus [Bun](https://bun.sh) for the lane runner, and `multi_agent = true` under `[features]` in `~/.codex/config.toml` if Codex is your parent. Sign in only to the CLIs whose subscriptions you actually have — missing families are fine now. + +## Keys for the gateway lanes + +DeepSeek and MiniMax have no login flow here; their lanes read an API key from your environment at spawn time. The runner never writes keys to disk or receipts, so the only question is how the env gets populated. Don't paste keys into `.zshrc` — store them encrypted and load on demand. macOS Keychain, built in and free: + +```zsh +# once: store each key (prompts for the value, nothing in shell history) +security add-generic-password -a "$USER" -s pstack-deepseek -w +security add-generic-password -a "$USER" -s pstack-minimax -w + +# in .zshrc: a function, not an export — keys enter env only when you call it +pstack-keys() { + export DEEPSEEK_API_KEY=$(security find-generic-password -a "$USER" -s pstack-deepseek -w) + export MINIMAX_API_KEY=$(security find-generic-password -a "$USER" -s pstack-minimax -w) +} +``` + +Daily flow: `pstack-keys -> claude -> /pstack:poteto-mode`. Alternatives, the threat model, and the spend-cap advice are in [LANES.md](LANES.md#storing-keys). Set spend caps on both provider dashboards; that is the real blast-radius control. + +## First-time setup: /setup-pstack + +```text +/pstack:setup-pstack +``` + +(Codex: `Use pstack:setup-pstack to configure pstack.`) + +Setup is assignment-first: pick which roles run on which families, answer one effort question per **assigned** family, and only assigned families get probed. Unassigned families are skipped, not errors. Every probe is a real one-turn run — a failed probe writes nothing. Three configurations that make sense: + +**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults. GPT-6 Sol writes code, Luna explores and verifies, and the panel spans three providers: + +```text +feature, refactoring: codex:gpt-6-sol@high +bug-fix: codex:gpt-6-sol@high +how explorer: codex:gpt-6-luna@high +swarm workers: codex:gpt-6-luna@high +arena runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh +``` + +**B. Hybrid saver** (Claude sub + two API keys) — frontier judgment, cheap volume: + +```text +feature, refactoring: deepseek:deepseek-flash@high +bug-fix: deepseek:deepseek-flash@high +judgment and prose: claude:fable@max +hardest tasks: claude:fable@max +swarm workers: deepseek:deepseek-flash@high +arena runners: claude:fable@max, deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +interrogate reviewers: claude:fable@max, deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +``` + +**C. Zero-subscription budget duo** (nothing but two keys) — start your parent session env-pointed at DeepSeek (walkthrough in [LANES.md](LANES.md#zero-subscription-walkthrough)), then assign everything across the two flex families: + +```text +arena runners: deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +interrogate reviewers: deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +``` + +Two labs are two distinct families, so panels keep real diversity without any override. A single-provider panel needs your explicit confirmation — by design. + +## Daily driving: the skills, with examples + +**poteto-mode** — the default entry point for any real task. It stays sticky across turns and pairs well with long autonomous sessions. + +```text +/pstack:poteto-mode + +Take ENG-142: saved reports lose their date-range filter after rename. +Repro is in the issue. Fix it, prove it in the running app, and prep the PR. +``` + +**interrogate** — multi-model review of a decision, design, or diff. Reviewers come from different families; the parent sorts their findings. + +```text +/pstack:interrogate + +Review this migration plan in docs/plans/report-store.md. Attack the +premise, the rollout order, and anything that loses data on rollback. +``` + +```mermaid +flowchart LR + Q[Decision or diff] --> A[Reviewer A
family 1] + Q --> B[Reviewer B
family 2] + Q --> C[Reviewer C
family 3] + A --> S[Parent synthesizes] + B --> S + C --> S + S --> O["consensus (2+ models) -> act on
lone findings -> consider
disagreements -> resolve explicitly"] +``` + +**arena** — N parallel attempts at the same task, an independent cross-judge, then graft the best parts onto a base. + +```text +/pstack:arena + +Implement the rate limiter from the spec in docs/spec.md. Run the +configured arena panel and keep the winner's tests regardless of base. +``` + +```mermaid +flowchart LR + T[Task] --> C1[Candidate 1] + T --> C2[Candidate 2] + T --> C3[Candidate 3] + C1 --> J[Cross-judge
different provider] + C2 --> J + C3 --> J + J --> G[Pick base + graft
best pieces] +``` + +**swarm** — same-shaped work fanned across N workers, one combined report. Good for sweeps: "apply this codemod across packages," "audit every endpoint for X." + +```text +/pstack:swarm + +Audit every handler under src/api/ for missing input validation. +One worker per file group, combined findings ranked by severity. +``` + +**architect** — competing designs from different families, scored by a judge on yet another family, before any code. + +```text +/pstack:architect + +Design the offline sync layer: local-first edits, conflict policy, +and migration from the current always-online store. +``` + +Worth knowing by name: `how` (explain how something works before touching it), `why` (root-cause an incident with your MCP context), `tdd`, `unslop` (de-slop prose and code), `fix-ci`, `babysit` (drive a PR to green). The 23 `principle-*` leaves are loaded by poteto-mode as needed — you rarely invoke them directly. + +## What actually happens on a gateway lane + +No new harness. The same runner that launches Codex and Grok lanes spawns the stock `claude` binary with swapped environment: + +```mermaid +sequenceDiagram + participant P as Parent session + participant R as pstack-runner + participant C as claude -p (subprocess) + participant D as DeepSeek / MiniMax API + P->>R: lane: deepseek:deepseek-flash@high + R->>R: guard: DEEPSEEK_API_KEY set?
config dir free of OAuth creds? + Note over R: refusal = unauthenticated receipt,
no subprocess ever spawned + R->>C: spawn with ANTHROPIC_BASE_URL,
ANTHROPIC_AUTH_TOKEN, isolated CLAUDE_CONFIG_DIR + C->>D: every model request in the agent loop + D-->>C: completions + C-->>R: JSON result + R-->>P: output file + receipt +``` + +The guard order matters: key check and OAuth check happen in-process **before** anything runs, so a claude.ai login can never be pointed at a third-party endpoint. Inherited `ANTHROPIC_*` values from your parent session are stripped before injection. + +## Reading receipts + +Every external lane writes a JSON receipt next to its output. The fields that matter: + +| Field | Meaning | +| --- | --- | +| `status` | `complete`, or a named dropout (`unauthenticated`, `unavailable-cli`, `timed-out`, ...) | +| `modelVerified` + `modelEvidence` | `provider-report` = the endpoint echoed the requested model (case-insensitive for gateways). `pinned-argv` = it didn't, but the argv pinned it — normal for Codex and sometimes gateways | +| `usage` | real token counts — trust these | +| `costUsd` | real for claude/grok subscription lanes; **always `null` on gateway lanes** (the CLI would price at Anthropic rates). Multiply `usage` by the [LANES.md](LANES.md) table instead | + +## Cost playbook + +- High-volume code-writing roles (`feature`, `bug-fix`, `swarm workers`) -> `deepseek:deepseek-flash` — cheapest tokens, near-free cache hits, and half price in the off-peak window. +- Long-context research and big-repo reading -> `minimax:MiniMax-M3` — 1M context. +- `judgment and prose` and `hardest tasks` -> your best frontier lane if you have one; this is the last role to economize. +- Panels: one frontier + two flex lanes gets you three-family diversity at a fraction of three subscriptions. + +## Troubleshooting + +| Symptom | Meaning | Fix | +| --- | --- | --- | +| Receipt `unauthenticated`, exit 77, "KEY is not set" | lane env missing | run your `pstack-keys` function (or export the key) in the shell that starts the parent | +| Receipt `unauthenticated`, "OAuth credentials found" | a claude.ai login sits in the lane's config dir | that's the leak guard working; remove the login from `~/.pstack-flex/` — never `claude login` there | +| Receipt `unauthenticated` after the model ran | the endpoint rejected the key (401) | check the key and the base URL against the provider's current guide | +| Exit 69 `unavailable-cli` | the `claude` binary isn't on PATH for the runner | install it or fix PATH | +| A panel ran with fewer lanes than configured | a lane dropped out with a named receipt | read that receipt; pstack proceeds N-1 and never silently substitutes a model | +| Everything gateway broke after a claude CLI update | Anthropic doesn't support third-party endpoints; compatibility can shift | pin the CLI version on machines that depend on gateway lanes; see [LANES.md](LANES.md#safety-and-policy) | + +## The GPT-6 Codex models + +`astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), and `luna` (`codex:gpt-6-luna@high`) are stock families and the first-run defaults: Astra on every panel, GPT-6 Sol on the solo code-writing roles, Luna on exploration and swarm work. They need only your Codex login. Each gets its own effort question and live probe. See [GPT-6 Codex families](LANES.md#gpt-6-codex-families). + +A sheet written before this release keeps its assignments. To move a role, run `/setup-pstack` and name it; every role you do not change keeps its descriptor, and `codex:gpt-5.6-sol` stays selectable. For example, this row keeps GPT-5.6 Sol on bug fixes while the rest of the sheet takes the new defaults: + +```text +bug-fix: codex:gpt-5.6-sol@max +``` + +## Selecting the additional gateway models + +Run `/setup-pstack` and assign `deepseek-pro` (`deepseek:deepseek-v4-pro@high`) or `minimax-preview` (`minimax:MiniMax-M3.1-Flash-Preview@high`) to named roles. Existing `deepseek` and `minimax` choices remain available. Each model has its own effort selection and live probe. MiniMax preview requires an eligible Token Plan key in `MINIMAX_API_KEY`; see [model choices and thinking controls](LANES.md#multiple-models-per-provider). No existing assignment changes until setup succeeds and you confirm the rendered sheet. diff --git a/docs/gateway-model-probes.md b/docs/gateway-model-probes.md new file mode 100644 index 00000000..ab03514a --- /dev/null +++ b/docs/gateway-model-probes.md @@ -0,0 +1,30 @@ +# Gateway model probe evidence + +Tracking: [issue #5](https://github.com/thisguymartin/pstack-flex/issues/5). + +## Candidate and scope + +- Source candidate: branch `flex/multiple-gateway-models`, based on `92dc0bc`, with uncommitted implementation changes. +- Implementation diff SHA-256 before this evidence file: `84b4c3ee7fc66902532e1d457048a487e6a63629b31d2d214e52585f72863c75`. +- Packaged version: 1.4.1. This candidate has not been installed as a plugin. +- Actual parent: Codex session, invoking the candidate's external runner with `--parent codex`. +- CLI: Claude Code 2.1.283. +- Each probe used `--effort high`, read-only mode, a separate empty synthetic workspace and isolated Claude configuration, and a synthetic text file. No repository or customer data was used in the prompt. +- Keys were supplied through hidden terminal input, injected into child environments, and were not included in commands, this repository, or evidence below. + +## Observed results + +Each probe exited 0, returned the exact requested marker, recorded `modelVerified: true` with `modelEvidence: provider-report`, and retained `costUsd: null`. + +| Requested model | Reported model | Elapsed milliseconds | Receipt status | +| --- | --- | --- | --- | +| `deepseek-flash` | `deepseek-flash` | 2684 | `complete` | +| `deepseek-v4-pro` | `deepseek-v4-pro` | 7737 | `complete` | +| `MiniMax-M3` | `MiniMax-M3` | 11416 | `complete` | +| `MiniMax-M3.1-Flash-Preview` | `MiniMax-M3.1-Flash-Preview` | 5090 | `complete` | + +These single short probes establish authentication, model selection, and successful completion through the runner. They do not rank coding quality or speed, prove hidden reasoning depth, or verify CLI request-body effort forwarding. The prompt included the expected marker, so completion does not independently prove a file tool was used. + +## Remaining release gate + +Install the exact candidate and run setup from both real Claude Code and Codex user surfaces. Verify independent model effort choices, per-model probes, mixed-provider panels, saved-sheet readback, and unchanged configuration on failed access. Record installed version, surface, action, and observed result before merge or rollout. Changing only the runner's `--parent` flag would not satisfy this gate. diff --git a/docs/reference.md b/docs/reference.md index 23015146..5b7ef3de 100644 --- a/docs/reference.md +++ b/docs/reference.md @@ -86,7 +86,7 @@ The Codex build shares one `skills/` tree with the Claude Code build. Nothing is - **Tool and built-in mapping.** Claude tool names and built-in skills resolve through [`codex-tools.md`](../plugins/pstack/skills/poteto-mode/references/codex-tools.md). Model execution resolves separately through [`provider-dispatch.md`](../plugins/pstack/skills/poteto-mode/references/provider-dispatch.md), so Codex can keep Sol native while invoking Claude and Grok externally. - **Subagents.** The `Agent` tool maps to Codex `spawn_agent` / `wait_agent`, enabled by `multi_agent = true`. Parallel fan-out is multiple `spawn_agent` calls in one turn. If the native Codex lane is unavailable, record that lane as a dropout; external Claude and Grok lanes still run, and no provider is silently substituted. There is no `poteto-agent` subagent type on Codex; route ad-hoc subagents by dispatching a `spawn_agent` told to read `poteto-mode` first. - **Auto-fire.** The `hooks/` SessionStart injection is Claude Code-only; Codex has no plugin hook runtime. Enter `pstack:poteto-mode` by name, or add a standing instruction to `~/.codex/AGENTS.md` if you want the same always-on routing. -- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-5.6 Sol max, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. In Codex, Sol uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; Sol and Grok use the runner. Children never detect the parent or reroute themselves. The `bug-fix`, `perf-issue`, and `hillclimb` roles stay on GPT-5.6 Sol max instead of upstream's Fable default because Sol costs less for these frequent delegated code roles. +- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-6 Astra high, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. The GPT-6 Astra, Sol, and Luna Codex families are stock: GPT-6 Sol high carries `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high carries `how explorer` and `swarm workers`; Astra high sits on every panel. GPT-5.6 Sol remains a selectable family. In Codex, every Codex family uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; the Codex families and Grok use the runner. Children never detect the parent or reroute themselves. The solo code roles stay on a Codex model instead of upstream's Fable default because it costs less for these frequent delegated code roles. Verified in fresh installed Claude Code and Codex sessions: the user-facing skills are discovered and namespaced under `pstack`; both parents fan out the frontier quad through the documented native/external route table, retain long-running handles without a default timeout, and cross-judge only after every candidate is terminal. The `principle-*` leaves remain available for `poteto-mode` to read by path. Claude honors their `user-invocable: false` metadata; Codex 0.149.0 does not ([#8](https://github.com/ericlitman/open-pstack/issues/8)). @@ -193,7 +193,7 @@ The port is editorial, not mechanical. Anywhere upstream pstack assumed Cursor-s | Cursor's `/goal` (standing objective across turns) | The program objective written into the run's standing orders and restated in the todolist | | The Cursor agent store (path in the system prompt) | `~/.claude/orchestrate//`, which survives the session restarts a multi-day program expects | | Model rule `~/.cursor/rules/pstack-models.mdc` | Override sheet `~/.claude/pstack-models.md`, included from `CLAUDE.md` | -| Multi-model panels (arena, architect, interrogate) | Provider dispatch restores the upstream frontier quad: `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Same-provider lanes stay native; external lanes use the bundled runner. | +| Multi-model panels (arena, architect, interrogate) | Provider dispatch owns the default panel: `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Same-provider lanes stay native; external lanes use the bundled runner. | ### Cross-vendor dispatch @@ -211,7 +211,7 @@ The earlier port collapsed panels to Claude-only models. The bundled runner rest - **`automations/benny/`** (upstream `0452e08`, the only pstack change between `e46364b` and v0.10.0) — a dormant Slack issue-triage and reproduce-and-fix automation pack built on Cursor's event-triggered automations. It registers no slash skills even upstream, so excluding it changes nothing about the ported plugin's behavior. Porting it would require Cursor's event-trigger runtime, Slack, and tracker plumbing that Open Pstack does not provide. - **`docs/guide/`** (upstream `02c03a9`, `0b7ef5b`, `424829e`) — the ten-chapter usage tutorial and its six screenshots (2.3 MB). It teaches pstack through Cursor's UI, sticky mode, and cloud agents, so a faithful port would be a rewrite rather than a sync, and none of it ships as skill content. Read it upstream at [cursor/plugins/pstack/docs/guide](https://github.com/cursor/plugins/tree/main/pstack/docs/guide); the concepts map through the substitution table above. - **`make-bot-ui`** (upstream `799151d`, relocated by `6fecddb`) uses Cursor routines, webhook events, hosted bot state, and Cursor UI primitives that have no shared Claude Code and Codex mapping. A provider-specific rewrite would be a separate feature, not an upstream sync. -- **Fable solo code defaults** (upstream `23a56e2`) move `bug-fix`, `perf-issue`, and `hillclimb` from GPT-5.6 Sol to Fable. Open Pstack keeps these frequent delegated code roles on `codex:gpt-5.6-sol@max` because Fable costs much more per task. +- **Fable solo code defaults** (upstream `23a56e2`) move `bug-fix`, `perf-issue`, and `hillclimb` from GPT-5.6 Sol to Fable. pstack-flex keeps these frequent delegated code roles on a Codex model (`codex:gpt-6-sol@high`) because Fable costs much more per task. - **Sticky mode** (upstream `#144`) — Cursor-only `mode`/`icon`/`color`/`reminder` frontmatter with no Claude Code equivalent. The port's 0.9.5 SessionStart hook is the analog and already carries the non-trivial / trivial / opt-out logic. - **`is_background: true` on `poteto-agent`** (upstream `99559f2`) — Cursor names this key differently. Claude-native frontier definitions use `background: true`; ad-hoc `poteto-agent` calls remain background dispatches at the call site. - **`cursor-team-kit` beyond the seven imported skills** — the rest either duplicate Claude Code built-ins (`verify-this` → the `verify` skill and built-in verification discipline; `check-compiler-errors` → LSP diagnostics; `control-cli`/`control-ui` → `run`/`verify`, already the substitution targets) or overlap skills this port ships (`loop-on-ci`, `review-and-ship`, `weekly-review` vs `babysit`, `fix-ci`, `make-pr-easy-to-review`, `what-did-i-get-done`). `pr-review-canvas` is Cursor-UI-specific. diff --git a/plugins/pstack/skills/architect/SKILL.md b/plugins/pstack/skills/architect/SKILL.md index f6893207..00031bc7 100644 --- a/plugins/pstack/skills/architect/SKILL.md +++ b/plugins/pstack/skills/architect/SKILL.md @@ -31,7 +31,7 @@ Skip Phase A only when the work is genuinely greenfield with no surrounding syst Run the **arena** skill with the design-sketch task and the Phase A grounding artifacts. Pass `references/runner-prompt.md` as each runner's prompt. Each candidate produces a design package shaped per `references/rationale-template.md`. -Use your configured architect runners (defaults `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`). +Use your configured architect runners (defaults `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`). Design it twice. Require at least two structurally distinct candidates before synthesis, even when the first looks sufficient. This is the **exhaust-the-design-space** principle skill made concrete. Whole-shape alternatives, not point fixes inside one shape. diff --git a/plugins/pstack/skills/arena/SKILL.md b/plugins/pstack/skills/arena/SKILL.md index 0e3b5f38..e03c7ae6 100644 --- a/plugins/pstack/skills/arena/SKILL.md +++ b/plugins/pstack/skills/arena/SKILL.md @@ -26,7 +26,7 @@ The N candidates will receive the same prompt, so the prompt is the contract. 1. State the artifact each candidate is producing. 2. Derive the rubric. State what success looks like for *this* task, then turn it into 3-6 concrete gradeable criteria. The rubric is the picker's tool in Phase D. Candidates only see the task. -3. Pick the runners. Use `arena runners` from the current harness's pstack model sheet when present. Otherwise default to one each on `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Spawn more when the arena covers multiple design directions. Same descriptor N times when the work is generation-bound rather than judgment-sensitive. +3. Pick the runners. Use `arena runners` from the current harness's pstack model sheet when present. Otherwise default to one each on `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Spawn more when the arena covers multiple design directions. Same descriptor N times when the work is generation-bound rather than judgment-sensitive. 4. Assign output paths. Each candidate writes to its own location (a git worktree where possible, otherwise `/tmp/arena-/candidate-/`), per the **separate-before-serializing-shared-state** principle skill. ## Phase B: Fan out diff --git a/plugins/pstack/skills/how/SKILL.md b/plugins/pstack/skills/how/SKILL.md index 84803c9d..7022727d 100644 --- a/plugins/pstack/skills/how/SKILL.md +++ b/plugins/pstack/skills/how/SKILL.md @@ -20,7 +20,7 @@ When in doubt, take the simple path. ## Step 2a. Explore (complex questions only) -Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Start all explorers in one fan-out phase through provider dispatch. Use your configured how-explorer descriptor (default `grok:grok-4.6@xhigh`) in `read-only` mode. A native lane uses the parent subagent primitive; an external lane uses the launcher directly. +Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Start all explorers in one fan-out phase through provider dispatch. Use your configured how-explorer descriptor (default `codex:gpt-6-luna@high`) in `read-only` mode. A native lane uses the parent subagent primitive; an external lane uses the launcher directly. Each explorer gets the prompt in `references/explorer-prompt.md` with its angle filled in. Then go to Step 3. diff --git a/plugins/pstack/skills/interrogate/SKILL.md b/plugins/pstack/skills/interrogate/SKILL.md index ee7514ff..5f388930 100644 --- a/plugins/pstack/skills/interrogate/SKILL.md +++ b/plugins/pstack/skills/interrogate/SKILL.md @@ -39,7 +39,7 @@ Start all reviewers in one fan-out phase. Use `interrogate reviewers` from the c | Subagent | Default model | |----------|---------------| | Reviewer A | `claude:fable@max` | -| Reviewer B | `codex:gpt-5.6-sol@max` | +| Reviewer B | `codex:gpt-6-astra@high` | | Reviewer C | `grok:grok-4.6@xhigh` | | Reviewer D | `claude:opus@xhigh` | diff --git a/plugins/pstack/skills/poteto-mode/SKILL.md b/plugins/pstack/skills/poteto-mode/SKILL.md index 5e51bfc6..67aa271a 100644 --- a/plugins/pstack/skills/poteto-mode/SKILL.md +++ b/plugins/pstack/skills/poteto-mode/SKILL.md @@ -89,7 +89,7 @@ Read the leaf skill in full for any principle you apply. Each entry names when i For `inherit-parent`, `auto`, or an unconfigured native ad-hoc helper, prefer `poteto-agent`. `/poteto-mode` and `poteto-agent` route through the same wrapper. A provider-qualified role instead follows provider dispatch: Claude's shipped frontier agent definitions select the model alias and requested effort, Codex passes both to `spawn_agent`, and external providers run through the deterministic launcher. Routed workflow skills set the task and access mode. Do not override their choices. -**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. Upstream defaults use Grok 4.6 xhigh for feature/refactoring, exploration, and swarm work; GPT-5.6 Sol max for bug fixes, performance work, hillclimbing, and tooling review; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the four-provider frontier panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel. +**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. First-run defaults use GPT-6 Sol high for feature/refactoring, bug fixes, performance work, and hillclimbing; GPT-6 Luna high for exploration and swarm work; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the Fable, GPT-6 Astra, Grok 4.6, Opus panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel. You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md index 18412149..ea0bc5d9 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md @@ -6,7 +6,7 @@ Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspender 1. Reproduce it yourself on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs) (Non-negotiables). Don't hand the repro to the user. A debug or instrumentation protocol that says to ask the user does not override this. You drive the instrumented runtime. Ask the user only with a stated, specific reason the control surface cannot reach the target, and only after driving it as far as it goes. Won't reproduce directly, force it: synthesize the trigger, tighten conditions, or instrument until it fires. 2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Claude Code's `loop` skill. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out. -3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write`, a dedicated worktree, and a specific scope. Review the diff. +3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope. Review the diff. 4. Verify on the same surface. The original repro now passes. "Inconclusive" or wrong-surface is not a pass. Flag it. Unit tests show branch behavior, not bug absence. 5. Stage the commits so the failing repro lands before the fix in git history. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path. Skip it when the test would be expensive, integration-heavy, or unclear. This is the canonical **sequence-verifiable-units** principle skill, the failing test first and the fix on top. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/feature.md b/plugins/pstack/skills/poteto-mode/playbooks/feature.md index 1cc22990..ab11a919 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/feature.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/feature.md @@ -9,7 +9,7 @@ - **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize. - **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants. - **Smallest safe decomposition.** If one worker is best, name why. -4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `grok:grok-4.6@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). Review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. +4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). Review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. 5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it. 6. Rebase into small, ordered commits. Stack follow-ups. Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md index c38089b7..bdde9697 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md @@ -9,7 +9,7 @@ Core discipline: one change, one measurement, keep or revert. Never stack untest 3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. Read it before each attempt. Keep it out of the tree (gitignored). 4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something". 5. Loop, one hypothesis per iteration: - - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill). + - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-6-sol@high`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill). - Measure before and after with the frozen harness, and run the regression gate. - Accept only when the metric moves past noise and the gate stays green. Otherwise revert the change in full. A tweak that "might help" is not kept. - One commit per accepted fix, staging only the files you changed (`git add `, never `-A`). Log the row either way, kept or reverted. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md index e08ac677..0f755d3e 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md @@ -13,7 +13,7 @@ - **Redundancy.** The wait hangs on one slow instance or attempt. Duplicate the work (replicas, hedged requests, speculative execution) and take the fastest result. The trace has to show the wait dominates and the system has headroom. - **Lazy evaluation.** Cost lands on results that are never used or not needed yet (eager init on the boot path, rendering offscreen items). Defer the work until first use. - **Scheduling.** The work must happen, but not during the interactive moment. Move it to where nobody is waiting: idle callbacks, a background warmup after boot, precompute before the user arrives, cleanup after the frame commits. The win is perceived latency, so measure the interactive path, not total work done. -3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace. +3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-6-sol@high`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace. Apply the **sequence-verifiable-units** principle skill, verifying each attempt before trying the next. 4. Parse and compare the artifacts (JSON to sqlite, diff). "Inconclusive" or wrong-surface is not a pass. Flag it. 5. Cite the measurement in the PR. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md index 5106c39a..ef73bce0 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md @@ -8,7 +8,7 @@ If the cleanup reveals a missing feature or a real bug, split it out and ship th 2. Name the structure the code is missing per **principle-model-the-domain**. Boring code stays when the shape is already clear and local. The reshape must delete branches or invalid states, not add indirection. 3. Name the target shape. State what the module layout, types, and call graph should be if built today (**principle-foundational-thinking**, **principle-redesign-from-first-principles**). If the target crosses a function boundary, run the **architect** skill for parallel design exploration of the shape before the move. 4. Subtract before you add. Delete dead code, collapse one-caller wrappers, drop redundant validators, and remove orphan references before introducing the new shape (**principle-subtract-before-you-add**). The smallest change that reaches the target shape ships (**principle-laziness-protocol**). A speculative cleanup that "might help" gets reverted. -5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `grok:grok-4.6@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). Review the diff yourself. +5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). Review the diff yourself. 6. Prove behavior is unchanged on the real artifact, not "it compiles" (**principle-prove-it-works**). For larger reshapes, run an equivalence check: a script that diffs old-vs-new outputs, a recorded baseline replayed against the new code, or a smoke run on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs). Own the verification yourself. Do not trust a delegate's "looks good" summary. 7. Confirm the change is worth keeping. The success measure is reduced reader load (**principle-minimize-reader-load**). If the diff does not lower reader load somewhere, revert it. 8. Rebase into small ordered commits. A subtraction commit, then the reshape, then any follow-on cleanup. Shape them with the **sequence-verifiable-units** principle skill, so each behavior-preserving slice stays green before the next. Run **Opening a PR**. diff --git a/plugins/pstack/skills/poteto-mode/references/codex-tools.md b/plugins/pstack/skills/poteto-mode/references/codex-tools.md index b967458d..1d8621c3 100644 --- a/plugins/pstack/skills/poteto-mode/references/codex-tools.md +++ b/plugins/pstack/skills/poteto-mode/references/codex-tools.md @@ -42,7 +42,7 @@ poteto-mode's Subagents section sets Claude-specific defaults (`subagent_type: " ## Models and providers -Do not replace every configured entry with a Codex model. `/setup-pstack` writes portable descriptors such as `claude:fable@max`, `codex:gpt-5.6-sol@max`, and `grok:grok-4.6@xhigh`. In a Codex parent, only `codex:*` is native. Route Claude and Grok descriptors through the external launcher exactly as `provider-dispatch.md` specifies. The current default panel intentionally keeps four-provider frontier diversity and contains no older GPT or Claude substitute. +Do not replace every configured entry with a Codex model. `/setup-pstack` writes portable descriptors such as `claude:fable@max`, `codex:gpt-5.6-sol@max`, and `grok:grok-4.6@xhigh`. In a Codex parent, only `codex:*` is native. Route Claude and Grok descriptors through the external launcher exactly as `provider-dispatch.md` specifies. The current default panel intentionally keeps four-provider frontier diversity and contains no older GPT or Claude substitute. pstack-flex gateway descriptors (`deepseek:*`, `minimax:*`) also always route through the external launcher in a Codex parent; they are never `spawn_agent` lanes. ## Claude built-in skills pstack references diff --git a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md index 74ab90da..bec620b8 100644 --- a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md +++ b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md @@ -14,11 +14,45 @@ pstack model choices are provider-qualified descriptors: | sol | gpt-5.6-sol-max | codex | gpt-5.6-sol | max | low medium high xhigh max | - | | grok | grok-4.6-fast-xhigh | grok | grok-4.6 | xhigh | low medium high xhigh max | - | | opus | opus | claude | opus | xhigh | low medium high xhigh max | opus | +| astra | - | codex | gpt-6-astra | high | low medium high xhigh max | - | +| sol-6 | - | codex | gpt-6-sol | high | low medium high xhigh max | - | +| luna | - | codex | gpt-6-luna | high | low medium high xhigh max | - | -The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. First-run requested efforts are the Default effort cell of each row. A Claude-native agent stem of `-` means the family has no Claude-native agent. Otherwise the shipped agent name is `pstack--`. +The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. First-run requested efforts are the Default effort cell of each row. A Claude-native agent stem of `-` means the family has no Claude-native agent. Otherwise the shipped agent name is `pstack--`. `-` in Upstream pstack choice means Cursor's pstack has no default for that family; the row is fork-owned. `fable` and `opus` are Claude Code's rolling aliases. Claude resolves each alias to the latest available family revision. A runner receipt keeps the requested alias in `model` and the concrete provider-reported revision in `reportedModel`; verification accepts only a numeric `claude-fable-*` or `claude-opus-*` revision from the matching family. +The `astra`, `sol-6`, and `luna` rows are the GPT-6 Codex families (pstack-flex addition). These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent. Each is its own family with its own requested effort and probe; `sol-6` is independent of `sol`, so an existing GPT-5.6 Sol assignment stays unchanged until setup reassigns the role. All Codex families count as one provider for panel diversity. + +## Default panel + +The first-run panel roles (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) use these four lanes, one per entry, at each family's default effort: + +`claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh` + +This line is the single source for the panel default. `setup-pstack`'s first-run sheet and the `arena`, `architect`, and `interrogate` skills copy it verbatim; the static invariant check fails when they drift. Solo code-writing roles (`feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb`) default to the `sol-6` row; exploration and swarm roles default to the `luna` row. + +## Flex model matrix + +pstack-flex addition. These lanes are additive. A flex lane runs the stock `claude` binary env-pointed at the provider's Anthropic-compatible endpoint, with the provider's own API key and an isolated `CLAUDE_CONFIG_DIR`, so it uses no Anthropic account, no claude.ai login, and no subscription. + +| Family | Provider | Model | Default effort | Selectable efforts | API key variable | Base URL default | +|---|---|---|---|---|---|---| +| deepseek | deepseek | deepseek-flash | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic | +| deepseek-pro | deepseek | deepseek-v4-pro | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic | +| minimax | minimax | MiniMax-M3 | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic | +| minimax-preview | minimax | MiniMax-M3.1-Flash-Preview | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic | + +A family identifies one `(provider, model)` pair, not an entire provider. The existing `deepseek` and `minimax` family names and descriptors remain valid. `deepseek-pro` and `minimax-preview` are additional choices with independent requested efforts. Multiple models from one provider still count as one provider for panel diversity. + +MiniMax preview requires Token Plan access; set `MINIMAX_API_KEY` to the eligible subscription key. A pay-as-you-go key is not proof of preview access. The preview always thinks and supports `low` through `max`; do not disable thinking. M3 thinking is off by default at the API and requires adaptive thinking to enable it; its effort flag does not imply preview-style depth control. Selectable efforts are runner requests, not a claim that every provider applies five distinct reasoning levels. Verify CLI forwarding and model access with live probes. Sources: [MiniMax models](https://platform.minimax.io/docs/guides/models-intro), [MiniMax thinking controls](https://platform.minimax.io/docs/api-reference/text-anthropic-api), [DeepSeek Anthropic compatibility](https://api-docs.deepseek.com/guides/anthropic_api) (checked 2026-09-27). + +Flex lanes have no Claude-native agent stem and always take the external runner in both parents. The base URL is a documented default; override it with `DEEPSEEK_BASE_URL` or `MINIMAX_BASE_URL`, and confirm it against the provider's current Claude Code guide during setup's live probe. The config dir defaults to `~/.pstack-flex/` (override: `PSTACK_FLEX__CONFIG_DIR`). Secrets stay in the environment: nothing in the sheet, the receipts, or this repository carries a key. + +Gateway receipt semantics differ from stock claude lanes in two documented ways. `costUsd` is always `null`: the claude CLI prices `total_cost_usd` at Anthropic rates, which would be fiction for third-party traffic; real prices live in [LANES.md](../../../../../docs/LANES.md), and token usage in the receipt stays accurate. Model verification accepts a case-insensitive matching provider report. A mismatched report fails the lane. When the endpoint reports no model, the receipt uses `modelEvidence: "pinned-argv"` and `modelVerified: false`. + +Panel diversity rule (pstack-flex): `arena runners` and `interrogate reviewers` must span at least two distinct providers. DeepSeek plus MiniMax satisfies it. A single-provider panel is written only after the operator explicitly confirms the reduced diversity during setup, and the setup report records that confirmation. The adversarial signal comes from model diversity, so treat the override as an exception, not a configuration style. + ## Read-time normalization Normalize configured descriptors before matching them to the matrix or choosing a route. If a provider-qualified Claude model starts with `claude-fable-` or `claude-opus-` and its remaining revision contains only digits and hyphens, replace that model component in memory with `fable` or `opus`. Preserve provider, effort, role, and lane order. Use only the normalized descriptor for native dispatch or runner argv. Never pass the versioned predecessor to Claude. @@ -31,10 +65,12 @@ This read-time rule makes an older installed sheet use the latest family revisio The top-level harness resolves the route once. A child receives an assigned provider, model, effort, access mode, prompt, working directory, and output path. A child never detects the harness, chooses a provider, or launches another model. Environment markers may corroborate the top-level harness before fan-out, but nested processes inherit parent markers and must not use them for routing. -| Parent | `claude:*` | `codex:*` | `grok:*` | -|---|---|---|---| -| Claude Code | native `Agent` | external runner | external runner | -| Codex | external runner | native `spawn_agent` | external runner | +| Parent | `claude:*` | `codex:*` | `grok:*` | `deepseek:*` | `minimax:*` | +|---|---|---|---|---|---| +| Claude Code | native `Agent` | external runner | external runner | external runner | external runner | +| Codex | external runner | native `spawn_agent` | external runner | external runner | external runner | + +Flex gateway descriptors are never native, even under a Claude Code parent: the gateway lane must run in its own process with injected endpoint, token, and isolated config dir, which the parent's native `Agent` primitive cannot provide. `inherit-parent` and `auto` remain aliases. They use the parent's current model and effort through its native subagent primitive. In a panel they still consume one lane, but they reduce provider diversity; say so in the synthesis record. @@ -54,7 +90,7 @@ The launcher lives at `skills/poteto-mode/scripts/runner/pstack-runner` under th ```text pstack-runner \ --parent \ - --provider \ + --provider \ --model \ --effort \ --mode \ @@ -67,6 +103,8 @@ pstack-runner \ Pass arguments as an argv array or quote every path. Never interpolate prompt text into a shell command. The launcher preflights the assigned CLI and authentication, invokes the model exactly once, disables recursive agents and ambient skill dispatch where the CLI supports it, restricts the built-in tool surface, and records the exact provider/model/effort flags. External lanes do not receive the parent's MCP surface. Keep MCP-dependent Why and Reflect roles on `inherit-parent` or `auto`. The launcher never falls back. +Gateway lanes (`deepseek`, `minimax`) run three checks before the model executes, all fail-closed. First, in-process: the lane's API key variable must be set, and the lane's isolated `CLAUDE_CONFIG_DIR` must be free of OAuth credentials — a `.credentials.json` carrying a claude.ai login, or one that cannot be parsed, refuses the lane with an `unauthenticated` receipt before any subprocess runs, so a claude.ai credential can never be sent to a third-party endpoint. Second, the spawned preflight is `claude --version`, which proves the binary executes; `claude auth status` is deliberately not used because its behavior under token auth is undocumented. Third, the one-shot invocation is the real authentication and model test; an endpoint authentication error classifies as `unauthenticated` like any other lane. + Grok authentication preflight has one bounded retry. If the first `grok models` result would be classified as unauthenticated, the runner waits five seconds and tries the same preflight once more. A second failure is terminal. The delay and second attempt share the runner's absolute deadline and cancellation latch, and the receipt keeps evidence from both attempts. Model execution is never retried. The parent tool sandbox still governs whether a subscribed child CLI can reach its credentials and network. Run setup's live probe from the actual parent profile. A blocked external CLI is a loud dropout, not a reason to elevate permissions or substitute a model silently. @@ -90,7 +128,7 @@ Success requires all of these: 1. Exit status `0`. 2. Receipt status `complete`. -3. Either `modelVerified: true` with `modelEvidence: "provider-report"`, or a Codex receipt with `reportedModel: null`, `modelVerified: false`, and `modelEvidence: "pinned-argv"`. For Claude's `fable` and `opus` aliases, the concrete provider report must belong to the requested family. Codex 0.149.0 accepts the exact `--model` argument but does not report the served model in its JSONL stream. +3. Either `modelVerified: true` with `modelEvidence: "provider-report"`, or a Codex receipt with `reportedModel: null`, `modelVerified: false`, and `modelEvidence: "pinned-argv"`, or a gateway (`deepseek`/`minimax`) receipt with `modelVerified: false` and `modelEvidence: "pinned-argv"` when the endpoint does not echo the requested slug. For Claude's `fable` and `opus` aliases, the concrete provider report must belong to the requested family. Codex 0.149.0 accepts the exact `--model` argument but does not report the served model in its JSONL stream. Gateway reports match case-insensitively because third-party endpoints are inconsistent about slug casing. 4. A non-empty output file. The receipt also carries elapsed time, token usage when the CLI exposes it, and cost when available. Keep it with the arena or review artifacts so parent-harness comparisons are evidence-based. diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts index 05f79b5a..e76b52dd 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts @@ -40,4 +40,28 @@ describe("runner CLI parsing", () => { "greater than zero" ); }); + + it("accepts gateway providers", () => { + const parsed = parseArgs([ + ...argv().map((value, index, all) => + all[index - 1] === "--provider" + ? "minimax" + : all[index - 1] === "--model" + ? "MiniMax-M3" + : value + ), + ]); + expect(parsed?.provider).toBe("minimax"); + expect(parsed?.model).toBe("MiniMax-M3"); + }); + + it("names the gateway providers in the provider rejection", () => { + expect(() => + parseArgs( + argv().map((value, index, all) => + all[index - 1] === "--provider" ? "gemini" : value + ) + ) + ).toThrow("deepseek, minimax"); + }); }); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts index 4fcce242..eb326f85 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts @@ -13,7 +13,7 @@ import { UsageError, } from "./types.ts"; -const HELP = `Usage: pstack-runner --parent --provider \\ +const HELP = `Usage: pstack-runner --parent --provider <${PROVIDERS.join("|")}> \\ --model --effort --mode \\ --prompt --cwd --output --receipt [--timeout ] diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts index ea697ff2..827a1b66 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts @@ -1,5 +1,5 @@ import { describe, expect, it } from "bun:test"; -import { invocationCommand } from "./commands.ts"; +import { invocationCommand, preflightCommand } from "./commands.ts"; import type { RunnerOptions } from "./types.ts"; function options(overrides: Partial = {}): RunnerOptions { @@ -146,6 +146,30 @@ describe("invocationCommand", () => { ); }); + it("runs gateway lanes with the exact claude argv for the lane's model", () => { + for (const [provider, model] of [ + ["deepseek", "deepseek-flash"], + ["minimax", "MiniMax-M3"], + ] as const) { + const gateway = invocationCommand(options({ provider, model })); + const claude = invocationCommand( + options({ provider: "claude", model }) + ); + expect(gateway.command).toBe("claude"); + expect(gateway.stdin).toBe("prompt"); + expect(gateway.args).toEqual(claude.args); + } + }); + + it("preflights gateway lanes with a version probe, not an auth check", () => { + for (const provider of ["deepseek", "minimax"] as const) { + const spec = preflightCommand(provider); + expect(spec.command).toBe("claude"); + expect(spec.args).toEqual(["--version"]); + expect(spec.stdin).toBe("none"); + } + }); + it("covers low, medium, and high for every external provider", () => { const cases = [ { diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts index 5f2b10c6..e61951d6 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts @@ -4,6 +4,7 @@ import type { Provider, RunnerOptions, } from "./types.ts"; +import { isGatewayProvider } from "./types.ts"; export interface CommandSpec { readonly command: string; @@ -12,6 +13,14 @@ export interface CommandSpec { } export function preflightCommand(provider: Provider): CommandSpec { + if (isGatewayProvider(provider)) { + // Gateway lanes run the claude binary with token auth against a + // third-party endpoint. `claude auth status` semantics under token + // auth are undocumented, so the preflight only proves the binary + // executes; credentials are checked in-process by the gateway guard + // and the one-shot invocation is the real auth test. + return { command: "claude", args: ["--version"], stdin: "none" }; + } switch (provider) { case "claude": return { @@ -63,33 +72,40 @@ function effortOverride(effort: Effort): string { return `model_reasoning_effort=${JSON.stringify(effort)}`; } +function claudeInvocation(options: RunnerOptions): CommandSpec { + return { + command: "claude", + args: [ + "-p", + "--model", + options.model, + "--effort", + options.effort, + "--permission-mode", + permissionMode(options.mode), + "--setting-sources", + "project", + "--strict-mcp-config", + "--tools", + claudeTools(options.mode), + "--no-session-persistence", + "--disable-slash-commands", + "--disallowed-tools", + claudeDeniedTools(options.mode), + "--output-format", + "json", + ], + stdin: "prompt", + }; +} + export function invocationCommand(options: RunnerOptions): CommandSpec { + // Gateway lanes use the same binary and argv as claude; the difference is + // injected environment (endpoint, token, isolated CLAUDE_CONFIG_DIR). + if (isGatewayProvider(options.provider)) return claudeInvocation(options); switch (options.provider) { case "claude": - return { - command: "claude", - args: [ - "-p", - "--model", - options.model, - "--effort", - options.effort, - "--permission-mode", - permissionMode(options.mode), - "--setting-sources", - "project", - "--strict-mcp-config", - "--tools", - claudeTools(options.mode), - "--no-session-persistence", - "--disable-slash-commands", - "--disallowed-tools", - claudeDeniedTools(options.mode), - "--output-format", - "json", - ], - stdin: "prompt", - }; + return claudeInvocation(options); case "codex": return { command: "codex", diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.test.ts new file mode 100644 index 00000000..755b70d3 --- /dev/null +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.test.ts @@ -0,0 +1,183 @@ +import { afterEach, beforeEach, describe, expect, it } from "bun:test"; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { homedir, tmpdir } from "node:os"; +import { join } from "node:path"; +import { + GATEWAY_INHERITED_CONFLICTS, + GATEWAY_SPECS, + gatewayConfigDir, + gatewayEnvironment, + gatewayGuard, +} from "./flex-providers.ts"; +import { GATEWAY_PROVIDERS } from "./types.ts"; + +let scratch = ""; + +beforeEach(() => { + scratch = mkdtempSync(join(tmpdir(), "flex-providers-")); +}); + +afterEach(() => { + rmSync(scratch, { recursive: true, force: true }); +}); + +describe("GATEWAY_SPECS", () => { + it("covers every gateway provider with an https default endpoint", () => { + for (const provider of GATEWAY_PROVIDERS) { + const spec = GATEWAY_SPECS[provider]; + expect(spec.apiKeyVar.length).toBeGreaterThan(0); + expect(spec.baseUrlDefault.startsWith("https://")).toBe(true); + } + }); +}); + +describe("gatewayConfigDir", () => { + it("defaults under the home directory per provider", () => { + expect(gatewayConfigDir("deepseek", {})).toBe( + join(homedir(), ".pstack-flex", "deepseek") + ); + expect(gatewayConfigDir("minimax", {})).toBe( + join(homedir(), ".pstack-flex", "minimax") + ); + }); + + it("honors the override variable and ignores blank overrides", () => { + expect( + gatewayConfigDir("deepseek", { PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: "/opt/lane" }) + ).toBe("/opt/lane"); + expect( + gatewayConfigDir("deepseek", { PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: " " }) + ).toBe(join(homedir(), ".pstack-flex", "deepseek")); + }); +}); + +describe("gatewayEnvironment", () => { + it("injects the full endpoint, token, model, and isolation map", () => { + const source = { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: scratch, + }; + expect(gatewayEnvironment("deepseek", "deepseek-flash", source)).toEqual({ + ANTHROPIC_BASE_URL: "https://api.deepseek.com/anthropic", + ANTHROPIC_AUTH_TOKEN: "sk-test", + ANTHROPIC_MODEL: "deepseek-flash", + ANTHROPIC_DEFAULT_OPUS_MODEL: "deepseek-flash", + ANTHROPIC_DEFAULT_SONNET_MODEL: "deepseek-flash", + ANTHROPIC_DEFAULT_HAIKU_MODEL: "deepseek-flash", + CLAUDE_CODE_SUBAGENT_MODEL: "deepseek-flash", + CLAUDE_CODE_ATTRIBUTION_HEADER: "0", + CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC: "1", + CLAUDE_CODE_MAX_CONTEXT_TOKENS: "128000", + CLAUDE_CONFIG_DIR: scratch, + }); + }); + + it("omits the context cap when the provider has no default", () => { + const env = gatewayEnvironment("minimax", "MiniMax-M3", { + MINIMAX_API_KEY: "mm-test", + PSTACK_FLEX_MINIMAX_CONFIG_DIR: scratch, + }); + expect(env.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined(); + expect(env.ANTHROPIC_BASE_URL).toBe("https://api.minimax.io/anthropic"); + expect(env.ANTHROPIC_MODEL).toBe("MiniMax-M3"); + }); + + it("honors base URL and context overrides", () => { + const env = gatewayEnvironment("deepseek", "deepseek-flash", { + DEEPSEEK_API_KEY: "sk-test", + DEEPSEEK_BASE_URL: "https://proxy.internal/anthropic", + DEEPSEEK_MAX_CONTEXT_TOKENS: "64000", + }); + expect(env.ANTHROPIC_BASE_URL).toBe("https://proxy.internal/anthropic"); + expect(env.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe("64000"); + }); + + it("never leaks a value from a non-token source variable", () => { + const env = gatewayEnvironment("deepseek", "deepseek-flash", { + DEEPSEEK_API_KEY: "sk-secret", + UNRELATED_SECRET: "do-not-copy", + }); + const values = Object.entries(env) + .filter(([key]) => key !== "ANTHROPIC_AUTH_TOKEN") + .map(([, value]) => value); + expect(values).not.toContain("sk-secret"); + expect(values).not.toContain("do-not-copy"); + }); + + it("lists every alternative Claude provider selector as an inherited conflict", () => { + const conflicts = new Set(GATEWAY_INHERITED_CONFLICTS); + for (const key of [ + "CLAUDE_CODE_USE_ANTHROPIC_AWS", + "CLAUDE_CODE_USE_BEDROCK", + "CLAUDE_CODE_USE_FOUNDRY", + "CLAUDE_CODE_USE_MANTLE", + "CLAUDE_CODE_USE_VERTEX", + "CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST", + "CLAUDE_CONFIG_DIR", + ]) { + expect(conflicts.has(key)).toBe(true); + } + }); +}); + +describe("gatewayGuard", () => { + it("refuses when the API key variable is missing or blank", () => { + expect(gatewayGuard("deepseek", {})?.message).toBe("DEEPSEEK_API_KEY is not set"); + expect(gatewayGuard("minimax", { MINIMAX_API_KEY: " " })?.message).toBe( + "MINIMAX_API_KEY is not set" + ); + }); + + it("passes when the config dir does not exist yet", () => { + expect( + gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: join(scratch, "never-created"), + }) + ).toBeNull(); + }); + + it("refuses an OAuth credentials file and cites the path, not the contents", () => { + const dir = join(scratch, "oauth"); + mkdirSync(dir); + const credentials = join(dir, ".credentials.json"); + writeFileSync( + credentials, + JSON.stringify({ claudeAiOauth: { accessToken: "oauth-secret" } }), + { mode: 0o600 } + ); + const refusal = gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: dir, + }); + expect(refusal?.message).toContain("OAuth credentials found"); + expect(refusal?.evidence).toBe(credentials); + expect(refusal?.evidence).not.toContain("oauth-secret"); + expect(refusal?.message).not.toContain("oauth-secret"); + }); + + it("refuses an unparseable credentials file", () => { + const dir = join(scratch, "garbage"); + mkdirSync(dir); + writeFileSync(join(dir, ".credentials.json"), "not json", { mode: 0o600 }); + const refusal = gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: dir, + }); + expect(refusal?.message).toContain("unknown credential state"); + }); + + it("passes a credentials file that carries no OAuth markers", () => { + const dir = join(scratch, "clean"); + mkdirSync(dir); + writeFileSync(join(dir, ".credentials.json"), JSON.stringify({ note: "empty" }), { + mode: 0o600, + }); + expect( + gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: dir, + }) + ).toBeNull(); + }); +}); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts new file mode 100644 index 00000000..1a247860 --- /dev/null +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts @@ -0,0 +1,135 @@ +import { existsSync, readFileSync } from "node:fs"; +import { homedir } from "node:os"; +import { join } from "node:path"; +import type { GatewayProvider } from "./types.ts"; + +// pstack-flex addition. Gateway providers run the stock `claude` binary +// against a third-party Anthropic-compatible endpoint. Everything a lane +// needs is injected as environment at spawn time; secrets come from the +// operator's environment and are never written to disk or receipts. + +export interface GatewaySpec { + readonly apiKeyVar: string; + readonly baseUrlDefault: string; + readonly baseUrlOverrideVar: string; + readonly configDirOverrideVar: string; + readonly maxContextTokensDefault: string | null; + readonly maxContextTokensOverrideVar: string; +} + +export const GATEWAY_SPECS: Record = { + deepseek: { + apiKeyVar: "DEEPSEEK_API_KEY", + baseUrlDefault: "https://api.deepseek.com/anthropic", + baseUrlOverrideVar: "DEEPSEEK_BASE_URL", + configDirOverrideVar: "PSTACK_FLEX_DEEPSEEK_CONFIG_DIR", + maxContextTokensDefault: "128000", + maxContextTokensOverrideVar: "DEEPSEEK_MAX_CONTEXT_TOKENS", + }, + minimax: { + apiKeyVar: "MINIMAX_API_KEY", + baseUrlDefault: "https://api.minimax.io/anthropic", + baseUrlOverrideVar: "MINIMAX_BASE_URL", + configDirOverrideVar: "PSTACK_FLEX_MINIMAX_CONFIG_DIR", + maxContextTokensDefault: null, + maxContextTokensOverrideVar: "MINIMAX_MAX_CONTEXT_TOKENS", + }, +}; + +// Provider selection and Claude configuration from the parent must not +// override the gateway's endpoint, token, or isolated config directory. +export const GATEWAY_INHERITED_CONFLICTS = [ + "CLAUDE_CODE_USE_ANTHROPIC_AWS", + "CLAUDE_CODE_USE_BEDROCK", + "CLAUDE_CODE_USE_FOUNDRY", + "CLAUDE_CODE_USE_MANTLE", + "CLAUDE_CODE_USE_VERTEX", + "CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST", + "CLAUDE_CODE_SUBAGENT_MODEL", + "CLAUDE_CODE_MAX_CONTEXT_TOKENS", + "CLAUDE_CONFIG_DIR", +] as const; + +function overridden(source: NodeJS.ProcessEnv, name: string): string | null { + const value = source[name]; + return value !== undefined && value.trim().length > 0 ? value : null; +} + +export function gatewayConfigDir( + provider: GatewayProvider, + source: NodeJS.ProcessEnv = process.env +): string { + return ( + overridden(source, GATEWAY_SPECS[provider].configDirOverrideVar) ?? + join(homedir(), ".pstack-flex", provider) + ); +} + +export function gatewayEnvironment( + provider: GatewayProvider, + model: string, + source: NodeJS.ProcessEnv = process.env +): NodeJS.ProcessEnv { + const spec = GATEWAY_SPECS[provider]; + const injected: NodeJS.ProcessEnv = { + ANTHROPIC_BASE_URL: overridden(source, spec.baseUrlOverrideVar) ?? spec.baseUrlDefault, + ANTHROPIC_MODEL: model, + ANTHROPIC_DEFAULT_OPUS_MODEL: model, + ANTHROPIC_DEFAULT_SONNET_MODEL: model, + ANTHROPIC_DEFAULT_HAIKU_MODEL: model, + CLAUDE_CODE_SUBAGENT_MODEL: model, + CLAUDE_CODE_ATTRIBUTION_HEADER: "0", + CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC: "1", + CLAUDE_CONFIG_DIR: gatewayConfigDir(provider, source), + }; + const token = overridden(source, spec.apiKeyVar); + if (token !== null) injected.ANTHROPIC_AUTH_TOKEN = token; + const maxContext = + overridden(source, spec.maxContextTokensOverrideVar) ?? spec.maxContextTokensDefault; + if (maxContext !== null) injected.CLAUDE_CODE_MAX_CONTEXT_TOKENS = maxContext; + return injected; +} + +export interface GatewayRefusal { + readonly message: string; + readonly evidence: string; +} + +// Runs in-process before any subprocess is spawned, so no request can leave +// the machine first. Refusals surface as `unauthenticated` receipts. +export function gatewayGuard( + provider: GatewayProvider, + source: NodeJS.ProcessEnv = process.env +): GatewayRefusal | null { + const spec = GATEWAY_SPECS[provider]; + if (overridden(source, spec.apiKeyVar) === null) { + return { + message: `${spec.apiKeyVar} is not set`, + evidence: `gateway lane ${provider} requires ${spec.apiKeyVar} in the environment`, + }; + } + const credentialsPath = join(gatewayConfigDir(provider, source), ".credentials.json"); + if (!existsSync(credentialsPath)) return null; + let raw: unknown; + try { + raw = JSON.parse(readFileSync(credentialsPath, "utf8")); + } catch { + return { + message: + "unreadable credentials file in gateway config dir; refusing to run with unknown credential state", + evidence: credentialsPath, + }; + } + const record = + raw !== null && typeof raw === "object" && !Array.isArray(raw) + ? (raw as Record) + : null; + if (record === null || "claudeAiOauth" in record || "accessToken" in record) { + return { + message: + "OAuth credentials found in gateway config dir; refusing to point a claude.ai login at a third-party endpoint", + evidence: credentialsPath, + }; + } + return null; +} diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts index e0d6df1e..3cedcb2f 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts @@ -1,7 +1,16 @@ import { describe, expect, it } from "bun:test"; import { readdirSync, readFileSync } from "node:fs"; import { join } from "node:path"; -import { EFFORTS, type Effort } from "./types.ts"; +import { parseArgs } from "./cli.ts"; +import { invocationCommand } from "./commands.ts"; +import { GATEWAY_SPECS } from "./flex-providers.ts"; +import { validateOptions } from "./run.ts"; +import { + EFFORTS, + GATEWAY_PROVIDERS, + type Effort, + type GatewayProvider, +} from "./types.ts"; const PLUGIN_ROOT = join(import.meta.dir, "../../../.."); const DISPATCH_PATH = join( @@ -21,7 +30,8 @@ const MATRIX_HEADER = [ "Claude-native agent stem", ] as const; -const FAMILY_ORDER = ["fable", "sol", "grok", "opus"] as const; +const FAMILY_ORDER = ["fable", "sol", "grok", "opus", "astra", "sol-6", "luna"] as const; +const GPT6_FAMILIES = ["astra", "sol-6", "luna"] as const; const PROVIDERS = ["claude", "codex", "grok"] as const; const DESCRIPTOR_RE = /(claude|codex|grok):[a-z0-9.-]+@(low|medium|high|xhigh|max)/g; @@ -51,12 +61,22 @@ const SHEET_ROLES = [ const SETUP_SECTION_ORDER = [ "### 2. Load current state", "### 3. Parse per-family efforts", - "### 4. Collect one requested effort per family", - "### 5. Probe the four requested pairs", + "### 4. Choose role assignments, then collect efforts", + "### 5. Probe the assigned pairs", "### 6. Render, preserving role families", "### 7. Confirm and commit", ] as const; +const FLEX_MATRIX_HEADER = [ + "Family", + "Provider", + "Model", + "Default effort", + "Selectable efforts", + "API key variable", + "Base URL default", +] as const; + interface MatrixRow { family: string; upstreamChoice: string; @@ -89,11 +109,15 @@ function asEffort(value: string): Effort { throw new Error(`not an effort: ${value}`); } -function parseModelMatrix(markdown: string): MatrixRow[] { +function parseModelMatrix( + markdown: string, + heading = "## Model matrix", + rowCount: number = FAMILY_ORDER.length +): MatrixRow[] { const lines = markdown.split(/\r?\n/); - const start = lines.findIndex((line) => line.trim() === "## Model matrix"); + const start = lines.findIndex((line) => line.trim() === heading); if (start < 0) { - throw new Error("missing ## Model matrix"); + throw new Error(`missing ${heading}`); } let end = lines.length; for (let i = start + 1; i < lines.length; i++) { @@ -106,9 +130,9 @@ function parseModelMatrix(markdown: string): MatrixRow[] { .slice(start + 1, end) .map((line) => line.trim()) .filter((line) => line.startsWith("|")); - if (table.length !== 6) { + if (table.length !== rowCount + 2) { throw new Error( - `model matrix must be header, separator, and 4 data rows, got ${table.length}` + `${heading} must be header, separator, and ${rowCount} data rows, got ${table.length}` ); } const header = splitRow(table[0]); @@ -159,10 +183,25 @@ function parseModelMatrix(markdown: string): MatrixRow[] { }); } -function defaultDescriptors(rows: MatrixRow[]): string[] { - return rows.map( - (row) => `${row.provider}:${row.model}@${row.defaultEffort}` - ); +function defaultDescriptor(row: MatrixRow): string { + return `${row.provider}:${row.model}@${row.defaultEffort}`; +} + +function parseDefaultPanel(markdown: string): string[] { + const lines = markdown.split(/\r?\n/); + const start = lines.findIndex((line) => line.trim() === "## Default panel"); + if (start < 0) { + throw new Error("missing ## Default panel"); + } + for (let i = start + 1; i < lines.length; i++) { + if (lines[i].startsWith("## ")) { + break; + } + if (lines[i].startsWith("`")) { + return lines[i].match(DESCRIPTOR_RE) ?? []; + } + } + throw new Error("## Default panel has no descriptor line"); } function parseFrontmatter(text: string): { @@ -198,9 +237,10 @@ function firstRunSheet(setup: string): string { } describe("model matrix", () => { - const rows = parseModelMatrix(readFileSync(DISPATCH_PATH, "utf8")); + const dispatch = readFileSync(DISPATCH_PATH, "utf8"); + const rows = parseModelMatrix(dispatch); const setup = readFileSync(SETUP_PATH, "utf8"); - const quad = defaultDescriptors(rows); + const panel = parseDefaultPanel(dispatch); it("owns the effort universe and first-run defaults", () => { expect([...EFFORTS]).toEqual(["low", "medium", "high", "xhigh", "max"]); @@ -220,6 +260,9 @@ describe("model matrix", () => { ["sol", "max"], ["grok", "xhigh"], ["opus", "xhigh"], + ["astra", "high"], + ["sol-6", "high"], + ["luna", "high"], ]); expect( rows @@ -275,6 +318,105 @@ describe("model matrix", () => { expect(shipped).toEqual([...expected].sort()); }); + it("ships the GPT-6 Codex families as stock rows and puts them in the first-run sheet", () => { + const gpt6Rows = rows.filter((row) => + (GPT6_FAMILIES as readonly string[]).includes(row.family) + ); + expect(gpt6Rows.map((row) => [row.family, row.model])).toEqual([ + ["astra", "gpt-6-astra"], + ["sol-6", "gpt-6-sol"], + ["luna", "gpt-6-luna"], + ]); + const sheet = firstRunSheet(setup); + for (const row of gpt6Rows) { + expect(row.upstreamChoice).toBe("-"); + expect(row.provider).toBe("codex"); + expect(row.defaultEffort).toBe("high"); + expect(row.selectableEfforts).toEqual([...EFFORTS]); + expect(row.claudeNativeAgentStem).toBeNull(); + expect(sheet).toContain(defaultDescriptor(row)); + } + expect(new Set(rows.map((row) => row.family)).size).toBe(rows.length); + expect(new Set(rows.map((row) => `${row.provider}:${row.model}`)).size) + .toBe(rows.length); + // Solo code roles ride the sol-6 row; exploration and swarm ride luna. + const sol6 = defaultDescriptor(rows.find((row) => row.family === "sol-6")!); + const luna = defaultDescriptor(rows.find((row) => row.family === "luna")!); + for (const role of ["feature, refactoring", "bug-fix", "perf-issue", "hillclimb"]) { + expect(sheet).toContain(`${role}: ${sol6}\n`); + } + for (const role of ["how explorer", "swarm workers"]) { + expect(sheet).toContain(`${role}: ${luna}\n`); + } + expect(setup).toContain("Its model matrices (stock and flex)"); + expect(setup).toContain("Read the model matrices, stock and flex."); + expect(setup).toContain("any stock or flex matrix family"); + expect(setup).toContain("Offer every stock family, including Astra, GPT-6 Sol, and Luna, when changing `architect runners`"); + expect(setup).toContain("Read each model, proposed effort, and selectable efforts from its row."); + expect(setup).toContain("outside the stock and flex matrix families"); + expect(setup).toContain( + "| Astra | Astra matrix row + selected effort | external runner | native `spawn_agent` |" + ); + expect(setup).toContain( + "| GPT-6 Sol | sol-6 matrix row + selected effort | external runner | native `spawn_agent` |" + ); + expect(setup).toContain( + "| Luna | Luna matrix row + selected effort | external runner | native `spawn_agent` |" + ); + expect(setup).toContain("each assigned Codex family gets a native `spawn_agent` probe"); + expect(setup).not.toContain("additional matrix"); + expect(dispatch).not.toContain("## Additional model matrix"); + expect(dispatch).toContain( + "These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent." + ); + }); + + it("owns the default panel: four lanes, three providers, matrix default efforts", () => { + expect(panel).toEqual([ + "claude:fable@max", + "codex:gpt-6-astra@high", + "grok:grok-4.6@xhigh", + "claude:opus@xhigh", + ]); + const byDescriptor = new Set(rows.map(defaultDescriptor)); + for (const descriptor of panel) { + expect(byDescriptor.has(descriptor)).toBe(true); + } + const providers = new Set(panel.map((descriptor) => descriptor.split(":")[0])); + expect(providers.size).toBeGreaterThanOrEqual(2); + }); + + it("passes each GPT-6 family's selected model and effort to the existing runner", () => { + for (const row of rows.filter((row) => (GPT6_FAMILIES as readonly string[]).includes(row.family))) { + for (const effort of row.selectableEfforts) { + const options = parseArgs([ + "--parent", "claude", + "--provider", row.provider, + "--model", row.model, + "--effort", effort, + "--mode", "read-only", + "--prompt", DISPATCH_PATH, + "--cwd", PLUGIN_ROOT, + "--output", join(PLUGIN_ROOT, `${row.family}-probe.md`), + "--receipt", join(PLUGIN_ROOT, `${row.family}-probe.json`), + ]); + if (options === null) { + throw new Error("model probe arguments must produce runner options"); + } + validateOptions(options); + expect(options.timeoutMs).toBeNull(); + const command = invocationCommand(options); + expect(command.command).toBe("codex"); + expect(command.args.slice(0, 5)).toEqual([ + "exec", "--model", row.model, + "--config", `model_reasoning_effort="${effort}"`, + ]); + expect(() => validateOptions({ ...options, parent: "codex" })) + .toThrow("provider codex is native to parent codex"); + } + } + }); + it("keeps setup's first-run default panel copy aligned with the matrix", () => { const sheet = firstRunSheet(setup); const roles = sheet @@ -295,7 +437,7 @@ describe("model matrix", () => { } expect(effort).toBe(row.defaultEffort); } - const expectedPanel = quad.join(", "); + const expectedPanel = panel.join(", "); for (const role of PANEL_ROLES) { const line = sheet .split("\n") @@ -317,7 +459,13 @@ describe("model matrix", () => { expect(setup).toContain("Do not invent a precedence rule."); expect(setup).toContain("Do not probe or write while any inconsistency is unresolved."); expect(setup).toContain("A failed probe writes nothing:"); - expect(setup).toContain("Run one probe per family"); + expect(setup).toContain("Run one probe per assigned family"); + expect(setup).toContain("There is no requirement to assign every matrix family."); + expect(setup).toContain("`architect runners` to keep at least two entries"); + expect(setup).toContain("span at least two distinct providers"); + expect(setup).toContain( + "A failed model demands explicit repair or role reassignment before saving." + ); expect(setup).toContain("normalized complete role map from step 2"); expect(setup).toContain("starts with `claude-fable-` or `claude-opus-`"); expect(setup).toContain("preserving the provider, effort, role, and lane order"); @@ -328,6 +476,64 @@ describe("model matrix", () => { expect(setup).toContain(""); }); + it("keeps the flex matrix additive, parseable, and aligned with the runner", () => { + const dispatch = readFileSync(DISPATCH_PATH, "utf8"); + const lines = dispatch.split(/\r?\n/); + const start = lines.findIndex((line) => line.trim() === "## Flex model matrix"); + expect(start).toBeGreaterThan(-1); + let end = lines.length; + for (let i = start + 1; i < lines.length; i++) { + if (lines[i].startsWith("## ")) { + end = i; + break; + } + } + const table = lines + .slice(start + 1, end) + .map((line) => line.trim()) + .filter((line) => line.startsWith("|")); + expect(table.length).toBeGreaterThan(2 + GATEWAY_PROVIDERS.length); + expect(splitRow(table[0]).join("|")).toBe(FLEX_MATRIX_HEADER.join("|")); + expect(isSeparator(splitRow(table[1]))).toBe(true); + const seen = new Set(); + const families = new Set(); + const pairs = new Set(); + for (const line of table.slice(2)) { + const cells = splitRow(line); + expect(cells.length).toBe(FLEX_MATRIX_HEADER.length); + const [family, provider, model, defaultEffortRaw, selectableRaw, keyVar, baseUrl] = + cells; + expect(GATEWAY_PROVIDERS as readonly string[]).toContain(provider); + const gateway = provider as GatewayProvider; + seen.add(gateway); + expect(/^[a-z0-9-]+$/.test(family)).toBe(true); + expect(families.has(family)).toBe(false); + families.add(family); + const pair = `${provider}:${model}`; + expect(pairs.has(pair)).toBe(false); + pairs.add(pair); + expect(/^[A-Za-z0-9.-]+$/.test(model)).toBe(true); + const selectable = selectableRaw.split(/\s+/).map(asEffort); + expect(selectable).toContain(asEffort(defaultEffortRaw)); + expect(keyVar).toBe(GATEWAY_SPECS[gateway].apiKeyVar); + expect(baseUrl).toBe(GATEWAY_SPECS[gateway].baseUrlDefault); + expect(baseUrl.startsWith("https://")).toBe(true); + } + expect([...seen]).toEqual([...GATEWAY_PROVIDERS]); + for (const pair of [ + "deepseek:deepseek-flash", + "deepseek:deepseek-v4-pro", + "minimax:MiniMax-M3", + "minimax:MiniMax-M3.1-Flash-Preview", + ]) expect(pairs.has(pair)).toBe(true); + expect(setup).toContain("Never group efforts or deduplicate probes by provider alone."); + expect(setup).toContain("Different models sharing a provider count as one provider"); + // The stock quad and first-run sheet must not carry flex descriptors: + // upstream's own checks parse descriptors with a lowercase-only, + // three-provider grammar and must never see a flex lane. + expect(firstRunSheet(setup)).not.toMatch(/deepseek:|minimax:/i); + }); + it("binds Claude-native dispatch to the matrix mapping", () => { const dispatch = readFileSync(DISPATCH_PATH, "utf8"); const nativeStart = dispatch.indexOf("## Native lanes"); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts index b4ebc041..270a5863 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts @@ -110,6 +110,38 @@ describe("parseProviderOutput", () => { expect(parsed.reportedModel).toBe("claude-fable-9-9"); }); + it("parses gateway output as claude-shaped JSON with cost forced null", () => { + const parsed = parseProviderOutput( + "minimax", + JSON.stringify({ + result: "GATEWAY_OK", + session_id: "mm-session", + usage: { input_tokens: 12, output_tokens: 5 }, + total_cost_usd: 0.42, + modelUsage: { "minimax-m3": {} }, + }), + "", + "MiniMax-M3" + ); + expect(parsed).toMatchObject({ + text: "GATEWAY_OK", + reportedModel: "minimax-m3", + sessionId: "mm-session", + usage: { inputTokens: 12, outputTokens: 5 }, + costUsd: null, + }); + }); + + it("matches gateway model slugs case-insensitively", () => { + expect(reportedModelMatches("minimax", "MiniMax-M3", "minimax-m3")).toBe(true); + expect(reportedModelMatches("deepseek", "deepseek-flash", "DeepSeek-Flash")).toBe(true); + expect( + reportedModelMatches("deepseek", "deepseek-flash", "deepseek-flash-0731") + ).toBe(true); + expect(reportedModelMatches("minimax", "MiniMax-M3", "some-other-model")).toBe(false); + expect(reportedModelMatches("claude", "MiniMax-M3", "minimax-m3")).toBe(false); + }); + it("matches only concrete Claude revisions from the requested rolling family", () => { expect(reportedModelMatches("claude", "fable", "claude-fable-9-9")).toBe(true); expect(reportedModelMatches("claude", "opus", "claude-opus-9")).toBe(true); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts index 81ed53d4..64d51b75 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts @@ -3,6 +3,7 @@ import type { ParsedOutput, Provider, } from "./types.ts"; +import { isGatewayProvider } from "./types.ts"; import { concreteModelMatchesRollingAlias, isRollingClaudeAlias, @@ -61,7 +62,11 @@ function modelFromUsage( ?? null; } -function parseClaude(stdout: string, requestedModel: string): ParsedOutput { +function parseClaude( + stdout: string, + requestedModel: string, + provider: Provider = "claude" +): ParsedOutput { let raw: unknown; try { raw = JSON.parse(stdout); @@ -77,7 +82,7 @@ function parseClaude(stdout: string, requestedModel: string): ParsedOutput { return { text, - reportedModel: modelFromUsage(value.modelUsage, "claude", requestedModel), + reportedModel: modelFromUsage(value.modelUsage, provider, requestedModel), sessionId: nullableString(value.session_id ?? value.sessionId), usage: normalizedUsage(value.usage), costUsd: finiteNumber(value.total_cost_usd) ?? null, @@ -163,6 +168,13 @@ export function parseProviderOutput( stderr: string, requestedModel: string ): ParsedOutput { + if (isGatewayProvider(provider)) { + // Gateway lanes emit claude-shaped JSON, but the CLI's + // total_cost_usd is computed at Anthropic rates and would be + // fiction for third-party traffic. Token usage stays; cost is null. + const parsed = parseClaude(stdout, requestedModel, provider); + return { ...parsed, costUsd: null }; + } switch (provider) { case "claude": return parseClaude(stdout, requestedModel); @@ -182,6 +194,13 @@ export function reportedModelMatches( if (provider === "claude" && isRollingClaudeAlias(requested)) { return concreteModelMatchesRollingAlias(requested, reported); } + if (isGatewayProvider(provider)) { + // Third-party endpoints are inconsistent about slug casing + // (e.g. MiniMax-M3 vs minimax-m3); compare case-insensitively. + const wanted = requested.toLowerCase(); + const got = reported.toLowerCase(); + return got === wanted || got.startsWith(`${wanted}-`); + } if (reported === requested || reported.startsWith(`${requested}-`)) { return true; } diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts index 20743b52..8c2c0914 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts @@ -25,7 +25,7 @@ import { appendFileSync, existsSync, unlinkSync, writeFileSync } from "node:fs"; const args = process.argv.slice(2); const name = process.argv[1].split("/").at(-1); const isPreflight = - (name === "claude" && args[0] === "auth") || + (name === "claude" && (args[0] === "auth" || args[0] === "--version")) || (name === "codex" && args[0] === "login") || (name === "grok" && args[0] === "models"); const stage = isPreflight ? "preflight" : "model"; @@ -61,6 +61,10 @@ if (name === "claude" && args[0] === "auth") { console.log(JSON.stringify({loggedIn:true})); process.exit(0); } +if (name === "claude" && args[0] === "--version") { + console.log("9.9.9 (fake)"); + process.exit(0); +} if (name === "codex" && args[0] === "login") { console.log("Logged in using ChatGPT"); process.exit(0); @@ -89,11 +93,18 @@ if (name === "grok" && args[0] === "models") { } const modelIndex = args.findIndex((value) => value === "--model"); const model = modelIndex >= 0 ? args[modelIndex + 1] : "unknown"; -const reportedModel = model === "fable" +const reportedModel = process.env.FAKE_REPORT_MODEL ?? (model === "fable" ? "claude-fable-9-9" : model === "opus" ? "claude-opus-9" - : model; + : model); +if (stage === "model" && process.env.FAKE_DUMP_ENV_PATH) { + writeFileSync(process.env.FAKE_DUMP_ENV_PATH, JSON.stringify(process.env)); +} +if (stage === "model" && process.env.FAKE_AUTH_ERROR === "1") { + console.error("API error: authentication_error - invalid api key"); + process.exit(1); +} if (process.env.FAKE_INVALID_MODEL === "1") { console.error("The requested model is not supported with this account."); process.exit(1); @@ -115,7 +126,7 @@ if (stage === "model" && process.env.FAKE_SELF_SIGNAL) { await Bun.sleep(5_000); } if (name === "claude") { - console.log(JSON.stringify({result:"CLAUDE_OK",session_id:"c1",usage:{input_tokens:10,output_tokens:2},total_cost_usd:0.01,modelUsage:{[reportedModel]:{}}})); + console.log(JSON.stringify({result:"CLAUDE_OK",session_id:"c1",usage:{input_tokens:10,output_tokens:2},total_cost_usd:0.01,...(process.env.FAKE_OMIT_MODEL_USAGE === "1" ? {} : {modelUsage:{[reportedModel]:{}}})})); } else if (name === "codex") { console.log(JSON.stringify({type:"thread.started",thread_id:"o1"})); console.log(JSON.stringify({type:"item.completed",item:{type:"agent_message",text:"CODEX_OK"}})); @@ -906,7 +917,237 @@ describe("runLane", () => { }); }); +describe("gateway lanes", () => { + const GATEWAY_TEST_KEYS = [ + "DEEPSEEK_API_KEY", + "MINIMAX_API_KEY", + "PSTACK_FLEX_DEEPSEEK_CONFIG_DIR", + "PSTACK_FLEX_MINIMAX_CONFIG_DIR", + "ANTHROPIC_API_KEY", + "ANTHROPIC_CUSTOM_HEADERS", + "CLAUDE_CODE_USE_BEDROCK", + "FAKE_DUMP_ENV_PATH", + "FAKE_AUTH_ERROR", + "FAKE_REPORT_MODEL", + "FAKE_OMIT_MODEL_USAGE", + ] as const; + + function gatewayOptions( + provider: "deepseek" | "minimax", + suffix: string + ): RunnerOptions { + return { + ...options(provider === "deepseek" ? "claude" : "codex", suffix), + provider, + parent: "claude", + model: provider === "deepseek" ? "deepseek-flash" : "MiniMax-M3", + effort: "high", + }; + } + + beforeEach(() => { + for (const key of GATEWAY_TEST_KEYS) delete process.env[key]; + process.env.DEEPSEEK_API_KEY = "sk-deepseek-test"; + process.env.MINIMAX_API_KEY = "sk-minimax-test"; + process.env.PSTACK_FLEX_DEEPSEEK_CONFIG_DIR = join(scratch, "flex-deepseek"); + process.env.PSTACK_FLEX_MINIMAX_CONFIG_DIR = join(scratch, "flex-minimax"); + }); + + afterEach(() => { + for (const key of GATEWAY_TEST_KEYS) delete process.env[key]; + }); + + it("refuses without spawning anything when the API key is missing", async () => { + delete process.env.DEEPSEEK_API_KEY; + process.env.FAKE_PREFLIGHT_STARTED_PATH = join(scratch, "preflight-started"); + process.env.FAKE_MODEL_STARTED_PATH = join(scratch, "model-started"); + const input = gatewayOptions("deepseek", "missing-key"); + const result = await runLane(input); + expect(result.exitCode).toBe(77); + const written = receipt(input.receiptPath); + expect(written.status).toBe("unauthenticated"); + expect(written.preflight.status).toBe("not-run"); + expect(written.error?.message).toBe("DEEPSEEK_API_KEY is not set"); + expect(existsSync(join(scratch, "preflight-started"))).toBe(false); + expect(existsSync(join(scratch, "model-started"))).toBe(false); + expect(existsSync(input.outputPath)).toBe(false); + }); + + it("refuses to run over an OAuth login without leaking its contents", async () => { + const dir = join(scratch, "flex-deepseek"); + mkdirSync(dir, { recursive: true }); + writeFileSync( + join(dir, ".credentials.json"), + JSON.stringify({ claudeAiOauth: { accessToken: "oauth-secret" } }), + { mode: 0o600 } + ); + const input = gatewayOptions("deepseek", "oauth-refused"); + const result = await runLane(input); + expect(result.exitCode).toBe(77); + const written = receipt(input.receiptPath); + expect(written.status).toBe("unauthenticated"); + expect(written.error?.message).toContain("OAuth credentials found"); + expect(written.error?.evidence).toBe(join(dir, ".credentials.json")); + expect(JSON.stringify(written)).not.toContain("oauth-secret"); + }); + + it("injects the gateway environment and never the parent's Anthropic identity", async () => { + process.env.ANTHROPIC_API_KEY = "parent-anthropic-secret"; + process.env.ANTHROPIC_CUSTOM_HEADERS = "Authorization: Bearer parent-header-secret"; + process.env.CLAUDE_CODE_USE_BEDROCK = "1"; + const dumpPath = join(scratch, "env-dump.json"); + process.env.FAKE_DUMP_ENV_PATH = dumpPath; + const input = gatewayOptions("deepseek", "env-dump"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const child = JSON.parse(readFileSync(dumpPath, "utf8")) as Record; + expect(child.ANTHROPIC_BASE_URL).toBe("https://api.deepseek.com/anthropic"); + expect(child.ANTHROPIC_AUTH_TOKEN).toBe("sk-deepseek-test"); + expect(child.ANTHROPIC_MODEL).toBe("deepseek-flash"); + expect(child.CLAUDE_CODE_SUBAGENT_MODEL).toBe("deepseek-flash"); + expect(child.CLAUDE_CODE_ATTRIBUTION_HEADER).toBe("0"); + expect(child.CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC).toBe("1"); + expect(child.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe("128000"); + expect(child.CLAUDE_CONFIG_DIR).toBe(join(scratch, "flex-deepseek")); + expect(child.ANTHROPIC_API_KEY).toBeUndefined(); + expect(child.ANTHROPIC_CUSTOM_HEADERS).toBeUndefined(); + expect(child.CLAUDE_CODE_USE_BEDROCK).toBeUndefined(); + expect(child.CLAUDECODE).toBeUndefined(); + }); + + it("completes with cost null and a verified provider report on the happy path", async () => { + const input = gatewayOptions("deepseek", "happy"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written.status).toBe("complete"); + expect(written.costUsd).toBeNull(); + expect(written.usage).toMatchObject({ inputTokens: 10, outputTokens: 2 }); + expect(written.modelVerified).toBe(true); + expect(written.modelEvidence).toBe("provider-report"); + expect(written.preflight.status).toBe("passed"); + expect(written.preflight.evidence).toBe( + "claude binary responded; gateway credentials verified in-process" + ); + expect(readFileSync(input.outputPath, "utf8")).toBe("CLAUDE_OK"); + }); + + for (const parent of ["claude", "codex"] as const) { + for (const [provider, model] of [ + ["deepseek", "deepseek-v4-pro"], + ["minimax", "MiniMax-M3.1-Flash-Preview"], + ] as const) { + it(`pins ${model} and effort through the ${parent} parent route`, async () => { + const dumpPath = join(scratch, "new-model-env.json"); + process.env.FAKE_DUMP_ENV_PATH = dumpPath; + process.env.FAKE_REPORT_MODEL = model.toLowerCase(); + const input: RunnerOptions = { + ...gatewayOptions(provider, "new-model"), parent, model, effort: "max", + }; + expect((await runLane(input)).exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written).toMatchObject({ + status: "complete", parent, provider, model, effort: "max", + modelVerified: true, modelEvidence: "provider-report", costUsd: null, + }); + expect(written.argv[written.argv.indexOf("--model") + 1]).toBe(model); + expect(written.argv[written.argv.indexOf("--effort") + 1]).toBe("max"); + const child = JSON.parse(readFileSync(dumpPath, "utf8")); + for (const key of [ + "ANTHROPIC_MODEL", "ANTHROPIC_DEFAULT_OPUS_MODEL", + "ANTHROPIC_DEFAULT_SONNET_MODEL", "ANTHROPIC_DEFAULT_HAIKU_MODEL", + "CLAUDE_CODE_SUBAGENT_MODEL", + ]) expect(child[key]).toBe(model); + }); + + it(`rejects a substituted ${model} in the ${parent} parent route`, async () => { + const input: RunnerOptions = { + ...gatewayOptions(provider, "substituted-model"), parent, model, + }; + process.env.FAKE_REPORT_MODEL = provider === "deepseek" ? "deepseek-flash" : "MiniMax-M3"; + expect((await runLane(input)).exitCode).toBe(65); + expect(receipt(input.receiptPath).status).toBe("malformed-output"); + expect(existsSync(input.outputPath)).toBe(false); + }); + } + } + + it("verifies a case-shifted served model for MiniMax", async () => { + process.env.FAKE_REPORT_MODEL = "minimax-m3"; + const input = gatewayOptions("minimax", "case-shift"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written.modelVerified).toBe(true); + expect(written.modelEvidence).toBe("provider-report"); + expect(written.reportedModel).toBe("minimax-m3"); + }); + + it("fails when the endpoint reports a different model", async () => { + process.env.FAKE_REPORT_MODEL = "unrelated-model"; + const input = gatewayOptions("minimax", "pinned"); + const result = await runLane(input); + expect(result.exitCode).toBe(65); + const written = receipt(input.receiptPath); + expect(written.status).toBe("malformed-output"); + expect(written.modelVerified).toBe(false); + expect(written.modelEvidence).toBeNull(); + expect(written.error?.message).toContain("requested model MiniMax-M3 was not reported"); + expect(existsSync(input.outputPath)).toBe(false); + }); + + it("uses the pinned argv only when the endpoint reports no model", async () => { + process.env.FAKE_OMIT_MODEL_USAGE = "1"; + const input = gatewayOptions("minimax", "unreported-model"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written.status).toBe("complete"); + expect(written.reportedModel).toBeNull(); + expect(written.modelVerified).toBe(false); + expect(written.modelEvidence).toBe("pinned-argv"); + }); + + it("classifies an endpoint authentication error as unauthenticated", async () => { + process.env.FAKE_AUTH_ERROR = "1"; + const input = gatewayOptions("deepseek", "endpoint-401"); + const result = await runLane(input); + expect(result.exitCode).toBe(77); + expect(receipt(input.receiptPath).status).toBe("unauthenticated"); + }); +}); + describe("childEnvironment", () => { + it("strips identity and Anthropic inheritance before gateway injection", () => { + const source = { + PATH: "/bin", + CLAUDECODE: "1", + CODEX_CI: "1", + ANTHROPIC_API_KEY: "parent-secret", + ANTHROPIC_BASE_URL: "https://api.anthropic.com", + ANTHROPIC_CUSTOM_HEADERS: "Authorization: Bearer parent-secret", + ANTHROPIC_BEDROCK_BASE_URL: "https://parent-bedrock.example", + CLAUDE_CODE_USE_BEDROCK: "1", + CLAUDE_CODE_USE_VERTEX: "1", + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: "/tmp/flex-deepseek", + KEEP_ME: "yes", + }; + const env = childEnvironment("deepseek", source, "deepseek-flash"); + expect(env.CLAUDECODE).toBeUndefined(); + expect(env.CODEX_CI).toBeUndefined(); + expect(env.ANTHROPIC_API_KEY).toBeUndefined(); + expect(env.ANTHROPIC_CUSTOM_HEADERS).toBeUndefined(); + expect(env.ANTHROPIC_BEDROCK_BASE_URL).toBeUndefined(); + expect(env.CLAUDE_CODE_USE_BEDROCK).toBeUndefined(); + expect(env.CLAUDE_CODE_USE_VERTEX).toBeUndefined(); + expect(env.ANTHROPIC_BASE_URL).toBe("https://api.deepseek.com/anthropic"); + expect(env.ANTHROPIC_AUTH_TOKEN).toBe("sk-test"); + expect(env.CLAUDE_CONFIG_DIR).toBe("/tmp/flex-deepseek"); + expect(env.KEEP_ME).toBe("yes"); + expect(env.PATH).toBe("/bin"); + }); + it("removes only inherited runtime identity needed to avoid nested detection", () => { const source = { PATH: "/bin", diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts index 054564a4..b3e2ef03 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts @@ -10,6 +10,11 @@ import { } from "node:fs"; import { dirname, resolve } from "node:path"; import { invocationCommand, preflightCommand, type CommandSpec } from "./commands.ts"; +import { + GATEWAY_INHERITED_CONFLICTS, + gatewayEnvironment, + gatewayGuard, +} from "./flex-providers.ts"; import { versionedClaudeAlias } from "./model-aliases.ts"; import { parseProviderOutput, reportedModelMatches } from "./parse-output.ts"; import type { @@ -18,7 +23,7 @@ import type { RunnerOptions, RunnerReceipt, } from "./types.ts"; -import { UsageError } from "./types.ts"; +import { isGatewayProvider, UsageError } from "./types.ts"; const ERROR_EVIDENCE_LIMIT = 4_000; const GROK_PREFLIGHT_RETRY_DELAY_MS = 5_000; @@ -131,7 +136,8 @@ const CLAUDE_IDENTITY = [ export function childEnvironment( provider: Provider, - source: NodeJS.ProcessEnv = process.env + source: NodeJS.ProcessEnv = process.env, + model: string = "" ): NodeJS.ProcessEnv { const result = { ...source }; const remove = provider === "claude" @@ -140,6 +146,13 @@ export function childEnvironment( ? CLAUDE_IDENTITY : [...CODEX_IDENTITY, ...CLAUDE_IDENTITY]; for (const key of remove) delete result[key]; + if (isGatewayProvider(provider)) { + for (const key of Object.keys(result)) { + if (key.startsWith("ANTHROPIC_")) delete result[key]; + } + for (const key of GATEWAY_INHERITED_CONFLICTS) delete result[key]; + Object.assign(result, gatewayEnvironment(provider, model, source)); + } return result; } @@ -352,6 +365,9 @@ async function waitForGrokPreflightRetry( function preflightPassed(provider: Provider, model: string, result: ProcessResult): boolean { if (result.exitCode !== 0 || result.timedOut) return false; + // `claude --version` succeeded; gateway credentials were already verified + // in-process by the gateway guard before any subprocess ran. + if (isGatewayProvider(provider)) return true; const combined = `${result.stdout}\n${result.stderr}`; switch (provider) { case "claude": { @@ -374,6 +390,9 @@ function preflightPassed(provider: Provider, model: string, result: ProcessResul } function successfulPreflightEvidence(provider: Provider, model: string): string { + if (isGatewayProvider(provider)) { + return "claude binary responded; gateway credentials verified in-process"; + } return provider === "grok" ? `authenticated; model ${model} available` : "authenticated"; @@ -396,6 +415,11 @@ function preflightFailureStatus( ): ReceiptStatus { const status = unavailableStatus(value); if (status !== "child-failed") return status; + if (isGatewayProvider(provider)) { + // The gateway preflight is a version probe, not an auth check; a + // failure here means the binary misbehaved, not that auth failed. + return "child-failed"; + } return provider === "grok" && !value.includes(model) ? "unavailable-model" : "unauthenticated"; @@ -457,6 +481,16 @@ function modelProof( modelEvidence: "pinned-argv", }; } + if (isGatewayProvider(provider) && reported === null) { + // Third-party Anthropic-compatible endpoints do not reliably echo the + // requested model slug. A reported mismatch is a failure, since some + // gateways silently substitute a default model for unknown slugs. + return { + reportedModel: null, + modelVerified: false, + modelEvidence: "pinned-argv", + }; + } return { reportedModel: reported, modelVerified: false, @@ -534,7 +568,7 @@ async function executeLane( ): Promise { const startedAt = new Date(started).toISOString(); const prompt = readFileSync(options.promptPath, "utf8"); - const env = childEnvironment(options.provider); + const env = childEnvironment(options.provider, process.env, options.model); const executable = Bun.which(invocation.command, { PATH: env.PATH, cwd: options.cwd, @@ -589,6 +623,37 @@ async function executeLane( return finishWithoutChild("timed-out", "before authentication preflight"); } + if (isGatewayProvider(options.provider)) { + const refusal = gatewayGuard(options.provider); + if (refusal !== null) { + const completed = Date.now(); + receipt = completeReceipt(options, { + status: "unauthenticated", + startedAt, + completedAt: new Date(completed).toISOString(), + elapsedMs: completed - started, + executable, + preflight: preflightState, + argv: [executable ?? invocation.command, ...invocation.args], + exitCode: null, + signal: null, + reportedModel: null, + modelVerified: false, + modelEvidence: null, + sessionId: null, + usage: null, + costUsd: null, + error: { + message: refusal.message, + evidence: refusal.evidence, + }, + }); + removeIfExists(options.outputPath); + writeReceipt(options.receiptPath, receipt); + return { exitCode: statusExitCode("unauthenticated"), receipt }; + } + } + if (executable === null) { const completed = Date.now(); receipt = completeReceipt(options, { diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts index 11c6dfb9..418b2bb6 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts @@ -1,10 +1,19 @@ export const PARENTS = ["claude", "codex"] as const; -export const PROVIDERS = ["claude", "codex", "grok"] as const; +// pstack-flex: gateway providers run the stock `claude` binary against a +// third-party Anthropic-compatible endpoint with injected environment. Adding +// one here requires a matching row in flex-providers.ts GATEWAY_SPECS. +export const GATEWAY_PROVIDERS = ["deepseek", "minimax"] as const; +export const PROVIDERS = ["claude", "codex", "grok", ...GATEWAY_PROVIDERS] as const; export const EFFORTS = ["low", "medium", "high", "xhigh", "max"] as const; export const ACCESS_MODES = ["read-only", "isolated-write"] as const; export type Parent = (typeof PARENTS)[number]; export type Provider = (typeof PROVIDERS)[number]; +export type GatewayProvider = (typeof GATEWAY_PROVIDERS)[number]; + +export function isGatewayProvider(provider: Provider): provider is GatewayProvider { + return (GATEWAY_PROVIDERS as readonly string[]).includes(provider); +} export type Effort = (typeof EFFORTS)[number]; export type AccessMode = (typeof ACCESS_MODES)[number]; diff --git a/plugins/pstack/skills/setup-pstack/SKILL.md b/plugins/pstack/skills/setup-pstack/SKILL.md index 9a0e7441..5290a7a7 100644 --- a/plugins/pstack/skills/setup-pstack/SKILL.md +++ b/plugins/pstack/skills/setup-pstack/SKILL.md @@ -1,11 +1,11 @@ --- name: setup-pstack -description: Configure pstack's provider-qualified models, per-family requested effort, and parent-owned routes per role. Verifies native and external Claude, Codex, and Grok lanes before writing the override sheet. Use for /setup-pstack, "configure pstack models", or changing pstack's model choices. +description: Configure pstack's provider-qualified models, per-family requested effort, and parent-owned routes per role. Verifies native and external Claude, Codex, Grok, DeepSeek, and MiniMax lanes before writing the override sheet. Use for /setup-pstack, "configure pstack models", or changing pstack's model choices. --- # Setup pstack -Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrix, descriptor grammar, and route table are the contract. Choose one requested effort per matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback. +Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrices (stock and flex), descriptor grammar, and route table are the contract. Role assignments are selected first; then choose one requested effort per assigned matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback. Claude Code writes `~/.claude/pstack-models.md` and loads it from `~/.claude/CLAUDE.md` with: @@ -35,28 +35,43 @@ Treat the normalized values as current role-to-family assignments. Overlay those ### 3. Parse per-family efforts -Read the model matrix. Every non-alias value must match `:@`. Map it to exactly one matrix family by `(provider, model)`, require its effort to appear in that row's Selectable efforts cell, and collect the effort. `inherit-parent` and `auto` rows carry no family effort. +Read the model matrices, stock and flex. Every non-alias value must match `:@`. Map it to exactly one matrix family by `(provider, model)`, require its effort to appear in that row's Selectable efforts cell, and collect the effort. `inherit-parent` and `auto` rows carry no family effort. An unmatched provider/model, out-of-domain effort, duplicate role, or unknown role is inconsistent state. Stop, show the conflicting rows verbatim, and ask for an explicit matrix family or alias replacement. If one or more families have mixed efforts, show every conflicting family and role row, then ask for one normalized effort per family from its Selectable efforts cell. Do not invent a precedence rule. Do not probe or write while any inconsistency is unresolved. +A family is a single `(provider, model)` matrix row. DeepSeek Flash and Pro have independent efforts, as do MiniMax M3 and M3.1 Flash Preview. Never group efforts or deduplicate probes by provider alone. + One distinct effort per family is the current value. A family with no non-alias occurrence is unassigned; use its matrix Default effort as the proposed value and label it unassigned rather than calling it current. -### 4. Collect one requested effort per family +### 4. Choose role assignments, then collect efforts + +Role assignments come first. Show the current role-to-family map (loaded and normalized from step 2, or the first-run map from step 7) and ask whether to keep it or change named roles. Keeping it is the default. A changed role may use any stock or flex matrix family, `inherit-parent`, or `auto`. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments. + +Offer every stock family, including Astra, GPT-6 Sol, and Luna, when changing `architect runners` or another configurable role. Read each model, proposed effort, and selectable efforts from its row. The Codex families are separate families even though they share the Codex provider; changing one family's effort does not change another's. GPT-6 Sol uses the `sol-6` family; the `sol` family keeps GPT-5.6 Sol for sheets that still assign it. -Ask exactly four effort questions, one each for Fable, Sol, Grok, and Opus. Name each model, its current or proposed value, and the Selectable efforts from its matrix row. Empty input keeps a current value or accepts the matrix proposal for an unassigned family. On a first run, state the four matrix defaults before asking. On a rerun, state the four parsed values without offering to reset customized role lanes. +The assigned families are exactly the matrix families that appear in the resulting role map. An unassigned family gets no effort question and no probe. There is no requirement to assign every matrix family. -### 5. Probe the four requested pairs +Then ask one effort question per assigned family. Name each model, its current or proposed value, and the Selectable efforts from its matrix row. Empty input keeps a current value or accepts the matrix proposal for a newly assigned family. On a first run, state the assigned families' matrix defaults before asking. On a rerun, state the parsed values without re-opening the role choices already made above. -Probe only the four selected `provider:model@effort` pairs. Run one probe per family, even when two families share a provider. Do not enumerate or offer older models as substitutes. A failed probe writes nothing: report the failing pair and provider, stop, and keep the active sheet plus parent integration bytes unchanged. A failed first run creates neither artifact. +### 5. Probe the assigned pairs + +Probe only the assigned families' selected `provider:model@effort` pairs. Run one probe per assigned family, even when two families share a provider. Do not enumerate or offer older models as substitutes. A failed probe writes nothing: report the failing pair and provider, stop, and keep the active sheet plus parent integration bytes unchanged. A failed model demands explicit repair or role reassignment before saving. A failed first run creates neither artifact. | Family | Pair source | Claude parent route | Codex parent route | Availability proof | |---|---|---|---|---| | Fable | Fable matrix row + selected effort | native Agent `pstack-fable-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe | | Sol | Sol matrix row + selected effort | `codex exec` | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| Astra | Astra matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| GPT-6 Sol | sol-6 matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| Luna | Luna matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | | Grok | Grok matrix row + selected effort | Grok CLI | Grok CLI | `grok models` must list the requested model; one-turn probe | | Opus | Opus matrix row + selected effort | native Agent `pstack-opus-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe | +| DeepSeek Flash / Pro | Each assigned DeepSeek flex row + selected effort | external runner | external runner | `DEEPSEEK_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | +| MiniMax M3 / M3.1 Flash Preview | Each assigned MiniMax flex row + selected effort | external runner | external runner | `MINIMAX_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | + +For MiniMax M3.1 Flash Preview, disclose the Token Plan requirement before probing. Use the eligible subscription key through `MINIMAX_API_KEY`; do not assume a working M3 key grants preview access. A failed preview probe must not silently select M3. Keep preview thinking enabled and verify requested effort forwarding; distinguish request evidence from hidden applied reasoning depth. -Use a tiny read-only probe that returns a unique marker. A login-status command alone proves credentials, not that the requested model and effort flags run. Record native and external results separately. Never call the external launcher for the parent's own provider. On a Claude parent, the Fable and Opus probes are one-turn runs of the mapped `pstack--` agent. On a Codex parent, the Sol probe is native `spawn_agent` with the selected `reasoning_effort`. Every other pair uses the external runner with the selected effort flag. +Use a tiny read-only probe that returns a unique marker. A login-status command alone proves credentials, not that the requested model and effort flags run. Record native and external results separately. Never call the external launcher for the parent's own provider. On a Claude parent, the Fable and Opus probes are one-turn runs of the mapped `pstack--` agent. On a Codex parent, each assigned Codex family gets a native `spawn_agent` probe with its matrix model and selected `reasoning_effort`. Every other pair, flex families always included, uses the external runner with the selected effort flag. A flex probe doubles as the base-URL confirmation: it proves the documented default (or the operator's override) actually serves the lane's model. Receipts and native transcripts prove the requested effort and the route. They do not prove a provider's hidden applied reasoning depth. There is no implicit timeout, weaker-model fallback, same-provider external fallback, or second mutable configuration source. @@ -67,11 +82,13 @@ Build the new sheet in memory. Do not write it yet. - First run: start from the complete role assignments in step 7. - Rerun: start from the normalized complete role map from step 2, preserving each loaded row's lane order and family (or alias) per lane. -After effort selection, ask whether to keep those role-to-family assignments or change named roles. Keeping them is the default. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments. A changed role may use one of the four probed matrix families, `inherit-parent`, or `auto`. +The role assignments were already chosen in step 4; do not re-open them here. Require every documented role to remain present and non-empty, `architect runners` to keep at least two entries, and the final role map to contain at least one assigned matrix family. There is no requirement to assign every matrix family. The sheet stores effort only in role descriptors, so an unassigned family's selection cannot persist without adding a second source of truth. + +Different models sharing a provider count as one provider, even when their efforts differ. -Require the final role map to contain at least one descriptor from each matrix family. The sheet stores effort only in role descriptors, so an unassigned family's selection cannot persist without adding a second source of truth. +Validate panel diversity: `arena runners` and `interrogate reviewers` must span at least two distinct providers. A single-provider panel is written only after the operator explicitly confirms the reduced diversity; record that confirmation in the setup report. -Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model other than the four matrix families, or a provider/model mismatch. +Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model outside the stock and flex matrix families, or a provider/model mismatch. ### 7. Confirm and commit @@ -88,33 +105,33 @@ After the operator confirms, write the in-memory render from step 6. Never paste Provider-qualified per-role choices. Read the installed pstack provider-dispatch reference before dispatching a configured role. Every documented role remains present. `inherit-parent` and `auto` use the parent model natively and still count as one panel lane. -feature, refactoring: grok:grok-4.6@xhigh -bug-fix: codex:gpt-5.6-sol@max -perf-issue: codex:gpt-5.6-sol@max -hillclimb: codex:gpt-5.6-sol@max +feature, refactoring: codex:gpt-6-sol@high +bug-fix: codex:gpt-6-sol@high +perf-issue: codex:gpt-6-sol@high +hillclimb: codex:gpt-6-sol@high judgment and prose: claude:fable@max hardest tasks: claude:fable@max -how explorer: grok:grok-4.6@xhigh +how explorer: codex:gpt-6-luna@high how explainer: claude:fable@max why investigators, synthesizer: inherit-parent reflect tooling, judgment, divergent, synthesizer: inherit-parent -arena runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh -arena cross-judge pool: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh -swarm workers: grok:grok-4.6@xhigh -architect runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh -interrogate reviewers: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh +arena runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh +arena cross-judge pool: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh +swarm workers: codex:gpt-6-luna@high +architect runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh +interrogate reviewers: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh ``` ### 8. Wire it in Render the parent integration in memory before either write. On Claude, the integration is the single `@~/.claude/pstack-models.md` include in `~/.claude/CLAUDE.md`. On Codex, it is the exact sheet bytes between one `` and `` pair in `~/.codex/AGENTS.md`. Replace that whole bounded block on a rerun. Insert one block at the end on first run. If either marker is missing, duplicated, or reversed, stop and report inconsistent state instead of guessing a boundary. -Snapshot every target's current bytes. Write the sheet and parent integration only after all four probes pass and the operator confirms. Read both targets back and compare them with the in-memory render. If either write or readback fails, restore every snapshot and report the failure. An unchanged rerun must produce byte-identical sheet and integration content after normalization. +Snapshot every target's current bytes. Write the sheet and parent integration only after every assigned family's probe passes and the operator confirms. Read both targets back and compare them with the in-memory render. If either write or readback fails, restore every snapshot and report the failure. An unchanged rerun must produce byte-identical sheet and integration content after normalization. Do not copy the model sheet between harnesses without rerunning the parent-specific probes; route availability can differ even on the same host. ### 9. Behavioral smoke -Before declaring setup complete, run one small read-only mixed panel from this parent: all four chosen descriptors, distinct output/receipt paths, and an independent cross-judge. Launch Claude-native agents and every external process in the background with retained handles, then drain them. Verify the native transcript entries and every external receipt. A structural config check or unit test is not a substitute. +Before declaring setup complete, run one small read-only mixed panel from this parent: every assigned family's chosen descriptor, distinct output/receipt paths, and an independent cross-judge when at least two providers are assigned. Launch Claude-native agents and every external process in the background with retained handles, then drain them. Verify the native transcript entries and every external receipt. A structural config check or unit test is not a substitute. Report the sheet path, parent route table, requested-effort probe results, smoke results, and external elapsed/token/cost receipts. Re-running this skill re-probes and updates the same sheet. Do not claim the provider exposed hidden applied-effort observability. diff --git a/plugins/pstack/skills/swarm/SKILL.md b/plugins/pstack/skills/swarm/SKILL.md index 92c8f48a..d010532d 100644 --- a/plugins/pstack/skills/swarm/SKILL.md +++ b/plugins/pstack/skills/swarm/SKILL.md @@ -23,7 +23,7 @@ Open a todolist with one entry per phase before launching anything. 1. State the done predicate and the artifact or report the swarm must return. 2. Choose the shape. Partition into slices, race N workers on identical briefs, or mix both. For a race or mixed shape, declare `first pass`, `rank all`, or `best-of` before spawning. 3. Set N from the user or derive it from the shape. N is total workers, not the number that run at once. -4. Pick the worker descriptor from `swarm workers` in the current harness's pstack model sheet when present. Otherwise use `grok:grok-4.6@xhigh`. For a model race, name each arm's descriptor up front. +4. Pick the worker descriptor from `swarm workers` in the current harness's pstack model sheet when present. Otherwise use `codex:gpt-6-luna@high`. For a model race, name each arm's descriptor up front. 5. Give each worker its own writable output when it writes. ## Phase B: Fan out diff --git a/tests/skill-collision-repro.sh b/tests/skill-collision-repro.sh index 22db0d8e..99bb7320 100755 --- a/tests/skill-collision-repro.sh +++ b/tests/skill-collision-repro.sh @@ -73,32 +73,15 @@ else fi # Static invariant (CHANGES maintenance note): provider-dispatch owns the default -# provider/model quad and the three panel skills plus setup-pstack copy it verbatim. +# panel quad (its "## Default panel" line) and the three panel skills plus setup-pstack copy it verbatim. setup="$repo/plugins/pstack/skills/setup-pstack/SKILL.md" dispatch="$repo/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md" quad_of() { { grep -oE '(claude|codex|grok):[a-z0-9.-]+@(low|medium|high|xhigh|max)' || true; } | tr '\n' ' ' | sed 's/ $//'; } canon_quad="$(awk ' - $0 == "## Model matrix" { in_matrix = 1; next } - in_matrix && /^## / { exit } - in_matrix && /^\|/ { - line = $0 - sub(/^\|/, "", line) - sub(/\|$/, "", line) - n = split(line, cells, "|") - for (i = 1; i <= n; i++) { - gsub(/^ +| +$/, "", cells[i]) - gsub(/`/, "", cells[i]) - } - family = cells[1] - if (family == "Family" || family ~ /^:?-+:?$/) next - provider = cells[3] - model = cells[4] - effort = cells[5] - if (out != "") out = out " " - out = out provider ":" model "@" effort - } - END { print out } -' "$dispatch")" + $0 == "## Default panel" { in_panel = 1; next } + in_panel && /^## / { exit } + in_panel && /^`/ { print; exit } +' "$dispatch" | quad_of)" quad_bad="" [ -n "$canon_quad" ] || quad_bad="could not read the canonical quad from $dispatch"$'\n' # Anchor on the quad's last slug rather than a hard-coded one, so a model swap in @@ -350,14 +333,14 @@ else fi sol_descriptor="$(awk -F '|' ' - $2 ~ /^[[:space:]]*sol[[:space:]]*$/ { + $2 ~ /^[[:space:]]*sol-6[[:space:]]*$/ { for (i = 4; i <= 6; i++) gsub(/^[[:space:]]+|[[:space:]]+$/, "", $i) print $4 ":" $5 "@" $6 } ' "$dispatch")" solo_code_bad="" if [ -z "$sol_descriptor" ]; then - solo_code_bad="could not read the sol row from $dispatch"$'\n' + solo_code_bad="could not read the sol-6 row from $dispatch"$'\n' fi for role in bug-fix perf-issue hillclimb; do setup_descriptor="$(sed -n "s/^${role}: //p" "$setup")" @@ -371,11 +354,11 @@ for role in bug-fix perf-issue hillclimb; do fi done if [ -n "$solo_code_bad" ]; then - note "FAIL: solo code roles must use the sol row:" + note "FAIL: solo code roles must use the sol-6 row:" note "$solo_code_bad" fail=1 else - note "ok: solo code roles stay on the sol row ($sol_descriptor)" + note "ok: solo code roles stay on the sol-6 row ($sol_descriptor)" fi codex_manifest="$plugin/.codex-plugin/plugin.json"