Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
"name": "pstack",
"source": "./plugins/pstack",
"description": "if you want to go fast, go deep first. pstack helps you write less, but higher quality code. rigorous agent workflows you can parallelize with confidence.",
"version": "1.4.1",
"version": "1.5.0",
"author": {
"name": "Lauren Tan (original)"
},
Expand Down
8 changes: 7 additions & 1 deletion CHANGES.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,12 @@

This port applies the Cursor → Claude Code substitutions in skill bodies. Earlier drafts left them flagged; this revision resolves them. A later pass added a Codex build that shares the same skills; see [Codex port](#codex-port) below.

## 1.5.0 syncs to Cursor pstack 0.15.5

Open Pstack 1.5.0 tracks Cursor pstack 0.15.5 at `12d587dfb20741cafc376c42c696c5f6e2a64487`. The first-run panel is now three models: `claude:opus@max`, `codex:gpt-5.6-sol@max`, and `grok:grok-4.7@xhigh`. Opus max takes judgment and prose, hardest tasks, and the How explainer. Grok 4.7 xhigh takes feature and refactoring work, the How explorer, and Swarm workers. Bug-fix, perf-issue, and hillclimb stay on Sol max. Why and Reflect stay on `inherit-parent`. Fable stays in the model matrix with its native agents, but no first-run role uses it. Setup asks about roles first, then asks efforts for and probes only the assigned families (the ordering follows PR #73 by @arjitj2). It drops and lists retired-role rows such as `how critics`.

Imported: the operator-neutral wording, tick status only on change, code-ready rounds with two or more audit lanes and one fix-forward, rebases at the code-ready report and at merge prep, `children.tsv`, the owner Babysit exception, countersign as approval, Architect reading `architect runners`, Swarm SHA and method briefs, Shipping build-noise lane reuse, decision-trail `start` rows and supersede-only corrections, the non-truncating `log.sh`, and the #419 cuts. Not applied: the five exclusions listed in `UPSTREAM.md`, the Cursor manifest, and `docs/guide/`. Existing sheets are not rewritten. Delete role lines and rerun setup to take the new defaults. A `grok:grok-4.6` row keeps running until setup asks for its replacement. With no sheet, the new defaults apply on update.

## 1.4.1 syncs to Cursor pstack 0.15.1

Open Pstack 1.4.1 tracks Cursor pstack 0.15.1 at `f8abeddd1862dc73704e3d719dd73df0d51b8c71`. Poteto-mode now requires each claim to include its evidence or a measured, inferred, or guess label in the same sentence. Agents also run any check they can run themselves instead of handing that check to the user. No playbook, model, runtime, or dependency changed.
Expand Down Expand Up @@ -213,7 +219,7 @@ pstack diverges from superpowers in one respect, and it is deliberate. superpowe

**Verified.** Codex discovers the skills and namespaces them under `pstack` (`pstack:poteto-mode` and so on) in a live session. Mapping resolution mid-task and `spawn_agent` fan-out follow the `superpowers` pattern and are worth confirming per session.

**Maintenance.** The open-pstack version string lives in `plugins/pstack/.claude-plugin/plugin.json`, `.claude-plugin/marketplace.json`, `plugins/pstack/.codex-plugin/plugin.json`, and the current-version row in `UPSTREAM.md`. A version bump must update all four. `tests/skill-collision-repro.sh` checks that they match. `.agents/plugins/marketplace.json` carries no version field. The canonical default panel quad is the model matrix in `provider-dispatch.md` (`provider:model@default` in family-row order). It is copied into the three panel skills (`arena`, `architect`, `interrogate`) and the `setup-pstack` first-run sheet. Keep those copies grep-identical when models change. The static test derives the quad from the matrix. After a sync that touches `skills/poteto-mode/scripts/`, run `bun install --frozen-lockfile`, `bun run test`, and `bun run typecheck` from that directory. `hooks/session-start-context.md` restates skill one-liners. Re-verify it whenever skill names or descriptions change. The package must not contain a `commands/` layer. Claude Code and Codex load the native `skills/` tree directly, and a command layer duplicates that inventory. The 23 `principle-*` leaves carry `user-invocable: false` to request exclusion from the user picker while `poteto-mode` reads them by path. Claude honors the metadata; Codex 0.149.0 currently does not ([#8](https://github.com/ericlitman/open-pstack/issues/8)). They must not carry `disable-model-invocation`, which would make them unreachable to the model. Re-run the behavioral mode of `tests/skill-collision-repro.sh` after Claude Code upgrades to check both model-initiated and user-initiated native skill invocation.
**Maintenance.** The open-pstack version string lives in `plugins/pstack/.claude-plugin/plugin.json`, `.claude-plugin/marketplace.json`, `plugins/pstack/.codex-plugin/plugin.json`, and the current-version row in `UPSTREAM.md`. A version bump must update all four. `tests/skill-collision-repro.sh` checks that they match. `.agents/plugins/marketplace.json` carries no version field. The canonical default panel is the `arena runners` row of the setup-pstack first-run sheet. It is copied into the other three panel rows and into `arena`, `architect`, and `interrogate`. Keep those copies grep-identical when models change. The static test reads the panel from that row, and `model-matrix.test.ts` pins its families and checks each effort against the matrix. After a sync that touches `skills/poteto-mode/scripts/`, run `bun install --frozen-lockfile`, `bun run test`, and `bun run typecheck` from that directory. `hooks/session-start-context.md` restates skill one-liners. Re-verify it whenever skill names or descriptions change. The package must not contain a `commands/` layer. Claude Code and Codex load the native `skills/` tree directly, and a command layer duplicates that inventory. The 23 `principle-*` leaves carry `user-invocable: false` to request exclusion from the user picker while `poteto-mode` reads them by path. Claude honors the metadata; Codex 0.149.0 currently does not ([#8](https://github.com/ericlitman/open-pstack/issues/8)). They must not carry `disable-model-invocation`, which would make them unreachable to the model. Re-run the behavioral mode of `tests/skill-collision-repro.sh` after Claude Code upgrades to check both model-initiated and user-initiated native skill invocation.

## 0.9.2 sync (against upstream `e46364b`)

Expand Down
1 change: 1 addition & 0 deletions NOTICE.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ This plugin is a port of upstream MIT-licensed work. All upstream copyright noti
| `plugins/pstack/skills/poteto-mode/playbooks/{shipping,babysit,autopilot-full,autopilot-stack,opening-a-pr,multi-phase-plan}.md`, `plugins/pstack/skills/poteto-mode/references/bugbot-triage.md`, `plugins/pstack/skills/poteto-mode/SKILL.md`, `plugins/pstack/skills/typescript-best-practices/{SKILL.md,references/patterns.md}`, `plugins/pstack/assets/logo.png` (v0.14.6 and v0.14.7 changes) | [cursor/plugins/pstack @ efa2a53](https://github.com/cursor/plugins/tree/efa2a531985e0a8084d36ff3cf87233be8a9f34b/pstack) | (c) 2026 Lauren Tan | MIT | [LICENSE](LICENSE) |
| `plugins/pstack/skills/` (0.15.0 prose changes and the new `principle-attack-the-premise` and `principle-test-behavior-not-implementation` leaves), `plugins/pstack/assets/logo.png`, `README-UPSTREAM.md` | [cursor/plugins/pstack @ 71ed0d1](https://github.com/cursor/plugins/tree/71ed0d1076fec562c1b74ee353121a8d00f75382/pstack) | (c) 2026 Lauren Tan | MIT | [LICENSE](LICENSE) |
| `plugins/pstack/skills/poteto-mode/SKILL.md` (0.15.1 reply-writing evidence rule) | [cursor/plugins/pstack @ f8abedd](https://github.com/cursor/plugins/tree/f8abeddd1862dc73704e3d719dd73df0d51b8c71/pstack) | (c) 2026 Lauren Tan | MIT | [LICENSE](LICENSE) |
| `plugins/pstack/skills/` (0.15.2 to 0.15.5 changes: three-model defaults, code-ready rounds, owner authority, prompt cuts, decision-trail runs, `show-me-your-work/scripts/log.sh`), `README-UPSTREAM.md` | [cursor/plugins/pstack @ 12d587d](https://github.com/cursor/plugins/tree/12d587dfb20741cafc376c42c696c5f6e2a64487/pstack) | (c) 2026 Lauren Tan | MIT | [LICENSE](LICENSE) |

## What changed in the port

Expand Down
8 changes: 5 additions & 3 deletions README-UPSTREAM.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,12 +22,12 @@ fork it. improve it. make it yours. PRs are welcome!

two steps:

1. run [`/setup-pstack`](./skills/setup-pstack/SKILL.md) and choose which models you want.
1. run [`/setup-pstack`](./skills/setup-pstack/SKILL.md), pick a reasoning budget, and choose which models you want.
2. use [`/poteto-mode`](./skills/poteto-mode/SKILL.md) whenever you're doing anything that requires rigor.

new here? the [pstack guide](./docs/guide/README.md) walks you through a first real task, from setup and prompting through verification and overnight runs.

that's it. the other skills are situational; the mode skill uses them for you as needed. out of the box the mode splits work by model strength: precisely-specified code, prose, and judgment go to fable 5.1, while fast mechanical code goes to grok. the default panel is fable 5.1 / sol / grok / opus 5. [`/setup-pstack`](./skills/setup-pstack/SKILL.md) changes any of it.
that's it. the other skills are situational; the mode skill uses them for you as needed. out of the box the mode splits work by model strength: code delegates (feature, refactoring, bug fix, perf, hillclimb) go to grok, while the hardest changes, prose, and judgment go to opus 5.5. the default panel is opus 5.5 / sol / grok. [`/setup-pstack`](./skills/setup-pstack/SKILL.md) changes any of it.

## usage

Expand Down Expand Up @@ -68,7 +68,7 @@ morning.
| [shipping](./skills/poteto-mode/playbooks/shipping.md) | independently verify a green stack, then land the contiguous verified run bottom-up through github by default or origin when available. |
| [autonomous run](./skills/poteto-mode/playbooks/autonomous-run.md) | drive a long task to completion without stopping. |
| [orchestrate](./skills/poteto-mode/playbooks/orchestrate.md) | a standing project handed to one coordinator chat: multi-day, many stacked prs, fleets of subagents. |
| [autopilot-full](./skills/poteto-mode/playbooks/autopilot-full.md) | run independent prs to merged with one owner per pr and root verification of each merge-ready head. |
| [autopilot-full](./skills/poteto-mode/playbooks/autopilot-full.md) | run independent prs to merged with one owner per pr and a root swarm verdict on each round, from the code-ready head on. |
| [autopilot-stack](./skills/poteto-mode/playbooks/autopilot-stack.md) | build and verify one linear base-branch stack for the operator to review and land. |
| [session pickup](./skills/poteto-mode/playbooks/session-pickup.md) | resume or take over a prior agent's in-flight work. |
| [pause safely](./skills/poteto-mode/playbooks/pause-safely.md) | suspend in-flight work cleanly so it can be resumed later. |
Expand Down Expand Up @@ -248,6 +248,8 @@ type [`/automate-me`](./skills/automate-me/SKILL.md). it mines your recent trans

models are configurable too. type [`/setup-pstack`](./skills/setup-pstack/SKILL.md). it detects the models you have access to and writes a small always-applied rule mapping each role (code, judgment, the review panels) to a model. every skill reads it and falls back to sensible defaults when the rule is absent, so you override only what you want.

a rule written before 0.15.3 pins the old default models. delete those role lines, or delete the file, then run `/setup-pstack` again. a rerun keeps any role whose model differs from the default.

## automations

pstack also ships a dormant [benny automation pack](./automations/benny/). benny triages slack issue reports, then reproduces and fixes confirmed bugs with real ui evidence. its files are not registered as slash skills.
Expand Down
10 changes: 6 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ pstack does not ask you to trust an agent on day one. It helps the agent leave e

## Install

You need a current Claude Code or Codex installation. For the full four-model review, install and sign in to the Claude Code, Codex, and Grok command-line tools. [Bun](https://bun.sh) runs the small local tool that starts models outside the app you are using. You can still use the core workflows with fewer models.
You need a current Claude Code or Codex installation. For the full three-model review, install and sign in to the Claude Code, Codex, and Grok command-line tools. [Bun](https://bun.sh) runs the small local tool that starts models outside the app you are using. You can still use the core workflows with fewer models.

### Claude Code

Expand Down Expand Up @@ -80,10 +80,12 @@ In Codex, ask:
Use pstack:setup-pstack to configure pstack.
```

Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-5.6 Sol, Grok 4.6, and Opus.
Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The default review panel uses Opus, GPT-5.6 Sol, and Grok 4.7, and setup probes only the models your roles use. Fable remains available.

An older model sheet starts using the rolling aliases in memory as soon as this release is installed. Run setup once after updating to persist that migration. It replaces versioned Fable and Opus entries while preserving every role assignment and effort selection.

A model sheet from an earlier release keeps its panel. To take the new defaults, delete those role lines and run setup again; setup fills missing roles from the defaults. A `grok:grok-4.6` entry keeps running until the next setup run asks you to replace it.

### 2. Use poteto-mode

Start any task that needs careful engineering with `poteto-mode`.
Expand Down Expand Up @@ -124,7 +126,7 @@ Plugin skills include `pstack:` in their name. In Claude Code, invoke a native s

Some pstack workflows use one model. Skills such as `architect`, `arena`, and `interrogate` can run several models in parallel. Each model run uses the subscription and token allowance of its own command-line tool.

`setup-pstack` lets you choose the models, one requested effort per model family, and how many run in parallel. A model from the app you are using runs inside that app. Other models run through their own command-line tools. Open Pstack does not quietly replace a failed model with a weaker one.
`setup-pstack` lets you choose the models, one requested effort per model family you assign, and how many run in parallel. A model from the app you are using runs inside that app. Other models run through their own command-line tools. Open Pstack does not quietly replace a failed model with a weaker one.

## Claude Code and Codex

Expand Down Expand Up @@ -153,7 +155,7 @@ This repository also keeps:

## Staying close to Lauren's pstack

Open Pstack 1.4.1 tracks pstack 0.15.1 at Cursor commit [`f8abeddd1862dc73704e3d719dd73df0d51b8c71`](https://github.com/cursor/plugins/commit/f8abeddd1862dc73704e3d719dd73df0d51b8c71).
Open Pstack 1.5.0 tracks pstack 0.15.5 at Cursor commit [`12d587dfb20741cafc376c42c696c5f6e2a64487`](https://github.com/cursor/plugins/commit/12d587dfb20741cafc376c42c696c5f6e2a64487).

The two projects have separate version numbers. The pstack version identifies Lauren's upstream content. The Open Pstack version identifies the Claude Code and Codex package built from it.

Expand Down
18 changes: 11 additions & 7 deletions UPSTREAM.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,17 +8,21 @@ open-pstack tracks [Cursor's pstack](https://github.com/cursor/plugins/tree/main
| --- | --- |
| Repository | `https://github.com/cursor/plugins.git` |
| Path | `pstack/` |
| Commit | `f8abeddd1862dc73704e3d719dd73df0d51b8c71` |
| Upstream version | `0.15.1` |
| open-pstack version | `1.4.1` |
| Commit | `12d587dfb20741cafc376c42c696c5f6e2a64487` |
| Upstream version | `0.15.5` |
| open-pstack version | `1.5.0` |

The table above is the current Cursor sync point. Open Pstack 1.4.1 imports this 0.15.1 sync. `README-UPSTREAM.md` preserves the upstream pstack README verbatim. `CHANGES.md` and `NOTICE.md` describe the adaptations and provenance.
The table above is the current Cursor sync point. Open Pstack 1.5.0 imports this 0.15.5 sync. `README-UPSTREAM.md` preserves the upstream pstack README verbatim. `CHANGES.md` and `NOTICE.md` describe the adaptations and provenance.

## Upstream-only exclusions

- Commits `799151d` and `6fecddb` add and relocate `make-bot-ui`. It depends on Cursor routines, webhook events, and UI primitives that Claude Code and Codex do not share.
- Four `disable-model-invocation: true` lines from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`. Poteto-mode invokes those skills by name, and the flag blocks that route on Claude Code.
- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on `codex:gpt-5.6-sol@max` for cost.
- The default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` from `23a56e2`, `889ec4b`, and `70b2dc8` are not applied. Those frequent code-writing roles stay on `codex:gpt-5.6-sol@max`.
- `5bf2b15`'s setup budget question, its `# budget` line, and its step down to a lower detected effort are not applied. Setup already asks one requested effort per assigned family, and the step-down would silently lower a requested effort.
- `12d587d`'s rule that reruns a rejected configured entry on its family default or the closest valid slug is not applied. An unavailable model stays a named dropout per `provider-dispatch.md`.
- The expected-runtime column in `70b2dc8`'s `children.tsv` and its expected-runtime stuck test are not applied. A lane is stuck only on affirmative failure evidence.
- The explicit Grok, Opus, and Sol defaults for the Why and Reflect roles are not applied. Those roles stay on `inherit-parent` because the external runner omits the parent's MCP servers.
- The Claude manifest does not take the logo field from `efa2a53` because Claude Code has no schema for it. The shared asset is exposed through the Codex manifest instead.

## Check for changes
Expand All @@ -33,8 +37,8 @@ Fetch and inspect only commits that touched pstack after the recorded sync point

```shell
git fetch cursor main
git log --oneline f8abeddd1862dc73704e3d719dd73df0d51b8c71..cursor/main -- pstack
git diff --stat f8abeddd1862dc73704e3d719dd73df0d51b8c71..cursor/main -- pstack
git log --oneline 12d587dfb20741cafc376c42c696c5f6e2a64487..cursor/main -- pstack
git diff --stat 12d587dfb20741cafc376c42c696c5f6e2a64487..cursor/main -- pstack
```

No output means the tracked pstack tree has not changed. This comparison does not need a polling service or generated mirror branch.
Expand Down
Loading
Loading