release: sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 - #98
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…milies Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
| 1. **Resolve the forge, repository identity, and both Git remotes.** GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, use `origin pr ...` for PR view, watch, and edit operations; use it for merge only when step 5's expected-head requirement is available. Otherwise stay on `gh` and record the fallback. Record the forge-reported base repository as canonical `<base-repo>`. On GitHub, split it into `<base-owner>` and `<base-name>`, pass `--repo "$base_repo"` to every `gh pr` command, and pass the two components to the GitHub watcher. Resolve `<head-remote>` from the branch's configured push remote, `remote.pushDefault`, branch remote, or sole unambiguous remote, in that order. Confirm its push URL names the PR head repository, and record that exact URL as `<head-url>` so the captured remote state and the guarded push address the same repository. Resolve `<base-remote>` independently by matching a fetch URL to `<base-repo>`, and confirm that URL before using it for trunk. The same remote name may fill both roles when its push URL matches the head repository and its fetch URL matches the base repository. Compare each URL by role instead of assuming the remote name identifies one repository. Capture every resolved or forge-reported value directly into a shell variable. Command examples use quoted lower-case variables such as `"$branch"` and `"$head_url"`. Never paste those values into shell source. Never guess or treat the forge name as a Git remote, and never require Graphite (`gt`). | ||
| 2. **Freeze and disarm the queue before verification.** Freeze an explicit bottom-to-top PR list. Confirm a same-repository stack against its base-branch chain. For fork heads, take the order from the verified local parent ancestry because every PR targets trunk and the forge bases do not encode the stack. Before launching any verifier, inspect every PR in the frozen list through the active forge, disarm every pre-existing merge-when-ready or auto-merge request, and confirm each request is off. On GitHub, query each PR through GraphQL for `id`, `headRefOid`, `baseRefName`, `autoMergeRequest`, and `mergeQueueEntry`. When `autoMergeRequest` is non-null, run `gh pr merge "$pr" --disable-auto --repo "$base_repo"`. When `mergeQueueEntry` is non-null, invoke the `dequeuePullRequest` mutation with `gh api graphql -F "id=$pr_node_id" -f query='mutation($id:ID!){dequeuePullRequest(input:{id:$id}){mergeQueueEntry{id}}}'`. The mutation takes the pull request node ID. Re-query and require that both `autoMergeRequest` and `mergeQueueEntry` are null. A null `autoMergeRequest` alone does not prove that the pull request is unarmed. On Origin, use its reported cancel operation and inspect every separately reported queue state. Stop if the active forge cannot confirm the whole list is unarmed. One subagent per PR, not batched, each in its own worktree, exercises the real surface against that PR's parent versus head. The bottom PR's patch base is trunk. Each child's patch base is the preceding PR's exact head, including when a fork child targets trunk at the forge. Each subagent returns `PASS`, `PASS+NOTES` or `FAIL` and posts that verdict on its own PR. Safe means a verdict from an agent that did not write the code. CI green is not a verdict, and an approving bot review is not a verdict. Walk up from the bottom and stop at the first PR without a passing verdict, where both `PASS` and `PASS+NOTES` pass. Report that ceiling and what breaks the chain. | ||
| 3. **Re-check that each verdict still describes the patch.** At the passing verdict, record `<verdict-sha>`, `<verdict-base-sha>`, and the stable `<verdict-patch-id>` of `<verdict-base-sha>..<verdict-sha>`. Initialize `<landing-sha>` and `<landing-base-sha>` to those head and base values. Do not replace a fork child's patch base with its forge base, because that forge base is trunk and includes its unmerged parents. Before landing, read the current published head through the active forge as `<current-head>` and require the local branch tip to match it. Calculate the stable patch ID of `<landing-base-sha>..<current-head>` and compare it with `<verdict-patch-id>`. Re-verify when the patch changed. When it did not, keep the code verdict, set `<landing-sha>` to `<current-head>`, and re-run mergeability and CI at that head. A patch-equivalent rebase can make `<landing-sha>` differ from `<verdict-sha>`. Keep both values. Never use matching commit messages or a green check from an older SHA as a substitute. | ||
| 3. **Re-check that each verdict still describes the patch.** At the passing verdict, record `<verdict-sha>`, `<verdict-base-sha>`, and the stable `<verdict-patch-id>` of `<verdict-base-sha>..<verdict-sha>`. Initialize `<landing-sha>` and `<landing-base-sha>` to those head and base values. Do not replace a fork child's patch base with its forge base, because that forge base is trunk and includes its unmerged parents. Before landing, read the current published head through the active forge as `<current-head>` and require the local branch tip to match it. Calculate the stable patch ID of `<landing-base-sha>..<current-head>` and compare it with `<verdict-patch-id>`. When the two patches differ only in tests, docs, or lint config, build what each lane ran. Build it twice at `<verdict-sha>` and once at `<current-head>`. A difference is noise if the two builds at `<verdict-sha>` also show it, or if it is an embedded commit SHA. Judge each difference, not each file, and report each kind of noise with its files. If only noise differs, that lane's result stays valid. Set `<landing-sha>` to `<current-head>`, and run checks and a review of the change fresh at that head. Do not reuse a lane result from a dev server or from anything else with no build output. Rerun that lane. Re-verify anything else when the patch changed. When it did not, keep the code verdict, set `<landing-sha>` to `<current-head>`, and re-run mergeability and CI at that head. A patch-equivalent rebase can make `<landing-sha>` differ from `<verdict-sha>`. Keep both values. Never use matching commit messages or a green check from an older SHA as a substitute. |
There was a problem hiding this comment.
Changed tests keep old verdicts
When a PR changes only tests, docs, or lint settings, Shipping can keep a test lane’s earlier pass if its build output differs only in ways it calls noise. A test-only edit can change what that lane checks without changing the built app. Fresh checks and a review do not replace the lane’s independent verdict, so rerun a lane when its inputs or checks change.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Not changing this. The rule builds what each lane ran. A test-only edit changes what a test lane runs, so its build differs beyond noise and the lane reruns. A lane with no build output is never reused. Only lanes whose inputs did not change keep their result. This text is upstream 12d587d's rule, taken verbatim per #88.
There was a problem hiding this comment.
You're right. The lane-specific build comparison means a test-only change changes the inputs for the test lane, so its build differs beyond the permitted noise and that lane is rerun. Lanes whose inputs and checks are unchanged can retain their results, while lanes without build output are never reused. I withdraw the finding; no change is needed.
| if [ ! -s "$logfile" ]; then | ||
| printf 'ts\tphase\tdecision\twhy\tevidence\tresult\n' >> "$logfile" | ||
| fi |
There was a problem hiding this comment.
Concurrent runs duplicate log headers
log.sh checks whether a log is empty before appending its header, but those are separate actions. If two runs start the same empty log together, both can append a header. Make header creation safe for concurrent first writes so the decision trail stays readable.
There was a problem hiding this comment.
Not changing this. log.sh is byte-identical to upstream 12d587d, taken verbatim per #88. Its comment accepts one stray header line as the cost, never lost rows. It replaces main's check-then-truncate version, which could drop rows. A concurrent first write to one log is outside the skill's usage (one log per effort).
There was a problem hiding this comment.
That context makes sense. Given the one-log-per-effort usage invariant, concurrent first writes are outside the supported usage, and the upstream change intentionally prioritizes avoiding row loss over eliminating a possible duplicate header. Since this is a verbatim upstream import and the comment explicitly documents that trade-off, I’m withdrawing this finding; no change is needed.
|
Queued — the merge queue status continues in this comment ↓. |
✅ Accepted riskAccepted by ericlitman: PSTACK-3 (authority 48fa77b) 🟡 Medium · Fence surviving writers before replacement 🟡 Medium · Standalone TDD no longer runs adjacent validation Advisories (1)
DetailsRun: @unfret review · @unfret status · @unfret help |
|
@unfret review |
1 similar comment
|
@unfret review |
There was a problem hiding this comment.
🛑 2 blocking findings
🟡 Medium · Fence surviving writers before replacement
🟡 Medium · Standalone TDD no longer runs adjacent validation
Check
| 7. **Stand down instantly on the operator's stop.** Her hold or stand-down reaches every owner as a zero-writes order immediately. Owners hold their briefs until she releases them. | ||
| 4. **Swarm-verify every round before its merge.** A round starts at the owner's code-ready head SHA and at each later push that changes the PR's patch. At that SHA, fan out parallel independent verifiers per the **swarm** skill and aggregate to one verdict. The merge needs a clean verdict from the round whose patch matches the merge-ready head. Audit the receipts in the merge-ready report before the verdict. The lanes: re-run the gates at that SHA. Prove the load-bearing behavior live on the real surface the change touches (the `run` skill for CLIs and TUIs, `verify` for UIs, or a named driver where neither fits). Audit the diff, distrusting the PR body. Run the audit as two or more review lanes with the full brief. Give each lane one main focus, such as consumer parity with trunk, lifetimes and races, or data and config safety. **Regression lane against trunk.** Run the same load-bearing scenario on current trunk. If trunk does not have the feature, record that fact and gate the behavior the diff adds plus the end state the user waits for instead of pretending trunk can produce it. The live lane is the floor, and a verdict without it is not clean. No merge without the root's clean verdict. When the lanes return, send every proven finding against the PR to the owner in one fix-forward. A defect that a lane filed as a note is a finding. For each behavior finding, ask for a red test that covers every site with the same defect. Where no test can show the defect, ask for a repro receipt instead. Add that defect to the next round's review brief. The new head gets a fresh swarm and a fresh verdict, except for lane results that stay valid under the patch-id rule in `playbooks/shipping.md`. | ||
| 5. **On a clean verdict the owner merges and takes the next item.** The owner merges only from a head freshly rebased onto trunk. Merge prep never comes before a round's lanes start, and it ends with a rebase onto current trunk right before the merge. After the merge-prep rebase, the owner reports the new head SHA. CI must pass on that head before the merge, and the patch-id rule decides whether the round's verdict still holds. A private-stack child reaches that state through Shipping step 4 after its parent merges. The merge-ready report records the verdict SHA and the current landing SHA. If trunk moves again before the merge, the patch ID rule in `playbooks/shipping.md` governs re-verification. A changed patch needs a new verdict. An unchanged patch keeps the verdict, but Shipping step 3 records the new `<landing-sha>` after current CI and mergeability pass. The owner lands its own PR only through Shipping step 5's server-enforced expected-head flow bound to `<landing-sha>`; never use a PR-number-only merge or unguarded auto-merge. The owner then picks up its next self-contained item from the queue. The operator's full-autonomy grant plus the root's clean verdict is the merge authorization that babysitting alone never has. Operator-named items stop at merge-ready and wait for the operator's click. | ||
| 6. **Run the root layer.** A genuinely new raise of a pinned gate or budget value (a limit CI only lets tighten) needs your fresh countersign, granted only after verifier proof. If the operator's grant or standing orders cover approvals, that countersign is the approval. The owner records it in the form that the tool's approval contract allows, with a pointer to the root's countersign. A lane checks the record against that countersign. The root never gives or bypasses an approval that the forge enforces. Absorbing values that already landed on main is drift, not a raise. Run an audit tick over all owners roughly every 30 minutes. The tick is an observation cadence, never a lease, deadline, or cancellation threshold. Arm each tick as a real `/loop` in dynamic mode, which schedules its own wake-up rather than blocking on a sleep. Never leave the cadence to memory or lossy completion notifications. At each tick, re-read this playbook from disk (`skills/poteto-mode/playbooks/autopilot-full.md` under the installed plugin), then re-read the standing objective. Audit the operation against both. Fix drift during that tick and treat it as urgent. Probe each owner through its retained handle with a generic liveness or status check, and collect the decision trails. Count commits, pushes, PR or check deltas, store reports, and a live retained process as evidence. Elapsed time or the absence of a new side effect alone never proves a lane is stuck; implementation runs can remain healthy for 90 minutes or much longer. Stand a lane down only on affirmative failure evidence such as a dead process, failed handle, explicit error, or a caller-supplied external deadline. Each tick also runs this stuck test over the parent's background task list and every owner's `children.tsv`. Cancel a stuck lane through its retained handle. Whether or not the cancel succeeds, the root has the owner record it as stuck in `children.tsv` and, if its work is still needed, replace it. Each replacement that fails gets the same steps. The root takes both steps when the owner cannot. A stall never proves or drops the work. When merges batch, run a retro pass and a post-merge bot-comment sweep. End the tick only when no delegated work is left, even after the last merge. |
There was a problem hiding this comment.
🟡 Medium · Fence surviving writers before replacement
If the runner is killed while its provider CLI remains alive, cancellation through the dead retained handle fails. This is reproducible with the shipped launcher, but step 6 now requires replacement even when cancellation fails. Reassigning the same PR leaves the original writer able to publish alongside its replacement. Separate worktrees do not isolate the shared PR branch, so late pushes can reject the replacement's push or change the head under verification. Confirm termination or fence the original writer's publication rights before replacement. Also reported for this defect: - Replacement can race a writer whose cancellation failed (plugins/pstack/skills/poteto-mode/playbooks/autopilot-full.md:10): Step 6 requires replacement even when cancellation fails. If a caller-supplied deadline expires while a writer remains alive and its cancellation fails, the replacement can write or push the same PR concurrently, causing conflicting edits, lease failures, or a blocked delivery run.
| ## If a Failing Test Is Impractical | ||
|
|
||
| Do not silently skip the regression step. Before fixing, explicitly explain why a failing test is impossible or not worth the cost, then choose the closest executable regression check available. Examples include a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check. | ||
| Use the closest executable regression check instead: a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check. |
There was a problem hiding this comment.
🟡 Medium · Standalone TDD no longer runs adjacent validation
This hunk deletes 7. Run nearby validation..., while the replacement requires only the focused regression check. When /tdd fixes a shared helper, that test can pass while an adjacent test, typecheck, or scenario fails; the workflow then reports a successful fix with a regression still present.
|
Unfret waiver for head ac4b5cf. Tracked by PSTACK-3 and #88. Eric authorized this waiver on 2026-10-01.
The advisory |
|
@Mergifyio queue |
Merge Queue Status
This pull request spent 18 seconds in the queue, including 2 seconds running CI. Required conditions to merge
|
Closes #88
Closes #72
Linear: PSTACK-3
What changed
Imports Cursor pstack 0.15.2 to 0.15.5 (
f8abedd..12d587d) and releases it as Open Pstack 1.5.0, following #88's per-hunk resolutions. Three commits, the same shape as #60 and #64:sync: port Cursor pstack 0.15.2-0.15.5 skill content: the Step 0 audit and merge reproduced Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 #88's counts (45 upstream paths: 7 verbatim, 15 clean merges, 36 conflict hunks in 20 files, 3 unmapped Cursor-only paths left unported). Every hunk is resolved from the port side plus upstream's new behavior. Model defaults are unchanged in this commit.defaults: Opus/Sol/Grok 4.7 first-run panel; setup probes assigned families: the first-run panel isclaude:opus@max,codex:gpt-5.6-sol@max,grok:grok-4.7@xhigh. Bug-fix, perf-issue and hillclimb stay on Sol max. Setup asks the role question first, then asks efforts for and probes only the assigned families (the Allow setup with only assigned model providers #73 ordering). Only the two existing tests Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 #88 names are edited.release: sync Cursor pstack 0.15.5 as Open Pstack 1.5.0: version 1.5.0 in the three manifests; UPSTREAM.md, CHANGES.md, NOTICE.md, README.md and docs/reference.md updated. Five upstream changes are recorded as exclusions in UPSTREAM.md.Acceptance (#88)
"All 45 upstream paths are accounted for… no conflict markers remain." Step 0 output:
verbatim 7, clean merge 15, needs review 20, removed 0, unmapped 3. The marker grep prints nothing.arena,interrogate,hillclimbandperf-issueresolve to the port side and are byte-identical tomainafter commit 1."No added line under
plugins/pstackintroduces a Cursor primitive, a rejected-slug fallback or an expected-runtime stuck test." The primitive grep prints nothing at each commit."Local behavior is intact."
tests/skill-collision-repro.shpasses with only the panel-derivation edit.check-plan.mjsis identical tomain(git diff --quietexits 0).git diff --stat origin/main...HEAD -- plugins/pstack/skills/poteto-mode/scripts ':!*model-matrix.test.ts'is empty."The CI commands, both
claude plugin validateruns andgit diff --checkpass." Run on the committed tree at each commit:claude plugin validate .andplugins/pstackgit diff --check, marker grep, primitive grepThe assertion count drops by 4 because the first-run sheet holds fewer descriptors, each checked against the matrix.
"UPSTREAM.md, the three manifests, README.md, docs/reference.md, CHANGES.md and NOTICE.md agree on 1.5.0, 0.15.5 and
12d587d." The static test's version check passes, and each file names12d587dfb20741cafc376c42c696c5f6e2a64487."Live checks C1-C3, X1 and X2 are recorded in the PR template." Recorded below.
Merge, tag and release. After Eric's go.
Verification
Live evidence:
Candidate commit
ac4b5cfd40c8d0a10fcbd20e9dce746ff99cbce1(PR head), run on pro14 on 2026-09-30. Each check ran in a fresh headless session with the check's answers in the prompt, as in #64.Recorded state, then restored. Before: Claude Code
pstack@open-pstack1.2.1 (legacypstack@pstack-claude0.9.12 disabled). Codex CLI 0.159.2 with[marketplaces.open-pstack] ref = "main"and 1.4.1 cached.~/.claude/pstack-models.md,~/.codex/pstack-models.mdand the~/.codex/AGENTS.mdblock were all sha256e01bece8…,~/.codex/AGENTS.mdwas32e5f6ac…and~/.codex/config.tomlwasc42bc86e…. After: the same versions, ref and four hashes, compared byte for byte. The Codex cache is back to 1.4.1 atde67e6b.Claude Code, installed version. Each session ran
claude -p --plugin-dir <detached worktree at ac4b5cf>/plugins/pstack. Its init event reportspstack@inline1.5.0 from that path, so the installed 1.2.1 did not load and nothing needed disabling. For C1 and C2,~/.claude/pstack-models.mdwas moved aside and put back afterward (hash unchanged)./pstack:setup-pstackwith no sheet. Answers: keep the roles, empty effort input, decline at confirmation.pstack:pstack-opus-maxreturned its marker. Sol went throughpstack-runner --provider codex --model gpt-5.6-sol --effort max, completed, with receiptmodelEvidence: pinned-argv. Grok went throughpstack-runner --provider grok --model grok-4.7 --effort xhigh, completed, withreportedModel: grok-4.7-buildandmodelVerified: true. After the decline,~/.claude/pstack-models.mdwas still absent./pstack:interrogateon a one-file diff in a scratch repo (a + b→a - b).pstack:pstack-opus-max, B was the Sol runner (receipt complete, pinned-argv) and C was the Grok runner (reportedModel: grok-4.7-build, provider-report). There was no Fable lane. All three flagged the sign flip./pstack:poteto-modewith the same request plus "State the protocol in full, step by step as the playbook defines it". It read the candidateautopilot-full.mdand stated code-ready rounds and swarm verdicts at each code-ready SHA,decisions.tsvandchildren.tsvper owner, the stuck test on affirmative evidence only with no expected runtime, and the merge-prep rebase with CI on the new head. The owner Babysit exception lives inbabysit.md, not in the Autopilot protocol, so a direct probe covered it: "As an Autopilot-full owner babysitting your own PR, a rebase is needed: rebase yourself or report upward?" It answered: rebase yourself, perbabysit.mdstep 4, published through the disarm and captured-SHA lease steps ofautopilot-full.mdstep 2. It wrote nothing (git statusclean).Codex, installed version.
~/.codex/config.tomland~/.codex/AGENTS.mdwere backed up to*.bak-live-gate. The ref was set tosync/PSTACK-3-cursor-pstack-0.15.5, followed bycodex plugin marketplace upgrade open-pstackandcodex plugin add pstack@open-pstack. That installed 1.5.0, with the marketplace root atac4b5cf.diff -rq ~/.codex/plugins/cache/open-pstack/pstack/1.5.0againstgit archive ac4b5cf plugins/pstackprinted nothing. X1 ran ascodex exec --sandbox danger-full-access, and X2 ran read-only.pstack:setup-pstackagainst the existing~/.codex/pstack-models.md, which matched the AGENTS.md block, so no seeding was needed. Answers: Grok family for unmatched rows, keep the roles, empty effort input, decline.how criticsas a dropped retired row. It reported the seven retainedgrok:grok-4.6rows as unmatched and asked for a replacement. After I answered Grok, it asked efforts only for the families the map uses (Fable, Sol, Grok and Opus, all xhigh in Eric's sheet). All four probes passed: Fableclaude-fable-5-1, Opusclaude-opus-5-5, Grokgrok-4.7-buildthrough the runner, and Sol native. The first external attempts failed only because the agent wrapped them in its own extrasandbox-execand moved the config dir; it corrected that and reran the same pairs. NoEPERMoccurred. After the decline, both hashes were unchanged.pstack:poteto-moderead the candidateautopilot-full.mdbut summarized it too thinly. The step-by-step prompt restated steps 1 to 7 from the 1.5.0 cache: code-ready rounds,children.tsvhandles and states, the affirmative-evidence stuck test, the merge-prep rebase and CI on the new head. The Babysit probe answered: rebase yourself, perbabysit.md:10, throughautopilot-full.mdstep 2's captured-SHA lease. It wrote nothing.One pre-existing wording ambiguity showed up, not introduced here: Codex's first X2 run read "Items the operator names" as every queued item, so both PRs would stop at merge-ready.
A pull request without live evidence remains a draft. Do not merge, tag, release, or roll it out.
🤖 Generated with Claude Code