Skip to content

release: sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 - #98

Merged
mergify[bot] merged 3 commits into
mainfrom
sync/PSTACK-3-cursor-pstack-0.15.5
Oct 1, 2026
Merged

mergify[bot] merged 3 commits into
mainfrom
sync/PSTACK-3-cursor-pstack-0.15.5

Conversation

@ericlitman

@ericlitman ericlitman commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Closes #88
Closes #72
Linear: PSTACK-3

What changed

Imports Cursor pstack 0.15.2 to 0.15.5 (f8abedd..12d587d) and releases it as Open Pstack 1.5.0, following #88's per-hunk resolutions. Three commits, the same shape as #60 and #64:

  1. sync: port Cursor pstack 0.15.2-0.15.5 skill content: the Step 0 audit and merge reproduced Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 #88's counts (45 upstream paths: 7 verbatim, 15 clean merges, 36 conflict hunks in 20 files, 3 unmapped Cursor-only paths left unported). Every hunk is resolved from the port side plus upstream's new behavior. Model defaults are unchanged in this commit.
  2. defaults: Opus/Sol/Grok 4.7 first-run panel; setup probes assigned families: the first-run panel is claude:opus@max, codex:gpt-5.6-sol@max, grok:grok-4.7@xhigh. Bug-fix, perf-issue and hillclimb stay on Sol max. Setup asks the role question first, then asks efforts for and probes only the assigned families (the Allow setup with only assigned model providers #73 ordering). Only the two existing tests Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 #88 names are edited.
  3. release: sync Cursor pstack 0.15.5 as Open Pstack 1.5.0: version 1.5.0 in the three manifests; UPSTREAM.md, CHANGES.md, NOTICE.md, README.md and docs/reference.md updated. Five upstream changes are recorded as exclusions in UPSTREAM.md.

Acceptance (#88)

  1. "All 45 upstream paths are accounted for… no conflict markers remain." Step 0 output: verbatim 7, clean merge 15, needs review 20, removed 0, unmapped 3. The marker grep prints nothing. arena, interrogate, hillclimb and perf-issue resolve to the port side and are byte-identical to main after commit 1.

  2. "No added line under plugins/pstack introduces a Cursor primitive, a rejected-slug fallback or an expected-runtime stuck test." The primitive grep prints nothing at each commit.

  3. "Local behavior is intact." tests/skill-collision-repro.sh passes with only the panel-derivation edit. check-plan.mjs is identical to main (git diff --quiet exits 0). git diff --stat origin/main...HEAD -- plugins/pstack/skills/poteto-mode/scripts ':!*model-matrix.test.ts' is empty.

  4. "The CI commands, both claude plugin validate runs and git diff --check pass." Run on the committed tree at each commit:

    Check 73ca5d4 3f2708b HEAD
    bun test + strict typecheck 158 pass, 652 assertions 158 pass, 648 assertions 158 pass, 648 assertions
    manifest JSON parse pass pass pass
    static invariants pass (old four-model panel) pass (three-model panel) pass (1.5.0 in all four places)
    claude plugin validate . and plugins/pstack pass pass pass
    git diff --check, marker grep, primitive grep clean clean clean

    The assertion count drops by 4 because the first-run sheet holds fewer descriptors, each checked against the matrix.

  5. "UPSTREAM.md, the three manifests, README.md, docs/reference.md, CHANGES.md and NOTICE.md agree on 1.5.0, 0.15.5 and 12d587d." The static test's version check passes, and each file names 12d587dfb20741cafc376c42c696c5f6e2a64487.

  6. "Live checks C1-C3, X1 and X2 are recorded in the PR template." Recorded below.

  7. Merge, tag and release. After Eric's go.

Verification

  • Bun tests, strict typecheck, static invariants, and plugin validation pass.
  • The exact candidate is installed in every affected harness.
  • The changed behavior passes from each real user surface.
  • The installed version, action, and observed result appear below.

Live evidence:

Candidate commit ac4b5cfd40c8d0a10fcbd20e9dce746ff99cbce1 (PR head), run on pro14 on 2026-09-30. Each check ran in a fresh headless session with the check's answers in the prompt, as in #64.

Recorded state, then restored. Before: Claude Code pstack@open-pstack 1.2.1 (legacy pstack@pstack-claude 0.9.12 disabled). Codex CLI 0.159.2 with [marketplaces.open-pstack] ref = "main" and 1.4.1 cached. ~/.claude/pstack-models.md, ~/.codex/pstack-models.md and the ~/.codex/AGENTS.md block were all sha256 e01bece8…, ~/.codex/AGENTS.md was 32e5f6ac… and ~/.codex/config.toml was c42bc86e…. After: the same versions, ref and four hashes, compared byte for byte. The Codex cache is back to 1.4.1 at de67e6b.

Claude Code, installed version. Each session ran claude -p --plugin-dir <detached worktree at ac4b5cf>/plugins/pstack. Its init event reports pstack@inline 1.5.0 from that path, so the installed 1.2.1 did not load and nothing needed disabling. For C1 and C2, ~/.claude/pstack-models.md was moved aside and put back afterward (hash unchanged).

Check Action Observed result
C1 /pstack:setup-pstack with no sheet. Answers: keep the roles, empty effort input, decline at confirmation. Pass. It proposed the 15-row first-run sheet from #88 exactly. It asked the role question first, then three effort questions (Opus max, Sol max, Grok xhigh) and none for Fable. Probes: native pstack:pstack-opus-max returned its marker. Sol went through pstack-runner --provider codex --model gpt-5.6-sol --effort max, completed, with receipt modelEvidence: pinned-argv. Grok went through pstack-runner --provider grok --model grok-4.7 --effort xhigh, completed, with reportedModel: grok-4.7-build and modelVerified: true. After the decline, ~/.claude/pstack-models.md was still absent.
C2 /pstack:interrogate on a one-file diff in a scratch repo (a + b → a - b). Pass. There were three lanes: Reviewer A was native pstack:pstack-opus-max, B was the Sol runner (receipt complete, pinned-argv) and C was the Grok runner (reportedModel: grok-4.7-build, provider-report). There was no Fable lane. All three flagged the sign flip.
C3 "full autopilot for these two items … state the protocol and wait" in a scratch repo. Pass on the rerun. The first attempt, with the literal prompt, loaded no skill: the model judged two one-line items too small for pstack. That routing comes from the session-start hook, which this PR does not change. The next run was /pstack:poteto-mode with the same request plus "State the protocol in full, step by step as the playbook defines it". It read the candidate autopilot-full.md and stated code-ready rounds and swarm verdicts at each code-ready SHA, decisions.tsv and children.tsv per owner, the stuck test on affirmative evidence only with no expected runtime, and the merge-prep rebase with CI on the new head. The owner Babysit exception lives in babysit.md, not in the Autopilot protocol, so a direct probe covered it: "As an Autopilot-full owner babysitting your own PR, a rebase is needed: rebase yourself or report upward?" It answered: rebase yourself, per babysit.md step 4, published through the disarm and captured-SHA lease steps of autopilot-full.md step 2. It wrote nothing (git status clean).

Codex, installed version. ~/.codex/config.toml and ~/.codex/AGENTS.md were backed up to *.bak-live-gate. The ref was set to sync/PSTACK-3-cursor-pstack-0.15.5, followed by codex plugin marketplace upgrade open-pstack and codex plugin add pstack@open-pstack. That installed 1.5.0, with the marketplace root at ac4b5cf. diff -rq ~/.codex/plugins/cache/open-pstack/pstack/1.5.0 against git archive ac4b5cf plugins/pstack printed nothing. X1 ran as codex exec --sandbox danger-full-access, and X2 ran read-only.

Check Action Observed result
X1 pstack:setup-pstack against the existing ~/.codex/pstack-models.md, which matched the AGENTS.md block, so no seeding was needed. Answers: Grok family for unmatched rows, keep the roles, empty effort input, decline. Pass. It listed how critics as a dropped retired row. It reported the seven retained grok:grok-4.6 rows as unmatched and asked for a replacement. After I answered Grok, it asked efforts only for the families the map uses (Fable, Sol, Grok and Opus, all xhigh in Eric's sheet). All four probes passed: Fable claude-fable-5-1, Opus claude-opus-5-5, Grok grok-4.7-build through the runner, and Sol native. The first external attempts failed only because the agent wrapped them in its own extra sandbox-exec and moved the config dir; it corrected that and reran the same pairs. No EPERM occurred. After the decline, both hashes were unchanged.
X2 C3 repeated in Codex. Pass on the rerun. The literal prompt with pstack:poteto-mode read the candidate autopilot-full.md but summarized it too thinly. The step-by-step prompt restated steps 1 to 7 from the 1.5.0 cache: code-ready rounds, children.tsv handles and states, the affirmative-evidence stuck test, the merge-prep rebase and CI on the new head. The Babysit probe answered: rebase yourself, per babysit.md:10, through autopilot-full.md step 2's captured-SHA lease. It wrote nothing.

One pre-existing wording ambiguity showed up, not introduced here: Codex's first X2 run read "Items the operator names" as every queued item, so both PRs would stop at merge-ready.

A pull request without live evidence remains a draft. Do not merge, tag, release, or roll it out.

🤖 Generated with Claude Code

ericlitman and others added 3 commits September 30, 2026 18:06
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…milies

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium risk] Updates plugin documentation and configuration for a new release.

The PR is not ready to merge because Shipping can keep an independent test verdict after the tests it covered change.

Findings

  1. P1 Changed tests keep old verdicts ▶
  2. P2 Concurrent runs duplicate log headers ▶

Summary

Open Pstack syncs Cursor pstack 0.15.5 and releases it as version 1.5.0. The update changes model setup, autonomous PR workflows, and decision-trail guidance.

  • The first-run model panel uses Opus, Sol, and Grok 4.7; setup probes only assigned model families.
  • PR playbooks add verification rounds for code-ready patches and later patch changes.
  • Decision logs mark where each agent run starts, and release files record the new version and upstream source.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A["Code-ready head"] --> B["Independent review lanes"]
  B --> C["Later head changes"]
  C --> D{"Patch still matches?"}
  D -->|"Yes"| E["Keep verdict"]
  D -->|"Tests, docs, or lint only"| F["Compare build outputs"]
  F --> G["Reuse or rerun lanes"]
  G --> H["Fresh checks, then merge"]
Loading

Reviews (1) · Last reviewed commit: "release: sync Cursor pstack 0.15.5 as Op..."

1. **Resolve the forge, repository identity, and both Git remotes.** GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, use `origin pr ...` for PR view, watch, and edit operations; use it for merge only when step 5's expected-head requirement is available. Otherwise stay on `gh` and record the fallback. Record the forge-reported base repository as canonical `<base-repo>`. On GitHub, split it into `<base-owner>` and `<base-name>`, pass `--repo "$base_repo"` to every `gh pr` command, and pass the two components to the GitHub watcher. Resolve `<head-remote>` from the branch's configured push remote, `remote.pushDefault`, branch remote, or sole unambiguous remote, in that order. Confirm its push URL names the PR head repository, and record that exact URL as `<head-url>` so the captured remote state and the guarded push address the same repository. Resolve `<base-remote>` independently by matching a fetch URL to `<base-repo>`, and confirm that URL before using it for trunk. The same remote name may fill both roles when its push URL matches the head repository and its fetch URL matches the base repository. Compare each URL by role instead of assuming the remote name identifies one repository. Capture every resolved or forge-reported value directly into a shell variable. Command examples use quoted lower-case variables such as `"$branch"` and `"$head_url"`. Never paste those values into shell source. Never guess or treat the forge name as a Git remote, and never require Graphite (`gt`).
2. **Freeze and disarm the queue before verification.** Freeze an explicit bottom-to-top PR list. Confirm a same-repository stack against its base-branch chain. For fork heads, take the order from the verified local parent ancestry because every PR targets trunk and the forge bases do not encode the stack. Before launching any verifier, inspect every PR in the frozen list through the active forge, disarm every pre-existing merge-when-ready or auto-merge request, and confirm each request is off. On GitHub, query each PR through GraphQL for `id`, `headRefOid`, `baseRefName`, `autoMergeRequest`, and `mergeQueueEntry`. When `autoMergeRequest` is non-null, run `gh pr merge "$pr" --disable-auto --repo "$base_repo"`. When `mergeQueueEntry` is non-null, invoke the `dequeuePullRequest` mutation with `gh api graphql -F "id=$pr_node_id" -f query='mutation($id:ID!){dequeuePullRequest(input:{id:$id}){mergeQueueEntry{id}}}'`. The mutation takes the pull request node ID. Re-query and require that both `autoMergeRequest` and `mergeQueueEntry` are null. A null `autoMergeRequest` alone does not prove that the pull request is unarmed. On Origin, use its reported cancel operation and inspect every separately reported queue state. Stop if the active forge cannot confirm the whole list is unarmed. One subagent per PR, not batched, each in its own worktree, exercises the real surface against that PR's parent versus head. The bottom PR's patch base is trunk. Each child's patch base is the preceding PR's exact head, including when a fork child targets trunk at the forge. Each subagent returns `PASS`, `PASS+NOTES` or `FAIL` and posts that verdict on its own PR. Safe means a verdict from an agent that did not write the code. CI green is not a verdict, and an approving bot review is not a verdict. Walk up from the bottom and stop at the first PR without a passing verdict, where both `PASS` and `PASS+NOTES` pass. Report that ceiling and what breaks the chain.
3. **Re-check that each verdict still describes the patch.** At the passing verdict, record `<verdict-sha>`, `<verdict-base-sha>`, and the stable `<verdict-patch-id>` of `<verdict-base-sha>..<verdict-sha>`. Initialize `<landing-sha>` and `<landing-base-sha>` to those head and base values. Do not replace a fork child's patch base with its forge base, because that forge base is trunk and includes its unmerged parents. Before landing, read the current published head through the active forge as `<current-head>` and require the local branch tip to match it. Calculate the stable patch ID of `<landing-base-sha>..<current-head>` and compare it with `<verdict-patch-id>`. Re-verify when the patch changed. When it did not, keep the code verdict, set `<landing-sha>` to `<current-head>`, and re-run mergeability and CI at that head. A patch-equivalent rebase can make `<landing-sha>` differ from `<verdict-sha>`. Keep both values. Never use matching commit messages or a green check from an older SHA as a substitute.
3. **Re-check that each verdict still describes the patch.** At the passing verdict, record `<verdict-sha>`, `<verdict-base-sha>`, and the stable `<verdict-patch-id>` of `<verdict-base-sha>..<verdict-sha>`. Initialize `<landing-sha>` and `<landing-base-sha>` to those head and base values. Do not replace a fork child's patch base with its forge base, because that forge base is trunk and includes its unmerged parents. Before landing, read the current published head through the active forge as `<current-head>` and require the local branch tip to match it. Calculate the stable patch ID of `<landing-base-sha>..<current-head>` and compare it with `<verdict-patch-id>`. When the two patches differ only in tests, docs, or lint config, build what each lane ran. Build it twice at `<verdict-sha>` and once at `<current-head>`. A difference is noise if the two builds at `<verdict-sha>` also show it, or if it is an embedded commit SHA. Judge each difference, not each file, and report each kind of noise with its files. If only noise differs, that lane's result stays valid. Set `<landing-sha>` to `<current-head>`, and run checks and a review of the change fresh at that head. Do not reuse a lane result from a dev server or from anything else with no build output. Rerun that lane. Re-verify anything else when the patch changed. When it did not, keep the code verdict, set `<landing-sha>` to `<current-head>`, and re-run mergeability and CI at that head. A patch-equivalent rebase can make `<landing-sha>` differ from `<verdict-sha>`. Keep both values. Never use matching commit messages or a green check from an older SHA as a substitute.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Changed tests keep old verdicts

When a PR changes only tests, docs, or lint settings, Shipping can keep a test lane’s earlier pass if its build output differs only in ways it calls noise. A test-only edit can change what that lane checks without changing the built app. Fresh checks and a review do not replace the lane’s independent verdict, so rerun a lane when its inputs or checks change.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this. The rule builds what each lane ran. A test-only edit changes what a test lane runs, so its build differs beyond noise and the lane reruns. A lane with no build output is never reused. Only lanes whose inputs did not change keep their result. This text is upstream 12d587d's rule, taken verbatim per #88.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. The lane-specific build comparison means a test-only change changes the inputs for the test lane, so its build differs beyond the permitted noise and that lane is rerun. Lanes whose inputs and checks are unchanged can retain their results, while lanes without build output are never reused. I withdraw the finding; no change is needed.

Comment on lines +21 to 23
if [ ! -s "$logfile" ]; then
printf 'ts\tphase\tdecision\twhy\tevidence\tresult\n' >> "$logfile"
fi

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Concurrent runs duplicate log headers

log.sh checks whether a log is empty before appending its header, but those are separate actions. If two runs start the same empty log together, both can append a header. Make header creation safe for concurrent first writes so the decision trail stays readable.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this. log.sh is byte-identical to upstream 12d587d, taken verbatim per #88. Its comment accepts one stray header line as the cost, never lost rows. It replaces main's check-then-truncate version, which could drop rows. A concurrent first write to one log is outside the skill's usage (one log per effort).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That context makes sense. Given the one-log-per-effort usage invariant, concurrent first writes are outside the supported usage, and the upstream change intentionally prioritizes avoiding row loss over eliminating a possible duplicate header. Since this is a verbatim upstream import and the comment explicitly documents that trade-off, I’m withdrawing this finding; no change is needed.

@ericlitman
ericlitman marked this pull request as ready for review September 30, 2026 22:30
@mergify

mergify Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Queued — the merge queue status continues in this comment ↓.

@unfret-eal

unfret-eal Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

✅ Accepted risk

Accepted by ericlitman: PSTACK-3 (authority 48fa77b)

🟡 Medium · Fence surviving writers before replacement

🟡 Medium · Standalone TDD no longer runs adjacent validation

Advisories (1)
  • Panel invariant still reports a provider-dispatch check it no longer runs
    Not blocking because: The script overstates its comparison coverage, but the required CI test suite independently checks setup against the provider matrix. The confirmed impact is misleading diagnostic text, not loss of matrix-drift enforcement.
Details

Run: GET /unfret/run/panel:ac4b5cfd40c8d0a10fcbd20e9dce746ff99cbce1:Sy5Ew4QWprz:V7XvBurIXLk.

@unfret review · @unfret status · @unfret help

@ericlitman

Copy link
Copy Markdown
Owner Author

@unfret review

1 similar comment
@ericlitman

Copy link
Copy Markdown
Owner Author

@unfret review

@unfret-eal unfret-eal Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛑 2 blocking findings
🟡 Medium · Fence surviving writers before replacement
🟡 Medium · Standalone TDD no longer runs adjacent validation
Check

7. **Stand down instantly on the operator's stop.** Her hold or stand-down reaches every owner as a zero-writes order immediately. Owners hold their briefs until she releases them.
4. **Swarm-verify every round before its merge.** A round starts at the owner's code-ready head SHA and at each later push that changes the PR's patch. At that SHA, fan out parallel independent verifiers per the **swarm** skill and aggregate to one verdict. The merge needs a clean verdict from the round whose patch matches the merge-ready head. Audit the receipts in the merge-ready report before the verdict. The lanes: re-run the gates at that SHA. Prove the load-bearing behavior live on the real surface the change touches (the `run` skill for CLIs and TUIs, `verify` for UIs, or a named driver where neither fits). Audit the diff, distrusting the PR body. Run the audit as two or more review lanes with the full brief. Give each lane one main focus, such as consumer parity with trunk, lifetimes and races, or data and config safety. **Regression lane against trunk.** Run the same load-bearing scenario on current trunk. If trunk does not have the feature, record that fact and gate the behavior the diff adds plus the end state the user waits for instead of pretending trunk can produce it. The live lane is the floor, and a verdict without it is not clean. No merge without the root's clean verdict. When the lanes return, send every proven finding against the PR to the owner in one fix-forward. A defect that a lane filed as a note is a finding. For each behavior finding, ask for a red test that covers every site with the same defect. Where no test can show the defect, ask for a repro receipt instead. Add that defect to the next round's review brief. The new head gets a fresh swarm and a fresh verdict, except for lane results that stay valid under the patch-id rule in `playbooks/shipping.md`.
5. **On a clean verdict the owner merges and takes the next item.** The owner merges only from a head freshly rebased onto trunk. Merge prep never comes before a round's lanes start, and it ends with a rebase onto current trunk right before the merge. After the merge-prep rebase, the owner reports the new head SHA. CI must pass on that head before the merge, and the patch-id rule decides whether the round's verdict still holds. A private-stack child reaches that state through Shipping step 4 after its parent merges. The merge-ready report records the verdict SHA and the current landing SHA. If trunk moves again before the merge, the patch ID rule in `playbooks/shipping.md` governs re-verification. A changed patch needs a new verdict. An unchanged patch keeps the verdict, but Shipping step 3 records the new `<landing-sha>` after current CI and mergeability pass. The owner lands its own PR only through Shipping step 5's server-enforced expected-head flow bound to `<landing-sha>`; never use a PR-number-only merge or unguarded auto-merge. The owner then picks up its next self-contained item from the queue. The operator's full-autonomy grant plus the root's clean verdict is the merge authorization that babysitting alone never has. Operator-named items stop at merge-ready and wait for the operator's click.
6. **Run the root layer.** A genuinely new raise of a pinned gate or budget value (a limit CI only lets tighten) needs your fresh countersign, granted only after verifier proof. If the operator's grant or standing orders cover approvals, that countersign is the approval. The owner records it in the form that the tool's approval contract allows, with a pointer to the root's countersign. A lane checks the record against that countersign. The root never gives or bypasses an approval that the forge enforces. Absorbing values that already landed on main is drift, not a raise. Run an audit tick over all owners roughly every 30 minutes. The tick is an observation cadence, never a lease, deadline, or cancellation threshold. Arm each tick as a real `/loop` in dynamic mode, which schedules its own wake-up rather than blocking on a sleep. Never leave the cadence to memory or lossy completion notifications. At each tick, re-read this playbook from disk (`skills/poteto-mode/playbooks/autopilot-full.md` under the installed plugin), then re-read the standing objective. Audit the operation against both. Fix drift during that tick and treat it as urgent. Probe each owner through its retained handle with a generic liveness or status check, and collect the decision trails. Count commits, pushes, PR or check deltas, store reports, and a live retained process as evidence. Elapsed time or the absence of a new side effect alone never proves a lane is stuck; implementation runs can remain healthy for 90 minutes or much longer. Stand a lane down only on affirmative failure evidence such as a dead process, failed handle, explicit error, or a caller-supplied external deadline. Each tick also runs this stuck test over the parent's background task list and every owner's `children.tsv`. Cancel a stuck lane through its retained handle. Whether or not the cancel succeeds, the root has the owner record it as stuck in `children.tsv` and, if its work is still needed, replace it. Each replacement that fails gets the same steps. The root takes both steps when the owner cannot. A stall never proves or drops the work. When merges batch, run a retro pass and a post-merge bot-comment sweep. End the tick only when no delegated work is left, even after the last merge.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium · Fence surviving writers before replacement

If the runner is killed while its provider CLI remains alive, cancellation through the dead retained handle fails. This is reproducible with the shipped launcher, but step 6 now requires replacement even when cancellation fails. Reassigning the same PR leaves the original writer able to publish alongside its replacement. Separate worktrees do not isolate the shared PR branch, so late pushes can reject the replacement's push or change the head under verification. Confirm termination or fence the original writer's publication rights before replacement. Also reported for this defect: - Replacement can race a writer whose cancellation failed (plugins/pstack/skills/poteto-mode/playbooks/autopilot-full.md:10): Step 6 requires replacement even when cancellation fails. If a caller-supplied deadline expires while a writer remains alive and its cancellation fails, the replacement can write or push the same PR concurrently, causing conflicting edits, lease failures, or a blocked delivery run.

## If a Failing Test Is Impractical

Do not silently skip the regression step. Before fixing, explicitly explain why a failing test is impossible or not worth the cost, then choose the closest executable regression check available. Examples include a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check.
Use the closest executable regression check instead: a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium · Standalone TDD no longer runs adjacent validation

This hunk deletes 7. Run nearby validation..., while the replacement requires only the focused regression check. When /tdd fixes a shared helper, that test can pass while an adjacent test, typecheck, or scenario fails; the workflow then reports a successful fix with a regression still present.

@ericlitman

Copy link
Copy Markdown
Owner Author

Unfret waiver for head ac4b5cf. Tracked by PSTACK-3 and #88. Eric authorized this waiver on 2026-10-01.

  • v1:2e0d4b6955e04b869983cb5acf6aac2e5cd58d2efd218048d44ee21085da4e3f, "Fence surviving writers before replacement". Class: fails-loudly.
    • The text in autopilot-full.md step 6 is upstream 12d587d's, and Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 #88 takes it verbatim.
    • When two writers share one PR branch, every push goes through the captured-SHA lease (autopilot-full.md step 2) or must fast-forward, so a stale writer's push is rejected.
    • A head that changes under verification cannot merge on the old verdict: Shipping step 5's expected-head merge and the step 3 patch-id rule send it to a fresh round.
    • The trigger needs a dead runner, a provider that is still alive and a failed cancel at the same time.
  • v1:f98f01d34841e571fc045f4fa40b00a8d822e626bc8eccc34892dca480741a03, "Standalone TDD no longer runs adjacent validation". Class: accepted by operator.
    • The removed step 7 is upstream b0b9c7a (#419), an intentional instruction cut. Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 #88 takes the tdd merge result with no edits, and Eric chose to keep upstream's cut.
    • The panel recorded that the wrong-result consequence is unproven.

The advisory v1:2f6c8062… (stale "across provider dispatch" wording in the static test's messages) stays advisory.

@ericlitman

Copy link
Copy Markdown
Owner Author

@Mergifyio queue

@mergify

mergify Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Merge Queue Status

  • ✅ Entered queue — 2026-10-01 13:08 UTC · Rule: default · triggered by @ericlitman with the @mergifyio queue command
  • ✅ Checks skipped · PR is already up-to-date
  • ✅ Merged — 2026-10-01 13:09 UTC · at 1c91e653dd0f2bfcd849f0cce27ae5c07a3a37fb

This pull request spent 18 seconds in the queue, including 2 seconds running CI.

Required conditions to merge
  • github-review-approved [🛡 GitHub repository ruleset rule Protect main]
  • any of [🛡 GitHub repository ruleset rule Protect main]:
    • check-success = @github-actions/verify
    • check-neutral = @github-actions/verify
    • check-skipped = @github-actions/verify
  • any of [🛡 GitHub repository ruleset rule Review gate]:
    • check-success = @unfret-eal/Unfret
    • check-neutral = @unfret-eal/Unfret
    • check-skipped = @unfret-eal/Unfret

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Sync Cursor pstack 0.15.5 as Open Pstack 1.5.0 Configure pstack with only the selected model providers

1 participant