Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

# Contributing

Boatstack is a generated content distribution. Propose changes to workflow semantics, templates, evidence rules, or generated presentation in [Intelligence Flow](https://github.com/operatorstack/intelligence-flow/tree/849cdf48cabc60e5331cc7e8b7668a9f72aff1e6/labs/12-product-engineering-loop).
Boatstack is a generated content distribution. Propose changes to workflow semantics, templates, evidence rules, or generated presentation in [Intelligence Flow](https://github.com/operatorstack/intelligence-flow/tree/b0e65b129cc80a5cb43543ac4b7f89563b64b45e/labs/12-product-engineering-loop).

The Boatstack repository receives product/runtime changes through a generated pull request. Review the PR's `UPSTREAM.json`, tests, adapter diff, and context-size change; do not hand-edit generated output on `main`. `.github/workflows` is the exception: it is Boatstack's executable control plane, excluded from scheduled projection and changed only through a separate manually reviewed Boatstack PR.

Expand Down
59 changes: 47 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,24 +8,48 @@

<p align="center"><strong>Build freely. Prove it. Ship.</strong></p>

## A delivery harness for AI coding agents
## Keep your software delivery process when you change coding agents

<!-- boatstack-claim:portable-product-flow -->An AI coding agent will guess a product decision, call the work "done" on one green check, and lose the reasoning the moment you change tools. Boatstack connects the work from an idea to a reviewed PR so none of that slips: the plan, decisions, gaps, tests, review findings, and project context stay connected along the way. Keep using Cursor, Codex, Claude Code, or Gemini CLI, with the models and specialist skills that fit the work.
Boatstack is a repository-local delivery harness for Cursor, Codex, Claude Code, and Gemini CLI.

The agent remains free to build. Before it says the work is done, Boatstack asks for the approval, tests, review, and recorded evidence appropriate to the change.
AI coding agents can write code quickly, but each tool brings its own planning flow, session state, and definition of “done.” Change agents and your delivery process often disappears with the chat.

**Your product development flow stays with the repository—not the coding agent.** Change tools without rebuilding how you ship or redefining what “done” means. Boatstack carries the workflow and saved project state—not an agent's private chat history or a command already in progress.
<!-- boatstack-claim:portable-product-flow -->Boatstack keeps the delivery process in the repository. Plans, product decisions, tests, review findings, accepted gaps, and completion evidence stay connected from idea to pull request, regardless of which agent or model performs the work. Use Cursor, Codex, Claude Code, or Gemini CLI. Boatstack keeps the same approval, testing, review, and shipping boundaries across them.

**Your product development flow stays with the repository, not the coding agent.** Change agents, models, or specialist skills without rebuilding how your team plans, verifies, reviews, and ships software.

The coding agent executes the work. Boatstack supervises the delivery. Your repository owns the policy and evidence.

<p align="center">
<img src="assets/boatstack-portability.svg" width="900" alt="Change the tools, keep the flow: coding agents, models, and specialist skills feed one repository-owned Boatstack flow that produces a reviewed pull request and useful context for the next feature">
</p>

| You change | Boatstack keeps |
| You change | You keep |
|---|---|
| Cursor, Codex, Claude Code, or Gemini CLI | The same path from planning through PR preparation |
| Lower-cost, general, or frontier model | The same approval, testing, and review requirements |
| React guidance, gstack, Spec Kit, or another skill | Human approval and evidence remain authoritative |
| Session, worktree, or feature | Durable decisions, gaps, evidence, and code state in the repository |
| Cursor, Codex, Claude Code, or Gemini CLI | The same path from approved plan to reviewed PR |
| Lower-cost, general, or frontier model | The same definition of done and proof requirements |
| React guidance, gstack, Spec Kit, or another skill | Human product decisions remain authoritative |
| Session, worktree, or feature | Decisions, open gaps, evidence, and verified delivery state |

## How it works

1. Save a plan in your coding agent.
2. Boatstack validates the plan and pauses for material product decisions.
3. The agent builds freely inside the approved scope.
4. Boatstack checks the promised outcomes against tests and evidence.
5. Review findings, risks, and accepted gaps become a focused PR brief.
6. The resulting context stays in the repository for the next feature.

## Each delivery makes the next one easier

Boatstack does not preserve an agent's private reasoning or replay old chats. It keeps the durable parts of delivery:
- approved product decisions
- unresolved gaps
- validation evidence
- review findings
- verified repository state

That means the next feature starts from recorded project knowledge instead of reconstructing intent from another agent session.

## Install with your coding agent

Expand Down Expand Up @@ -76,7 +100,16 @@ Invoke `/repair` in Claude Code, Cursor, or Gemini CLI, or `$boatstack repair` i

Receipts remain as history; published corrections become linked deliveries.

## Features
## What you get

- **Change coding agents without changing how you ship.**
- **Resume work without reconstructing the previous chat.**
- **Stop agents from guessing material product decisions.**
- **Require evidence for every outcome the change claims to deliver.**
- **Create reviewer-ready PRs from the actual scope, changes, risks, and validation.**

<details>
<summary>Technical Features</summary>

- **A guided path from idea to PR.** `/auto-plan` starts a one-action-at-a-time delivery flow.
- **Instant orientation after a break.** `boatstack next` reconstructs the verified stage without treating chat or a running process as workflow evidence, so you resume in seconds instead of re-reading history.
Expand All @@ -90,6 +123,8 @@ Receipts remain as history; published corrections become linked deliveries.
- **Portable across your AI stack.** Hosts, models, and skills share one repository-owned delivery contract.
- **Repository-friendly maintenance.** Worktrees restore runtime; updates stay in separate infrastructure PRs.

</details>

### Optional changelog

It is disabled by default. Enable it in `.boatstack-project.json`:
Expand Down Expand Up @@ -123,9 +158,9 @@ Boatstack is a repository-local delivery harness.

This does not mean every model performs equally. [See the evidence and paired evaluation design](docs/why-these-steps.md#model-choice-and-budget).

## Why these steps?
## Built from failures observed in real coding work

They derive from coding failures observed in benchmark and product work—not guesses. Each link explains what happened, what Boatstack does, and whether that behavior has actually been tested.
They derive from coding failures observed in benchmark and product work—not guesses. When a failure reveals a reusable delivery problem rather than a project-specific mistake, Boatstack turns it into a boundary future runs can enforce. Each link explains what happened, what Boatstack does, and whether that behavior has actually been tested.

| What happened | What Boatstack does | Current evidence |
|---|---|---|
Expand Down
19 changes: 11 additions & 8 deletions UPSTREAM.json
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@
},
"files": {
".gitignore": "a7079e923a776f14f1bb3a6aa0a11a133a8e1dfb35af020f327623357b7e3957",
"CONTRIBUTING.md": "038386f8aa7deeee581c2b5edbfdc3911aea09e7b6e7a5e1ae798e9fa0bcc951",
"README.md": "a90ba7af7a9cbf71c858d99187f66e0f3db305b2a76fc98ba5625f78ac7c23a4",
"CONTRIBUTING.md": "9cf8d9e3f15f2247c5967a189d649392d3edc18d92d2964a30d5069a489ebb52",
"README.md": "8d481f8e395346400726d02f760f831a8b11062de18b7a76fe4cf00e5e12ca08",
"assets/boatstack-journey.svg": "e465befc50c8ce30f3e07e8fd97012931beeb053392c8fbf38ad645023b3cc63",
"assets/boatstack-mark.svg": "be1f984da1bfa69fa5d1f986d8343d21f7e20921b71db888c928b4d2e54b09b5",
"assets/boatstack-portability.svg": "66dfdfa85db857b3bd18b32047a6975f1fbbfc4dc091158e8277193f9969a346",
Expand All @@ -36,6 +36,8 @@
"boatstack/changelog_test.go": "ce792f23a7fe1e09fb3096cd1314130a6ab69321d4877b12a8e994027541baf7",
"boatstack/cmd/boatstack-helper/main.go": "3504ea23b4792ac1c3cc1aacbdc245be61217378add98ee52e3dbb342c0240b1",
"boatstack/cmd/boatstack-helper/main_test.go": "ff73003b6a5157202fa09ddf1129fb13c3d79702b2e05a8721ce5a11bf5ab779",
"boatstack/decision.go": "257ca328da6ae19ab252f10ee5d06bd7daf49dd8141d083ab1b32f106ea7a94c",
"boatstack/decision_test.go": "1a92ff832610f9559bd47ccac7fc1755a8b4f8261c35bc72a092830dff05f7c0",
"boatstack/delivery.go": "6ff71b6f4ae4f85a184edaf453b5933a79366e36137802fda056e58f83fe319c",
"boatstack/delivery_test.go": "e0323d4e2ef9c74a07c799cf42df91c90a43d61a0fdfddd8b10c5e3e27d5e492",
"boatstack/evidence.go": "497a31e6ff632cb1d7c3adfc9f269af3f6aa84e948dd5d417c162767542a27df",
Expand All @@ -55,8 +57,8 @@
"boatstack/next_test.go": "57828f76257ad7344c383084a97399f6523cbbced28ac2bbe4c14b9da84ccf32",
"boatstack/plan.go": "4ff1b6bf8b187aa26e4851d30c32496a8e979814a5ce2b4b2a1f1d656f7735d0",
"boatstack/plan_test.go": "878dd9086bb583328a7eba1cf45318baeef3d55693be745d5520f07c6ace5e3e",
"boatstack/plan_validation.go": "e601fdcd6ac141aec72819bd95f8411505b81320abbd966e97dfa6cf6d8b771b",
"boatstack/plan_validation_test.go": "7c00279ee271dabaf262e7d6c54d45ee070d1f41ba4ae9d1b5dd146b31ae4070",
"boatstack/plan_validation.go": "848f895e323ae8a428e57f57233e068a8d71c2a7c8716d2ff60e60d23188fd1b",
"boatstack/plan_validation_test.go": "4094b5fdbcb89d7a5a16e4a7595ade88a1f02a13df7824d4c37019b8ca01d860",
"boatstack/planning.go": "4b86ae9dc16393f099ca26812bf3e62909fee42a541b33c33275cc16a80262aa",
"boatstack/planning_test.go": "6b156a64182ed76d4c3d392b4c5a26abe5d8b81cea27ee12ac7c4627c827e186",
"boatstack/pr.go": "28b9c7bd41c0cbe3ec8d05787d868db36b424ecd3489678524fd0f4878262d78",
Expand Down Expand Up @@ -89,10 +91,10 @@
"docs/account-recovery-walkthrough.md": "676034974594a7d1a559b24dbed31d7ccc429eb81404b203ca07bbdaa19ec3d3",
"docs/benchmark-corpus-audit.md": "f2d206fe8579a514f9da82b2c96c19b343ac004be67617e1bd34f0f8e0e5e6c6",
"docs/benchmark-submission-audit.md": "9518abdd17690729c6423f87cab20418ed47b0915b5faa44b9ef975e9e9c3b79",
"docs/evidence-engineered-coding.md": "f1a9de99e5d708c469887a3e723a1112df79b504f6732e7631086d2f2ee4ffa6",
"docs/evidence-engineered-coding.md": "4996b3e4639827757734a6fef96a2d6f29ea83d01426292254d1efd73861ab85",
"docs/generated-files.md": "136422baf0c7fc2bd5100cfe0ebdb3d9d0705dfd7e7d54bf745dd1037e63492c",
"docs/getting-started.md": "eacc814fdffdfa3c7d8052b7cd99a79c04da5c75d88d8b44f3fb68d9afec0316",
"docs/public-claims.json": "960de2b346a95f996795420e27dfda0eb3e991433d98581eee2357b0c8b33186",
"docs/public-claims.json": "2f18d8de941840fd6a39661db60120ddfd9e3be391875570086050f94c3d4502",
"docs/public-surface.md": "713f7a050b5f339cf948299103ef3800417dccfecf2cc1a4166397ea6f978907",
"docs/research-and-design.md": "d65c66e323037bda5d45aacef5d48afa6bf93da55901378891d235aca3a5684f",
"docs/safety.md": "7b9b5c515d36e683767ec8d3d9d6d119ac93650b2f629d351deadd4c600ed6a6",
Expand All @@ -106,7 +108,7 @@
"labs/diagram-json/compiled/evidence.md": "1ba1c989ade070a8ef9a508fbd788d100d7292f2dbacbb2bce895468019f619d",
"labs/diagram-json/compiled/tasks.json": "88f60851abf79d851e9fccc754ff3040034ae595306bc87d64784c19eb403e71",
"labs/diagram-json/compiled/test-matrix.json": "424657ff505768e50fa113801fd8363364a18269d5297480907a993d44063a39",
"labs/diagram-json/plan.lock.json": "2884dad1e921170c8ee4c8ce279e5dc9ad45e6388ce4dfae0e33db497c002f0c",
"labs/diagram-json/plan.lock.json": "e5541208e20e872bc371459e349656bced442c4de4213f091f108260f961cf82",
"labs/diagram-json/plan.md": "3cc4f533b8d69386deff16b3a594a3ba09d4c0c3db636cccd8c4380084ce6a51",
"labs/diagram-json/questions.md": "74733b015002c8a6777c558e7e997fa48c94850b9bd39054fe9366c97ecf728d",
"labs/diagram-json/request.md": "0808fc41c36779c404f4a3a121167da6e76cac56df526e70f9ed6d3e0d4c02ed",
Expand Down Expand Up @@ -148,14 +150,15 @@
"release-notes/2026-07-20-speak-software-standards.md": "a8890ed7eb38868bdf3035572d71785ca89abd3ffc3a314a5dc5315150afe397",
"release-notes/2026-07-21-blueprint-diagrams-value-first-readme.md": "11eb70cb814a99606b4f7761401d3670074fe6bd1af8557f95f573b04a3195af",
"release-notes/2026-07-21-e2e-architecture-grounding.md": "7fa7e99fc6fd4e0688009bb736e6506830bf0216f3fa19757cbd9f5679c7ac3b",
"release-notes/2026-07-21-plan-decision-operator.md": "c2a7416ef17a6042583dd60f1f4e847cb1e0f06b28dbde885479efd566bb5cae",
"release-notes/2026-07-21-prevent-hallucinated-approver-names.md": "a5fd08bc3d8b983340a2ddd67b71c78b2b34a915fdfe6eb7cdffd0f1d52e3427",
"release-notes/2026-07-21-prevent-worktree-dirty-state.md": "d3ebac81a14565461fd7f3c5bd520d881d70b584408ba35a1022cd917c7bf132",
"release-notes/2026-07-21-value-translation-readme.md": "8dd16fd08c1591667a1074fc6825dcbf58beda0faa18ede1526647267418c8ea"
},
"generator": "operatorstack/intelligence-flow:boatstack-distribution",
"schema_version": 1,
"source": {
"commit": "849cdf48cabc60e5331cc7e8b7668a9f72aff1e6",
"commit": "b0e65b129cc80a5cb43543ac4b7f89563b64b45e",
"path": "labs/12-product-engineering-loop",
"repository": "operatorstack/intelligence-flow"
}
Expand Down
94 changes: 94 additions & 0 deletions boatstack/decision.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
package boatstack

type DecisionOperator string

const (
OperatorInfer DecisionOperator = "infer"
OperatorQuery DecisionOperator = "query"
OperatorVerify DecisionOperator = "verify"
OperatorReject DecisionOperator = "reject"
OperatorEscalate DecisionOperator = "escalate"
)

type EvidenceLevel string

const (
EvidenceVerified EvidenceLevel = "verified"
EvidenceSupported EvidenceLevel = "supported"
EvidenceAbsent EvidenceLevel = "absent"
EvidenceConflicting EvidenceLevel = "conflicting"
)

type PremiseStatus string

const (
PremiseUnknown PremiseStatus = "unknown"
PremiseValid PremiseStatus = "valid"
PremiseInvalid PremiseStatus = "invalid"
)

type PlanDecisionInput struct {
DecisionKind string
IsMaterial bool
RepositoryEvidence []EvidenceRecord
EvidenceLevel EvidenceLevel
PremiseStatus PremiseStatus
}

type DecisionResolution struct {
Operator DecisionOperator
RuleID string
Reason string
Evidence []EvidenceRecord
EvidenceLevel EvidenceLevel
PremiseStatus PremiseStatus
}

func ResolvePlanDecision(input PlanDecisionInput) DecisionResolution {
resolution := DecisionResolution{
Evidence: input.RepositoryEvidence,
EvidenceLevel: input.EvidenceLevel,
PremiseStatus: input.PremiseStatus,
}

if input.PremiseStatus == PremiseInvalid {
resolution.Operator = OperatorReject
resolution.RuleID = "invalid-premise-rejected"
resolution.Reason = "planning premise is not supported"
return resolution
}

if input.EvidenceLevel == EvidenceConflicting {
resolution.Operator = OperatorEscalate
resolution.RuleID = "conflicting-evidence-escalated"
resolution.Reason = "repository evidence conflicts"
return resolution
}

if input.EvidenceLevel == EvidenceVerified {
// Assuming verified evidence is sufficient to resolve the decision
resolution.Operator = OperatorInfer
resolution.RuleID = "verified-evidence-inferred"
resolution.Reason = "verified repository evidence resolves the decision"
return resolution
}

if input.EvidenceLevel == EvidenceSupported {
resolution.Operator = OperatorVerify
resolution.RuleID = "supported-evidence-requires-verification"
resolution.Reason = "evidence is supported but requires independent verification"
return resolution
}

if input.IsMaterial && input.EvidenceLevel == EvidenceAbsent {
resolution.Operator = OperatorQuery
resolution.RuleID = "material-intent-requires-human"
resolution.Reason = "material product intent requires human input"
return resolution
}

resolution.Operator = OperatorEscalate
resolution.RuleID = "unresolved-uncertainty-escalated"
resolution.Reason = "unresolved uncertainty requires escalation"
return resolution
}
70 changes: 70 additions & 0 deletions boatstack/decision_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
package boatstack

import "testing"

func TestResolvePlanDecision(t *testing.T) {
tests := []struct {
name string
input PlanDecisionInput
expected DecisionOperator
}{
{
name: "invalid premise rejects",
input: PlanDecisionInput{
PremiseStatus: PremiseInvalid,
},
expected: OperatorReject,
},
{
name: "conflicting evidence escalates",
input: PlanDecisionInput{
PremiseStatus: PremiseValid,
EvidenceLevel: EvidenceConflicting,
},
expected: OperatorEscalate,
},
{
name: "verified evidence infers",
input: PlanDecisionInput{
PremiseStatus: PremiseValid,
EvidenceLevel: EvidenceVerified,
},
expected: OperatorInfer,
},
{
name: "supported evidence verifies",
input: PlanDecisionInput{
PremiseStatus: PremiseValid,
EvidenceLevel: EvidenceSupported,
},
expected: OperatorVerify,
},
{
name: "absent evidence for material intent queries",
input: PlanDecisionInput{
PremiseStatus: PremiseValid,
IsMaterial: true,
EvidenceLevel: EvidenceAbsent,
},
expected: OperatorQuery,
},
{
name: "unknown state escalates",
input: PlanDecisionInput{
PremiseStatus: PremiseUnknown,
EvidenceLevel: EvidenceAbsent,
IsMaterial: false,
},
expected: OperatorEscalate,
},
}

for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
resolution := ResolvePlanDecision(tt.input)
if resolution.Operator != tt.expected {
t.Errorf("expected operator %s, got %s", tt.expected, resolution.Operator)
}
})
}
}
Loading
Loading