Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 28 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,10 @@

Agent Action Stack is a thin orchestrator. It does not re-implement the libraries. It runs them in a fixed order so a visitor can see how they compose.

The testbench pass gates orchestration for a response fixture. It is not a
signed authorization over the rail's separately constructed synthetic proposal.
The same-case workflow below binds the rail outcome to its review and replay.

![Reference workflow from policy evaluation through recourse-gated action and outcome verification, with an optional dispute evidence simulation. Policy failure stops execution.](.github/assets/project-overview.svg)

On policy failure the stack stops. On a clean `settled` outcome, MandateBound is skipped unless you pass `--dispute`.
Expand Down Expand Up @@ -37,9 +41,18 @@ Private repositories are never cloned or modified.

```bash
npm run bootstrap
npm run demo
node ./bin/aas.mjs demo --fault duplicate --prove rail
```

This primary demonstration compensates a duplicate synthetic refund, verifies
its rail receipt, and records a MandateBound review of those same bytes. Expect
`act_outcome: compensated`, `prove_mode: rail-review`, and
`flow: decide -> act -> prove`. A recorded handoff does not establish source
truth or legal effect. Follow the export/replay commands below to review it
without executing the action again.

For a clean settlement that needs no review, run `npm run demo`.

Optional, only for the browser test suite: Playwright needs a Chromium
binary. After `npm install`, run `npx playwright install chromium` once
before `npm run test:browser`.
Expand Down Expand Up @@ -69,10 +82,10 @@ Fail closed at decide:
npm run demo:fail
```

Force the dispute path via a compensated rail outcome:
Review a compensated rail outcome using the same case:

```bash
npm run demo:dispute
node ./bin/aas.mjs demo --fault duplicate --prove rail
```

Expected flow line:
Expand All @@ -81,12 +94,16 @@ Expected flow line:
flow: decide -> act -> prove
```

Review the same case instead of simulating one:
The separate canned simulation remains available explicitly:

```bash
node ./bin/aas.mjs demo --fault duplicate --prove rail
node ./bin/aas.mjs demo --fault duplicate --prove simulate
```

`npm run demo:dispute` retains this simulation behavior for compatibility. Its
MandateBound scenario is unrelated to the rail case and does not produce a
same-case handoff. CLI defaults are unchanged.

The rail-review path persists the act-stage rail bundle, verifies it with the
rail's own verifier, and binds it into a MandateBound review record for the
same action id and digests. The review records the rail's verdict without
Expand Down Expand Up @@ -182,6 +199,9 @@ same-case rail review; every result and export stays tied to its run id.
long-running process. The server binds only to `127.0.0.1` on port
8787 by default (`AAS_GUI_PORT` selects another loopback port), requires the
exact loopback Host and same-origin boundary, and uses POST for a run.
Choose the **Duplicate compensation and review** preset for the
primary handoff workflow. Applying it only prepares the controls; **Run stack**
starts the synthetic action.

## Tests

Expand All @@ -193,6 +213,9 @@ npm run check
`npm test` is the unit suite (orchestrator and GUI models). `npm run
integration` proves the pinned components from a clean checkout, and
`npm run example:review-handoff` runs the integrator example.
Installation, bootstrap, tests, and replay can write dependencies, caches, or
temporary files. See the [local write targets](docs/architecture.md#local-write-targets)
before running them in an existing checkout.

Real browser workflow tests drive the GUI through actual clicks, file
selection, and asynchronous responses with Playwright (Chromium only, to
Expand Down
4 changes: 2 additions & 2 deletions bin/aas-gui.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -319,14 +319,14 @@ export function renderPage() {
<title>Agent Action Stack</title>
<style>html{color-scheme:light}body{font:16px/1.55 system-ui,sans-serif;max-width:1000px;margin:32px auto;padding:0 20px;color:#17202a;background:#f7f9fc}h1{font-size:2.2rem;line-height:1.2}h2{font-size:1.35rem}.state,.panel{background:white;border:1px solid #d7e0ea;border-radius:12px;padding:20px}label{display:inline-block;margin:6px 12px 6px 0}input[type=search]{padding:9px;max-width:100%;box-sizing:border-box}button{background:#183f71;color:white;border:1px solid #183f71;border-radius:6px}a{color:#164d8e}button:focus-visible,a:focus-visible,input:focus-visible,select:focus-visible{outline:3px solid #b46b00;outline-offset:3px}.boundary{border-left:4px solid #183f71;padding:12px 16px;background:#eaf1fa}select,input[type=file]{max-width:100%;box-sizing:border-box}.panel,.boundary{overflow-wrap:anywhere}@media(max-width:600px){body{margin:16px auto;padding:0 12px}.state,.panel{padding:14px}label{display:block}button{min-height:44px}pre{font-size:13px}}button{padding:10px 14px;margin:4px 0;cursor:pointer}button:disabled{cursor:wait;opacity:.6}select{padding:9px;margin:4px}pre{background:#f3f5f7;padding:16px;overflow:auto;border-radius:6px}.state{margin:16px 0}.download{display:none}.panel{margin:16px 0}.error{color:#7a1f1f}</style></head>
<body><main><h1>Agent Action Stack</h1><p>Run the local decide, act, and prove flow using the reviewed component lock.</p>
<p class="boundary">Synthetic local demo only. No real account operations. Policy evaluation, rail receipt verification, and MandateBound recording remain separate authorities. Source truth is unknown; legal effect is not determined.</p>
<p class="boundary">Synthetic local demo only. No real account operations. The testbench response check gates the run; it is not a signed authorization over the rail proposal. Rail receipt verification and MandateBound recording remain separate authorities. Source truth is unknown; legal effect is not determined.</p>
<div class="state"><label>Scenario <select id="scenario"><option value="settled">Clean settlement</option><option value="refusal">Policy refusal</option><option value="compensated">Duplicate compensation and review</option><option value="review">Settled action review</option></select></label> <button id="apply-scenario">Apply scenario</button>
<p id="scenario-note">Choose a scenario or configure the options below. Applying a scenario only changes controls.</p>
<label>Response <select id="response"><option value="pass">pass</option><option value="fail">fail</option></select></label>
<label>Fault <select id="fault"><option value="none">none</option><option value="duplicate">duplicate</option></select></label>
<label>Domain <select id="domain"><option value="refund">refund</option><option value="inventory">inventory allocation</option></select></label>
<label><input id="dispute" type="checkbox"> force dispute proof</label>
<label>Prove <select id="prove"><option value="simulate">simulation</option><option value="rail">same-case rail review</option></select></label>
<label>Prove <select id="prove"><option value="simulate">separate canned simulation</option><option value="rail">same-case rail review</option></select></label>
<br><button id="run">Run stack</button>
<a id="download" class="download" download="agent-action-stack-run.json">Download run bundle</a></div>
<div class="panel" id="summary" aria-live="polite"></div>
Expand Down
30 changes: 26 additions & 4 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,11 @@ The `runDemo` function in `bin/aas.mjs` is the orchestrator. It runs three
stages in order. Each stage returns a child result captured from a spawned
process, and the orchestrator records the stage status into the report.

The decide stage evaluates a response fixture. Its pass gates whether the act
stage runs, but it is not a signed authorization over the rail proposal. The
act stage constructs its own synthetic action inside the rail demo. Same-case
binding begins with that rail bundle and its subsequent review and replay.

```
decide (Constitutional Agent Testbench, Python)
|
Expand Down Expand Up @@ -108,7 +113,24 @@ else: `test` (unit suite plus syntax check plus GUI smoke), `integration`
or release step. It does not write to any registry, package index, or
hosted target. `contents: read` is the only permission requested.

The same boundary holds locally: `npm test`, `npm run integration`,
`npm run example:review-handoff`, and `npm run test:browser` are read-only
with respect to anything outside `.out/`. The orchestrator writes only to
`.out/` for run bundles and to `.out/latest.json` for the latest pointer.
## Local write targets

Local verification does write files. Use a disposable checkout for integration
and browser tests. The commands do not publish, deploy, or operate real accounts.

| Command | Local writes |
| --- | --- |
| `npm ci --ignore-scripts` | Root `node_modules/` and npm's configured cache/log directory |
| `npm run bootstrap` | Pinned public clones under `deps/`, MandateBound `node_modules/` and `dist/`, and npm cache/log files |
| `aas demo`, GUI **Run stack**, integrator examples | `.out/runs/`, `.out/latest.json`, temporary handoff directories under the OS temporary directory, and Python bytecode caches under `deps/constitutional-agent-testbench/src/constitutional_agent_testbench/__pycache__/` unless bytecode writing is disabled |
| `npm test`, `npm run gui:smoke` | Test fixtures, temporary case stores, and child-process scratch files under the OS temporary directory where needed |
| `npm run integration` | Root install, dependency bootstrap/build, the demo writes above (including Python bytecode), and temporary copied verifier runtimes; invokes npm and Git |
| Playwright install and `npm run test:browser` | Configured browser cache, `test-results/`, `playwright-report/`, temporary test case stores, and the GUI **Run stack** writes above (including `.out/` and Python bytecode) |
| `aas runs`, `cases`, `compare`, `inspect`, `latest` | Read saved cases only; shell redirection can write the printed report |
| `aas replay`, `verify`, GUI verification | Read the case store or import, write temporary verifier inputs, and remove scratch files afterward; do not execute an action or alter saved cases |
| `aas export --out` | Writes the named export; existing files require explicit `--overwrite` |
| `aas prune --keep` | Deletes eligible old runs under the selected output root; `--dry-run` previews without deletion |

Saved-case commands honor `--root`. npm and Playwright honor their own cache
configuration, and temporary directories use the operating system's configured
temporary location. Interrupted processes can leave temporary files behind.
19 changes: 13 additions & 6 deletions docs/release-readiness.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,13 +2,20 @@

This checklist describes the evidence required before a public versioning update.

Current requirements: Node.js 22.12+ and Python 3.11+; the exact component pins
come from `stack-lock.json`. CI includes unit, full-stack integration, and real
Chromium browser workflows. The dated candidate records below are historical
evidence for their stated commits, not current runtime or test-count claims.
Local checks may install dependencies and write caches or temporary files; see
[local write targets](architecture.md#local-write-targets).

## Source and dependency gates

- [ ] Review the final raw tree and changed-file list.
- [ ] Confirm `stack-lock.json` still contains the approved public URLs and commits:
- Constitutional Agent Testbench: `16b2faa71b0f92b9afa15b13afad8c48da8132f4`
- Consequence Rail: `6c61e9fdcd1a4701afad1d2371abcb3f13bbab57`
- MandateBound: `e526c4c32ac61571757a98ca1a69189821c3dce7`
- Consequence Rail: `9f60ab3223970c22371c20d3584e8330674d997c` ([PR #67](https://github.com/EauDoon/consequence-rail/pull/67))
- MandateBound: `b51fe137958afe26eee052c5a129e5481ccae560` ([PR #102](https://github.com/EauDoon/mandatebound/pull/102))
- [ ] Run bootstrap from a clean workspace and verify detached, clean, exact dependency checkouts.
- [ ] Confirm no private repository, credential, or production endpoint is referenced.

Expand All @@ -24,8 +31,8 @@ This checklist describes the evidence required before a public versioning update
## Integration evidence, 2026-09-06 candidate

Candidate: this branch at the pin-update commit (`chore/update-stack-pins-202609`;
orchestrator base `a0e25144aaebff2885b67eb6b5b355c37a167f37`). Component pins
are the reviewed merge SHAs above: testbench PR #20 (`16b2faa7`), rail PR #20
orchestrator base `a0e25144aaebff2885b67eb6b5b355c37a167f37`). That candidate's
component pins were testbench PR #20 (`16b2faa7`), rail PR #20
(`64cb30`, includes the unknown-recourse receipt-refusal fix), MandateBound PR
#26 (`3682a24`, includes the fast-uri advisory fix). No component was upgraded
past its reviewed merge; the rail SHA supersedes the earlier `7cf59e79`
Expand All @@ -45,8 +52,8 @@ Clean-checkout proof (`/tmp/aas-clean`, fresh clone, no `deps/`, `dist/`, or
modifications, HEAD equal to the pinned full SHA; remotes are the three
public `EauDoon` repository URLs; MandateBound `npm ci` reports
0 vulnerabilities and its `tsc` build produces `dist/cli.js`.
- `npm test`: 51 passed, 0 failed (includes the lock-provenance fixture that
asserts the three pins above).
- `npm test`: 51 passed, 0 failed (includes the lock-provenance fixture for
that candidate's three pins).
- `npm run gui:smoke`: passed.
- `node ./bin/aas.mjs demo`: pass path, `decide passed`, `act settled`,
`CLOSED`, `flow: decide -> act`; bundle records stage artifacts plus
Expand Down
26 changes: 25 additions & 1 deletion docs/stack-lock.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,30 @@ is bumped only when the shape changes in a way that requires loader
changes; existing tools keep reading older versions until the bump lands
across all consumers.

The integration proof also runs `test/component-compatibility.test.mjs` against
the prepared components. It checks canonical bytes independently of producer
round trips, both currency validation boundaries, and refusal of re-signed
synthetic evidence whose currency contradicts the proposal. The historical
`fixtures/legacy-rail-review.json` was exported using Rail `6c61e9f` and
MandateBound `e526c4c`; it pins replay compatibility for an ordinary synthetic
refund. Its public demonstration signatures establish no real-world provenance.

An older artifact with numeric-looking object keys may contain signatures made
with the former Rail canonical ordering. The current verifier does not try that
obsolete ordering after verification fails. Preserve the original artifact and
its recorded producer revision for historical inspection; do not rewrite its
signatures or describe a newly generated case as the same evidence. A successful
legacy fixture replay establishes compatibility for that fixture, not every
previously accepted artifact or an alternate canonical profile.

The MandateBound pin also corrects U+2028/U+2029 bytes under its existing
RFC8785 profile. Legacy proofs containing those separators can fail current
integrity checks. Keep original bytes and the exact producer commit for
historical replay; version labels alone do not identify the affected behavior.
There is no alternate-byte verification or automatic migration. Follow the
[producer's compatibility guidance](https://github.com/EauDoon/mandatebound/blob/b51fe137958afe26eee052c5a129e5481ccae560/docs/PROTOCOL.md)
before reissuing affected artifacts.

## Who can update

The lock is owned by the Agent Action Stack maintainers. Updates land via
Expand All @@ -79,4 +103,4 @@ or in this policy document.
Real connectors, real secrets, real payment or merchant integrations;
changes that would read or write private repositories; policy or rail
rules that should live in their owning library. The lock pins reviewable
public artifacts. Anything else belongs in the sibling that owns it.
public artifacts. Anything else belongs in the sibling that owns it.
6 changes: 5 additions & 1 deletion examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,10 @@ One journey across both synthetic domains (`--domain refund|inventory`):
The pass path exits 0 after every binding verifies. The refusal path exits 1
after the policy refuses, before anything executes.

This is the primary same-case demonstration. The separate `--prove simulate`
mode runs an unrelated canned dispute scenario. The testbench pass is a gate
over a response fixture, not a signed authorization of the rail proposal.

What it establishes: the policy gate passed for this response, the rail
produced this outcome for this action, recourse was reserved before the
effect, the rail's verifier accepts the persisted bytes under the synthetic
Expand Down Expand Up @@ -76,7 +80,7 @@ treats an observed effect as proof of external truth.
connector's synthetic demo key. Verifier trust is supplied by the caller;
nothing embedded in a bundle is trusted for its own integrity.
- Caller-owned anchors (expected digests, expected action identity) are
separate from exported untrusted material. Rethem recomputation always runs
separate from exported untrusted material. Digest recomputation always runs
over the actual bytes.
- Source truth is never established. A recorded review proves the handoff
and the digest binding, not the underlying external state.
Expand Down
12 changes: 9 additions & 3 deletions examples/connector-conformance.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -129,15 +129,21 @@ function main() {
await inventory.execute(inventoryProposal, inventoryProposal.idempotency_key);
const first = await inventory.remediate(inventoryProposal, { connector_commitment: reservation }, "remedy:a");
const onHandAfterFirst = inventory.inventory.get("sku_demo_1");
const second = await inventory.remediate(inventoryProposal, { connector_commitment: reservation }, "remedy:b");
const replay = await inventory.remediate(inventoryProposal, { connector_commitment: reservation }, "remedy:a");
let secondError = null;
try { await inventory.remediate(inventoryProposal, { connector_commitment: reservation }, "remedy:b"); }
catch (error) { secondError = error.code; }
out({
first: first.status, second: second.status,
first: first.status, sameResult: JSON.stringify(first) === JSON.stringify(replay),
recourse: inventory.recourseStatus(reservation.reservation_token).status,
secondError,
restoredOnce: inventory.inventory.get("sku_demo_1") === onHandAfterFirst,
onHand: inventory.inventory.get("sku_demo_1"),
});
`);
check(RULES[4],
remedy.first === "remediated" && remedy.second === "failed" && remedy.restoredOnce === true,
remedy.first === "remediated" && remedy.sameResult === true && remedy.recourse === "consumed"
&& remedy.secondError === "RECOURSE_NOT_ACTIVE" && remedy.restoredOnce === true,
`the remedy must reverse once and refuse a second restoration (saw ${JSON.stringify(remedy)})`);

const remedyStatus = run(`
Expand Down
Loading
Loading