Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -247,6 +247,17 @@ jobs:
env:
PS_REQUIRE_CANARY: '1'

# After the build, because it drives the built CLI against a project that links the built package:
# the listener reporter is COPIED into `dist/` rather than bundled, so a rename or a missed copy
# step is invisible to every source test and shows up only here.
#
# `PS_REQUIRE_RUNTIME_CHECK` turns a skipped run into a failure, for the same reason as the canary:
# a verification check that skips reads exactly like one that passed.
- name: The runtime traversal check works through the built CLI
run: npx vitest run tests/protect/runtime-check-built.test.ts
env:
PS_REQUIRE_RUNTIME_CHECK: '1'

- name: Verify package contents
run: npm pack --dry-run

Expand Down
43 changes: 42 additions & 1 deletion AGENT-INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,46 @@ It is server-only. Never put it in the widget tag, client bundles, or public env

`setup` performs both steps automatically. The explicit commands are for manual setup or repair. If verification reports a generic or existing framework seam, complete the printed source edit and re-run `--check`; do not report protection as active until it exits successfully.

`--check` reads the app's source. It can establish that the guard is imported and called on a request
path; it cannot establish that a request ever reaches it — an app can wire the guard onto one server
and serve traffic from another, and that passes. To settle the difference there is an opt-in check
that **starts the application**:

```
npx @patchstack/connect protect --check --runtime
```

It launches the project's entry with `node`, moves the HTTP listeners **that process** opens to an
ephemeral loopback port, sends one request per listener carrying a per-run challenge, and reports
whether the scaffolded guard seam answered it. Exit `0` runtime traversal reached the seam, `1` a
listener answered and the seam did not, `2` it could not be established — neither a pass nor a
failure, with the structural checks still standing on their own.

Exit `2` is the answer for everything this cannot speak for, and the reason is always printed. The
common one is an entry that needs the project's own toolchain (a TypeScript entry, a framework
launcher, a watcher, another runtime, anything reached through a package manager), which this never
installs, builds or invents. The others are about scope: **the run answers for one process, one
thread, and one discovery window.** If the app attempts to start another process, the launch is
refused and the answer is `2`. A child can daemonize after it starts without declaring that in its
launch options, so allowing it would make the end-of-run process-group cleanup a claim the verifier
cannot establish. The app sees `EPERM`. A worker thread is also `2`: it inherits the listener
handling, but it cannot report back, so its listeners can be neither counted nor asked. So is a
listener that bound an address other than loopback, one that cannot be probed, and anything the app
opens after the discovery window has closed — the app is asked to stop and its acknowledgement is
what closes that window, so a run that never gets one is `2` as well. An inherited `NODE_OPTIONS`
that would run code before the listener handling is in place — a `--require` or `--import` in your
environment — is `2` too, and is refused before the app is launched rather than after.

A worker handed a replacement environment that does not preserve the propagated `NODE_OPTIONS` is
refused outright, because it would not load the listener handling. The app sees `EPERM`, and the run
reports `2`.

What a pass says is exactly: **runtime traversal reached the scaffolded guard seam.** It does not say
rules were delivered, that the deployed app is wired, or that ordinary traffic is blocked.

Nothing else runs the application. `protect`, `protect --check`, `setup`, `guide`, `scan`, `status`
and `mark-build` only read and write files.

5. **Commit** `.patchstackrc.json`, the updated `package.json`, the guard/framework source changes, and the layout/HTML file carrying the widget tag (and the production marker, when `scan` wrote one into a JSX root), so every developer and CI run reports to the same site.

**Do not commit `.patchstackrc.local.json`.** That file holds the API key issued at provision; the scan writes it and adds it to `.gitignore`, and tells you if it could not. `.patchstackrc.json` holds only the site UUID and settings, and the UUID is public by design — it ships in the widget tag in served HTML.
Expand Down Expand Up @@ -388,7 +428,8 @@ Two more endpoints the package can call, for completeness:
## Verifying the install

- `npx @patchstack/connect status` re-prints the site UUID and dashboard URL, and checks whether the site still exists on Patchstack (`Site status: active / removed / could not be verified`).
- `npx @patchstack/connect protect --check` verifies the runtime guard is connected to the request path.
- `npx @patchstack/connect protect --check` verifies from the source that the runtime guard is connected to the request path. It does not run the app.
- `npx @patchstack/connect protect --check --runtime` additionally **starts the app** on a loopback port and sends it one request, to establish that a request reaches the guard seam. Opt-in, and the only command that runs the application; exit `0`/`1`/`2` as described in step 4.
- Load the site in a browser — the "Report a vulnerability" button should appear. Refresh a page that was already open before the tag was added: the button only loads with the page.
- On the deployed site, the button appears only after a deploy that includes these source changes.

Expand Down
6 changes: 4 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,10 +61,12 @@ The one-line install prompt in `README.md` ("Install prompt (for AI coding tools
Invariants when touching it:

- The prompt appears in three places that must stay identical: `README.md`, `GETTING-STARTED.md` (the teammate-facing flow), and `field-test/prompt.txt` — `prompt.txt` is the artifact the harness tests.
- Any change to the prompt, the `guide` checklist output (`src/guide.ts`), or `AGENT-INSTALL.md` must pass `node field-test/run.mjs --persona hostile --rounds 3` before shipping. Agents audit the shipped docs, so inaccuracies in `AGENT-INSTALL.md` cost trust and cause refusals. **A round that never unpacked the tarball is VOID, not a pass and not a failure**: the shipped docs were never on disk, so no audit of them can have happened, and its scorecard is identical to a doc regression's. Unpacked means a non-empty `node_modules/@patchstack/connect/AGENT-INSTALL.md` — a dependency declaration in `package.json` is not an install. The harness retries void rounds within a bounded budget and exits 2 when every round was void — read that as "re-run", never as "the docs are fine". A doc change also wants a re-run after publication, when the published tarball actually carries it.
- **When the hostile run is the gate depends on what changed, because the fixture installs the PUBLISHED tarball.** A change to the prompt itself must pass `node field-test/run.mjs --persona hostile --rounds 3` *before shipping*: the prompt the harness sends comes from `field-test/prompt.txt` in this checkout, so a pre-publication run tests the real artifact. A change to `AGENT-INSTALL.md` or the `guide` checklist output (`src/guide.ts`) cannot be gated that way — the docs the agent audits are unpacked from the registry, so a pre-publication run audits the *previous* text and a green result says nothing about the change. Those changes require the same hostile run *immediately after the release that carries them*, and until then the local deterministic gates (the disclosure tests, `capabilities:check`) are what stand behind them. Ship such a change with the outstanding run named explicitly rather than reporting the gate as met. This split exists only because the harness has no local-registry mode; give it one and both become pre-publication gates.
- Agents audit the shipped docs, so inaccuracies in `AGENT-INSTALL.md` cost trust and cause refusals. **A round that never unpacked the tarball is VOID, not a pass and not a failure**: the shipped docs were never on disk, so no audit of them can have happened, and its scorecard is identical to a doc regression's. Unpacked means a non-empty `node_modules/@patchstack/connect/AGENT-INSTALL.md` — a dependency declaration in `package.json` is not an install. The harness retries void rounds within a bounded budget and exits 2 when every round was void — read that as "re-run", never as "the docs are fine".
- `hostile` measures whether the PROMPT survives pressure. It is a poor instrument for doc accuracy — void rounds are its modal outcome — so a docs-only verification wants a persona that reliably installs (`standard`, or `lovable`). Neither persona escapes the published-tarball problem above.
- Don't add reassurance language ("it's safe", "nothing is executed remotely") — agents flag it as a manipulation signal. Don't ask the agent to "follow the guide/instructions it prints" unbounded — name the concrete steps instead.
- A new real-world refusal report becomes a persona in `field-test/personas/`, written in your own words as a synthetic reconstruction, so the regression stays covered without publishing someone else's material.
- The fixture installs the *published* package, so an unpublished `guide`/CLI change can't be exercised end-to-end — publish first, or accept the run validates only the prompt shape.
- The fixture installs the *published* package, so an unpublished `guide`/CLI change can't be exercised end-to-end — publish first, or accept the run validates only the prompt shape. The free stubs (`field-test/README.md`) exercise the harness's own paths without an agent, and are worth running when the real gate is deferred: they show the harness is sound, and nothing about the docs.

## Comments and Explanations

Expand Down
14 changes: 11 additions & 3 deletions MAINTAINING.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,19 +22,27 @@ The deep "why" — the AI-agent refusal modes each clause guards against — liv

The prompt deliberately contains no model-authored verification step. In staged-command UIs, assistants cannot observe an `npm view` command until the user runs it; asking them to verify first caused fabricated registry findings and stopped the flow. Instead, the prompt records the user's explicit authorization for the disclosure widget and production-only build hooks before asking for one bounded `setup` command. Bolt's dependency-first fallback splits declaration from execution without adding hidden `postinstall` behavior. Keep `@patchstack/connect` discoverable in Patchstack's official product documentation so web-verifying platforms can independently establish its provenance.

## Before shipping a 🔴 or 🟠 change
## Verification timing for a 🔴 or 🟠 change

Run the gate:
For a prompt change, run the gate before shipping:

```bash
node field-test/run.mjs --persona hostile --rounds 3
```

It runs a real AI agent through the full install in a throwaway fixture against a mocked API and scores the outcome. Read each round's `report.md` (the HESITATIONS / DECISION ANALYSIS sections) — anything an agent pauses on is a future refusal. See [`field-test/README.md`](field-test/README.md) for the improve-and-retest loop, the safety model, and what to do when the gate is red for environmental reasons (stale published docs, release freshness).

For an `AGENT-INSTALL.md` or `src/guide.ts` change, the fixture installs the published tarball and
therefore cannot exercise unpublished text. Run the deterministic disclosure and capability checks
before merge, ship with the hostile field test recorded as outstanding, and run it immediately after the
release carries the change. That remains a prompt-survival gate. Add `standard` or `lovable` when the
question is document accuracy; those personas install more reliably than `hostile`.

What it establishes is **structural regression coverage**: the prompt still survives pressures that once broke it. The personas are synthetic — written from our own analysis, not from any platform's policy text — so a green run is not evidence that a live platform accepts the prompt, and should not be reported as though it were.

Caveat: the fixture installs the **published** package, so an unpublished `guide`/CLI change can't be exercised end-to-end — publish first, or accept that the run validates only the prompt shape.
This split is a limitation of the current harness. A local-registry mode would let every artifact be
tested before publication; until one exists, do not describe a run against the previous tarball as a
gate on unpublished docs.

## Don'ts (these are refusal triggers, not style nits)

Expand Down
70 changes: 70 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,13 @@ patchstack-connect protect Install/reconcile the always-
guard. Auto-wires supported server stacks;
use --check to verify or --demo for local rules.
Also run by setup; never run by scan/guide/mark-build.
--check reads your source and never runs the app.
--check --runtime STARTS THE APP on a loopback
port and sends it one request, to establish that
a request reaches the guard seam. Opt-in, and the
only mode that runs the app. Exit 0 traversed,
1 a listener answered instead of the guard,
2 could not be established (see below).
patchstack-connect map [--dir p] [--out f] [--upload]
Print a JSON map of this project's attack
surface: server entry points, the inputs each
Expand Down Expand Up @@ -103,6 +110,69 @@ Options (for demo and demo-guide):
(default: http://localhost:3000/api/tasks)
```

### Verifying the guard at runtime (opt-in)

`protect --check` reads the app's source. That establishes the guard is imported and called on a
request path — not that a request ever reaches it. An app can wire the guard onto one server and serve
its traffic from another, and the structural check passes.

`protect --check --runtime` settles that one question by **starting the application**:

```
npx @patchstack/connect protect --check --runtime
```

It runs the project's entry with `node`, moves the HTTP listeners **that process** opens to an ephemeral
loopback port (so a port already in use is not a failure), sends one request per listener carrying a
challenge generated for that run, and reports whether the scaffolded guard seam answered it. The child
is started in its own process group and killed with it — including if you interrupt the command. A
listener it cannot probe — a Unix socket, a file descriptor, a handed-over handle, an HTTP/2 server — is
reported and prevented from binding at all, rather than opened on the verifier's behalf.

**One process is the scope — one thread of it, and one discovery window.** If the app attempts to start
another process, the launch is refused and the answer is `2`, whatever that process is. A child can
daemonize after it starts without declaring that in its launch options, so allowing it would make the
end-of-run process-group cleanup a claim the verifier cannot establish. The app sees the launch fail
with `EPERM`. A **worker thread** is also `2`: it inherits the listener handling, but a worker has no
channel back, so its listeners can be neither counted nor asked. A worker handed a replacement
environment that does not preserve the propagated `NODE_OPTIONS` is refused, because it would not load
that handling at all.

Everything else the run finds also ends it this way: a listener that bound an address other than
loopback, a listener it cannot probe, and anything the app opens **after the discovery window closes** —
a second listener appearing while the first is still being asked cannot join a set that is already being
answered from. Closing that window is a handshake: the app is asked to stop opening listeners and its
acknowledgement is what proves nothing is still in flight, so a run that never gets one reports `2`
rather than passing. An inherited `NODE_OPTIONS` is checked before anything is launched, too: Node reads
that variable ahead of the command line, so a `--require` or `--import` sitting in your environment would
run before the listener handling was in place, and the run reports `2` rather than starting the app with
less containment than it claims. Recognised flags are passed through, with the reporter first.

A pass says exactly this: **runtime traversal reached the scaffolded guard seam.** It does not say
rules were delivered, that the deployed app is wired, or that ordinary traffic is blocked. The
challenge is generated per run, so a fixed response or a reflected header cannot answer it — but the
challenge does reach the whole app process, so this establishes traversal in a cooperating app rather
than against an app written to answer for itself.

| Exit | Meaning |
|---|---|
| `0` | A request reached the scaffolded guard seam. |
| `1` | A listener answered and the seam did not — or the structural checks failed, in which case the app is not started at all. |
| `2` | It could not be established. Neither a pass nor a failure; the structural checks still stand. |

Exit `2` is the common answer for entries this deliberately will not start. It runs `node <file>` on a
file the project already has, and nothing else — no package-manager scripts, no `node_modules/.bin`, no
build, no install. So a TypeScript entry, a framework launcher (`next start`), a watcher (`nodemon`),
another runtime (`bun`), or a script wrapped in an environment shim all report unavailable with the
reason printed. To make such a project verifiable, point a `start` script at a built, directly loadable
file — the check reports which entry it used and where it came from.

Windows reports exit `2` without starting anything: the cleanup this relies on is a POSIX process
group, and a verification that can leave a server running is worse than an unanswered question.

No other command runs your application. `protect`, `protect --check`, `setup`, `guide`, `scan`,
`status` and `mark-build` only read and write files.

## Configuration

Precedence (highest wins):
Expand Down
7 changes: 7 additions & 0 deletions scripts/copy-protect-templates.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -18,3 +18,10 @@ copyFileSync('src/protect/protect.d.ts', 'dist/protect.d.ts');
// so there is nothing per-format to express.
copyFileSync('src/protect/protect.d.ts', 'dist/protect.d.cts');
console.log('copied protect types -> dist/protect.d.ts, dist/protect.d.cts');

// The listener reporter `protect --check --runtime` preloads into the app it starts. Copied rather than
// bundled: it is loaded by path into ANOTHER process, as CommonJS, so it has to exist as a file that
// `node --require` can take. The built-CLI runtime test asserts the CLI can find it after a build.
mkdirSync('dist/protect/runtime', { recursive: true });
copyFileSync('src/protect/install/runtime/report-listeners.cjs', 'dist/protect/runtime/report-listeners.cjs');
console.log('copied listener reporter -> dist/protect/runtime/report-listeners.cjs');
Loading