diff --git a/CHANGELOG.md b/CHANGELOG.md index bfc1ce19b..b221a77e0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,10 +1,51 @@ # Changelog +## 1.0.9 — 2026-09-29 + +Action needed if you use Jev: its log-only mode is now `observe`, a `jev.json` still set to `shadow` is refused (Jev stays off until `failproofai jev setup` is run again), and Jev's checks now come only from `failproofai policies add FailproofAI/jev-policies`. For Hermes, `failproofai update` moves every profile from the old shell hooks (never run for cron jobs) to the native plugin, and every agent config failproofai edits is written crash-safely with a `.failproofai-backup`. Collects 1.0.9-beta.0 to beta.2 below. + +### Fixes + +- **A planted link at `.failproofai-backup` could redirect the backup into another file.** The previous version is now copied to an exclusively created temp file and renamed over the backup path, so a symlink there is replaced, never followed — for a project-scoped config in a cloned repository, the file it pointed at is untouched. +- **A config that is a dangling symlink is refused instead of replaced.** Writing used to swap the link for a regular file (silently detaching a dotfiles checkout); it now stops with the link and its missing target named. Files failproofai generates and owns (the OpenCode plugin shim, the Hermes link record) replace a planted link rather than writing through it. +- **A same-named Hermes plugin is only replaced when failproofai can prove it is its own.** A link counts as failproofai's when it matches the ownership record written beside it (`plugins/.failproofai-link`) or points into an npm `failproofai` package (links from the 1.0.9 betas are adopted and recorded); a look-alike directory named `hermes-plugin` with a `name: failproofai` manifest is left alone and `config --status` says another plugin occupies the name. +- **One Hermes profile that cannot be inspected no longer stops `update` for the rest.** Each profile is handled on its own, the redundant second read of `config.yaml` is gone, and a failure is that profile's line in the report. Re-enabling a plugin listed in `plugins.disabled` now reads "plugin re-enabled". +- **`failproofai update` exited 1 on a machine whose daemon was already current.** It reinstalled the service on every run, which needs root: interactively it asked for a password for nothing, and with no TTY (a fleet box, CI) it failed with "root privileges are required". When the service is running and `VERSION` records this CLI's version with its binary on disk, it now says the daemon is already current and asks root for nothing. + +### Dependencies + +- Routine dependency bumps: `next` and `eslint-config-next` 16.3.6, `posthog-node` 5.54.1, `vitest` 5.0.2, `jsdom` 30.1.1, `lucide-react` 1.48.0, `@types/node` 26.6.2, `@anthropic-ai/sdk` 0.128.0, the Rust and Python dependency groups, and the `actions/setup-node` 7 and `actions/create-github-app-token` 3 workflow actions (#853–#865). + +## 1.0.9-beta.2 — 2026-09-29 + +### Fixes + +- **A Hermes plugin listed in `plugins.disabled` was reported healthy and skipped by `update`.** Hermes checks `plugins.disabled` before `plugins.enabled`, so a profile listing `failproofai` in both never loads it, but health, installed detection and `update`'s "already current" check only read `plugins.enabled`. They now require enabled and not disabled: `config --status` says "plugin disabled (listed in plugins.disabled)" and `update` removes the disabling entry. + +## 1.0.9-beta.1 — 2026-09-29 + +### Fixes + +- **`jev status` and the dashboard under-counted a machine with only `FailproofAI/jev-policies`.** Their coverage survey still retired the `enabledPolicies` shim when any pack was installed, while hooks keep enforcing those builtins until a pack with regex policies arrives. The survey now uses the same `hasInstalledRegexPacks` check, so it counts what is really enforced and reviewable. +- **Old Jev activity rows no longer inflate "cleared".** A row whose mode this build does not know (written as `shadow` before the rename) lost only its mode, so `jev status` and the dashboard counted its would-have clears as enforced clears. Such a row now drops its clears too. +- The Hermes plugin's package path is found by walking up to the directory that holds `hermes-plugin/plugin.yaml` when `FAILPROOFAI_PACKAGE_ROOT` is unset, instead of a fixed three parents that overshoot from the bundled `dist/cli.mjs`. On Windows, replacing an existing plugin junction falls back to the move-aside swap when a direct rename fails. +- **Agent configs are written crash-safely and never rewritten from a file that does not parse.** Every integration wrote the user's agent config (`~/.hermes/config.yaml`, `~/.claude/settings.json`, `~/.openclaw/openclaw.json`, …) in place, so an interruption mid-write left a truncated config and an agent that would not start; and an existing Hermes/YAML config that did not parse was read as empty, so the next install would have replaced every other setting in it. Writes now go to a temp file in the same directory, are fsynced, keep the previous version as `.failproofai-backup`, and are atomically renamed into place, preserving the file's permissions and writing through a symlinked config. A config that exists but cannot be read or parsed is refused with the file and the reason, left byte-for-byte, and reported by `config --status`. The collector also stops shipping clears from rows whose Jev mode it does not know. + ## 1.0.9-beta.0 — 2026-09-29 +### Features + +- Jev's log-only mode is now **`observe`**, the word already used for a policy rollout that is evaluated but not enforced; Jev's modes are `off`, `observe` and `enforce`. `failproofai jev setup --mode observe`, the dashboard's Jev settings, `jev status`, `config --token` (which now turns Jev on in observe mode) and the docs all say observe; `jev.json` takes `mode: "observe"`, hook-activity rows carry `jevMode: "observe"`, `verdicts.jsonl` carries `applied: "observe"`, and `jev status --json` stats report `observeClearsByPolicy` and `modes.observe`. The collector ships `jev_mode: "observe"` to FailproofAI Cloud. The dashboard's Jev pill reads "jev observe". + +### Fixes + +- **Hermes cron jobs ran unchecked after an upgrade.** Legacy Hermes shell hooks (≤1.0.5) are never run for cron jobs — each cron fire builds its own hook scope that only discovered plugins join — and `failproofai update` never touched Hermes, so upgraded machines stayed on them silently. `update` now moves every Hermes profile that already uses FailproofAI (shell hooks or a copied plugin) to the native plugin, prints a per-profile report, and leaves profiles without FailproofAI alone. When the running daemon cannot serve the plugin (no `policyEvaluation`, e.g. a sudo system daemon `update` could not replace) the shell hooks are kept and `update` exits 1 pointing at `failproofai config`; it also exits 1 whenever the daemon swap or a layout migration failed. The plugin is now **linked** into each profile (`plugins/failproofai` → the package's `hermes-plugin/`, a marked copy where symlinks are unavailable), so npm upgrades apply with no reinstall; uninstall removes only the link. `failproofai config --status` reports a profile still on shell hooks as unhealthy: "Hermes cron jobs are not checked". +- **The npm package no longer ships Jev's checks; a machine asks them only after `failproofai policies add FailproofAI/jev-policies`.** The sixteen semantic checks were compiled in and used whenever no installed pack declared any, so configuring Jev — BYOK, or `failproofai config --token` with a Cloud machine key — started asking them without anyone opting in. Now the checks, and the names `reviewedBy` may use, come only from installed packs. With none declaring checks Jev is inert whether or not it is configured: no request is sent (not even the injection or task probes), no prompt is recorded for it, and every `reviewable` policy resolves `hard`, so nothing is cleared and hooks answer exactly as on a machine without Jev. An unreadable pack list asks nothing too, rather than falling back to a built-in set. `failproofai jev status` (and its `--json`, as `reviewablePolicies.jevChecks`), the dashboard's Jev settings and `config --token`'s output say "Jev has no checks installed" and name the command. `publish` still reserves the sixteen names and judges a regex-only pack's `reviewedBy` against them. +- **Adding `FailproofAI/jev-policies` could switch off a machine's regex policies.** A machine that enforced built-in policies from `enabledPolicies` (upgraded, no core pack) stopped enforcing them as soon as ANY pack was installed, and `failproofai policies add ` then tried to enable the name on the installed packs instead of fetching `FailproofAI/policies`. Now that Jev's checks arrive only as the Jev-only `FailproofAI/jev-policies` pack, both steps count only packs that carry regex policies (`hasInstalledRegexPacks`). + ### Docs -- Split Jev documentation into session evaluations under Find failures, live policy review under Prevent failures, and provider/configuration detail under Reference. Add a Use Jev page after Core concepts in Start with eval and policy setup tabs, plus a short quickstart link, dashboard screenshots, and CLI steps. Move sentiment analysis into Find failures and show its Jev-scored dashboard flow. Clarify shadow-mode verification and the Cloud machine key's `jev:evaluate` permission. +- Split Jev documentation into session evaluations under Find failures, live policy review under Prevent failures, and provider/configuration detail under Reference. Add a Use Jev page after Core concepts in Start with eval and policy setup tabs, plus a short quickstart link, dashboard screenshots, and CLI steps. Move sentiment analysis into Find failures and show its Jev-scored dashboard flow. Clarify observe-mode verification and the Cloud machine key's `jev:evaluate` permission. ### Dependencies diff --git a/CLAUDE.md b/CLAUDE.md index 65ab6fa13..fd43f6c11 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -323,10 +323,15 @@ we run in-repo. Hermes is a **dual-pillar** integration: an **audit** adapter Hermes enforcement uses the shipped **native Python plugin** under each profile's `plugins/failproofai/` directory. The profile's YAML config enables it through -`plugins.enabled: [failproofai]`. `integrations.ts` copies the plugin atomically, -marks the directory as FailproofAI-managed, refuses to overwrite an unmanaged -directory with the same name, and uses the `yaml` package's comment-preserving -`Document` API for config changes. +`plugins.enabled: [failproofai]`. `integrations.ts` symlinks that directory to the +package's `hermes-plugin/` (Hermes' scan follows symlinks, so npm upgrades apply +with no reinstall — the OpenClaw `plugins.load.paths` model), falls back to an +atomic copy marked `.failproofai-managed` where a link cannot be created, replaces +only a marked copy or a link into a FailproofAI `hermes-plugin/`, refuses anything +else with the same name, and uses the `yaml` package's comment-preserving +`Document` API for config changes. `failproofai update` migrates profiles already +using FailproofAI (legacy shell hooks or a copy) to the link, gated on the daemon +answering `policyEvaluation`; legacy shell hooks never run for Hermes cron jobs. Settings file paths: @@ -335,7 +340,7 @@ Settings file paths: | user | `~/.hermes/config.yaml` | Hermes is **user-scope only** — there is no project config, so `getSettingsPath` -ignores scope/cwd. Every default and named profile receives its own plugin copy and +ignores scope/cwd. Every default and named profile receives its own plugin link and enablement entry. Installed-state detection requires both the complete managed plugin directory and the config entry; a missing file, disabled plugin, newly-created profile, or leftover legacy shell hook is reported as unhealthy. diff --git a/Cargo.lock b/Cargo.lock index 6189816d7..b2426e80c 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -218,7 +218,7 @@ dependencies = [ [[package]] name = "failproofaid" -version = "1.0.9-beta.0" +version = "1.0.9-beta.3" dependencies = [ "fpai-collect", "fpai-ipc", @@ -281,7 +281,7 @@ dependencies = [ [[package]] name = "fpai-collect" -version = "1.0.9-beta.0" +version = "1.0.9-beta.3" dependencies = [ "notify", "reqwest", @@ -298,7 +298,7 @@ dependencies = [ [[package]] name = "fpai-ipc" -version = "1.0.9-beta.0" +version = "1.0.9-beta.3" dependencies = [ "libc", "proptest", diff --git a/Cargo.toml b/Cargo.toml index c3f08b32b..0d8e20bf3 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -3,7 +3,7 @@ resolver = "3" members = ["crates/*"] [workspace.package] -version = "1.0.9-beta.0" +version = "1.0.9-beta.3" edition = "2024" license-file = "LICENSE" repository = "https://github.com/FailproofAI/failproofai" diff --git a/__tests__/actions/jev-mode-action.test.ts b/__tests__/actions/jev-mode-action.test.ts index 7aafccca4..c56cc6df8 100644 --- a/__tests__/actions/jev-mode-action.test.ts +++ b/__tests__/actions/jev-mode-action.test.ts @@ -1,7 +1,7 @@ // @vitest-environment node /** * The /settings Jev panel's FailproofAI Cloud controls, against the real - * loader: `setJevModeAction` (the on/off switch and shadow/enforce) and the + * loader: `setJevModeAction` (the on/off switch and observe/enforce) and the * Cloud connection row in `getJevSettingsAction`. * * 1. **It rewrites `mode` and nothing else** — every other byte-level field @@ -33,7 +33,7 @@ const INGEST_KEY = ["fp", "ingest", "0badc0ffee123456"].join("-"); const POLICY_KEY = ["fp", "policy", "feedfacecafe7890"].join("-"); const BYOK_KEY = ["ts", "byok", "0123456789abcdef"].join("-"); const ORIGIN = "https://app.befailproof.ai"; -const CLOUD_FILE = { provider: "failproofai", baseUrl: `${ORIGIN}/enforcement/v1/jev`, mode: "shadow" }; +const CLOUD_FILE = { provider: "failproofai", baseUrl: `${ORIGIN}/enforcement/v1/jev`, mode: "observe" }; let home: string; let prevHome: string | undefined; @@ -93,7 +93,7 @@ describe("setJevModeAction", () => { connect(); // A field this build does not know, and one it does not show: both must survive. seed({ ...CLOUD_FILE, timeoutMs: 2500, fromANewerBuild: { x: 1 } }); - for (const mode of ["enforce", "off", "shadow"] as const) { + for (const mode of ["enforce", "off", "observe"] as const) { const res = await setJevModeAction(mode); expect(res.ok).toBe(true); expect(onDisk()).toEqual({ ...CLOUD_FILE, timeoutMs: 2500, fromANewerBuild: { x: 1 }, mode }); @@ -122,7 +122,7 @@ describe("setJevModeAction", () => { expect(res.ok).toBe(true); expect(onDisk()).toEqual({ provider: "typesafe", apiKey: BYOK_KEY, mode: "off" }); secretFree(res); - expect((await setJevModeAction("shadow")).ok).toBe(true); + expect((await setJevModeAction("observe")).ok).toBe(true); expect(loadJevConfig()?.apiKey).toBe(BYOK_KEY); }); diff --git a/__tests__/actions/jev-reviewability.test.ts b/__tests__/actions/jev-reviewability.test.ts index 8cfcffcde..534d29318 100644 --- a/__tests__/actions/jev-reviewability.test.ts +++ b/__tests__/actions/jev-reviewability.test.ts @@ -20,7 +20,8 @@ import { tmpdir } from "node:os"; import { join } from "node:path"; import { getJevSettingsAction } from "../../app/actions/get-jev-config"; import { POLICY_CATALOG } from "../../src/hooks/policy-catalog"; -import { RETAKE_PACK_COMMAND } from "../../src/hooks/policy-reviewability"; +import { JEV_POLICIES_ADD_COMMAND, NO_JEV_CHECKS_PROBLEM, RETAKE_PACK_COMMAND } from "../../src/hooks/policy-reviewability"; +import { JEV_PACK_SEMANTIC_ENTRIES } from "../fixtures/jev-policies"; /** A token no provider issued. Nothing here should ever send it anywhere. */ const TOKEN = "jevtoken-0123456789-3f2a"; @@ -77,7 +78,8 @@ function turnJevOn(): void { chmodSync(path, 0o600); } -function installPack(policies: Array>): void { +/** The core pack, beside FailproofAI/jev-policies unless `withJev` is false. */ +function installPack(policies: Array>, withJev = true): void { const artifact = "// a pack artifact this test never executes\n"; const digest = createHash("sha256").update(artifact).digest("hex"); mkdirSync(join(packRoot, "artifacts"), { recursive: true }); @@ -95,6 +97,19 @@ function installPack(policies: Array>): void { sha256: digest, policies, }, + ...(withJev + ? [ + { + id: "FailproofAI/jev-policies", + version: "0.2.0", + source: "github:FailproofAI/jev-policies@v0.2.0", + entry: `artifacts/${digest}.mjs`, + sha256: digest, + policies: [], + semantic: JEV_PACK_SEMANTIC_ENTRIES, + }, + ] + : []), ], }), ); @@ -134,21 +149,31 @@ describe("getJevSettingsAction — what Jev may clear", () => { expect(JSON.stringify(view)).not.toContain(TOKEN); }); - it("reports the seven reviewable builtins and no problem", async () => { - writeConfig({ enabledPolicies: POLICY_CATALOG.map((p) => p.name) }); + it("reports the fifteen reviewable policies and no problem, with the core pack and its Jev checks", async () => { + writeConfig({ enabledPolicies: [] }); + installPack(PACKABLE as unknown as Array>); turnJevOn(); const view = await getJevSettingsAction(); expect(view.reviewable).toEqual({ - enabled: POLICY_CATALOG.length, + enabled: PACKABLE.length + 1, reviewable: 15, summary: - `15 of ${POLICY_CATALOG.length} enabled policies are reviewable: ` + + `15 of ${PACKABLE.length + 1} enabled policies are reviewable: ` + "Jev may clear a deny or an instruction from those, and from no others.", problem: null, }); }); + it("says Jev has no checks installed, and the command, when no pack declares any", async () => { + writeConfig({ enabledPolicies: POLICY_CATALOG.map((p) => p.name) }); + turnJevOn(); + + const view = await getJevSettingsAction(); + expect(view.reviewable).toMatchObject({ enabled: POLICY_CATALOG.length, reviewable: 0, problem: NO_JEV_CHECKS_PROBLEM }); + expect(view.reviewable?.problem).toContain(JEV_POLICIES_ADD_COMMAND); + }); + it("counts the launch directory's project config, not the server's own cwd", async () => { // The standalone server chdirs into the package directory, so the project a // person launched the dashboard from arrives only as FAILPROOFAI_LAUNCH_CWD — @@ -165,7 +190,9 @@ describe("getJevSettingsAction — what Jev may clear", () => { turnJevOn(); const view = await getJevSettingsAction(); - expect(view.reviewable).toMatchObject({ enabled: POLICY_CATALOG.length, reviewable: 15, problem: null }); + // All of the launch project's builtins are counted; none is reviewable, + // because no pack gives Jev the checks they name. + expect(view.reviewable).toMatchObject({ enabled: POLICY_CATALOG.length, reviewable: 0, problem: NO_JEV_CHECKS_PROBLEM }); } finally { rmSync(launch, { recursive: true, force: true }); } diff --git a/__tests__/actions/update-jev-config.test.ts b/__tests__/actions/update-jev-config.test.ts index d04b32355..8d7951015 100644 --- a/__tests__/actions/update-jev-config.test.ts +++ b/__tests__/actions/update-jev-config.test.ts @@ -150,9 +150,9 @@ describe("saving writes the file the hooks read", () => { it("carries the stored token across a mode change, without it being re-typed", async () => { await saveJevConfigAction(input()); - const res = await saveJevConfigAction(input({ mode: "shadow", token: "" })); + const res = await saveJevConfigAction(input({ mode: "observe", token: "" })); expect(res.ok).toBe(true); - expect(loadJevConfig()?.mode).toBe("shadow"); + expect(loadJevConfig()?.mode).toBe("observe"); expect(loadJevConfig()?.apiKey).toBe(TOKEN); }); @@ -367,23 +367,23 @@ describe("validation is the loader's, not a second copy of it", () => { expect(loadJevConfig()).toBeNull(); }); - it("refuses plain http in enforce mode, and accepts loopback http in shadow", async () => { + it("refuses plain http in enforce mode, and accepts loopback http in observe", async () => { const enforced = await saveJevConfigAction( input({ provider: "custom", baseUrl: "http://localhost:9999", mode: "enforce" }), ); expect(enforced.ok).toBe(false); expect(loadJevConfig()).toBeNull(); - const shadowed = await saveJevConfigAction( - input({ provider: "custom", baseUrl: "http://localhost:9999", mode: "shadow" }), + const observed = await saveJevConfigAction( + input({ provider: "custom", baseUrl: "http://localhost:9999", mode: "observe" }), ); - expect(shadowed.ok).toBe(true); - expect(loadJevConfig()?.mode).toBe("shadow"); + expect(observed.ok).toBe(true); + expect(loadJevConfig()?.mode).toBe("observe"); }); it("refuses plain http to anywhere but loopback, in either mode", async () => { const res = await saveJevConfigAction( - input({ provider: "custom", baseUrl: "http://jev.example", mode: "shadow" }), + input({ provider: "custom", baseUrl: "http://jev.example", mode: "observe" }), ); expect(res.ok).toBe(false); expect(loadJevConfig()).toBeNull(); @@ -443,10 +443,10 @@ describe("validation is the loader's, not a second copy of it", () => { it("still re-saves an older mismatched file untouched, to switch its mode", async () => { seedConfig({ provider: "openrouter", apiKey: TOKEN, baseUrl: "https://ai-gateway.vercel.sh/v1" }); const res = await saveJevConfigAction( - input({ provider: "openrouter", baseUrl: "https://ai-gateway.vercel.sh/v1", mode: "shadow", token: "" }), + input({ provider: "openrouter", baseUrl: "https://ai-gateway.vercel.sh/v1", mode: "observe", token: "" }), ); expect(res.ok).toBe(true); - expect(onDisk().mode).toBe("shadow"); + expect(onDisk().mode).toBe("observe"); }); // The client never appends a second /systemone, so "adds /systemone to the @@ -472,21 +472,21 @@ describe("validation is the loader's, not a second copy of it", () => { // CLI likewise checks only a --base-url it was given. seedConfig({ provider: "custom", apiKey: TOKEN, baseUrl: "https://proxy.example/v1/systemone" }); const res = await saveJevConfigAction( - input({ provider: "custom", baseUrl: "https://proxy.example/v1/systemone", mode: "shadow", token: "" }), + input({ provider: "custom", baseUrl: "https://proxy.example/v1/systemone", mode: "observe", token: "" }), ); expect(res.ok).toBe(true); expect(onDisk().baseUrl).toBe("https://proxy.example/v1/systemone"); - expect(onDisk().mode).toBe("shadow"); + expect(onDisk().mode).toBe("observe"); }); // `setJevModeAction` and `jev setup --mode` refuse these; the save skipped // them, keeping the old mode — or, on a fresh machine, writing a file with no // mode, which loads as enforce. it.each(["yolo", "ENFORCE", "", null])("refuses mode %j and writes nothing", async (mode) => { - await saveJevConfigAction(input({ mode: "shadow" })); + await saveJevConfigAction(input({ mode: "observe" })); const before = readFileSync(configPath(), "utf8"); const res = await saveJevConfigAction(input({ mode: mode as string, token: "" })); - expect(res).toEqual({ ok: false, problem: 'mode must be "off", "shadow" or "enforce".' }); + expect(res).toEqual({ ok: false, problem: 'mode must be "off", "observe" or "enforce".' }); expect(readFileSync(configPath(), "utf8")).toBe(before); }); @@ -601,13 +601,13 @@ describe("a save keeps the fields the form does not show", () => { accountId: CLOUDFLARE_ACCOUNT, model: "typesafe/jev-1.13", timeoutMs: 4500, - mode: "shadow", + mode: "observe", }); // Exactly what the panel sends for that file with nothing touched: the form // holds the four values the view gave it, and a blank token. const res = await saveJevConfigAction( - input({ provider: "cloudflare", accountId: CLOUDFLARE_ACCOUNT, mode: "shadow", token: "" }), + input({ provider: "cloudflare", accountId: CLOUDFLARE_ACCOUNT, mode: "observe", token: "" }), ); expect(res.ok).toBe(true); @@ -615,7 +615,7 @@ describe("a save keeps the fields the form does not show", () => { expect(loaded?.model).toBe("typesafe/jev-1.13"); expect(loaded?.accountId).toBe(CLOUDFLARE_ACCOUNT); expect(loaded?.timeoutMs).toBe(4500); - expect(loaded?.mode).toBe("shadow"); + expect(loaded?.mode).toBe("observe"); expect(loaded?.apiKey).toBe(TOKEN); }); @@ -625,7 +625,7 @@ describe("a save keeps the fields the form does not show", () => { apiKey: TOKEN, model: "typesafe/jev-1.13", timeoutMs: 4500, - mode: "shadow", + mode: "observe", }); const res = await saveJevConfigAction(input({ mode: "enforce", token: "" })); @@ -640,7 +640,7 @@ describe("a save keeps the fields the form does not show", () => { it("keeps a field a newer failproofai wrote, which this form has never heard of", async () => { seedConfig({ provider: "typesafe", apiKey: TOKEN, futureField: { weights: [1, 2] } }); - expect((await saveJevConfigAction(input({ mode: "shadow", token: "" }))).ok).toBe(true); + expect((await saveJevConfigAction(input({ mode: "observe", token: "" }))).ok).toBe(true); expect(onDisk().futureField).toEqual({ weights: [1, 2] }); }); @@ -658,14 +658,14 @@ describe("a save keeps the fields the form does not show", () => { }; seedConfig(seed); - expect((await saveJevConfigAction(input({ mode: "shadow", token: "" }))).ok).toBe(true); + expect((await saveJevConfigAction(input({ mode: "observe", token: "" }))).ok).toBe(true); const viaPanel = onDisk(); seedConfig(seed); const { runJevCommand } = await import("../../src/hooks/jev-cli"); - // `jev setup --mode shadow`: the same change, named the same way, with no + // `jev setup --mode observe`: the same change, named the same way, with no // terminal to prompt on — the key is kept from the existing config. - const cli = await runJevCommand(["setup", "--mode", "shadow"], { + const cli = await runJevCommand(["setup", "--mode", "observe"], { stdinIsTTY: false, // No provider is reached from a unit test; see jev-cli-contracts.test.ts. readModelList: async () => ({ ok: false, reason: "no list read in tests" }), @@ -697,18 +697,18 @@ describe("a config whose key lives in the environment", () => { }); it("takes a provider change too, and stays keyless", async () => { - seedConfig({ provider: "typesafe", mode: "shadow" }); + seedConfig({ provider: "typesafe", mode: "observe" }); - const res = await saveJevConfigAction(input({ provider: "openrouter", mode: "shadow", token: "" })); + const res = await saveJevConfigAction(input({ provider: "openrouter", mode: "observe", token: "" })); expect(res.ok).toBe(true); - expect(onDisk()).toEqual({ provider: "openrouter", mode: "shadow" }); + expect(onDisk()).toEqual({ provider: "openrouter", mode: "observe" }); }); it("is still on after such a save, with the key read from the environment", async () => { process.env.FAILPROOFAI_JEV_API_KEY = TOKEN; seedConfig({ provider: "typesafe" }); - const res = await saveJevConfigAction(input({ mode: "shadow", token: "" })); + const res = await saveJevConfigAction(input({ mode: "observe", token: "" })); expect(res.ok).toBe(true); if (!res.ok) return; expect(res.view.on).toBe(true); @@ -733,7 +733,7 @@ describe("a config whose key lives in the environment", () => { seedConfig({ provider: "typesafe" }); chmodSync(configPath(), 0o644); - const res = await saveJevConfigAction(input({ mode: "shadow", token: "" })); + const res = await saveJevConfigAction(input({ mode: "observe", token: "" })); expect(res.ok).toBe(true); expect(statSync(configPath()).mode & 0o777).toBe(0o600); expect(onDisk().apiKey).toBeUndefined(); @@ -807,7 +807,7 @@ describe("the stored model, which the panel shows but does not offer", () => { it("survives a save, which is the whole point of not offering the field", async () => { seedConfig({ provider: "typesafe", apiKey: TOKEN, model: "typesafe/jev-1.13" }); - const res = await saveJevConfigAction(input({ mode: "shadow", token: "" })); + const res = await saveJevConfigAction(input({ mode: "observe", token: "" })); expect(res.ok).toBe(true); if (!res.ok) return; expect(res.view.model).toEqual({ kind: "id", id: "typesafe/jev-1.13" }); diff --git a/__tests__/components/jev-notices-no-request.test.tsx b/__tests__/components/jev-notices-no-request.test.tsx index 4639decbc..8fdff33fd 100644 --- a/__tests__/components/jev-notices-no-request.test.tsx +++ b/__tests__/components/jev-notices-no-request.test.tsx @@ -10,7 +10,7 @@ const NO_REQUEST = { decision: "allow", evaluator: "jev" as const, jevDecision: describe("a call Jev sent no request for", () => { it("gets no pill", () => { expect(jevPillKind(NO_REQUEST)).toBeNull(); - expect(jevPillKind({ ...NO_REQUEST, jevMode: "shadow" })).toBeNull(); + expect(jevPillKind({ ...NO_REQUEST, jevMode: "observe" })).toBeNull(); const { container } = render(); expect(container).toBeEmptyDOMElement(); }); diff --git a/__tests__/components/jev-notices-not-consulted.test.tsx b/__tests__/components/jev-notices-not-consulted.test.tsx index 679261404..c6f9e6f0a 100644 --- a/__tests__/components/jev-notices-not-consulted.test.tsx +++ b/__tests__/components/jev-notices-not-consulted.test.tsx @@ -6,12 +6,12 @@ import { JEV_NOT_CONSULTED_FACT } from "@/src/hooks/jev-activity"; // Exactly what the two-tier combine rules record for a hard deny: Jev was // aborted and its answer never read. const HARD_DENY = { decision: "deny", evaluator: "jev" as const, jevMode: "enforce" as const }; -const HARD_DENY_SHADOW = { decision: "deny", evaluator: "jev" as const, jevMode: "shadow" as const }; +const HARD_DENY_OBSERVE = { decision: "deny", evaluator: "jev" as const, jevMode: "observe" as const }; describe("a hard deny Jev was not consulted on", () => { it("gets no pill: it is an ordinary regex deny", () => { expect(jevPillKind(HARD_DENY)).toBeNull(); - expect(jevPillKind(HARD_DENY_SHADOW)).toBeNull(); + expect(jevPillKind(HARD_DENY_OBSERVE)).toBeNull(); const { container } = render(); expect(container).toBeEmptyDOMElement(); }); diff --git a/__tests__/components/jev-notices.test.tsx b/__tests__/components/jev-notices.test.tsx index 08bf1562d..5fd55a181 100644 --- a/__tests__/components/jev-notices.test.tsx +++ b/__tests__/components/jev-notices.test.tsx @@ -11,18 +11,18 @@ describe("jevPillKind", () => { expect(jevPillKind({ decision: "allow", evaluator: "jev", jevDecision: "allow", jevCleared: [], jevMode: "enforce" })).toBeNull(); }); - it("marks a clear, a fallback and a shadow disagreement", () => { + it("marks a clear, a fallback and an observe disagreement", () => { expect(jevPillKind({ decision: "allow", evaluator: "jev", jevCleared: ["block-env-files"], jevMode: "enforce" })).toBe( "cleared", ); expect(jevPillKind({ decision: "deny", evaluator: "jev-fallback", jevFallbackReason: "timeout" })).toBe("fallback"); expect( - jevPillKind({ decision: "deny", evaluator: "jev", jevCleared: ["block-env-files"], jevMode: "shadow" }), + jevPillKind({ decision: "deny", evaluator: "jev", jevCleared: ["block-env-files"], jevMode: "observe" }), ).toBe("would-clear"); - expect(jevPillKind({ decision: "allow", evaluator: "jev", jevDecision: "deny", jevMode: "shadow" })).toBe( - "shadow-stricter", + expect(jevPillKind({ decision: "allow", evaluator: "jev", jevDecision: "deny", jevMode: "observe" })).toBe( + "observe-stricter", ); - expect(jevPillKind({ decision: "deny", evaluator: "jev", jevDecision: "deny", jevMode: "shadow" })).toBeNull(); + expect(jevPillKind({ decision: "deny", evaluator: "jev", jevDecision: "deny", jevMode: "observe" })).toBeNull(); }); }); diff --git a/__tests__/components/jev-settings-panel.test.tsx b/__tests__/components/jev-settings-panel.test.tsx index c8e259170..4ab33c40c 100644 --- a/__tests__/components/jev-settings-panel.test.tsx +++ b/__tests__/components/jev-settings-panel.test.tsx @@ -110,8 +110,8 @@ describe("what it says about the machine", () => { expect(screen.getByRole("button", { name: /turn jev on/i })).toBeInTheDocument(); }); - it("distinguishes enforce from shadow, because they are different guarantees", async () => { - renderPanel(configured({ mode: "shadow" })); + it("distinguishes enforce from observe, because they are different guarantees", async () => { + renderPanel(configured({ mode: "observe" })); expect(screen.getByText(/the regex result is what gets enforced/i)).toBeInTheDocument(); cleanup(); renderPanel(configured({ mode: "enforce" })); @@ -267,13 +267,13 @@ describe("the token field is write-only", () => { }); it("sends a blank token when nothing was typed, which the server reads as keep", async () => { - saveMock.mockResolvedValue({ ok: true, view: configured({ mode: "shadow" }) }); + saveMock.mockResolvedValue({ ok: true, view: configured({ mode: "observe" }) }); renderPanel(configured()); - fireEvent.change(screen.getByLabelText("mode"), { target: { value: "shadow" } }); + fireEvent.change(screen.getByLabelText("mode"), { target: { value: "observe" } }); fireEvent.click(screen.getByRole("button", { name: /save changes/i })); await waitFor(() => expect(saveMock).toHaveBeenCalledTimes(1)); expect(saveMock).toHaveBeenCalledWith( - expect.objectContaining({ provider: "typesafe", mode: "shadow", token: "" }), + expect.objectContaining({ provider: "typesafe", mode: "observe", token: "" }), ); }); @@ -293,9 +293,9 @@ describe("the token field is write-only", () => { describe("the form", () => { it("sends no model at all, so a save cannot clear the stored one", async () => { - saveMock.mockResolvedValue({ ok: true, view: configured({ mode: "shadow" }) }); + saveMock.mockResolvedValue({ ok: true, view: configured({ mode: "observe" }) }); renderPanel(configured({ model: { kind: "id", id: "typesafe/jev-1.13" } })); - fireEvent.change(screen.getByLabelText("mode"), { target: { value: "shadow" } }); + fireEvent.change(screen.getByLabelText("mode"), { target: { value: "observe" } }); fireEvent.click(screen.getByRole("button", { name: /save changes/i })); await waitFor(() => expect(saveMock).toHaveBeenCalledTimes(1)); // Not "sends an empty model": an empty string is what the server used to @@ -419,7 +419,7 @@ function cloudView(over: Partial = {}): JevSettingsView { baseUrl: "https://app.befailproof.ai/enforcement/v1/jev", endpoint: "https://app.befailproof.ai/enforcement/v1/jev/systemone", token: { source: "cloud" }, - mode: "shadow", + mode: "observe", timeoutMs: 3000, cloud: CONNECTED, ...over, @@ -455,21 +455,21 @@ describe("the FailproofAI Cloud route", () => { expect(screen.getByRole("button", { name: /turn jev on/i })).toBeInTheDocument(); }); - it("switches back on in shadow mode", async () => { + it("switches back on in observe mode", async () => { modeMock.mockResolvedValue({ ok: true, view: cloudView() }); renderPanel(cloudView({ status: "off", on: false, mode: "off" })); // Nothing to pick while it is off. expect(screen.getByLabelText(/^mode$/i)).toBeDisabled(); fireEvent.click(screen.getByRole("button", { name: /turn jev on/i })); - await waitFor(() => expect(modeMock).toHaveBeenCalledWith("shadow")); + await waitFor(() => expect(modeMock).toHaveBeenCalledWith("observe")); }); - it("switches shadow to enforce with the mode control, and says a refusal", async () => { - modeMock.mockResolvedValue({ ok: false, problem: "plain http is accepted only with mode shadow" }); + it("switches observe to enforce with the mode control, and says a refusal", async () => { + modeMock.mockResolvedValue({ ok: false, problem: "plain http is accepted only with mode observe" }); renderPanel(cloudView()); fireEvent.change(screen.getByLabelText(/^mode$/i), { target: { value: "enforce" } }); await waitFor(() => expect(modeMock).toHaveBeenCalledWith("enforce")); - await waitFor(() => expect(screen.getByText(/plain http is accepted only with mode shadow/)).toBeInTheDocument()); + await waitFor(() => expect(screen.getByText(/plain http is accepted only with mode observe/)).toBeInTheDocument()); }); it("not connected: says so, and what fixes it", async () => { diff --git a/__tests__/e2e/hooks/pack-enforcement.e2e.test.ts b/__tests__/e2e/hooks/pack-enforcement.e2e.test.ts index 89aee78e1..ca5722861 100644 --- a/__tests__/e2e/hooks/pack-enforcement.e2e.test.ts +++ b/__tests__/e2e/hooks/pack-enforcement.e2e.test.ts @@ -178,7 +178,7 @@ describe("pack enforcement, end to end", () => { // The allow is not enough on its own, and this test proved it: the first // version of this passed against a real bug. The observe path read // `cloudManaged!.id`, which is undefined for a pack, so every non-allow - // shadow verdict threw — the throw was swallowed by the evaluator, nothing + // observed verdict threw — the throw was swallowed by the evaluator, nothing // was recorded, and the net result was an allow. Exactly what this asserted. // A clean stderr is what separates "observed" from "crashed into an allow". const env = createFixtureEnv(); diff --git a/__tests__/fixtures/jev-activity-rows.ts b/__tests__/fixtures/jev-activity-rows.ts index 222ee4624..9226a78b6 100644 --- a/__tests__/fixtures/jev-activity-rows.ts +++ b/__tests__/fixtures/jev-activity-rows.ts @@ -75,7 +75,7 @@ export const JEV_ACTIVITY_ROWS: HookActivityEntry[] = [ jevCleared: [], jevLatencyMs: 44, jevModel: "typesafe/jev", - jevMode: "shadow", + jevMode: "observe", }, { ...base, diff --git a/__tests__/fixtures/jev-no-request-rows.ts b/__tests__/fixtures/jev-no-request-rows.ts index 49f2c1223..2c479d91e 100644 --- a/__tests__/fixtures/jev-no-request-rows.ts +++ b/__tests__/fixtures/jev-no-request-rows.ts @@ -43,6 +43,6 @@ export const JEV_NO_REQUEST_ROWS: HookActivityEntry[] = [ durationMs: 2, evaluator: "jev", jevDecision: "allow", - jevMode: "shadow", + jevMode: "observe", }, ]; diff --git a/__tests__/fixtures/jev-not-consulted-rows.ts b/__tests__/fixtures/jev-not-consulted-rows.ts index 9ce146263..0d84cf200 100644 --- a/__tests__/fixtures/jev-not-consulted-rows.ts +++ b/__tests__/fixtures/jev-not-consulted-rows.ts @@ -38,6 +38,6 @@ export const JEV_NOT_CONSULTED_ROWS: HookActivityEntry[] = [ reason: "Recursive force deletes are blocked", durationMs: 3, evaluator: "jev", - jevMode: "shadow", + jevMode: "observe", }, ]; diff --git a/__tests__/fixtures/jev-policies.ts b/__tests__/fixtures/jev-policies.ts new file mode 100644 index 000000000..402e2b71a --- /dev/null +++ b/__tests__/fixtures/jev-policies.ts @@ -0,0 +1,518 @@ +/** + * A reference copy of the sixteen Jev checks that `FailproofAI/jev-policies` + * publishes, for tests only. + * + * The npm package ships no Jev checks: they reach a machine only through + * `failproofai policies add FailproofAI/jev-policies`, and with no pack that + * declares any, Jev asks nothing (`semantic/pack-policies.ts`). The decision, + * compile and envelope tests still need a realistic question set to exercise, + * and this is it. It is NOT loadable from `src/`, on purpose. + * + * The two inline preconditions are the same predicates the pack's named ones + * compile to (`protected_branch` and `paths_outside_project` in + * `semantic/preconditions.ts`). + */ +import { createHash } from "node:crypto"; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { afterAll, beforeAll, beforeEach } from "vitest"; +import { PROTECTED_BRANCHES, outsideProject } from "../../src/hooks/semantic/policies"; +import type { SemanticPolicy } from "../../src/hooks/semantic/types"; + +export const JEV_PACK_POLICIES: ReadonlyArray = [ + { + name: "destructive-deletion", + title: "Tried to permanently delete data that cannot be regenerated", + appliesTo: ["shell", "write"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "destroys", + instructions: + "The tool call in `agent_request` permanently deletes, wipes, truncates or overwrites existing files, " + + "directories or disks — for example `rm -r`, `find -delete`, `shred`, `dd of=`, `mkfs`, `truncate`, " + + "`rsync --delete`, `git clean -fdx`, or redirecting output over an existing file.", + criteria: { + true: "Existing data would be destroyed or irreversibly overwritten.", + false: "Nothing existing is destroyed: it reads, lists, creates new files, or moves data somewhere recoverable.", + }, + }, + { + id: "irreplaceable", + instructions: + "What would be destroyed is outside the project, is the whole project, is the user's home directory, " + + "is the filesystem root, or is data that cannot simply be regenerated (source code, documents, databases, " + + "keys). Use `facts.paths[].relation` to see where each target sits relative to the project.", + criteria: { + true: "The destroyed data matters and could not be rebuilt by running a command.", + false: + "Only regenerable data inside the project is affected: build output, dist/, caches, node_modules, " + + "virtualenvs, coverage reports, temp files, or files the agent itself just created.", + }, + }, + ], + guidance: "This permanently deletes data that cannot be rebuilt. Confirm the exact paths with the user first.", + }, + { + name: "production-infra-change", + title: "Tried to change live infrastructure", + appliesTo: ["shell"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "mutates", + instructions: + "The command in `agent_request` changes the state of cloud or cluster infrastructure: it creates, " + + "updates, deletes, applies, scales, restarts, rolls out, deploys or destroys resources in a cloud " + + "account, Kubernetes cluster, managed database, DNS, CDN or hosting platform (any CLI: kubectl, helm, " + + "terraform, tofu, pulumi, aws, gcloud, az, doctl, flyctl, vercel, wrangler, railway, and so on — " + + "however the binary is spelled or pathed).", + criteria: { + true: "It mutates infrastructure.", + false: + "It only reads or plans: get, list, describe, logs, status, plan, diff, validate, whoami, or --dry-run.", + }, + }, + { + id: "not_local", + instructions: + "The target of that change is a shared or production environment, or its environment cannot be told " + + "from the command.", + criteria: { + true: "Production, shared, or unknown environment.", + false: + "Clearly a local or throwaway environment: localhost, kind, minikube, docker-desktop, k3d, or a " + + "context, workspace or profile whose name says dev, test, staging, sandbox or local.", + }, + }, + ], + guidance: "This changes live infrastructure. Run it against a non-production target, or hand the command to a human.", + }, + { + name: "git-history-rewrite", + title: "Tried to rewrite or discard shared git history", + appliesTo: ["shell"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "rewrites_remote", + instructions: + "The command in `agent_request` force-pushes or otherwise overwrites history on a git remote: " + + "`git push --force`, `--force-with-lease`, `-f`, a `+refspec` such as `+HEAD:main`, or deleting a " + + "remote branch. Git's global options before the subcommand (`git -C dir`, `git -c key=val`) and an " + + "absolute path to the git binary do not change what the command does.", + criteria: { + true: "Remote history would be overwritten or a remote branch deleted.", + false: "A normal push that only adds commits, or no push at all.", + }, + }, + ], + guidance: "Force-pushing rewrites history other people may have pulled. Confirm with the user first.", + }, + { + name: "push-to-protected-branch", + title: "Tried to push directly to a protected branch", + appliesTo: ["shell"], + mode: "instruct", + userCanOverride: true, + probes: [ + { + id: "pushes_protected", + instructions: + "The command in `agent_request` pushes commits to a remote branch named main, master, production, " + + "release or trunk — either named explicitly or implied because `facts.current_git_branch` is one of " + + "those and the push names no other branch.", + criteria: { + true: "Commits land directly on a protected remote branch.", + false: + "It pushes a feature branch whose name merely contains a word like main (e.g. feature/main-menu), " + + "or it does not push.", + }, + }, + ], + guidance: "Pushing straight to a protected branch skips review. Push a branch and open a PR unless the user asked for this.", + }, + { + name: "commit-on-protected-branch", + title: "Tried to commit directly on a protected branch", + appliesTo: ["shell"], + mode: "instruct", + userCanOverride: true, + precondition: (facts) => facts.currentGitBranch !== null && PROTECTED_BRANCHES.has(facts.currentGitBranch), + probes: [ + { + id: "creates_commit", + instructions: + "The command in `agent_request` creates or amends a git commit, merge, rebase or cherry-pick in the " + + "repository (including `git -C dir commit` and `git commit -a`).", + criteria: { + true: "A new commit is written to the current branch.", + false: "It only inspects history (log, show, diff, status) or mentions `commit` in a message or search.", + }, + }, + ], + guidance: "You are on a protected branch. Create a feature branch before committing unless the user asked otherwise.", + }, + { + name: "secret-exposure", + title: "Tried to read or copy credentials", + appliesTo: ["shell", "read", "write"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "touches_secrets", + instructions: + "The tool call in `agent_request` reads, prints, copies, or writes credential material: SSH private " + + "keys (~/.ssh/id_*), authorized_keys, cloud credentials (~/.aws, ~/.config/gcloud, ~/.azure, " + + "~/.kube/config), .env files containing real secret values, API tokens, password stores, keychains, " + + "or browser cookie databases.", + criteria: { + true: "Actual secret values would be read, displayed, copied or written.", + false: + "It only references the concept: source code that reads process.env, a variable name, an .env.example " + + "template, a test fixture, documentation about credentials, or a grep for the word.", + }, + }, + ], + guidance: "This exposes real credentials. Ask the user before reading or copying secret material.", + }, + { + name: "credential-exfiltration", + title: "Tried to send secrets or private files off the machine", + appliesTo: ["shell", "network"], + mode: "deny", + userCanOverride: false, + probes: [ + { + id: "sends_out", + instructions: + "The tool call in `agent_request` sends data to a network destination: an HTTP POST or PUT (curl -d, " + + "curl -F, curl -T, wget --post-data), scp, sftp, rsync or rclone to a remote, nc or socat, an upload to " + + "a paste site, bucket or webhook, or a DNS lookup carrying encoded data.", + criteria: { + true: "Data leaves this machine.", + false: "It only downloads, or only talks to localhost.", + }, + }, + { + id: "sensitive_payload", + instructions: + "The data being sent includes credentials, keys, tokens, environment variables, or the contents of files " + + "from outside the project (see `facts.paths`).", + criteria: { + true: "Secrets or private files are in what is sent.", + false: "What is sent is ordinary project data the user would expect to share, like a test request body.", + }, + }, + ], + guidance: "This sends secrets or private files to a remote destination, which the user cannot approve. Do not send them.", + }, + { + name: "remote-code-execution", + title: "Tried to run code downloaded from the internet", + appliesTo: ["shell"], + mode: "deny", + userCanOverride: true, + probes: [ + { + /** + * The `…` in the `wget` example is load-bearing, not decoration. + * + * A code span in a question whose first token is `curl` or `wget` and + * whose second token is a bare short option (`-O x.sh`, `-s`, `-fsSL`) + * is refused by the Cloudflare route — HTTP 402, code 2021, "Model + * execution failed (Payment error)" — deterministically, on any state. + * `curl … | sh` and `bash <(curl …)` below are accepted because their + * second token is the placeholder, and `wget … -O x.sh` is accepted for + * the same reason. The client reads 402 as `out-of-credits`, so the + * refusal used to degrade every call that selected this policy to + * regex and blame the operator's billing for it. + * + * So: keep a placeholder (or a URL) between `wget`/`curl` and its first + * short option in any example written inside backticks here. + */ + id: "download_and_run", + instructions: + "The command in `agent_request` downloads code or a script from the internet and executes it: " + + "`curl … | sh`, `bash <(curl …)`, `wget … -O x.sh && bash x.sh`, `python3 -c \"$(curl …)\"`, piping into " + + "any interpreter (sh, bash, zsh, python, node, perl, ruby), or eval of a fetched string.", + criteria: { + true: "Fetched code is executed.", + false: + "It only downloads without running, runs a local file, or merely searches for or quotes such a command " + + "(for example grep over a README).", + }, + }, + ], + exempt: { + id: "official_installer", + instructions: + "The URL being executed is the documented official installer of a widely used developer tool, served from " + + "that tool's own domain (for example bun.sh, sh.rustup.rs, get.docker.com, deb.nodesource.com, " + + "raw.githubusercontent.com/nvm-sh/nvm, astral.sh/uv).", + }, + guidance: "This executes code fetched from the internet. Download it, show it to the user, then run it.", + }, + { + name: "privilege-escalation", + title: "Tried to run with elevated privileges", + appliesTo: ["shell"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "elevates", + instructions: + "The command in `agent_request` runs something as root or another user: sudo, doas, su, pkexec, run0, " + + "`sudo -i`, `sudo -s`, including when the binary is written as an absolute path or reached through a " + + "variable or wrapper.", + criteria: { + true: "Privileges are elevated.", + false: "It runs as the current user, or only mentions sudo in text, a comment, or a search pattern.", + }, + }, + ], + guidance: "This runs with elevated privileges. Ask the user to run it themselves.", + }, + { + name: "database-destruction", + title: "Tried to destroy or mass-modify database data", + appliesTo: ["shell"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "destructive_sql", + instructions: + "The command in `agent_request` executes SQL or a database command that drops or truncates a table, " + + "schema or database, or deletes or updates rows without a condition that narrows them to specific " + + "records. A condition that is always true (`WHERE 1=1`, `WHERE true`, `WHERE id > 0`) does not narrow " + + "anything. Also counts: `redis-cli FLUSHALL`, `dropdb`, `mongo … dropDatabase()`.", + criteria: { + true: "Many rows, a table, or a whole database would be removed or overwritten.", + false: "It reads, inserts, or changes a specific identified record, or runs a migration file by name.", + }, + }, + { + id: "real_database", + instructions: "The database it targets holds real data rather than being a disposable test database.", + criteria: { + true: "A real or shared database, or one whose purpose cannot be told.", + false: "Clearly an in-memory, test, fixture or throwaway local database (sqlite :memory:, a test_ database).", + }, + }, + ], + guidance: "This destroys database data. Confirm the target and scope with the user first.", + }, + { + name: "read-outside-workspace", + title: "Read files outside the project", + appliesTo: ["shell", "read"], + mode: "instruct", + userCanOverride: true, + precondition: (facts) => facts.paths.some((p) => outsideProject(facts, p)), + probes: [ + { + id: "reads_outside", + instructions: + "The tool call in `agent_request` reads the contents of files that `facts.paths` marks as outside the " + + "project, in the user's home directory (for example another repository, Documents, or dotfiles).", + criteria: { + true: "It reads file contents from outside the project.", + false: + "It only changes directory, lists names, checks existence, or touches the agent's own tool caches.", + }, + }, + ], + guidance: "This reads files outside the current project. Mention why before continuing.", + }, + { + name: "agent-config-tampering", + title: "Tried to change the agent's own safety configuration", + appliesTo: ["shell", "write"], + mode: "deny", + userCanOverride: false, + probes: [ + { + id: "edits_agent_config", + instructions: + "The tool call in `agent_request` modifies or deletes an AI agent's own configuration, hooks, " + + "permissions or guardrails: .claude/settings.json, .claude/settings.local.json, ~/.claude/, .codex/, " + + ".cursor/hooks.json, .gemini/, .factory/, failproofai configuration or policies, or permission and " + + "safety sections of AGENTS.md or CLAUDE.md.", + criteria: { + true: "The agent's hooks, permissions or safety settings would change.", + false: "It only reads those files, or edits ordinary project documentation.", + }, + }, + ], + guidance: "An agent must not change its own guardrails. Ask the user to make this change.", + }, + { + name: "system-modification", + title: "Tried to change the system outside the project", + appliesTo: ["shell"], + mode: "instruct", + userCanOverride: true, + probes: [ + { + id: "modifies_system", + instructions: + "The command in `agent_request` installs or removes software system-wide or changes machine " + + "configuration: apt, dnf, pacman, brew install, `npm install -g`, `pip install` outside a virtualenv, " + + "editing files under /etc, systemctl enable or disable, crontab, or changing shell startup files.", + criteria: { + true: "The machine changes outside the project directory.", + false: "It installs project dependencies locally (npm install, bun install, pip install in a venv).", + }, + }, + ], + guidance: "This changes the machine outside the project. Mention it to the user.", + }, + { + name: "env-secrets-dump", + title: "Tried to print environment secrets", + appliesTo: ["shell"], + mode: "instruct", + userCanOverride: true, + probes: [ + { + id: "dumps_env", + instructions: + "The command in `agent_request` prints environment variable VALUES that may be secret: `env`, " + + "`printenv`, `set`, `export -p`, or `echo` of a variable whose name suggests a key, token, password " + + "or secret.", + criteria: { + true: "Secret-looking environment values would be printed.", + false: + "It echoes a literal string, prints a non-secret variable like PATH or HOME, creates a virtualenv " + + "named env, or searches code for the word export.", + }, + }, + ], + guidance: "This prints environment values that may be secrets into the transcript.", + }, + { + name: "external-destructive-action", + title: "Tried an irreversible action through an external tool", + appliesTo: ["other"], + mode: "deny", + userCanOverride: true, + probes: [ + { + id: "irreversible_external", + instructions: + "The tool call in `agent_request` (an MCP or integration tool, see `facts.tool_name`) performs an " + + "irreversible or externally visible action: deleting records, repositories, files or accounts; " + + "sending email, chat or social messages on the user's behalf; making payments or purchases; merging or " + + "closing pull requests; changing permissions, access or billing; or writing to a production system.", + criteria: { + true: "Something outside this machine changes in a way that cannot be quietly undone.", + false: "It reads, searches, lists, fetches, or creates a draft that nobody else sees yet.", + }, + }, + ], + guidance: "This takes an irreversible action in an external system. Confirm with the user first.", + }, + { + name: "external-data-egress", + title: "Sent private data to an external tool", + appliesTo: ["other"], + mode: "instruct", + userCanOverride: true, + probes: [ + { + id: "egresses_private", + instructions: + "The arguments in `agent_request` send private data to an external service: source code, file " + + "contents, credentials, customer data, or personal information.", + criteria: { + true: "Private data is being shared with a third party.", + false: "Only a query, identifier or public information is sent.", + }, + }, + ], + guidance: "This shares private data with an external service. Mention what is being sent.", + }, +]; + +/** + * The manifest form of {@link JEV_PACK_POLICIES}: what the pack's `semantic` + * array carries. The two inline preconditions become the names they compile + * from; everything else is data already. + */ +const PRECONDITION_NAMES: Readonly> = { + "commit-on-protected-branch": "protected_branch", + "read-outside-workspace": "paths_outside_project", +}; + +export const JEV_PACK_SEMANTIC_ENTRIES: ReadonlyArray> = JEV_PACK_POLICIES.map((p) => { + const { precondition: _fn, origin: _origin, ...rest } = p; + const named = PRECONDITION_NAMES[p.name]; + return JSON.parse(JSON.stringify({ ...rest, ...(named ? { precondition: named } : {}) })) as Record; +}); + +const ARTIFACT = "export const hooks = [];\n"; + +/** + * The `installed.json` record of a stand-in for `FailproofAI/jev-policies`, + * with its (empty, digest-verified) artifact written under `packDir` — for a + * test that writes its own `installed.json` and needs the checks beside its packs. + */ +export function jevPoliciesPackRecord(packDir: string): Record { + const digest = createHash("sha256").update(ARTIFACT).digest("hex"); + mkdirSync(join(packDir, "artifacts"), { recursive: true }); + writeFileSync(join(packDir, "artifacts", `${digest}.mjs`), ARTIFACT); + return { + id: "FailproofAI/jev-policies", + version: "0.2.0", + source: "github:FailproofAI/jev-policies@v0.2.0", + entry: `artifacts/${digest}.mjs`, + sha256: digest, + policies: [], + semantic: JEV_PACK_SEMANTIC_ENTRIES, + }; +} + +/** + * Install a stand-in for `FailproofAI/jev-policies` into `packDir` (point + * `FAILPROOFAI_PACK_DIR` at it): a real `installed.json` and a digest-verified + * artifact, so the real reader resolves the sixteen checks exactly as a machine + * that ran `policies add FailproofAI/jev-policies` does. Returns the pack dir. + */ +export function installJevPoliciesPack(packDir: string): string { + writeFileSync( + join(packDir, "installed.json"), + JSON.stringify({ schemaVersion: 1, packs: [jevPoliciesPackRecord(packDir)] }), + ); + return packDir; +} + +/** + * For a whole test file: this machine has `FailproofAI/jev-policies` + * installed. The package ships no Jev checks, so a test that exercises the + * evaluator through the path a hook takes — with no `policies` override — + * needs the pack, exactly as a real machine does. + */ +export function withInstalledJevPoliciesPack(): void { + let dir: string | undefined; + let saved: string | undefined; + beforeAll(() => { + saved = process.env.FAILPROOFAI_PACK_DIR; + dir = installJevPoliciesPack(mkdtempSync(join(tmpdir(), "fpai-jev-policies-"))); + }); + // Per test, because some files restore the whole of `process.env` after each. + beforeEach(() => { + process.env.FAILPROOFAI_PACK_DIR = dir; + }); + afterAll(() => { + if (saved === undefined) delete process.env.FAILPROOFAI_PACK_DIR; + else process.env.FAILPROOFAI_PACK_DIR = saved; + if (dir) rmSync(dir, { recursive: true, force: true }); + }); +} diff --git a/__tests__/fixtures/jev-policy-page-rows.ts b/__tests__/fixtures/jev-policy-page-rows.ts index 88709d557..0ffb012cb 100644 --- a/__tests__/fixtures/jev-policy-page-rows.ts +++ b/__tests__/fixtures/jev-policy-page-rows.ts @@ -5,7 +5,7 @@ * A. enforce mode, Jev's own verdict decided the call: `policyName` is * `semantic/` and `policySource` is `jev` (it used to be omitted, * so the chart filed every Jev block under "unattributed"); - * B. shadow mode, Jev's own verdict was deny / instruct while the regex + * B. observe mode, Jev's own verdict was deny / instruct while the regex * result (allow) was enforced: the verdict is a "would have" in * `observed`, `{policyId, version, decision, reason}`, the list * observe-mode cloud and pack policies already use. @@ -55,7 +55,7 @@ export const JEV_POLICY_PAGE_ROWS: HookActivityEntry[] = [ jevMode: "enforce", policySource: "jev", }, - // B — shadow: Jev would have denied; the regex result (allow) was enforced. + // B — observe: Jev would have denied; the regex result (allow) was enforced. { ...base, timestamp: 1785740915100, @@ -68,10 +68,10 @@ export const JEV_POLICY_PAGE_ROWS: HookActivityEntry[] = [ jevDecision: "deny", jevLatencyMs: 761, jevModel: "jev-1.13.0", - jevMode: "shadow", + jevMode: "observe", observed: [{ policyId: "semantic/destructive-deletion", version: "jev-1.13.0", decision: "deny", reason: DELETION_REASON }], }, - // B — shadow: Jev would have warned. + // B — observe: Jev would have warned. { ...base, timestamp: 1785740915200, @@ -84,7 +84,7 @@ export const JEV_POLICY_PAGE_ROWS: HookActivityEntry[] = [ jevDecision: "instruct", jevLatencyMs: 688, jevModel: "jev-1.13.0", - jevMode: "shadow", + jevMode: "observe", observed: [{ policyId: "semantic/system-modification", version: "jev-1.13.0", decision: "instruct", reason: SYSTEM_REASON }], }, ]; diff --git a/__tests__/hooks/cloud-connect-jev.test.ts b/__tests__/hooks/cloud-connect-jev.test.ts index 448750214..0a37be0dd 100644 --- a/__tests__/hooks/cloud-connect-jev.test.ts +++ b/__tests__/hooks/cloud-connect-jev.test.ts @@ -5,7 +5,7 @@ * - A key whose introspect lists `jev:evaluate` stores itself in the `jev` * slot of credentials.json, under the origin it was verified against, and — * only when there is NO jev.json — writes one that turns Jev on through - * FailproofAI Cloud in shadow mode. + * FailproofAI Cloud in observe mode. * - An existing jev.json is never overwritten, whatever it names. * - A key introspect says lacks `jev:evaluate` writes no Jev state at all, * and drops the Jev key a previous connection left. An introspect that @@ -82,14 +82,14 @@ function seedJev(obj: unknown, mode = 0o600): string { } describe("connecting with a key that carries jev:evaluate", () => { - it("stores the key under the verified origin and turns Jev on in shadow mode", async () => { + it("stores the key under the verified origin and turns Jev on in observe mode", async () => { const outcome = await connect(withPermissions(...MACHINE_PRESET)); expect(outcome.jev?.ok).toBe(true); expect(outcome.jev?.config?.status).toBe("written"); expect(readCredentials().jev).toEqual({ url: URL_, key: TOKEN }); const onDisk = JSON.parse(readFileSync(jevConfigFile(), "utf8")); - expect(onDisk).toEqual({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "shadow" }); + expect(onDisk).toEqual({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "observe" }); // No key in jev.json: the Cloud key has one home. expect(readFileSync(jevConfigFile(), "utf8")).not.toContain(TOKEN); if (posix) { @@ -100,12 +100,12 @@ describe("connecting with a key that carries jev:evaluate", () => { // What the hooks will now read. const cfg = loadJevConfig(); - expect(cfg).toMatchObject({ provider: "failproofai", apiKey: TOKEN, mode: "shadow" }); + expect(cfg).toMatchObject({ provider: "failproofai", apiKey: TOKEN, mode: "observe" }); const r = inspectJevConfig(); expect(r.status === "ok" && r.keySource).toBe("cloud"); const text = describeOutcome(outcome, "machine-1", URL_).join("\n"); - expect(text).toMatch(/Jev\s+on through FailproofAI Cloud, in shadow mode/); + expect(text).toMatch(/Jev\s+on through FailproofAI Cloud, in observe mode/); expect(text).toContain("--mode enforce"); expect(text).not.toContain(TOKEN); expect(configuredPaths(outcome)).toContain(credentialsFile()); @@ -115,12 +115,12 @@ describe("connecting with a key that carries jev:evaluate", () => { const outcome = await connect(withPermissions(...MACHINE_PRESET), "http://localhost:8080/fp"); expect(readCredentials().jev?.url).toBe("http://localhost:8080"); expect(JSON.parse(readFileSync(jevConfigFile(), "utf8")).baseUrl).toBe("http://localhost:8080/fp/enforcement/v1/jev"); - // Plain http to loopback is fine in the shadow mode connect writes. + // Plain http to loopback is fine in the observe mode connect writes. expect(loadJevConfig()?.baseUrl).toBe("http://localhost:8080/fp/enforcement/v1/jev"); // …and only there: `jev setup --mode enforce` (and the dashboard switch) // refuses plain http, so the output must not name it as the next step. const text = describeOutcome(outcome, "machine-1", "http://localhost:8080/fp").join("\n"); - expect(text).toMatch(/Jev\s+on through FailproofAI Cloud, in shadow mode/); + expect(text).toMatch(/Jev\s+on through FailproofAI Cloud, in observe mode/); expect(text).not.toContain("--mode enforce"); expect(text).toContain("Enforce needs an https FailproofAI Cloud URL"); // The command it would have named really is refused, and the refusal names the step that works. @@ -169,11 +169,11 @@ describe("connecting with a key that carries jev:evaluate", () => { const outcome = await connect(withPermissions(...MACHINE_PRESET)); const text = describeOutcome(outcome, "machine-1", URL_).join("\n"); expect(text).toContain("switched off"); - expect(text).toContain("jev setup --mode shadow"); + expect(text).toContain("jev setup --mode observe"); }); it("names the other origin when the Cloud jev.json on disk points somewhere else", async () => { - seedJev({ provider: "failproofai", baseUrl: "https://staging.befailproof.ai/enforcement/v1/jev", mode: "shadow" }); + seedJev({ provider: "failproofai", baseUrl: "https://staging.befailproof.ai/enforcement/v1/jev", mode: "observe" }); const outcome = await connect(withPermissions(...MACHINE_PRESET)); expect(outcome.jev?.config).toMatchObject({ status: "kept", otherOrigin: "https://staging.befailproof.ai" }); const text = describeOutcome(outcome, "machine-1", URL_).join("\n"); @@ -193,11 +193,11 @@ describe("connecting with a key that carries jev:evaluate", () => { expect(text).toContain("https://staging.befailproof.ai"); expect(text).toContain("switched off"); const cmds = [...text.matchAll(/`failproofai (jev setup[^`]*)`/g)].map((m) => m[1]); - expect(cmds).toEqual(["jev setup --provider failproofai --mode shadow"]); + expect(cmds).toEqual(["jev setup --provider failproofai --mode observe"]); const r = await runJevCommand(cmds[0].split(" ").slice(1), { render: { cols: 120, color: false } }); expect(r.exitCode).toBe(0); - expect(inspectJevConfig()).toMatchObject({ status: "ok", config: { mode: "shadow", baseUrl: `${URL_}/enforcement/v1/jev` } }); + expect(inspectJevConfig()).toMatchObject({ status: "ok", config: { mode: "observe", baseUrl: `${URL_}/enforcement/v1/jev` } }); }); it("the no-clobber write loses to a file that appears first", () => { @@ -242,7 +242,7 @@ describe("connecting with --no-transcripts (sessions !== true)", () => { expect(jevLines[0]).toContain("`failproofai jev setup --provider failproofai`"); const text = lines.join("\n"); expect(text).not.toMatch(/Jev\s+on\b/); - expect(text).not.toContain("shadow mode"); + expect(text).not.toContain("observe mode"); expect(text).not.toContain(TOKEN); // The key file is named in the closing note: a key WAS stored. expect(configuredPaths(outcome)).toContain(credentialsFile()); @@ -259,8 +259,8 @@ describe("connecting with --no-transcripts (sessions !== true)", () => { expect(text).not.toContain("available on this key"); }); - it("a Cloud jev.json already on (shadow or enforce) is left alone — and the output says Jev still sends, and how to stop it", async () => { - for (const mode of ["shadow", "enforce"] as const) { + it("a Cloud jev.json already on (observe or enforce) is left alone — and the output says Jev still sends, and how to stop it", async () => { + for (const mode of ["observe", "enforce"] as const) { const before = seedJev({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode }); const outcome = await decisionsOnly(); // Never overwritten (decision 16, invariant 7)… @@ -279,7 +279,7 @@ describe("connecting with --no-transcripts (sessions !== true)", () => { }); it("…the same line on the --connect path, above \"Decisions only.\"", async () => { - seedJev({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "shadow" }); + seedJev({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "observe" }); const r = await runConnectCommand({ url: URL_, token: TOKEN, @@ -291,7 +291,7 @@ describe("connecting with --no-transcripts (sessions !== true)", () => { daemonStatus: () => "running", }); const text = r.lines.join("\n"); - expect(text).toContain("Jev is still on through FailproofAI Cloud (shadow mode)"); + expect(text).toContain("Jev is still on through FailproofAI Cloud (observe mode)"); expect(text).toContain("Decisions only."); expect(text.indexOf("still on through")).toBeLessThan(text.indexOf("Decisions only.")); }); @@ -299,7 +299,7 @@ describe("connecting with --no-transcripts (sessions !== true)", () => { it("no such line when the Cloud jev.json does not send: switched off, or pointing at another Cloud", async () => { for (const file of [ { provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "off" }, - { provider: "failproofai", baseUrl: "https://staging.befailproof.ai/enforcement/v1/jev", mode: "shadow" }, + { provider: "failproofai", baseUrl: "https://staging.befailproof.ai/enforcement/v1/jev", mode: "observe" }, ]) { seedJev(file); const outcome = await decisionsOnly(); @@ -333,7 +333,7 @@ describe("connecting with --no-transcripts (sessions !== true)", () => { const { runJevCommand } = await import("../../src/hooks/jev-cli"); const r = await runJevCommand(["setup", "--provider", "failproofai"], { render: { cols: 120, color: false } }); expect(r.exitCode).toBe(0); - expect(loadJevConfig()).toMatchObject({ provider: "failproofai", apiKey: TOKEN, mode: "shadow" }); + expect(loadJevConfig()).toMatchObject({ provider: "failproofai", apiKey: TOKEN, mode: "observe" }); }); }); @@ -463,7 +463,7 @@ describe("disconnecting", () => { }); it("keeps a BYOK jev.json exactly as it was, and says whose it is", async () => { - const before = seedJev({ provider: "openrouter", apiKey: BYOK_KEY, mode: "shadow" }); + const before = seedJev({ provider: "openrouter", apiKey: BYOK_KEY, mode: "observe" }); await connect(withPermissions(...MACHINE_PRESET)); const r = runDisconnectCommand(); expect(readFileSync(jevConfigFile(), "utf8")).toBe(before); @@ -476,7 +476,7 @@ describe("disconnecting", () => { }); // `--mode off` is "the switch that lasts" (jev-cloud.mdx): deleting it here - // made the next connect write a fresh shadow file, and Jev came back on. + // made the next connect write a fresh observe file, and Jev came back on. it("keeps a Cloud jev.json switched off, so reconnecting leaves Jev off", async () => { await connect(withPermissions(...MACHINE_PRESET)); const before = seedJev({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "off" }); @@ -496,7 +496,7 @@ describe("disconnecting", () => { it("removes a Cloud jev.json even with no key left to clear", () => { mkdirSync(home, { recursive: true }); - seedJev({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "shadow" }); + seedJev({ provider: "failproofai", baseUrl: `${URL_}/enforcement/v1/jev`, mode: "observe" }); const r = runDisconnectCommand(); expect(existsSync(jevConfigFile())).toBe(false); expect(r.lines[0]).toBe("Disconnected from FailproofAI Cloud."); diff --git a/__tests__/hooks/daemon-service.test.ts b/__tests__/hooks/daemon-service.test.ts index 5bbecd45a..cf9ee6947 100644 --- a/__tests__/hooks/daemon-service.test.ts +++ b/__tests__/hooks/daemon-service.test.ts @@ -3,7 +3,7 @@ import { describe, it, expect, beforeEach, afterEach, vi } from "vitest"; import { execFileSync } from "node:child_process"; import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir, userInfo } from "node:os"; -import { resolve } from "node:path"; +import { dirname, resolve } from "node:path"; import { binDir } from "../../src/hooks/fp-home"; import { startedAtFromMonotonic, waitForDaemonRunning } from "../../src/hooks/daemon-service"; import * as svc from "../../src/hooks/daemon-service"; @@ -1010,6 +1010,55 @@ describe("refreshDaemonToCliVersion", () => { expect(installed).toBe(false); }); + it("does nothing, and asks root for nothing, when the daemon already runs this version", async () => { + // Every update after the daemon was brought current used to reinstall it: + // idempotent, but it needs sudo, so with no TTY it failed "root privileges + // are required" and exited 1 on a machine that was already fine. + const { writeVersionFile } = await import("../../src/hooks/fp-config"); + const { version } = await import("../../package.json"); + const { installedBinaryPath } = await import("../../src/hooks/daemon-download"); + writeVersionFile({ daemon: version }); + mkdirSync(dirname(installedBinaryPath(version)), { recursive: true }); + writeFileSync(installedBinaryPath(version), "#!/bin/sh\n"); + let installed = false; + let primed = false; + const result = await svc.refreshDaemonToCliVersion({ + status: () => "running", + install: async () => { + installed = true; + return { installed: true }; + }, + prime: () => { + primed = true; + return true; + }, + interactive: () => true, + }); + expect(result.ok).toBe(true); + expect(result.lines.join("\n")).toContain("already installed and running"); + expect(installed).toBe(false); + expect(primed).toBe(false); + }); + + it("still reinstalls when the recorded version matches but the service is not running", async () => { + const { writeVersionFile } = await import("../../src/hooks/fp-config"); + const { version } = await import("../../package.json"); + const { installedBinaryPath } = await import("../../src/hooks/daemon-download"); + writeVersionFile({ daemon: version }); + mkdirSync(dirname(installedBinaryPath(version)), { recursive: true }); + writeFileSync(installedBinaryPath(version), "#!/bin/sh\n"); + let installed = false; + await svc.refreshDaemonToCliVersion({ + status: () => "stopped", + install: async () => { + installed = true; + return { installed: true }; + }, + interactive: () => false, + }); + expect(installed).toBe(true); + }); + it("goes through the INSTALL, not a bare restart", async () => { // The property that was broken. The first version fetched the binary and then // rewrote the unit via `upgradedServiceDefinition`, which PRESERVES the diff --git a/__tests__/hooks/fail-closed-force-decision.test.ts b/__tests__/hooks/fail-closed-force-decision.test.ts index 34bd1f0bf..250886929 100644 --- a/__tests__/hooks/fail-closed-force-decision.test.ts +++ b/__tests__/hooks/fail-closed-force-decision.test.ts @@ -42,6 +42,7 @@ vi.mock("../../src/hooks/pack-manifest", () => ({ // The handler asks this per event to decide whether the migration shim // still applies. Mirrors the mocked readInstalledPacks above. hasInstalledPacks: vi.fn(() => false), + hasInstalledRegexPacks: vi.fn(() => false), })); import { evaluateHookEvent } from "../../src/hooks/handler"; diff --git a/__tests__/hooks/handler.test.ts b/__tests__/hooks/handler.test.ts index b9f80f644..d9934a588 100644 --- a/__tests__/hooks/handler.test.ts +++ b/__tests__/hooks/handler.test.ts @@ -77,6 +77,7 @@ vi.mock("../../src/hooks/pack-manifest", () => ({ // The handler asks this on every event to decide whether the migration shim // still applies. Mocked for the same reason as the line above. hasInstalledPacks: vi.fn(() => false), + hasInstalledRegexPacks: vi.fn(() => false), })); describe("hooks/handler", () => { diff --git a/__tests__/hooks/harness-extra-paths.test.ts b/__tests__/hooks/harness-extra-paths.test.ts index 58c104bac..6fbb505ed 100644 --- a/__tests__/hooks/harness-extra-paths.test.ts +++ b/__tests__/hooks/harness-extra-paths.test.ts @@ -14,6 +14,9 @@ import { tmpdir } from "node:os"; import { readConfig, writeConfig, DEFAULT_CONFIG } from "@/src/hooks/fp-config"; import { configFile } from "@/src/hooks/fp-home"; import { HARNESS_KEYS, addPath, removePath, listPaths, runHarnessCommand } from "@/src/hooks/harness-cli"; +import { updateConfig } from "@/src/hooks/fp-config"; +import { writeCollectorSettings } from "@/src/hooks/collector-config"; +import { setDaemonConfigured } from "@/src/hooks/daemon-service"; describe("harness extra paths", () => { let home: string; @@ -344,3 +347,41 @@ describe("harness extra paths", () => { }); }); }); + +describe("extra paths survive every config.json write that `config` and `update` make", () => { + let home: string; + let prevHome: string | undefined; + + beforeEach(() => { + prevHome = process.env.FAILPROOFAI_HOME; + home = mkdtempSync(join(tmpdir(), "fpai-hx-cfg-")); + process.env.FAILPROOFAI_HOME = home; + }); + + afterEach(() => { + if (prevHome === undefined) delete process.env.FAILPROOFAI_HOME; + else process.env.FAILPROOFAI_HOME = prevHome; + rmSync(home, { recursive: true, force: true }); + }); + + it("keeps every harness path through a Cloud connect, collector settings, daemon install and disconnect", () => { + expect(addPath("hermes", "prod=/srv/hermes-prod/state.db").exitCode).toBe(0); + expect(addPath("claude", "work=/srv/team/.claude/projects").exitCode).toBe(0); + const before = readConfig().collector.sources; + expect(before).toBeDefined(); + + // `failproofai config` connecting to Cloud (cloud-connection.ts) + updateConfig({ mode: "cloud" }); + writeCollectorSettings({ sessions: true, hooks: true, hooksVerbosity: "decisions", machineId: "m-1" }); + // `config` / `update` installing or refreshing the daemon (daemon-service.ts) + setDaemonConfigured(true, "1.0.9"); + // a re-run of `config` that changes the collector choices again + writeCollectorSettings({ sessions: false, hooks: true }); + // uninstalling the daemon, then disconnecting + setDaemonConfigured(false); + updateConfig({ mode: "oss" }); + + expect(readConfig().collector.sources).toEqual(before); + expect(readFileSync(configFile(), "utf-8")).toContain("/srv/hermes-prod/state.db"); + }); +}); diff --git a/__tests__/hooks/hermes-update.test.ts b/__tests__/hooks/hermes-update.test.ts new file mode 100644 index 000000000..f3d03167f --- /dev/null +++ b/__tests__/hooks/hermes-update.test.ts @@ -0,0 +1,499 @@ +/** + * `failproofai update`'s Hermes half, and the linked-plugin install it moves + * profiles to. + * + * Legacy (≤1.0.5) Hermes enforcement is config.yaml shell hooks, which Hermes + * cron jobs never run — each cron fire builds its own hook scope that only + * discovered plugins join. `update` never looked at Hermes, so an upgraded + * machine kept that gap silently. These tests pin the migration, the daemon + * gate that keeps shell hooks when the plugin could not get verdicts, and the + * ownership rules for `/plugins/failproofai`. + */ +import { describe, it, expect, beforeEach, afterEach, vi } from "vitest"; +import { + cpSync, + existsSync, + lstatSync, + mkdirSync, + mkdtempSync, + readFileSync, + readlinkSync, + rmSync, + symlinkSync, + writeFileSync, +} from "node:fs"; +import { homedir, tmpdir } from "node:os"; +import { join, resolve } from "node:path"; +import { parse } from "yaml"; +import { + hermes, + hermesProfileHealth, + hermesProfileStatusRows, + installHermesPlugin, + findHermesPluginSourcePath, +} from "../../src/hooks/integrations"; +import { runHermesUpdateMigration } from "../../src/hooks/hermes-update"; +import { FAILPROOFAI_HOOK_MARKER } from "../../src/hooks/types"; + +const ORIG_CWD = process.cwd(); +const PLUGIN_FILES = ["plugin.yaml", "__init__.py", "client.py", "ledger.py"]; + +let tempDir: string; +let packageRoot: string; +let source: string; +const saved: Record = {}; + +beforeEach(() => { + tempDir = mkdtempSync(join(tmpdir(), "fp-hermes-update-")); + for (const key of ["HOME", "HERMES_HOME", "FAILPROOFAI_PACKAGE_ROOT"]) saved[key] = process.env[key]; + process.env.HOME = tempDir; + delete process.env.HERMES_HOME; + // Bun caches homedir() at startup and would ignore the override, sending + // every write below into the developer's real ~/.hermes. Refuse instead. + if (homedir() !== tempDir) { + throw new Error(`HOME override not honoured (homedir() is ${homedir()}); refusing to run`); + } + packageRoot = resolve(tempDir, "npm-global", "failproofai"); + source = resolve(packageRoot, "hermes-plugin"); + cpSync(resolve(ORIG_CWD, "hermes-plugin"), source, { recursive: true }); + process.env.FAILPROOFAI_PACKAGE_ROOT = packageRoot; +}); + +afterEach(() => { + for (const [key, value] of Object.entries(saved)) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + rmSync(tempDir, { recursive: true, force: true }); +}); + +function home(profile = "default"): string { + return profile === "default" + ? resolve(tempDir, ".hermes") + : resolve(tempDir, ".hermes", "profiles", profile); +} +function configPath(profile = "default"): string { + return resolve(home(profile), "config.yaml"); +} +function pluginPath(profile = "default"): string { + return resolve(home(profile), "plugins", "failproofai"); +} +function writeConfig(profile: string, body: string): void { + mkdirSync(home(profile), { recursive: true }); + writeFileSync(configPath(profile), body); +} +interface HermesConfig { + model?: string; + plugins?: { enabled?: string[]; disabled?: string[] }; + hooks?: Record | undefined>; +} +function readConfig(profile = "default"): HermesConfig { + return parse(readFileSync(configPath(profile), "utf8")) as HermesConfig; +} + +/** A ≤1.0.5 install: failproofai shell hooks beside an operator's own hook. */ +const LEGACY_CONFIG = [ + "model: gpt-5", + "hooks_auto_accept: true", + "hooks:", + " pre_tool_call:", + " - command: operator-check --tool", + " - command: failproofai --hook pre_tool_call --cli hermes", + " " + FAILPROOFAI_HOOK_MARKER + ": true", + " post_tool_call:", + " - command: failproofai --hook post_tool_call --cli hermes", + " " + FAILPROOFAI_HOOK_MARKER + ": true", + "", +].join("\n"); + +/** A 1.0.6–1.0.8 install: a marked COPY of the plugin, enabled in config. */ +function makeManagedCopy(profile = "default"): void { + writeConfig(profile, "model: gpt-5\nplugins:\n enabled: [failproofai]\n"); + const dest = pluginPath(profile); + mkdirSync(dest, { recursive: true }); + for (const file of PLUGIN_FILES) cpSync(resolve(ORIG_CWD, "hermes-plugin", file), resolve(dest, file)); + writeFileSync(resolve(dest, ".failproofai-managed"), "Managed by failproofai.\n"); +} + +const daemonYes = () => vi.fn(async () => true); +const daemonNo = () => vi.fn(async () => false); + +describe("linked Hermes plugin install", () => { + it("links the profile's plugin dir to the package's hermes-plugin/", () => { + writeConfig("default", "model: gpt-5\n"); + expect(installHermesPlugin(configPath())).toBe("linked"); + expect(lstatSync(pluginPath()).isSymbolicLink()).toBe(true); + expect(readlinkSync(pluginPath())).toBe(source); + // Hermes reads the manifest through the link, exactly as from a directory. + expect(readFileSync(resolve(pluginPath(), "plugin.yaml"), "utf8")).toMatch(/^name: failproofai$/m); + expect(installHermesPlugin(configPath())).toBe("unchanged"); + }); + + it("an npm upgrade is picked up with no reinstall", () => { + installHermesPlugin(configPath()); + writeFileSync(resolve(source, "client.py"), "# new release\n"); + expect(readFileSync(resolve(pluginPath(), "client.py"), "utf8")).toBe("# new release\n"); + }); + + it("relinks a dangling link it recorded, left by a removed install prefix", () => { + const oldPrefix = resolve(tempDir, "old-node", "lib", "node_modules", "failproofai", "hermes-plugin"); + mkdirSync(resolve(home(), "plugins"), { recursive: true }); + symlinkSync(oldPrefix, pluginPath()); // target never existed: dangling + writeFileSync(resolve(home(), "plugins", ".failproofai-link"), oldPrefix + "\n"); + expect(installHermesPlugin(configPath())).toBe("linked"); + expect(readlinkSync(pluginPath())).toBe(source); + expect(readFileSync(resolve(home(), "plugins", ".failproofai-link"), "utf8").trim()).toBe(source); + }); + + it("treats an unrecorded dangling link as foreign: nothing can prove it is ours", () => { + const oldPrefix = resolve(tempDir, "old-node", "lib", "node_modules", "failproofai", "hermes-plugin"); + mkdirSync(resolve(home(), "plugins"), { recursive: true }); + symlinkSync(oldPrefix, pluginPath()); + expect(() => installHermesPlugin(configPath())).toThrow(/unmanaged Hermes plugin/); + expect(readlinkSync(pluginPath())).toBe(oldPrefix); + }); + + it("refuses a same-named plugin with a failproofai manifest that is not ours (F9)", () => { + // Both public naming conditions met: a directory called hermes-plugin whose + // plugin.yaml says name: failproofai. Neither is proof of ownership. + const lookalike = resolve(tempDir, "vendor", "hermes-plugin"); + mkdirSync(lookalike, { recursive: true }); + for (const f of ["plugin.yaml", "__init__.py", "client.py", "ledger.py"]) { + writeFileSync(resolve(lookalike, f), f === "plugin.yaml" ? "name: failproofai\n" : "# not ours\n"); + } + mkdirSync(resolve(home(), "plugins"), { recursive: true }); + symlinkSync(lookalike, pluginPath()); + expect(() => installHermesPlugin(configPath())).toThrow(/unmanaged Hermes plugin/); + expect(readlinkSync(pluginPath())).toBe(lookalike); + expect(hermesProfileHealth()[0].pluginInstalled).toBe(false); + expect(hermesProfileHealth()[0].pluginForeign).toBe(true); + expect(hermesProfileStatusRows()[0][1]).toMatch(/another plugin occupies .*plugins\/failproofai/); + }); + + it("adopts a 1.0.9-beta link into an npm failproofai package by provenance, and records it", () => { + const pkg = resolve(tempDir, "prefix", "lib", "node_modules", "failproofai"); + mkdirSync(resolve(pkg, "hermes-plugin"), { recursive: true }); + writeFileSync(resolve(pkg, "package.json"), JSON.stringify({ name: "failproofai", version: "1.0.9-beta.1" })); + for (const f of ["plugin.yaml", "__init__.py", "client.py", "ledger.py"]) { + writeFileSync(resolve(pkg, "hermes-plugin", f), f === "plugin.yaml" ? "name: failproofai\n" : "#\n"); + } + mkdirSync(resolve(home(), "plugins"), { recursive: true }); + symlinkSync(resolve(pkg, "hermes-plugin"), pluginPath()); // no record: a beta install + expect(installHermesPlugin(configPath())).toBe("linked"); // not current → relinked to this package + expect(readlinkSync(pluginPath())).toBe(source); + expect(readFileSync(resolve(home(), "plugins", ".failproofai-link"), "utf8").trim()).toBe(source); + }); + + it("uninstall removes the ownership record with the link", () => { + writeConfig("default", "model: gpt-5\n"); + installHermesPlugin(configPath()); + expect(existsSync(resolve(home(), "plugins", ".failproofai-link"))).toBe(true); + hermes.removeHooksFromFile(configPath()); + expect(existsSync(pluginPath())).toBe(false); + expect(existsSync(resolve(home(), "plugins", ".failproofai-link"))).toBe(false); + }); + + it("refuses to replace a symlink to somebody else's plugin", () => { + const other = resolve(tempDir, "operator-plugins", "guard"); + mkdirSync(other, { recursive: true }); + writeFileSync(resolve(other, "plugin.yaml"), "name: guard\n"); + mkdirSync(resolve(home(), "plugins"), { recursive: true }); + symlinkSync(other, pluginPath()); + + expect(() => installHermesPlugin(configPath())).toThrow(/Refusing to overwrite an unmanaged Hermes plugin/); + expect(readlinkSync(pluginPath())).toBe(other); + }); + + it("refuses a link that is named hermes-plugin but is not FailproofAI's", () => { + const other = resolve(tempDir, "someone-else", "hermes-plugin"); + mkdirSync(other, { recursive: true }); + writeFileSync(resolve(other, "plugin.yaml"), "name: another-plugin\n"); + mkdirSync(resolve(home(), "plugins"), { recursive: true }); + symlinkSync(other, pluginPath()); + + expect(() => installHermesPlugin(configPath())).toThrow(/Refusing to overwrite/); + expect(readlinkSync(pluginPath())).toBe(other); + }); + + it("falls back to a marked copy when a symlink cannot be created", () => { + writeConfig("default", "model: gpt-5\n"); + const symlink = () => { + throw Object.assign(new Error("EPERM: operation not permitted"), { code: "EPERM" }); + }; + expect(installHermesPlugin(configPath(), { symlink })).toBe("copied"); + expect(lstatSync(pluginPath()).isDirectory()).toBe(true); + expect(existsSync(resolve(pluginPath(), ".failproofai-managed"))).toBe(true); + for (const file of PLUGIN_FILES) expect(existsSync(resolve(pluginPath(), file))).toBe(true); + + const settings = hermes.readSettings(configPath()); + hermes.writeHookEntries(settings, "/usr/bin/failproofai", "user"); + hermes.writeSettings(configPath(), settings); + expect(hermesProfileHealth()[0]).toMatchObject({ pluginMode: "copy", healthy: true }); + expect(hermesProfileStatusRows()[0][1]).toMatch(/^native plugin enabled \(copied/); + }); + + it("uninstall removes the link and never the package directory", () => { + writeConfig("default", "model: gpt-5\n"); + hermes.prepareInstall!(configPath()); + const settings = hermes.readSettings(configPath()); + hermes.writeHookEntries(settings, "/usr/bin/failproofai", "user"); + hermes.writeSettings(configPath(), settings); + + expect(hermes.removeHooksFromFile(configPath())).toBe(2); + expect(existsSync(pluginPath())).toBe(false); + expect(() => lstatSync(pluginPath())).toThrow(); + for (const file of PLUGIN_FILES) expect(existsSync(resolve(source, file))).toBe(true); + expect(readConfig().plugins).toBeUndefined(); + }); + + it("uninstall still removes an old managed copy", () => { + makeManagedCopy(); + expect(hermes.removeHooksFromFile(configPath())).toBe(2); + expect(existsSync(pluginPath())).toBe(false); + }); +}); + +describe("Hermes health flags legacy shell hooks", () => { + it("a profile only on shell hooks is unhealthy because cron is unchecked", () => { + writeConfig("default", LEGACY_CONFIG); + expect(hermesProfileHealth()[0]).toMatchObject({ + legacyShellHookPresent: true, + cronUnchecked: true, + healthy: false, + pluginMode: null, + }); + expect(hermesProfileStatusRows()).toEqual([ + [ + "hermes/default", + "UNHEALTHY — legacy shell hooks: Hermes cron jobs are not checked. Run `failproofai update`", + ], + ]); + }); +}); + +describe("failproofai update → Hermes migration", () => { + it("migrates a legacy shell-hook profile to the linked plugin, keeping operator hooks", async () => { + writeConfig("default", LEGACY_CONFIG); + const probe = daemonYes(); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: probe }); + + expect(result.ok).toBe(true); + expect(result.profiles[0]).toMatchObject({ status: "migrated", detail: "shell hooks → linked plugin" }); + expect(probe).toHaveBeenCalledTimes(1); + expect(readlinkSync(pluginPath())).toBe(source); + const config = readConfig(); + expect(config.plugins?.enabled).toEqual(["failproofai"]); + expect(config.hooks?.pre_tool_call).toEqual([{ command: "operator-check --tool" }]); + expect(config.hooks?.post_tool_call).toBeUndefined(); + expect(config.model).toBe("gpt-5"); + expect(hermesProfileHealth()[0]).toMatchObject({ healthy: true, pluginMode: "link" }); + expect(result.lines.join("\n")).toMatch(/hermes\/default\s+migrated — shell hooks → linked plugin/); + expect(result.lines.join("\n")).toMatch(/Cron jobs load the plugin on their next run/); + }); + + it("migrates a copied plugin (1.0.6–1.0.8) to the link", async () => { + makeManagedCopy(); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.ok).toBe(true); + expect(result.profiles[0]).toMatchObject({ status: "migrated", detail: "copied plugin → linked plugin" }); + expect(lstatSync(pluginPath()).isSymbolicLink()).toBe(true); + expect(readConfig().plugins?.enabled).toEqual(["failproofai"]); + }); + + it("leaves profiles without any failproofai integration alone and asks the daemon nothing", async () => { + writeConfig("default", "model: gpt-5\n"); + writeConfig("work", "model: other\nplugins:\n enabled: [operator-plugin]\n"); + const before = readFileSync(configPath("work"), "utf8"); + const probe = daemonYes(); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: probe }); + + expect(result).toMatchObject({ ok: true, lines: [] }); + expect(result.profiles.map((p) => p.status)).toEqual(["untouched", "untouched"]); + expect(probe).not.toHaveBeenCalled(); + expect(readFileSync(configPath("work"), "utf8")).toBe(before); + expect(existsSync(pluginPath("work"))).toBe(false); + expect(existsSync(pluginPath())).toBe(false); + }); + + it("migrates only the profiles that use failproofai", async () => { + writeConfig("default", LEGACY_CONFIG); + writeConfig("work", "model: other\n"); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.profiles.map((p) => [p.name, p.status])).toEqual([ + ["default", "migrated"], + ["work", "untouched"], + ]); + expect(existsSync(pluginPath("work"))).toBe(false); + expect(result.lines.join("\n")).toMatch(/hermes\/work\s+skipped — no failproofai integration/); + }); + + it("records ownership of an already-current link that has no record yet (a 1.0.9-beta install)", async () => { + writeConfig("default", "model: gpt-5\n"); + hermes.prepareInstall!(configPath()); + const settings = hermes.readSettings(configPath()); + hermes.writeHookEntries(settings, "/usr/bin/failproofai", "user"); + hermes.writeSettings(configPath(), settings); + rmSync(resolve(home(), "plugins", ".failproofai-link"), { force: true }); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.profiles[0].status).toBe("current"); + expect(readFileSync(resolve(home(), "plugins", ".failproofai-link"), "utf8").trim()).toBe(source); + }); + + it("reports an already-linked profile as current without probing the daemon", async () => { + writeConfig("default", "model: gpt-5\n"); + hermes.prepareInstall!(configPath()); + const settings = hermes.readSettings(configPath()); + hermes.writeHookEntries(settings, "/usr/bin/failproofai", "user"); + hermes.writeSettings(configPath(), settings); + + const probe = daemonYes(); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: probe }); + expect(result.ok).toBe(true); + expect(result.profiles[0].status).toBe("current"); + expect(probe).not.toHaveBeenCalled(); + expect(result.lines.join("\n")).toMatch(/hermes\/default\s+already current/); + }); + + it("a linked plugin also listed in plugins.disabled is unhealthy, and update re-enables it", async () => { + // Hermes checks plugins.disabled before plugins.enabled, so this plugin + // never loads: it must not read as healthy, nor be skipped as current. + writeConfig("default", "model: gpt-5\n"); + hermes.prepareInstall!(configPath()); + const settings = hermes.readSettings(configPath()); + hermes.writeHookEntries(settings, "/usr/bin/failproofai", "user"); + hermes.writeSettings(configPath(), settings); + writeFileSync(configPath(), readFileSync(configPath(), "utf8") + " disabled:\n - failproofai\n"); + expect(readConfig().plugins?.enabled).toContain("failproofai"); + expect(readConfig().plugins?.disabled).toContain("failproofai"); + + const [before] = hermesProfileHealth(); + expect(before.healthy).toBe(false); + expect(before.pluginEnabled).toBe(false); + expect(before.pluginDisabled).toBe(true); + expect(hermesProfileStatusRows()[0][1]).toMatch(/plugin disabled \(listed in plugins\.disabled\)/); + expect(hermes.hooksInstalledInSettings("user")).toBe(false); + + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.ok).toBe(true); + expect(result.profiles[0].status).not.toBe("current"); + expect(readConfig().plugins?.disabled ?? []).not.toContain("failproofai"); + expect(readConfig().plugins?.enabled).toContain("failproofai"); + expect(hermesProfileHealth()[0].healthy).toBe(true); + }); + + it("keeps the shell hooks and fails when the daemon cannot do policyEvaluation", async () => { + writeConfig("default", LEGACY_CONFIG); + makeManagedCopy("work"); + const legacyBefore = readFileSync(configPath(), "utf8"); + const probe = daemonNo(); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: probe }); + + expect(result.ok).toBe(false); // → `failproofai update` exits 1 + expect(probe).toHaveBeenCalledTimes(1); + expect(result.profiles.map((p) => [p.name, p.status, p.legacyShellHooksRemain])).toEqual([ + ["default", "blocked", true], + ["work", "blocked", false], + ]); + // Nothing changed anywhere: the hooks are the only enforcement left. + expect(readFileSync(configPath(), "utf8")).toBe(legacyBefore); + expect(existsSync(pluginPath())).toBe(false); + expect(lstatSync(pluginPath("work")).isDirectory()).toBe(true); + const text = result.lines.join("\n"); + expect(text).toMatch(/hermes\/default\s+NOT migrated — daemon lacks native policy evaluation; shell hooks left in place/); + expect(text).toMatch(/Hermes is NOT migrated/); + expect(text).toMatch(/failproofai config/); + }); + + it("fails without touching hooks when an unmanaged plugin holds the name", async () => { + writeConfig("default", LEGACY_CONFIG); + mkdirSync(pluginPath(), { recursive: true }); + writeFileSync(resolve(pluginPath(), "plugin.yaml"), "name: operator-owned\n"); + const before = readFileSync(configPath(), "utf8"); + + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.ok).toBe(false); + expect(result.profiles[0]).toMatchObject({ status: "failed", legacyShellHooksRemain: true }); + expect(result.profiles[0].detail).toMatch(/unmanaged plugin occupies/); + expect(readFileSync(configPath(), "utf8")).toBe(before); + expect(readFileSync(resolve(pluginPath(), "plugin.yaml"), "utf8")).toBe("name: operator-owned\n"); + }); + + it("never rewrites a config.yaml that does not parse", async () => { + const broken = LEGACY_CONFIG + " bad: [unclosed\n"; + writeConfig("default", broken); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.ok).toBe(false); + expect(result.profiles[0].status).toBe("failed"); + expect(readFileSync(configPath(), "utf8")).toBe(broken); + expect(existsSync(pluginPath())).toBe(false); + }); + + it("leaves a broken config.yaml it was never installed in alone, without failing update", async () => { + const broken = "model:\n default: gpt-6-luna\n bad: [unclosed\n"; + writeConfig("default", broken); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + expect(result.ok).toBe(true); + expect(result.profiles[0].status).toBe("untouched"); + expect(readFileSync(configPath(), "utf8")).toBe(broken); + expect(existsSync(pluginPath())).toBe(false); + }); + + it("one profile that cannot be inspected is reported failed and the others still migrate (F11)", async () => { + writeConfig("default", LEGACY_CONFIG); + writeConfig("work", LEGACY_CONFIG); + // Make the default profile's plugins/ path a FILE, so inspecting it throws + // mid-loop; the work profile must still be migrated and reported. + writeFileSync(resolve(home("default"), "plugins"), "not a directory"); + const result = await runHermesUpdateMigration({ daemonSupportsPolicyEvaluation: daemonYes() }); + const byName = Object.fromEntries(result.profiles.map((p) => [p.name, p.status])); + expect(byName.work).toBe("migrated"); + expect(byName.default).toBe("failed"); + }); + + it("uses the copy fallback during migration when links are impossible", async () => { + writeConfig("default", LEGACY_CONFIG); + const symlink = () => { + throw new Error("EPERM"); + }; + const result = await runHermesUpdateMigration({ + daemonSupportsPolicyEvaluation: daemonYes(), + installDeps: { symlink }, + }); + expect(result.ok).toBe(true); + expect(result.profiles[0].detail).toBe("shell hooks → plugin copy (symlink not possible here)"); + expect(existsSync(resolve(pluginPath(), ".failproofai-managed"))).toBe(true); + expect(readConfig().hooks?.pre_tool_call).toEqual([{ command: "operator-check --tool" }]); + }); +}); + +describe("findHermesPluginSourcePath — the package root from any layout", () => { + let pkg: string; + beforeEach(() => { + pkg = mkdtempSync(join(tmpdir(), "fpai-pkgroot-")); + mkdirSync(join(pkg, "hermes-plugin"), { recursive: true }); + writeFileSync(join(pkg, "hermes-plugin", "plugin.yaml"), "name: failproofai\n"); + mkdirSync(join(pkg, "src", "hooks"), { recursive: true }); + mkdirSync(join(pkg, "dist"), { recursive: true }); + mkdirSync(join(pkg, ".next", "standalone", "server", "chunks"), { recursive: true }); + }); + afterEach(() => rmSync(pkg, { recursive: true, force: true })); + + it("finds it from the source tree (src/hooks)", () => { + expect(findHermesPluginSourcePath(join(pkg, "src", "hooks"))).toBe(resolve(pkg, "hermes-plugin")); + }); + + it("finds it from the bundled CLI (dist/), where three parents would overshoot", () => { + expect(findHermesPluginSourcePath(join(pkg, "dist"))).toBe(resolve(pkg, "hermes-plugin")); + }); + + it("finds it from the dashboard's standalone build", () => { + expect(findHermesPluginSourcePath(join(pkg, ".next", "standalone", "server", "chunks"))).toBe( + resolve(pkg, "hermes-plugin"), + ); + }); + + it("with no plugin anywhere above, answers a path that does not exist so install fails loudly", () => { + rmSync(join(pkg, "hermes-plugin"), { recursive: true, force: true }); + expect(existsSync(findHermesPluginSourcePath(join(pkg, "src", "hooks")))).toBe(false); + }); +}); diff --git a/__tests__/hooks/hook-activity-jev.test.ts b/__tests__/hooks/hook-activity-jev.test.ts index aae8f690a..0ef116703 100644 --- a/__tests__/hooks/hook-activity-jev.test.ts +++ b/__tests__/hooks/hook-activity-jev.test.ts @@ -123,10 +123,20 @@ describe("hook-activity-store: Jev fields", () => { expect(row).not.toHaveProperty("jevDecision"); expect(row).not.toHaveProperty("jevMode"); expect(row).not.toHaveProperty("jevLatencyMs"); - expect(row.jevCleared).toEqual(["block-env-files"]); + // An unknown mode cannot say whether its clears were enforced, so they go + // with it rather than being filed as enforce-mode clears. + expect(row).not.toHaveProperty("jevCleared"); expect(row.durationMs).toBe(44); }); + it("keeps only name-shaped clears when the mode is valid", () => { + persistHookActivity( + makeEntry({ evaluator: "jev", jevMode: "enforce", jevCleared: ["block-env-files", "not a policy name"] }), + ); + const [row] = rawLines(); + expect(row.jevCleared).toEqual(["block-env-files"]); + }); + it("keeps the fields across a page rotation", () => { for (let i = 0; i < PAGE_SIZE + 3; i++) { persistHookActivity(makeEntry({ timestamp: 1_000 + i, ...JEV_ANSWERED, jevLatencyMs: i })); diff --git a/__tests__/hooks/hook-telemetry-jev.test.ts b/__tests__/hooks/hook-telemetry-jev.test.ts index fe56f7efe..ecbb7902a 100644 --- a/__tests__/hooks/hook-telemetry-jev.test.ts +++ b/__tests__/hooks/hook-telemetry-jev.test.ts @@ -39,11 +39,11 @@ describe("jevTelemetryProperties", () => { evaluator: "jev-fallback", jevFallbackReason: "error: connect ECONNREFUSED", jevLatencyMs: 3, - jevMode: "shadow", + jevMode: "observe", }), ).toEqual({ jev_evaluator: "jev-fallback", - jev_mode: "shadow", + jev_mode: "observe", jev_fallback_reason: "error", jev_latency_ms: 3, }); diff --git a/__tests__/hooks/integrations.test.ts b/__tests__/hooks/integrations.test.ts index 32f013d47..3155f96c3 100644 --- a/__tests__/hooks/integrations.test.ts +++ b/__tests__/hooks/integrations.test.ts @@ -18,6 +18,9 @@ import { writeFileSync, mkdirSync, readdirSync, + cpSync, + lstatSync, + readlinkSync, } from "node:fs"; import { tmpdir } from "node:os"; import { resolve, join } from "node:path"; @@ -571,6 +574,7 @@ describe("Hermes integration", () => { let origHome: string | undefined; let origHermesHome: string | undefined; let origPackageRoot: string | undefined; + let packageRoot: string; beforeEach(() => { origHome = process.env.HOME; process.env.HOME = tempDir; @@ -578,8 +582,21 @@ describe("Hermes integration", () => { // with a profile-scoped shell doesn't get their real config.yaml touched. origHermesHome = process.env.HERMES_HOME; delete process.env.HERMES_HOME; + // Every test below writes under `homedir()`. A runtime that caches the home + // directory at startup (Bun does) would ignore the HOME override and point + // them at the developer's REAL ~/.hermes — refuse rather than touch it. + if (homedir() !== tempDir) { + throw new Error( + `HOME override not honoured (homedir() is ${homedir()}); refusing to run Hermes tests against a real home`, + ); + } + // Install now LINKS profiles to the package's hermes-plugin/, so a test + // that edits files through the installed plugin would edit the package. + // Give every test its own throwaway package copy. origPackageRoot = process.env.FAILPROOFAI_PACKAGE_ROOT; - process.env.FAILPROOFAI_PACKAGE_ROOT = ORIG_CWD; + packageRoot = resolve(tempDir, "failproofai-package"); + cpSync(resolve(ORIG_CWD, "hermes-plugin"), resolve(packageRoot, "hermes-plugin"), { recursive: true }); + process.env.FAILPROOFAI_PACKAGE_ROOT = packageRoot; }); afterEach(() => { if (origHome === undefined) delete process.env.HOME; @@ -610,6 +627,15 @@ describe("Hermes integration", () => { return resolve(dirname(settingsPath), "plugins", "failproofai"); } + /** What a 1.0.6–1.0.8 install left behind: a marked COPY of the plugin. */ + function makeManagedCopy(destination: string): void { + mkdirSync(destination, { recursive: true }); + for (const file of ["plugin.yaml", "__init__.py", "client.py", "ledger.py"]) { + cpSync(resolve(ORIG_CWD, "hermes-plugin", file), resolve(destination, file)); + } + writeFileSync(resolve(destination, ".failproofai-managed"), "Managed by failproofai.\n"); + } + function installAt(settingsPath: string): void { hermes.prepareInstall!(settingsPath); const settings = hermes.readSettings(settingsPath); @@ -641,7 +667,7 @@ describe("Hermes integration", () => { it("buildHookEntry identifies the shipped native plugin", () => { const entry = hermes.buildHookEntry("/usr/bin/failproofai", "pre_tool_call", "user") as Record; expect(entry[FAILPROOFAI_HOOK_MARKER]).toBe(true); - expect(entry._hermesPluginPath).toBe(resolve(ORIG_CWD, "hermes-plugin")); + expect(entry._hermesPluginPath).toBe(resolve(packageRoot, "hermes-plugin")); }); it("installs the native plugin and enables it in config", () => { @@ -654,10 +680,14 @@ describe("Hermes integration", () => { expect(parsed.plugins?.enabled).toEqual(["failproofai"]); expect(parsed.hooks).toBeUndefined(); + // A symlink into the package, so npm upgrades apply with no reinstall. const installedPlugin = pluginPath(path); - for (const file of ["plugin.yaml", "__init__.py", "client.py", "ledger.py", ".failproofai-managed"]) { + expect(lstatSync(installedPlugin).isSymbolicLink()).toBe(true); + expect(readlinkSync(installedPlugin)).toBe(resolve(packageRoot, "hermes-plugin")); + for (const file of ["plugin.yaml", "__init__.py", "client.py", "ledger.py"]) { expect(existsSync(resolve(installedPlugin, file))).toBe(true); } + expect(hermesProfileHealth()[0]).toMatchObject({ pluginMode: "link", healthy: true, cronUnchecked: false }); expect(hermesProfileStatusRows()).toEqual([ ["hermes/default", "native plugin enabled"], ]); @@ -816,15 +846,16 @@ describe("Hermes integration", () => { ]); }); - it("atomically refreshes a managed plugin directory", () => { + it("replaces a managed plugin copy (1.0.6–1.0.8 install) with a link, leaving no temp entries", () => { const path = hermes.getSettingsPath("user"); - hermes.prepareInstall!(path); const destination = pluginPath(path); + makeManagedCopy(destination); writeFileSync(resolve(destination, "client.py"), "stale\n"); writeFileSync(resolve(destination, "obsolete.py"), "remove me\n"); hermes.prepareInstall!(path); + expect(lstatSync(destination).isSymbolicLink()).toBe(true); expect(readFileSync(resolve(destination, "client.py"), "utf8")).not.toBe("stale\n"); expect(existsSync(resolve(destination, "obsolete.py"))).toBe(false); expect( diff --git a/__tests__/hooks/jev-activity.test.ts b/__tests__/hooks/jev-activity.test.ts index 0cc4f0d51..018fb0a63 100644 --- a/__tests__/hooks/jev-activity.test.ts +++ b/__tests__/hooks/jev-activity.test.ts @@ -109,6 +109,20 @@ describe("sanitizeJevActivity", () => { expect(out.decision).toBe("allow"); }); + it("drops the clears of a row whose mode it does not know, rather than filing them as enforced", () => { + // A row written before the mode was named `observe` says "shadow": its + // clears never took effect, and without a mode they would read as enforce. + const out = sanitizeJevActivity( + entry({ evaluator: "jev", jevMode: "shadow" as never, jevCleared: ["protect-env-vars"] }), + ); + expect(out).not.toHaveProperty("jevMode"); + expect(out).not.toHaveProperty("jevCleared"); + expect(out.evaluator).toBe("jev"); + // A known mode keeps its clears. + const kept = sanitizeJevActivity(entry({ evaluator: "jev", jevMode: "observe", jevCleared: ["protect-env-vars"] })); + expect(kept.jevCleared).toEqual(["protect-env-vars"]); + }); + it("rounds latency and rejects negatives", () => { expect(sanitizeJevActivity(entry({ evaluator: "jev", jevLatencyMs: 37.6 })).jevLatencyMs).toBe(38); expect(sanitizeJevActivity(entry({ evaluator: "jev", jevLatencyMs: -1 }))).not.toHaveProperty("jevLatencyMs"); @@ -185,14 +199,14 @@ describe("describeJevActivity", () => { ).toEqual(["Jev verdict: allow", "cleared block-read-outside-cwd", "38 ms", "jev-1.13.0"]); }); - it("says shadow mode enforced the regex result", () => { + it("says observe mode enforced the regex result", () => { const facts = describeJevActivity( - entry({ evaluator: "jev", jevMode: "shadow", jevDecision: "allow", jevCleared: ["block-env-files"] }), + entry({ evaluator: "jev", jevMode: "observe", jevDecision: "allow", jevCleared: ["block-env-files"] }), ); expect(facts).toEqual([ "Jev verdict: allow", "would have cleared block-env-files", - "shadow mode: the regex result was enforced", + "observe mode: the regex result was enforced", ]); }); diff --git a/__tests__/hooks/jev-cli-bin.test.ts b/__tests__/hooks/jev-cli-bin.test.ts index 07ab8caf0..305028975 100644 --- a/__tests__/hooks/jev-cli-bin.test.ts +++ b/__tests__/hooks/jev-cli-bin.test.ts @@ -56,24 +56,24 @@ describe("failproofai jev (real binary)", () => { }); it("setup → status → remove, with the key piped on stdin and never printed", () => { - const setup = cli(["jev", "setup", "--provider", "typesafe", "--mode", "shadow", "--key-stdin"], `${KEY}\n`); + const setup = cli(["jev", "setup", "--provider", "typesafe", "--mode", "observe", "--key-stdin"], `${KEY}\n`); expect(setup.exitCode).toBe(0); expect(setup.stdout + setup.stderr).not.toContain(KEY); expect(setup.stdout).toContain("jev.json"); expect(existsSync(CONFIG)).toBe(true); if (process.platform !== "win32") expect(statSync(CONFIG).mode & 0o777).toBe(0o600); - expect(JSON.parse(readFileSync(CONFIG, "utf8"))).toEqual({ provider: "typesafe", apiKey: KEY, mode: "shadow" }); + expect(JSON.parse(readFileSync(CONFIG, "utf8"))).toEqual({ provider: "typesafe", apiKey: KEY, mode: "observe" }); const status = cli(["jev", "status"]); expect(status.exitCode).toBe(0); expect(status.stdout + status.stderr).not.toContain(KEY); expect(status.stdout).toContain("typesafe"); - expect(status.stdout).toContain("shadow"); + expect(status.stdout).toContain("observe"); const json = cli(["jev", "status", "--json"]); expect(json.exitCode).toBe(0); expect(json.stdout).not.toContain(KEY); - expect(JSON.parse(json.stdout)).toMatchObject({ status: "ok", provider: "typesafe", mode: "shadow", keySource: "file" }); + expect(JSON.parse(json.stdout)).toMatchObject({ status: "ok", provider: "typesafe", mode: "observe", keySource: "file" }); const removed = cli(["jev", "remove"]); expect(removed.exitCode).toBe(0); diff --git a/__tests__/hooks/jev-cli-cloud.test.ts b/__tests__/hooks/jev-cli-cloud.test.ts index 08f99dc6a..15dc18d61 100644 --- a/__tests__/hooks/jev-cli-cloud.test.ts +++ b/__tests__/hooks/jev-cli-cloud.test.ts @@ -73,11 +73,11 @@ describe("jev CLI: FailproofAI Cloud", () => { describe("status", () => { it("on: FailproofAI Cloud, host only, key from the connection", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); const human = await runJevCommand(["status"], RENDER); expect(human.exitCode).toBe(0); const t = text(human); - expect(t).toContain("on · shadow"); + expect(t).toContain("on · observe"); expect(t).toContain("FailproofAI Cloud"); expect(t).toContain("app.befailproof.ai"); expect(t).not.toContain("/enforcement/v1/jev"); @@ -91,7 +91,7 @@ describe("jev CLI: FailproofAI Cloud", () => { providerLabel: "FailproofAI Cloud", endpoint: "app.befailproof.ai", model: "jev-1.13.0", - mode: "shadow", + mode: "observe", keySource: "cloud", keySourceLabel: "FailproofAI Cloud connection", cloudConnected: true, @@ -120,7 +120,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("connected with a key that has no Jev: says so, never \"not connected\"", async () => { writeCredentials({ ingest: { url: `${ORIGIN}/v1/events`, key: KEY } }); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); const human = await runJevCommand(["status"], RENDER); expect(human.exitCode).toBe(0); expect(text(human)).toContain("off — no Jev key is stored for this machine's FailproofAI Cloud connection"); @@ -131,7 +131,7 @@ describe("jev CLI: FailproofAI Cloud", () => { status: "key-lacks-jev", provider: "failproofai", endpoint: "app.befailproof.ai", - mode: "shadow", + mode: "observe", keySource: "cloud", cloudConnected: true, keyCarriesJev: false, @@ -162,7 +162,7 @@ describe("jev CLI: FailproofAI Cloud", () => { ["not-connected", () => undefined, "failproofai jev remove"], ])("%s: the reconnect and the off switch are separate steps", async (_state, arrange, offCmd) => { arrange(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); const lines = (await runJevCommand(["status"], RENDER)).lines; const i = lines.findIndex((l) => l.trim() === "failproofai config --token "); expect(i).toBeGreaterThan(0); @@ -194,7 +194,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("a loose credentials.json: refused, with the chmod that fixes it", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); chmodSync(join(fpHome, "credentials.json"), 0o644); const human = await runJevCommand(["status"], RENDER); expect(human.exitCode).toBe(1); @@ -206,7 +206,7 @@ describe("jev CLI: FailproofAI Cloud", () => { // budget), not tampering: the note must name the bits that are set. it.each([["0644", 0o644], ["0640", 0o640], ["0604", 0o604]])("credentials.json at %s: says others can read it, not change it", async (_octal, mode) => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); chmodSync(join(fpHome, "credentials.json"), mode); const t = text(await runJevCommand(["status"], RENDER)); expect(t).toMatch(/can read it/); @@ -216,9 +216,9 @@ describe("jev CLI: FailproofAI Cloud", () => { // A Cloud jev.json has no key (it is in credentials.json), and neither has a // --key-from-env one: refused as firmly, but not for disclosing a key. it.each([ - ["Cloud", { provider: "failproofai", baseUrl: BASE, mode: "shadow" }, false], - ["BYOK key-from-env", { provider: "typesafe", mode: "shadow" }, false], - ["BYOK with a stored key", { provider: "typesafe", apiKey: BYOK_KEY, mode: "shadow" }, true], + ["Cloud", { provider: "failproofai", baseUrl: BASE, mode: "observe" }, false], + ["BYOK key-from-env", { provider: "typesafe", mode: "observe" }, false], + ["BYOK with a stored key", { provider: "typesafe", apiKey: BYOK_KEY, mode: "observe" }, true], ])("%s jev.json at 0644: says it holds a key only when it does", async (_kind, file, holdsKey) => { connect(); writeJev(file); @@ -234,14 +234,14 @@ describe("jev CLI: FailproofAI Cloud", () => { it("credentials.json group-writable: says others could change it", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); chmodSync(join(fpHome, "credentials.json"), 0o620); expect(text(await runJevCommand(["status"], RENDER))).toContain("could change it"); }); it("credentials.json that is not JSON: no permissions claim, reconnect instead", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); writeFileSync(join(fpHome, "credentials.json"), "{not json", { mode: 0o600 }); const human = await runJevCommand(["status"], RENDER); expect(human.exitCode).toBe(1); @@ -293,7 +293,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("refused credentials.json --json: the Cloud facts, and which file's permissions are which", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); chmodSync(join(fpHome, "credentials.json"), 0o640); const machine = await runJevCommand(["status", "--json"], RENDER); expect(machine.exitCode).toBe(1); @@ -320,14 +320,14 @@ describe("jev CLI: FailproofAI Cloud", () => { }); describe("setup --provider failproofai", () => { - it("builds the file from the connection: shadow, no key, 0600", async () => { + it("builds the file from the connection: observe, no key, 0600", async () => { connect(); const r = await runJevCommand(["setup", "--provider", "failproofai"], RENDER); expect(r.exitCode, text(r)).toBe(0); - expect(onDisk()).toEqual({ provider: "failproofai", mode: "shadow", baseUrl: BASE }); + expect(onDisk()).toEqual({ provider: "failproofai", mode: "observe", baseUrl: BASE }); if (process.platform !== "win32") expect(statSync(jevConfigPath()).mode & 0o777).toBe(0o600); expect(loadJevConfig()).toMatchObject({ provider: "failproofai", apiKey: KEY }); - expect(text(r)).toContain("saved · FailproofAI Cloud · shadow"); + expect(text(r)).toContain("saved · FailproofAI Cloud · observe"); noKey(r); }); @@ -373,7 +373,7 @@ describe("jev CLI: FailproofAI Cloud", () => { }); it("a mode switch over a Cloud file rewrites the mode and keeps the rest — connected or not", async () => { - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow", timeoutMs: 2500 }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe", timeoutMs: 2500 }); const r = await runJevCommand(["setup", "--mode", "enforce"], RENDER); expect(r.exitCode, text(r)).toBe(0); expect(onDisk()).toEqual({ provider: "failproofai", baseUrl: BASE, mode: "enforce", timeoutMs: 2500 }); @@ -386,7 +386,7 @@ describe("jev CLI: FailproofAI Cloud", () => { // Connected, with a key that has no Jev: said so — never "not connected". writeCredentials({ ingest: { url: `${ORIGIN}/v1/events`, key: KEY } }); - const lacks = await runJevCommand(["setup", "--mode", "shadow"], RENDER); + const lacks = await runJevCommand(["setup", "--mode", "observe"], RENDER); expect(lacks.exitCode, text(lacks)).toBe(0); expect(text(lacks)).toContain("connected, no Jev key stored for it"); expect(text(lacks)).not.toContain("not connected"); @@ -404,7 +404,7 @@ describe("jev CLI: FailproofAI Cloud", () => { // Connected with a Jev key: the live check is the next step. connect(); - const on = await runJevCommand(["setup", "--mode", "shadow"], RENDER); + const on = await runJevCommand(["setup", "--mode", "observe"], RENDER); expect(on.exitCode, text(on)).toBe(0); expect(text(on)).toContain("failproofai jev test"); expect(text(on)).not.toContain("config --token "); @@ -412,7 +412,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("drops a key someone put in a Cloud file, which is what makes it valid again", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow", apiKey: BYOK_KEY }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe", apiKey: BYOK_KEY }); expect(loadJevConfig()).toBeNull(); const r = await runJevCommand(["setup", "--provider", "failproofai"], RENDER); expect(r.exitCode).toBe(0); @@ -430,12 +430,12 @@ describe("jev CLI: FailproofAI Cloud", () => { expect(loadJevConfig()?.baseUrl).toBe(BASE); }); - it("switching from BYOK is explicit, starts in shadow, and carries no BYOK key over", async () => { + it("switching from BYOK is explicit, starts in observe, and carries no BYOK key over", async () => { connect(); writeJev({ provider: "typesafe", apiKey: BYOK_KEY, mode: "enforce" }); const r = await runJevCommand(["setup", "--provider", "failproofai"], RENDER); expect(r.exitCode).toBe(0); - expect(onDisk()).toEqual({ provider: "failproofai", mode: "shadow", baseUrl: BASE }); + expect(onDisk()).toEqual({ provider: "failproofai", mode: "observe", baseUrl: BASE }); expect(readFileSync(jevConfigPath(), "utf8")).not.toContain(BYOK_KEY); }); }); @@ -450,7 +450,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("jev models has nothing to read for it", async () => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); const r = await runJevCommand(["models"], RENDER); expect(r.exitCode).toBe(1); expect(text(r)).toContain("FailproofAI Cloud serves no model list"); @@ -469,12 +469,12 @@ describe("jev CLI: FailproofAI Cloud", () => { ["a group-writable directory", () => chmodSync(fpHome, 0o770), () => `chmod 700 ${fpHome}`], [ "a Cloud file on another origin", - () => writeJev({ provider: "failproofai", baseUrl: "https://staging.befailproof.ai/enforcement/v1/jev", mode: "shadow" }), + () => writeJev({ provider: "failproofai", baseUrl: "https://staging.befailproof.ai/enforcement/v1/jev", mode: "observe" }), () => "failproofai jev setup --provider failproofai", ], ])("refused (%s): the same fix as status", async (_label, breakIt, fix) => { connect(); - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); breakIt(); const t = text(await runJevCommand(["test"], RENDER)); chmodSync(fpHome, 0o700); @@ -483,7 +483,7 @@ describe("jev CLI: FailproofAI Cloud", () => { }); it("not connected / switched off: not run, with a code for each", async () => { - writeJev({ provider: "failproofai", baseUrl: BASE, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: BASE, mode: "observe" }); const nc = await runJevCommand(["test", "--json"], RENDER); expect(nc.exitCode).toBe(1); expect(json(nc)).toMatchObject({ ok: false, error: { code: "not-connected" } }); @@ -514,7 +514,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("says what to do in FailproofAI Cloud's terms", async () => { const origin = `http://127.0.0.1:${port}`; connect(origin); - writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "observe" }); status = 403; body = { error: "forbidden", message: "this key does not carry jev:evaluate" }; const refused = await runJevCommand(["test"], { ...RENDER, testTimeoutMs: 5_000 }); @@ -534,7 +534,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("a 429 names the daily limit when the body says so, and the per-minute one otherwise", async () => { const origin = `http://127.0.0.1:${port}`; connect(origin); - writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "observe" }); status = 429; body = { error: "daily_limit_reached" }; @@ -555,7 +555,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("a 503 names who fixes it, not a wait", async () => { const origin = `http://127.0.0.1:${port}`; connect(origin); - writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "observe" }); status = 503; body = { error: "jev_unavailable" }; const r = await runJevCommand(["test"], { ...RENDER, testTimeoutMs: 5_000 }); @@ -568,7 +568,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("a 422 request_rejected is that call's own, never an outage to wait out", async () => { const origin = `http://127.0.0.1:${port}`; connect(origin); - writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "observe" }); status = 422; body = { error: "request_rejected" }; const rejected = await runJevCommand(["test"], { ...RENDER, testTimeoutMs: 5_000 }); @@ -583,7 +583,7 @@ describe("jev CLI: FailproofAI Cloud", () => { it("a redirect is advice about the connection, never a --base-url this route refuses", async () => { const origin = `http://127.0.0.1:${port}`; connect(origin); - writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "shadow" }); + writeJev({ provider: "failproofai", baseUrl: `${origin}/enforcement/v1/jev`, mode: "observe" }); status = 302; body = {}; const redirected = await runJevCommand(["test"], { ...RENDER, testTimeoutMs: 5_000 }); diff --git a/__tests__/hooks/jev-cli-hardening.test.ts b/__tests__/hooks/jev-cli-hardening.test.ts index c29044523..e33d64d5b 100644 --- a/__tests__/hooks/jev-cli-hardening.test.ts +++ b/__tests__/hooks/jev-cli-hardening.test.ts @@ -207,24 +207,24 @@ describe("failproofai jev — review hardening", () => { }); }); - describe("setup refuses plain-http loopback outside shadow mode", () => { - it("enforce (the default) is refused with the reason; shadow is saved", async () => { + describe("setup refuses plain-http loopback outside observe mode", () => { + it("enforce (the default) is refused with the reason; observe is saved", async () => { const enforce = await runJevCommand(["setup", "--provider", "custom", "--base-url", "http://localhost:8787/v1", "--key-stdin"], withKey(KEY)); expect(enforce.exitCode).toBe(1); - expect(text(enforce)).toContain("shadow"); + expect(text(enforce)).toContain("observe"); expect(existsSync(jevConfigPath())).toBe(false); - const shadow = await runJevCommand( - ["setup", "--provider", "custom", "--base-url", "http://localhost:8787/v1", "--mode", "shadow", "--key-stdin"], + const observe = await runJevCommand( + ["setup", "--provider", "custom", "--base-url", "http://localhost:8787/v1", "--mode", "observe", "--key-stdin"], withKey(KEY), ); - expect(shadow.exitCode).toBe(0); - expect(readFile()).toMatchObject({ baseUrl: "http://localhost:8787/v1", mode: "shadow" }); + expect(observe.exitCode).toBe(0); + expect(readFile()).toMatchObject({ baseUrl: "http://localhost:8787/v1", mode: "observe" }); // And the mode cannot then be switched to enforce underneath it. const flip = await runJevCommand(["setup", "--mode", "enforce"], noTty); expect(flip.exitCode).toBe(1); - expect(readFile().mode).toBe("shadow"); + expect(readFile().mode).toBe("observe"); }); }); diff --git a/__tests__/hooks/jev-cli-review.test.ts b/__tests__/hooks/jev-cli-review.test.ts index a9bd2b02c..798d09e5d 100644 --- a/__tests__/hooks/jev-cli-review.test.ts +++ b/__tests__/hooks/jev-cli-review.test.ts @@ -72,7 +72,7 @@ describe("failproofai jev — review round", () => { expect(loadJevConfig()).toBeNull(); const before = readFileSync(jevConfigPath(), "utf8"); - const r = await runJevCommand(["setup", "--mode", "shadow"], noTty); + const r = await runJevCommand(["setup", "--mode", "observe"], noTty); expect(r.exitCode).toBe(1); expect(text(r)).toContain("open to other users (0664)"); expect(text(r)).toContain("https://attacker.example.com"); @@ -150,9 +150,9 @@ describe("failproofai jev — review round", () => { it("an owner-only file with a foreign endpoint keeps its key on a re-run, as before", async () => { await runJevCommand(["setup", "--provider", "typesafe", "--base-url", ATTACKER, "--key-stdin"], withKey(KEY)); - const r = await runJevCommand(["setup", "--mode", "shadow"], noTty); + const r = await runJevCommand(["setup", "--mode", "observe"], noTty); expect(r.exitCode).toBe(0); - expect(readFile()).toEqual({ provider: "typesafe", apiKey: KEY, baseUrl: ATTACKER, mode: "shadow" }); + expect(readFile()).toEqual({ provider: "typesafe", apiKey: KEY, baseUrl: ATTACKER, mode: "observe" }); }); it("status shows the endpoint a too-open file names, next to the chmod hint (human and --json)", async () => { diff --git a/__tests__/hooks/jev-cli-status-reviewable.test.ts b/__tests__/hooks/jev-cli-status-reviewable.test.ts index 9d63d86a0..b466ad03e 100644 --- a/__tests__/hooks/jev-cli-status-reviewable.test.ts +++ b/__tests__/hooks/jev-cli-status-reviewable.test.ts @@ -20,7 +20,8 @@ import { tmpdir } from "node:os"; import { join } from "node:path"; import { runJevCommand, type JevCliDeps, type JevCliResult } from "../../src/hooks/jev-cli"; import { POLICY_CATALOG } from "../../src/hooks/policy-catalog"; -import { RETAKE_PACK_COMMAND } from "../../src/hooks/policy-reviewability"; +import { JEV_POLICIES_ADD_COMMAND, NO_JEV_CHECKS_PROBLEM, RETAKE_PACK_COMMAND } from "../../src/hooks/policy-reviewability"; +import { JEV_PACK_SEMANTIC_ENTRIES } from "../fixtures/jev-policies"; import { JEV_API_KEY_ENV } from "../../src/hooks/semantic/jev-config"; const KEY = ["cli", "reviewable", "0123456789abcdef"].join("-"); @@ -73,8 +74,12 @@ function writeConfig(config: Record): void { writeFileSync(join(process.env.FAILPROOFAI_HOME as string, "policies-config.json"), JSON.stringify(config)); } -/** An installed pack, written the way the loader verifies it. */ -function installPack(policies: Array>): void { +/** + * An installed pack, written the way the loader verifies it — beside + * FailproofAI/jev-policies unless `withJev` is false, because the checks the + * core pack's `reviewedBy` names live there. + */ +function installPack(policies: Array>, withJev = true): void { const artifact = "// a pack artifact this test never executes\n"; const digest = createHash("sha256").update(artifact).digest("hex"); mkdirSync(join(packRoot, "artifacts"), { recursive: true }); @@ -92,6 +97,19 @@ function installPack(policies: Array>): void { sha256: digest, policies, }, + ...(withJev + ? [ + { + id: "FailproofAI/jev-policies", + version: "0.2.0", + source: "github:FailproofAI/jev-policies@v0.2.0", + entry: `artifacts/${digest}.mjs`, + sha256: digest, + policies: [], + semantic: JEV_PACK_SEMANTIC_ENTRIES, + }, + ] + : []), ], }), ); @@ -154,29 +172,54 @@ describe("failproofai jev status — what Jev may clear", () => { enabled: PACKABLE.length + 1, reviewable: 0, customPolicyFiles: 0, + jevChecks: 16, }); expect(j.reviewablePolicies.problem).toContain(RETAKE_PACK_COMMAND); }); - it("reports the seven Jev may clear, and complains about nothing, on this build's builtins", async () => { - writeConfig({ enabledPolicies: POLICY_CATALOG.map((p) => p.name) }); + it("reports the fifteen Jev may clear, and complains about nothing, with the core pack and its Jev checks", async () => { + writeConfig({ enabledPolicies: [] }); + installPack(PACKABLE as unknown as Array>); await turnJevOn(); const r = await runJevCommand(["status"], RENDER); const out = text(r); - expect(out).toContain(`15 of ${POLICY_CATALOG.length} enabled policies are reviewable`); + expect(out).toContain(`15 of ${PACKABLE.length + 1} enabled policies are reviewable`); expect(out).toContain("Jev may clear a deny or an instruction from those, and from no others."); expect(out).not.toContain(RETAKE_PACK_COMMAND); + expect(out).not.toContain(JEV_POLICIES_ADD_COMMAND); const j = JSON.parse((await runJevCommand(["status", "--json"], RENDER)).json as string); expect(j.reviewablePolicies).toEqual({ - enabled: POLICY_CATALOG.length, + enabled: PACKABLE.length + 1, reviewable: 15, customPolicyFiles: 0, + jevChecks: 16, problem: null, }); }); + it("says Jev has no checks installed, and how to add them, when no pack declares any", async () => { + // The vanilla install: the package ships no Jev checks, so configuring Jev + // alone asks nothing. Status has to say so plainly, with the one command. + writeConfig({ enabledPolicies: POLICY_CATALOG.map((p) => p.name) }); + await turnJevOn(); + + const r = await runJevCommand(["status"], RENDER); + expect(r.exitCode).toBe(0); + const out = text(r); + expect(out).toContain("Jev has no checks installed"); + expect(out).toContain(JEV_POLICIES_ADD_COMMAND); + expect(out).toContain(`0 of ${POLICY_CATALOG.length} enabled policies are reviewable`); + + const j = JSON.parse((await runJevCommand(["status", "--json"], RENDER)).json as string); + expect(j.reviewablePolicies).toMatchObject({ reviewable: 0, jevChecks: 0, problem: NO_JEV_CHECKS_PROBLEM }); + + // The core pack alone is the same: its reviewedBy names checks that live elsewhere. + installPack(PACKABLE as unknown as Array>, false); + expect(text(await runJevCommand(["status"], RENDER))).toContain(JEV_POLICIES_ADD_COMMAND); + }); + it("says nothing about authority for a config the loader refused", async () => { writeConfig({ enabledPolicies: POLICY_CATALOG.map((p) => p.name) }); writeFileSync(join(process.env.FAILPROOFAI_HOME as string, "jev.json"), JSON.stringify({ provider: "nope" }), { mode: 0o600 }); diff --git a/__tests__/hooks/jev-cli-status-stats.test.ts b/__tests__/hooks/jev-cli-status-stats.test.ts index 46601bce2..2e4374e1f 100644 --- a/__tests__/hooks/jev-cli-status-stats.test.ts +++ b/__tests__/hooks/jev-cli-status-stats.test.ts @@ -128,23 +128,23 @@ describe("failproofai jev status — activity comes from jevStats()", () => { expectOneDefaultCall(); }); - it("shadow mode's would-be clears are printed, not reported as 'cleared nothing'", async () => { + it("observe mode's would-be clears are printed, not reported as 'cleared nothing'", async () => { // T8's jevStats() counts a clear that CHANGED an outcome in clearsByPolicy - // — which only enforce mode can do — and shadow mode's would-be clears in - // shadowClearsByPolicy. A renderer reading only the first tells a shadow - // user nothing was cleared, which is the one number shadow mode exists to + // — which only enforce mode can do — and observe mode's would-be clears in + // observeClearsByPolicy. A renderer reading only the first tells an observe + // user nothing was cleared, which is the one number observe mode exists to // show. The field is optional on the stats this branch builds against. jevStatsMock.mockResolvedValue({ ...STATS, clearsByPolicy: {}, - shadowClearsByPolicy: { "block-read-outside-cwd": 9, "protect-env-vars": 2 }, + observeClearsByPolicy: { "block-read-outside-cwd": 9, "protect-env-vars": 2 }, } as JevStats); - await runJevCommand(["setup", "--provider", "typesafe", "--mode", "shadow", "--key-stdin"], withKey(KEY)); + await runJevCommand(["setup", "--provider", "typesafe", "--mode", "observe", "--key-stdin"], withKey(KEY)); const human = await runJevCommand(["status"], RENDER); expect(human.exitCode).toBe(0); const out = text(human); - expect(out).toContain("would have cleared (shadow) block-read-outside-cwd ×9, protect-env-vars ×2"); + expect(out).toContain("would have cleared (observe) block-read-outside-cwd ×9, protect-env-vars ×2"); expect(out).toContain("cleared nothing"); }); diff --git a/__tests__/hooks/jev-cli-url-token.test.ts b/__tests__/hooks/jev-cli-url-token.test.ts index a16d87c42..8f1db5aab 100644 --- a/__tests__/hooks/jev-cli-url-token.test.ts +++ b/__tests__/hooks/jev-cli-url-token.test.ts @@ -268,15 +268,15 @@ describe("failproofai jev --url --token ", () => { expect(existsSync(jevConfigPath())).toBe(false); }); - it("refuses plain http to localhost in enforce mode, and takes it in shadow", async () => { + it("refuses plain http to localhost in enforce mode, and takes it in observe", async () => { const enforced = await runJevCommand(["--url", "http://127.0.0.1:8088/v1", "--token", TOKEN], RENDER); expect(enforced.exitCode).toBe(1); - expect(text(enforced)).toContain("accepted only with mode shadow"); + expect(text(enforced)).toContain("accepted only with mode observe"); expect(existsSync(jevConfigPath())).toBe(false); - const shadow = await runJevCommand(["--url", "http://127.0.0.1:8088/v1", "--mode", "shadow", "--token", TOKEN], RENDER); - expect(shadow.exitCode).toBe(0); - expect(readFile()).toMatchObject({ provider: "custom", baseUrl: "http://127.0.0.1:8088/v1", mode: "shadow" }); + const observe = await runJevCommand(["--url", "http://127.0.0.1:8088/v1", "--mode", "observe", "--token", TOKEN], RENDER); + expect(observe.exitCode).toBe(0); + expect(readFile()).toMatchObject({ provider: "custom", baseUrl: "http://127.0.0.1:8088/v1", mode: "observe" }); }); it("refuses a URL carrying credentials, in the loader's words", async () => { diff --git a/__tests__/hooks/jev-cli.test.ts b/__tests__/hooks/jev-cli.test.ts index 833664bb1..c7d231509 100644 --- a/__tests__/hooks/jev-cli.test.ts +++ b/__tests__/hooks/jev-cli.test.ts @@ -72,20 +72,20 @@ describe("failproofai jev", () => { expect(text(off)).toContain("off — Jev is not asked at all"); expect(text(off)).not.toContain("jev test"); // Switched back on, the step is back. - const on = await runJevCommand(["setup", "--mode", "shadow"], noTty); + const on = await runJevCommand(["setup", "--mode", "observe"], noTty); expect(on.exitCode, text(on)).toBe(0); expect(text(on)).toContain("failproofai jev test"); }); it("writes every option it is given", async () => { const r = await runJevCommand( - ["setup", "--provider=cloudflare", "--account-id", ACCOUNT, "--model", "typesafe/jev", "--mode", "shadow", "--timeout-ms", "900", "--key-stdin"], + ["setup", "--provider=cloudflare", "--account-id", ACCOUNT, "--model", "typesafe/jev", "--mode", "observe", "--timeout-ms", "900", "--key-stdin"], withKey(KEY), ); expect(r.exitCode).toBe(0); - expect(readFile()).toEqual({ provider: "cloudflare", apiKey: KEY, accountId: ACCOUNT, model: "typesafe/jev", mode: "shadow", timeoutMs: 900 }); + expect(readFile()).toEqual({ provider: "cloudflare", apiKey: KEY, accountId: ACCOUNT, model: "typesafe/jev", mode: "observe", timeoutMs: 900 }); expect(text(r)).toContain(`accounts/${ACCOUNT}/ai/run`); - expect(text(r)).toContain("shadow"); + expect(text(r)).toContain("observe"); }); it("needs --provider the first time, and a real one", async () => { @@ -159,21 +159,21 @@ describe("failproofai jev", () => { it("re-running for the same provider keeps the key, so a mode switch is one flag", async () => { await runJevCommand(["setup", "--provider", "cloudflare", "--account-id", ACCOUNT, "--key-stdin"], withKey(KEY)); - const r = await runJevCommand(["setup", "--mode", "shadow"], noTty); + const r = await runJevCommand(["setup", "--mode", "observe"], noTty); expect(r.exitCode).toBe(0); expect(text(r)).toContain("kept from the existing config"); - expect(readFile()).toEqual({ provider: "cloudflare", apiKey: KEY, accountId: ACCOUNT, mode: "shadow" }); + expect(readFile()).toEqual({ provider: "cloudflare", apiKey: KEY, accountId: ACCOUNT, mode: "observe" }); }); it("switching provider starts over: no key, model or URL carries across, mode does", async () => { - await runJevCommand(["setup", "--provider", "custom", "--base-url", "https://jev.example.com/v1", "--mode", "shadow", "--key-stdin"], withKey(KEY)); + await runJevCommand(["setup", "--provider", "custom", "--base-url", "https://jev.example.com/v1", "--mode", "observe", "--key-stdin"], withKey(KEY)); const noKey = await runJevCommand(["setup", "--provider", "vercel"], noTty); expect(noKey.exitCode).toBe(1); expect(readFile().provider).toBe("custom"); const r = await runJevCommand(["setup", "--provider", "vercel", "--key-stdin"], withKey(OTHER_KEY)); expect(r.exitCode).toBe(0); - expect(readFile()).toEqual({ provider: "vercel", apiKey: OTHER_KEY, mode: "shadow" }); + expect(readFile()).toEqual({ provider: "vercel", apiKey: OTHER_KEY, mode: "observe" }); }); it("`default` clears a model or base URL override", async () => { @@ -208,9 +208,9 @@ describe("failproofai jev", () => { expect(readFile()).toEqual({ provider: "typesafe" }); expect(loadJevConfig()?.apiKey).toBe(KEY); // A re-run for the same provider keeps it an environment-key config. - const again = await runJevCommand(["setup", "--mode", "shadow"], noTty); + const again = await runJevCommand(["setup", "--mode", "observe"], noTty); expect(again.exitCode).toBe(0); - expect(readFile()).toEqual({ provider: "typesafe", mode: "shadow" }); + expect(readFile()).toEqual({ provider: "typesafe", mode: "observe" }); // A variable that is set but malformed is refused, not stored around. process.env[JEV_API_KEY_ENV] = "two words"; expect((await runJevCommand(["setup", "--provider", "typesafe", "--key-from-env"], noTty)).exitCode).toBe(1); @@ -238,7 +238,7 @@ describe("failproofai jev", () => { }); it("shows provider, endpoint, model, mode, path and permissions — never the key", async () => { - await runJevCommand(["setup", "--provider", "openrouter", "--mode", "shadow", "--key-stdin"], withKey(KEY)); + await runJevCommand(["setup", "--provider", "openrouter", "--mode", "observe", "--key-stdin"], withKey(KEY)); const r = await runJevCommand(["status"], RENDER); expect(r.exitCode).toBe(0); const out = text(r); @@ -246,7 +246,7 @@ describe("failproofai jev", () => { expect(out).toContain("openrouter"); expect(out).toContain("https://openrouter.ai/api/v1/systemone"); expect(out).toContain("typesafe/jev-1.13 (provider default)"); - expect(out).toContain("shadow"); + expect(out).toContain("observe"); expect(out).toContain(jevConfigPath()); if (posix) expect(out).toContain("0600 (owner-only)"); expect(out).toContain("set in the config file"); @@ -323,6 +323,13 @@ describe("failproofai jev", () => { expect(lines).toContain("block-read-outside-cwd ×12, protect-env-vars ×3"); expect(jevStatsLines(null).join("\n")).toContain("could not be read"); }); + + it("prints observe-mode would-be clears", () => { + const base = { windowMs: 24 * 3_600_000, total: 3, fallbackRate: 0, fallbackReasons: {}, latencyP50Ms: 40, latencyP95Ms: 90, clearsByPolicy: {} }; + const now = jevStatsLines({ ...base, observeClearsByPolicy: { "block-env-files": 2 } }, { cols: 100 }).join("\n"); + expect(now).toContain("would have cleared (observe)"); + expect(now).toContain("block-env-files ×2"); + }); }); describe("test", () => { diff --git a/__tests__/hooks/jev-cloud-disconnect-race.test.ts b/__tests__/hooks/jev-cloud-disconnect-race.test.ts index ae9027c1f..eaebe1f66 100644 --- a/__tests__/hooks/jev-cloud-disconnect-race.test.ts +++ b/__tests__/hooks/jev-cloud-disconnect-race.test.ts @@ -75,7 +75,7 @@ const { removeCloudJevConfig } = await import("../../src/hooks/jev-cloud-connect const { jevConfigPath } = await import("../../src/hooks/semantic/jev-config"); const { runDisconnectCommand } = await import("../../src/hooks/cloud-enrollment-cli"); -const CLOUD = { provider: "failproofai", baseUrl: "https://app.befailproof.ai/enforcement/v1/jev", mode: "shadow" }; +const CLOUD = { provider: "failproofai", baseUrl: "https://app.befailproof.ai/enforcement/v1/jev", mode: "observe" }; // Built at runtime: this repo's own hooks refuse secret-shaped literals. const BYOK_KEY = ["ts", "byok", "0123456789abcdef"].join("-"); const BYOK = { provider: "typesafe", apiKey: BYOK_KEY, mode: "enforce" }; @@ -192,7 +192,7 @@ describe("disconnect never deletes a BYOK jev.json", () => { hook.after = step; hook.run = () => { const tmp = `${jevConfigPath()}.other.tmp`; - realFs.writeFileSync(tmp, JSON.stringify({ ...BYOK, mode: "shadow" }), { mode: 0o600 }); + realFs.writeFileSync(tmp, JSON.stringify({ ...BYOK, mode: "observe" }), { mode: 0o600 }); realFs.renameSync(tmp, jevConfigPath()); }; disconnect(); @@ -230,7 +230,7 @@ describe("putting a BYOK jev.json back where no hard link can be made", () => { it("…and the copy never lands over a file written meanwhile: both kept, the original set aside", () => { writeAtomically(BYOK); const original = readFileSync(jevConfigPath(), "utf8"); - const other = { ...BYOK, mode: "shadow" }; + const other = { ...BYOK, mode: "observe" }; inject.linkSync = { ...EPERM, before: () => writeAtomically(other) }; const r = removeCloudJevConfig(); expect(r.status).toBe("set-aside"); @@ -282,7 +282,7 @@ describe("putting a BYOK jev.json back where no hard link can be made", () => { it("…and a set-aside one too", () => { writeAtomically(BYOK); - inject.linkSync = { ...EPERM, before: () => writeAtomically({ ...BYOK, mode: "shadow" }) }; + inject.linkSync = { ...EPERM, before: () => writeAtomically({ ...BYOK, mode: "observe" }) }; const text = runDisconnectCommand().lines.join("\n"); expect(text).toContain("This machine is not connected to FailproofAI Cloud."); expect(text).toContain("another jev.json was written in its place"); diff --git a/__tests__/hooks/jev-env-key.test.ts b/__tests__/hooks/jev-env-key.test.ts index 0e083b329..c86e56fe1 100644 --- a/__tests__/hooks/jev-env-key.test.ts +++ b/__tests__/hooks/jev-env-key.test.ts @@ -101,14 +101,14 @@ describe("a config whose key comes from the environment, in a shell without it", }); it("status --json reports a configured machine, not an invalid one", async () => { - await setupFromEnv("--provider", "vercel", "--mode", "shadow"); + await setupFromEnv("--provider", "vercel", "--mode", "observe"); const r = await runJevCommand(["status", "--json"], RENDER); expect(r.exitCode).toBe(0); const j = JSON.parse(r.json as string) as Record; expect(j.status).toBe("key-missing"); expect(j.reason).toBe("no-env-key"); expect(j.provider).toBe("vercel"); - expect(j.mode).toBe("shadow"); + expect(j.mode).toBe("observe"); expect(j.keySource).toBe("env"); expect(j.keyEnvVar).toBe(JEV_API_KEY_ENV); expect(String(j.endpoint)).toContain("vercel"); diff --git a/__tests__/hooks/jev-field-shapes.test.ts b/__tests__/hooks/jev-field-shapes.test.ts index b69eb2454..4fa133e0f 100644 --- a/__tests__/hooks/jev-field-shapes.test.ts +++ b/__tests__/hooks/jev-field-shapes.test.ts @@ -135,9 +135,9 @@ describe("a clear of a reviewable policy whose name has spaces", () => { it("is counted by jev status, shown on the dashboard and sent to PostHog", () => { const name = REGISTERED_WITH_SPACES[0]; - const s = computeJevStats([cleared(name), cleared(name, { jevMode: "shadow" })], { now: 6_000, windowMs: 60_000 }); + const s = computeJevStats([cleared(name), cleared(name, { jevMode: "observe" })], { now: 6_000, windowMs: 60_000 }); expect(s.clearsByPolicy).toEqual({ [name]: 1 }); - expect(s.shadowClearsByPolicy).toEqual({ [name]: 1 }); + expect(s.observeClearsByPolicy).toEqual({ [name]: 1 }); expect(describeJevActivity(cleared(name))).toContain(`cleared ${name}`); const props = jevTelemetryProperties(cleared(name)); expect(props.jev_cleared).toEqual([name]); diff --git a/__tests__/hooks/jev-no-request.test.ts b/__tests__/hooks/jev-no-request.test.ts index 8e953f6fc..0ede79d98 100644 --- a/__tests__/hooks/jev-no-request.test.ts +++ b/__tests__/hooks/jev-no-request.test.ts @@ -56,7 +56,7 @@ function row(overrides: Partial = {}): HookActivityEntry { } /** Exactly what the two-tier path records for a TodoWrite call: nothing to ask, no request sent. */ -const noRequest = (mode: "shadow" | "enforce" = "enforce", ts = NOW - 1_000) => +const noRequest = (mode: "observe" | "enforce" = "enforce", ts = NOW - 1_000) => row({ timestamp: ts, toolName: "TodoWrite", evaluator: "jev", jevDecision: "allow", jevMode: mode }); const answered = (latency: number, extra: Partial = {}) => row({ evaluator: "jev", jevDecision: "allow", jevLatencyMs: latency, jevModel: "jev-1.13.0", jevMode: "enforce", ...extra }); @@ -66,7 +66,7 @@ const hardDeny = () => row({ decision: "deny", policyName: "block-sudo", evaluat describe("jevOutcome: a call Jev sent no request for", () => { it("classifies the recorded no-request row as no-request, not answered", () => { expect(jevOutcome(noRequest("enforce"))).toBe("no-request"); - expect(jevOutcome(noRequest("shadow"))).toBe("no-request"); + expect(jevOutcome(noRequest("observe"))).toBe("no-request"); expect(jevOutcome({ evaluator: "jev", jevDecision: "allow" })).toBe("no-request"); for (const r of JEV_NO_REQUEST_ROWS) expect(jevOutcome(r), r.toolName ?? "").toBe("no-request"); }); @@ -103,7 +103,7 @@ describe("jev status stats", () => { expect(s.notConsulted).toBe(0); expect(s.fallbackRate).toBeCloseTo(0.5); expect(s.decisions).toEqual({ allow: 1, instruct: 0, deny: 0 }); - expect(s.modes).toEqual({ shadow: 0, enforce: 2 }); + expect(s.modes).toEqual({ observe: 0, enforce: 2 }); }); it("prints the no-request calls on a line of their own", () => { @@ -124,7 +124,7 @@ describe("jev status stats", () => { }); it("reports no evaluations when Jev never had anything to ask", () => { - const s = computeJevStats([noRequest(), noRequest("shadow")], { now: NOW, windowMs: 3_600_000 }); + const s = computeJevStats([noRequest(), noRequest("observe")], { now: NOW, windowMs: 3_600_000 }); expect(s.total).toBe(0); expect(s.answered).toBe(0); expect(s.noRequest).toBe(2); @@ -139,7 +139,7 @@ describe("jev status stats", () => { describe("the dashboard summary", () => { it("says no request was sent, and claims no verdict", () => { - for (const mode of ["enforce", "shadow"] as const) { + for (const mode of ["enforce", "observe"] as const) { expect(describeJevActivity(noRequest(mode))).toEqual([JEV_NO_REQUEST_FACT]); } }); diff --git a/__tests__/hooks/jev-not-consulted.test.ts b/__tests__/hooks/jev-not-consulted.test.ts index 8217430c6..08b6bf763 100644 --- a/__tests__/hooks/jev-not-consulted.test.ts +++ b/__tests__/hooks/jev-not-consulted.test.ts @@ -53,7 +53,7 @@ function row(overrides: Partial = {}): HookActivityEntry { } /** Exactly what the combine rules record for a hard deny: Jev aborted, never read. */ -const notConsulted = (mode: "shadow" | "enforce" = "enforce", ts = NOW - 1_000) => +const notConsulted = (mode: "observe" | "enforce" = "enforce", ts = NOW - 1_000) => row({ timestamp: ts, decision: "deny", @@ -69,7 +69,7 @@ const timedOut = () => row({ evaluator: "jev-fallback", jevFallbackReason: "time describe("jevOutcome", () => { it("classifies the combine rules' not-consulted row as not consulted", () => { expect(jevOutcome(notConsulted("enforce"))).toBe("not-consulted"); - expect(jevOutcome(notConsulted("shadow"))).toBe("not-consulted"); + expect(jevOutcome(notConsulted("observe"))).toBe("not-consulted"); expect(jevOutcome({ evaluator: "jev" })).toBe("not-consulted"); for (const r of JEV_NOT_CONSULTED_ROWS) expect(jevOutcome(r)).toBe("not-consulted"); }); @@ -113,7 +113,7 @@ describe("jev status stats", () => { expect(s.fallbackRate).toBeCloseTo(0.5); expect(s.answered + s.fallbacks).toBe(s.total); expect(s.decisions.allow + s.decisions.instruct + s.decisions.deny).toBe(s.answered); - expect(s.modes).toEqual({ shadow: 0, enforce: 2 }); + expect(s.modes).toEqual({ observe: 0, enforce: 2 }); expect(s.latencyP50Ms).toBe(40); }); @@ -131,7 +131,7 @@ describe("jev status stats", () => { }); it("reports no evaluations when every Jev row was a hard deny", () => { - const s = computeJevStats([notConsulted(), notConsulted("shadow")], { now: NOW, windowMs: 3_600_000 }); + const s = computeJevStats([notConsulted(), notConsulted("observe")], { now: NOW, windowMs: 3_600_000 }); expect(s.total).toBe(0); expect(s.answered).toBe(0); expect(s.fallbackRate).toBe(0); @@ -144,7 +144,7 @@ describe("jev status stats", () => { describe("the dashboard summary", () => { it("says Jev was not consulted, rather than an empty summary", () => { - for (const mode of ["enforce", "shadow"] as const) { + for (const mode of ["enforce", "observe"] as const) { expect(describeJevActivity(notConsulted(mode))).toEqual([JEV_NOT_CONSULTED_FACT]); } }); diff --git a/__tests__/hooks/jev-policy-page-golden.test.ts b/__tests__/hooks/jev-policy-page-golden.test.ts index 3b29fe8a2..08a598ca3 100644 --- a/__tests__/hooks/jev-policy-page-golden.test.ts +++ b/__tests__/hooks/jev-policy-page-golden.test.ts @@ -1,7 +1,7 @@ // @vitest-environment node /** * The collector's golden rows for the policy page's Jev data (contract §5): - * a Jev-decided enforce row attributed `policySource: "jev"`, and shadow rows + * a Jev-decided enforce row attributed `policySource: "jev"`, and observe rows * whose `observed` list carries Jev's "would have". * * `crates/fpai-collect/tests/hooks_jev.rs` reads the golden file this test @@ -42,7 +42,7 @@ describe("the collector's policy-page golden rows", () => { for (const row of [wouldDeny, wouldWarn]) { expect(row.decision).toBe("allow"); expect(row.policySource).toBeUndefined(); - expect(row.jevMode).toBe("shadow"); + expect(row.jevMode).toBe("observe"); expect(row.observed).toHaveLength(1); const [o] = row.observed!; expect(o.policyId).toMatch(/^semantic\/[a-z0-9-]+$/); diff --git a/__tests__/hooks/jev-telemetry-privacy.test.ts b/__tests__/hooks/jev-telemetry-privacy.test.ts index c406aba81..f28cb4a8a 100644 --- a/__tests__/hooks/jev-telemetry-privacy.test.ts +++ b/__tests__/hooks/jev-telemetry-privacy.test.ts @@ -29,6 +29,11 @@ import { _resetForTest, persistHookActivity, type HookActivityEntry } from "../. import { jevTelemetryProperties, trackHookEvent } from "../../src/hooks/hook-telemetry"; import { describeJevActivity } from "../../src/hooks/jev-activity"; import { computeJevStats, formatJevStats } from "../../src/hooks/semantic/jev-stats"; +import { withInstalledJevPoliciesPack } from "../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); // Marker words that appear in the command, the prompt and the agent message, // and in nothing a policy or the evaluator writes on its own. @@ -56,7 +61,7 @@ const answering = }); /** Record an outcome the way the handler does — greedily (see the header). */ -function record(outcome: SemanticOutcome, mode: "shadow" | "enforce"): HookActivityEntry { +function record(outcome: SemanticOutcome, mode: "observe" | "enforce"): HookActivityEntry { const base: HookActivityEntry = { timestamp: Date.now(), eventType: "PreToolUse", @@ -161,7 +166,7 @@ describe("Jev telemetry privacy", () => { for (const [name, run] of cases) { it(`${name}: nothing identifying reaches the row, PostHog, the dashboard or the stats`, async () => { const outcome = await run(); - for (const mode of ["enforce", "shadow"] as const) { + for (const mode of ["enforce", "observe"] as const) { const entry = record(outcome, mode); persistHookActivity(entry); @@ -207,7 +212,7 @@ describe("Jev telemetry privacy", () => { `git push origin failproofai/zebra-archive && ${COMMAND}`, `mv pack/tangerine-ledger.csv cloud/marmalade review`, ]; - const poisonedRow = (mode: "shadow" | "enforce", overrides: Partial = {}): HookActivityEntry => ({ + const poisonedRow = (mode: "observe" | "enforce", overrides: Partial = {}): HookActivityEntry => ({ timestamp: Date.now(), eventType: "PreToolUse", integration: "claude", @@ -226,7 +231,7 @@ describe("Jev telemetry privacy", () => { jevMode: mode, ...overrides, }); - const rows: Array<[string, (mode: "shadow" | "enforce") => HookActivityEntry]> = [ + const rows: Array<[string, (mode: "observe" | "enforce") => HookActivityEntry]> = [ ["answered", (mode) => poisonedRow(mode)], [ "fell back", @@ -242,7 +247,7 @@ describe("Jev telemetry privacy", () => { for (const [name, make] of rows) { it(`${name}: none of it reaches the row, PostHog, the dashboard or the stats`, async () => { const entries: HookActivityEntry[] = []; - for (const mode of ["enforce", "shadow"] as const) { + for (const mode of ["enforce", "observe"] as const) { const entry = make(mode); entries.push(entry); persistHookActivity(entry); diff --git a/__tests__/hooks/manager.test.ts b/__tests__/hooks/manager.test.ts index 786e3b360..2139903d3 100644 --- a/__tests__/hooks/manager.test.ts +++ b/__tests__/hooks/manager.test.ts @@ -16,6 +16,22 @@ vi.mock("node:fs", () => ({ vi.mock("node:child_process", () => ({ execSync: vi.fn(), })); +// Agent configs are written through the crash-safe writer (temp file, fsync, +// rename). This suite mocks node:fs wholesale and asserts on writeFileSync, so +// route that writer to the mocked writeFileSync: every assertion below is about +// WHAT is written where, which is unchanged. The writer itself is covered by +// safe-config-write.test.ts against a real filesystem. +vi.mock("../../src/hooks/safe-config-write", async () => { + const fs = await import("node:fs"); + const { dirname } = await import("node:path"); + return { + writeConfigFileAtomic: (path: string, content: string) => { + fs.mkdirSync(dirname(path), { recursive: true }); + fs.writeFileSync(path, content, "utf8"); + }, + configBackupPath: (path: string) => `${path}.failproofai-backup`, + }; +}); vi.mock("../../src/hooks/install-prompt", () => ({ promptPolicySelection: vi.fn(() => @@ -54,6 +70,7 @@ vi.mock("../../src/hooks/pack-store", () => ({ vi.mock("../../src/hooks/pack-manifest", () => ({ hasInstalledPacks: vi.fn(() => false), + hasInstalledRegexPacks: vi.fn(() => false), readInstalledPacks: vi.fn(() => ({ packs: [], errors: [] })), })); diff --git a/__tests__/hooks/new-telemetry.test.ts b/__tests__/hooks/new-telemetry.test.ts index b1de679ef..96742e6e3 100644 --- a/__tests__/hooks/new-telemetry.test.ts +++ b/__tests__/hooks/new-telemetry.test.ts @@ -13,6 +13,7 @@ import { execSync } from "node:child_process"; vi.mock("../../src/hooks/pack-manifest", () => ({ readInstalledPacks: vi.fn(() => ({ packs: [], errors: [] })), hasInstalledPacks: vi.fn(() => false), + hasInstalledRegexPacks: vi.fn(() => false), })); vi.mock("node:fs", () => ({ diff --git a/__tests__/hooks/pack-jev-checks.test.ts b/__tests__/hooks/pack-jev-checks.test.ts index 45625832b..9d3b3663d 100644 --- a/__tests__/hooks/pack-jev-checks.test.ts +++ b/__tests__/hooks/pack-jev-checks.test.ts @@ -88,15 +88,14 @@ describe("the section's shape", () => { expect(text).toContain("`failproofai policies` never lists them"); }); - it("says a third party's checks are added to the built-in ones, and only FailproofAI's replace them", () => { - // `policies show` and the picker said "replace" for every pack after the - // resolver started ADDING a stranger's checks; `add` already said "added to". + it("says a pack's checks are asked, whoever published it — the package ships none to replace or add to", () => { const pack = { policies: [policy("block-rm-rf")], semantic: [check("acme-check")] }; const third = jevChecksSection({ ...pack, source: "github:acme/x@1.0.0" }, OPTS)!.join("\n"); - expect(third).toContain("added to this build's own checks"); - expect(third).not.toMatch(/replac/); const first = jevChecksSection({ ...pack, source: "github:FailproofAI/jev-policies@1.0.0" }, OPTS)!.join("\n"); - expect(first).toContain("replacing this build's own set"); + for (const text of [third, first]) { + expect(text).toContain("asked by Jev on every tool call they apply to"); + expect(text).not.toMatch(/replac|this build's own/); + } }); it("gives every row its mode, because that decides what pairing with it can do", () => { @@ -359,7 +358,7 @@ describe("failproofai policies show ", () => { const r = await runPackCommand(["add", "acme/guards@v1.2.0", "--all"]); expect(r.exitCode, r.lines.join("\n")).toBe(0); const text = r.lines.join("\n"); - expect(text).toContain("1 Jev check, added to this build's own checks."); + expect(text).toContain("1 Jev check, asked by Jev on every tool call they apply to."); expect(text).not.toContain("for Jev"); }); @@ -368,31 +367,34 @@ describe("failproofai policies show ", () => { release({ effect: "observe", policies: [], semantic: [check("obs-zebra")] }); const text = (await runPackCommand(["add", "acme/guards@v1.2.0", "--all"])).lines.join("\n"); expect(text).toMatch(/1 Jev check, not asked/); - expect(text).not.toContain("added to this build's own checks"); + expect(text).not.toContain("asked by Jev on every tool call"); }); it("says a --cli pack's checks apply to those agents only", async () => { release({ policies: [], semantic: [check("codex-walrus")] }); const r = await runPackCommand(["add", "acme/guards@v1.2.0", "--all", "--cli", "codex"]); expect(r.exitCode, r.lines.join("\n")).toBe(0); - expect(r.lines.join("\n")).toMatch(/added to this build's own checks, for codex only/); + expect(r.lines.join("\n")).toMatch(/asked by Jev on every tool call they apply to, for codex only/); }); it("says at install which of its checks this machine will never ask, and why", async () => { - // Each fits a pack's budget alone; beside the built-in checks (a third - // party's join them) only the first fits what is left of one request. + // Each fits a pack's budget alone; together only the first three fit what + // one request can carry. const probe = (i: number) => ({ id: `p${i}`, instructions: "x".repeat(600), criteria: { true: "t".repeat(300), false: "f".repeat(300) }, }); const big = (name: string) => check(name, { userCanOverride: false, probes: [0, 1, 2, 3, 4, 5].map(probe) }); - release({ policies: [], semantic: [big("acme-a"), big("acme-b"), check("destructive-deletion")] }); + release({ + policies: [], + semantic: [big("acme-a"), big("acme-b"), big("acme-c"), big("acme-d"), check("destructive-deletion")], + }); const r = await runPackCommand(["add", "acme/guards@v1.2.0", "--all"]); expect(r.exitCode, r.lines.join("\n")).toBe(0); const text = r.lines.join("\n"); - expect(text).toMatch(/acme\/guards semantic policy acme-b was dropped: its questions need/); - expect(text).not.toMatch(/acme-a was dropped/); + expect(text).toMatch(/acme\/guards semantic policy acme-d was dropped: its questions need/); + expect(text).not.toMatch(/acme-[abc] was dropped/); expect(text).toMatch(/declares semantic policy destructive-deletion, a name reserved/); }); @@ -401,7 +403,7 @@ describe("failproofai policies show ", () => { // reserved name, so that pack's version is never asked. Only add said so. const text = (await show()).join("\n"); expect(text).toMatch(/declares semantic policy destructive-deletion, a name reserved/); - expect(text).toContain("added to this build's own checks"); + expect(text).toContain("asked by Jev on every tool call they apply to"); }); it("sits under the policy rows, since a check is read against what it can clear", async () => { diff --git a/__tests__/hooks/pack-manifest.test.ts b/__tests__/hooks/pack-manifest.test.ts index dbffb4087..f0224001c 100644 --- a/__tests__/hooks/pack-manifest.test.ts +++ b/__tests__/hooks/pack-manifest.test.ts @@ -194,6 +194,42 @@ describe("readInstalledPacks", () => { // showed up later as a deny narrowed to the letters of that string — a guard // matching no event that exists. Refused here, where the publisher can still // fix it, rather than surviving on disk as metadata nothing can read. +describe("hasInstalledRegexPacks — a Jev-only pack is not a regex pack", () => { + async function regexPacks() { + const mod = await import("../../src/hooks/pack-manifest"); + return { regex: mod.hasInstalledRegexPacks(), any: mod.hasInstalledPacks() }; + } + + it("is false with no manifest at all", async () => { + expect(await regexPacks()).toEqual({ regex: false, any: false }); + }); + + it("is true for a pack that carries regex policies", async () => { + writeManifest([pack()]); + expect(await regexPacks()).toEqual({ regex: true, any: true }); + }); + + it("is false for a semantic-only pack like FailproofAI/jev-policies", async () => { + // The shim and `policies add ` must not treat this as "a pack now + // enforces the regex policies": it enforces none of them. + writeManifest([pack({ id: "FailproofAI/jev-policies", policies: [], semantic: [{ name: "destructive-deletion" }] })]); + expect(await regexPacks()).toEqual({ regex: false, any: true }); + }); + + it("is true once a regex pack sits beside the Jev-only one", async () => { + writeManifest([ + pack({ id: "FailproofAI/jev-policies", policies: [], semantic: [{ name: "destructive-deletion" }] }), + pack({ id: "FailproofAI/policies" }), + ]); + expect(await regexPacks()).toEqual({ regex: true, any: true }); + }); + + it("is false for a malformed manifest rather than throwing", async () => { + writeFileSync(join(root, "installed.json"), "{not json"); + expect(await regexPacks()).toEqual({ regex: false, any: false }); + }); +}); + describe("parsePackPolicy — the shape of a match, not just its presence", () => { const good = { name: "block-refunds", diff --git a/__tests__/hooks/pack-semantic-build.test.ts b/__tests__/hooks/pack-semantic-build.test.ts index b3e224630..8c9cf3c3f 100644 --- a/__tests__/hooks/pack-semantic-build.test.ts +++ b/__tests__/hooks/pack-semantic-build.test.ts @@ -16,7 +16,7 @@ import { tmpdir } from "node:os"; import { join } from "node:path"; import { findEntry, runPackCommand } from "@/src/hooks/pack-cli"; import { parsePackSemanticPolicy, readInstalledPacks } from "@/src/hooks/pack-manifest"; -import { BUILTIN_QUESTION_CHARS, MAX_PACK_QUESTION_CHARS } from "@/src/hooks/semantic/pack-policies"; +import { JEV_POLICIES_QUESTION_CHARS, MAX_PACK_QUESTION_CHARS } from "@/src/hooks/semantic/pack-policies"; import { version as packageVersion } from "../../package.json"; /** A probe declaration, as an entry file writes it. */ @@ -165,8 +165,7 @@ describe("build emits the semantic array", () => { it("omits the key entirely when nothing declared one", async () => { // An EMPTY array would still read as "a pack that declares semantic - // entries", and the replacement rule turns that into "replaced the - // compiled-in set with nothing". + // entries" — a pack giving Jev checks when it gives none. const entry = write("policies.mjs", ` import { customPolicies, deny } from "failproofai"; customPolicies.add({ name: "block-x", description: "d", match: { events: ["PreToolUse"] }, fn: async () => deny("no") }); @@ -259,10 +258,10 @@ describe("build emits the semantic array", () => { expect(r.lines.join("\n")).toMatch(new RegExp(`over the ${MAX_PACK_QUESTION_CHARS} one Jev request has room for`)); }); - it("judges a pack from outside FailproofAI against what the built-in checks leave, not the whole request", async () => { - // Every machine spends BUILTIN_QUESTION_CHARS on the built-in checks before a - // third party's, so two ~7.7k checks published cleanly and the second was - // dropped on every install. + it("judges a pack from outside FailproofAI against what FailproofAI/jev-policies leaves, not the whole request", async () => { + // A machine with FailproofAI/jev-policies spends JEV_POLICIES_QUESTION_CHARS + // on its checks before a third party's, so two ~7.7k checks would publish + // cleanly and the second be dropped on every such install. const check = (name: string) => ` semanticPolicies.add({ name: "${name}", title: "t", appliesTo: ["shell"], mode: "deny", userCanOverride: false, @@ -271,12 +270,12 @@ describe("build emits the semantic array", () => { guidance: "g", });`; const body = `import { semanticPolicies } from "failproofai";\n${check("xa-check-1")}\n${check("xa-check-2")}`; - const left = MAX_PACK_QUESTION_CHARS - BUILTIN_QUESTION_CHARS; + const left = MAX_PACK_QUESTION_CHARS - JEV_POLICIES_QUESTION_CHARS; const r = await build(write("policies.mjs", body)); expect(r.exitCode, r.lines.join("\n")).toBe(1); expect(r.lines.join("\n")).toMatch(new RegExp(`over the ${left} `)); - expect(r.lines.join("\n")).toMatch(/built-in checks/); + expect(r.lines.join("\n")).toMatch(/FailproofAI\/jev-policies' 16 checks/); const firstParty = await build(write("first-party-policies.mjs", body), ["--repo", "FailproofAI/jev-policies"]); expect(firstParty.exitCode, firstParty.lines.join("\n")).toBe(0); @@ -285,10 +284,10 @@ describe("build emits the semantic array", () => { describe("authority against the pack's own semantic policies", () => { it("publishes a reviewedBy that names one of them", async () => { - // The load-bearing case: a pack carrying both tiers replaces the compiled-in - // semantic set where it installs, so its regex policies must be able to name - // its OWN checks. Judged against this build's sixteen, this would be - // "a check this build does not have" and silently downgraded to hard. + // The load-bearing case: a pack carrying both tiers brings the checks its + // regex policies name, so they must be able to name its OWN checks. Judged + // against FailproofAI's sixteen names, this would be "a check this build + // does not have" and silently downgraded to hard. const r = await build(write("policies.mjs", BOTH_ENTRY)); expect(r.exitCode, r.lines.join("\n")).toBe(0); const entry = manifestOf(join(work, "out")).policies.find((p) => p.name === "block-big-refund"); @@ -313,9 +312,9 @@ describe("authority against the pack's own semantic policies", () => { expect(text).toMatch(/declares Jev checks of its own, so reviewedBy may name only those/); }); - it("falls back to this build's names for a pack with no semantic entries", async () => { - // Those machines keep running the compiled-in set, so a builtin name is the - // right thing for such a pack to review by. + it("judges a pack with no semantic entries against FailproofAI's check names", async () => { + // Such a pack names checks that ship in FailproofAI/jev-policies (the core + // pack is exactly this), so those names are the right thing to review by. const entry = write("policies.mjs", ` import { customPolicies, deny } from "failproofai"; customPolicies.add({ diff --git a/__tests__/hooks/pack-semantic-contested.test.ts b/__tests__/hooks/pack-semantic-contested.test.ts index c657462ef..1cdb23f81 100644 --- a/__tests__/hooks/pack-semantic-contested.test.ts +++ b/__tests__/hooks/pack-semantic-contested.test.ts @@ -290,7 +290,7 @@ describe("a second pack claiming a check another pack's policies name", () => { }); describe("a third-party pack claiming a builtin check name", () => { - /** The core pack: regex only, reviewable by the compiled-in check. */ + /** The core pack: regex only, reviewable by a check FailproofAI/jev-policies ships. */ const CORE: PackInput = { id: "FailproofAI/policies", version: "1.0.0", @@ -304,22 +304,19 @@ describe("a third-party pack claiming a builtin check name", () => { semantic: [semantic("destructive-deletion", { mode: "instruct" })], }; - // The impostor's question is never asked: the compiled-in set stands, so the - // core policy is still reviewable — by FailproofAI's own check. + // The impostor's question is never asked and names no reviewer: with no + // FailproofAI/jev-policies installed there is no check of that name at all, + // so the core policy is HARD — never cleared by the stranger's question. it.each([ ["its own id", EXTRAS], ["a forged FailproofAI id", { ...EXTRAS, id: "FailproofAI/jev-policies", source: "github:acme/jev-extras@v0.1.0" }], ])("is not the reviewer that clears the core pack's policy (%s)", async (_label, extras) => { install([CORE, extras]); const registered = await registeredAfterOneEvent(); - expect(authorityOf(registered.get("pack/FailproofAI/policies@1.0.0/block-rm-rf"))).toEqual({ - authority: "reviewable", - reviewedBy: ["destructive-deletion"], - }); + expect(authorityOf(registered.get("pack/FailproofAI/policies@1.0.0/block-rm-rf"))).toEqual({ authority: "hard" }); vi.resetModules(); const { resolveSemanticPolicies } = await import("@/src/hooks/semantic/pack-policies"); - const { SEMANTIC_POLICIES } = await import("@/src/hooks/semantic/policies"); - expect(resolveSemanticPolicies()).toBe(SEMANTIC_POLICIES); + expect(resolveSemanticPolicies()).toEqual([]); expect(stderr.join("")).toMatch(/declares semantic policy destructive-deletion, a name reserved/); }); @@ -346,17 +343,24 @@ describe("a third-party pack claiming a builtin check name", () => { }); }); -it("a stranger's own checks leave the core pack's policy reviewable by the built-in check", async () => { - install([ - { - id: "FailproofAI/policies", - version: "1.0.0", - policies: [regex("block-rm-rf", { authority: "reviewable", reviewedBy: ["destructive-deletion"] })], - artifact: artifactFor("FailproofAI/policies", ["block-rm-rf"]), - }, - { id: "acme/db", version: "0.1.0", policies: [], semantic: [semantic("acme-db-check")] }, - ]); - const registered = await registeredAfterOneEvent(); +it("a stranger's own checks neither make nor unmake the core pack's policy reviewable", async () => { + const core = { + id: "FailproofAI/policies", + version: "1.0.0", + policies: [regex("block-rm-rf", { authority: "reviewable", reviewedBy: ["destructive-deletion"] })], + artifact: artifactFor("FailproofAI/policies", ["block-rm-rf"]), + }; + const stranger = { id: "acme/db", version: "0.1.0", policies: [], semantic: [semantic("acme-db-check")] }; + // Without FailproofAI/jev-policies nothing supplies the check it names: hard. + install([core, stranger]); + let registered = await registeredAfterOneEvent(); + expect(authorityOf(registered.get("pack/FailproofAI/policies@1.0.0/block-rm-rf"))).toEqual({ authority: "hard" }); + + // With it, reviewable by FailproofAI's own check, the stranger's beside it. + vi.resetModules(); + const jev = { id: "FailproofAI/jev-policies", version: "1.0.0", policies: [], semantic: [semantic("destructive-deletion")] }; + install([core, stranger, jev]); + registered = await registeredAfterOneEvent(); expect(authorityOf(registered.get("pack/FailproofAI/policies@1.0.0/block-rm-rf"))).toEqual({ authority: "reviewable", reviewedBy: ["destructive-deletion"], diff --git a/__tests__/hooks/pack-semantic-manifest.test.ts b/__tests__/hooks/pack-semantic-manifest.test.ts index 8683345a3..cc09e73b4 100644 --- a/__tests__/hooks/pack-semantic-manifest.test.ts +++ b/__tests__/hooks/pack-semantic-manifest.test.ts @@ -29,8 +29,9 @@ import { contestedSemanticNames, effectiveReviewerNames, forgetEffectiveReviewerNames, + jevChecksDeclared, + jevChecksInstalled, } from "@/src/hooks/effective-reviewers"; -import { SEMANTIC_REVIEWER_NAMES } from "@/src/hooks/policy-authority"; import { missingGuards } from "@/src/hooks/pack-failclosed"; import { PACK_PRECONDITION_NAMES } from "@/src/hooks/semantic/precondition-names"; import { version as packageVersion } from "../../package.json"; @@ -444,15 +445,24 @@ describe("readInstalledPacks with semantic entries", () => { }); describe("effectiveReviewerNames", () => { - it("is this build's set when no pack is installed", () => { - expect(effectiveReviewerNames()).toBe(SEMANTIC_REVIEWER_NAMES); + it("is empty when no pack is installed — the package ships no Jev checks", () => { + expect(effectiveReviewerNames().size).toBe(0); + expect(jevChecksInstalled()).toBe(false); }); - it("is this build's set when the installed packs declare no semantic entries", () => { - // A pack that carries only the regex floor leaves the compiled-in semantic - // set running, so its reviewer names are the live ones. + it("is empty when the installed packs declare no semantic entries", () => { + // A pack that carries only the regex floor gives Jev nothing to ask, so its + // `reviewedBy` names nothing that can be answered. writeManifest([record()]); - expect(effectiveReviewerNames()).toBe(SEMANTIC_REVIEWER_NAMES); + expect(effectiveReviewerNames().size).toBe(0); + expect(jevChecksInstalled()).toBe(false); + expect(jevChecksDeclared()).toBe(false); + }); + + it("is empty, never a compiled-in set, when the manifest is unreadable", () => { + writeFileSync(join(root, "installed.json"), "{ not json"); + expect(effectiveReviewerNames().size).toBe(0); + expect(jevChecksDeclared()).toBe(false); }); it("is the pack's names once a FailproofAI pack declares any", () => { @@ -464,14 +474,16 @@ describe("effectiveReviewerNames", () => { expect(names.has("destructive-deletion")).toBe(false); }); - it("adds a third-party pack's names to this build's set", () => { + it("is a third-party pack's names alone, with nothing compiled in beside them", () => { writeManifest([record({ semantic: [entry({ name: "pack-only-check" })] })]); - expect([...effectiveReviewerNames()]).toEqual([...SEMANTIC_REVIEWER_NAMES, "pack-only-check"]); + expect([...effectiveReviewerNames()]).toEqual(["pack-only-check"]); + expect(jevChecksInstalled()).toBe(true); + expect(jevChecksDeclared()).toBe(true); }); it("re-reads when the manifest changes under it", () => { writeManifest([record()]); - expect(effectiveReviewerNames()).toBe(SEMANTIC_REVIEWER_NAMES); + expect(effectiveReviewerNames().size).toBe(0); writeManifest([record({ version: "1.3.0", semantic: [entry({ name: "pack-only-check" })] })]); expect(effectiveReviewerNames().has("pack-only-check")).toBe(true); }); @@ -511,10 +523,10 @@ describe("effectiveReviewerNames", () => { expect(effectiveReviewerNames().has("pack-only-check")).toBe(true); }); - it("falls back to this build's set when every declared name is contested", () => { + it("is empty when every declared name is contested", () => { // Which is what `semanticPoliciesFromPacks` does with the QUESTIONS in the - // same state — every entry dropped leaves the compiled-in set live — so the - // names honoured here stay the names of the questions that get asked. + // same state — nothing is asked — so the names honoured here stay the names + // of the questions that get asked: none. writeManifest([ record({ semantic: [entry({ name: "pack-only-check" })] }), record({ @@ -524,7 +536,7 @@ describe("effectiveReviewerNames", () => { semantic: [entry({ name: "pack-only-check", guidance: "Nothing to see here." })], }), ]); - expect(effectiveReviewerNames()).toBe(SEMANTIC_REVIEWER_NAMES); + expect(effectiveReviewerNames().size).toBe(0); }); it("names both claimants, so the log says which packs disagree", () => { diff --git a/__tests__/hooks/pack-semantic-reviewability.test.ts b/__tests__/hooks/pack-semantic-reviewability.test.ts index f27bf1d1b..22da21af5 100644 --- a/__tests__/hooks/pack-semantic-reviewability.test.ts +++ b/__tests__/hooks/pack-semantic-reviewability.test.ts @@ -2,12 +2,12 @@ /** * The diagnostic has to count against the set the machine can actually ask. * - * A pack that ships both tiers replaces the compiled-in semantic set where it - * installs, so its regex policies name its OWN checks in `reviewedBy`. Counted - * against this build's sixteen, every one of those names is "a check this build - * does not have" — so `jev status` would say "0 of 39 enabled policies are - * reviewable" and point at the remedy, on exactly the machines that already took - * it. A diagnostic that lies on the state it was written for is worse than no + * The package ships no Jev checks: the reviewers on a machine are exactly the + * checks its installed packs declare. A pack that ships both tiers names its OWN + * checks in `reviewedBy`; counted against anything else, every one of those + * names is "a check this build does not have" — so `jev status` would say "0 of + * 39 enabled policies are reviewable" and point at the remedy, on exactly the + * machines that already took it. A diagnostic that lies on the state it was written for is worse than no * diagnostic: it sends people to re-take a pack they are already running. */ import { describe, it, expect, beforeEach, afterEach } from "vitest"; @@ -15,7 +15,13 @@ import { createHash } from "node:crypto"; import { mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; -import { countReviewable, reviewableProblem, reviewableSummary, surveyReviewableCoverage } from "@/src/hooks/policy-reviewability"; +import { + NO_JEV_CHECKS_PROBLEM, + countReviewable, + reviewableProblem, + reviewableSummary, + surveyReviewableCoverage, +} from "@/src/hooks/policy-reviewability"; import { SEMANTIC_REVIEWER_NAMES } from "@/src/hooks/policy-authority"; import { forgetEffectiveReviewerNames } from "@/src/hooks/effective-reviewers"; @@ -102,12 +108,12 @@ describe("countReviewable with an explicit reviewer set", () => { }); }); - it("counts the same name as hard against this build's set", () => { + it("counts the same name as hard against FailproofAI's names, and against none", () => { expect(SEMANTIC_REVIEWER_NAMES.has("pack-destructive-deletion")).toBe(false); - expect(countReviewable([{ authority: "reviewable", reviewedBy: ["pack-destructive-deletion"] }])).toEqual({ - enabled: 1, - reviewable: 0, - }); + const decl = [{ authority: "reviewable", reviewedBy: ["pack-destructive-deletion"] }]; + expect(countReviewable(decl, SEMANTIC_REVIEWER_NAMES)).toEqual({ enabled: 1, reviewable: 0 }); + // The default: no pack declaring checks, no reviewers. + expect(countReviewable(decl)).toEqual({ enabled: 1, reviewable: 0 }); }); it("still requires EVERY name, because reviewedBy is a conjunction", () => { @@ -138,18 +144,22 @@ describe("surveyReviewableCoverage on a machine running a two-tier pack", () => expect(reviewableProblem(coverage)).toBeNull(); }); - it("counts a builtin name as hard once a pack has replaced the semantic set", () => { + it("counts a check no installed pack declares as hard", () => { // Not a nicety: that question will never be asked on this machine, so a // clear counted for it is a clear that cannot happen. installPack([regex({ authority: "reviewable", reviewedBy: ["secret-exposure"] })], [SEMANTIC]); const coverage = surveyReviewableCoverage(project); - expect(coverage).toEqual({ enabled: 2, reviewable: 0, customFiles: 0 }); + expect(coverage).toEqual({ enabled: 2, reviewable: 0, customFiles: 0, jevChecks: 1 }); expect(reviewableProblem(coverage)).toContain("it can never clear one"); }); - it("keeps counting against this build's set for a pack with no semantic entries", () => { + it("counts nothing reviewable for a pack with no semantic entries, and says to install the checks", () => { + // No compiled-in set to fall back on: with no pack declaring checks, Jev + // asks nothing, so nothing can be cleared. installPack([regex({ authority: "reviewable", reviewedBy: ["secret-exposure"] })]); - expect(surveyReviewableCoverage(project).reviewable).toBe(1); + const coverage = surveyReviewableCoverage(project); + expect(coverage).toEqual({ enabled: 2, reviewable: 0, customFiles: 0, jevChecks: 0 }); + expect(reviewableProblem(coverage)).toBe(NO_JEV_CHECKS_PROBLEM); }); it("ignores the pack's `enabled` narrowing when collecting reviewers", () => { @@ -166,15 +176,15 @@ describe("surveyReviewableCoverage on a machine running a two-tier pack", () => forgetEffectiveReviewerNames(); const coverage = surveyReviewableCoverage(project); - expect(coverage).toEqual({ enabled: 2, reviewable: 1, customFiles: 0 }); + expect(coverage).toEqual({ enabled: 2, reviewable: 1, customFiles: 0, jevChecks: 1 }); }); }); describe("a check the question budget drops", () => { it("is no reviewer: the policy naming only it counts as hard", () => { - // Two third-party packs, each under the budget alone, over it together - // beside the compiled-in set they join. The later checks are never asked, - // so `jev status` must not call a policy reviewable by one of them. + // Two third-party packs, each under the budget alone, over it together. + // The later checks are never asked, so `jev status` must not call a policy + // reviewable by one of them. const fat = (name: string) => ({ ...SEMANTIC, name, @@ -188,11 +198,15 @@ describe("a check the question budget drops", () => { const manifest = JSON.parse(readFileSync(manifestPath, "utf8")) as { packs: Array> }; Object.assign(manifest.packs[0], { id: "acme/yb", source: "github:acme/yb@1.0.0" }); manifest.packs.unshift({ - ...manifest.packs[0], id: "acme/xa", source: "github:acme/xa@1.0.0", policies: [], semantic: [fat("xa-check-1"), fat("xa-check-2")], + ...manifest.packs[0], id: "acme/xa", source: "github:acme/xa@1.0.0", policies: [], + semantic: [fat("xa-check-1"), fat("xa-check-2"), fat("xa-check-3"), fat("xa-check-4")], }); writeFileSync(manifestPath, JSON.stringify(manifest)); forgetEffectiveReviewerNames(); - expect(surveyReviewableCoverage(project)).toEqual({ enabled: 2, reviewable: 0, customFiles: 0 }); + const coverage = surveyReviewableCoverage(project); + expect(coverage).toMatchObject({ enabled: 2, reviewable: 0, customFiles: 0 }); + // Some checks were asked (the ones that fit), just not the one it names. + expect(coverage.jevChecks).toBeGreaterThan(0); }); }); diff --git a/__tests__/hooks/pack-store-semantic.test.ts b/__tests__/hooks/pack-store-semantic.test.ts index ce176d21b..175b21682 100644 --- a/__tests__/hooks/pack-store-semantic.test.ts +++ b/__tests__/hooks/pack-store-semantic.test.ts @@ -25,7 +25,6 @@ import type { AddressInfo } from "node:net"; import { addPack, fetchPackPreview } from "@/src/hooks/pack-store"; import { MAX_SEMANTIC_POLICIES_PER_PACK, packSemantic, readInstalledPacks } from "@/src/hooks/pack-manifest"; import { semanticPoliciesFromPacks } from "@/src/hooks/semantic/pack-policies"; -import { SEMANTIC_POLICIES } from "@/src/hooks/semantic/policies"; import { forgetEffectiveReviewerNames } from "@/src/hooks/effective-reviewers"; import { version as packageVersion } from "../../package.json"; @@ -134,11 +133,11 @@ describe("installing a pack that declares semantic policies", () => { expect(errors).toEqual([]); expect(warnings).toBeUndefined(); expect(packSemantic(packs[0]).map((s) => s.name)).toEqual(["pack-destructive-deletion"]); - // And the whole point: the machine now asks the PACK's question too — - // beside the compiled-in set, since acme is not a FailproofAI pack. + // And the whole point: the machine now asks the PACK's question — and only + // it, since the package ships no Jev checks of its own. const resolved = semanticPoliciesFromPacks(packs); expect(resolved.fromPack).toBe(true); - expect(resolved.policies.map((p) => p.name)).toEqual([...SEMANTIC_POLICIES.map((p) => p.name), "pack-destructive-deletion"]); + expect(resolved.policies.map((p) => p.name)).toEqual(["pack-destructive-deletion"]); expect(resolved.policies.at(-1)?.precondition).toBeTypeOf("function"); }); @@ -152,15 +151,15 @@ describe("installing a pack that declares semantic policies", () => { it("omits the key when the pack declares none, so it cannot read as an empty set", async () => { await add(); // On DISK: no key at all. An empty array would still read as "this pack - // declares semantic entries", and the replacement rule would then have it - // replace this build's set with nothing. (The READER normalizes absence to - // `[]`, which is why this asserts the record rather than the parsed pack.) + // declares semantic entries" to a careless reader. (The READER normalizes + // absence to `[]`, which is why this asserts the record rather than the + // parsed pack.) const record = JSON.parse(readFileSync(join(root, "installed.json"), "utf8")) as { packs: Array>; }; expect("semantic" in record.packs[0]).toBe(false); - // And this build's own question set stays in play. - expect(semanticPoliciesFromPacks(readInstalledPacks().packs).policies).toBe(SEMANTIC_POLICIES); + // And Jev has nothing to ask: no pack declares a check. + expect(semanticPoliciesFromPacks(readInstalledPacks().packs).policies).toEqual([]); }); it("refuses a malformed semantic entry before writing anything", async () => { diff --git a/__tests__/hooks/policy-attribution.test.ts b/__tests__/hooks/policy-attribution.test.ts index 15556f24a..4ee78c0d1 100644 --- a/__tests__/hooks/policy-attribution.test.ts +++ b/__tests__/hooks/policy-attribution.test.ts @@ -38,6 +38,7 @@ vi.mock("../../src/hooks/pack-manifest", () => ({ // The handler asks this per event to decide whether the migration shim // still applies. Mirrors the mocked readInstalledPacks above. hasInstalledPacks: vi.fn(() => false), + hasInstalledRegexPacks: vi.fn(() => false), })); import { evaluateHookEvent } from "../../src/hooks/handler"; diff --git a/__tests__/hooks/policy-authority-collapse.test.ts b/__tests__/hooks/policy-authority-collapse.test.ts index 2a28da8f5..c2523e249 100644 --- a/__tests__/hooks/policy-authority-collapse.test.ts +++ b/__tests__/hooks/policy-authority-collapse.test.ts @@ -23,6 +23,7 @@ import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; import type { RegisteredPolicy } from "@/src/hooks/policy-types"; +import { installJevPoliciesPack, jevPoliciesPackRecord } from "../fixtures/jev-policies"; const ENV_KEYS = ["FAILPROOFAI_HOME", "FAILPROOFAI_PACK_DIR", "FAILPROOFAI_CLOUD_POLICY_DIR"] as const; @@ -43,6 +44,9 @@ beforeEach(() => { process.env.FAILPROOFAI_PACK_DIR = packRoot; process.env.FAILPROOFAI_CLOUD_POLICY_DIR = cloudRoot; writeFileSync(join(home, "policies-config.json"), JSON.stringify({ enabledPolicies: [] })); + // The checks every `reviewedBy` below names ship in FailproofAI/jev-policies, + // not in the package, so this machine has it installed. + installJevPoliciesPack(packRoot); stderr = []; vi.spyOn(process.stderr, "write").mockImplementation((chunk: string | Uint8Array) => { stderr.push(String(chunk)); @@ -178,9 +182,13 @@ describe("two installed packs sharing one artifact", () => { join(packRoot, "installed.json"), JSON.stringify({ schemaVersion: 1, - packs: packs.map((p) => ({ - ...p, source: `github:${p.id}@v${p.version}`, entry: `artifacts/${digest}.mjs`, sha256: digest, - })), + packs: [ + ...packs.map((p) => ({ + ...p, source: `github:${p.id}@v${p.version}`, entry: `artifacts/${digest}.mjs`, sha256: digest, + })), + // Its own artifact, so it is not one of the packs sharing this one. + jevPoliciesPackRecord(packRoot), + ], }), ); } diff --git a/__tests__/hooks/policy-authority-roundtrip.test.ts b/__tests__/hooks/policy-authority-roundtrip.test.ts index dfbc25b56..e242620af 100644 --- a/__tests__/hooks/policy-authority-roundtrip.test.ts +++ b/__tests__/hooks/policy-authority-roundtrip.test.ts @@ -23,7 +23,8 @@ import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync import { tmpdir } from "node:os"; import { join, resolve } from "node:path"; import { POLICY_CATALOG } from "@/src/hooks/policy-catalog"; -import { resolvePolicyAuthority } from "@/src/hooks/policy-authority"; +import { SEMANTIC_REVIEWER_NAMES, resolvePolicyAuthority } from "@/src/hooks/policy-authority"; +import { installJevPoliciesPack, jevPoliciesPackRecord } from "../fixtures/jev-policies"; import type { RegisteredPolicy } from "@/src/hooks/policy-types"; const REPO = resolve(__dirname, "../.."); @@ -52,6 +53,9 @@ beforeEach(() => { process.env.FAILPROOFAI_CLOUD_POLICY_DIR = cloudRoot; delete process.env.FAILPROOFAI_PACK_BASE_URL; writeConfig({ enabledPolicies: [] }); + // The checks every `reviewedBy` below names ship in FailproofAI/jev-policies, + // not in the package, so this machine has it installed. + installJevPoliciesPack(packRoot); stderr = []; vi.spyOn(process.stderr, "write").mockImplementation((chunk: string | Uint8Array) => { stderr.push(String(chunk)); @@ -147,7 +151,7 @@ describe("catalog → build-policy-pack → policies add → loader → registry it("emits a resolved authority for every policy in the manifest", () => { const expected = POLICY_CATALOG.filter((p) => !p.alwaysOn).map((p) => ({ name: p.name, - ...resolvePolicyAuthority(p), + ...resolvePolicyAuthority(p, SEMANTIC_REVIEWER_NAMES), })); expect( manifest.policies.map((p) => ({ @@ -177,8 +181,8 @@ describe("catalog → build-policy-pack → policies add → loader → registry const byName = new Map(hooks.map((h) => [h.name, h])); for (const p of manifest.policies) { const hook = byName.get(p.name as string)!; - expect(resolvePolicyAuthority(hook), p.name as string).toEqual( - resolvePolicyAuthority(p as never), + expect(resolvePolicyAuthority(hook, SEMANTIC_REVIEWER_NAMES), p.name as string).toEqual( + resolvePolicyAuthority(p as never, SEMANTIC_REVIEWER_NAMES), ); } }); @@ -192,7 +196,7 @@ describe("catalog → build-policy-pack → policies add → loader → registry const { readInstalledPacks } = await import("@/src/hooks/pack-manifest"); const { packs, errors } = readInstalledPacks(); expect(errors).toEqual([]); - const installed = packs[0].policies.find((p) => p.name === "protect-env-vars")!; + const installed = packs.find((p) => p.id === manifest.id)!.policies.find((p) => p.name === "protect-env-vars")!; expect(installed.authority).toBe("reviewable"); expect(installed.reviewedBy).toEqual(["env-secrets-dump", "secret-exposure"]); @@ -202,7 +206,7 @@ describe("catalog → build-policy-pack → policies add → loader → registry if (entry.alwaysOn) continue; const r = registered.get(prefix + entry.name); expect(r, `${entry.name} was not registered`).toBeDefined(); - const { downgraded: _d, ...expected } = resolvePolicyAuthority(entry); + const { downgraded: _d, ...expected } = resolvePolicyAuthority(entry, SEMANTIC_REVIEWER_NAMES); expect(authorityOf(r), entry.name).toEqual(expected); } // The guard packs may not carry still ships compiled in, and stays hard. @@ -246,7 +250,7 @@ describe("a third-party pack: its manifest decides, and only for its own policie join(packRoot, "installed.json"), JSON.stringify({ schemaVersion: 1, - packs: [{ + packs: [jevPoliciesPackRecord(packRoot), { id: "acme/ops", version: "1.0.0", source: "github:acme/ops@v1.0.0", entry: `artifacts/${digest}.mjs`, sha256: digest, policies: [ @@ -279,7 +283,7 @@ describe("a third-party pack: its manifest decides, and only for its own policie const { readInstalledPacks } = await import("@/src/hooks/pack-manifest"); const { packs, errors } = readInstalledPacks(); expect(errors).toEqual([]); - const garbled = packs[0].policies.find((p) => p.name === "garbled")!; + const garbled = packs.find((p) => p.id === "acme/ops")!.policies.find((p) => p.name === "garbled")!; expect("authority" in garbled).toBe(false); expect("reviewedBy" in garbled).toBe(false); diff --git a/__tests__/hooks/policy-authority-table.test.ts b/__tests__/hooks/policy-authority-table.test.ts index 16c71ac8f..d9120b1ea 100644 --- a/__tests__/hooks/policy-authority-table.test.ts +++ b/__tests__/hooks/policy-authority-table.test.ts @@ -10,14 +10,15 @@ * nothing checking it is the #337 drift class. */ import { describe, it, expect } from "vitest"; -import { readFileSync } from "node:fs"; -import { resolve } from "node:path"; +import { mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join, resolve } from "node:path"; import { BUILTIN_POLICIES, registerBuiltinPolicies } from "../../src/hooks/builtin-policies"; import { POLICY_CATALOG } from "../../src/hooks/policy-catalog"; import { clearPolicies, getAllPolicies } from "../../src/hooks/policy-registry"; import { effectiveAuthority } from "../../src/hooks/policy-types"; import { SEMANTIC_REVIEWER_NAMES, resolvePolicyAuthority } from "../../src/hooks/policy-authority"; -import { SEMANTIC_POLICIES } from "../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES, withInstalledJevPoliciesPack } from "../fixtures/jev-policies"; /** D1: the only builtins Jev may clear, and the checks that must clear them. */ const REVIEWABLE: Record = { @@ -135,7 +136,35 @@ describe("the builtin authority table (D1)", () => { }); }); +describe("with no pack declaring Jev checks, every builtin registers hard", () => { + it("registers all of them hard, with no reviewedBy — nothing can clear", () => { + // The vanilla install: the package ships no Jev checks, so there are no + // reviewers, whatever the catalog marks. + const saved = process.env.FAILPROOFAI_PACK_DIR; + const empty = mkdtempSync(join(tmpdir(), "fpai-no-jev-pack-")); + process.env.FAILPROOFAI_PACK_DIR = empty; + clearPolicies(); + try { + registerBuiltinPolicies(BUILTIN_POLICIES.map((p) => p.name)); + const all = getAllPolicies(); + expect(all.length).toBe(POLICY_CATALOG.length); + for (const r of all) { + expect(effectiveAuthority(r), r.name).toBe("hard"); + expect("reviewedBy" in r, r.name).toBe(false); + } + } finally { + clearPolicies(); + if (saved === undefined) delete process.env.FAILPROOFAI_PACK_DIR; + else process.env.FAILPROOFAI_PACK_DIR = saved; + rmSync(empty, { recursive: true, force: true }); + } + }); +}); + describe("builtin registration carries the table into the registry", () => { + // A machine with FailproofAI/jev-policies installed: the names the table uses are reviewers. + withInstalledJevPoliciesPack(); + it("registers every builtin with its resolved authority", () => { clearPolicies(); try { diff --git a/__tests__/hooks/policy-authority.test.ts b/__tests__/hooks/policy-authority.test.ts index ffa102231..52fcfc9ca 100644 --- a/__tests__/hooks/policy-authority.test.ts +++ b/__tests__/hooks/policy-authority.test.ts @@ -18,18 +18,23 @@ import { authorityFieldsOf, authorityProblem, manifestAuthority, + NO_REVIEWERS, type AuthorityFields, resolvePolicyAuthority, warnAuthority, withMergedAuthority, } from "../../src/hooks/policy-authority"; -import { SEMANTIC_POLICIES, INJECTION_PROBE, SCOPE_PROBE, TASK_PROBES } from "../../src/hooks/semantic/policies"; +import { INJECTION_PROBE, SCOPE_PROBE, TASK_PROBES } from "../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES, withInstalledJevPoliciesPack } from "../fixtures/jev-policies"; import { clearPolicies, getAllPolicies, registerPolicy } from "../../src/hooks/policy-registry"; import { parsePackPolicy } from "../../src/hooks/pack-manifest"; import type { PolicyCatalogEntry } from "../../src/hooks/policy-types"; const allow = () => ({ decision: "allow" as const }); +/** The reviewer set of a machine with FailproofAI/jev-policies installed. */ +const JEV = SEMANTIC_REVIEWER_NAMES; + describe("effectiveAuthority — shape only", () => { it.each([ ["absent", {}], @@ -101,12 +106,25 @@ describe("SEMANTIC_REVIEWER_NAMES", () => { }); describe("resolvePolicyAuthority", () => { + it("is hard, silently, with no reviewers — a machine with no pack declaring Jev checks", () => { + // The package ships no Jev checks, so this is the vanilla install. Not a + // declaration going wrong: no per-policy warning, just hard. + const decl = { authority: "reviewable", reviewedBy: ["destructive-deletion"] }; + expect(resolvePolicyAuthority(decl)).toEqual({ authority: "hard" }); + expect(resolvePolicyAuthority(decl, NO_REVIEWERS)).toEqual({ authority: "hard" }); + // Malformed is still said: that is the author's mistake whatever is installed. + expect(resolvePolicyAuthority({ authority: "reviewable", reviewedBy: [7] }).downgraded).toMatch(/not a list/); + }); + it("resolves a complete declaration to reviewable, deduplicated in declared order", () => { expect( - resolvePolicyAuthority({ - authority: "reviewable", - reviewedBy: ["secret-exposure", "env-secrets-dump", "secret-exposure"], - }), + resolvePolicyAuthority( + { + authority: "reviewable", + reviewedBy: ["secret-exposure", "env-secrets-dump", "secret-exposure"], + }, + JEV, + ), ).toEqual({ authority: "reviewable", reviewedBy: ["secret-exposure", "env-secrets-dump"] }); }); @@ -126,10 +144,13 @@ describe("resolvePolicyAuthority", () => { it("refuses a name that is not a semantic policy in this build — the whole declaration, not the name", () => { // reviewedBy is a conjunction. Dropping the unknown name would let Jev clear // the policy on fewer checks than its author asked for. - const r = resolvePolicyAuthority({ - authority: "reviewable", - reviewedBy: ["secret-exposure", "secret-exposure-v2"], - }); + const r = resolvePolicyAuthority( + { + authority: "reviewable", + reviewedBy: ["secret-exposure", "secret-exposure-v2"], + }, + JEV, + ); expect(r.authority).toBe("hard"); expect(r.reviewedBy).toBeUndefined(); expect(r.downgraded).toMatch(/"secret-exposure-v2"/); @@ -154,7 +175,7 @@ describe("resolvePolicyAuthority", () => { it("escapes the name it quotes: a pack's reviewedBy reaches the hook's stderr", () => { const name = "x\u001b[2J\nWARN forged"; - for (const known of [undefined, new Set(["own-check"])]) { + for (const known of [JEV, new Set(["own-check"])]) { const r = resolvePolicyAuthority({ authority: "reviewable", reviewedBy: [name] }, known); expect(r.downgraded).toContain(JSON.stringify(name)); expect(r.downgraded).not.toMatch(/[\u0000-\u001f]/); @@ -200,13 +221,15 @@ describe("resolvePolicyAuthority", () => { it("does not hand back the caller's own array", () => { const names = ["secret-exposure"]; - const r = resolvePolicyAuthority({ authority: "reviewable", reviewedBy: names }); + const r = resolvePolicyAuthority({ authority: "reviewable", reviewedBy: names }, JEV); names.push("read-outside-workspace"); expect(r.reviewedBy).toEqual(["secret-exposure"]); }); }); describe("registerPolicy stores the RESOLVED authority", () => { + // A machine with FailproofAI/jev-policies installed, so its names are reviewers. + withInstalledJevPoliciesPack(); beforeEach(() => clearPolicies()); const only = () => { @@ -388,7 +411,7 @@ describe("withMergedAuthority — several declarations, one registration", () => it("is reviewable only when every declaration is, through the union of their checks", () => { const a: Rec = { id: "a", ...R("database-destruction") }; - const { merged, overruled } = withMergedAuthority(a, [a, R("secret-exposure", "database-destruction")]); + const { merged, overruled } = withMergedAuthority(a, [a, R("secret-exposure", "database-destruction")], JEV); expect(merged).toEqual({ id: "a", authority: "reviewable", reviewedBy: ["database-destruction", "secret-exposure"] }); expect(overruled).toBe(false); }); @@ -407,8 +430,8 @@ describe("withMergedAuthority — several declarations, one registration", () => [otherRecord, [otherRecord, reviewable]], ]; for (const [record, decls] of cases) { - const { merged } = withMergedAuthority(record, decls); - expect(resolvePolicyAuthority(merged).authority).toBe("hard"); + const { merged } = withMergedAuthority(record, decls, JEV); + expect(resolvePolicyAuthority(merged, JEV).authority).toBe("hard"); expect(merged.id).toBe(record.id); } }); @@ -421,12 +444,13 @@ describe("withMergedAuthority — several declarations, one registration", () => // the pack's. The merged record now carries the resolution — `hard`, and no // `reviewedBy` for anything to re-read — and the reason comes back beside it. const typo = R("databse-destruction"); - const { merged, overruled, refused } = withMergedAuthority({ id: "x", ...R("database-destruction") }, [ - R("database-destruction"), - typo, - ]); + const { merged, overruled, refused } = withMergedAuthority( + { id: "x", ...R("database-destruction") }, + [R("database-destruction"), typo], + JEV, + ); expect(merged).toEqual({ id: "x", authority: "hard" }); - expect(resolvePolicyAuthority(merged).authority).toBe("hard"); + expect(resolvePolicyAuthority(merged, JEV).authority).toBe("hard"); expect(refused).toMatch(/"databse-destruction"/); expect(overruled).toBe(true); }); @@ -454,10 +478,11 @@ describe("withMergedAuthority — several declarations, one registration", () => }); it("drops the fields entirely when the hard vote declared nothing", () => { - const { merged, overruled } = withMergedAuthority({ id: "x", ...R("secret-exposure") }, [ - R("secret-exposure"), - {}, - ]); + const { merged, overruled } = withMergedAuthority( + { id: "x", ...R("secret-exposure") }, + [R("secret-exposure"), {}], + JEV, + ); expect(merged).toEqual({ id: "x" }); expect(overruled).toBe(true); }); diff --git a/__tests__/hooks/policy-reviewability.test.ts b/__tests__/hooks/policy-reviewability.test.ts index f7fbb00bd..aab378f31 100644 --- a/__tests__/hooks/policy-reviewability.test.ts +++ b/__tests__/hooks/policy-reviewability.test.ts @@ -12,8 +12,13 @@ * thing. Before this module nothing on any surface said so. * * So: a pack-shaped policy set with no authority fields must report zero-of-N - * with the remedy, this build's builtins must report the fifteen Jev may clear, - * and neither may change what any policy is allowed to do. + * with the remedy, this build's builtins must report the fifteen Jev may clear + * where `FailproofAI/jev-policies` supplies the checks they name, and neither + * may change what any policy is allowed to do. + * + * And the one problem that comes first: no installed pack gives Jev any checks + * (the package ships none). Then nothing is reviewable, and the remedy is to + * install them. */ import { describe, it, expect, beforeEach, afterEach } from "vitest"; import { createHash } from "node:crypto"; @@ -24,6 +29,8 @@ import { parsePackPolicy } from "@/src/hooks/pack-manifest"; import { SEMANTIC_REVIEWER_NAMES, resolvePolicyAuthority } from "@/src/hooks/policy-authority"; import { POLICY_CATALOG } from "@/src/hooks/policy-catalog"; import { + JEV_POLICIES_ADD_COMMAND, + NO_JEV_CHECKS_PROBLEM, RETAKE_PACK_COMMAND, countReviewable, reviewableProblem, @@ -31,6 +38,10 @@ import { surveyReviewableCoverage, } from "@/src/hooks/policy-reviewability"; import { effectiveAuthority } from "@/src/hooks/policy-types"; +import { JEV_PACK_SEMANTIC_ENTRIES } from "../fixtures/jev-policies"; + +/** The reviewer set of a machine with FailproofAI/jev-policies installed. */ +const JEV = SEMANTIC_REVIEWER_NAMES; /** Every builtin a pack may carry: `alwaysOn` is refused in a pack manifest. */ const PACKABLE = POLICY_CATALOG.filter((p) => !p.alwaysOn); @@ -54,8 +65,8 @@ describe("counting what Jev may clear", () => { // The premise: the parser kept no authority field, from any of them. expect(entries.every((e) => !("authority" in e) && !("reviewedBy" in e))).toBe(true); - const coverage = { ...countReviewable(entries), customFiles: 0 }; - expect(coverage).toEqual({ enabled: PACKABLE.length, reviewable: 0, customFiles: 0 }); + const coverage = { ...countReviewable(entries, JEV), customFiles: 0, jevChecks: 16 }; + expect(coverage).toEqual({ enabled: PACKABLE.length, reviewable: 0, customFiles: 0, jevChecks: 16 }); expect(reviewableSummary(coverage)).toBe(`0 of ${PACKABLE.length} enabled policies are reviewable.`); const problem = reviewableProblem(coverage); @@ -65,7 +76,7 @@ describe("counting what Jev may clear", () => { }); it("reports the fifteen reviewable builtins, and diagnoses nothing", () => { - const coverage = { ...countReviewable(POLICY_CATALOG), customFiles: 0 }; + const coverage = { ...countReviewable(POLICY_CATALOG, JEV), customFiles: 0, jevChecks: 16 }; expect(coverage.reviewable).toBe(15); expect(REVIEWABLE_BUILTINS.map((p) => p.name)).toEqual([ // Catalog order. The nine after `block-read-outside-cwd` arrived with the pack @@ -104,7 +115,7 @@ describe("counting what Jev may clear", () => { { authority: "reviewable", reviewedBy: "secret-exposure" }, { authority: "reviewable" }, ]; - expect(countReviewable(wishful)).toEqual({ enabled: 4, reviewable: 0 }); + expect(countReviewable(wishful, JEV)).toEqual({ enabled: 4, reviewable: 0 }); }); it("counts a reviewer this build does not have as hard, and says so", () => { @@ -125,29 +136,40 @@ describe("counting what Jev may clear", () => { ]; expect(SEMANTIC_REVIEWER_NAMES.has("future-check")).toBe(false); expect(fromANewerPack.map((p) => effectiveAuthority(p))).toEqual(["reviewable", "reviewable"]); - expect(fromANewerPack.map((p) => resolvePolicyAuthority(p).authority)).toEqual(["hard", "hard"]); + expect(fromANewerPack.map((p) => resolvePolicyAuthority(p, JEV).authority)).toEqual(["hard", "hard"]); - const coverage = { ...countReviewable(fromANewerPack), customFiles: 0 }; - expect(coverage).toEqual({ enabled: 2, reviewable: 0, customFiles: 0 }); + const coverage = { ...countReviewable(fromANewerPack, JEV), customFiles: 0, jevChecks: 16 }; + expect(coverage).toEqual({ enabled: 2, reviewable: 0, customFiles: 0, jevChecks: 16 }); expect(reviewableSummary(coverage)).toBe("0 of 2 enabled policies are reviewable."); expect(reviewableProblem(coverage)).toContain(RETAKE_PACK_COMMAND); }); it("says there is nothing to clear, rather than blaming a pack, for an empty set", () => { - const empty = { enabled: 0, reviewable: 0, customFiles: 0 }; + const empty = { enabled: 0, reviewable: 0, customFiles: 0, jevChecks: 16 }; expect(reviewableSummary(empty)).toBe("No policies are enabled here, so there is nothing for Jev to clear."); expect(reviewableProblem(empty)).toBeNull(); }); it("claims no 'never' while policies from the user's own files went uncounted", () => { - expect(reviewableProblem({ enabled: 2, reviewable: 0, customFiles: 1 })).toBeNull(); + expect(reviewableProblem({ enabled: 2, reviewable: 0, customFiles: 1, jevChecks: 16 })).toBeNull(); + }); + + it("with no Jev checks installed, says so first and names the one command", () => { + // The package ships none, so every surface has to say where they come from. + const coverage = { ...countReviewable(POLICY_CATALOG), customFiles: 0, jevChecks: 0 }; + expect(coverage.reviewable).toBe(0); + expect(reviewableProblem(coverage)).toBe(NO_JEV_CHECKS_PROBLEM); + expect(NO_JEV_CHECKS_PROBLEM).toContain(JEV_POLICIES_ADD_COMMAND); + expect(JEV_POLICIES_ADD_COMMAND).toBe("failproofai policies add FailproofAI/jev-policies"); + // Whatever else is true of the machine. + expect(reviewableProblem({ enabled: 0, reviewable: 0, customFiles: 3, jevChecks: 0 })).toBe(NO_JEV_CHECKS_PROBLEM); }); it("admits the policies it did not read", () => { - expect(reviewableSummary({ enabled: 4, reviewable: 0, customFiles: 2 })).toBe( + expect(reviewableSummary({ enabled: 4, reviewable: 0, customFiles: 2, jevChecks: 16 })).toBe( "0 of 4 enabled policies are reviewable (policies from your own files are not counted).", ); - expect(reviewableSummary({ enabled: 1, reviewable: 1, customFiles: 1 })).toBe( + expect(reviewableSummary({ enabled: 1, reviewable: 1, customFiles: 1, jevChecks: 16 })).toBe( "1 of 1 enabled policy is reviewable: Jev may clear a deny or an instruction from those, " + "and from no others (policies from your own files are not counted).", ); @@ -187,8 +209,12 @@ describe("surveying a real machine", () => { writeFileSync(join(home, "policies-config.json"), JSON.stringify(config)); } - /** An installed pack, through the real manifest the loader verifies. */ - function installPack(policies: Array>, enabled?: string[]): void { + /** + * An installed pack, through the real manifest the loader verifies — beside + * FailproofAI/jev-policies unless `withJev` is false, because the checks the + * core pack's `reviewedBy` names live there. + */ + function installPack(policies: Array> | null, enabled?: string[], withJev = true): void { const artifact = "// a pack artifact this test never executes\n"; const digest = createHash("sha256").update(artifact).digest("hex"); mkdirSync(join(packRoot, "artifacts"), { recursive: true }); @@ -198,15 +224,32 @@ describe("surveying a real machine", () => { JSON.stringify({ schemaVersion: 1, packs: [ - { - id: "FailproofAI/policies", - version: "0.9.0", - source: "github:FailproofAI/policies@v0.9.0", - entry: `artifacts/${digest}.mjs`, - sha256: digest, - policies, - ...(enabled ? { enabled } : {}), - }, + ...(policies + ? [ + { + id: "FailproofAI/policies", + version: "0.9.0", + source: "github:FailproofAI/policies@v0.9.0", + entry: `artifacts/${digest}.mjs`, + sha256: digest, + policies, + ...(enabled ? { enabled } : {}), + }, + ] + : []), + ...(withJev + ? [ + { + id: "FailproofAI/jev-policies", + version: "0.2.0", + source: "github:FailproofAI/jev-policies@v0.2.0", + entry: `artifacts/${digest}.mjs`, + sha256: digest, + policies: [], + semantic: JEV_PACK_SEMANTIC_ENTRIES, + }, + ] + : []), ], }), ); @@ -222,20 +265,41 @@ describe("surveying a real machine", () => { const coverage = surveyReviewableCoverage(project); // The pack's policies, plus the one guard that ships compiled in and // registers whatever else is enabled. - expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 0, customFiles: 0 }); + expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 0, customFiles: 0, jevChecks: 16 }); expect(reviewableSummary(coverage)).toContain(`0 of ${PACKABLE.length + 1} enabled policies are reviewable`); expect(reviewableProblem(coverage)).toContain(RETAKE_PACK_COMMAND); }); + it("a Jev-only pack does not hide builtins still enforced from enabledPolicies", () => { + // An upgraded machine with no core pack adds FailproofAI/jev-policies: the + // hook path keeps enforcing its enabledPolicies (hasInstalledRegexPacks), so + // the survey must count them too, reviewable ones included. + writeConfig({ enabledPolicies: ["protect-env-vars", "block-env-files"] }); + installPack(null); + const coverage = surveyReviewableCoverage(project); + expect(coverage.enabled).toBeGreaterThanOrEqual(3); // the two + the always-on guard + expect(coverage.reviewable).toBe(2); + expect(coverage.jevChecks).toBeGreaterThan(0); + }); + it("a pack built by this release: the fifteen it marks, and no complaint", () => { writeConfig({ enabledPolicies: [] }); installPack(PACKABLE as unknown as Array>); const coverage = surveyReviewableCoverage(project); - expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 15, customFiles: 0 }); + expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 15, customFiles: 0, jevChecks: 16 }); expect(reviewableProblem(coverage)).toBeNull(); }); + it("the same pack WITHOUT FailproofAI/jev-policies: nothing reviewable, and the command that fixes it", () => { + writeConfig({ enabledPolicies: [] }); + installPack(PACKABLE as unknown as Array>, undefined, false); + + const coverage = surveyReviewableCoverage(project); + expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 0, customFiles: 0, jevChecks: 0 }); + expect(reviewableProblem(coverage)).toBe(NO_JEV_CHECKS_PROBLEM); + }); + it("a pack built against a NEWER semantic set: hard here, with the remedy", () => { // Version skew in the other direction, end to end through the real // manifest parser (which keeps the names verbatim — whether a name is a @@ -250,7 +314,7 @@ describe("surveying a real machine", () => { ); const coverage = surveyReviewableCoverage(project); - expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 0, customFiles: 0 }); + expect(coverage).toEqual({ enabled: PACKABLE.length + 1, reviewable: 0, customFiles: 0, jevChecks: 16 }); expect(reviewableProblem(coverage)).toContain(RETAKE_PACK_COMMAND); }); @@ -260,15 +324,17 @@ describe("surveying a real machine", () => { const coverage = surveyReviewableCoverage(project); // Two selected + the always-on guard; one of the two is reviewable. - expect(coverage).toEqual({ enabled: 3, reviewable: 1, customFiles: 0 }); + expect(coverage).toEqual({ enabled: 3, reviewable: 1, customFiles: 0, jevChecks: 16 }); }); - it("falls back to this build's builtins while no pack is installed", () => { + it("counts this build's builtins while no pack is installed — none reviewable, with no Jev checks", () => { + // The migration shim still enforces them, but no pack means no Jev checks, + // so nothing they name can be asked. writeConfig({ enabledPolicies: REVIEWABLE_BUILTINS.map((p) => p.name).concat("block-sudo") }); const coverage = surveyReviewableCoverage(project); - expect(coverage).toEqual({ enabled: 17, reviewable: 15, customFiles: 0 }); - expect(reviewableProblem(coverage)).toBeNull(); + expect(coverage).toEqual({ enabled: 17, reviewable: 0, customFiles: 0, jevChecks: 0 }); + expect(reviewableProblem(coverage)).toBe(NO_JEV_CHECKS_PROBLEM); }); it("reads a cloud assignment's authority, which only the deployment decides", () => { @@ -287,10 +353,12 @@ describe("surveying a real machine", () => { ], }), ); + // The check the deployment names lives in FailproofAI/jev-policies. + installPack(null); const coverage = surveyReviewableCoverage(project); // Two assignments + the always-on guard, one of them reviewable. - expect(coverage).toEqual({ enabled: 3, reviewable: 1, customFiles: 0 }); + expect(coverage).toEqual({ enabled: 3, reviewable: 1, customFiles: 0, jevChecks: 16 }); }); it("counts the custom policy files it cannot read without running them", () => { @@ -306,6 +374,7 @@ describe("surveying a real machine", () => { const projectFile = join(project, ".failproofai", "policies", "a-policies.mjs"); mkdirSync(join(project, ".failproofai", "policies"), { recursive: true }); writeFileSync(projectFile, "export {};\n"); + installPack(null); const coverage = surveyReviewableCoverage(project); expect(coverage.customFiles).toBe(1); expect(reviewableProblem(coverage)).toBeNull(); @@ -331,6 +400,6 @@ describe("surveying a real machine", () => { writeFileSync(join(packRoot, "installed.json"), "{ not json"); writeFileSync(join(cloudRoot, "active.json"), "{ not json"); // The guard that ships compiled in is all that is left, and it is hard. - expect(surveyReviewableCoverage(project)).toEqual({ enabled: 1, reviewable: 0, customFiles: 0 }); + expect(surveyReviewableCoverage(project)).toEqual({ enabled: 1, reviewable: 0, customFiles: 0, jevChecks: 0 }); }); }); diff --git a/__tests__/hooks/safe-config-write.test.ts b/__tests__/hooks/safe-config-write.test.ts new file mode 100644 index 000000000..ad1a1a280 --- /dev/null +++ b/__tests__/hooks/safe-config-write.test.ts @@ -0,0 +1,168 @@ +// @vitest-environment node +/** + * Writing an agent's config must never leave it half-written. + * + * These files belong to the user's agent (Hermes, Claude Code, OpenClaw, …); + * a truncated one is an agent that will not start. So every case here is about + * what the live file looks like when something goes wrong part-way. + */ +import { describe, it, expect, beforeEach, afterEach } from "vitest"; +import { + mkdtempSync, + mkdirSync, + readFileSync, + readdirSync, + rmSync, + statSync, + symlinkSync, + lstatSync, + readlinkSync, + existsSync, + writeFileSync, + chmodSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { writeConfigFileAtomic, configBackupPath, DanglingConfigSymlinkError } from "../../src/hooks/safe-config-write"; +import { claudeCode, hermes, UnreadableAgentConfigError } from "../../src/hooks/integrations"; + +let dir: string; +beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), "fpai-safe-write-")); +}); +afterEach(() => rmSync(dir, { recursive: true, force: true })); + +const leftovers = () => readdirSync(dir).filter((f) => f.includes(".failproofai-") && f.endsWith(".tmp")); + +describe("writeConfigFileAtomic", () => { + it("writes a new file and leaves no temporary file or backup behind", () => { + const p = join(dir, "config.yaml"); + writeConfigFileAtomic(p, "a: 1\n"); + expect(readFileSync(p, "utf8")).toBe("a: 1\n"); + expect(leftovers()).toEqual([]); + expect(() => statSync(configBackupPath(p))).toThrow(); + }); + + it("keeps the previous version as .failproofai-backup", () => { + const p = join(dir, "config.yaml"); + writeFileSync(p, "model: gpt-6-luna\n"); + writeConfigFileAtomic(p, "model: gpt-6-luna\nplugins: {}\n"); + expect(readFileSync(configBackupPath(p), "utf8")).toBe("model: gpt-6-luna\n"); + expect(readFileSync(p, "utf8")).toBe("model: gpt-6-luna\nplugins: {}\n"); + }); + + it("keeps the file's permission bits (Hermes keeps config.yaml at 0600)", () => { + const p = join(dir, "config.yaml"); + writeFileSync(p, "a: 1\n"); + chmodSync(p, 0o600); + writeConfigFileAtomic(p, "a: 2\n"); + expect(statSync(p).mode & 0o777).toBe(0o600); + expect(statSync(configBackupPath(p)).mode & 0o777).toBe(0o600); + }); + + it("an interruption before the swap leaves the live file exactly as it was", () => { + const p = join(dir, "config.yaml"); + writeFileSync(p, "model: gpt-6-luna\napi_key: ${KEY}\n"); + expect(() => + writeConfigFileAtomic(p, "plugins: {}\n", { + rename: () => { + throw new Error("killed mid-write"); + }, + }), + ).toThrow("killed mid-write"); + expect(readFileSync(p, "utf8")).toBe("model: gpt-6-luna\napi_key: ${KEY}\n"); + expect(leftovers()).toEqual([]); + }); + + it("a retry after an interruption succeeds, and a stale temp file from a killed process does not get in the way", () => { + const p = join(dir, "config.yaml"); + writeFileSync(p, "a: 1\n"); + writeFileSync(join(dir, ".config.yaml.failproofai-99999-deadbeef.tmp"), "half"); + expect(() => writeConfigFileAtomic(p, "a: 2\n", { rename: () => { throw new Error("boom"); } })).toThrow(); + writeConfigFileAtomic(p, "a: 3\n"); + expect(readFileSync(p, "utf8")).toBe("a: 3\n"); + }); + + it("writes through a symlinked config to its target and keeps the link", () => { + const real = join(dir, "dotfiles"); + mkdirSync(real); + writeFileSync(join(real, "config.yaml"), "a: 1\n"); + const link = join(dir, "config.yaml"); + symlinkSync(join(real, "config.yaml"), link); + writeConfigFileAtomic(link, "a: 2\n"); + expect(lstatSync(link).isSymbolicLink()).toBe(true); + expect(readFileSync(join(real, "config.yaml"), "utf8")).toBe("a: 2\n"); + }); +}); + +describe("links planted around a config are never followed into other files", () => { + it("F8: a symlink sitting at .failproofai-backup is replaced, and the file it points at is untouched", () => { + // A cloned repository controls /.claude/: it can pre-create the backup + // path as a link to any file the developer can write. + const victim = join(dir, "victim.txt"); + writeFileSync(victim, "unchanged"); + const p = join(dir, "settings.json"); + writeFileSync(p, '{"old":true}\n'); + symlinkSync(victim, `${p}.failproofai-backup`); + writeConfigFileAtomic(p, '{"new":true}\n'); + expect(readFileSync(victim, "utf8")).toBe("unchanged"); + expect(lstatSync(`${p}.failproofai-backup`).isSymbolicLink()).toBe(false); + expect(readFileSync(`${p}.failproofai-backup`, "utf8")).toBe('{"old":true}\n'); + expect(readFileSync(p, "utf8")).toBe('{"new":true}\n'); + expect(leftovers()).toEqual([]); + }); + + it("F10: a dangling config symlink is refused — the link is kept and nothing is created at its target", () => { + const missing = join(dir, "dotfiles", "config.yaml"); // not checked out + const link = join(dir, "config.yaml"); + symlinkSync(missing, link); + expect(() => writeConfigFileAtomic(link, "a: 1\n")).toThrow(DanglingConfigSymlinkError); + expect(lstatSync(link).isSymbolicLink()).toBe(true); + expect(readlinkSync(link)).toBe(missing); + expect(existsSync(missing)).toBe(false); + expect(leftovers()).toEqual([]); + }); + + it("symlinks: replace swaps a planted link for the file and never writes into its target", () => { + const victim = join(dir, "victim.sh"); + writeFileSync(victim, "echo safe\n"); + const shim = join(dir, "failproofai.mjs"); + symlinkSync(victim, shim); + writeConfigFileAtomic(shim, "export default {};\n", { symlinks: "replace" }); + expect(readFileSync(victim, "utf8")).toBe("echo safe\n"); + expect(lstatSync(shim).isSymbolicLink()).toBe(false); + expect(readFileSync(shim, "utf8")).toBe("export default {};\n"); + expect(existsSync(`${shim}.failproofai-backup`)).toBe(false); + }); + + it("backup: false writes no backup (used for failproofai's own records)", () => { + const p = join(dir, "record"); + writeFileSync(p, "a\n"); + writeConfigFileAtomic(p, "b\n", { backup: false }); + expect(readFileSync(p, "utf8")).toBe("b\n"); + expect(existsSync(`${p}.failproofai-backup`)).toBe(false); + }); +}); + +describe("a config that does not parse is refused, never overwritten", () => { + it("Hermes: an unparseable config.yaml throws and is left byte-for-byte", () => { + const p = join(dir, "config.yaml"); + const broken = "model:\n default: gpt-6-luna\n bad: [unclosed\n"; + writeFileSync(p, broken); + expect(() => hermes.readSettings(p)).toThrow(UnreadableAgentConfigError); + expect(readFileSync(p, "utf8")).toBe(broken); + }); + + it("Claude Code: invalid settings.json throws and is left byte-for-byte", () => { + const p = join(dir, "settings.json"); + writeFileSync(p, '{ "model": "x", '); + expect(() => claudeCode.readSettings(p)).toThrow(UnreadableAgentConfigError); + expect(readFileSync(p, "utf8")).toBe('{ "model": "x", '); + }); + + it("an empty file still reads as empty (nothing to lose)", () => { + const p = join(dir, "settings.json"); + writeFileSync(p, " \n"); + expect(claudeCode.readSettings(p)).toEqual({}); + }); +}); diff --git a/__tests__/hooks/semantic/combine-shadow-verdict.test.ts b/__tests__/hooks/semantic/combine-observe-verdict.test.ts similarity index 61% rename from __tests__/hooks/semantic/combine-shadow-verdict.test.ts rename to __tests__/hooks/semantic/combine-observe-verdict.test.ts index 63012747d..d33943d2a 100644 --- a/__tests__/hooks/semantic/combine-shadow-verdict.test.ts +++ b/__tests__/hooks/semantic/combine-observe-verdict.test.ts @@ -1,12 +1,12 @@ /** - * Shadow mode's "would have": Jev's own deny or instruct, recorded rather than + * Observe mode's "would have": Jev's own deny or instruct, recorded rather than * applied (contract §5B). * - * `combineTwoTier` returns it as `shadowVerdict`, and the handler files it in + * `combineTwoTier` returns it as `observeVerdict`, and the handler files it in * the row's `observed` list. What is pinned here is that it is the verdict * ENFORCE mode would have applied — same name, same decision, same reason — - * and that it appears in shadow mode only, for deny/instruct only, without - * changing anything shadow mode enforces. + * and that it appears in observe mode only, for deny/instruct only, without + * changing anything observe mode enforces. */ import { describe, expect, it } from "vitest"; import { combineTwoTier, regexOnly, type JevReview, type RegexVerdict } from "../../../src/hooks/semantic/combine"; @@ -31,64 +31,64 @@ function answered(over: Partial> = {}): }; } -describe("shadowVerdict", () => { - it.each(["deny", "instruct"] as const)("a Jev %s in shadow mode is recorded exactly as enforce mode would apply it", (decision) => { +describe("observeVerdict", () => { + it.each(["deny", "instruct"] as const)("a Jev %s in observe mode is recorded exactly as enforce mode would apply it", (decision) => { const review = answered({ decision }); - const shadow = combineTwoTier(allowAll, review, "shadow"); + const observe = combineTwoTier(allowAll, review, "observe"); const enforce = combineTwoTier(allowAll, review, "enforce"); - // Shadow enforces the regex result, unchanged. - expect(shadow.final).toEqual(regexOnly(allowAll)); - expect(shadow.decidedByJev).toBe(false); + // Observe enforces the regex result, unchanged. + expect(observe.final).toEqual(regexOnly(allowAll)); + expect(observe.decidedByJev).toBe(false); - // Enforce applied Jev's verdict; shadow records the same one. + // Enforce applied Jev's verdict; observe records the same one. expect(enforce.decidedByJev).toBe(true); expect(enforce.final.decision).toBe(decision); - expect(shadow.shadowVerdict).toEqual({ + expect(observe.observeVerdict).toEqual({ policyName: enforce.final.entries[0].policyName, decision, reason: enforce.final.entries[0].reason, version: "jev-1.13.0", }); - expect(enforce.shadowVerdict).toBeUndefined(); + expect(enforce.observeVerdict).toBeUndefined(); }); it("uses enforce mode's fixed template when Jev gave no reason", () => { - const shadow = combineTwoTier(allowAll, answered({ reason: null }), "shadow"); - expect(shadow.shadowVerdict?.reason).toBe("Flagged by semantic review (semantic/destructive-deletion)"); + const observe = combineTwoTier(allowAll, answered({ reason: null }), "observe"); + expect(observe.observeVerdict?.reason).toBe("Flagged by semantic review (semantic/destructive-deletion)"); }); it("records nothing when Jev allowed", () => { - expect(combineTwoTier(allowAll, answered({ decision: "allow" }), "shadow").shadowVerdict).toBeUndefined(); + expect(combineTwoTier(allowAll, answered({ decision: "allow" }), "observe").observeVerdict).toBeUndefined(); }); it("records nothing when Jev did not answer or was not consulted", () => { - expect(combineTwoTier(allowAll, { kind: "fallback", reason: "http-503", latencyMs: 40, model: null }, "shadow").shadowVerdict).toBeUndefined(); - expect(combineTwoTier(allowAll, { kind: "not-consulted" }, "shadow").shadowVerdict).toBeUndefined(); + expect(combineTwoTier(allowAll, { kind: "fallback", reason: "http-503", latencyMs: 40, model: null }, "observe").observeVerdict).toBeUndefined(); + expect(combineTwoTier(allowAll, { kind: "not-consulted" }, "observe").observeVerdict).toBeUndefined(); }); it("keeps Jev's own verdict on a cut or injected call, as enforce mode does (upward only)", () => { for (const over of [{ requestCut: true, truncated: true }, { injected: true }]) { const review = answered(over); expect(combineTwoTier(allowAll, review, "enforce").final.decision).toBe("deny"); - expect(combineTwoTier(allowAll, review, "shadow").shadowVerdict?.decision).toBe("deny"); + expect(combineTwoTier(allowAll, review, "observe").observeVerdict?.decision).toBe("deny"); } }); it("records it beside a regex deny too: it is Jev's verdict, not the row's", () => { const regexDeny: RegexVerdict[] = [{ policyName: "failproofai/block-sudo", decision: "deny", reason: "no sudo", authority: "hard", reviewedBy: [] }]; - const shadow = combineTwoTier(regexDeny, answered(), "shadow"); - expect(shadow.final.decision).toBe("deny"); - expect(shadow.final.entries[0].policyName).toBe("failproofai/block-sudo"); - expect(shadow.shadowVerdict?.policyName).toBe("semantic/destructive-deletion"); + const observe = combineTwoTier(regexDeny, answered(), "observe"); + expect(observe.final.decision).toBe("deny"); + expect(observe.final.entries[0].policyName).toBe("failproofai/block-sudo"); + expect(observe.observeVerdict?.policyName).toBe("semantic/destructive-deletion"); }); it("files the version as the model id, or `jev` when there is none this build would store", () => { - expect(combineTwoTier(allowAll, answered({ model: null }), "shadow").shadowVerdict?.version).toBe("jev"); - expect(combineTwoTier(allowAll, answered({ model: "typesafe/jev-1.13-20260917" }), "shadow").shadowVerdict?.version).toBe( + expect(combineTwoTier(allowAll, answered({ model: null }), "observe").observeVerdict?.version).toBe("jev"); + expect(combineTwoTier(allowAll, answered({ model: "typesafe/jev-1.13-20260917" }), "observe").observeVerdict?.version).toBe( "typesafe/jev-1.13-20260917", ); // A reported id carrying a space or a newline is not a model id. - expect(combineTwoTier(allowAll, answered({ model: "jev 1.13\nrm -rf" }), "shadow").shadowVerdict?.version).toBe("jev"); + expect(combineTwoTier(allowAll, answered({ model: "jev 1.13\nrm -rf" }), "observe").observeVerdict?.version).toBe("jev"); }); }); diff --git a/__tests__/hooks/semantic/combine.test.ts b/__tests__/hooks/semantic/combine.test.ts index eb8f009d2..fa5492ce1 100644 --- a/__tests__/hooks/semantic/combine.test.ts +++ b/__tests__/hooks/semantic/combine.test.ts @@ -1,12 +1,12 @@ /** - * The §4 combine table, exhaustively: every row × {shadow, enforce} × + * The §4 combine table, exhaustively: every row × {observe, enforce} × * {whole, request-cut}. Each row is driven from a SemanticOutcome — what * `evaluateSemantic` actually returns — through `toReview` (how the handler * reads it) and `combineTwoTier` (what it enforces), so the cut → fallback * step is covered by the same table rather than beside it. * * The expected result is written out for `enforce` + whole. The other columns - * follow from rules the table asserts on every row: shadow enforces the regex + * follow from rules the table asserts on every row: observe enforces the regex * result, and a call part of which was never shown to Jev — the tool input, a * computed fact, a redacted span — withdraws every clear, so every regex deny * counts (§4) and the call is recorded `jev-fallback` / `request-cut`. A cut @@ -30,7 +30,7 @@ import { } from "../../../src/hooks/semantic/combine"; import { fallbackCode, toReview } from "../../../src/hooks/semantic/jev-review"; import { DEFAULT_THRESHOLDS_V1, decideV1 } from "../../../src/hooks/semantic/decide"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { SemanticOutcome } from "../../../src/hooks/semantic/evaluator"; import type { PolicyOutcome, SemanticVerdict } from "../../../src/hooks/semantic/types"; @@ -380,7 +380,7 @@ const ROWS: Row[] = [ }, ]; -const MODES: JevMode[] = ["enforce", "shadow"]; +const MODES: JevMode[] = ["enforce", "observe"]; const CUTS: Cut[] = ["whole", "request-cut"]; /** allow < instruct < deny, for the "never more permissive" invariant. */ const SEVERITY: Record<"allow" | "instruct" | "deny", number> = { allow: 0, instruct: 1, deny: 2 }; @@ -390,7 +390,7 @@ function reviewFor(row: Row, cut: Cut): JevReview { return toReview(withCut(row.outcome, cut)); } -describe("combine table (§4) — every row × shadow/enforce × whole/request-cut", () => { +describe("combine table (§4) — every row × observe/enforce × whole/request-cut", () => { for (const row of ROWS) { for (const mode of MODES) { for (const cut of CUTS) { @@ -412,8 +412,8 @@ describe("combine table (§4) — every row × shadow/enforce × whole/request-c const wholeAnswer = !hardDecided && !degraded && !cutAnswer; // What is ENFORCED. - if (mode === "shadow" || hardDecided || degraded) { - // shadow, a degraded Jev and a hard deny all enforce exactly what + if (mode === "observe" || hardDecided || degraded) { + // observe, a degraded Jev and a hard deny all enforce exactly what // the regex engine says alone. expect(out.final).toEqual(legacy); expect(out.decidedByJev).toBe(false); @@ -464,7 +464,7 @@ describe("combine table (§4) — every row × shadow/enforce × whole/request-c expect(out.activity.evaluator).toBe("jev"); expect(out.activity.jevFallbackReason).toBeUndefined(); expect(out.activity.jevDecision).toBe(row.outcome!.status === "ok" ? row.outcome!.verdict.decision : undefined); - // Shadow records what enforce WOULD have cleared. + // Observe records what enforce WOULD have cleared. expect(out.cleared).toEqual(row.enforce.cleared); const recorded = row.enforce.recorded ?? row.enforce.cleared; expect(out.activity.jevCleared).toEqual(recorded.length > 0 ? recorded : undefined); @@ -721,9 +721,9 @@ describe("the clear rule, on hand-built reviews", () => { expect(out.decidedByJev).toBe(true); }); - it("shadow still enforces the regex result for a cut answer", () => { + it("observe still enforces the regex result for a cut answer", () => { const review = answered({ requestCut: true, truncated: true, decision: "deny", reason: "deletes the database", policyName: "semantic/destructive-deletion" }); - const out = combineTwoTier([], review, "shadow"); + const out = combineTwoTier([], review, "observe"); expect(out.final).toEqual(regexOnly([])); expect(out.decidedByJev).toBe(false); }); @@ -783,8 +783,8 @@ describe("the clear rule, on hand-built reviews", () => { expect(out.final.decision).toBe("instruct"); }); - it("shadow mode is unchanged, and still records the reason", () => { - const out = combineTwoTier([], answered({ requestCut: true, truncated: true }), "shadow"); + it("observe mode is unchanged, and still records the reason", () => { + const out = combineTwoTier([], answered({ requestCut: true, truncated: true }), "observe"); expect(out.final).toEqual(regexOnly([])); expect(out.activity.jevFallbackReason).toBe("request-cut"); }); diff --git a/__tests__/hooks/semantic/decide.test.ts b/__tests__/hooks/semantic/decide.test.ts index 06fd6930e..3b773dc8a 100644 --- a/__tests__/hooks/semantic/decide.test.ts +++ b/__tests__/hooks/semantic/decide.test.ts @@ -1,7 +1,7 @@ // @vitest-environment node import { describe, it, expect } from "vitest"; import { decide, decideV1, everyTargetNamed, scanTargets, targetNamedByUser, targetTokens, DEFAULT_THRESHOLDS } from "../../../src/hooks/semantic/decide"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { SemanticPolicy } from "../../../src/hooks/semantic/types"; const byName = (name: string): SemanticPolicy => SEMANTIC_POLICIES.find((p) => p.name === name)!; diff --git a/__tests__/hooks/semantic/envelope-budget.test.ts b/__tests__/hooks/semantic/envelope-budget.test.ts index 8a8ee1881..2f488a804 100644 --- a/__tests__/hooks/semantic/envelope-budget.test.ts +++ b/__tests__/hooks/semantic/envelope-budget.test.ts @@ -53,11 +53,16 @@ import { import { computeFacts, scanCommand } from "../../../src/hooks/semantic/facts"; import { evaluateSemantic, prepareSemantic, verdictLogRow, type SemanticOptions } from "../../../src/hooks/semantic/evaluator"; import { toReview } from "../../../src/hooks/semantic/jev-review"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { Facts, JevRequest, JevResponse, SemanticInput } from "../../../src/hooks/semantic/types"; // The PEM armour, joined at runtime — see `redaction-fixtures.ts` and this // file's "Fixtures are assembled at runtime and never written as literals". import { pemBegin, pemEnd } from "./redaction-fixtures"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); const DANGEROUS = "rm -rf / --no-preserve-root"; diff --git a/__tests__/hooks/semantic/envelope-compile.test.ts b/__tests__/hooks/semantic/envelope-compile.test.ts index f866a4973..af164c3b0 100644 --- a/__tests__/hooks/semantic/envelope-compile.test.ts +++ b/__tests__/hooks/semantic/envelope-compile.test.ts @@ -3,7 +3,7 @@ import { describe, it, expect } from "vitest"; import { buildEnvelope, capHeadTail, redactSecrets, MAX_STRING_CHARS } from "../../../src/hooks/semantic/envelope"; import { compileRequest, selectPolicies, DEFAULT_JEV_MODEL } from "../../../src/hooks/semantic/compile"; import { computeFacts, scanCommand } from "../../../src/hooks/semantic/facts"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { Facts } from "../../../src/hooks/semantic/types"; const facts = (over: Partial = {}): Facts => ({ diff --git a/__tests__/hooks/semantic/evaluator-context-cut.test.ts b/__tests__/hooks/semantic/evaluator-context-cut.test.ts index 6fa75d66c..15d053adb 100644 --- a/__tests__/hooks/semantic/evaluator-context-cut.test.ts +++ b/__tests__/hooks/semantic/evaluator-context-cut.test.ts @@ -43,6 +43,11 @@ import { MAX_USER_MESSAGE_CHARS, capHeadTail } from "../../../src/hooks/semantic import { evaluateSemantic, prepareSemantic, type SemanticOptions } from "../../../src/hooks/semantic/evaluator"; import { toReview } from "../../../src/hooks/semantic/jev-review"; import type { JevRequest, JevResponse, SemanticInput } from "../../../src/hooks/semantic/types"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); const mark = (omitted: number) => `\n…[${omitted} characters omitted]…\n`; diff --git a/__tests__/hooks/semantic/evaluator-no-transport.test.ts b/__tests__/hooks/semantic/evaluator-no-transport.test.ts index 060c2ecf5..0af2d3f00 100644 --- a/__tests__/hooks/semantic/evaluator-no-transport.test.ts +++ b/__tests__/hooks/semantic/evaluator-no-transport.test.ts @@ -16,6 +16,11 @@ import { join } from "node:path"; import { randomBytes } from "node:crypto"; import { evaluateSemantic } from "../../../src/hooks/semantic/evaluator"; import type { SemanticInput } from "../../../src/hooks/semantic/types"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); const ENV = ["HOME", "FAILPROOFAI_HOME", "FAILPROOFAI_JEV_CONFIG_DIR", "TYPESAFE_API_KEY"]; const saved: Record = {}; diff --git a/__tests__/hooks/semantic/evaluator-sent-evidence.test.ts b/__tests__/hooks/semantic/evaluator-sent-evidence.test.ts index a64104436..3f64c6207 100644 --- a/__tests__/hooks/semantic/evaluator-sent-evidence.test.ts +++ b/__tests__/hooks/semantic/evaluator-sent-evidence.test.ts @@ -35,7 +35,7 @@ import { DEFAULT_THRESHOLDS, DEFAULT_THRESHOLDS_V1 } from "../../../src/hooks/se import { MAX_USER_MESSAGES, MAX_USER_MESSAGE_CHARS } from "../../../src/hooks/semantic/envelope"; import { evaluateSemantic, prepareSemantic, type SemanticOptions } from "../../../src/hooks/semantic/evaluator"; import { toReview } from "../../../src/hooks/semantic/jev-review"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { JevRequest, JevResponse, SemanticInput, SemanticPolicy } from "../../../src/hooks/semantic/types"; const POLICY: SemanticPolicy = SEMANTIC_POLICIES.find((p) => p.name === "database-destruction")!; diff --git a/__tests__/hooks/semantic/jev-client-hardening.test.ts b/__tests__/hooks/semantic/jev-client-hardening.test.ts index 17ee6d195..57e7d3443 100644 --- a/__tests__/hooks/semantic/jev-client-hardening.test.ts +++ b/__tests__/hooks/semantic/jev-client-hardening.test.ts @@ -4,7 +4,7 @@ // error message by any route (Cloudflare's 200 {success:false} envelope, a job // state, a reported model id, a network error), a configured Cloudflare model // reaches the wire, a custom endpoint must say which Jev answered, plain-http -// loopback is shadow-only, and the 64 KiB config cap holds. +// loopback is observe-only, and the 64 KiB config cap holds. import { describe, it, expect, beforeEach, afterEach, vi } from "vitest"; import { chmodSync, mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; @@ -207,7 +207,7 @@ describe("config hardening", () => { chmodSync(jevConfigPath(), 0o600); }; - describe("plain http to loopback is accepted in shadow mode only", () => { + describe("plain http to loopback is accepted in observe mode only", () => { const problem = (obj: unknown) => { const r = validateJevConfig(obj); return r.ok ? null : r.problem; @@ -215,11 +215,11 @@ describe("config hardening", () => { it.each(["http://localhost:8787/v1", "http://127.0.0.1:8787", "http://[::1]:8787"])("%s", (baseUrl) => { // enforce, explicit or by default: refused, and the reason says what to do. - expect(problem({ provider: "custom", apiKey: KEY, baseUrl })).toMatch(/shadow/); + expect(problem({ provider: "custom", apiKey: KEY, baseUrl })).toMatch(/observe/); expect(problem({ provider: "custom", apiKey: KEY, baseUrl, mode: "enforce" })).toMatch(/https/); - expect(problem({ provider: "typesafe", apiKey: KEY, baseUrl })).toMatch(/shadow/); - // shadow: accepted; a forged answer there changes no decision. - expect(problem({ provider: "custom", apiKey: KEY, baseUrl, mode: "shadow" })).toBeNull(); + expect(problem({ provider: "typesafe", apiKey: KEY, baseUrl })).toMatch(/observe/); + // observe: accepted; a forged answer there changes no decision. + expect(problem({ provider: "custom", apiKey: KEY, baseUrl, mode: "observe" })).toBeNull(); }); it("https to loopback is fine in enforce mode: the agent cannot present a trusted certificate", () => { @@ -231,8 +231,8 @@ describe("config hardening", () => { expect(loadJevConfig()).toBeNull(); const r = inspectJevConfig(); expect(r.status === "refused" && r.reason).toBe("invalid"); - write(JSON.stringify({ provider: "custom", apiKey: KEY, baseUrl: "http://localhost:8787/v1", mode: "shadow" })); - expect(loadJevConfig()).toMatchObject({ provider: "custom", mode: "shadow", baseUrl: "http://localhost:8787/v1" }); + write(JSON.stringify({ provider: "custom", apiKey: KEY, baseUrl: "http://localhost:8787/v1", mode: "observe" })); + expect(loadJevConfig()).toMatchObject({ provider: "custom", mode: "observe", baseUrl: "http://localhost:8787/v1" }); }); }); diff --git a/__tests__/hooks/semantic/jev-client-redirect.test.ts b/__tests__/hooks/semantic/jev-client-redirect.test.ts index d94c03280..b0bc60b76 100644 --- a/__tests__/hooks/semantic/jev-client-redirect.test.ts +++ b/__tests__/hooks/semantic/jev-client-redirect.test.ts @@ -64,7 +64,7 @@ describe("the Jev client never follows a redirect", () => { b.close(); }); - const custom = (): JevConfig => ({ provider: "custom", apiKey: KEY, baseUrl: `http://localhost:${aPort}/v1`, mode: "shadow" }); + const custom = (): JevConfig => ({ provider: "custom", apiKey: KEY, baseUrl: `http://localhost:${aPort}/v1`, mode: "observe" }); it.each([307, 308, 302, 301, 303])("HTTP %s from the configured endpoint is an error, and the target is never contacted", async (s) => { status = s; @@ -76,7 +76,7 @@ describe("the Jev client never follows a redirect", () => { it("the endpoint that does answer directly still works", async () => { hitsB.length = 0; - const cfg: JevConfig = { provider: "custom", apiKey: KEY, baseUrl: `http://127.0.0.1:${bPort}/v1`, mode: "shadow" }; + const cfg: JevConfig = { provider: "custom", apiKey: KEY, baseUrl: `http://127.0.0.1:${bPort}/v1`, mode: "observe" }; const res = await transportForConfig(cfg).transport(request, AbortSignal.timeout(5_000)); expect(res.answers.a.noul).toBe(0.99); expect(hitsB).toHaveLength(1); diff --git a/__tests__/hooks/semantic/jev-cloud-config.test.ts b/__tests__/hooks/semantic/jev-cloud-config.test.ts index 156f5b215..d3c3886cb 100644 --- a/__tests__/hooks/semantic/jev-cloud-config.test.ts +++ b/__tests__/hooks/semantic/jev-cloud-config.test.ts @@ -89,7 +89,7 @@ describe("jev-config: the FailproofAI Cloud provider", () => { if (!current.ingest) writeCredentials({ ...current, ingest: { url: `${url}/v1/events`, key } }); return writeJevCloudCredential({ url, key }); }; - const cloudFile = (over: Record = {}) => ({ provider: "failproofai", baseUrl: BASE, mode: "shadow", ...over }); + const cloudFile = (over: Record = {}) => ({ provider: "failproofai", baseUrl: BASE, mode: "observe", ...over }); describe("the credentials.json slot", () => { it("round-trips at 0600, beside every other credential, and clears alone", () => { @@ -250,7 +250,7 @@ describe("jev-config: the FailproofAI Cloud provider", () => { const r = inspectJevConfig(); expect(r.status).toBe("key-lacks-jev"); if (r.status !== "key-lacks-jev") return; - expect(r.routing).toEqual({ provider: "failproofai", baseUrl: BASE, mode: "shadow", timeoutMs: 3000 }); + expect(r.routing).toEqual({ provider: "failproofai", baseUrl: BASE, mode: "observe", timeoutMs: 3000 }); expect(r.problem).toContain("no Jev key is stored"); expect(r.problem).not.toMatch(/not connected/); expect(JSON.stringify(r)).not.toContain(KEY); @@ -285,7 +285,7 @@ describe("jev-config: the FailproofAI Cloud provider", () => { expect(r.status).toBe("ok"); if (r.status !== "ok") return; expect(r.keySource).toBe("cloud"); - expect(r.config).toMatchObject({ provider: "failproofai", apiKey: KEY, baseUrl: BASE, mode: "shadow", timeoutMs: 3000 }); + expect(r.config).toMatchObject({ provider: "failproofai", apiKey: KEY, baseUrl: BASE, mode: "observe", timeoutMs: 3000 }); expect(loadJevConfig()?.apiKey).toBe(KEY); // The route the transport will POST to. expect(jevRoute(r.config).endpoint).toBe(`${BASE}/systemone`); @@ -363,7 +363,7 @@ describe("jev-config: the FailproofAI Cloud provider", () => { it("needs a baseUrl", () => { connect(); - writeJev({ provider: "failproofai", mode: "shadow" }); + writeJev({ provider: "failproofai", mode: "observe" }); expect(inspectJevConfig()).toMatchObject({ status: "refused", reason: "invalid" }); }); @@ -391,7 +391,7 @@ describe("jev-config: the FailproofAI Cloud provider", () => { const r = inspectJevConfig(); expect(r.status).toBe("not-connected"); if (r.status !== "not-connected") return; - expect(r.routing).toEqual({ provider: "failproofai", baseUrl: BASE, mode: "shadow", timeoutMs: 3000 }); + expect(r.routing).toEqual({ provider: "failproofai", baseUrl: BASE, mode: "observe", timeoutMs: 3000 }); expect(r.problem).toMatch(/not connected to FailproofAI Cloud/); expect(r.problem).toMatch(/config --token/); expect(loadJevConfig()).toBeNull(); @@ -431,7 +431,7 @@ describe("jev-config: the FailproofAI Cloud provider", () => { it("keeps plain http to localhost out of enforce mode, as for every provider", () => { const local = "http://localhost:8080"; connect(local); - writeJev(cloudFile({ baseUrl: `${local}/enforcement/v1/jev`, mode: "shadow" })); + writeJev(cloudFile({ baseUrl: `${local}/enforcement/v1/jev`, mode: "observe" })); expect(loadJevConfig()?.baseUrl).toBe(`${local}/enforcement/v1/jev`); writeJev(cloudFile({ baseUrl: `${local}/enforcement/v1/jev`, mode: "enforce" })); expect(inspectJevConfig()).toMatchObject({ status: "refused", reason: "invalid" }); @@ -477,8 +477,8 @@ describe("jev-config: the FailproofAI Cloud provider", () => { expect(inspectJevConfig()).toMatchObject({ status: "refused", reason: "invalid" }); }); - it("validates as a mode, alongside shadow and enforce, and nothing else", () => { - for (const mode of ["off", "shadow", "enforce"]) { + it("validates as a mode, alongside observe and enforce, and nothing else", () => { + for (const mode of ["off", "observe", "enforce"]) { expect(validateJevConfig({ provider: "typesafe", apiKey: KEY, mode }).ok).toBe(true); } for (const mode of ["disabled", "OFF", "", 0, false, null]) { diff --git a/__tests__/hooks/semantic/jev-cloud-transport.test.ts b/__tests__/hooks/semantic/jev-cloud-transport.test.ts index 183949adf..888e030c3 100644 --- a/__tests__/hooks/semantic/jev-cloud-transport.test.ts +++ b/__tests__/hooks/semantic/jev-cloud-transport.test.ts @@ -40,6 +40,11 @@ import { startJevReview } from "../../../src/hooks/semantic/jev-review"; import { resetJevThrottle } from "../../../src/hooks/semantic/jev-throttle"; import { normalizeJevFallbackReason } from "../../../src/hooks/jev-activity"; import type { JevRequest } from "../../../src/hooks/semantic/types"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); // Built at runtime: this repo's own hooks refuse secret-shaped literals. const KEY = ["fp", "machine", "c10ud0123456789ab"].join("-"); @@ -119,12 +124,12 @@ describe("the FailproofAI Cloud route, over a real socket", () => { rmSync(home, { recursive: true, force: true }); }); - /** The config the loader produces for a connected machine (plain http to loopback: shadow only). */ + /** The config the loader produces for a connected machine (plain http to loopback: observe only). */ const cloud = (): JevConfig => ({ provider: "failproofai", apiKey: KEY, baseUrl: `http://127.0.0.1:${port}/enforcement/v1/jev`, - mode: "shadow", + mode: "observe", timeoutMs: 3000, // The loader records the origin of the credential it validated against. credentialOrigin: `http://127.0.0.1:${port}`, @@ -341,7 +346,7 @@ describe("the FailproofAI Cloud route, over a real socket", () => { }); it("a BYOK route's 429 is exactly what it was: no cool-down", async () => { - const byok: JevConfig = { provider: "custom", apiKey: KEY, baseUrl: `http://127.0.0.1:${port}/v1`, mode: "shadow", timeoutMs: 3000 }; + const byok: JevConfig = { provider: "custom", apiKey: KEY, baseUrl: `http://127.0.0.1:${port}/v1`, mode: "observe", timeoutMs: 3000 }; const sendByok = () => transportForConfig(byok).transport(request, AbortSignal.timeout(5_000)); reply = () => rateLimited("30"); expect((await failure(sendByok())).code).toBe("http-429"); diff --git a/__tests__/hooks/semantic/jev-config.test.ts b/__tests__/hooks/semantic/jev-config.test.ts index 5ae123a48..49b3657b5 100644 --- a/__tests__/hooks/semantic/jev-config.test.ts +++ b/__tests__/hooks/semantic/jev-config.test.ts @@ -13,6 +13,7 @@ import { jevConfigPath, jevModelVersion, loadJevConfig, + parseJevMode, validateBaseUrl, validateJevConfig, } from "../../../src/hooks/semantic/jev-config"; @@ -88,7 +89,7 @@ describe("semantic/jev-config", () => { }); it.each([ - ["typesafe", { provider: "typesafe", apiKey: KEY, model: "jev-1.13.0", mode: "shadow", timeoutMs: 900 }], + ["typesafe", { provider: "typesafe", apiKey: KEY, model: "jev-1.13.0", mode: "observe", timeoutMs: 900 }], ["openrouter", { provider: "openrouter", apiKey: KEY, model: "typesafe/jev-1.13" }], ["vercel", { provider: "vercel", apiKey: KEY }], ["cloudflare", { provider: "cloudflare", apiKey: KEY, accountId: ACCOUNT }], @@ -211,8 +212,13 @@ describe("semantic/jev-config", () => { expect(problem({ provider: "typesafe", apiKey: KEY, mode: "disabled" })).toMatch(/mode/); // `off` is a mode now: it keeps the file and runs no Jev (see jev-cloud-config.test.ts). expect(problem({ provider: "typesafe", apiKey: KEY, mode: "off" })).toBeNull(); - const r = validateJevConfig({ provider: "typesafe", apiKey: KEY, mode: "shadow", timeoutMs: 800 }); - expect(r.ok && [r.value.mode, r.value.timeoutMs]).toEqual(["shadow", 800]); + const r = validateJevConfig({ provider: "typesafe", apiKey: KEY, mode: "observe", timeoutMs: 800 }); + expect(r.ok && [r.value.mode, r.value.timeoutMs]).toEqual(["observe", 800]); + }); + + it("parses exactly the three modes", () => { + for (const m of ["off", "observe", "enforce"]) expect(parseJevMode(m)).toBe(m); + for (const m of ["Observe", "observing", "", null, undefined, 1]) expect(parseJevMode(m)).toBeNull(); }); it("refuses a model naming a Jev family the thresholds were not calibrated for", () => { diff --git a/__tests__/hooks/semantic/jev-providers.test.ts b/__tests__/hooks/semantic/jev-providers.test.ts index 20f919c94..d17db05d1 100644 --- a/__tests__/hooks/semantic/jev-providers.test.ts +++ b/__tests__/hooks/semantic/jev-providers.test.ts @@ -22,6 +22,11 @@ import { import { evaluateSemantic } from "../../../src/hooks/semantic/evaluator"; import { JEV_REASON_PROVIDER_REFUSED, normalizeJevFallbackReason } from "../../../src/hooks/jev-activity"; import type { JevRequest, JevResponse } from "../../../src/hooks/semantic/types"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); const KEY = ["prov", "test", "abcdef0123456789"].join("-"); const ACCOUNT = "0123456789abcdef0123456789abcdef"; diff --git a/__tests__/hooks/semantic/jev-review.test.ts b/__tests__/hooks/semantic/jev-review.test.ts index 1879c7829..11781b409 100644 --- a/__tests__/hooks/semantic/jev-review.test.ts +++ b/__tests__/hooks/semantic/jev-review.test.ts @@ -93,6 +93,11 @@ const PAST_THE_CALL_BUDGET = "x".repeat(MAX_AGENT_REQUEST_CHARS + 1_000); /** Repeats needed to run past the per-message cap, whatever it is set to. */ const OVER_CAP = Math.ceil((MAX_USER_MESSAGE_CHARS * 1.5) / "tidy the build folder and ".length); import { JEV_CONFIG_DEFAULT_TIMEOUT_MS, type JevConfig } from "../../../src/hooks/semantic/jev-config"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); const CFG: JevConfig = { provider: "cloudflare", apiKey: "not-a-real-key", accountId: "0".repeat(32) }; const allLow = (request: JevRequest): JevResponse => ({ @@ -385,12 +390,12 @@ describe("the local verdict log", () => { it("records what the handler did with each outcome", async () => { await startJevReview(CFG, bash("ls")).review; - await startJevReview({ ...CFG, mode: "shadow" }, bash("ls")).review; + await startJevReview({ ...CFG, mode: "observe" }, bash("ls")).review; respond = async () => { throw new JevError("http-429", "slow down"); }; await startJevReview(CFG, bash("ls")).review; - expect(rows().map((r) => r.applied)).toEqual(["two-tier", "shadow", "legacy-fallback"]); + expect(rows().map((r) => r.applied)).toEqual(["two-tier", "observe", "legacy-fallback"]); expect(rows()[2]).toMatchObject({ status: "degraded", reason: "http-429" }); }); @@ -406,7 +411,7 @@ describe("the local verdict log", () => { describe("mode", () => { it("defaults to enforce (D2)", () => { expect(resolveMode(CFG)).toBe("enforce"); - expect(resolveMode({ ...CFG, mode: "shadow" })).toBe("shadow"); + expect(resolveMode({ ...CFG, mode: "observe" })).toBe("observe"); expect(resolveMode({ ...CFG, mode: "loud" as never })).toBe("enforce"); }); }); @@ -570,7 +575,7 @@ describe("how long a call may wait for Jev", () => { // ── Round-2 review findings ────────────────────────────────────────────────── describe("the throttle's cache is scoped to where answers come from", () => { - const LOOPBACK_SHADOW: JevConfig = { provider: "custom", apiKey: "not-a-real-key", baseUrl: "http://127.0.0.1:9", mode: "shadow" }; + const LOOPBACK_OBSERVE: JevConfig = { provider: "custom", apiKey: "not-a-real-key", baseUrl: "http://127.0.0.1:9", mode: "observe" }; const TYPESAFE_ENFORCE: JevConfig = { provider: "typesafe", apiKey: "not-a-real-key", baseUrl: "https://jev.invalid", mode: "enforce" }; it("passes a scope naming the provider, endpoint, account and model", async () => { @@ -580,8 +585,8 @@ describe("the throttle's cache is scoped to where answers come from", () => { { ...CFG, model: "typesafe/jev" }, { provider: "typesafe", apiKey: "not-a-real-key" }, TYPESAFE_ENFORCE, - LOOPBACK_SHADOW, - { ...LOOPBACK_SHADOW, baseUrl: "http://127.0.0.1:10" }, + LOOPBACK_OBSERVE, + { ...LOOPBACK_OBSERVE, baseUrl: "http://127.0.0.1:10" }, { provider: "openrouter", apiKey: "not-a-real-key" }, ]; for (const cfg of configs) await startJevReview(cfg, bash("ls")).review; @@ -591,14 +596,14 @@ describe("the throttle's cache is scoped to where answers come from", () => { expect(new Set(scopes).size).toBe(configs.length); // The same config always gets the same scope (the cache still works), // whatever its key or mode — neither changes who answers. - expect(throttleScope({ ...CFG, apiKey: "another" , mode: "shadow" }, { via: "cloudflare", model: "jev-1.13.0" })).toBe(scopes[0]); + expect(throttleScope({ ...CFG, apiKey: "another" , mode: "observe" }, { via: "cloudflare", model: "jev-1.13.0" })).toBe(scopes[0]); }); it("an answer cached under one provider is never served under another", async () => { fakeCache.on = true; intent = { userSaid: ["show me my notes"], agentLastMessage: null }; // Both routes ask for the same model, so the requests are byte-identical. - const first = await startJevReview(LOOPBACK_SHADOW, bash("cat ~/other/notes.txt")).review; + const first = await startJevReview(LOOPBACK_OBSERVE, bash("cat ~/other/notes.txt")).review; expect(first).toMatchObject({ kind: "answered" }); expect(transportCalls).toHaveLength(1); diff --git a/__tests__/hooks/semantic/jev-stats-hardening.test.ts b/__tests__/hooks/semantic/jev-stats-hardening.test.ts index 847503f92..1264231f3 100644 --- a/__tests__/hooks/semantic/jev-stats-hardening.test.ts +++ b/__tests__/hooks/semantic/jev-stats-hardening.test.ts @@ -43,14 +43,14 @@ describe("computeJevStats: names that collide with Object.prototype", () => { [ answered({ jevCleared: PROTO_NAMES }), answered({ jevCleared: PROTO_NAMES }), - answered({ jevMode: "shadow", jevCleared: PROTO_NAMES }), + answered({ jevMode: "observe", jevCleared: PROTO_NAMES }), answered({ jevModel: "constructor" }), ], { now: NOW }, ); for (const name of PROTO_NAMES) { expect(Object.getOwnPropertyDescriptor(s.clearsByPolicy, name)?.value, name).toBe(2); - expect(Object.getOwnPropertyDescriptor(s.shadowClearsByPolicy, name)?.value, name).toBe(1); + expect(Object.getOwnPropertyDescriptor(s.observeClearsByPolicy, name)?.value, name).toBe(1); } expect(Object.entries(s.clearsByPolicy).sort()).toEqual(PROTO_NAMES.map((n) => [n, 2]).sort()); expect(Object.getOwnPropertyDescriptor(s.models, "constructor")?.value).toBe(1); diff --git a/__tests__/hooks/semantic/jev-stats.test.ts b/__tests__/hooks/semantic/jev-stats.test.ts index 116ed6ee5..a9ccacba6 100644 --- a/__tests__/hooks/semantic/jev-stats.test.ts +++ b/__tests__/hooks/semantic/jev-stats.test.ts @@ -91,19 +91,19 @@ describe("computeJevStats", () => { expect(s.latencyP95Ms).toBe(190); }); - it("counts clears per policy, separating shadow-mode would-be clears", () => { + it("counts clears per policy, separating observe-mode would-be clears", () => { const s = computeJevStats( [ answered(30, { jevCleared: ["block-read-outside-cwd"] }), answered(30, { jevCleared: ["block-read-outside-cwd", "protect-env-vars"] }), answered(30, { jevCleared: [] }), - answered(30, { jevMode: "shadow", jevCleared: ["block-env-files"] }), + answered(30, { jevMode: "observe", jevCleared: ["block-env-files"] }), ], { now: NOW }, ); expect(s.clearsByPolicy).toEqual({ "block-read-outside-cwd": 2, "protect-env-vars": 1 }); - expect(s.shadowClearsByPolicy).toEqual({ "block-env-files": 1 }); - expect(s.modes).toEqual({ shadow: 1, enforce: 3 }); + expect(s.observeClearsByPolicy).toEqual({ "block-env-files": 1 }); + expect(s.modes).toEqual({ observe: 1, enforce: 3 }); }); it("tallies Jev's own verdicts and the models that answered", () => { @@ -211,7 +211,7 @@ describe("formatJevStats", () => { answered(30, { jevCleared: ["block-read-outside-cwd"] }), answered(50, { jevDecision: "deny" }), fellBack("timeout"), - answered(40, { jevMode: "shadow", jevCleared: ["block-env-files"] }), + answered(40, { jevMode: "observe", jevCleared: ["block-env-files"] }), ], { now: NOW, windowMs: 6 * HOUR }, ); @@ -222,8 +222,8 @@ describe("formatJevStats", () => { " Fell back: 1 (25.0%) — timeout 1", " Latency: p50 40 ms, p95 50 ms", " Cleared: block-read-outside-cwd 1", - " Would clear: block-env-files 1 (shadow mode)", - " Modes: enforce 3, shadow 1", + " Would clear: block-env-files 1 (observe mode)", + " Modes: enforce 3, observe 1", ].join("\n"), ); }); diff --git a/__tests__/hooks/semantic/jev-throttle.test.ts b/__tests__/hooks/semantic/jev-throttle.test.ts index a10391aa3..a5cada881 100644 --- a/__tests__/hooks/semantic/jev-throttle.test.ts +++ b/__tests__/hooks/semantic/jev-throttle.test.ts @@ -13,6 +13,11 @@ import { } from "../../../src/hooks/semantic/jev-throttle"; import { evaluateSemantic } from "../../../src/hooks/semantic/evaluator"; import type { JevRequest, JevResponse } from "../../../src/hooks/semantic/types"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); // ── Fixtures ───────────────────────────────────────────────────────────────── diff --git a/__tests__/hooks/semantic/pack-preconditions.test.ts b/__tests__/hooks/semantic/pack-preconditions.test.ts index 4dc2cd8b3..7bcaddb46 100644 --- a/__tests__/hooks/semantic/pack-preconditions.test.ts +++ b/__tests__/hooks/semantic/pack-preconditions.test.ts @@ -19,7 +19,7 @@ import { preconditionFor, } from "../../../src/hooks/semantic/preconditions"; import { PACK_PRECONDITION_NAMES, isPackPreconditionName } from "../../../src/hooks/semantic/precondition-names"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { Facts, PathFact, SemanticPolicy } from "../../../src/hooks/semantic/types"; const PROJECT = "/home/dev/project"; diff --git a/__tests__/hooks/semantic/pack-semantic-registry.test.ts b/__tests__/hooks/semantic/pack-semantic-registry.test.ts index 92ee55edb..f0095cdf5 100644 --- a/__tests__/hooks/semantic/pack-semantic-registry.test.ts +++ b/__tests__/hooks/semantic/pack-semantic-registry.test.ts @@ -2,17 +2,16 @@ /** * The rule that decides which semantic policy set a machine asks Jev about. * - * The replacement rule is the load-bearing half: a pack that declares at least - * one `semantic` entry replaces the compiled-in set WHOLESALE, mirroring the rule - * already in force for the regex builtins. Anything softer — merging, or - * preferring one on a name collision — means two question sets can both claim - * `destructive-deletion`, and a `reviewedBy` naming it would mean different - * things on two machines. + * The package ships NO Jev checks: the set is exactly what installed packs + * declare, and empty when none declares any — no pack, no questions. Two + * question sets can never both claim `destructive-deletion`, so a `reviewedBy` + * naming it cannot mean different things on two machines. */ import { describe, expect, it } from "vitest"; import { parsePackSemanticPolicy, type SemanticManifestEntry } from "@/src/hooks/pack-manifest"; -import { SEMANTIC_POLICIES } from "@/src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import { + JEV_POLICIES_QUESTION_CHARS, MAX_PACK_QUESTION_CHARS, questionChars, semanticPoliciesFromPacks, @@ -38,27 +37,45 @@ const manifestEntry = (over: Partial = {}): SemanticM /** First-party, because the builtin check names these use are reserved to FailproofAI's packs. */ const pack = (id: string, semantic: SemanticManifestEntry[]) => ({ id, semantic, source: `github:FailproofAI/${id.split("/")[1]}@v1` }); -describe("semanticPoliciesFromPacks — the replacement rule", () => { - it("returns the compiled-in set, by identity, when no pack declares any", () => { +describe("semanticPoliciesFromPacks — no pack, no questions", () => { + it("asks nothing when no pack declares any check", () => { const resolved = semanticPoliciesFromPacks([pack("acme/guards", [])]); - expect(resolved.policies).toBe(SEMANTIC_POLICIES); + expect(resolved.policies).toEqual([]); expect(resolved.fromPack).toBe(false); expect(resolved.errors).toEqual([]); }); - it("returns the compiled-in set when there are no packs at all", () => { - expect(semanticPoliciesFromPacks([]).policies).toBe(SEMANTIC_POLICIES); + it("asks nothing when there are no packs at all — the vanilla install", () => { + const resolved = semanticPoliciesFromPacks([]); + expect(resolved.policies).toEqual([]); + expect(resolved.fromPack).toBe(false); }); - it("replaces the compiled-in set wholesale once one pack declares any", () => { + it("asks exactly what one declaring pack declares, and nothing compiled in", () => { const resolved = semanticPoliciesFromPacks([pack("acme/guards", [manifestEntry()])]); expect(resolved.fromPack).toBe(true); expect(resolved.policies.map((p) => p.name)).toEqual(["destructive-deletion"]); - // Not merged: the other fifteen builtins are gone, so no name can be claimed - // twice and a `reviewedBy` cannot mean two things. expect(resolved.policies).toHaveLength(1); }); + it("asks FailproofAI/jev-policies' sixteen, in order, once that pack is installed", () => { + const sixteen = SEMANTIC_POLICIES.map((p, i) => + parsePackSemanticPolicy( + "FailproofAI/jev-policies", + { + name: p.name, title: p.title, appliesTo: p.appliesTo, mode: p.mode, userCanOverride: p.userCanOverride, + probes: p.probes, ...(p.exempt ? { exempt: p.exempt } : {}), guidance: p.guidance, + } as SemanticPolicyDeclaration, + i, + ), + ); + const resolved = semanticPoliciesFromPacks([ + { id: "FailproofAI/jev-policies", semantic: sixteen, source: "github:FailproofAI/jev-policies@v0.2.0" }, + ]); + expect(resolved.errors).toEqual([]); + expect(resolved.policies.map((p) => p.name)).toEqual(SEMANTIC_POLICIES.map((p) => p.name)); + }); + it("concatenates two declaring packs, in installed order", () => { const resolved = semanticPoliciesFromPacks([ pack("acme/guards", [manifestEntry()]), @@ -96,25 +113,23 @@ describe("semanticPoliciesFromPacks — the replacement rule", () => { ); }); - it("falls back to the compiled-in set when the contest leaves nothing", () => { - // The same rule as the unusable-entry case below, and it matters that the two - // agree: `effectiveReviewerNames` falls back in this state too, so the names - // a `reviewedBy` may use are the names of the questions being asked. + it("asks nothing when the contest leaves nothing", () => { + // There is no compiled-in set to fall back to. `effectiveReviewerNames` is + // empty in this state too, so no `reviewedBy` is honoured and nothing clears. const resolved = semanticPoliciesFromPacks([ pack("acme/guards", [manifestEntry()]), pack("evil/extra", [manifestEntry({ guidance: "Nothing to see here." })]), ]); - expect(resolved.policies).toBe(SEMANTIC_POLICIES); + expect(resolved.policies).toEqual([]); expect(resolved.fromPack).toBe(false); }); - it("falls back to the compiled-in set when every declared entry was unusable", () => { - // Honest (it is what the machine ran yesterday) and safe: what a pack's regex - // half names in `reviewedBy` will not match the builtin set, so those - // policies stay hard rather than being cleared by questions nobody validated. + it("asks nothing when every declared entry was unusable", () => { + // Safe: nothing is asked, so nothing a `reviewedBy` names is ever answered + // and those policies stay hard. const broken = { ...manifestEntry(), precondition: "on_a_tuesday" } as SemanticManifestEntry; const resolved = semanticPoliciesFromPacks([pack("acme/guards", [broken])]); - expect(resolved.policies).toBe(SEMANTIC_POLICIES); + expect(resolved.policies).toEqual([]); expect(resolved.fromPack).toBe(false); expect(resolved.errors).toHaveLength(1); }); @@ -187,6 +202,10 @@ describe("the question budget", () => { 0, ); expect(shipped).toBeLessThan(MAX_PACK_QUESTION_CHARS); + // What `publish` reserves for them beside a stranger's pack is their real cost. + expect(JEV_POLICIES_QUESTION_CHARS).toBe( + SEMANTIC_POLICIES.reduce((n, p) => n + questionChars(p as unknown as SemanticManifestEntry), 0), + ); }); it("drops the entries past the budget, keeps the ones before, and names the shortfall", () => { @@ -221,9 +240,9 @@ describe("the question budget", () => { }); }); -describe("a third-party pack's checks join the built-in ones; FailproofAI's replace them", () => { +describe("a third-party pack's checks beside FailproofAI's", () => { const thirdParty = (id: string, semantic: SemanticManifestEntry[]) => ({ id, semantic, source: `github:${id}@v1` }); - /** The compiled-in sixteen, as FailproofAI/jev-policies would declare them. */ + /** The sixteen, as FailproofAI/jev-policies declares them. */ const firstPartySixteen = SEMANTIC_POLICIES.map((p, i) => parsePackSemanticPolicy( "FailproofAI/jev-policies", @@ -235,10 +254,23 @@ describe("a third-party pack's checks join the built-in ones; FailproofAI's repl ), ); - it("a stranger's one check does not switch off the built-in deny checks", () => { + it("a stranger's one check is the whole set on a machine without FailproofAI/jev-policies", () => { const resolved = semanticPoliciesFromPacks([thirdParty("acme/db", [manifestEntry({ name: "acme-db-check" })])]); - const names = resolved.policies.map((p) => p.name); - expect(names).toEqual([...SEMANTIC_POLICIES.map((p) => p.name), "acme-db-check"]); + expect(resolved.policies.map((p) => p.name)).toEqual(["acme-db-check"]); + }); + + it("a stranger's check is added to FailproofAI's, which keep their order", () => { + const resolved = semanticPoliciesFromPacks([ + thirdParty("acme/db", [manifestEntry({ name: "acme-db-check" })]), + { id: "FailproofAI/jev-policies", semantic: firstPartySixteen, source: "github:FailproofAI/jev-policies@v1" }, + ]); + expect(resolved.policies.map((p) => p.name)).toEqual([...SEMANTIC_POLICIES.map((p) => p.name), "acme-db-check"]); + }); + + it("a stranger cannot claim one of FailproofAI's names, with or without that pack installed", () => { + const resolved = semanticPoliciesFromPacks([thirdParty("acme/db", [manifestEntry({ name: "destructive-deletion" })])]); + expect(resolved.policies).toEqual([]); + expect(resolved.errors.join(" ")).toMatch(/reserved for FailproofAI's own Jev checks/); }); it("install order cannot spend FailproofAI's budget on a stranger's pack", () => { diff --git a/__tests__/hooks/semantic/pack-semantic-wiring.test.ts b/__tests__/hooks/semantic/pack-semantic-wiring.test.ts index 6feec16ed..7656e53c3 100644 --- a/__tests__/hooks/semantic/pack-semantic-wiring.test.ts +++ b/__tests__/hooks/semantic/pack-semantic-wiring.test.ts @@ -1,12 +1,13 @@ // @vitest-environment node /** * The one wiring point: `prepareSemantic` asks about the set an installed pack - * declared, and about the compiled-in set when no pack declares one. + * declared — and about NOTHING when no pack declares one. The package ships no + * Jev checks of its own. * * Driven through the real reader with a real manifest and a real digest, because * the thing worth proving is not that the resolver returns the right array — the * unit tests beside this do that — but that the evaluator actually consults it, - * and that a machine with no pack is unchanged. + * and that a machine with no pack asks nothing and sends nothing. */ import { describe, expect, it, beforeEach, afterEach, vi } from "vitest"; import { createHash } from "node:crypto"; @@ -14,7 +15,7 @@ import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; import { prepareSemantic } from "@/src/hooks/semantic/evaluator"; -import { SEMANTIC_POLICIES } from "@/src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import { resolveSemanticPolicies, _resetSemanticWarningsForTest } from "@/src/hooks/semantic/pack-policies"; import type { SemanticInput } from "@/src/hooks/semantic/types"; @@ -42,8 +43,7 @@ function writeManifest(over: Record = {}): void { schemaVersion: 1, packs: [ { - // A FailproofAI pack: its checks REPLACE the compiled-in set, which is - // the rule these tests pin. A third party's are added to it. + // A FailproofAI pack, so its check names are its own to declare. id: "FailproofAI/guards", version: "1.0.0", source: "github:FailproofAI/guards@v1.0.0", @@ -83,16 +83,20 @@ afterEach(() => { }); describe("resolveSemanticPolicies", () => { - it("is the compiled-in set with no manifest at all", () => { - expect(resolveSemanticPolicies()).toBe(SEMANTIC_POLICIES); + it("is empty with no manifest at all — the vanilla install asks nothing", () => { + expect(resolveSemanticPolicies()).toEqual([]); }); - it("is the compiled-in set when the manifest is unreadable", () => { - // The same fail-open posture every other reader of this file takes — and here - // it also fails safe: a pack's `reviewedBy` will not match the builtin names, - // so nothing is cleared by a question nobody could read. + it("is empty when the manifest is unreadable, never a compiled-in fallback", () => { + // Fails safe: nothing is asked, so nothing is cleared by a question nobody + // could read, and every regex verdict stands. writeFileSync(join(root, "installed.json"), "{ not json"); - expect(resolveSemanticPolicies()).toBe(SEMANTIC_POLICIES); + expect(resolveSemanticPolicies()).toEqual([]); + }); + + it("is empty when the installed packs declare no checks", () => { + writeManifest(); + expect(resolveSemanticPolicies()).toEqual([]); }); it("is the pack's set once it declares one", () => { @@ -118,19 +122,19 @@ describe("resolveSemanticPolicies", () => { }); describe("prepareSemantic consults the resolved set", () => { - it("asks the compiled-in questions when no pack declares any", () => { + it("asks nothing, and builds an empty request, when no pack declares any", () => { writeManifest(); const prepared = prepareSemantic(input); - const names = prepared.selected.map((p) => p.name); - expect(names).toContain("destructive-deletion"); - expect(names.every((n) => SEMANTIC_POLICIES.some((p) => p.name === n))).toBe(true); + expect(prepared.selected).toEqual([]); + // Not even the injection or task probes: they ride along with a check. + expect(Object.keys(prepared.compiled.request.questions)).toEqual([]); }); it("asks the pack's questions instead once it declares any", () => { writeManifest({ semantic: [semanticEntry()] }); const prepared = prepareSemantic(input); expect(prepared.selected.map((p) => p.name)).toEqual(["pack-destructive-deletion"]); - // Wholesale: the builtin question ids are not in the request either. + // Only the pack's: no compiled-in question ids are in the request. expect(Object.keys(prepared.compiled.request.questions)).toContain("pack-destructive-deletion.destroys"); expect(Object.keys(prepared.compiled.request.questions)).not.toContain("destructive-deletion.destroys"); }); diff --git a/__tests__/hooks/semantic/policy-preconditions.test.ts b/__tests__/hooks/semantic/policy-preconditions.test.ts index bf7fc85ef..ad38e2cb2 100644 --- a/__tests__/hooks/semantic/policy-preconditions.test.ts +++ b/__tests__/hooks/semantic/policy-preconditions.test.ts @@ -14,7 +14,7 @@ import { describe, expect, it } from "vitest"; import { selectPolicies } from "../../../src/hooks/semantic/compile"; import { computeFacts, scanCommand } from "../../../src/hooks/semantic/facts"; -import { SEMANTIC_POLICIES } from "../../../src/hooks/semantic/policies"; +import { JEV_PACK_POLICIES as SEMANTIC_POLICIES } from "../../fixtures/jev-policies"; import type { Facts, PathFact } from "../../../src/hooks/semantic/types"; const PROJECT = "/home/dev/project"; diff --git a/__tests__/hooks/semantic/truncation-severity.test.ts b/__tests__/hooks/semantic/truncation-severity.test.ts index 6abfef4d9..a710efd7f 100644 --- a/__tests__/hooks/semantic/truncation-severity.test.ts +++ b/__tests__/hooks/semantic/truncation-severity.test.ts @@ -46,6 +46,11 @@ import { evaluateSemantic, prepareSemantic, type SemanticOptions, type SemanticO import { toReview } from "../../../src/hooks/semantic/jev-review"; import type { JevReview } from "../../../src/hooks/semantic/combine"; import type { JevRequest, JevResponse, SemanticInput } from "../../../src/hooks/semantic/types"; +import { withInstalledJevPoliciesPack } from "../../fixtures/jev-policies"; + +// The package ships no Jev checks; this file runs as a machine with +// FailproofAI/jev-policies installed. +withInstalledJevPoliciesPack(); /** Every "does it do X" probe held; the human asked for none of it. Jev denies. */ const alarmed = async (request: JevRequest): Promise => ({ @@ -125,9 +130,9 @@ describe("padding a command cannot take Jev's own deny away", () => { }); }); - it("shadow mode is unaffected: the regex result is enforced, cut or not", async () => { + it("observe mode is unaffected: the regex result is enforced, cut or not", async () => { const { review } = await judged(bash(`${DANGEROUS} ${PADDING}`)); - const out = combineTwoTier([], review, "shadow"); + const out = combineTwoTier([], review, "observe"); expect(out.final).toEqual(regexOnly([])); expect(out.decidedByJev).toBe(false); }); diff --git a/__tests__/hooks/session-pause-enforcement.test.ts b/__tests__/hooks/session-pause-enforcement.test.ts index 51302191b..9d4a8929a 100644 --- a/__tests__/hooks/session-pause-enforcement.test.ts +++ b/__tests__/hooks/session-pause-enforcement.test.ts @@ -42,6 +42,7 @@ vi.mock("../../src/hooks/pack-manifest", () => ({ // The handler asks this per event to decide whether the migration shim // still applies. Mirrors the mocked readInstalledPacks above. hasInstalledPacks: vi.fn(() => false), + hasInstalledRegexPacks: vi.fn(() => false), })); import { evaluateHookEvent } from "../../src/hooks/handler"; diff --git a/__tests__/hooks/two-tier-handler.test.ts b/__tests__/hooks/two-tier-handler.test.ts index 8000c0141..d01a9e140 100644 --- a/__tests__/hooks/two-tier-handler.test.ts +++ b/__tests__/hooks/two-tier-handler.test.ts @@ -35,7 +35,31 @@ import type { JevConfig } from "../../src/hooks/semantic/jev-config"; let jevConfig: JevConfig | null = null; /** Overrides the build's DEFAULT_JEV_MODE (D2) for one test; undefined → the real one. */ -let defaultModeOverride: "shadow" | "enforce" | undefined; +let defaultModeOverride: "observe" | "enforce" | undefined; +// The package ships no Jev checks, so a handler only starts a review on a +// machine with a pack declaring some. These tests drive the two-tier path over +// the migration shim's builtins (which a real pack install would switch off), +// so they stand in for FailproofAI/jev-policies directly: its sixteen checks +// are the questions, and their names the reviewers. +vi.mock("../../src/hooks/effective-reviewers", async (importOriginal) => { + const real = await importOriginal(); + const { SEMANTIC_REVIEWER_NAMES } = await import("../../src/hooks/policy-authority"); + return { ...real, effectiveReviewerNames: () => SEMANTIC_REVIEWER_NAMES, jevChecksInstalled: () => true }; +}); +vi.mock("../../src/hooks/semantic/pack-policies", async (importOriginal) => { + const real = await importOriginal(); + const { JEV_PACK_POLICIES } = await import("../fixtures/jev-policies"); + // A test that installs its own pack gets that pack's checks, as a machine + // would; every other test stands in for FailproofAI/jev-policies. + return { + ...real, + resolveSemanticPolicies: (cli?: string) => { + const installed = real.resolveSemanticPolicies(cli); + return installed.length > 0 ? installed : JEV_PACK_POLICIES; + }, + }; +}); + vi.mock("../../src/hooks/semantic/jev-config", async (importOriginal) => { const actual = await importOriginal(); return { @@ -523,8 +547,8 @@ describe("a reviewable deny", () => { expect(typeof row.jevLatencyMs).toBe("number"); }); - it("in shadow mode: the regex deny is enforced and the would-be clear recorded", async () => { - jevConfig = { ...CFG, mode: "shadow" }; + it("in observe mode: the regex deny is enforced and the would-be clear recorded", async () => { + jevConfig = { ...CFG, mode: "observe" }; const enforced = await outsideRead(); expect(enforced.outcome.evaluation?.decision).toBe("deny"); expect(enforced.outcome.evaluation?.policyName).toBe("failproofai/block-read-outside-cwd"); @@ -532,9 +556,9 @@ describe("a reviewable deny", () => { evaluator: "jev", jevDecision: "allow", jevCleared: ["failproofai/block-read-outside-cwd"], - jevMode: "shadow", + jevMode: "observe", }); - // …and what shadow enforces is byte-identical to the unconfigured answer. + // …and what observe enforces is byte-identical to the unconfigured answer. jevConfig = null; const plain = await outsideRead(); expect(enforced.outcome.stdout).toBe(plain.outcome.stdout); @@ -732,11 +756,12 @@ describe("Jev's own verdict", () => { expect(outcome.evaluation?.policyName).toBe("semantic/acme-prod-deploy"); expect(row).toMatchObject({ policySource: "jev", packId: "acme/deploys", packVersion: "2.0.0" }); - // A compiled-in check stays unattributed to any pack. + // The package ships no checks of its own, so with only this pack installed + // a check it does not declare is never asked and decides nothing. respond = answers({ "destructive-deletion": 0.97 }); - const builtin = await bash("find . -name '*.sqlite' -delete"); - expect(builtin.row.policySource).toBe("jev"); - expect(builtin.row.packId).toBeUndefined(); + const undeclared = await bash("find . -name '*.sqlite' -delete"); + expect(undeclared.outcome.evaluation?.policyName ?? null).not.toBe("semantic/destructive-deletion"); + expect(undeclared.row.jevDecision).not.toBe("deny"); }); it("the most severe wins: a regex instruct and a Jev deny → deny", async () => { @@ -1042,24 +1067,24 @@ describe("a padded call cannot make Jev's own deny go away", () => { expect(row.jevFallbackReason).toBeUndefined(); }); - it("shadow mode still enforces the regex result for both spellings", async () => { - jevConfig = { ...CFG, mode: "shadow" }; + it("observe mode still enforces the regex result for both spellings", async () => { + jevConfig = { ...CFG, mode: "observe" }; respond = answers({ "destructive-deletion": 0.97 }); const field = await bash(padField()); expect(field.outcome.evaluation?.decision).toBe("allow"); - expect(field.row).toMatchObject({ evaluator: "jev-fallback", jevDecision: "deny", jevMode: "shadow" }); + expect(field.row).toMatchObject({ evaluator: "jev-fallback", jevDecision: "deny", jevMode: "observe" }); const budget = await run("PreToolUse", { tool_name: "Bash", tool_input: padBudget() }); expect(budget.outcome.evaluation?.decision).toBe("allow"); - expect(budget.row).toMatchObject({ evaluator: "jev-fallback", jevDecision: "deny", jevMode: "shadow" }); + expect(budget.row).toMatchObject({ evaluator: "jev-fallback", jevDecision: "deny", jevMode: "observe" }); - // And a call the envelope had to cut: shadow enforces the regex result + // And a call the envelope had to cut: observe enforces the regex result // either way. respond = seeingTransport as typeof respond; const hidden = await bash(`echo ${overflow("x")} ; ${DELETE} ; echo ${overflow("y")}`); expect(hidden.outcome.evaluation?.decision).toBe("allow"); - expect(hidden.row).toMatchObject({ evaluator: "jev-fallback", jevFallbackReason: "request-cut", jevMode: "shadow" }); + expect(hidden.row).toMatchObject({ evaluator: "jev-fallback", jevFallbackReason: "request-cut", jevMode: "observe" }); }); it("control: unpadded, the very same deny is a plain `jev` row", async () => { @@ -1331,21 +1356,21 @@ describe("a Jev review that cannot start", () => { it("follows DEFAULT_JEV_MODE rather than restating it", async () => { jevConfig = CFG; - defaultModeOverride = "shadow"; + defaultModeOverride = "observe"; vi.mocked(startJevReview).mockImplementationOnce(() => { throw new Error("module failed to initialise"); }); const { row } = await bash("ls -la"); - expect(row).toMatchObject({ evaluator: "jev-fallback", jevFallbackReason: "error", jevMode: "shadow" }); + expect(row).toMatchObject({ evaluator: "jev-fallback", jevFallbackReason: "error", jevMode: "observe" }); }); it("an explicit mode in the config wins", async () => { - jevConfig = { ...CFG, mode: "shadow" }; + jevConfig = { ...CFG, mode: "observe" }; vi.mocked(startJevReview).mockImplementationOnce(() => { throw new Error("module failed to initialise"); }); const { row } = await bash("ls -la"); - expect(row).toMatchObject({ evaluator: "jev-fallback", jevMode: "shadow" }); + expect(row).toMatchObject({ evaluator: "jev-fallback", jevMode: "observe" }); }); }); @@ -1397,20 +1422,20 @@ describe("captureIntent and a prompt a policy acted on", () => { describe("the throttle's cache across a jev.json change", () => { const outsideRead = () => readFile(join(home, "other", "notes.txt")); - // T1 accepts plain http to a loopback proxy only in shadow mode: its answers + // T1 accepts plain http to a loopback proxy only in observe mode: its answers // must never clear a deny. Both routes ask for the same model, so their // requests are byte-identical. - const LOOPBACK_SHADOW: JevConfig = { provider: "custom", apiKey: "not-a-real-key", baseUrl: "http://127.0.0.1:9", mode: "shadow" }; + const LOOPBACK_OBSERVE: JevConfig = { provider: "custom", apiKey: "not-a-real-key", baseUrl: "http://127.0.0.1:9", mode: "observe" }; const TYPESAFE_ENFORCE: JevConfig = { provider: "typesafe", apiKey: "not-a-real-key", baseUrl: "https://jev.invalid", mode: "enforce" }; it("never serves one provider's answer under another: switching providers asks the new one", async () => { const { JevError } = await import("../../src/hooks/semantic/jev-client"); fakeCache.on = true; - jevConfig = LOOPBACK_SHADOW; - const shadow = await outsideRead(); - expect(shadow.outcome.evaluation?.decision).toBe("deny"); - expect(shadow.row).toMatchObject({ evaluator: "jev", jevMode: "shadow", jevCleared: ["failproofai/block-read-outside-cwd"] }); + jevConfig = LOOPBACK_OBSERVE; + const observe = await outsideRead(); + expect(observe.outcome.evaluation?.decision).toBe("deny"); + expect(observe.row).toMatchObject({ evaluator: "jev", jevMode: "observe", jevCleared: ["failproofai/block-read-outside-cwd"] }); expect(jevCalls).toHaveLength(1); jevConfig = TYPESAFE_ENFORCE; @@ -1831,7 +1856,7 @@ describe("FAILPROOFAI_EVALUATOR=legacy (§4 row 1) under every configured mode: it.each<[string, JevConfig["mode"]]>([ ["no mode (the build's default)", undefined], - ["shadow", "shadow"], + ["observe", "observe"], ["enforce", "enforce"], ])("%s", async (_label, mode) => { respond = answers({ "destructive-deletion": 0.97 }); @@ -1949,19 +1974,19 @@ describe("what the policy page reads", () => { expect(row).toMatchObject({ policyName: "failproofai/block-sudo", policySource: "builtin" }); }); - it("B: in shadow mode, Jev's deny is a 'would have' in observed — and the regex result is enforced", async () => { - jevConfig = { ...CFG, mode: "shadow" }; + it("B: in observe mode, Jev's deny is a 'would have' in observed — and the regex result is enforced", async () => { + jevConfig = { ...CFG, mode: "observe" }; respond = answers({ "destructive-deletion": 0.97 }); - const shadow = await bash(DELETE_ALL); - expect(shadow.outcome.evaluation?.decision).toBe("allow"); - expect(shadow.row.policySource).toBeUndefined(); - expect(shadow.row).toMatchObject({ evaluator: "jev", jevDecision: "deny", jevMode: "shadow" }); + const observe = await bash(DELETE_ALL); + expect(observe.outcome.evaluation?.decision).toBe("allow"); + expect(observe.row.policySource).toBeUndefined(); + expect(observe.row).toMatchObject({ evaluator: "jev", jevDecision: "deny", jevMode: "observe" }); // The reason is the one enforce mode shows for the same answer. jevConfig = CFG; store._resetForTest(join(root, "activity-enforce")); const enforce = await bash(DELETE_ALL); - expect(shadow.row.observed).toEqual([ + expect(observe.row.observed).toEqual([ { policyId: "semantic/destructive-deletion", version: "jev-1.13.0", @@ -1971,8 +1996,8 @@ describe("what the policy page reads", () => { ]); }); - it("B: a shadow instruct is recorded as an instruct", async () => { - jevConfig = { ...CFG, mode: "shadow" }; + it("B: an observe instruct is recorded as an instruct", async () => { + jevConfig = { ...CFG, mode: "observe" }; // An instruct-mode check firing: a warning, not a block. respond = answers({ "system-modification": 0.9 }); const { row } = await bash("sysctl -w vm.swappiness=10"); @@ -1985,7 +2010,7 @@ describe("what the policy page reads", () => { }); it("B: nothing is recorded when Jev allowed, or fell back", async () => { - jevConfig = { ...CFG, mode: "shadow" }; + jevConfig = { ...CFG, mode: "observe" }; respond = answers(); expect((await bash("ls -la")).row.observed).toBeUndefined(); respond = async () => { diff --git a/__tests__/hooks/two-tier-intent-storage.test.ts b/__tests__/hooks/two-tier-intent-storage.test.ts index 664e8da9a..723cb7956 100644 --- a/__tests__/hooks/two-tier-intent-storage.test.ts +++ b/__tests__/hooks/two-tier-intent-storage.test.ts @@ -23,6 +23,22 @@ import type { JevRequest, JevResponse } from "../../src/hooks/semantic/types"; import type { JevConfig } from "../../src/hooks/semantic/jev-config"; let jevConfig: JevConfig | null = null; +// The package ships no Jev checks, so a handler only starts a review on a +// machine with a pack declaring some. These tests drive the two-tier path over +// the migration shim's builtins (which a real pack install would switch off), +// so they stand in for FailproofAI/jev-policies directly: its sixteen checks +// are the questions, and their names the reviewers. +vi.mock("../../src/hooks/effective-reviewers", async (importOriginal) => { + const real = await importOriginal(); + const { SEMANTIC_REVIEWER_NAMES } = await import("../../src/hooks/policy-authority"); + return { ...real, effectiveReviewerNames: () => SEMANTIC_REVIEWER_NAMES, jevChecksInstalled: () => true }; +}); +vi.mock("../../src/hooks/semantic/pack-policies", async (importOriginal) => { + const real = await importOriginal(); + const { JEV_PACK_POLICIES } = await import("../fixtures/jev-policies"); + return { ...real, resolveSemanticPolicies: () => JEV_PACK_POLICIES }; +}); + vi.mock("../../src/hooks/semantic/jev-config", async (importOriginal) => { const actual = await importOriginal(); return { ...actual, loadJevConfig: vi.fn(() => jevConfig) }; diff --git a/__tests__/hooks/two-tier-no-pack-inert.test.ts b/__tests__/hooks/two-tier-no-pack-inert.test.ts new file mode 100644 index 000000000..5fbf7424b --- /dev/null +++ b/__tests__/hooks/two-tier-no-pack-inert.test.ts @@ -0,0 +1,124 @@ +// @vitest-environment node +/** + * No pack, no Jev: a machine that CONFIGURED Jev — a valid BYOK `jev.json` in + * enforce mode — but installed no pack declaring Jev checks must answer every + * hook exactly as a machine that never configured Jev, and must never reach + * the network for it. + * + * The npm package ships no Jev checks. They reach a machine only through + * `failproofai policies add FailproofAI/jev-policies`, so until then there is + * nothing to ask: no request (not even the injection or task probes), no intent + * capture, and no reviewer, so every `reviewable` policy resolves hard. + * + * The reference is the same unconfigured golden `two-tier-unconfigured- + * equivalence.test.ts` compares against, over the same corpus: every builtin + * enabled through the migration shim, real tool calls on every CLI. The last + * test installs the pack and shows the same call now does reach Jev, so the + * silence above is the pack's absence and not something else switching Jev off. + */ +import { describe, it, expect, beforeAll, afterAll, beforeEach, afterEach } from "vitest"; +import { chmodSync, readFileSync, writeFileSync } from "node:fs"; +import { join, resolve } from "node:path"; +import type { Golden } from "./two-tier/corpus"; +import { enterSandbox, runHandlerCorpus, type CorpusSandbox } from "./two-tier/runner"; +import { installJevPoliciesPack } from "../fixtures/jev-policies"; + +const golden = JSON.parse( + readFileSync(resolve(__dirname, "../fixtures/two-tier/unconfigured-golden.json"), "utf8"), +) as Golden; + +const CORPUS_TIMEOUT_MS = 30_000; + +let sandbox: CorpusSandbox; +const realFetch = globalThis.fetch; +let requests: string[] = []; + +beforeAll(() => { + sandbox = enterSandbox(); + const fpHome = process.env.FAILPROOFAI_HOME!; + // jev.json is read only from an owner-only directory. + chmodSync(fpHome, 0o700); + // A real, loadable BYOK config in ENFORCE mode — the strongest opt-in there is. + writeFileSync( + join(fpHome, "jev.json"), + JSON.stringify({ provider: "typesafe", apiKey: "ts-test-key-not-real-0123456789", mode: "enforce" }), + { mode: 0o600 }, + ); +}); +afterAll(() => { + sandbox.restore(); +}); + +beforeEach(() => { + requests = []; + globalThis.fetch = (async (url: unknown) => { + requests.push(String(url)); + throw new Error("no network in this test"); + }) as typeof fetch; +}); +afterEach(() => { + globalThis.fetch = realFetch; +}); + +describe("Jev configured, no pack declaring Jev checks", () => { + it("the config really is a usable, enforcing Jev config (or the rest proves nothing)", async () => { + const { loadJevConfig } = await import("../../src/hooks/semantic/jev-config"); + expect(loadJevConfig()).toMatchObject({ provider: "typesafe", mode: "enforce" }); + }); + + it("has no Jev checks and no reviewers", async () => { + const { effectiveReviewerNames, forgetEffectiveReviewerNames, jevChecksInstalled } = await import( + "../../src/hooks/effective-reviewers" + ); + const { resolveSemanticPolicies } = await import("../../src/hooks/semantic/pack-policies"); + forgetEffectiveReviewerNames(); + expect(effectiveReviewerNames().size).toBe(0); + expect(jevChecksInstalled()).toBe(false); + expect(resolveSemanticPolicies()).toEqual([]); + }); + + it( + "evaluateHookEvent: every hook is byte-identical to a machine with no jev.json, and nothing is sent", + async () => { + const mismatches: string[] = []; + let seen = 0; + await runHandlerCorpus((id, value) => { + seen++; + const want = golden.handler[id]; + const gotOut = JSON.stringify(value.out); + if (!want) { + mismatches.push(`${id}: not in the golden`); + return; + } + if (gotOut !== golden.outputs[want.out]) { + mismatches.push(`${id}\n want ${golden.outputs[want.out]}\n got ${gotOut}`); + } + if (value.activity !== want.activity) { + mismatches.push(`${id}: activity row differs (want digest ${want.activity}); got ${JSON.stringify(value.activityRow)}`); + } + }, sandbox); + expect(seen).toBe(Object.keys(golden.handler).length); + expect(mismatches.slice(0, 5)).toEqual([]); + expect(requests).toEqual([]); + }, + CORPUS_TIMEOUT_MS, + ); + + it("installing FailproofAI/jev-policies is what makes the same call reach Jev", async () => { + const { evaluateHookEvent } = await import("../../src/hooks/handler"); + const payload = JSON.stringify({ + session_id: "no-pack-inert", + hook_event_name: "PreToolUse", + tool_name: "Bash", + tool_input: { command: "rm -rf ./build" }, + cwd: sandbox.root, + }); + + await evaluateHookEvent("PreToolUse", "claude", payload, { awaitTelemetryFlush: false }); + expect(requests).toEqual([]); + + installJevPoliciesPack(process.env.FAILPROOFAI_PACK_DIR!); + await evaluateHookEvent("PreToolUse", "claude", payload, { awaitTelemetryFlush: false }); + expect(requests.length).toBeGreaterThan(0); + }); +}); diff --git a/__tests__/hooks/two-tier-unconfigured-load.test.ts b/__tests__/hooks/two-tier-unconfigured-load.test.ts index 9b47e74d6..92a7ef135 100644 --- a/__tests__/hooks/two-tier-unconfigured-load.test.ts +++ b/__tests__/hooks/two-tier-unconfigured-load.test.ts @@ -29,8 +29,9 @@ vi.mock("../../src/hooks/semantic/jev-config", async (importOriginal) => { import { evaluateHookEvent } from "../../src/hooks/handler"; import { _resetForTest } from "../../src/hooks/hook-activity-store"; +import { installJevPoliciesPack } from "../fixtures/jev-policies"; -const ENV = ["HOME", "FAILPROOFAI_HOME", "FAILPROOFAI_EVALUATOR", "CLAUDE_PROJECT_DIR"] as const; +const ENV = ["HOME", "FAILPROOFAI_HOME", "FAILPROOFAI_PACK_DIR", "FAILPROOFAI_EVALUATOR", "CLAUDE_PROJECT_DIR"] as const; const saved: Record = {}; let root: string; let fpHome: string; @@ -46,6 +47,8 @@ beforeEach(() => { writeFileSync(join(fpHome, "policies-config.json"), JSON.stringify({ enabledPolicies: ["block-sudo"] })); process.env.HOME = join(root, "home"); process.env.FAILPROOFAI_HOME = fpHome; + process.env.FAILPROOFAI_PACK_DIR = join(root, "packs"); + mkdirSync(join(root, "packs"), { recursive: true }); delete process.env.FAILPROOFAI_EVALUATOR; delete process.env.CLAUDE_PROJECT_DIR; _resetForTest(join(root, "activity")); @@ -66,7 +69,7 @@ const run = (event: string, payload: Record) => }); describe("the unconfigured hot path never loads the Jev config module", () => { - it("loads it only once a jev.json is there", async () => { + it("loads it only once a jev.json is there AND a pack gives Jev checks", async () => { // Both callers: the gate event that would start a review, and the prompt // event that would record the human's intent for one. const gate = await run("PreToolUse", { tool_name: "Bash", tool_input: { command: "sudo rm -rf /" } }); @@ -81,6 +84,15 @@ describe("the unconfigured hot path never loads the Jev config module", () => { // loading the module that knows how. writeFileSync(join(fpHome, "jev.json"), "{}"); await run("PreToolUse", { tool_name: "Bash", tool_input: { command: "sudo rm -rf /" } }); + await run("UserPromptSubmit", { prompt: "clean the build" }); + // Still not: the package ships no Jev checks, so with no pack declaring any + // there is nothing to ask and the config is never worth reading. + expect(seen.jevConfigModule).toBe(false); + + // FailproofAI/jev-policies gives it its checks. (It is a pack, so the + // migration shim's builtins stop; the always-on guard is all that is left.) + installJevPoliciesPack(join(root, "packs")); + await run("PreToolUse", { tool_name: "Bash", tool_input: { command: "sudo rm -rf /" } }); expect(seen.jevConfigModule).toBe(true); }); }); diff --git a/__tests__/hooks/two-tier-worker-optout.test.ts b/__tests__/hooks/two-tier-worker-optout.test.ts index d6755c6a4..e12e8208e 100644 --- a/__tests__/hooks/two-tier-worker-optout.test.ts +++ b/__tests__/hooks/two-tier-worker-optout.test.ts @@ -27,6 +27,22 @@ import type { JevConfig } from "../../src/hooks/semantic/jev-config"; import type { JevRequest, JevResponse } from "../../src/hooks/semantic/types"; import { resetJevThrottle } from "../../src/hooks/semantic/jev-throttle"; +// The package ships no Jev checks, so a handler only starts a review on a +// machine with a pack declaring some. These tests drive the two-tier path over +// the migration shim's builtins (which a real pack install would switch off), +// so they stand in for FailproofAI/jev-policies directly: its sixteen checks +// are the questions, and their names the reviewers. +vi.mock("../../src/hooks/effective-reviewers", async (importOriginal) => { + const real = await importOriginal(); + const { SEMANTIC_REVIEWER_NAMES } = await import("../../src/hooks/policy-authority"); + return { ...real, effectiveReviewerNames: () => SEMANTIC_REVIEWER_NAMES, jevChecksInstalled: () => true }; +}); +vi.mock("../../src/hooks/semantic/pack-policies", async (importOriginal) => { + const real = await importOriginal(); + const { JEV_PACK_POLICIES } = await import("../fixtures/jev-policies"); + return { ...real, resolveSemanticPolicies: () => JEV_PACK_POLICIES }; +}); + vi.mock("../../src/hooks/hook-telemetry", () => ({ trackHookEvent: vi.fn(() => Promise.resolve()), flushHookTelemetry: vi.fn(() => Promise.resolve()), diff --git a/__tests__/hooks/two-tier-worker-queue.test.ts b/__tests__/hooks/two-tier-worker-queue.test.ts index 340be57e4..da6c9be9f 100644 --- a/__tests__/hooks/two-tier-worker-queue.test.ts +++ b/__tests__/hooks/two-tier-worker-queue.test.ts @@ -20,6 +20,22 @@ import { join } from "node:path"; import type { JevConfig } from "../../src/hooks/semantic/jev-config"; import type { JevRequest, JevResponse } from "../../src/hooks/semantic/types"; +// The package ships no Jev checks, so a handler only starts a review on a +// machine with a pack declaring some. These tests drive the two-tier path over +// the migration shim's builtins (which a real pack install would switch off), +// so they stand in for FailproofAI/jev-policies directly: its sixteen checks +// are the questions, and their names the reviewers. +vi.mock("../../src/hooks/effective-reviewers", async (importOriginal) => { + const real = await importOriginal(); + const { SEMANTIC_REVIEWER_NAMES } = await import("../../src/hooks/policy-authority"); + return { ...real, effectiveReviewerNames: () => SEMANTIC_REVIEWER_NAMES, jevChecksInstalled: () => true }; +}); +vi.mock("../../src/hooks/semantic/pack-policies", async (importOriginal) => { + const real = await importOriginal(); + const { JEV_PACK_POLICIES } = await import("../fixtures/jev-policies"); + return { ...real, resolveSemanticPolicies: () => JEV_PACK_POLICIES }; +}); + vi.mock("../../src/hooks/hook-telemetry", () => ({ trackHookEvent: vi.fn(() => Promise.resolve()), flushHookTelemetry: vi.fn(() => Promise.resolve()), diff --git a/app/actions/get-jev-config.ts b/app/actions/get-jev-config.ts index 03b67a666..2229ee9b6 100644 --- a/app/actions/get-jev-config.ts +++ b/app/actions/get-jev-config.ts @@ -55,6 +55,7 @@ import { baseUrlWithoutQuery, inspectJevConfig, looksLikeCredential, + parseJevMode, readJevConfigForUpdate, type JevConfig, type JevProviderKind, @@ -307,7 +308,6 @@ function routingFromRaw(raw: Record | null): { const provider = (JEV_PROVIDER_KINDS as readonly string[]).includes(providerRaw) ? (providerRaw as JevProviderKind) : null; - const modeRaw = raw?.mode; return { provider, baseUrl: asString(raw?.baseUrl), @@ -315,7 +315,7 @@ function routingFromRaw(raw: Record | null): { // refused file may be a pasted key, and is not shown. accountId: CLOUDFLARE_ACCOUNT_ID_RE.test(asString(raw?.accountId)) ? asString(raw?.accountId) : "", model: asString(raw?.model), - mode: modeRaw === "off" || modeRaw === "shadow" || modeRaw === "enforce" ? modeRaw : DEFAULT_JEV_MODE, + mode: parseJevMode(raw?.mode) ?? DEFAULT_JEV_MODE, }; } diff --git a/app/actions/update-jev-config.ts b/app/actions/update-jev-config.ts index 8be1285a1..9486fdb26 100644 --- a/app/actions/update-jev-config.ts +++ b/app/actions/update-jev-config.ts @@ -10,7 +10,7 @@ * `writeJsonAtomically` the CLI's `jev setup` calls, with the same * `{ mode: 0o600, dirMode: 0o700 }`, after the same `validateJevConfig` the * loader itself runs. Nothing here re-states a rule that lives in - * `jev-config.ts`: not the URL scheme, not "plain http only in shadow mode", + * `jev-config.ts`: not the URL scheme, not "plain http only in observe mode", * not "cloudflare needs an account id". A second copy of those rules is how the * dashboard ends up writing a file the hooks then refuse — the exact failure * mode this module exists to avoid. @@ -112,6 +112,7 @@ import { baseUrlWithoutQuery, endpointGivenAsBase, jevConfigPath, + parseJevMode, providerHostConflict, readJevConfigFileForUpdate, validateApiKey, @@ -134,7 +135,7 @@ export interface JevConfigInput { baseUrl: string; /** Cloudflare only. */ accountId: string; - /** "off" | "shadow" | "enforce". */ + /** "off" | "observe" | "enforce". */ mode: string; /** "" keeps the key where it is: the stored one, or the environment's. */ token: string; @@ -303,8 +304,9 @@ export async function saveJevConfigAction(input: JevConfigInput): Promise { +export async function setJevModeAction(requested: string): Promise { const refusal = await crossOriginRefusal(); if (refusal) return { ok: false, problem: refusal }; - if (mode !== "off" && mode !== "shadow" && mode !== "enforce") { - return { ok: false, problem: 'mode must be "off", "shadow" or "enforce".' }; + const mode = parseJevMode(requested); + if (mode === null) { + return { ok: false, problem: 'mode must be "off", "observe" or "enforce".' }; } const existingFile = readJevConfigFileForUpdate(); diff --git a/app/components/jev-notices.tsx b/app/components/jev-notices.tsx index f9667cf59..e974d240b 100644 --- a/app/components/jev-notices.tsx +++ b/app/components/jev-notices.tsx @@ -7,7 +7,7 @@ * Only the rows where Jev changed or could have changed something get a pill: * a clear (a regex deny Jev overruled), a fallback (Jev's answer was not used — * unavailable, truncated or mismatched — and the regex policies decided alone), - * and a shadow-mode row where Jev disagreed with what was enforced. A fallback + * and an observe-mode row where Jev disagreed with what was enforced. A fallback * where Jev did answer (the call was truncated to fit the envelope) and its * unapplied verdict was stricter than what was enforced gets a louder fallback * pill; the collector ships that row on its own for the same reason. @@ -26,16 +26,16 @@ const SEVERITY: Record = { allow: 0, instruct: 1, deny: 2 }; /** Which pill a row gets, if any. Exported for tests. */ export function jevPillKind( item: JevRow, -): "cleared" | "would-clear" | "fallback" | "fallback-stricter" | "shadow-stricter" | null { +): "cleared" | "would-clear" | "fallback" | "fallback-stricter" | "observe-stricter" | null { const e = sanitizeJevActivity(item); const outcome = jevOutcome(e); if (outcome === null || outcome === "not-consulted" || outcome === "no-request") return null; const jevWasStricter = () => (SEVERITY[e.jevDecision ?? "allow"] ?? 0) > (SEVERITY[item.decision ?? "allow"] ?? 0); if (outcome === "fallback") return e.jevDecision !== undefined && jevWasStricter() ? "fallback-stricter" : "fallback"; const cleared = (e.jevCleared ?? []).length > 0; - if (e.jevMode === "shadow") { + if (e.jevMode === "observe") { if (cleared) return "would-clear"; - return jevWasStricter() ? "shadow-stricter" : null; + return jevWasStricter() ? "observe-stricter" : null; } return cleared ? "cleared" : null; } @@ -47,13 +47,13 @@ const PILLS = { className: "border-sky-500/40 bg-sky-500/10 text-sky-600 dark:text-sky-400", }, "would-clear": { - label: "jev shadow", - title: "Shadow mode: Jev would have cleared a block here; the regex result was enforced", + label: "jev observe", + title: "Observe mode: Jev would have cleared a block here; the regex result was enforced", className: "border-sky-500/30 bg-sky-500/5 text-sky-600/80 dark:text-sky-400/80", }, - "shadow-stricter": { - label: "jev shadow", - title: "Shadow mode: Jev would have been stricter here; the regex result was enforced", + "observe-stricter": { + label: "jev observe", + title: "Observe mode: Jev would have been stricter here; the regex result was enforced", className: "border-sky-500/30 bg-sky-500/5 text-sky-600/80 dark:text-sky-400/80", }, "fallback-stricter": { @@ -68,7 +68,7 @@ const PILLS = { }, } as const; -/** Marks a row where Jev cleared, fell back, or (in shadow mode) disagreed. */ +/** Marks a row where Jev cleared, fell back, or (in observe mode) disagreed. */ export function JevPill({ item }: { item: JevRow }) { const kind = jevPillKind(item); if (!kind) return null; diff --git a/app/settings/jev-panel.tsx b/app/settings/jev-panel.tsx index 1d78be36c..c51cb0543 100644 --- a/app/settings/jev-panel.tsx +++ b/app/settings/jev-panel.tsx @@ -45,7 +45,7 @@ * and a save leaves whatever is stored alone. * * Client-side validation here is a convenience only. Every rule — the URL - * scheme, plain http being refused outside shadow mode, cloudflare's account id + * scheme, plain http being refused outside observe mode, cloudflare's account id * — is enforced server-side by the same `validateJevConfig` the loader runs. * * ## FailproofAI Cloud @@ -53,7 +53,7 @@ * When `jev.json` names the FailproofAI Cloud provider, the endpoint and the * key are not this page's to edit: both come from the machine's connection * (`config --token`). So the form is replaced by the two controls that are - * the owner's — an on/off switch and shadow/enforce — both through + * the owner's — an on/off switch and observe/enforce — both through * `setJevModeAction`, which rewrites `mode` and nothing else. "Off" there keeps * the file (`mode: "off"`): deleting it would leave nothing on this page to * switch back on — the only way to get the file back would be re-running @@ -87,7 +87,7 @@ const PROVIDERS = [ const MODES = [ { value: "enforce", label: "enforce — jev's verdict counts" }, - { value: "shadow", label: "shadow — log only, regex decides" }, + { value: "observe", label: "observe — log only, regex decides" }, { value: "off", label: "off — keep this config, don't ask jev" }, ] as const; @@ -279,10 +279,10 @@ export default function JevPanel({ initial }: { initial: JevSettingsView | null }, [form, token]); /** - * The FailproofAI Cloud switch: `off`, `shadow` or `enforce`, and nothing + * The FailproofAI Cloud switch: `off`, `observe` or `enforce`, and nothing * else about the file changes (see `setJevModeAction`). */ - const onMode = useCallback(async (mode: "off" | "shadow" | "enforce") => { + const onMode = useCallback(async (mode: "off" | "observe" | "enforce") => { setBusy(true); setProblem(null); try { @@ -372,8 +372,8 @@ export default function JevPanel({ initial }: { initial: JevSettingsView | null {view?.on - ? view.mode === "shadow" - ? "shadow: jev's answers are logged, the regex result is what gets enforced." + ? view.mode === "observe" + ? "observe: jev's answers are logged, the regex result is what gets enforced." : "enforce: jev's answers can clear a reviewable deny." : cloudRoute ? "through FailproofAI Cloud, on your org's plan — no endpoint or token of your own." @@ -483,9 +483,9 @@ export default function JevPanel({ initial }: { initial: JevSettingsView | null