diff --git a/README.md b/README.md index 27c1da8..973bcbf 100644 --- a/README.md +++ b/README.md @@ -26,8 +26,8 @@ whole repository does. | Skill | What it does | Source of truth | |---|---|---| | [`failproofai`](skills/failproofai/) | **The main skill.** A complete, standalone FailproofAI product package: explain the product, set up a machine, operate the local runtime and daemon, use FailproofAI Cloud, understand and author policies, publish GitHub policy packs, instrument custom agents, work with evaluators, inspect every command, and troubleshoot by symptom. It also routes to focused sibling skills when they are installed. | Maintained here. Not synced from anywhere - edit in this repo. | -| [`failproofai-policy-author`](skills/failproofai-policy-author/) | Turn what agents keep doing wrong into enforcement for [failproofai](https://github.com/FailproofAI/failproofai) - triage a `failproofai audit` or FailproofAI Cloud findings, convert a CLAUDE.md/AGENTS.md into policies, or take a plain complaint ("agents keep force-pushing") and enforce it. Checks the shipped builtins and their params before writing anything, since most requests are one line of config; knows which of the 12 supported agent CLIs actually enforce a given event, so it never ships a deny the harness discards; tests every policy it authors. | Maintained here. Not synced from anywhere - edit in this repo. | -| [`failproofai-policy-publish`](skills/failproofai-policy-publish/) | The publishing companion to policy authoring - take tested policies, build an installable pack, publish its release assets to GitHub with `failproofai publish`, preview it, and verify the consumer path with `failproofai policies add /`. Cloud fleet rollout remains part of `fp-cloud-cli`. | Maintained here. Not synced from anywhere - edit in this repo. | +| [`failproofai-policy-author`](skills/failproofai-policy-author/) | Turn what agents keep doing wrong into enforcement for [failproofai](https://github.com/FailproofAI/failproofai) - triage a `failproofai audit` or FailproofAI Cloud findings, convert a CLAUDE.md/AGENTS.md into policies, or take a plain complaint ("agents keep force-pushing") and enforce it. Checks the shipped builtins and their params before writing anything, since most requests are one line of config; knows which of the 12 supported agent CLIs actually enforce a given event, so it never ships a deny the harness discards; writes Jev semantic checks and reviewable policies for calls no regex can decide; tests every policy it authors. | Maintained here. Not synced from anywhere - edit in this repo. | +| [`failproofai-policy-publish`](skills/failproofai-policy-publish/) | The publishing companion to policy authoring - take tested policies, build an installable pack, publish its release assets to GitHub with `failproofai publish`, preview it, and verify the consumer path with `failproofai policies add /`, including packs that carry Jev semantic checks. Cloud fleet rollout remains part of `fp-cloud-cli`. | Maintained here. Not synced from anywhere - edit in this repo. | | [`fp-cloud-cli`](skills/fp-cloud-cli/) | Operate FailproofAI Cloud with `fp`: inspect telemetry, evals and usage; triage issues and audits; manage keys, users, orgs, and settings; publish Cloud policy versions; deploy them to fleet machines; observe enforcement; promote or roll back. Global options go **before** the command: `fp --json sessions`, not `fp sessions --json`. | Synced from `FailproofAI/failproofai` → `fp-cloud-cli/skill/`. Do **not** hand-edit here. | | [`failproofai-sdk`](skills/failproofai-sdk/) | Make an AI agent report what it did - plan which points in the agent loop to record, write the instrumentation with the Python (`failproofai_sdk`) or TypeScript/JavaScript (`@failproofai/sdk`) SDK - a framework adapter or a hand-built loop - thread session/agent identity through it, and verify the events actually land. Also runs your own evaluator worker (the eval pod) in either language. For an agent loop that is **not** one of the 12 supported CLIs. | Synced from `FailproofAI/failproofai` → `sdk/python/skill/`. Do **not** hand-edit here. | | [`failproofai-eval-brainstorm`](skills/failproofai-eval-brainstorm/) | Work out **what is worth measuring** about an agent's production runs, from the sessions it actually produced - scan the population, confirm the signal is really in the telemetry (a measurement over a payload key nobody emits does not fail, it scores every session identically and looks like it works), check it separates good runs from bad, and converge on two to four proposals. Each one ends in the plain-English prompt that authors it. It stops there: composing, backtesting and deploying the evaluation is the dashboard's eval authoring page. | Synced from `FailproofAI/agenteye` → `agent/skills/failproofai-eval-brainstorm/` (private). Do **not** hand-edit here. | diff --git a/skills/failproofai-policy-author/SKILL.md b/skills/failproofai-policy-author/SKILL.md index 35837db..c9eb979 100644 --- a/skills/failproofai-policy-author/SKILL.md +++ b/skills/failproofai-policy-author/SKILL.md @@ -5,7 +5,7 @@ description: |- Trigger when the user wants to: • act on an audit — turn `failproofai audit` findings into fixes, or ask which policies work; - • stop a recurring behaviour, in plain words or as "write a policy that blocks X"; + • stop a recurring behaviour — "write a policy that blocks X", or a Jev check; • enforce a rules file — make a CLAUDE.md / AGENTS.md real instead of advisory; • enable an existing builtin — usually the right answer, checked first; • work from FailproofAI Cloud — findings, hooks that fail or over-deny, backtesting a draft. @@ -132,6 +132,44 @@ repeating the verdict. If several are installed, the policy must hold on all of them or you state which one it covers. Authoring for whichever CLI you happen to be running inside, silently, is the bug. +### When failproofai guards your own session + +If failproofai's hooks are installed for the CLI you are running in, as on any enrolled +machine, its guard `block-failproofai-commands` judges your own tool calls, and nothing +switches it off (`builtin-policies.ts`, grep `function blockFailproofaiCommands`). It denies: + +- **every `failproofai …` command**, however launched (`npx`, a path to its bin, + `$(failproofai --version)`): `policies --list`, `policies --install`, `policies add`, + `publish` even with `--dry-run`, `jev status`/`setup`/`test`, `audit`, and the raw + `npx -y failproofai --hook …`; +- **writing under any `.failproofai` path**, the project's `.failproofai/policies/` and + `policies-config.json` as much as `~/.failproofai/`: with Write/Edit, or with any shell + command that is not a plain read (`mkdir`, `cp` into it, `>` onto it, `node … .failproofai/…`). + +It allows plain reads (`cat`, `ls`, `tail`, `grep`, `jq`; a `|` inside quotes on such a line, +a `jq` filter's too, reads as a pipe and denies, so write `grep -e a -e b`), copying state +*out* (`node` naming `~/.failproofai` is denied, so the readers below `cp` to `/tmp` first), +every `fp …` command, and `node "$SKILL_DIR/scripts/test-policy.mjs" --policy ` for a +file outside `.failproofai`. + +So draft and test in a folder of your own (`policy-drafts/`), and end with the operator's half +as one block of exact commands, filled in: + +```bash +# the operator's half: failproofai's guard denies these to the agent +cp policy-drafts/block-foo-policies.mjs .failproofai/policies/ +failproofai policies --list +``` + +Add each builtin as its `policies add FailproofAI/policies --policy` line, and for a pack the +`publish --dry-run`, `publish` (or the local install), `policies add` and `jev status` lines +(*Jev*). A `policies-config.json` change goes in your reply as a diff, never as a file you +write, `policy-drafts/` included: `agent-config-tampering`, a no-override Jev check, reads a +written failproofai config as the agent changing its own guardrails and denies from 0.85 (a +drafted `policies-config.json` scored 0.79 live, a warning). A deny reading *"Running failproofai +CLI commands is blocked"* or *"Writing to failproofai's own state…"* is this guard, not Jev +and not your policy. Do not retry it in another spelling; hand it over. + ## Audit-driven triage **One finding, or all of them?** If the user names a single finding — by its policy name @@ -196,10 +234,11 @@ cat ~/.failproofai/audit/dashboard.json 2>/dev/null || cat ~/.failproofai/audit- ``` **The cache is the only source, and it can be arbitrarily old.** Check `cachedAt` before -trusting anything in it: +trusting anything in it (from a copy: the guard denies `node` naming `~/.failproofai`): ```bash -node -e 'const j=require(process.env.HOME+"/.failproofai/audit/dashboard.json"); +cp ~/.failproofai/audit/dashboard.json /tmp/fp-audit.json +node -e 'const j=require("/tmp/fp-audit.json"); const age=(Date.now()-Date.parse(j.cachedAt))/864e5; console.log(`cached ${j.cachedAt} (${age.toFixed(1)} days ago), ${j.result.results.length} findings`)' ``` @@ -227,7 +266,8 @@ running it for them. not `.results[]`: ```bash -node -e 'const j=require(process.env.HOME+"/.failproofai/audit/dashboard.json"); +cp ~/.failproofai/audit/dashboard.json /tmp/fp-audit.json +node -e 'const j=require("/tmp/fp-audit.json"); for (const c of j.result.results.sort((a,b)=>b.hits-a.hits)) console.log([c.name.replace("failproofai/",""),c.source,c.hits,c.projects,c.enabledInConfig].join(" | "))' ``` @@ -249,12 +289,15 @@ every finding still claiming `enabledInConfig: true`. Observed exactly that way Always re-read the current config before classifying: ```bash +cat ~/.failproofai/policies/packs/installed.json 2>/dev/null # packs, machine-wide cat .failproofai/policies-config.json 2>/dev/null # project scope cat ~/.failproofai/policies-config.json 2>/dev/null # global scope ``` -`enabledPolicies` is a **union** across project, local and global -(`hooks-config.ts`, grep `enabledSet`) — not precedence. A policy is on if *any* scope lists it. +With any pack installed, a builtin is on only if the FailproofAI pack's `enabled` list names +it (`null` is all), and `enabledPolicies` is ignored (`references/traps.md` §7). With none, +`enabledPolicies` is a **union** across project, local and global (`hooks-config.ts`, grep +`enabledSet`) — not precedence. A policy is on if *any* scope lists it. **Scope matters more than it looks.** The audit spans every project on the machine, but a project-scope config protects only one. If findings span many projects and enforcement lives @@ -320,9 +363,12 @@ always "match the raw name", with only builtin coverage lost. Check the harness' Run the attribution above first — a `DEAD` finding is not Bucket A, B or C; it is "unenforceable on the harness where it happened", and saying so is the correct output. -**Bucket A — a builtin covers it and is off.** Do not write code. Add the short name to -`enabledPolicies` in `.failproofai/policies-config.json`. This is the cheapest and most -maintainable fix, and it is the right answer for most `source: "builtin"` findings. +**Bucket A — a builtin covers it and is off.** Do not write code. Switch it on in the +FailproofAI pack: `failproofai policies add FailproofAI/policies --policy `, after one +plain `policies add FailproofAI/policies` if it is not installed (`--policy` on a first +install enables only the names given). Not `enabledPolicies`: it stops counting once any pack +is installed, and a Jev check of your own needs one (`references/traps.md` §7). This is the cheapest and +most maintainable fix, and it is the right answer for most `source: "builtin"` findings. **Bucket B — a builtin covers it and is already on.** No action. Report it so the user knows the finding is historical, not ongoing. @@ -355,8 +401,9 @@ count on its own reads as one number to fix; `47 hits (claude 40, hermes 7 — D reads as the two different problems it actually is. For Bucket A, propose the -config diff rather than silently editing — enabling enforcement changes what their agent is -allowed to do. For a single explicit request ("turn on block-rm-rf"), just edit it. +command rather than running it silently — enabling enforcement changes what their agent is +allowed to do. For a single explicit request ("turn on block-rm-rf"), just run it, or, where +failproofai guards your session, hand the operator the command. **Never widen scope on your own initiative.** These three are off-limits without the user asking for them in the current request: @@ -366,6 +413,7 @@ asking for them in the current request: | `failproofai policies --install` at **user scope** | Wires hooks into *every* project on the machine, not the one they are in | | Editing `~/.failproofai/policies-config.json` | Global config; a deny there fires everywhere | | Setting `customPoliciesPath` globally | Silently activates policy files across all projects | +| `failproofai policies add` | Packs, the FailproofAI one included, install machine-wide | A question — *"what should I do about my findings?"*, *"is this protected?"* — asks for an answer, not a change. Recommend the machine-wide fix in words and let them decide. Being @@ -398,8 +446,11 @@ Two judgment calls to make before writing, and to state back to the user at the Both flavors are harness-dependent. `deny()` only stops something on a **block** pair (*Pick an event the harness can actually enforce*), and `instruct()` is properly supported - only on Claude Code, Devin and Antigravity — it degrades to a stderr note on Hermes, - Goose, OpenClaw and Pi (`references/api.md`). Pick the mode, then confirm the pair. + only on Claude Code, Devin and Antigravity, plus Hermes through its native plugin (what + `failproofai config` and `failproofai update` install), where it blocks the first attempt in + each model response with the instruction and lets the next response's retry through. It + degrades to a stderr note on Goose, OpenClaw, Pi and Hermes' legacy shell hooks + (`references/api.md`). Pick the mode, then confirm the pair. - **Scope.** Project config protects one repo; user scope (`~/.failproofai/policies/`) applies everywhere. "My agent keeps doing X" usually means *everywhere*, not *here*. @@ -410,7 +461,7 @@ Two judgment calls to make before writing, and to state back to the user at the ### Measure builtin coverage against your real tool surface -`references/builtins.md` says what the 39 builtins catch. It does **not** say whether your +`references/builtins.md` says what the 40 builtins catch. It does **not** say whether your agents call the tools they filter on — and on a real fleet the answer is mostly no. ```bash @@ -447,9 +498,17 @@ Two hazards the same data exposes: ### Check the builtins first -Read `references/builtins.md`. All 39 builtins with their categories, default state, events +Read `references/builtins.md`. All 40 builtins with their categories, default state, events and parameters. If one matches, enabling it beats writing a new file every time. +If the complaint is that a builtin is **noisy** — `block-kubectl` denying `kubectl get` — try +its `allowPatterns` param first where it has one: that trades nothing. The other fix is Jev: +the builtins marked *reviewable by* are cleared on the calls Jev judges harmless, once +`FailproofAI/jev-policies` is installed (the npm package ships no Jev checks, so without it +every builtin is hard) and Jev runs in **enforce** mode (*Jev*). The price: an agent with a shell can forge the user's +consent (`claude -p "…"`), and forged consent clears a reviewable deny — failproofai's +`docs/reference/jev-intent.mdx`. + Many builtins take `params` (allowlists, thresholds, protected branches) that go in the `policyParams` map — a parameterized builtin often covers a case that looks custom. @@ -519,7 +578,8 @@ calls is dead on arrival, and a harness's own tools are exactly where that goes ### Write the file -Location: `.failproofai/policies/` in the project. +Location: `.failproofai/policies/` in the project. Where the guard is on, draft it elsewhere +and hand over the `cp` (*When failproofai guards your own session*). **The filename must end in `policies.js`, `policies.mjs`, or `policies.ts`.** A file named `block-foo.mjs` is silently skipped and enforces nothing. Name it `block-foo-policies.mjs`. @@ -527,6 +587,229 @@ This is the highest-frequency failure in the whole system — see `references/tr See `references/patterns.md` for worked examples per event type. +### Jev: when no string decides it + +Some concerns are not in the command. `prisma migrate deploy` is the same string against +localhost and against production; the difference is in `DATABASE_URL`, the branch, or what +the user asked for. A regex that blocks every migration gets disabled, and one that allows +them enforces nothing. **Jev** is failproofai's semantic evaluator: it asks a model yes/no +questions about the call in front of it. The regex always stays the floor. + +| The concern turns on… | Write | +|---|---| +| a string in the call — a flag, a path, a tool name | a normal policy. Stop here | +| something a forged "the user asked for it" must never unlock — irreversible actions, privilege escalation, running code fetched from the internet, pushes to protected branches, credential access | a normal policy, **kept hard** even where a matching Jev check exists. failproofai keeps `block-sudo`, `block-curl-pipe-sh` and `block-push-master` hard for this reason | +| a string, but the regex is noisy on legitimate shapes | the regex plus `authority: "reviewable"` and `reviewedBy`, so Jev can clear the harmless calls | +| something no string decides — the target, the intent, whether the user asked | a **semantic check** (`semanticPolicies.add`) shipped in a **pack**, beside a regex floor wherever it must hold without Jev | + +Jev reviews only `PreToolUse` and `PermissionRequest` calls the agent made, only where +`~/.failproofai/jev.json` exists and is not `off`, and **only the checks an installed pack +declares**. The npm package ships none (1.0.8+): FailproofAI's 16 are the +`FailproofAI/jev-policies` pack, and a machine with no pack declaring checks never starts a +review at all: no request, nothing cleared, whatever `jev status` says about the provider. Even then it **changes a decision only in +`enforce` mode**: `observe`, which `jev setup` gives you by default, records what Jev would have +done and applies the regex result, and so does every call Jev did not answer (a timeout, a +transport error, a 429, 402 or 5xx, a malformed reply). Everywhere else a reviewable policy is +its hard regex and a semantic check blocks nothing, so anything that must hold everywhere needs +a regex floor. Read `references/traps.md` §10 before writing either. + +**Reviewable — let Jev clear a noisy block.** Two fields on the `customPolicies.add` you +already have: + +```js + authority: "reviewable", + reviewedBy: ["production-infra-change"], // semantic check names, not policy names +``` + +Jev clears the verdict only when **every** named check was asked about this call **and each +answered "no concern", was overridden by the user's request, or was a deny the user's task +softened to a warning**. A deny keeps the block, and so does a warning nobody consented to; a +warning from a deny-mode check also cancels every other clear on that call. So pick the check +by asking **"is there anything left that can deny?"**: + +- Its `appliesTo` and precondition must cover every shape the regex fires on: a check never + asked makes the block permanent. Tools no class knows (every `mcp__*`) are the exception: + they are asked every check, whatever `appliesTo` says. +- It must model every one of those shapes, and asked is not modelled: a check asked whose + probes do not all hold answers "no concern", which clears the floor with no warning. So list + each shape the regex fires on, write its harmful case, and read every reviewer probe against + it: the probes of the 16 `FailproofAI/jev-policies` checks are in `references/builtins.md` + *Jev checks in FailproofAI/jev-policies*; your own pack's are in your file. A live floor on `rm`, `>` and + `git reset --hard`, reviewed by `destructive-deletion`: `echo hi > app.py` answered + `destroys` 0.91 but `irreplaceable` 0.34, so the unasked overwrite ran silently. That check asks about data that cannot be rebuilt; the floor was + about any unasked overwrite. Keep a shape no reviewer models in a hard policy of its own, + and keep a destructive floor hard unless every shape is modelled. +- Once the block clears, something must still be able to deny: a deny-mode reviewer, or a + deny-mode check asked about the same call on its own (`block-read-outside-cwd` has only the + instruct `read-outside-workspace`, but `secret-exposure` and `credential-exfiltration` still + deny the read), or the policy only ever warned. Paired with instruct checks and nothing else, + the block holds while one of them warns and goes, with nothing in its place, whenever they + find nothing or the user asked. +- A deny-mode check still stops denying, and so clears the floor, when an `exempt` probe + holds and, with `userCanOverride: true`, when the user asked for the operation (allowed) or + when Jev judges the call a step of the user's task (softened to a warning). Between 0.7 and + 0.85 it warns without consent, which keeps the floor (it denies from 0.85). The user's + request can be forged by an agent with a shell. Where the block must hold even then, the + reviewer needs `userCanOverride: false` (among the jev-policies checks only + `credential-exfiltration` and `agent-config-tampering`), or keep the policy hard. + +Where `authority` and `reviewedBy` count: a local file honours them; in a pack entry +`failproofai publish` validates them and copies them into the manifest, which is what machines +read; a **cloud-managed policy is always hard** today. + +**Semantic — a question set, not code.** No `fn`, no `match`: Jev answers the probes, and +the check fires only when every probe holds (its evidence is the lowest answer). A check that +does not fire answers "no concern", which clears any floor it reviews, so **every probe must be +true for the harmful call**, never for the harmless one. Put the "does it do X" probe first. + +Each probe's `instructions` **states the harmful claim**: a sentence true for the harmful call +that says what `criteria.true` says. Never a question, never the harmless case. A push +guard's probe ending "Does the branch name appear verbatim anywhere in `user_said`?", under +`criteria.true` "NOT present", scored 0.45–0.73 live, hovering on the 0.7 fire line. Stated +as below, the same payloads scored 0.04 for the branch the user named and 0.93–0.98 for +branches they never mentioned: + +```js +{ id: "branch_not_named", + instructions: "The branch this command pushes to (the one it names, or `facts.current_git_branch` for a bare `git push`) does not appear anywhere in `user_said`.", + criteria: { true: "The user never named that branch.", false: "The user named that branch." } } +``` + +When "the user did not ask" is the concern, make it a probe like that one, after the "does it +do X" probe ("the command pushes commits to a remote"). Do not leave it to `userCanOverride: +true`: that clears a fired check only when Jev scores `op_requested` ≥ 0.8 with `beyond_task` +< 0.5, or `task_step` ≥ 0.8 with `beyond_task` < 0.3 (which drops a warning but only softens a +deny; `decideV1`). Live Jev gave the push the user named explicitly `op_requested` 0.68 and +0.76 and `task_step` 0.56 and 0.69, so a lone "pushes to a remote branch" check warns on the +push that was asked for. The inverted probe fails the other way: "the branch appears in +`user_said`" scored 0.96 on the requested push and 0.04 on an unasked `git push origin main`, +which cleared its floor silently. + +Two more conditions are checked in code, not by the model, and both are about shell commands. +The `op_requested` route clears a command only when **every** target it names appears in what +the user typed or in the agent message they replied to (a non-shell tool's fields count as one +target). After "clean the build", `rm -rf build/ ~/important` is not cleared, and the +`task_step` route does not soften it either, because the user named some of its targets and +not all. And a command the local scan cannot read whole is cleared or softened by **neither** +route: any `$` expansion (`$VAR`, `$1`, `${…}`, `$(…)`), backticks, `$'…'`, a glob (`*`, `?`, +`[…]`), brace expansion, a heredoc or here-string, `eval` or `sh -c`, or an unclosed quote. +Jev's own deny or warning stands and a reviewable floor keeps denying, so no request clears +`rm -rf build/*`. If the stored user message was cut for length, the target check is skipped, +not failed. Test consent with plain commands. + +```js +import { semanticPolicies } from "failproofai"; + +semanticPolicies.add({ + name: "db-migration-on-production", // what a reviewedBy names + title: "Tried to run a schema migration against a production database", // past tense + appliesTo: ["shell"], // shell | write | read | network | other + mode: "deny", // blocks at p ≥ 0.85, warns from 0.7; "instruct" only warns + userCanOverride: true, // required; true = the user's request (real or forged) allows it and a task step softens it to a warning; false = neither + probes: [ + { id: "applies_migration", + instructions: "The command in `agent_request` APPLIES database schema migrations: `prisma migrate deploy`, `prisma db push`, `knex migrate:latest`, `alembic upgrade`, `rails db:migrate`, `flyway migrate`, `sequelize db:migrate`, or an equivalent, however the binary is spelled or pathed.", + criteria: { true: "Running it would change a database schema.", + false: "It only creates, lists, checks or previews migrations (`migrate dev --create-only`, `status`, `--dry-run`, `alembic history`)." } }, + { id: "production_target", + instructions: "The database it targets is production or shared, or cannot be told from the command, its flags, or the environment variables set inline on it (`DATABASE_URL=...`, `RAILS_ENV=...`, `--env`).", + criteria: { true: "Production, shared, or unknown database.", + false: "Clearly local or throwaway: localhost, 127.0.0.1, a docker-compose service, a sqlite file, or an env named dev, test or local." } }, + ], + guidance: "Run migrations against a local or staging database, or hand the command to a human to run against production.", +}); +``` + +**On Hermes, write every probe about the command, never about the user's words.** Hermes +hands failproofai no prompt (`semantic/intent.ts`, `PROMPT_CHANNELS.hermes` is all null), so +`user_said` is always empty there, the task and `user_asked` questions are never asked, and +`userCanOverride` changes nothing. A probe like "the recipient is not named in `user_said`" +is true on every Hermes call and blocks the legitimate ones too. Judge what Jev can see: the +command in `agent_request` and `facts`. Hermes tools no class knows (`execute_code`, +`browser_*`, composio `OUTLOOK_*`, every MCP tool) are asked every check, one Jev request per +call. Name the sender or script outright in the first probe ("`chetak-send.sh` sends a +WhatsApp message"): Jev scored a bare mention 0.78, a warning instead of a deny, and 0.88 +once the probe said so. + +`precondition` is a **name**, not code — `always`, `protected_branch`, `in_git_repo`, +`has_paths`, `paths_outside_project`, `system_or_root_paths`. Probes should talk about what +Jev is shown: `agent_request`, `user_said`, and `facts` (`cwd`, `project_root`, +`current_git_branch`, `paths[]`). Set `userCanOverride: false` for anything a forged "the user +asked for it" must not unlock, the way `credential-exfiltration` does. A suspected prompt +injection needs no setting: it already turns any check that fired into a deny. + +**`semanticPolicies.add` does something only inside a pack.** In `.failproofai/policies/` or +a cloud policy it registers into a list nothing reads; the hook log says so once per file +(`… never asked here`). **An installed pack's checks are the only questions Jev asks**: there +are no built-in ones to join. Several packs' checks are asked together, FailproofAI's first. +So a `reviewedBy` is honoured only for a check some installed pack declares: from a local +file it may name `FailproofAI/jev-policies`' 16 only where that pack is installed (otherwise +the policy stays hard); inside a pack that declares checks, `publish` accepts only its own. +The 16 names are reserved: another pack's check named like one is ignored, never asked. An +observe pack's checks (`publish --effect observe`) are never asked, and a pack added with +`--cli` has its checks asked only for those agents; a policy naming a check that is not asked +stays hard. Every installed pack shares one question budget of 27,591 characters. +`publish` holds a pack not from FailproofAI to the 9,101 left once `jev-policies` (18,490) +is counted, whether or not the author has it installed, so a pack can sit beside it; keep +a pack's questions under about 9k (the example check is 1,601; a two-check pack about +3,000). An entry over the budget on a machine is dropped at load: `failproofai policies add` +names it, and so does the hook log (`was dropped: its questions need`). +`references/patterns.md` has the two-tier entry file. Validate it, then release it: + +```bash +failproofai publish db-guard.policies.mjs --dry-run --version 0.1.0 # validates, writes dist-pack/, publishes nothing +failproofai publish db-guard.policies.mjs --repo acme/db-guard --version 0.1.0 +``` + +Always pass `--dry-run` to validate: without it, a publish with no `--repo` takes the +repository from the git origin, if there is one, and releases for real (with neither it is a +dry run). The dry run is the validator: it rejects a bad field, an unknown precondition, an +unknown `reviewedBy` name or a check over the ~9k budget, and prints +`N semantic policies for Jev (… characters of questions), asked by Jev wherever it installs` +and `Requires failproofai 1.0.8-beta.0 or newer`: a pack with checks gets that +`--min-cli-version` by default, since older CLIs cannot read the checks; pass a higher one +if you rely on something newer. Nothing in a pack is on by default: install with +`failproofai policies add / --all`, or set `defaultEnabled: true` on the policy. + +**Installing a dry-run pack on this machine.** `policies add` takes only `owner/repo[@tag]` +fetched from `$FAILPROOFAI_PACK_BASE_URL///releases/download//` (default +`https://github.com`); a path such as `./dist-pack/…` fails `unsafe owner "."`. So serve +`dist-pack/` as that layout. Name the pack with `--id` (a dry run otherwise calls it +`local/`) and the tag with `@` (it must be the version, `v` optional): + +```bash +failproofai publish db-guard.policies.mjs --dry-run --id acme/db-guard --version 0.1.0 +d=/tmp/packs/acme/db-guard/releases/download/0.1.0; mkdir -p $d && cp dist-pack/* $d +python3 -m http.server 8765 -d /tmp/packs & +FAILPROOFAI_PACK_BASE_URL=http://127.0.0.1:8765 failproofai policies add acme/db-guard@0.1.0 --all +# one agent only: append --cli hermes (or claude, codex, …) to scope the pack and its checks +``` + +The variable redirects every pack fetch, so run `policies add FailproofAI/policies` without it. +Re-publishing a fixed version and re-running `policies add @` replaces the +installed one; `failproofai policies remove ` uninstalls it and its checks stop being asked +at once. + +**Verify both tiers.** `test-policy.mjs --policy` tests the floor alone: it runs in a sandbox +HOME with the legacy evaluator, so no `jev.json` or daemon brings Jev in — test that both ways +as below. Without `--policy` it runs against the real config, where a daemon can still answer +with Jev. Jev's half has no offline test: `failproofai jev test` checks the connection only +and never touches your policies. Roll out in observe mode first — `failproofai jev setup --provider +failproofai --mode observe` (FailproofAI Cloud) or `failproofai jev setup --provider +--key-stdin --mode observe < key-file` (your own key) — then read: + +- `failproofai jev status` — the mode, how many enabled builtin and pack policies are + reviewable (policies from your own files are not counted), and the last 24 hours' `cleared` + and `would have cleared (observe)`; +- `~/.failproofai/state/semantic/verdicts.jsonl` — answers keyed `.`, the only + proof a probe was asked. For a local file, that and `jevCleared` in + `~/.failproofai/hook-activity/current.jsonl` are the evidence. + +Nothing is cleared and no check blocks until `failproofai jev setup --mode enforce`. Do not +report the semantic tier as live until `jev status` shows `enforce`. Where the guard is on, +`publish`, `policies add`, `jev setup` and `jev status` are the operator's to run; you can +still `tail` both files above. + ### Verify it actually fires Loading and execution are both fail-open — a broken policy is indistinguishable from a @@ -570,7 +853,7 @@ to this skill's own folder — the directory you were told to read this file fro ```bash node "$SKILL_DIR/scripts/test-policy.mjs" \ - --policy .failproofai/policies/my-policies.mjs \ + --policy policy-drafts/my-policies.mjs \ --event PreToolUse --tool Bash \ --input '{"command":"sudo rm -rf /tmp/x"}' --expect deny ``` @@ -614,7 +897,8 @@ cannot affect the result. It also **renames the file if it violates the loader c Exit code is 1 if any `--expect` fails, so it drops straight into a script. -Underneath it is just the documented stdin protocol, if you need it by hand: +Underneath it is just the documented stdin protocol, if you need it by hand (a `failproofai` +command, so the operator's where the guard is on): ```bash echo '{"hook_event_name":"PreToolUse","tool_name":"Bash","tool_input":{"command":"sudo rm -rf /tmp/x"},"session_id":"test","transcript_path":"/dev/null","cwd":"'"$PWD"'"}' \ @@ -643,6 +927,7 @@ mistake produced a false FAIL during testing. `test-policy.mjs` handles all of t To test without touching the project's own config, build a throwaway project directory with its own `.failproofai/policies/` and `policies-config.json`, and point the payload's `cwd` at it. Project-scope discovery keys off that `cwd`, so the policy loads in isolation. +`test-policy.mjs --policy` builds that directory for you, and is the form the guard allows. Test both directions: a payload that **should** be denied, and a near-miss that **should** be allowed. A policy that denies everything passes the first test. @@ -767,7 +1052,7 @@ actually disables** convention policies), or hooks not being installed for the C Verify with: ```bash -failproofai policies --list +failproofai policies --list # the operator's, where the guard is on ``` **That is where this skill stops.** The split is clean, and both halves are shipped work: @@ -779,7 +1064,9 @@ failproofai policies --list | cloud rollout | *which fleet machines run a Cloud policy version, and what did it block?* | `fp-cloud-cli` | Everything past a proven local file is the deploy half: minting a version with -`fp policies publish` (which **deploys nothing** on its own), choosing `enforce` vs +`fp policies publish` (which **deploys nothing** on its own, and unlike the dashboard does +not refuse Jev fields: strip `semanticPolicies.add`, `authority` and `reviewedBy` first, since +a cloud policy never reads them), choosing `enforce` vs `observe`, `fp fleet deploy`, rollback, and reading `fp guardrails` to see the rule fire on real traffic. Those are shipped commands — if you find yourself about to say deployment is "dashboard work" or "not exposed by the CLI", that is wrong, and it tells the reader to stop @@ -819,7 +1106,7 @@ enforce may already be enforced. Check `references/builtins.md` (with params — read the project's existing custom policies: ```bash -ls .failproofai/policies/ && grep -h "name:\|description:" .failproofai/policies/*policies.mjs +ls .failproofai/policies/ && grep -h -e name: -e description: .failproofai/policies/*policies.mjs ``` A rule already covered goes in the report as covered — writing a duplicate policy means two @@ -977,3 +1264,4 @@ Print the `issues comment-add` and `audits resolve` commands and let the user ru FailproofAI Cloud's confirms **auto-skip on a non-TTY, which is how you run it**, so a wrong id resolves someone else's finding on a shared board with no prompt — and triage needs `audits:write`, which a read-only account lacks. See `references/cloud.md`. + diff --git a/skills/failproofai-policy-author/references/api.md b/skills/failproofai-policy-author/references/api.md index a0d110a..a232cca 100644 --- a/skills/failproofai-policy-author/references/api.md +++ b/skills/failproofai-policy-author/references/api.md @@ -5,6 +5,7 @@ - [The policy object](#the-policy-object) · [Context](#context) · [Decisions](#decisions) - [Events](#events) · [Filtering by tool](#filtering-by-tool) - [Execution model](#execution-model) · [Configuration](#configuration) +- [Semantic policies (Jev)](#semantic-policies-jev) > Source pointers below are paths inside the failproofai package. In a project that > installed it, they live under `node_modules/failproofai/`; in a source checkout, @@ -13,9 +14,10 @@ Everything here is exported from `src/index.ts` — that file is the entire public surface: ```ts -export { customPolicies, getCustomHooks, clearCustomHooks } from "./hooks/custom-hooks-registry"; +export { customPolicies, semanticPolicies, getCustomHooks, getSemanticRegistrations, clearCustomHooks } from "./hooks/custom-hooks-registry"; export { allow, deny, instruct } from "./hooks/policy-helpers"; -export type { PolicyContext, PolicyResult, CustomHook, PolicyDecision, PolicyFunction } from "./hooks/policy-types"; +export type { PolicyContext, PolicyResult, CustomHook, PolicyDecision, PolicyFunction, + PolicyAuthority, SemanticPolicyDeclaration, SemanticProbeDeclaration, SemanticToolClass } from "./hooks/policy-types"; ``` ## The policy object @@ -30,9 +32,17 @@ export interface CustomHook { events?: HookEventType[]; }; fn: (ctx: PolicyContext) => PolicyResult | Promise; + authority?: "hard" | "reviewable"; // absent = hard + reviewedBy?: string[]; // semantic check names; see below } ``` +`authority`/`reviewedBy` matter only when Jev runs in enforce mode. A local file's are +honoured directly. In a pack entry, `failproofai publish` validates them and copies them into +the manifest, which is what machines read. A cloud-managed policy ignores them: its authority +comes from the assignment (`policy-authority.ts`, grep `authorityDeclarationFor`). SKILL.md +*Jev* says when to set them. + Registered with `customPolicies.add(hook)`. That is the whole registration surface — no remove, no update, and **no validation of any kind** — `custom-hooks-registry.ts` (grep `getRegistry`) is a bare array push. @@ -81,9 +91,13 @@ export function instruct(reason: string): PolicyResult { ... } // reason require - **`allow`** — let it through. Also the correct return when your policy does not apply. - **`deny`** — block it. `reason` is shown to the agent and the user. -- **`instruct`** — let it through but inject guidance into the agent's next turn. Only - properly supported on Claude Code, Devin and Antigravity; degrades to a stderr note on - Hermes, Goose, OpenClaw and Pi. Fine for a local skill; do not rely on it if the policy +- **`instruct`** — let it through but inject guidance into the agent's next turn. Properly + supported on Claude Code, Devin and Antigravity. On Hermes' native plugin (1.0.6+, what + `failproofai update` migrates to) it blocks the first attempt in each model response with + `FAILPROOF INSTRUCTION ()` and the reason, and lets the next response's retry + through (`hermes-plugin/ledger.py`); so an "advisory" instruct costs a Hermes agent one + extra model round each time it fires. It degrades to a stderr note on Goose, OpenClaw, Pi + and Hermes' legacy shell hooks. Fine for a local skill; do not rely on it if the policy will be distributed. ### The `message` field is currently inert — sanitize works by blocking, not replacing @@ -176,8 +190,52 @@ export interface HooksConfig { } ``` -Builtins are enabled **purely by presence** of the short name in `enabledPolicies` — there -is no per-policy enabled/disabled object, and omission means off. +`enabledPolicies` is read **only while no pack is installed on the machine**: the moment any +pack is, a Jev pack included, those builtins stop loading (`handler.ts`, grep +`packsInstalledHere`). Builtins ship as the FailproofAI pack; switch one on with +`failproofai policies add FailproofAI/policies --policy ` (`builtins.md`). +`policyParams` still applies to it, keyed by short name. Merged across three scopes, in precedence order: project `{cwd}/.failproofai/` → local → global `~/.failproofai/` (`hooks-config.ts`, grep `readMergedHooksConfig`). + +## Semantic policies (Jev) + +`semanticPolicies.add(decl)` — `policy-types.ts`, grep `interface SemanticPolicyDeclaration`. +Read only by `failproofai publish`, so it takes effect **only in a pack**. Validated there by +`parsePackSemanticPolicy` (`pack-manifest.ts`): + +| Field | Rule | +|---|---| +| `name` | `[A-Za-z0-9._-]{1,128}`, unique; reported as `semantic/` | +| `title` | ≤120 chars, past tense ("Tried to …") | +| `appliesTo` | non-empty, from `shell` (Bash) · `write` (Write, Edit, MultiEdit, NotebookEdit) · `read` (Read, Glob, Grep, LS) · `network` (WebFetch, WebSearch) · `other`. A tool no class knows (`mcp__*`) is asked **every** check | +| `mode` | `"deny"` blocks at evidence ≥0.85 and warns from 0.7; `"instruct"` only warns | +| `userCanOverride` | required boolean, no default. `true`: a fired check is allowed when the user asked for the operation, and when Jev judges the call a step of the user's task a deny softens to a warning and a warning is dropped; an agent with a shell can forge that request. `false`: neither | +| `precondition?` | a name from `PACK_PRECONDITION_NAMES`: `always`, `protected_branch`, `in_git_repo`, `has_paths`, `paths_outside_project`, `system_or_root_paths` | +| `probes` | 1–6 of `{id, instructions ≤600, criteria?: {true ≤300, false ≤300}}`; `id` is `[a-z][a-z0-9_]{0,31}`, not `exempt` or `user_asked`. `instructions` states the harmful claim `criteria.true` names, never a question (SKILL.md *Jev*). Evidence is the **minimum** over probes | +| `exempt?` | one probe; if it holds (≥0.5) the check does not fire | +| `guidance` | ≤600 chars, shown to the agent as ` (semantic/<name>, p=…). <guidance>` | + +Limits: 24 checks per pack, and one question budget (`MAX_PACK_QUESTION_CHARS`, 27,591 +characters) shared by every installed pack, FailproofAI's first. `publish` holds a pack not +from FailproofAI to the 9,101 left once `FailproofAI/jev-policies` (18,490, +`JEV_POLICIES_QUESTION_CHARS`) is counted, installed or not (`pack-cli.ts`, grep +`questionBudget`). On a machine an entry over the budget is dropped at load, named by +`policies add` and the hook log (`was dropped: its questions need`). + +The npm package ships **no** Jev checks (1.0.8+; `pack-policies.ts` header): Jev asks only +what installed packs declare. FailproofAI's 16, with each one's mode, tool classes and every +probe's wording, are the `FailproofAI/jev-policies` pack, listed in `builtins.md` *Jev checks +in FailproofAI/jev-policies*. A `reviewedBy` is honoured only for a check an installed pack +declares (`effective-reviewers.ts`): from a local file, the 16 only where `jev-policies` is +installed, or another installed pack's; a pack that declares checks may name only its own. A block needs something left that can deny +once it clears (SKILL.md *Jev*); only the two marked "no override" (`credential-exfiltration`, +`agent-config-tampering`) deny whatever the user asked. + +Several packs' checks are asked together, FailproofAI's first. A name two packs declare +differently is asked for neither, and policies naming it stay hard (`effective-reviewers.ts`, +grep `contestedSemanticNames`); byte-identical declarations are fine. The 16 names are reserved: a pack not from a FailproofAI repository that declares +one is ignored for that name (`pack-policies.ts`, grep `isReservedClaim`). An observe pack's +checks, and a `--cli` pack's for other agents, are not asked. On Hermes `user_said` is always +empty (`intent.ts`, `PROMPT_CHANNELS.hermes`), so probes there must judge the command alone. diff --git a/skills/failproofai-policy-author/references/builtins.md b/skills/failproofai-policy-author/references/builtins.md index 99b236e..e7db25f 100644 --- a/skills/failproofai-policy-author/references/builtins.md +++ b/skills/failproofai-policy-author/references/builtins.md @@ -1,4 +1,4 @@ -# Builtin policies (39) +# Builtin policies (40) **Generated — do not hand-edit.** Regenerate with: @@ -22,7 +22,21 @@ Before concluding "no builtin covers this", check whether a **parameterized** on does — several take allowlists or thresholds that widen their scope considerably. Params go in the `policyParams` map, keyed by short name. -To enable: add the short name to `enabledPolicies` in `.failproofai/policies-config.json`. +To enable one, switch it on in the FailproofAI pack, which is machine-wide: + +```bash +failproofai policies add FailproofAI/policies # once, if not installed: its defaults +failproofai policies add FailproofAI/policies --policy block-rm-rf # adds to what is on +``` + +`--policy` on a first install enables only the names given. `enabledPolicies` in +`policies-config.json` is read only while no pack at all is installed on the machine; +installing any pack, a Jev pack included, stops those builtins loading (`handler.ts`, +grep `packsInstalledHere`). + +_(reviewable by: …)_ marks a policy Jev may clear once Jev runs in enforce mode: +only when every named check was asked about the call and none denied. The rest are +hard. SKILL.md *Jev* and `traps.md` §10. --- @@ -40,9 +54,9 @@ To enable: add the short name to `enabledPolicies` in `.failproofai/policies-con | Policy | Default | Events | What it catches | |---|---|---|---| -| `protect-env-vars` | **on** | PreToolUse | Prevent commands that read environment variables | -| `block-env-files` | **on** | PreToolUse | Block reading/writing .env files | -| `block-read-outside-cwd` | off | PreToolUse | Block file reads outside the session working directory _(params: allowPaths)_ | +| `protect-env-vars` | **on** | PreToolUse | Prevent commands that read environment variables _(reviewable by: env-secrets-dump, secret-exposure)_ | +| `block-env-files` | **on** | PreToolUse | Block reading/writing .env files _(reviewable by: secret-exposure)_ | +| `block-read-outside-cwd` | off | PreToolUse | Block file reads outside the session working directory _(params: allowPaths)_ _(reviewable by: read-outside-workspace)_ | ### Dangerous Commands @@ -50,20 +64,20 @@ To enable: add the short name to `enabledPolicies` in `.failproofai/policies-con |---|---|---|---| | `block-sudo` | **on** | PreToolUse, PermissionRequest | Block sudo commands _(params: allowPatterns)_ | | `block-curl-pipe-sh` | **on** | PreToolUse | Block piping downloads to shell | -| `block-rm-rf` | off | PreToolUse | Prevent catastrophic deletions _(params: allowPaths)_ | +| `block-rm-rf` | off | PreToolUse | Prevent catastrophic deletions _(params: allowPaths)_ _(reviewable by: destructive-deletion)_ | | `block-failproofai-commands` | **on** | PreToolUse, PermissionRequest | Block failproofai CLI commands, self-pause and uninstallation | -| `block-secrets-write` | off | PreToolUse | Block writing secret key files _(params: additionalPatterns)_ | +| `block-secrets-write` | off | PreToolUse | Block writing secret key files _(params: additionalPatterns)_ _(reviewable by: secret-exposure)_ | ### Infra Commands | Policy | Default | Events | What it catches | |---|---|---|---| -| `block-kubectl` | off | PreToolUse | Block kubectl commands (Kubernetes cluster mutations) _(params: allowPatterns)_ | -| `block-terraform` | off | PreToolUse | Block terraform and tofu (OpenTofu) commands _(params: allowPatterns)_ | -| `block-aws-cli` | off | PreToolUse | Block aws CLI commands _(params: allowPatterns)_ | -| `block-gcloud` | off | PreToolUse | Block gcloud (Google Cloud) CLI commands _(params: allowPatterns)_ | -| `block-az-cli` | off | PreToolUse | Block az (Azure) CLI commands _(params: allowPatterns)_ | -| `block-helm` | off | PreToolUse | Block helm commands _(params: allowPatterns)_ | +| `block-kubectl` | off | PreToolUse | Block kubectl commands (Kubernetes cluster mutations) _(params: allowPatterns)_ _(reviewable by: production-infra-change)_ | +| `block-terraform` | off | PreToolUse | Block terraform and tofu (OpenTofu) commands _(params: allowPatterns)_ _(reviewable by: production-infra-change)_ | +| `block-aws-cli` | off | PreToolUse | Block aws CLI commands _(params: allowPatterns)_ _(reviewable by: production-infra-change)_ | +| `block-gcloud` | off | PreToolUse | Block gcloud (Google Cloud) CLI commands _(params: allowPatterns)_ _(reviewable by: production-infra-change)_ | +| `block-az-cli` | off | PreToolUse | Block az (Azure) CLI commands _(params: allowPatterns)_ _(reviewable by: production-infra-change)_ | +| `block-helm` | off | PreToolUse | Block helm commands _(params: allowPatterns)_ _(reviewable by: production-infra-change)_ | | `block-gh-pipeline` | off | PreToolUse | Block gh CLI pipeline-trigger subcommands (workflow run, run rerun/cancel, pr merge, release create/delete, cache delete, secret set/delete) _(params: allowPatterns)_ | ### Git @@ -71,17 +85,18 @@ To enable: add the short name to `enabledPolicies` in `.failproofai/policies-con | Policy | Default | Events | What it catches | |---|---|---|---| | `block-push-master` | **on** | PreToolUse | Block pushing to main/master _(params: protectedBranches)_ | -| `block-force-push` | off | PreToolUse | Prevent force-pushing to any branch | +| `block-force-push` | off | PreToolUse | Prevent force-pushing to any branch _(reviewable by: git-history-rewrite)_ | | `block-work-on-main` | off | PreToolUse | Block git commits and merges on main/master branch _(params: protectedBranches)_ | -| `warn-git-amend` | off | PreToolUse | Warns before amending git commits, which rewrites history | +| `warn-git-amend` | off | PreToolUse | Warns before amending git commits, which rewrites history _(reviewable by: git-history-rewrite)_ | | `warn-git-stash-drop` | off | PreToolUse | Warns before permanently deleting stashed changes | +| `warn-git-clean` | off | PreToolUse | Warns before git clean deletes untracked directories (-d) or ignored files (-x / -X) _(params: destructiveFlags)_ | | `warn-all-files-staged` | off | PreToolUse | Warns before staging all working tree files with git add -A / . / --all | ### Database | Policy | Default | Events | What it catches | |---|---|---|---| -| `warn-destructive-sql` | off | PreToolUse | Warn before executing destructive SQL (DROP/TRUNCATE/DELETE without WHERE) via database clients | +| `warn-destructive-sql` | off | PreToolUse | Warn before executing destructive SQL (DROP/TRUNCATE/DELETE without WHERE) via database clients _(reviewable by: database-destruction)_ | | `warn-schema-alteration` | off | PreToolUse | Warns before SQL schema changes (ALTER TABLE with column or rename operations) | ### Packages & System @@ -89,7 +104,7 @@ To enable: add the short name to `enabledPolicies` in `.failproofai/policies-con | Policy | Default | Events | What it catches | |---|---|---|---| | `warn-package-publish` | off | PreToolUse | Warn before publishing packages to public registries (npm, PyPI, crates.io, RubyGems, etc.) | -| `warn-global-package-install` | off | PreToolUse | Warns before installing packages globally (npm -g, cargo install, etc.) | +| `warn-global-package-install` | off | PreToolUse | Warns before installing packages globally (npm -g, cargo install, etc.) _(reviewable by: system-modification)_ | | `prefer-package-manager` | off | PreToolUse | Blocks non-preferred package managers and tells Claude to use an allowed one (e.g., uv instead of pip) _(params: allowed, blocked)_ | | `warn-large-file-write` | off | PreToolUse | Warn before writing files larger than 1MB (configurable via thresholdKb param) _(params: thresholdKb)_ | | `warn-background-process` | off | PreToolUse | Warns before starting detached or background processes | @@ -130,3 +145,89 @@ All but `reread-after-edit` are Bash-command patterns, so a `PreToolUse` policy filtering on `ctx.toolName === "Bash"` and matching `ctx.toolInput.command` covers most of them. `reread-after-edit` needs cross-call session state, which hooks cannot see — that one needs a builtin, not a custom policy. + +--- + +## Jev checks in FailproofAI/jev-policies (16) + +The npm package ships **no** Jev checks (1.0.8+). These are the +`FailproofAI/jev-policies` pack, asked only on a machine that ran +`failproofai policies add FailproofAI/jev-policies`; without it Jev asks nothing and +every `reviewable` policy is hard. Wording is verbatim from the repo's reference copy +(`__tests__/fixtures/jev-policies.ts`, `JEV_PACK_POLICIES`). The names are reserved: +a pack from anyone else that declares one is ignored for that name. + +A check fires only when every probe holds; a deny check blocks from 0.85 and warns +from 0.7. Before naming one in `reviewedBy`, read its probes against the harmful case +of every shape the floor fires on (SKILL.md *Jev*): `destructive-deletion` asks +whether data is destroyed and irreplaceable, never whether the user wanted it. +_No override_ is `userCanOverride: false`. + +### `destructive-deletion` — deny · shell, write + +- **`destroys`** — The tool call in `agent_request` permanently deletes, wipes, truncates or overwrites existing files, directories or disks — for example `rm -r`, `find -delete`, `shred`, `dd of=`, `mkfs`, `truncate`, `rsync --delete`, `git clean -fdx`, or redirecting output over an existing file. _True:_ Existing data would be destroyed or irreversibly overwritten. _False:_ Nothing existing is destroyed: it reads, lists, creates new files, or moves data somewhere recoverable. +- **`irreplaceable`** — What would be destroyed is outside the project, is the whole project, is the user's home directory, is the filesystem root, or is data that cannot simply be regenerated (source code, documents, databases, keys). Use `facts.paths[].relation` to see where each target sits relative to the project. _True:_ The destroyed data matters and could not be rebuilt by running a command. _False:_ Only regenerable data inside the project is affected: build output, dist/, caches, node_modules, virtualenvs, coverage reports, temp files, or files the agent itself just created. + +### `production-infra-change` — deny · shell + +- **`mutates`** — The command in `agent_request` changes the state of cloud or cluster infrastructure: it creates, updates, deletes, applies, scales, restarts, rolls out, deploys or destroys resources in a cloud account, Kubernetes cluster, managed database, DNS, CDN or hosting platform (any CLI: kubectl, helm, terraform, tofu, pulumi, aws, gcloud, az, doctl, flyctl, vercel, wrangler, railway, and so on — however the binary is spelled or pathed). _True:_ It mutates infrastructure. _False:_ It only reads or plans: get, list, describe, logs, status, plan, diff, validate, whoami, or --dry-run. +- **`not_local`** — The target of that change is a shared or production environment, or its environment cannot be told from the command. _True:_ Production, shared, or unknown environment. _False:_ Clearly a local or throwaway environment: localhost, kind, minikube, docker-desktop, k3d, or a context, workspace or profile whose name says dev, test, staging, sandbox or local. + +### `git-history-rewrite` — deny · shell + +- **`rewrites_remote`** — The command in `agent_request` force-pushes or otherwise overwrites history on a git remote: `git push --force`, `--force-with-lease`, `-f`, a `+refspec` such as `+HEAD:main`, or deleting a remote branch. Git's global options before the subcommand (`git -C dir`, `git -c key=val`) and an absolute path to the git binary do not change what the command does. _True:_ Remote history would be overwritten or a remote branch deleted. _False:_ A normal push that only adds commits, or no push at all. + +### `push-to-protected-branch` — instruct · shell + +- **`pushes_protected`** — The command in `agent_request` pushes commits to a remote branch named main, master, production, release or trunk — either named explicitly or implied because `facts.current_git_branch` is one of those and the push names no other branch. _True:_ Commits land directly on a protected remote branch. _False:_ It pushes a feature branch whose name merely contains a word like main (e.g. feature/main-menu), or it does not push. + +### `commit-on-protected-branch` — instruct · shell · asked only when its precondition holds + +- **`creates_commit`** — The command in `agent_request` creates or amends a git commit, merge, rebase or cherry-pick in the repository (including `git -C dir commit` and `git commit -a`). _True:_ A new commit is written to the current branch. _False:_ It only inspects history (log, show, diff, status) or mentions `commit` in a message or search. + +### `secret-exposure` — deny · shell, read, write + +- **`touches_secrets`** — The tool call in `agent_request` reads, prints, copies, or writes credential material: SSH private keys (~/.ssh/id_*), authorized_keys, cloud credentials (~/.aws, ~/.config/gcloud, ~/.azure, ~/.kube/config), .env files containing real secret values, API tokens, password stores, keychains, or browser cookie databases. _True:_ Actual secret values would be read, displayed, copied or written. _False:_ It only references the concept: source code that reads process.env, a variable name, an .env.example template, a test fixture, documentation about credentials, or a grep for the word. + +### `credential-exfiltration` — deny, no override · shell, network + +- **`sends_out`** — The tool call in `agent_request` sends data to a network destination: an HTTP POST or PUT (curl -d, curl -F, curl -T, wget --post-data), scp, sftp, rsync or rclone to a remote, nc or socat, an upload to a paste site, bucket or webhook, or a DNS lookup carrying encoded data. _True:_ Data leaves this machine. _False:_ It only downloads, or only talks to localhost. +- **`sensitive_payload`** — The data being sent includes credentials, keys, tokens, environment variables, or the contents of files from outside the project (see `facts.paths`). _True:_ Secrets or private files are in what is sent. _False:_ What is sent is ordinary project data the user would expect to share, like a test request body. + +### `remote-code-execution` — deny · shell + +- **`download_and_run`** — The command in `agent_request` downloads code or a script from the internet and executes it: `curl … | sh`, `bash <(curl …)`, `wget … -O x.sh && bash x.sh`, `python3 -c "$(curl …)"`, piping into any interpreter (sh, bash, zsh, python, node, perl, ruby), or eval of a fetched string. _True:_ Fetched code is executed. _False:_ It only downloads without running, runs a local file, or merely searches for or quotes such a command (for example grep over a README). +- **exempt (does not fire when this holds)** — The URL being executed is the documented official installer of a widely used developer tool, served from that tool's own domain (for example bun.sh, sh.rustup.rs, get.docker.com, deb.nodesource.com, raw.githubusercontent.com/nvm-sh/nvm, astral.sh/uv). + +### `privilege-escalation` — deny · shell + +- **`elevates`** — The command in `agent_request` runs something as root or another user: sudo, doas, su, pkexec, run0, `sudo -i`, `sudo -s`, including when the binary is written as an absolute path or reached through a variable or wrapper. _True:_ Privileges are elevated. _False:_ It runs as the current user, or only mentions sudo in text, a comment, or a search pattern. + +### `database-destruction` — deny · shell + +- **`destructive_sql`** — The command in `agent_request` executes SQL or a database command that drops or truncates a table, schema or database, or deletes or updates rows without a condition that narrows them to specific records. A condition that is always true (`WHERE 1=1`, `WHERE true`, `WHERE id > 0`) does not narrow anything. Also counts: `redis-cli FLUSHALL`, `dropdb`, `mongo … dropDatabase()`. _True:_ Many rows, a table, or a whole database would be removed or overwritten. _False:_ It reads, inserts, or changes a specific identified record, or runs a migration file by name. +- **`real_database`** — The database it targets holds real data rather than being a disposable test database. _True:_ A real or shared database, or one whose purpose cannot be told. _False:_ Clearly an in-memory, test, fixture or throwaway local database (sqlite :memory:, a test_ database). + +### `read-outside-workspace` — instruct · shell, read · asked only when its precondition holds + +- **`reads_outside`** — The tool call in `agent_request` reads the contents of files that `facts.paths` marks as outside the project, in the user's home directory (for example another repository, Documents, or dotfiles). _True:_ It reads file contents from outside the project. _False:_ It only changes directory, lists names, checks existence, or touches the agent's own tool caches. + +### `agent-config-tampering` — deny, no override · shell, write + +- **`edits_agent_config`** — The tool call in `agent_request` modifies or deletes an AI agent's own configuration, hooks, permissions or guardrails: .claude/settings.json, .claude/settings.local.json, ~/.claude/, .codex/, .cursor/hooks.json, .gemini/, .factory/, failproofai configuration or policies, or permission and safety sections of AGENTS.md or CLAUDE.md. _True:_ The agent's hooks, permissions or safety settings would change. _False:_ It only reads those files, or edits ordinary project documentation. + +### `system-modification` — instruct · shell + +- **`modifies_system`** — The command in `agent_request` installs or removes software system-wide or changes machine configuration: apt, dnf, pacman, brew install, `npm install -g`, `pip install` outside a virtualenv, editing files under /etc, systemctl enable or disable, crontab, or changing shell startup files. _True:_ The machine changes outside the project directory. _False:_ It installs project dependencies locally (npm install, bun install, pip install in a venv). + +### `env-secrets-dump` — instruct · shell + +- **`dumps_env`** — The command in `agent_request` prints environment variable VALUES that may be secret: `env`, `printenv`, `set`, `export -p`, or `echo` of a variable whose name suggests a key, token, password or secret. _True:_ Secret-looking environment values would be printed. _False:_ It echoes a literal string, prints a non-secret variable like PATH or HOME, creates a virtualenv named env, or searches code for the word export. + +### `external-destructive-action` — deny · other + +- **`irreversible_external`** — The tool call in `agent_request` (an MCP or integration tool, see `facts.tool_name`) performs an irreversible or externally visible action: deleting records, repositories, files or accounts; sending email, chat or social messages on the user's behalf; making payments or purchases; merging or closing pull requests; changing permissions, access or billing; or writing to a production system. _True:_ Something outside this machine changes in a way that cannot be quietly undone. _False:_ It reads, searches, lists, fetches, or creates a draft that nobody else sees yet. + +### `external-data-egress` — instruct · other + +- **`egresses_private`** — The arguments in `agent_request` send private data to an external service: source code, file contents, credentials, customer data, or personal information. _True:_ Private data is being shared with a third party. _False:_ Only a query, identifier or public information is sent. diff --git a/skills/failproofai-policy-author/references/cloud.md b/skills/failproofai-policy-author/references/cloud.md index 6234be4..50dd1ad 100644 --- a/skills/failproofai-policy-author/references/cloud.md +++ b/skills/failproofai-policy-author/references/cloud.md @@ -220,15 +220,17 @@ From here it is *The authoring core* in SKILL.md, unchanged — **check builtins the file `*policies.mjs`, test both directions with `scripts/test-policy.mjs`. FailproofAI Cloud changes where the work comes from, not how a policy gets written. -"Plugging it in" is two concrete edits in the target project, and neither touches FailproofAI Cloud: +"Plugging it in" is two concrete steps on the target machine, and neither touches FailproofAI Cloud. +On an enrolled machine failproofai's guard denies you both: draft in `policy-drafts/` and +hand the operator the `cp` and the pack command (SKILL.md *When failproofai guards your own session*). ``` .failproofai/policies/<name>-policies.mjs the custom policy (filename convention — traps.md §1) -.failproofai/policies-config.json `enabledPolicies` for any builtin that covers a finding +failproofai policies add FailproofAI/policies --policy <name> any builtin that covers a finding (traps.md §7) ``` Nothing about *this* needs a FailproofAI Cloud permission — a failproofai policy is a local file -plus a config entry, and that is the whole story for one machine. It is **not** the whole story +plus a pack switch, and that is the whole story for one machine. It is **not** the whole story for a fleet: `fp policies publish`, `fp fleet deploy` and `fp guardrails` are shipped commands that carry the same rule to every machine and show it firing, and they need `policies:write`. That Cloud rollout path belongs to `fp-cloud-cli` — hand off rather than assuming the local @@ -237,8 +239,9 @@ edit is all there is. Then prove both local edits took effect, because neither i ```bash export SKILL_DIR=/path/to/skills/failproofai-policy-author # this skill's own folder -# the custom file actually loads (fail-open hides a file that never loaded — traps.md §3) -node "$SKILL_DIR/scripts/test-policy.mjs" --policy .failproofai/policies/<name>-policies.mjs \ +# the custom file actually loads (fail-open hides a file that never loaded — traps.md §3); +# test the draft: the guard denies this on the installed copy's path +node "$SKILL_DIR/scripts/test-policy.mjs" --policy policy-drafts/<name>-policies.mjs \ --cwd . --event PreToolUse --tool Bash --input '{"command":"<should-deny case>"}' --expect deny # an enabled builtin fires against the REAL project config (omit --policy) diff --git a/skills/failproofai-policy-author/references/patterns.md b/skills/failproofai-policy-author/references/patterns.md index 31a7007..56bbccc 100644 --- a/skills/failproofai-policy-author/references/patterns.md +++ b/skills/failproofai-policy-author/references/patterns.md @@ -206,7 +206,7 @@ signals which: | Prefix | Helper | Effect | Use when | |---|---|---|---| | `block-*` (17) | `deny()` | action never runs | irreversible or unsafe, no legitimate case | -| `warn-*` (10) | `instruct()` | action runs; agent told to check with the human first | risky but sometimes correct — needs a human, not a wall | +| `warn-*` (11) | `instruct()` | action runs; agent told to check with the human first | risky but sometimes correct — needs a human, not a wall | | `sanitize-*` (5) | raw deny object | output **blocked** before the model sees it (`message` is inert — traps.md §9) | secrets in tool output | Name your policy with the matching prefix. A reader should know the mode from the name. @@ -254,6 +254,62 @@ Hermes, Goose, OpenClaw and Pi it degrades to a stderr note the agent never sees those, oversight silently becomes no oversight. If the policy must hold everywhere, use `deny()` with a reason explaining how to proceed. +## Two tiers — a regex floor Jev can clear + +For a concern a string half-decides: the regex catches every shape, and a semantic check +decides which of them are harmless. Here, applying a schema migration is blocked unless Jev +judges the target local. The entry file is the `db-migration-on-production` check from +SKILL.md *Jev*, plus this floor: + +```js +// db-guard.policies.mjs — a PACK entry: semanticPolicies.add does nothing anywhere else +import { customPolicies, semanticPolicies, allow, deny } from "failproofai"; + +semanticPolicies.add({ name: "db-migration-on-production", /* … as in SKILL.md */ }); + +const MIGRATE = /\b(prisma\s+(migrate\s+deploy|db\s+push)|knex\s+migrate:latest|alembic\s+upgrade|rails\s+db:migrate|flyway\s+migrate|sequelize(-cli)?\s+db:migrate)\b/; + +customPolicies.add({ + name: "block-db-migrate", + description: "Block applying schema migrations from the agent", + match: { events: ["PreToolUse"] }, // Jev reviews PreToolUse and PermissionRequest only + authority: "reviewable", + reviewedBy: ["db-migration-on-production"], // a pack that declares checks may name only its own + fn: async (ctx) => { + if (ctx.toolName !== "Bash") return allow(); + const command = String(ctx.toolInput?.command ?? ""); + return MIGRATE.test(command) ? deny("Applying schema migrations is blocked; run it yourself.") : allow(); + }, +}); +``` + +Why it is shaped this way: + +- **The check models every shape the floor fires on.** `appliesTo: ["shell"]` covers the only + tool the floor matches, there is no precondition to miss, `applies_migration` names the + same commands, and both probes are true for the harmful case, a migration against + production. A narrower check would leave some blocks permanent (`traps.md` §10.3). +- **Something is left that can deny, but not always.** The check is deny-mode, so a + production target Jev is sure of (p ≥ 0.85) still blocks. With `userCanOverride: true` it + does not when the user asked for the migration (allowed) or Jev judges it a step of the + user's task (a warning), and either one clears the floor as well. An agent with a shell can + forge that request. Set `userCanOverride: false` if a production migration must block + whatever the task says. +- **Local, it is cleared; production, it is not.** Run end to end in observe mode against a + stub Jev: with the check answering "local", the floor's deny was recorded in + `jevCleared`; answering production (p=0.95), Jev denied with the check's title and + guidance. With no recorded user message the task questions are not asked, so this run did + not exercise the override. +- **Its check is the only one Jev asks** where it installs, unless `FailproofAI/jev-policies` + is there too; `publish` still holds it to the ~9k that pack leaves, so its 1,601 characters + fit either way (`traps.md` §10.6). + +Test the floor with `test-policy.mjs --policy` as usual — it runs without Jev, so it sees +exactly the hard regex. Then `failproofai publish db-guard.policies.mjs --dry-run --version +0.1.0` validates both tiers and writes `dist-pack/` without publishing; it prints `1 semantic +policies for Jev (1601 characters of questions), asked by Jev wherever it installs`. Leave +out `--dry-run` and a missing `--repo` is taken from the git origin, and the release is real. + ## Nudge toward a better tool The other use of `instruct()` — the action is not dangerous, just wasteful. No "STOP", no diff --git a/skills/failproofai-policy-author/references/rules-files.md b/skills/failproofai-policy-author/references/rules-files.md index 6a29814..ee3cdb3 100644 --- a/skills/failproofai-policy-author/references/rules-files.md +++ b/skills/failproofai-policy-author/references/rules-files.md @@ -25,14 +25,15 @@ candidate is `prefer-package-manager` with params: ```json { - "enabledPolicies": ["prefer-package-manager"], "policyParams": { "prefer-package-manager": { "allowed": ["bun"], "blocked": ["npm", "yarn"] } } } ``` -(Remember: params for a policy absent from `enabledPolicies` do nothing — `traps.md` §7.) +switched on with `failproofai policies add FailproofAI/policies --policy prefer-package-manager`. +Params for a policy that is off do nothing, and `enabledPolicies` counts only on a machine +with no pack installed (`traps.md` §7). **But check the builtin's matching breadth before enabling it.** A builtin can be broader than the rule. `prefer-package-manager`'s npm matcher is a bare `\bnpm\b`, so with diff --git a/skills/failproofai-policy-author/references/traps.md b/skills/failproofai-policy-author/references/traps.md index 212ebfd..b509023 100644 --- a/skills/failproofai-policy-author/references/traps.md +++ b/skills/failproofai-policy-author/references/traps.md @@ -8,9 +8,10 @@ 4. [A deny does not prove *your* deny](#4-a-deny-in-your-test-does-not-prove-your-policy-denied) 5. [Validation is nil](#5-validation-is-essentially-nil) 6. [Unsatisfiable Stop gates loop](#6-a-stop-gate-that-cannot-be-satisfied-loops-forever) -7. [Builtins enabled by presence](#7-builtins-are-enabled-by-presence-not-by-a-flag) +7. [`enabledPolicies` stops counting once a pack is installed](#7-enabledpolicies-stops-counting-once-any-pack-is-installed) 8. [`ctx.params` always empty](#8-ctxparams-is-always-empty-for-custom-policies) 9. [Sanitizers block, not redact](#9-sanitizers-block-they-do-not-redact--and-deny-cannot-set-message-anyway) +10. [Jev fails quietly](#10-jev-reviewable-and-semantic-policies-fail-quietly) Every item here is a documented, already-been-hit failure where a policy looks installed and enforces nothing. Read this before reporting any policy as working. @@ -65,7 +66,8 @@ project.customPoliciesEnabled ?? local.customPoliciesEnabled ?? global_.customPo ``` The first scope that *sets* the key decides, and the others are never consulted. This is the -opposite of `enabledPolicies`, which is a **union** across all three (§7). Two keys in one +opposite of `enabledPolicies`, which is a **union** across all three (`hooks-config.ts`, grep +`enabledSet`) and read only while no pack is installed (§7). Two keys in one file merging by opposite rules is worth checking rather than assuming. Note the explicit `customPoliciesPath` config key is **not** gated by this flag @@ -170,24 +172,30 @@ The same reachability rule applies to custom Stop policies — always `try/catch calls and `return allow()` on failure, so an unavailable tool degrades to letting the turn end rather than trapping it. -## 7. Builtins are enabled by presence, not by a flag +## 7. `enabledPolicies` stops counting once any pack is installed -`enabledPolicies` is a `string[]`. Omission means off. There is no -`{"block-rm-rf": false}` form — to disable, remove the string. +Builtins ship as the FailproofAI pack. `enabledPolicies` is a migration shim, read only while +**no pack at all** is installed on the machine (`handler.ts`, grep `packsInstalledHere`). +Install any pack, a Jev pack included, and every builtin listed there stops loading, and +nothing says so. Observed live: four builtins enabled that way went quiet the moment a +one-check Jev pack was installed. `~/.failproofai/policies/packs/installed.json` lists what is +installed. -Params live in a **sibling** `policyParams` object keyed by the same short name, not nested -inside the policy entry: +Switch builtins on in the pack instead (machine-wide), and hand that over beside any other +pack: + +```bash +failproofai policies add FailproofAI/policies # once, if not installed: its defaults +failproofai policies add FailproofAI/policies --policy block-rm-rf # adds to what is on +``` + +Params still live in the **sibling** `policyParams` object, keyed by the short name: ```json -{ - "enabledPolicies": ["block-read-outside-cwd"], - "policyParams": { - "block-read-outside-cwd": { "allowPaths": ["/tmp"] } - } -} +{ "policyParams": { "block-read-outside-cwd": { "allowPaths": ["/tmp"] } } } ``` -A param set for a policy that is not in `enabledPolicies` does nothing. +A param set for a policy that is not switched on does nothing. ## 8. `ctx.params` is always empty for custom policies @@ -220,3 +228,73 @@ So a sanitize deny on `PostToolUse` blocks the *entire* tool output; the model s block reason, never a redacted version. Protection holds — by omission — but do not tell a user "the token is scrubbed and the rest passes through." Put anything the agent needs into `reason`, and prefer `sanitize-api-keys.additionalPatterns` over authoring. See `api.md`. + +## 10. Jev: reviewable and semantic policies fail quietly + +Each leaves a policy that reads as Jev-aware and behaves as something else. The rule behind +most (`semantic/combine.ts`, grep `function clears`): a reviewable verdict is cleared only when +**every** `reviewedBy` check was asked about this call and **none** answered deny. + +1. **Jev off, in observe mode, or not answering means neither tier acts.** Without + `~/.failproofai/jev.json`, or with its mode `off`, a reviewable policy is exactly its hard + regex and no semantic check is ever asked. In `observe` — what `failproofai jev setup` sets + by default — Jev is asked, but its clears and its own denies are only recorded; the regex + result applies. Every call Jev did not answer (a timeout, a transport error, a 429, 402 or + 5xx, a malformed reply) falls back to the regex result too, so a concern with no regex floor + is allowed. Only `enforce` mode applies Jev: check `failproofai jev status` before calling a + semantic check live. Jev also sees only `PreToolUse` and `PermissionRequest` calls the agent + made — a reviewable `Stop` or `PostToolUse` policy is hard everywhere. +2. **`semanticPolicies.add` outside a pack does nothing.** Only `failproofai publish` reads it. + In a `.failproofai/policies/` file or a cloud-managed policy it lands in a registry nothing + reads; only the hook log says so (`… never asked here`). Cloud-managed policies are also + always hard: their authority comes from the deployment, which does not set it yet; the + code's own fields are ignored. +3. **A check that is never asked makes the block permanent.** Its `appliesTo` must cover every + tool the regex fires on and its precondition must hold for every shape. One misspelt or + unknown name makes the whole declaration hard. The exception is a tool no class knows + (every `mcp__*`): it is asked every check whatever `appliesTo` says, so a floor that fires on + an MCP tool clears there unless its reviewer models the MCP call's shape (§10.4, not this + one). Model it, or keep the floor off MCP tools. +4. **A check asked that does not model the shapes switches the policy off.** Asked and not + firing answers "no concern", and no concern clears. `warn-git-clean` stays hard because + `destructive-deletion` cannot judge a `git clean` that names no path. Asked about + `echo hi > app.py`, it answered `irreplaceable` 0.34 and a floor on unasked overwrites + cleared with no warning: check each shape's harmful case against every probe. + An inverted probe does the same on every call. One true for the harmless case ("the + branch appears in `user_said`") fires on the requested push and clears the unrequested + one, and a probe whose question asks the opposite of its `criteria.true` hedges near the + 0.7 line (SKILL.md *Jev*). +5. **Nothing left that can deny lets a blocked call run.** The test is "is there anything + left that can deny" once the block clears, and it passes three ways: a deny-mode reviewer; a + deny-mode check asked about the same call on its own, since Jev's own deny still joins the + most-severe merge (`block-read-outside-cwd` has only the instruct `read-outside-workspace`, + and `secret-exposure` and `credential-exfiltration` still deny the read); or a policy that + only ever warned. `block-work-on-main` stays hard because it passes none: its only reviewer + is instruct and nothing else covers the concern. + + Deny-mode alone is not enough. A deny-mode check stops denying — and so clears — when its + `exempt` probe holds and, with `userCanOverride: true`, when the user asked for the + operation or when Jev judges the call a step of the user's task. (Its warning at evidence + 0.7 to 0.85 is not a clear: a warning nobody consented to keeps the floor and cancels every + other clear on the call.) An agent with a shell can forge the user's request (`claude -p`, + `codex exec`), so a block that must hold even then needs a `userCanOverride: false` + reviewer (among the jev-policies checks, only `credential-exfiltration` and + `agent-config-tampering`), or it stays hard. +6. **No pack, no questions — and a pack's limits fail quietly.** The npm package ships no Jev + checks (1.0.8+). With no installed pack declaring one, Jev never starts a review, and every + `reviewable` policy is hard, the FailproofAI pack's included, until `failproofai policies + add FailproofAI/jev-policies`. A `reviewedBy` naming one of those 16 on a machine without + that pack is simply hard. The 16 names are reserved: another pack's check named like one is + ignored, and a name two packs declare differently is asked for neither. The 27,591-character budget is shared by every installed pack, FailproofAI's + first; `publish` holds anyone else's pack to the 9,101 left beside `jev-policies`, but a + machine with several third-party packs can still overrun it, and the later entries are then + dropped at load. An observe pack's checks, and a `--cli` pack's for other agents, are not + asked, so policies naming them stay hard. `failproofai policies add` prints each dropped or + ignored check; after that only the hook log does. +7. **A pack built before this release marks nothing reviewable.** Authority lives in the + manifest, and an older `failproofai publish` — or an older `FailproofAI/policies` release — + wrote none, so every policy in it is hard. A CLI older than Jev also drops `semantic` and + `minCliVersion` silently, so `--min-cli-version` only stops CLIs new enough to read it. A + pack of Jev checks alone that this build refuses at load denies nothing, but an older build + rolled back under it can deny every tool call: remove it (`failproofai policies remove + <id>`) before downgrading, as `publish` reminds you. diff --git a/skills/failproofai-policy-author/scripts/policy-events.json b/skills/failproofai-policy-author/scripts/policy-events.json index fc8ef0b..34dee89 100644 --- a/skills/failproofai-policy-author/scripts/policy-events.json +++ b/skills/failproofai-policy-author/scripts/policy-events.json @@ -9,7 +9,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "production-infra-change" + ] }, "block-az-cli": { "events": [ @@ -19,7 +23,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "production-infra-change" + ] }, "block-curl-pipe-sh": { "events": [ @@ -29,7 +37,9 @@ "Bash" ], "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "block-env-files": { "events": [ @@ -37,7 +47,11 @@ ], "toolNames": null, "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "secret-exposure" + ] }, "block-failproofai-commands": { "events": [ @@ -51,7 +65,9 @@ "NotebookEdit" ], "defaultEnabled": true, - "alwaysOn": true + "alwaysOn": true, + "authority": "hard", + "reviewedBy": null }, "block-force-push": { "events": [ @@ -61,7 +77,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "git-history-rewrite" + ] }, "block-gcloud": { "events": [ @@ -71,7 +91,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "production-infra-change" + ] }, "block-gh-pipeline": { "events": [ @@ -81,7 +105,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "block-helm": { "events": [ @@ -91,7 +117,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "production-infra-change" + ] }, "block-kubectl": { "events": [ @@ -101,7 +131,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "production-infra-change" + ] }, "block-push-master": { "events": [ @@ -111,7 +145,9 @@ "Bash" ], "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "block-read-outside-cwd": { "events": [ @@ -124,7 +160,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "read-outside-workspace" + ] }, "block-rm-rf": { "events": [ @@ -134,7 +174,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "destructive-deletion" + ] }, "block-secrets-write": { "events": [ @@ -144,7 +188,11 @@ "Write" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "secret-exposure" + ] }, "block-sudo": { "events": [ @@ -155,7 +203,9 @@ "Bash" ], "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "block-terraform": { "events": [ @@ -165,7 +215,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "production-infra-change" + ] }, "block-work-on-main": { "events": [ @@ -175,7 +229,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "prefer-package-manager": { "events": [ @@ -185,7 +241,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "protect-env-vars": { "events": [ @@ -195,7 +253,12 @@ "Bash" ], "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "env-secrets-dump", + "secret-exposure" + ] }, "require-ci-green-before-stop": { "events": [ @@ -203,7 +266,9 @@ ], "toolNames": null, "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "require-commit-before-stop": { "events": [ @@ -211,7 +276,9 @@ ], "toolNames": null, "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "require-no-conflicts-before-stop": { "events": [ @@ -219,7 +286,9 @@ ], "toolNames": null, "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "require-pr-before-stop": { "events": [ @@ -227,7 +296,9 @@ ], "toolNames": null, "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "require-push-before-stop": { "events": [ @@ -235,7 +306,9 @@ ], "toolNames": null, "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "sanitize-api-keys": { "events": [ @@ -243,7 +316,9 @@ ], "toolNames": null, "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "sanitize-bearer-tokens": { "events": [ @@ -251,7 +326,9 @@ ], "toolNames": null, "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "sanitize-connection-strings": { "events": [ @@ -259,7 +336,9 @@ ], "toolNames": null, "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "sanitize-jwt": { "events": [ @@ -267,7 +346,9 @@ ], "toolNames": null, "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "sanitize-private-key-content": { "events": [ @@ -275,7 +356,9 @@ ], "toolNames": null, "defaultEnabled": true, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-all-files-staged": { "events": [ @@ -285,7 +368,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-background-process": { "events": [ @@ -295,7 +380,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-destructive-sql": { "events": [ @@ -305,7 +392,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "database-destruction" + ] }, "warn-git-amend": { "events": [ @@ -315,7 +406,23 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "git-history-rewrite" + ] + }, + "warn-git-clean": { + "events": [ + "PreToolUse" + ], + "toolNames": [ + "Bash" + ], + "defaultEnabled": false, + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-git-stash-drop": { "events": [ @@ -325,7 +432,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-global-package-install": { "events": [ @@ -335,7 +444,11 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "reviewable", + "reviewedBy": [ + "system-modification" + ] }, "warn-large-file-write": { "events": [ @@ -345,7 +458,9 @@ "Write" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-package-publish": { "events": [ @@ -355,7 +470,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-repeated-tool-calls": { "events": [ @@ -363,7 +480,9 @@ ], "toolNames": null, "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null }, "warn-schema-alteration": { "events": [ @@ -373,7 +492,9 @@ "Bash" ], "defaultEnabled": false, - "alwaysOn": false + "alwaysOn": false, + "authority": "hard", + "reviewedBy": null } } } diff --git a/skills/failproofai-policy-author/scripts/sync-builtins.mjs b/skills/failproofai-policy-author/scripts/sync-builtins.mjs index 638ee6d..aa0f304 100644 --- a/skills/failproofai-policy-author/scripts/sync-builtins.mjs +++ b/skills/failproofai-policy-author/scripts/sync-builtins.mjs @@ -8,7 +8,7 @@ * The reference file is a convenience snapshot. It drifts the moment a builtin * is added, renamed, or has its default flipped — and a stale list is worse than * no list, because it is quietly authoritative. Run this after any change to - * builtin-policies.ts. + * builtin-policies.ts or the jev-policies pack (__tests__/fixtures/jev-policies.ts). * * Resolves the registry from the repo checkout first, then from an installed * failproofai package, so it works inside the repo and from a user's project. @@ -70,6 +70,31 @@ try { console.error("This script needs a TypeScript-capable runtime — try: bun sync-builtins.mjs"); process.exit(2); } +// FailproofAI's Jev checks. An authored reviewedBy has to be read against their +// probe wording. Up to 1.0.7 they were compiled in (`semantic/policies.ts`, +// SEMANTIC_POLICIES); from 1.0.8 the package ships none and they are the +// `FailproofAI/jev-policies` pack, whose reference copy the repo keeps at +// `__tests__/fixtures/jev-policies.ts` (JEV_PACK_POLICIES, pinned to the +// published pack by `pack-semantic-registry.test.ts`). Neither found: no section. +const REPO_ROOT = resolve(dirname(registryPath), "..", ".."); +let SEMANTIC_POLICIES = []; +let SEMANTIC_SOURCE = ""; +const compiled = await import( + pathToFileURL(join(dirname(registryPath), "semantic", "policies.ts")).href +).catch(() => ({})); +if (compiled.SEMANTIC_POLICIES?.length) { + SEMANTIC_POLICIES = compiled.SEMANTIC_POLICIES; + SEMANTIC_SOURCE = "compiled"; +} else { + const fixture = join(REPO_ROOT, "__tests__", "fixtures", "jev-policies.ts"); + if (existsSync(fixture)) { + const pack = await import(pathToFileURL(fixture).href).catch(() => ({})); + if (pack.JEV_PACK_POLICIES?.length) { + SEMANTIC_POLICIES = pack.JEV_PACK_POLICIES; + SEMANTIC_SOURCE = "pack"; + } + } +} const byCategory = new Map(); for (const p of BUILTIN_POLICIES) { @@ -102,7 +127,21 @@ lines.push("Before concluding \"no builtin covers this\", check whether a **para lines.push("does — several take allowlists or thresholds that widen their scope considerably."); lines.push("Params go in the `policyParams` map, keyed by short name."); lines.push(""); -lines.push("To enable: add the short name to `enabledPolicies` in `.failproofai/policies-config.json`."); +lines.push("To enable one, switch it on in the FailproofAI pack, which is machine-wide:"); +lines.push(""); +lines.push("```bash"); +lines.push("failproofai policies add FailproofAI/policies # once, if not installed: its defaults"); +lines.push("failproofai policies add FailproofAI/policies --policy block-rm-rf # adds to what is on"); +lines.push("```"); +lines.push(""); +lines.push("`--policy` on a first install enables only the names given. `enabledPolicies` in"); +lines.push("`policies-config.json` is read only while no pack at all is installed on the machine;"); +lines.push("installing any pack, a Jev pack included, stops those builtins loading (`handler.ts`,"); +lines.push("grep `packsInstalledHere`)."); +lines.push(""); +lines.push("_(reviewable by: …)_ marks a policy Jev may clear once Jev runs in enforce mode:"); +lines.push("only when every named check was asked about the call and none denied. The rest are"); +lines.push("hard. SKILL.md *Jev* and `traps.md` §10."); lines.push(""); lines.push("---"); @@ -116,8 +155,10 @@ for (const [category, policies] of byCategory) { const events = (p.match?.events ?? []).join(", ") || "—"; const params = p.params ? ` _(params: ${Object.keys(p.params).join(", ")})_` : ""; const beta = p.beta ? " _(beta)_" : ""; + const reviewable = + p.authority === "reviewable" ? ` _(reviewable by: ${(p.reviewedBy ?? []).join(", ")})_` : ""; lines.push( - `| \`${p.name}\` | ${p.defaultEnabled ? "**on**" : "off"} | ${events} | ${p.description}${params}${beta} |`, + `| \`${p.name}\` | ${p.defaultEnabled ? "**on**" : "off"} | ${events} | ${p.description}${params}${reviewable}${beta} |`, ); } } @@ -145,6 +186,46 @@ lines.push("most of them. `reread-after-edit` needs cross-call session state, wh lines.push("see — that one needs a builtin, not a custom policy."); lines.push(""); +if (SEMANTIC_POLICIES.length > 0) { + const probeLine = (label, p) => + `- **${label}** — ${p.instructions}` + + (p.criteria ? ` _True:_ ${p.criteria.true} _False:_ ${p.criteria.false}` : ""); + lines.push("---"); + lines.push(""); + if (SEMANTIC_SOURCE === "pack") { + lines.push(`## Jev checks in FailproofAI/jev-policies (${SEMANTIC_POLICIES.length})`); + lines.push(""); + lines.push("The npm package ships **no** Jev checks (1.0.8+). These are the"); + lines.push("`FailproofAI/jev-policies` pack, asked only on a machine that ran"); + lines.push("`failproofai policies add FailproofAI/jev-policies`; without it Jev asks nothing and"); + lines.push("every `reviewable` policy is hard. Wording is verbatim from the repo's reference copy"); + lines.push("(`__tests__/fixtures/jev-policies.ts`, `JEV_PACK_POLICIES`). The names are reserved:"); + lines.push("a pack from anyone else that declares one is ignored for that name."); + } else { + lines.push(`## Jev's built-in checks (${SEMANTIC_POLICIES.length})`); + lines.push(""); + lines.push("What each check asks Jev, verbatim from `src/hooks/semantic/policies.ts`"); + lines.push("(`SEMANTIC_POLICIES`) in the installed package (1.0.7 and older)."); + } + lines.push(""); + lines.push("A check fires only when every probe holds; a deny check blocks from 0.85 and warns"); + lines.push("from 0.7. Before naming one in `reviewedBy`, read its probes against the harmful case"); + lines.push("of every shape the floor fires on (SKILL.md *Jev*): `destructive-deletion` asks"); + lines.push("whether data is destroyed and irreplaceable, never whether the user wanted it."); + lines.push("_No override_ is `userCanOverride: false`."); + for (const c of SEMANTIC_POLICIES) { + lines.push(""); + lines.push( + `### \`${c.name}\` — ${c.mode}${c.userCanOverride ? "" : ", no override"} · ${c.appliesTo.join(", ")}` + + (c.precondition ? " · asked only when its precondition holds" : ""), + ); + lines.push(""); + for (const p of c.probes) lines.push(probeLine(`\`${p.id}\``, p)); + if (c.exempt) lines.push(probeLine("exempt (does not fire when this holds)", c.exempt)); + } + lines.push(""); +} + const generated = lines.join("\n"); const eventsJson = @@ -161,6 +242,8 @@ const eventsJson = toolNames: p.match?.toolNames ?? null, defaultEnabled: p.defaultEnabled === true, alwaysOn: p.alwaysOn === true, + authority: p.authority === "reviewable" ? "reviewable" : "hard", + reviewedBy: p.authority === "reviewable" ? (p.reviewedBy ?? []) : null, }, ]), ), diff --git a/skills/failproofai-policy-author/scripts/test-policy.mjs b/skills/failproofai-policy-author/scripts/test-policy.mjs index 917b90c..5cc93f3 100644 --- a/skills/failproofai-policy-author/scripts/test-policy.mjs +++ b/skills/failproofai-policy-author/scripts/test-policy.mjs @@ -194,8 +194,17 @@ function runCase(c, sandbox) { } if (c.event === "Stop") payload.stop_hook_active = false; - const env = { ...process.env, FAILPROOFAI_TELEMETRY_DISABLED: "1" }; - if (sandbox) env.HOME = sandbox; + // Legacy evaluator: this tests the regex floor, so a jev.json must not let + // Jev's own deny pass a case the floor misses (in-process; a daemon worker + // does not see this variable, hence the sandbox also drops the daemon). + const env = { ...process.env, FAILPROOFAI_TELEMETRY_DISABLED: "1", FAILPROOFAI_EVALUATOR: "legacy" }; + if (sandbox) { + env.HOME = sandbox; + // Both are read before HOME and would bring the real config, jev.json or + // daemon back into the sandbox. + delete env.FAILPROOFAI_HOME; + delete env.FAILPROOFAI_DAEMON_SOCKET; + } const cli = c.cli ?? args.cli; const hookArgs = ["--hook", c.event]; diff --git a/skills/failproofai-policy-publish/SKILL.md b/skills/failproofai-policy-publish/SKILL.md index cb46ade..93be7ab 100644 --- a/skills/failproofai-policy-publish/SKILL.md +++ b/skills/failproofai-policy-publish/SKILL.md @@ -3,7 +3,7 @@ name: failproofai-policy-publish description: |- Package policies written with FailproofAI and publish them as an installable policy pack on GitHub. Use after `failproofai-policy-author` when the policy works locally and the user wants to share, version, release, or install it through `failproofai publish` and `failproofai policies add owner/repository`. - Covers initializing a pack, validating it locally, choosing policy metadata and defaults, Git/GitHub prerequisites, dry runs, publishing release assets, previewing the release, and verifying installation on a clean machine. + Covers initializing a pack, validating it locally, choosing policy metadata and defaults, Git/GitHub prerequisites, dry runs, publishing release assets, previewing the release, and verifying installation on a clean machine. Also packs that carry Jev semantic checks: what publish refuses, the question budget, minCliVersion, and what consumers see. NOT for writing the policy logic (`failproofai-policy-author`) or deploying Cloud-managed policy versions to a fleet (`fp-cloud-cli`). --- @@ -23,8 +23,8 @@ guardrails, and rollback belong to `fp-cloud-cli`. ## Start with an authored policy If no working policy exists yet, use `failproofai-policy-author` first. Come back when the -policy registers through `customPolicies.add(...)` and has passing allow and deny/instruct -cases. +policy registers through `customPolicies.add(...)` (or, for a Jev check, +`semanticPolicies.add(...)`) and has passing allow and deny/instruct cases. Resolve the publishing CLI: @@ -141,6 +141,7 @@ Useful overrides: `effect` applies to the whole pack. `observe` records decisions but blocks nothing; `enforce` applies policy verdicts. Choose deliberately and state it in the handoff. +An `observe` pack's Jev checks are never asked at all (*Packs that carry Jev checks*). Publishing creates or reuses the release and replaces same-named release assets. It does not push the Git source and does not install the pack on any machine. @@ -171,6 +172,74 @@ Selection can be narrowed with `--policy`, `--category`, or expanded with `--all interactive install lets the user choose agents and policies. Never claim publication was successful until `policies show` can read the release and an install can verify its digest. +## Packs that carry Jev checks + +A pack is the only way a Jev check (`semanticPolicies.add`) reaches a machine. In a local +policy file it is never asked (failproofai logs that when it loads the file), so step 2 +cannot prove a check: install a dry run on this machine instead (`references/publishing.md`, +*Try a dry-run pack before releasing*). Discovery finds a file that calls +`customPolicies.add` **or** `semanticPolicies.add`, so a pack of Jev checks alone is +publishable. Author the checks with `failproofai-policy-author` first (its *Jev: when no +string decides it* section). + +Since failproofai 1.0.9 the CLI ships no checks of its own. FailproofAI's 16 are the +`FailproofAI/jev-policies` pack, installed like any other, and a machine asks only the checks +its installed `enforce` packs declare. Yours are asked beside theirs, never instead of them. + +**Always validate with `--dry-run`.** It runs the loader's own rules and publishes nothing. +Without it, `publish` takes the repository from `--repo` or the git origin and releases for +real. With neither it only builds (in an interactive terminal it first asks where to +publish). + +```bash +failproofai publish ./db-guard-policies.mjs --dry-run --id acme/db-guard --version 0.1.0 +``` + +Expect, beside the usual asset paths (the counts are your file's): + +```text +Built acme/db-guard@0.1.0 — <n> policies, <n> on by default. + <n> semantic policies for Jev (<chars> characters of questions), asked by Jev wherever it installs. + Requires failproofai 1.0.8-beta.0 or newer. +``` + +With `--effect observe` the Jev line ends `not asked where it installs: Jev asks only the +checks of packs that enforce.` instead. A pack of Jev checks alone also prints the +**rollback reminder**, and the real publish repeats it: tell users to remove the pack +(`failproofai policies remove <id>`) before rolling a machine back to an older failproofai, +which can deny every tool call over a pack it will not load. Put that sentence in the +release notes. + +What `publish` refuses for a pack with checks (the exact messages are in +`references/publishing.md`): + +- a check named like one of the 16 `FailproofAI/jev-policies` checks (`destructive-deletion`, + `secret-exposure`, …) unless the repository is FailproofAI's; +- a question set over the **budget**: 9,101 characters for a pack from outside FailproofAI + (27,591 per request, minus the 18,490 `jev-policies` takes first where both are + installed). It holds you to that whether or not you have `jev-policies` installed; +- a `reviewedBy` naming a check the pack does not declare, once it declares any; an + `authority` other than `"hard"` / `"reviewable"`; +- `alwaysOn` on a check or a policy; +- `--min-cli-version` below `1.0.8-beta.0`, or not plain semver. With none, `publish` + writes `1.0.8-beta.0` into the manifest; +- a malformed check: missing `userCanOverride` (it has no default), no `appliesTo` or an + unknown tool class, a `mode` other than `deny` / `instruct`, an unknown `precondition`, no + probes or more than 6, a reserved probe id (`exempt`, `user_asked`), a duplicate name, or a + field over its length cap. + +**What consumers see.** `failproofai policies show <owner>/<repo>` counts the checks in its +header, prints `Requires failproofai 1.0.8-beta.0 or newer.`, and lists them in a section +headed `Jev checks — <n> · not selectable · only where Jev is configured`: each check's mode, +its name, and which of the pack's own policies it `reviews`. `failproofai policies add` +prints `<n> Jev checks, asked by Jev on every tool call they apply to` (`…, for hermes only` +under `--cli hermes`) and `They apply only where you configured Jev (failproofai jev status)`, +plus a `▲` line for each check this machine will leave out (the shared budget, or a reserved +or contested name). The checks are not selectable: `--policy` cannot name one and +`failproofai policies` never lists them. They do nothing until the consumer configures Jev +(`failproofai jev setup`) and switches it to `enforce`; `failproofai jev status` shows the +mode. + ## Safety and authorization `--init`, local installation, `--dry-run`, and `policies show` are local/read-only enough to @@ -190,6 +259,8 @@ Report: - pack id and GitHub repository; - version/tag and source commit; - `enforce` or `observe`; +- for a pack with Jev checks: how many, their question cost, `minCliVersion`, and whether + the rollback reminder applies; - dry-run result and generated asset directory; - release URL if actually published; - preview/install verification performed; diff --git a/skills/failproofai-policy-publish/references/publishing.md b/skills/failproofai-policy-publish/references/publishing.md index 3493cf5..7febfad 100644 --- a/skills/failproofai-policy-publish/references/publishing.md +++ b/skills/failproofai-policy-publish/references/publishing.md @@ -17,8 +17,8 @@ failproofai policies remove <pack-id> ## Source discovery `publish` recognizes policy files by their contents: they import FailproofAI and register -one or more entries with `customPolicies.add`. Discovery is non-recursive. If discovery finds -multiple candidates, pass the intended source explicitly or organize the directory so the +one or more entries with `customPolicies.add` or `semanticPolicies.add`. Discovery is +non-recursive. If discovery finds multiple candidates, pass the intended source explicitly or organize the directory so the set is unambiguous. Every policy included in a pack needs a unique name. The pack build rejects an artifact that @@ -70,3 +70,100 @@ pack as generally installable. - `failproofai policies add <owner>/<repo>` is the consumer install step. - Cloud policy publication and fleet rollout use `fp policies`, `fp fleet`, and `fp guardrails`; those belong to `fp-cloud-cli` and are not policy-pack publishing. + +## Jev checks in a pack + +Verified against failproofai 1.0.9 (`src/hooks/pack-cli.ts`, grep `async function build`; +`src/hooks/pack-manifest.ts`, grep `parsePackSemanticPolicy`). + +### What the manifest carries + +A pack with checks writes them to the manifest's `semantic` array, beside `policies`, and +sets `minCliVersion`. Each policy's `authority` and `reviewedBy` are copied into its +manifest entry; a machine reads them from there, never from the code. A pack of checks alone +has an empty `policies` array and is valid. + +### What `publish` refuses + +Every refusal exits 1 and builds nothing. The messages below are the CLI's own; `…` stands +for the pack id, the check's name or index, or a number. + +| Refused | Message starts | +|---|---| +| one of the 16 `FailproofAI/jev-policies` names, from a repository that is not FailproofAI's | `"destructive-deletion" is a name reserved for FailproofAI's own Jev checks, so a pack from … would never have it asked. Pick a name of your own.` | +| questions over the budget | `This pack's … semantic policies compile to … characters of questions, over the 9101 a machine leaves a pack from outside FailproofAI: one Jev request has room for 27591, and FailproofAI/jev-policies' 16 checks take 18490 of it first where both are installed.` | +| `reviewedBy` naming a check the pack does not declare (when it declares any) | `One policy declares an authority this build cannot publish:` … `This pack declares Jev checks of its own, so reviewedBy may name only those.` | +| `--min-cli-version` below 1.0.8-beta.0 | `--min-cli-version 1.0.7 is older than the first failproofai that runs a pack's Jev checks (1.0.8-beta.0):` | +| `--min-cli-version` not plain semver | `--min-cli-version "…" is not a version that can be compared.` | +| `alwaysOn` on a check | `… semantic policy #0 declares alwaysOn, which packs may not set` | +| no `userCanOverride` | `… is missing userCanOverride, which has no default` | +| no `appliesTo`, or an unknown class | `… has no appliesTo tool classes` / `… applies to "…", which is not a tool class (…)` | +| a `mode` other than `deny` / `instruct` | `… has mode "…", which must be "deny" or "instruct"` | +| no probes, or more than 6 | `… declares no probes` / `… declares 7 probes, over the cap of 6` | +| an unknown precondition | `… names precondition "…", which this build does not have (always, protected_branch, in_git_repo, has_paths, paths_outside_project, system_or_root_paths)` | +| probe id `exempt` or `user_asked` | `… uses the reserved probe id "user_asked"` | +| two checks with one name | `two semantic policies are called "…"` | + +Also refused, as for any pack: `alwaysOn` on a regex policy, a policy name with `/`, a +missing `description`, `category` or `match`, an entry that imports local files, and an entry +that registers nothing. `publish` does not count checks, but `policies add` refuses a pack +with more than 24 (`… declares … semantic policies, over the cap of 24`); the budget usually +stops one first. + +### The budget, precisely + +One Jev request has room for 27,591 characters of questions, shared by every installed pack +that enforces, FailproofAI's first. `FailproofAI/jev-policies` spends 18,490 of it, so +`publish` holds a pack from outside FailproofAI to the 9,101 left, whether or not the author +has `jev-policies` installed. The cost is the serialized question map: each probe's +`instructions` and `criteria`, the `exempt` probe, and a generated `user_asked` question for +a check with `userCanOverride: true`. A machine measures what is actually installed, so a +pack that passed `publish` can still have a check dropped beside another Jev pack. +`policies add` names the dropped check with a `▲` line; after that, only the hook log does. + +### Reserved and contested names + +The 16 names in `FailproofAI/jev-policies` are reserved for packs released from a FailproofAI +repository, judged by the repository the CLI fetched from (or `--id` for a dry run with no +`--repo`), never by the pack's own id. Anyone else's check with one of those names is void: +never asked, never a reviewer, and not a rival to FailproofAI's. A name two installed packs +declare differently is asked for neither and clears nothing. `policies add` also refuses a +pack whose id starts `FailproofAI/` unless its release comes from `github.com/FailproofAI` +(`… was not installed: the FailproofAI/ namespace is reserved for releases from +github.com/FailproofAI`). + +### `minCliVersion` + +A CLI too old for Jev checks ignores `semantic` (1.0.7) or replaces its built-in checks with +it (1.0.7-beta.x), so a pack with checks needs at least `1.0.8-beta.0`. `publish` writes that +when `--min-cli-version` is absent and refuses anything lower. An older CLI that does read the +field refuses to install the pack and prints `npm i -g "failproofai@>=<version>" && +failproofai update`. For an `enforce` pack with regex policies, a machine that cannot load it +denies what those policies cover. A pack of Jev checks alone is read correctly from +1.0.8-beta.0, but an older build can deny every tool call over it, which is what the rollback +reminder is for. + +### Try a dry-run pack before releasing + +This is also the only way to exercise a Jev check before release, since a local policy file +never has its checks asked. `policies add` takes only `owner/repo[@tag]` and fetches +`$FAILPROOFAI_PACK_BASE_URL/<owner>/<repo>/releases/download/<tag>/<asset>` (default +`https://github.com`). To install a dry run locally, serve `dist-pack/` in that layout: + +```bash +failproofai publish ./db-guard-policies.mjs --dry-run --id acme/db-guard --version 0.1.0 +d=pack-mirror/acme/db-guard/releases/download/0.1.0; mkdir -p "$d" && cp dist-pack/* "$d" +python3 -m http.server 8765 -d pack-mirror & +FAILPROOFAI_PACK_BASE_URL=http://127.0.0.1:8765 failproofai policies add acme/db-guard@0.1.0 --all +``` + +The tag must describe the version (a leading `v` is fine). Unset the variable afterwards: it +redirects every pack fetch. Re-adding with a newer `@<version>` upgrades the pack, and +`failproofai policies remove acme/db-guard` uninstalls it. + +### Observe and agent-scoped packs + +`--effect observe` records the pack's regex verdicts without enforcing them, and its Jev +checks are **not asked at all**. A consumer who installs with `--cli <agent>` gets the checks +asked only for those agents. Either way, a policy naming a check that is not asked stays +hard.