Skip to content

Unify the request guards with thinkwatch-core's rule model - #72

Merged
fylorn merged 8 commits into
devfrom
feat/guard-unify
Oct 3, 2026
Merged

fylorn merged 8 commits into
devfrom
feat/guard-unify

Conversation

@fylorn

@fylorn fylorn commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

The enterprise backend half of the guard unification: outbound redaction, tool-call inspection and the content filter now run on thinkwatch-core's shared rule model (tw_guard::policy, view, trial, content::screen, redact::flow), the same one the desktop gateway uses. The console side is on guard-unify-web and is merged separately.

Pinned to ThinkWatch-Core v0.58.0 (c2a7bc6), the release that carries ThinkWatchProject/ThinkWatch-Core#268; the layer-one crates move from v0.55.0.

Settings

  • Three keys, one policy object each, the shape tw_guard::policy defines: security.redact, security.inspect_tools, security.content. Saved through PATCH /api/admin/settings, checked by core's own validation (400 with its message). Seeds ship {} (the factory policy: observe).
  • The old keys (security.content_filter_patterns, security.pii_redactor_patterns, security.hidden_text, security.tool_inspection) and models.output_guardrails are converted at boot, in one transaction that also deletes them, under an advisory lock (common::guard_policy::legacy). The conversion keeps behaviour: content rules identical to a built-in rule become that rule; a list with rules runs in enforce with unlisted built-ins switched off; hidden_text maps onto unicode-tags / bidi-controls (block with an empty list still enforces); the four seeded PII patterns become built-in rules, the rest custom rules labelled by their old prefix; a byte cap of N becomes ceil(N / 4) output tokens. Unit tests cover every mapping; an integration test runs it against a real database twice.

Gateway

  • Content filter: content::screen on the caller's own bytes; refuse (403, masked quote), strip (the request continues as the stripped body, decoded again — redaction, every hop and the audit row see it), or record. The Responses WebSocket goes through the same pipeline.
  • Outbound redaction: flow::look on the whole request body, flow::replace on every hop with the same ledger, restoration by the shared engine; <<TW_LABEL_n>> placeholders. pii_redactor.rs and its own traversal are gone.
  • Tool-call inspection: ToolPolicy; calls are judged as the client receives them — converted and with placeholders restored — which is what core's new secret-to-unknown-host rule needs.
  • One audit event per hit: gateway.content_{flagged,stripped,blocked}, gateway.redaction_{flagged,replaced}, gateway.tool_call_{flagged,blocked}. Excerpts (and the content refusal quote) are masked with the redaction rules first.
  • output_guardrails.rs is gone. A model's max_output_tokens lowers a larger ask and fills a missing one with tw_dialect::params::cap_max_output_tokens, forwarded or converted.

Console API

  • GET /api/admin/security → tw_guard::view::detail (settings:read).
  • POST /api/admin/security/{guard}/test → tw_guard::trial::run (pii_redactor:read / content_filter:read).
  • Writing security.redact takes pii_redactor:write; security.content and security.inspect_tools take content_filter:write.
  • Removed: /api/admin/settings/content-filter/test, /content-filter/presets, /pii-redactor/test, /tool-inspection/rules, /tool-inspection/test.
  • Models API: max_output_tokens (1–2147483647, null clears) replaces output_guardrails.

CHANGELOG has the upgrade notes under Unreleased.

Tests

cargo fmt --check, cargo clippy --workspace --all-targets -D warnings and the unit tests pass locally; the integration suite passed in full locally (379 tests) and runs in CI on every push. The web side came in from guard-unify-web, and dev was merged in for the clippy fix (#73).

🤖 Generated with Claude Code

fylorn and others added 5 commits October 3, 2026 02:52
Outbound redaction, tool-call inspection and the content filter now run
on the shared model in tw-guard: the policy shape, built-in catalogs,
validation, rule view and trial, the content `screen` and the redaction
`flow`. The enterprise copies of each are gone.

- Settings: `security.redact`, `security.inspect_tools` and
  `security.content` hold one policy each, validated by core and seeded
  as the factory policy (observe). The old keys and
  `models.output_guardrails` are converted once at boot, in one
  transaction, keeping what they did, then removed.
- Content filter: rules refuse (403), strip the matched text (the
  request continues as the stripped body, decoded again) or record;
  code-point rules; hidden characters are built-in rules now.
- Redaction: the whole request body is searched, every hop goes out
  through the same ledger, placeholders are `<<TW_LABEL_n>>`.
- Tool calls are judged as the client receives them, restored.
- One audit event per hit; excerpts and refusal quotes are masked with
  the redaction rules before they are written.
- Models: `max_output_tokens` caps the output a request may ask for,
  forwarded or converted, replacing the output length guardrail.
- Console API: `GET /api/admin/security`,
  `POST /api/admin/security/{guard}/test`; the five old content-filter,
  PII and tool-inspection endpoints are removed. Writing a guard policy
  takes `pii_redactor:write` or `content_filter:write`.

Core is pinned to `feat/guard-unify` by rev until it is released.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… model

The security page shows the three request guards - outbound redaction,
tool-call inspection and the content filter - one tab each, in the shape
the shared guard model gives both products:

- A mode card per guard: Off / Observe / a third mode named for what it
  does there (Replace, Cut off, Enforce), what the current mode does,
  and the cost of the third mode, said before anyone switches to it.
- A rule table grouped as the server lists it (content: hidden
  characters, instruction override, identity and prompts, Chinese
  phrasing, custom) with what each rule matches, what it does
  (redaction: the placeholder a hit becomes) and a switch.
- Dialogs to create and edit custom rules (content rules match by text,
  regex or code points, and refuse, delete or record; redaction rules
  name their placeholder, SECRET unless changed), to view a built-in
  rule and change its action, and to test a sample against a guard:
  the hits, the text as the third mode would send it, and whether the
  content filter would refuse the request.

Every change writes the guard's whole policy object to its
security.redact / security.inspect_tools / security.content settings
key, built from the rule view GET /api/admin/security returns; samples
are tried through POST /api/admin/security/{guard}/test. Mode and
switch changes show at once and offer an undo. Writes stay behind the
existing permissions: pii_redactor:* for redaction, content_filter:*
for the other two.

The hidden-character card, the content-filter presets and the old test
sandbox are gone: hidden characters are content rules now.

The model editor's output guardrails give way to "Max output tokens"
(blank for no limit), sent as max_output_tokens; the model drawer
shows it.

The types for the new endpoints are provisional (src/lib/security-types.ts)
until the backend's OpenAPI schema carries them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checked against tw-guard's view.rs and trial.rs at the pinned core
commit and the handlers in this branch:

- Tool-call rules implemented in code carry `{ kind: "builtin", check }`
  instead of a regex. The two shipped ones - secret-to-unknown-host
  (cuts) and upload-file-to-host (records) - get names and reasons in
  both languages, and the match column and the rule dialog say what the
  check looks for. They have no pattern, so they cannot be copied as a
  custom rule. The column header is "Match" for every guard now that
  not every tool rule is a regex.
- A single rule is tried with the action chosen in its dialog (the test
  request's new `action`, never sent for redaction), so the marks, the
  text the third mode would send and the refusal all follow the choice;
  the dialogs show what the server answers instead of guessing.
- The page needs `settings:read`, as GET /api/admin/security does;
  writes and tests keep the guards' own permissions.
- The types say where they come from and how the server fills them
  (`why` absent rather than null, `trial` as the rule of a tried
  pattern, `output` null when the request would be refused).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README (both languages) still described PII patterns, hidden
character detection set to warn and a content filter that blocks, warns
or logs. It now describes the three guards as they ship: three modes
each with the third named for what it does, every rule visible and
switchable on the security page, redaction over the whole request with
`<<TW_LABEL_n>>` placeholders, content rules that match phrases,
regexes or code points and refuse, delete or record, hidden characters
as content rules, the two built-in exfiltration rules for tool calls,
and a model's maximum output tokens in place of the output length
guardrail. The web README lists the security page and corrects what the
Settings security tab holds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fylorn and others added 3 commits October 3, 2026 04:16
ThinkWatch-Core#268 is merged (72954bf) and its branch is gone, so the
layer-one crates follow main by rev until the release tag exists. The
layer-one code is the same as at d86320b.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The guard-unify work is released in ThinkWatch-Core v0.58.0, so the
layer-one crates move from the main commit (72954bf) to the tag. Their
code is unchanged between the two.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first v0.58.0 tag (d691c20) was deleted unreleased; the tag now
points at c2a7bc6. The layer-one crates are the same in both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn marked this pull request as ready for review October 3, 2026 01:18
@fylorn
fylorn merged commit 8d8d96f into dev Oct 3, 2026
6 checks passed
@fylorn
fylorn deleted the feat/guard-unify branch October 3, 2026 08:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant