Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
138 changes: 138 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,144 @@ target.

## [Unreleased]

The request guards — outbound redaction, tool-call inspection and the
content filter — now share their rule model with the desktop gateway:
the same policy shape, built-in rule catalog, validation, rule view and
sample trial, from thinkwatch-core. Each guard has three modes (off,
observe, and one named for what it does: replace, cut off, enforce), and
the content filter can delete what a rule matches as well as refuse or
record it, and match by code point. Hidden characters become content
filter rules, and the per-model output length guardrail becomes a cap on
the output tokens a request may ask for. Settings saved by an earlier
version are converted at the first start: read the first section before
deploying.

### Read before upgrading

- **The old guard settings are converted at the first start, and
behave as before.** `security.content_filter_patterns`,
`security.hidden_text`, `security.pii_redactor_patterns` and
`security.tool_inspection` become `security.content`,
`security.redact` and `security.inspect_tools`, and are deleted, in
one transaction during the boot migration; a second start finds
nothing to convert. A content rule identical to a built-in rule
becomes that rule, switched on, any other a custom rule; a list with
rules in it runs in enforce mode with every built-in rule it did not
name switched off. The four seeded PII patterns become the built-in
rules for the same data (`cn-resident-id`, `bank-card`, `email`,
`cn-mobile-phone`), any other pattern a custom rule whose label is
its old placeholder prefix. A model's `output_guardrails` length cap
becomes `max_output_tokens` (below), and the column is dropped. **Stop
every replica of the previous version before the first new one
starts**: a replica still running 2.2 finds its settings gone (it
then filters and redacts nothing) and can no longer rebuild its
router once the column is dropped. To see what was converted, read
the three keys from Settings or `system_settings` afterwards; the
start-up log lists them too.
- **Placeholders are written `<<TW_EMAIL_1>>`, not `{{EMAIL_1}}`.** The
label of a custom rule is upper case letters, digits and
underscores (an old prefix is converted: `REDACTED-SSN` →
`REDACTED_SSN`); the built-in identity number and bank card rules
use `ID_NUMBER` and `CARD_NUMBER`. Anything that looked for the old
form in answers or logs needs the new one.
- **Redaction searches the whole request**, not only the user's
messages: the system prompt, earlier answers and tool-call arguments
are redacted too. Base64 payloads (images, files, signatures) are
still left alone. Rules run on the request as it is sent, as JSON,
where a custom pattern's match ends at a quote or a backslash: a
pattern written to match across a `"` in the decoded text needs
rewriting.
- **Built-in credential rules start replacing on deployments that were
redacting.** API keys and tokens with a known prefix, private keys,
JWTs and connection-string passwords are built-in rules that ship
switched on. A deployment whose PII list had patterns in it runs
redaction in enforce mode after the upgrade, so these values are now
replaced as well. Switch the ones you do not want off on the
console's security page. With an empty PII list, redaction
converts to observe mode: it records what it finds and changes
nothing.
- **Tool calls are judged as the client receives them, and two built-in
rules are new.** Inspection now reads a tool call converted to the
caller's format and with redacted values restored — what the client
would run — where it used to read the placeholders. The new
`secret-to-unknown-host` rule cuts (in enforce mode) a call that sends
a recognised API key or private key to a host that is neither local
nor the key's own provider; `upload-file-to-host` records a call that
uploads a local file to an outside host. A deployment running
tool-call inspection in enforce mode starts cutting the first; add it
to `disable` if that is not wanted.
- **"Warn" and "log" are one action now, "record only"**, and hidden
characters are content filter rules: `unicode-tags` and
`bidi-controls`, plus `zero-width` and `private-use`, which ship off.
`security.hidden_text: block` converts to those two rules refusing,
`warn` and `log` to recording, `off` to switching them off.
- **The output length guardrail is replaced by a model's maximum output
tokens.** A cap of N bytes on the answer converts to `ceil(N / 4)`
output tokens. The answer is no longer measured or cut: a request
asking for more tokens than the cap is lowered to it, and one asking
for none gets it, in whichever field its API uses; the upstream stops
there. The model API's `output_guardrails` field is gone;
`max_output_tokens` (1 to 2147483647, `null` for no limit) replaces
it.
- **A new installation observes by default.** Every guard starts in
observe mode, with only the built-in rules that rarely misfire
switched on (personal data such as e-mail addresses and phone
numbers ships off). Nothing is refused, replaced or deleted until a
guard is switched to its third mode.
- **A content filter refusal is `403`**, with the error type of the
caller's API (`permission_error` for OpenAI-style APIs). Keyword and
regex rules used to refuse with `400`.
- **Guard policies are changed with their own permissions.** Writing
`security.redact` through `PATCH /api/admin/settings` takes
`pii_redactor:write`, `security.content` and `security.inspect_tools`
take `content_filter:write`; `settings:write` no longer covers them.
The seeded `admin` and `super_admin` roles hold both.
- **Console API changes.** `GET /api/admin/security` lists each guard's
mode and every rule, and `POST /api/admin/security/{guard}/test` tries
a sample; they replace `/api/admin/settings/content-filter/test`,
`/content-filter/presets`, `/pii-redactor/test`,
`/tool-inspection/rules` and `/tool-inspection/test`, which are gone.
- **Audit events.** Every guard hit writes one event:
`gateway.content_flagged`, `gateway.content_stripped` and
`gateway.content_blocked`; `gateway.redaction_flagged` and
`gateway.redaction_replaced`; `gateway.tool_call_flagged` and
`gateway.tool_call_blocked` as before. `gateway.hidden_text_flagged`
and `gateway.hidden_text_blocked` are gone; hidden characters are
content events. With `audit.body_redact_pii` on, captured bodies are
redacted with the outbound redaction rules, built-in ones included,
whatever the redaction mode.

### Added

- **Deleting what a content rule matches.** A content rule can refuse
the request, delete the matched text from the caller's messages and
tool results and send the rest, or only record. Text deleted joins
back what it separated, so the request is checked again afterwards.
- **Code point rules.** A content rule can match characters by code
point (`U+200B`, `U+E0000–U+E007F`), for invisible characters a
keyword cannot be written for.
- **Every rule visible and switchable**, built-in and custom, in each
guard, with what it does in the third mode and what it did out of the
box; a sample can be tried against one rule, an unsaved one, or all
of them.

### Changed

- **Core crates at ThinkWatch-Core v0.58.0.** `tw-dialect`, `tw-guard`,
`tw-breaker` and `tw-bedrock` move from v0.55.0; the shared guard model
described above comes with them.

### Fixed

- **A credential in a matched tool call no longer reaches the audit
log.** The excerpt of a tool call that inspection cut or recorded —
and of a content filter hit — is masked with the redaction rules
before it is written; a key the model echoed, or one restored from a
placeholder, used to be stored as it was.
- **A request's audit events and its log row carry the same id** when
the caller sends no `x-trace-id`. The log row of a request that went
through used to carry a second, unrelated id.

## [2.2.0] — 2026-10-01

This release fixes authorization. The gateways never checked
Expand Down
26 changes: 14 additions & 12 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

8 changes: 4 additions & 4 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -70,10 +70,10 @@ opt-level = 3
# never re-exported through a local shim. And the reverse: something only
# this side uses (the at-rest crypto, IMDSv2 credentials, the gateway error)
# lives here, not in core.
tw-bedrock = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.55.0" }
tw-breaker = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.55.0" }
tw-dialect = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.55.0" }
tw-guard = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.55.0" }
tw-bedrock = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.58.0" }
tw-breaker = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.58.0" }
tw-dialect = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.58.0" }
tw-guard = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.58.0" }

# Web framework
axum = { version = "0.8", features = ["macros", "ws"] }
Expand Down
13 changes: 8 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,12 +38,12 @@
## Highlights

- **MCP tool calls run as the real user.** Each user connects their own GitHub, Notion, Linear, Slack or Atlassian account through OAuth or a personal token, so the upstream's own audit log shows who acted. Tokens are encrypted at rest, tool lists are cached per user, and each tool can be granted per role and per API key.
- **Security guards on every request.** PII such as emails, phone numbers and card numbers is replaced with placeholders before a request goes upstream and restored in the answer, including streamed ones. Tool calls in model responses are checked against rules for dangerous commands, and hidden Unicode characters and prompt-injection phrases in requests are logged or refused.
- **Security guards on every request.** Outbound redaction replaces credentials and personal data anywhere in a request with placeholders such as `<<TW_EMAIL_1>>` before it goes upstream, and restores them in the answer, streamed ones included. Tool-call inspection checks the tool calls in each response for dangerous commands, and the content filter looks for prompt-injection phrases and hidden characters in what the caller sent, then refuses the request, deletes them or records them.
- **Identity from the organization's directory.** Sign-in works through any OIDC provider (Zitadel, Okta, Azure AD and others), with optional TOTP. Five built-in roles, from Super Admin to Viewer, and custom roles decide who may use which models, tools and admin pages.
- **One key for AI and MCP.** Users receive `tw-` virtual keys that can be scoped to the AI gateway, the MCP gateway or both. Keys are stored only as hashes and rotate with a grace period.
- **Rate limits and budgets.** Sliding windows from one minute to one week limit requests or tokens, and daily, weekly or monthly budgets cap spending. Both attach to users, API keys or roles, and rate limits apply to MCP tool calls as well as model requests.
- **Cost accounting that finance can use.** Spend is reported by model, user, provider and cost center, with CSV chargeback reports and a month-end forecast. Per-model weights make expensive models count for more against the same quota.
- **Audit trail in ClickHouse.** Every model request and tool call is recorded with user, parameters, response, latency and errors, and request bodies can be PII-redacted before storage (off by default). Events can be forwarded to a SIEM over Syslog, Kafka (through a REST proxy) or signed webhooks.
- **Audit trail in ClickHouse.** Every model request and tool call is recorded with user, parameters, response, latency and errors, and captured bodies can be redacted with the outbound redaction rules before storage (off by default). Events can be forwarded to a SIEM over Syslog, Kafka (through a REST proxy) or signed webhooks.
- **One endpoint for every client.** OpenAI Chat Completions, OpenAI Responses, Anthropic Messages and Gemini requests are served on one port and converted to whatever the upstream speaks. Routing spreads traffic by weight, latency or health, and a circuit breaker takes failing upstreams out of rotation.

## Quick start
Expand Down Expand Up @@ -82,9 +82,12 @@ The gateway (port `3000`) is the only part that clients need to reach. The conso
- Responses from servers that use per-user credentials are cached per user and account, never shared.

**Security guards**
- Tool-call inspection starts in observe mode: hits are recorded, and nothing is cut off until enforce mode is chosen. Built-in rules can be switched off or re-graded, and custom rules added.
- Hidden-character detection defaults to warn; it covers Unicode tag characters and bidirectional overrides in the caller's messages and tool results.
- The content filter ships with rules for common prompt-injection phrases, each set to block, warn or log. PII patterns are editable in the console.
- There are three guards, each with three modes: off, observe, and one named for what it does — replace (outbound redaction), cut off (tool-call inspection) and enforce (content filter). A new installation starts all three in observe mode: hits go to the audit log and nothing is changed until a guard is switched to its third mode.
- Every rule is listed on the console's security page, built-in and custom. Built-in rules can be switched on or off, tool-call and content rules can take another action, custom rules can be added, and a sample can be tried against one rule or a whole guard first.
- Outbound redaction searches the whole request, system prompt and earlier answers included, but not base64 payloads. A match becomes `<<TW_LABEL_n>>` — `SECRET` for credentials, `ID_NUMBER`, `CARD_NUMBER`, `EMAIL` and `PHONE` for personal data, a label of its own for a custom rule — and is restored in the answer. E-mail addresses and phone numbers ship switched off.
- A content rule matches a phrase, a regular expression or code points (`U+200B`, `U+E0000–U+E007F`), and either refuses the request, deletes what it matched from the caller's messages and tool results, or records only. Hidden characters are content rules: Unicode tag characters and bidirectional controls ship on, zero-width and private-use characters off.
- A tool-call rule cuts the response at the call or records it. Besides the dangerous-command rules, two built-in rules catch a credential sent to an unknown host and a local file uploaded to an external host.
- A model's maximum output tokens, set on the Models page, caps `max_tokens` on every request to that model; it replaces the old output length guardrail.

**Limits and budgets**
- Request-count limits are checked before the request; token limits and budgets are counted after the response, so one request can cross a budget before the next is refused.
Expand Down
Loading
Loading