Developer tooling for the Routeplane AI Gateway — a neutral, multi-provider, OpenAI-compatible proxy with sovereign routing, governance, and agentic security.
This is a pnpm + Turborepo monorepo. All three packages are published with npm provenance attestations, so every release is publicly verifiable back to the exact repository state and CI run that built it.
| Package | What it is |
|---|---|
@routeplane/sdk |
TypeScript SDK — drop-in OpenAI client + zero-dependency core client |
@routeplane/cli |
rp command-line interface |
@routeplane/mcp-server |
Model Context Protocol server — 40 gateway operations as tools |
Routeplane is a subclass of the official openai client. Point your existing OpenAI code at the gateway, change nothing else, and get multi-provider fallback, sovereign routing, and FinOps attribution for free:
import { Routeplane } from '@routeplane/sdk';
const client = new Routeplane({
apiKey: process.env.ROUTEPLANE_API_KEY!, // rp_...
provider: 'openai,anthropic', // try OpenAI, fall back to Anthropic
residency: 'IN', // keep regulated data in-region
useCase: 'support-bot', // FinOps cost attribution
});
// Exactly the OpenAI SDK you already know:
const completion = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: 'Say hello in one word.' }],
});
console.log(completion.choices[0]?.message.content);
// Steer a single request and read what the gateway decided:
const withMeta = await client.createChatCompletion(
{ model: 'gpt-4o-mini', messages: [{ role: 'user', content: 'Ping' }] },
{ strategy: 'cost', idempotencyKey: 'ping-1' },
);
console.log(withMeta.routeplane.provider); // which provider actually served it
console.log(withMeta.routeplane.cache); // 'hit' | 'miss' | 'bypass'If you don't want the openai peer dependency, import the core client and the typed header builder from @routeplane/sdk/core:
import { RouteplaneCoreClient, createHeaders } from '@routeplane/sdk/core';
const rp = new RouteplaneCoreClient({ apiKey: process.env.ROUTEPLANE_API_KEY! });
// Prompt management, logs, FinOps, cache, feedback — the non-OpenAI surfaces:
const usage = await rp.finops.usage(); // recent process-local usage
// Build the x-routeplane-* headers yourself for any transport:
const headers = createHeaders({ provider: 'gemini', strategy: 'latency', residency: 'IN' });For durable FinOps history, use rp.finops.usageDailyReport({ from, to })
with inclusive UTC dates (YYYY-MM-DD). The managed gateway requires the
finops_export and telemetry_durable capabilities and an armed durable reader;
this route is absent from Community Edition. The server defaults to the last
14 days and clamps the window to 92 days. The SDK returns the complete report.
The current server contract includes typed cost_evidence at totals, daily rows, model rows,
and key rows. This field is optional in SDK types so older gateways remain
compatible: omission means unknown evidence and stays absent. Its nested fields
are required, including explicit nulls. A complete estimated zero is numeric
zero; unpriced or unavailable totals carry null. Mixed priced/unpriced coverage
can expose only the named estimated_priced_subtotal, including a valid zero
subtotal. Legacy cost_micro_usd and pricing remain independent and can differ
from typed evidence; do not substitute either for a missing typed total.
Typed amounts are exact integer micro-currency with scale 6 and their stated
currency, within JavaScript's safe integer range. They represent recorded chat
estimates under receipt-anchored rate evidence, not provider-reported, billed,
settled, or reconciled charges. Non-chat requests are explicitly not covered.
Scope, freshness, coverage, component omissions, snapshots, and truncation are
returned with the amount; null complete_through and retention boundaries do
not assert complete history. The SDK uses its existing JSON transport without
changing response values.
These are TypeScript response types; they do not add runtime validation or
normalization of server responses. This source contract update becomes available
in published SDK packages through a deliberate release.
Use the gateway-generated req_... identifier from the response's
x-routeplane-request-id (or its x-routeplane-trace-id alias), not the
provider's completion body ID:
await rp.feedback.create({ requestId: gatewayRequestId, score: 1 });The existing arguments now serialize as {"trace_id":"req_...","value":1}.
Scores must be integers from −10 through 10, without rescaling. Fractional,
nonfinite, out-of-range and non-number values are rejected before dispatch.
Omitted, null and empty comments are omitted; every nonempty comment, including
whitespace-only text, is explicitly rejected as unsupported. The helper still
returns Promise<void>; acknowledgement proves neither target existence nor
durable storage.
The existing CLI uses the same contract:
rp feedback --request-id req_from_gateway_response_header --score 1It prints acknowledgement only after success. Invalid scores and unsupported
comments exit nonzero without sending feedback. --comment remains recognized,
but only an empty value is supported by this legacy endpoint.
Wrap every agent tool call in the gateway's default-deny policy boundary. A refusal is a
verdict rather than a thrown error, so authorization reads as a branch, not a try/catch:
const verdict = await rp.mcp.authorizeToolCall({
agentId: 'support-bot',
server: 'github',
tool: 'create_issue',
arguments: { repo: 'acme/api', title: 'Flaky test' },
});
if (verdict.outcome === 'deny') throw new Error(verdict.reason);
// ... make the tool call, then screen the result before the model sees it:
const inspection = await rp.mcp.inspectToolResult(result);rp.mcp covers the whole surface: the two enforcement points (authorizeToolCall,
inspectToolResult), run governance (runStep, listRuns), sampling defense
(evaluateSampling), the human-in-the-loop queue (listPendingHitl, hitlStatus,
approveHitl, denyHitl), signed action receipts (issueReceipt, verifyReceipt), the
anomaly operator surface (anomalyStatus, clearAnomaly), and the enforcement-event feed
(securityEvents).
Reading a verdict is fail-closed: only an explicit allow is an allow. An empty body, an
unknown outcome, or a proxy's error page all read as a deny, so a response the client
cannot parse can never fall through as permission granted. (A 4xx carrying no verdict at
all still throws — that is a malformed request, and turning your own bug into a policy deny
would hide it.)
All of it is gated on the tenant's AgenticSecurity entitlement. The gateway hides the
surface rather than refusing it, so an un-entitled key gets RouteplaneError with status
404 — not a 403.
The same surfaces are available from the CLI (rp agents runs | events | pending | approve | deny) and as MCP-server tools.
rp eval scores your outputs against a suite file and fails the build when quality drops. It
is deterministic — no model is called, so a run costs nothing per case and scores
identically every time. A gate that is itself non-deterministic is not a gate.
rp eval run eval-suite.jsonEVALUATOR N PASSED SKIPPED MEAN THRESHOLD GATE
valid_json 3 3 1.000 —
canonical_json_match 2 1 1 0.500 0.950 FAIL
trajectory_superset_match 1 1 2 1.000 1.000 PASS
Exit codes are the point: 0 all thresholds met, 1 a threshold was missed, 2 the suite file or the request was rejected. A build gate has to tell "quality dropped" from "the config is wrong" — collapsing both into one code trains people to ignore the signal.
Two behaviours worth knowing:
- A missing input is a skip, not a failure. A case with no
referencecannot be exact-matched; it is reported as skipped and excluded from the means. Counting it as 0 would blend "we did not check this" into "this was wrong" and read as a regression that never happened. - An evaluator with no threshold is reported but never fails the run. Reporting and gating are separate decisions; conflating them makes people delete checks instead of fixing them.
trajectory_match compares the tool calls an agent made against a reference run — message
content is never compared, so a reworded explanation is the same trajectory. superset is
the one that catches a dropped step:
{
"type": "trajectory_match",
"match_mode": "superset",
"overrides": { "issue_refund": { "type": "on_keys", "paths": ["order_id"] } }
}The per-tool on_keys override is what makes this usable on a real agent: pin the order id,
ignore the timestamp that changes every run.
rp eval rubrics lists the built-in judge rubrics. Those are scored by an operator-armed
evaluation run on the gateway, not by rp eval.
From the SDK:
const report = await client.evaluations.score(cases, [{ type: 'valid_json' }]);Runnable snippets live in examples/:
| File | Shows |
|---|---|
basic.ts |
Minimal Routeplane client — a drop-in OpenAI subclass |
headers-only.ts |
Stock openai SDK + createHeaders for per-request steering |
streaming.ts |
Streaming with the gateway's decision metadata |
vercel-ai-sdk.ts |
Vercel AI SDK (@ai-sdk/openai) integration |
resources.ts |
Non-OpenAI surfaces — status, logs, FinOps, prompts, cache |
agentic-security.ts |
Guarding an agent loop — tool-call authorization, result inspection, run breakers, receipts |
eval-suite.json |
A rp eval suite — JSON/text checks, a trajectory comparison, and gate thresholds |
See examples/ for more.
pnpm install
pnpm build # turbo build across all packages
pnpm test # vitest via turbo
pnpm lint # tsc --noEmit type-checkRequires Node.js >= 18 (native fetch).