Skip to content

Repository files navigation

routeplane-devtools

Developer tooling for the Routeplane AI Gateway — a neutral, multi-provider, OpenAI-compatible proxy with sovereign routing, governance, and agentic security.

This is a pnpm + Turborepo monorepo. All three packages are published with npm provenance attestations, so every release is publicly verifiable back to the exact repository state and CI run that built it.

Package What it is
@routeplane/sdk TypeScript SDK — drop-in OpenAI client + zero-dependency core client
@routeplane/cli rp command-line interface
@routeplane/mcp-server Model Context Protocol server — 40 gateway operations as tools

The 30-second wow

Routeplane is a subclass of the official openai client. Point your existing OpenAI code at the gateway, change nothing else, and get multi-provider fallback, sovereign routing, and FinOps attribution for free:

import { Routeplane } from '@routeplane/sdk';

const client = new Routeplane({
  apiKey: process.env.ROUTEPLANE_API_KEY!, // rp_...
  provider: 'openai,anthropic',            // try OpenAI, fall back to Anthropic
  residency: 'IN',                         // keep regulated data in-region
  useCase: 'support-bot',                  // FinOps cost attribution
});

// Exactly the OpenAI SDK you already know:
const completion = await client.chat.completions.create({
  model: 'gpt-4o-mini',
  messages: [{ role: 'user', content: 'Say hello in one word.' }],
});
console.log(completion.choices[0]?.message.content);

// Steer a single request and read what the gateway decided:
const withMeta = await client.createChatCompletion(
  { model: 'gpt-4o-mini', messages: [{ role: 'user', content: 'Ping' }] },
  { strategy: 'cost', idempotencyKey: 'ping-1' },
);
console.log(withMeta.routeplane.provider); // which provider actually served it
console.log(withMeta.routeplane.cache);    // 'hit' | 'miss' | 'bypass'

Zero-dependency core

If you don't want the openai peer dependency, import the core client and the typed header builder from @routeplane/sdk/core:

import { RouteplaneCoreClient, createHeaders } from '@routeplane/sdk/core';

const rp = new RouteplaneCoreClient({ apiKey: process.env.ROUTEPLANE_API_KEY! });

// Prompt management, logs, FinOps, cache, feedback — the non-OpenAI surfaces:
const usage = await rp.finops.usage(); // recent process-local usage

// Build the x-routeplane-* headers yourself for any transport:
const headers = createHeaders({ provider: 'gemini', strategy: 'latency', residency: 'IN' });

For durable FinOps history, use rp.finops.usageDailyReport({ from, to }) with inclusive UTC dates (YYYY-MM-DD). The managed gateway requires the finops_export and telemetry_durable capabilities and an armed durable reader; this route is absent from Community Edition. The server defaults to the last 14 days and clamps the window to 92 days. The SDK returns the complete report.

The current server contract includes typed cost_evidence at totals, daily rows, model rows, and key rows. This field is optional in SDK types so older gateways remain compatible: omission means unknown evidence and stays absent. Its nested fields are required, including explicit nulls. A complete estimated zero is numeric zero; unpriced or unavailable totals carry null. Mixed priced/unpriced coverage can expose only the named estimated_priced_subtotal, including a valid zero subtotal. Legacy cost_micro_usd and pricing remain independent and can differ from typed evidence; do not substitute either for a missing typed total.

Typed amounts are exact integer micro-currency with scale 6 and their stated currency, within JavaScript's safe integer range. They represent recorded chat estimates under receipt-anchored rate evidence, not provider-reported, billed, settled, or reconciled charges. Non-chat requests are explicitly not covered. Scope, freshness, coverage, component omissions, snapshots, and truncation are returned with the amount; null complete_through and retention boundaries do not assert complete history. The SDK uses its existing JSON transport without changing response values. These are TypeScript response types; they do not add runtime validation or normalization of server responses. This source contract update becomes available in published SDK packages through a deliberate release.

Legacy feedback

Use the gateway-generated req_... identifier from the response's x-routeplane-request-id (or its x-routeplane-trace-id alias), not the provider's completion body ID:

await rp.feedback.create({ requestId: gatewayRequestId, score: 1 });

The existing arguments now serialize as {"trace_id":"req_...","value":1}. Scores must be integers from −10 through 10, without rescaling. Fractional, nonfinite, out-of-range and non-number values are rejected before dispatch. Omitted, null and empty comments are omitted; every nonempty comment, including whitespace-only text, is explicitly rejected as unsupported. The helper still returns Promise<void>; acknowledgement proves neither target existence nor durable storage.

The existing CLI uses the same contract:

rp feedback --request-id req_from_gateway_response_header --score 1

It prints acknowledgement only after success. Invalid scores and unsupported comments exit nonzero without sending feedback. --comment remains recognized, but only an empty value is supported by this legacy endpoint.

Agentic security

Wrap every agent tool call in the gateway's default-deny policy boundary. A refusal is a verdict rather than a thrown error, so authorization reads as a branch, not a try/catch:

const verdict = await rp.mcp.authorizeToolCall({
  agentId: 'support-bot',
  server: 'github',
  tool: 'create_issue',
  arguments: { repo: 'acme/api', title: 'Flaky test' },
});
if (verdict.outcome === 'deny') throw new Error(verdict.reason);

// ... make the tool call, then screen the result before the model sees it:
const inspection = await rp.mcp.inspectToolResult(result);

rp.mcp covers the whole surface: the two enforcement points (authorizeToolCall, inspectToolResult), run governance (runStep, listRuns), sampling defense (evaluateSampling), the human-in-the-loop queue (listPendingHitl, hitlStatus, approveHitl, denyHitl), signed action receipts (issueReceipt, verifyReceipt), the anomaly operator surface (anomalyStatus, clearAnomaly), and the enforcement-event feed (securityEvents).

Reading a verdict is fail-closed: only an explicit allow is an allow. An empty body, an unknown outcome, or a proxy's error page all read as a deny, so a response the client cannot parse can never fall through as permission granted. (A 4xx carrying no verdict at all still throws — that is a malformed request, and turning your own bug into a policy deny would hide it.)

All of it is gated on the tenant's AgenticSecurity entitlement. The gateway hides the surface rather than refusing it, so an un-entitled key gets RouteplaneError with status 404 — not a 403.

The same surfaces are available from the CLI (rp agents runs | events | pending | approve | deny) and as MCP-server tools.

Evals in CI

rp eval scores your outputs against a suite file and fails the build when quality drops. It is deterministic — no model is called, so a run costs nothing per case and scores identically every time. A gate that is itself non-deterministic is not a gate.

rp eval run eval-suite.json
EVALUATOR                  N  PASSED  SKIPPED  MEAN   THRESHOLD  GATE
valid_json                 3  3                1.000             —
canonical_json_match       2  1       1        0.500  0.950      FAIL
trajectory_superset_match  1  1       2        1.000  1.000      PASS

Exit codes are the point: 0 all thresholds met, 1 a threshold was missed, 2 the suite file or the request was rejected. A build gate has to tell "quality dropped" from "the config is wrong" — collapsing both into one code trains people to ignore the signal.

Two behaviours worth knowing:

  • A missing input is a skip, not a failure. A case with no reference cannot be exact-matched; it is reported as skipped and excluded from the means. Counting it as 0 would blend "we did not check this" into "this was wrong" and read as a regression that never happened.
  • An evaluator with no threshold is reported but never fails the run. Reporting and gating are separate decisions; conflating them makes people delete checks instead of fixing them.

trajectory_match compares the tool calls an agent made against a reference run — message content is never compared, so a reworded explanation is the same trajectory. superset is the one that catches a dropped step:

{
  "type": "trajectory_match",
  "match_mode": "superset",
  "overrides": { "issue_refund": { "type": "on_keys", "paths": ["order_id"] } }
}

The per-tool on_keys override is what makes this usable on a real agent: pin the order id, ignore the timestamp that changes every run.

rp eval rubrics lists the built-in judge rubrics. Those are scored by an operator-armed evaluation run on the gateway, not by rp eval.

From the SDK:

const report = await client.evaluations.score(cases, [{ type: 'valid_json' }]);

Examples

Runnable snippets live in examples/:

File Shows
basic.ts Minimal Routeplane client — a drop-in OpenAI subclass
headers-only.ts Stock openai SDK + createHeaders for per-request steering
streaming.ts Streaming with the gateway's decision metadata
vercel-ai-sdk.ts Vercel AI SDK (@ai-sdk/openai) integration
resources.ts Non-OpenAI surfaces — status, logs, FinOps, prompts, cache
agentic-security.ts Guarding an agent loop — tool-call authorization, result inspection, run breakers, receipts
eval-suite.json A rp eval suite — JSON/text checks, a trajectory comparison, and gate thresholds

See examples/ for more.

Development

pnpm install
pnpm build   # turbo build across all packages
pnpm test    # vitest via turbo
pnpm lint    # tsc --noEmit type-check

Requires Node.js >= 18 (native fetch).

License

Apache-2.0

About

Official developer tooling for Routeplane AI Gateway — TypeScript SDK, CLI, and MCP Server

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages