Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
66 changes: 63 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,9 +74,10 @@ headers are typed.

## 3. Full — the `Routeplane` client with rich metadata

`Routeplane` subclasses `openai.OpenAI`, so everything works exactly as before —
but it wires up auth for you, lets you set default routing once, and can parse
the gateway's response headers into a typed `RouteplaneMeta`.
`Routeplane` subclasses `openai.OpenAI` and keeps its resource/request API. It
wires up auth, lets you set default routing once, and parses the gateway's
response headers into a typed `RouteplaneMeta`. Chat response tier handling is
provider-neutral at runtime; see the served-tier contract below.

```python
from routeplane import Routeplane
Expand Down Expand Up @@ -213,6 +214,65 @@ parses the gateway's `x-routeplane-*` response headers: `provider`, `trace_id`,
`request_id`, `cache`, `guardrails`, `hedged`, `shed`, `budget_remaining`,
`budget_warning`, `compliance_warning`, `pii_masked`, `idempotent_replayed`.

### Actual served tier

`create_with_meta()` returns a `RouteplaneChatCompletion`, and `stream_with_meta()`
yields `RouteplaneChatCompletionChunk` objects, on both sync and async clients.
These exported models retain the installed OpenAI response fields and validation
while widening only returned `service_tier` to an optional strict string. Native
labels such as `standard`, `priority`, `batch`, and future provider labels keep
their exact spelling. Missing or `null` tier evidence remains `None`; serialization
with `exclude_unset=True` preserves omission versus explicit `null`. Non-string,
non-null tiers raise `APIResponseValidationError` even in default tolerant mode.
Unrelated response fields continue to follow the vendor's default or opt-in
`_strict_response_validation=True` behavior.

```python
from routeplane import Routeplane, RouteplaneChatCompletion

client = Routeplane(api_key="rp_your_gateway_key")
completion, meta = client.create_with_meta(
model="gpt-4o-mini", messages=[{"role": "user", "content": "Hello!"}],
)
assert isinstance(completion, RouteplaneChatCompletion)
print(completion.service_tier) # provider-reported evidence, possibly None

with client.stream_with_meta(
model="gpt-4o-mini", messages=[{"role": "user", "content": "Hello!"}],
) as stream:
for chunk in stream:
# Continue past a finish_reason: tier-only metadata may arrive afterward.
# choices may be empty, and usage may be omitted on that metadata chunk.
if chunk.service_tier is not None:
print(chunk.service_tier)
```

The returned tier is never inferred from a requested tier, renamed, or treated as
proof of a free response, tariff eligibility, complete cost, or billing authority.
Streaming keeps the vendor SSE decoder and `[DONE]` termination; no aggregation or
synthetic usage is added.

The inherited `client.chat.completions` resource signatures still expose the
**OpenAI vendor model types**. Routeplane's exact-model response-processing hook
also uses the neutral models at runtime for those calls, including raw responses,
but it does not widen the inherited static facade. Use the `*_with_meta` helpers
for a typed neutral result. On OpenAI versions that support custom raw parsing,
an explicit neutral model is also available:

```python
raw = client.chat.completions.with_raw_response.create(
model="gpt-4o-mini", messages=[{"role": "user", "content": "Hello!"}],
)
completion = raw.parse(to=RouteplaneChatCompletion)
```

OpenAI 1.0.0 has no `parse(to=...)` argument. The helpers and the runtime hook
support that advertised dependency floor using its existing response processor;
`raw.parse()` receives the neutral runtime model there too. A stock `openai.OpenAI`
client and custom response-model targets retain their own validation behavior.
Stream cleanup follows the installed vendor version; OpenAI 1.0.0 itself has no
`Stream.close()` or `AsyncStream.close()` method.

## Legacy feedback

Use the gateway-generated request ID from response metadata, not the provider's
Expand Down
3 changes: 3 additions & 0 deletions src/routeplane/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,12 +45,15 @@
TimeseriesData,
UsageData,
)
from .responses import RouteplaneChatCompletion, RouteplaneChatCompletionChunk

__all__ = [
"Routeplane",
"AsyncRouteplane",
"RouteplaneStream",
"AsyncRouteplaneStream",
"RouteplaneChatCompletion",
"RouteplaneChatCompletionChunk",
"headers",
"RouteplaneMeta",
"RouteplaneRateLimits",
Expand Down
23 changes: 19 additions & 4 deletions src/routeplane/async_client.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@

import httpx
import openai
from openai.types.chat import ChatCompletion, ChatCompletionChunk
from openai._types import ResponseT

from ._streaming import AsyncRouteplaneStream
from .client import DEFAULT_BASE_URL
Expand All @@ -35,6 +35,7 @@
ResidencyResource,
StatusResource,
)
from .responses import RouteplaneChatCompletion, RouteplaneChatCompletionChunk, _response_type

__all__ = ["AsyncRouteplane"]

Expand Down Expand Up @@ -139,18 +140,32 @@ def meta_from_headers(headers: Mapping[str, str]) -> RouteplaneMeta:
"""
return RouteplaneMeta.from_headers(headers)

async def create_with_meta(self, **kwargs: Any) -> Tuple[ChatCompletion, RouteplaneMeta]:
def _process_response_data(
self, *, data: object, cast_to: type[ResponseT], response: Any
) -> ResponseT:
# SDK versions use httpx or httpx2; forward the opaque response unchanged.
return super()._process_response_data(
data=data,
cast_to=_response_type(cast_to=cast_to, data=data, response=response),
response=response,
)

async def create_with_meta(
self, **kwargs: Any
) -> Tuple[RouteplaneChatCompletion, RouteplaneMeta]:
"""Chat completion that also returns the gateway :class:`RouteplaneMeta`.

::

completion, meta = await client.create_with_meta(model="gpt-4o", messages=[...])
"""
raw = await self.chat.completions.with_raw_response.create(**kwargs)
completion = cast(ChatCompletion, raw.parse())
completion = cast(RouteplaneChatCompletion, raw.parse())
return completion, RouteplaneMeta.from_headers(raw.headers)

async def stream_with_meta(self, **kwargs: Any) -> AsyncRouteplaneStream[ChatCompletionChunk]:
async def stream_with_meta(
self, **kwargs: Any
) -> AsyncRouteplaneStream[RouteplaneChatCompletionChunk]:
"""Streaming chat completion that also exposes ``meta`` on the stream.

::
Expand Down
25 changes: 19 additions & 6 deletions src/routeplane/client.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
"""The :class:`Routeplane` client — a thin subclass of ``openai.OpenAI``.

Everything the OpenAI SDK can do works unchanged (``client.chat.completions``,
``client.embeddings``, streaming, retries). The subclass adds three things:
The OpenAI resource API remains available (``client.chat.completions``,
``client.embeddings``, streaming, retries). Chat responses use provider-neutral
returned-tier models at runtime; inherited resource annotations stay vendor-owned.
The subclass adds three things:

1. gateway auth + default ``x-routeplane-*`` routing headers, so callers don't
repeat them on every request;
Expand All @@ -17,7 +19,7 @@

import httpx
import openai
from openai.types.chat import ChatCompletion, ChatCompletionChunk
from openai._types import ResponseT

from ._streaming import RouteplaneStream
from .headers import headers as build_headers
Expand All @@ -35,6 +37,7 @@
ResidencyResource,
StatusResource,
)
from .responses import RouteplaneChatCompletion, RouteplaneChatCompletionChunk, _response_type

__all__ = ["Routeplane"]

Expand Down Expand Up @@ -153,7 +156,17 @@ def meta_from_headers(headers: Mapping[str, str]) -> RouteplaneMeta:
"""
return RouteplaneMeta.from_headers(headers)

def create_with_meta(self, **kwargs: Any) -> Tuple[ChatCompletion, RouteplaneMeta]:
def _process_response_data(
self, *, data: object, cast_to: type[ResponseT], response: Any
) -> ResponseT:
# SDK versions use httpx or httpx2; forward the opaque response unchanged.
return super()._process_response_data(
data=data,
cast_to=_response_type(cast_to=cast_to, data=data, response=response),
response=response,
)

def create_with_meta(self, **kwargs: Any) -> Tuple[RouteplaneChatCompletion, RouteplaneMeta]:
"""Chat completion that also returns the gateway :class:`RouteplaneMeta`.

::
Expand All @@ -162,10 +175,10 @@ def create_with_meta(self, **kwargs: Any) -> Tuple[ChatCompletion, RouteplaneMet
print(meta.provider, meta.cache)
"""
raw = self.chat.completions.with_raw_response.create(**kwargs)
completion = cast(ChatCompletion, raw.parse())
completion = cast(RouteplaneChatCompletion, raw.parse())
return completion, RouteplaneMeta.from_headers(raw.headers)

def stream_with_meta(self, **kwargs: Any) -> RouteplaneStream[ChatCompletionChunk]:
def stream_with_meta(self, **kwargs: Any) -> RouteplaneStream[RouteplaneChatCompletionChunk]:
"""Streaming chat completion that also exposes ``meta`` on the stream.

``meta`` is populated from the response headers, which arrive before the
Expand Down
67 changes: 67 additions & 0 deletions src/routeplane/responses.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
"""OpenAI-shaped responses with provider-neutral actual served-tier evidence."""

from __future__ import annotations

from typing import Any, Optional, TypeVar, cast

from openai import APIResponseValidationError
from openai.types.chat import ChatCompletion, ChatCompletionChunk
from pydantic import StrictStr

__all__ = ["RouteplaneChatCompletion", "RouteplaneChatCompletionChunk"]


class RouteplaneChatCompletion(ChatCompletion):
"""A completion whose returned tier is an exact provider string or ``None``.

All other fields and validation come from the installed OpenAI model. The
label is evidence of what served the request, not a tariff or cost claim.
"""

# Intentionally widen the vendor literal union at this one wire-contract seam.
# The 1.0.0 model predates this field, so the assignment ignore is unused there.
service_tier: Optional[StrictStr] = None # type: ignore[assignment, unused-ignore]


class RouteplaneChatCompletionChunk(ChatCompletionChunk):
"""A stream chunk with the same provider-neutral returned-tier contract."""

# Same deliberate widening, including the dependency floor without this field.
service_tier: Optional[StrictStr] = None # type: ignore[assignment, unused-ignore]


_T = TypeVar("_T")


def _response_type(*, cast_to: type[_T], data: object, response: Any) -> type[_T]:
"""Select the neutral model and check only tier even in tolerant mode.

The client's existing response processor still controls validation of every
other field. Exact class matching leaves custom models and other APIs alone.
This hook is present at the supported OpenAI 1.0.0 dependency floor, where
raw-response ``parse(to=...)`` is not available yet.

The opaque vendor response is forwarded unchanged: supported SDK versions
use unrelated httpx/httpx2 classes and export no shared response type.
Data, model selection, and returned models retain their typed contracts.
"""
neutral: type[object]
if cast_to is ChatCompletion:
neutral = RouteplaneChatCompletion
elif cast_to is ChatCompletionChunk:
neutral = RouteplaneChatCompletionChunk
elif cast_to in (RouteplaneChatCompletion, RouteplaneChatCompletionChunk):
neutral = cast_to
else:
return cast_to

if isinstance(data, dict):
tier = data.get("service_tier")
if tier is not None and not isinstance(tier, str):
raise APIResponseValidationError(
response=response,
body=data,
message="Expected response service_tier to be a string or null.",
)

return cast("type[_T]", neutral)
Loading
Loading