Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion eval/harbor/RUN_LUNA_TB21.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ clawcodex Harbor adapter.
| | |
|---|---|
| id | `openai/gpt-5.6-luna` (a `-pro` variant exists with identical specs) |
| context | 1,050,000 tokens (registered as 1,048,576 = 2^20, matching the sibling gpt-5.6 rows; under-reading is the safe direction) |
| context | 1,050,000 advertised; 922K max input (registered as 872,000 since 2026-09-24, the smaller of the API and ChatGPT-subscription input limits; runs past ~822K now compact where older builds overflowed, so such runs are not like-for-like with earlier baselines) |
| max output | 128,000 tokens |
| price | $0.10/M in, $0.60/M out — doubling to $0.20 / $0.90 above 272K prompt tokens |
| reasoning | `reasoning_effort` supported; verified honored, not just accepted — reasoning tokens rise low 148 → medium 154 → high 266 → max 516 on a fixed prompt |
Expand Down
27 changes: 15 additions & 12 deletions eval/harbor/clawcodex_agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -113,19 +113,22 @@
path ``openrouter/openai/gpt-5.6-luna`` takes, and OpenRouter accepts
``max`` for that model.
- **First-party ``--provider openai``** — reasoning models go over the
Responses API, which tops out at ``xhigh``: ``max`` is degraded to
``xhigh`` rather than sent (the API rejects ``max`` outright with
"Supported values are: 'none', 'low', 'medium', 'high', and 'xhigh'").
So ``--ak effort=max`` against ``openai/gpt-5.6-luna`` really runs at
``xhigh``, while the same nominal setting on OpenRouter sends ``max``.
Non-reasoning models (gpt-4o, …) stay on Chat Completions and have the
effort field STRIPPED, since they reject it as an unknown argument.
Responses API. GPT-6 (astra/sol/luna) accepts ``max``; every other
model tops out at ``xhigh``, so ``max`` is degraded to ``xhigh`` rather
than sent (gpt-5.6-luna rejects it outright with "Supported values are:
'none', 'low', 'medium', 'high', and 'xhigh'"). So ``--ak effort=max``
against ``openai/gpt-5.6-luna`` really runs at ``xhigh``, while the same
nominal setting on OpenRouter sends ``max``. Non-reasoning models
(gpt-4o, …) stay on Chat Completions and have the effort field
STRIPPED, since they reject it as an unknown argument.
- **ChatGPT subscription** (``subscription=true`` with
``--model openai/…``) clamps ``xhigh`` and ``max`` down to ``high``,
because that backend advertises only low/medium/high. So the same
``effort=max`` runs at three different levels depending on route:
``max`` on OpenRouter, ``xhigh`` on an OpenAI key, ``high`` on a
ChatGPT plan.
``--model openai/…``) clamps PER MODEL to the levels the login's Codex
catalog advertises (static fallback when uncached), probed 2026-09-24:
gpt-6-* and gpt-5.6-* take ``max``; gpt-5.5 stops at ``xhigh``; older
models stop at ``high``. So ``effort=max`` on gpt-5.6-luna runs at
``max`` on OpenRouter and on a ChatGPT plan, but ``xhigh`` on an OpenAI
key. Builds before 2026-09-24 clamped every subscription model to
``high``.

REQUIRES a clawcodex build from 2026-07-31 or later. Before that, effort
was emitted ONLY on the Anthropic branch, so ``--effort`` was a silent
Expand Down
53 changes: 45 additions & 8 deletions src/models/configs.py
Original file line number Diff line number Diff line change
Expand Up @@ -722,30 +722,67 @@ class ModelConfig:
# entries so they never take that path. As with the Meta entry above,
# max_output_tokens is not sent on the wire for OpenAI providers — its
# live effect is the auto-compact output reservation (clamped at 20K).
#
# --- GPT-6 -------------------------------------------------------------
# GPT-6 (Astra / Sol / Luna). developers.openai.com model pages
# (2026-09-24): 1,050,000 context, 922K max INPUT, 128K max output. The
# ChatGPT subscription is tighter still: its Codex catalog gives every
# gpt-6 model ``max_context_window: 872000``. ``context_window`` here
# drives the auto-compact threshold, and OpenAI's context-overflow error
# is not one reactive compaction recognises, so over-estimating is a
# hard failure while under-estimating only compacts early. Hence 872K —
# the smaller of the two real input limits, safe on both backends.
# Placed before every gpt-5.x row for the same prefix-fallback reason as
# GPT-5.6 below: each has base "gpt-6", so an unlisted variant
# (``gpt-6-sol-pro``) lands here rather than on the 272K catch-all.
"gpt-6-astra": ModelConfig(
model_id="gpt-6-astra",
display_name="GPT-6 Astra",
context_window=872_000,
max_output_tokens=128_000,
),
"gpt-6-sol": ModelConfig(
model_id="gpt-6-sol",
display_name="GPT-6 Sol",
context_window=872_000,
max_output_tokens=128_000,
),
"gpt-6-luna": ModelConfig(
model_id="gpt-6-luna",
display_name="GPT-6 Luna",
context_window=872_000,
max_output_tokens=128_000,
),
# GPT-5.6 (Sol / Terra / Luna — three durable capability tiers on one
# generation, 1.05M context each). These keys sit BEFORE "gpt-5.5"
# generation, 1.05M advertised context each). These keys sit BEFORE "gpt-5.5"
# deliberately: ``get_model_config``'s prefix fallback walks in insertion
# order and each of these has base "gpt-5.6" (rsplit on the last "-"), so
# putting them first is what lets an unlisted variant like
# ``gpt-5.6-sol-pro`` resolve to 1.05M instead of falling through to the
# ``gpt-5.6-sol-pro`` resolve to the 5.6 window instead of falling through to the
# 272K catch-all. The bare ``gpt-5.6`` alias is deliberately NOT here —
# see below.
# Window is 872K, not the advertised 1.05M, for the same reason as the
# GPT-6 rows above: the public API caps INPUT at 922K
# (developers.openai.com/api/docs/models/gpt-5.6-sol) and the ChatGPT
# subscription catalog caps it at ``max_context_window: 872000``. At
# 1.05M auto-compact fired near 998K — past both limits, into an
# overflow error that reactive compaction does not recognise.
"gpt-5.6-sol": ModelConfig(
model_id="gpt-5.6-sol",
display_name="GPT-5.6 Sol",
context_window=1_048_576,
context_window=872_000,
max_output_tokens=128_000,
),
"gpt-5.6-terra": ModelConfig(
model_id="gpt-5.6-terra",
display_name="GPT-5.6 Terra",
context_window=1_048_576,
context_window=872_000,
max_output_tokens=128_000,
),
"gpt-5.6-luna": ModelConfig(
model_id="gpt-5.6-luna",
display_name="GPT-5.6 Luna",
context_window=1_048_576,
context_window=872_000,
max_output_tokens=128_000,
),
# The same model as the row above, under its OpenRouter id. This table is
Expand All @@ -770,7 +807,7 @@ class ModelConfig:
"openai/gpt-5.6-luna": ModelConfig(
model_id="openai/gpt-5.6-luna",
display_name="GPT-5.6 Luna",
context_window=1_048_576,
context_window=872_000,
max_output_tokens=128_000,
),
"gpt-5.5": ModelConfig(
Expand All @@ -781,14 +818,14 @@ class ModelConfig:
),
# ``gpt-5.6`` is OpenAI's alias for Sol. Its base is "gpt" (not
# "gpt-5.6"), so it would become the catch-all for EVERY unknown gpt id if
# it preceded "gpt-5.5" — handing them a 1.05M window. Over-estimating
# it preceded "gpt-5.5" — handing them an 872K window. Over-estimating
# overflows the context; under-estimating only compacts early, so the
# catch-all must stay on the 272K entry. Exact lookups are unaffected by
# position: ``get_model_config`` checks for an exact key first.
"gpt-5.6": ModelConfig(
model_id="gpt-5.6",
display_name="GPT-5.6 (Sol)",
context_window=1_048_576,
context_window=872_000,
max_output_tokens=128_000,
),
"gpt-5.4": ModelConfig(
Expand Down
7 changes: 6 additions & 1 deletion src/providers/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -89,8 +89,13 @@ class ProviderInfo(_ProviderInfoOptional):
"default_base_url": "https://api.openai.com/v1",
"default_model": "gpt-5.4",
"available_models": [
# https://developers.openai.com/api/docs/models (2026-09-06)
# https://developers.openai.com/api/docs/models (2026-09-24)
# GPT-6 — Astra is the frontier tier, Sol the flagship, Luna the
# cheap high-volume tier. All three are also served by the
# ChatGPT subscription (Codex catalog, client_version >= 0.155).
"gpt-6-astra",
"gpt-6-sol",
"gpt-6-luna",
# GPT-5.6 — Sol / Terra / Luna are
# durable capability tiers rather than a size ladder: Sol is the
# flagship, Terra balances capability against cost, Luna is the
Expand Down
50 changes: 41 additions & 9 deletions src/providers/effort_options.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,22 +31,33 @@
AUTO = "auto"


def effort_options(provider_name: str | None, model: str | None) -> dict[str, object]:
def effort_options(
provider_name: str | None,
model: str | None,
*,
openai_subscription: bool | None = None,
) -> dict[str, object]:
"""The effort levels ``model`` accepts under ``provider_name``.

Returns ``{"supported": bool, "levels": [...]}``. ``supported`` False
means the model takes no effort parameter at all — the picker skips its
third step rather than offering a list that cannot be applied. ``levels``
never includes ``auto``; the caller prepends it, since "let the provider
decide" is meaningful exactly when some real level is also on offer.

``openai_subscription`` says whether OpenAI requests ride the ChatGPT
login rather than an API key — the two backends accept different levels
for the same model. ``None`` infers it from stored credentials.
"""
canonical = _canonical(provider_name)

if canonical == "anthropic":
return _anthropic_options(model)

if canonical == "openai":
return _openai_options(model)
if openai_subscription is None:
openai_subscription = _openai_subscription_in_use()
return _openai_options(model, subscription=openai_subscription)

# Every other provider: the codebase carries no per-model effort table,
# and the OpenAI-compatible paths pass the value through as a body field
Expand Down Expand Up @@ -92,18 +103,39 @@ def _anthropic_options(model: str | None) -> dict[str, object]:
}


def _openai_options(model: str | None) -> dict[str, object]:
def _openai_subscription_in_use() -> bool:
"""Mirror ``OpenAIProvider.__init__``: a configured key wins, otherwise a
stored ChatGPT login is used. (The base-URL guard is not mirrored; a
proxy with no key and a stale login is an edge the wire clamp absorbs.)"""
try:
from src.auth.openai_subscription import load_credentials
from src.providers import resolve_api_key

return not resolve_api_key("openai") and load_credentials() is not None
except Exception: # noqa: BLE001 — a picker query must not fail on this
return False


def _openai_options(model: str | None, *, subscription: bool = False) -> dict[str, object]:
"""OpenAI: the Responses API rejects a reasoning block outright on
non-reasoning models, so those get no third step at all."""
from src.providers.openai_responses import OPENAI_REASONING_EFFORTS, supports_reasoning
non-reasoning models, so those get no third step at all. The rest get the
per-model ladder the request path clamps to, so every offered level is
sent as-is rather than silently downgraded."""
from src.providers.openai_responses import api_effort_levels, supports_reasoning

if not supports_reasoning(model or ""):
return {"supported": False, "levels": []}

# Intersect rather than pass through: OPENAI_REASONING_EFFORTS carries
# ``none``, which ``/effort`` would reject, and omits ``max``, which
# OpenAI 400s on for the same model OpenRouter tolerates it for.
if subscription:
from src.providers.openai_subscription_models import get_subscription_effort_levels

accepted = get_subscription_effort_levels(model or "")
else:
accepted = api_effort_levels(model or "")

# Intersect rather than pass through: the OpenAI lists carry ``none``
# (and ``minimal``), which ``/effort`` would reject.
return {
"supported": True,
"levels": [lvl for lvl in LADDER if lvl in OPENAI_REASONING_EFFORTS],
"levels": [lvl for lvl in LADDER if lvl in accepted],
}
58 changes: 31 additions & 27 deletions src/providers/openai_provider.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@
RESPONSES_ITEM_BLOCK_TYPE,
INCLUDE_ENCRYPTED_REASONING,
SUBSCRIPTION_MODELS,
clamp_effort,
normalize_openai_effort,
supports_reasoning,
build_usage_dict,
Expand All @@ -73,10 +74,9 @@

logger = logging.getLogger(__name__)

_REASONING_EFFORTS = ("minimal", "low", "medium", "high", "xhigh")


def _subscription_reasoning_effort(requested: str | None = None) -> str:
def _subscription_reasoning_effort(
requested: str | None = None, model: str = "",
) -> str | None:
"""Reasoning effort for subscription requests.

Precedence: the session's ``/effort`` setting (arrives as
Expand All @@ -86,25 +86,27 @@ def _subscription_reasoning_effort(requested: str | None = None) -> str:
default, transform.ts:1176, and the backend's own
default_reasoning_level).

``xhigh``/``max`` clamp to ``high`` HERE, and only here: this is the
ChatGPT-subscription backend (chatgpt.com/backend-api/codex), whose
general gpt-5.x models advertise low/medium/high and reject higher
tiers (probed 2026-07-25). That is narrower than the public API —
developers.openai.com/api/docs/guides/reasoning lists none | minimal |
low | medium | high | xhigh | max and notes support varies by model —
and narrower than what a gateway may accept (``openai/gpt-5.6-luna``
via OpenRouter takes both ``xhigh`` and ``max``, probed 2026-07-31,
with reasoning-token counts rising monotonically across the ladder).
So the clamp is a property of THIS backend, not of the level names;
the generic OpenAI-compatible path deliberately does not clamp.
The requested level is clamped to what THIS model takes on the ChatGPT
backend (chatgpt.com/backend-api/codex), which varies per model: probed
2026-09-24, gpt-6-* and gpt-5.6-* accept ``max``, gpt-5.5 accepts
``xhigh`` but 400s on ``max``. The per-model list comes from this login's
cached model catalog (``supported_reasoning_levels``), falling back to a
static table — see ``get_subscription_effort_levels``. This used to clamp
``xhigh``/``max`` to ``high`` for every model (true of the gpt-5.x models
probed 2026-07-25), which silently capped gpt-6 two notches low.

The generic OpenAI-compatible path deliberately does not clamp: a gateway
(``openai/gpt-5.6-luna`` via OpenRouter) takes both ``xhigh`` and ``max``.
"""
from .openai_subscription_models import get_subscription_effort_levels

levels = get_subscription_effort_levels(model)
for candidate in (requested, os.environ.get("CLAWCODEX_OPENAI_REASONING_EFFORT")):
effort = (candidate or "").strip().lower()
if effort in ("xhigh", "max"):
return "high"
if effort in _REASONING_EFFORTS:
return effort
return "medium"
if (candidate or "").strip():
effort = clamp_effort(candidate, levels)
if effort:
return effort
return clamp_effort("medium", levels)


class _HttpxStreamHolder:
Expand Down Expand Up @@ -487,16 +489,18 @@ def _subscription_request_body(
# model rather than only the reasoning ones.
if supports_reasoning(model):
if self._subscription_active:
# The ChatGPT backend advertises only low/medium/high and
# rejects higher tiers, so it keeps its own clamp.
# The ChatGPT backend's levels vary per model, so it keeps
# its own catalog-driven clamp.
effort = _subscription_reasoning_effort(
(kwargs.get("extra_body") or {}).get("reasoning_effort")
(kwargs.get("extra_body") or {}).get("reasoning_effort"),
model,
)
else:
# The public API accepts xhigh; only ``max`` is unsupported,
# and it degrades rather than failing the request.
# The public API accepts xhigh everywhere and ``max`` on
# gpt-6 only; elsewhere ``max`` degrades rather than failing.
effort = normalize_openai_effort(
(kwargs.get("extra_body") or {}).get("reasoning_effort")
(kwargs.get("extra_body") or {}).get("reasoning_effort"),
model,
)
if effort:
body["reasoning"] = {"effort": effort, "summary": "auto"}
Expand Down
Loading
Loading