Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,32 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Fixed

- **Delegating to a subagent no longer takes minutes at `/effort max`.** A
subagent's requests carried no effort of their own, so every subagent fell
back to the saved `settings.effort`: the built-in Explore agent — the fast,
read-only search agent — reasoned at `max` on every turn; an `effort:` in an
agent definition's frontmatter was parsed and then ignored; and a
session-only level (headless `--effort`, or `/effort` on a client that does
not save preferences) never reached its subagents. A subagent now runs at
its definition's `effort`, else the level its parent runs at (the
reference's precedence). Explore caps its level at `low` — a ceiling, so a
session with no effort configured still sends none. On the ChatGPT
subscription (`gpt-6-astra`, `/effort max`), a "what is this repo" Explore
delegation fell from ~121 s to ~57 s (three runs each). This applies where
the wire takes an effort level; on first-party Anthropic, Explore runs on
`claude-haiku-4-5`, which takes none. Custom agents without an `effort:`
keep the session's level.
- **Subagent tier aliases resolve on the ChatGPT subscription.** `haiku` — the
tier the Explore agent asks for — now runs on `gpt-6-luna` (`sonnet` →
`gpt-6-sol`, `opus`/`fable` → `gpt-6-astra`) when your plan's model catalog
lists it. Before, every alias fell back to the session model (`gpt-6-astra`
in the reported session); with the model left to the agent definition the
same delegation takes ~33 s. API-key and custom-endpoint OpenAI sessions
keep inheriting the session model, as do an explicit `model: "inherit"` and
agents that name no model.

## [1.7.0] - 2026-09-23

### Added
Expand Down
19 changes: 16 additions & 3 deletions src/agent/agent_definitions.py
Original file line number Diff line number Diff line change
Expand Up @@ -165,10 +165,23 @@ def _explore_system_prompt(**_kwargs: Any) -> str:
# ch08 round-4 (critic M1) — Explore is the fast/cheap read-only agent;
# TS exploreAgent.ts:77 runs it on Haiku. get_agent_model resolves this
# per provider via the ``subagent_tier_models`` tables (anthropic →
# claude-haiku-4-5, deepseek → deepseek-v4-flash) and inherits on
# providers without a haiku-class mapping, so it is cross-provider
# safe.
# claude-haiku-4-5, deepseek → deepseek-flash, openai on the ChatGPT
# subscription → gpt-6-luna) and inherits the session model everywhere
# else (API-key openai, custom endpoints, providers without a table).
model="haiku",
# A CEILING, not a level of its own (run_agent.resolve_subagent_effort):
# Explore never reasons harder than low, and never adds an effort field
# to a session that configured none. Without it the fast agent reasoned
# at the session's /effort max on every search turn — a quick repo
# survey took ~121 s on gpt-6-astra vs ~57 s at low (same task and
# tools, three runs each, 2026-09-24). Only wires that take an effort
# level are affected: on first-party Anthropic Explore runs on
# claude-haiku-4-5, which takes none (its cost is its thinking budget).
# TS goes further — runAgent.ts:720-725 runs every non-fork subagent
# with thinking DISABLED — which this port deliberately does not copy:
# custom agents doing long proofs rely on the session level. A
# user/project ``Explore`` definition replaces this one.
effort="low",
get_system_prompt=_explore_system_prompt,
)

Expand Down
42 changes: 35 additions & 7 deletions src/agent/agent_model.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,12 +10,13 @@

* ``subagent_tier_models`` — the resolution targets for bare tier aliases
(anthropic haiku → ``claude-haiku-4-5``; deepseek haiku →
``deepseek-v4-flash``). This is the port of the TS reference's
``deepseek-flash``; openai haiku → ``gpt-6-luna``, on the ChatGPT
subscription only). This is the port of the TS reference's
per-provider ``getDefault{Opus,Sonnet,Haiku}Model()`` functions.
* ``subagent_model`` — the model subagents run on when neither the Agent
tool call nor the agent definition names one (anthropic:
``claude-haiku-4-5``, the cheapest current-gen tier; deepseek:
``deepseek-v4-flash``), overridable per provider via
``deepseek-flash``), overridable per provider via
``providers.<id>.subagent_model`` in config.json. NOTE this is a
DELIBERATE DIVERGENCE from both references, by explicit user directive
(cheap fan-outs): TS ``getDefaultSubagentModel()`` returns ``'inherit'``
Expand All @@ -38,7 +39,9 @@
the session model rather than 400-ing the request, and an unspecified model
inherits. A custom Anthropic-compatible endpoint (proxy / self-hosted) also
inherits rather than trusting the first-party table, mirroring TS
``checkIsClaudeNativeProvider``.
``checkIsClaudeNativeProvider``. The openai table applies only on the
ChatGPT subscription, the one route whose availability gate reads the
account's own catalog; API-key and custom-endpoint openai sessions inherit.

Deliberate asymmetry: a KNOWN alias whose canonical target is retired
degrades to inherit through the availability gate (``h35``,
Expand Down Expand Up @@ -118,19 +121,44 @@ def _is_custom_anthropic(session_provider: Any) -> bool:
return False


def _is_openai_without_live_catalog(session_provider: Any) -> bool:
"""An openai provider that is NOT on the ChatGPT subscription route.

The openai tier targets are trusted only where the availability gate is
a real per-account check: the subscription's Codex catalog
(``openai_subscription_models``). Everywhere else
``get_available_models()`` is the static registry list, so the gate
passes every listed id — an API key whose project may not use
``gpt-6-luna``, and a custom base URL (LiteLLM, vLLM, Azure,
``$OPENAI_BASE_URL``) that serves none of them, would both be sent a
model they cannot serve. Those keep the reference behavior and inherit
(TS agent.ts:97-112 inherits haiku/sonnet on non-Claude-native
providers). The subscription only ever activates against the first-party
endpoint (``OpenAIProvider.__init__``). Read with an ``isinstance``
check, like the agent-server's effort-options probe: a mock answers any
attribute with a truthy object.
"""
if _provider_id(session_provider) != "openai":
return False
active = getattr(_unwrap(session_provider), "_subscription_active", None)
return not (isinstance(active, bool) and active)


def _provider_info_row(session_provider: Any) -> dict[str, Any]:
"""The session provider's PROVIDER_INFO row, or ``{}``.

Empty when the provider carries no ``provider_id`` (unregistered) or is
an anthropic provider on a custom endpoint — both mean "no subagent
table", so every lookup falls through to the reference (inherit)
behavior.
Empty when the provider carries no ``provider_id`` (unregistered), is an
anthropic provider on a custom endpoint, or is an openai provider off
the subscription route — each means "no subagent table", so every lookup
falls through to the reference (inherit) behavior.
"""
provider_id = _provider_id(session_provider)
if not provider_id:
return {}
if _is_custom_anthropic(session_provider):
return {}
if _is_openai_without_live_catalog(session_provider):
return {}
try:
from src.providers import PROVIDER_INFO

Expand Down
43 changes: 43 additions & 0 deletions src/agent/run_agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -179,6 +179,45 @@ def _build_permission_context(
)


def resolve_subagent_effort(
agent_definition: AgentDefinition,
parent_context: ToolContext,
) -> str | None:
"""The reasoning-effort level a subagent's query runs at.

TS runAgent.ts:514-518: ``agentDefinition.effort ?? state.effortValue``
— the definition's own level, else the SESSION's. Here the session level
is whatever the spawning query runs at (``ToolContext.thinking_effort``,
captured by query() at turn entry). ``None`` leaves the wire boundary's
fallback to the persisted ``settings.effort`` — the same place the
parent's own requests land when it has no explicit level.

A BUILT-IN definition's level is a ceiling on the level already in
force, not a level of its own: with nothing configured it must keep
sending nothing, because the field's mere presence is a hard 400 on some
wires (Groq's Llama models, older Grok) where the parent sends none. A
user/project/plugin definition's level is the author's explicit choice
and is used as written (TS parity).

A definition level off the ladder (frontmatter also accepts integers,
TS's numeric effort, which no wire here takes) is treated as unset so
it falls through to the parent's level rather than to the settings
fallback.
"""
from ..query.query import VALID_THINKING_EFFORT_LEVELS, resolve_thinking_effort

parent = getattr(parent_context, "thinking_effort", None)
own = (agent_definition.effort or "").strip().lower()
if own not in VALID_THINKING_EFFORT_LEVELS:
return parent
if not is_built_in_agent(agent_definition):
return own
in_force = resolve_thinking_effort(parent, None, clamp_xhigh=False)
if in_force is None:
return parent
return min(own, in_force, key=VALID_THINKING_EFFORT_LEVELS.index)


def filter_incomplete_tool_calls(messages: list[Message]) -> list[Message]:
"""Remove assistant messages that contain incomplete tool_use blocks.

Expand Down Expand Up @@ -436,6 +475,10 @@ async def run_agent(params: RunAgentParams) -> AsyncGenerator[Message, None]:
abort_controller=abort_controller,
query_source=effective_query_source,
max_turns=max_turns,
# Unset, every subagent fell back to the persisted settings.effort:
# an ``effort:`` in the definition was dead, and a session-only level
# (headless --effort, an unsaved /effort) never reached the subagent.
thinking_effort=resolve_subagent_effort(agent_def, params.parent_context),
)

terminal = TerminalHolder()
Expand Down
18 changes: 18 additions & 0 deletions src/providers/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,24 @@ class ProviderInfo(_ProviderInfoOptional):
"label": "OpenAI GPT",
"default_base_url": "https://api.openai.com/v1",
"default_model": "gpt-5.4",
# Tier aliases → GPT-6 tiers, applied ONLY on the ChatGPT
# subscription route (see agent_model._is_openai_without_live_catalog):
# that is the one place the availability gate checks the account's
# own catalog. An API key's gate is this static list and a custom
# base URL serves none of these, so both keep inheriting — which is
# also what TS does for Explore on OpenAI-shaped providers
# (agent.ts:97-112); this row is a deliberate divergence for the
# subscription. Without it Explore's ``haiku`` inherited the session
# model: ~57 s for a repo survey on gpt-6-astra vs ~33 s on
# gpt-6-luna, both at effort low (measured 2026-09-24). Deliberately
# NO ``subagent_model``: an agent with no model keeps the session
# model on this provider.
"subagent_tier_models": {
"fable": "gpt-6-astra",
"opus": "gpt-6-astra",
"sonnet": "gpt-6-sol",
"haiku": "gpt-6-luna",
},
"available_models": [
# https://developers.openai.com/api/docs/models (2026-09-24)
# GPT-6 — Astra is the frontier tier, Sol the flagship, Luna the
Expand Down
10 changes: 10 additions & 0 deletions src/query/query.py
Original file line number Diff line number Diff line change
Expand Up @@ -1738,6 +1738,16 @@ async def query(
params.tool_use_context.rendered_system_prompt = params.system_prompt
except Exception: # noqa: BLE001 — read-only stub context
pass
# Same capture for the effort level: run_agent gives a subagent its
# definition's ``effort`` else THIS value (TS runAgent.ts:514-518).
# Without it a subagent's query carried no effort at all and fell back
# to the persisted ``settings.effort`` — a session-only level (headless
# ``--effort``, an unsaved ``/effort``) never reached it. Unconditional:
# ``None`` is meaningful (the parent itself runs on that fallback).
try:
params.tool_use_context.thinking_effort = params.thinking_effort
except Exception: # noqa: BLE001 — read-only stub context
pass

# How much of the session-lifetime outbox predates THIS query. Entries
# below the mark belong to earlier prompts and must not count as "the
Expand Down
10 changes: 10 additions & 0 deletions src/tool_system/context.py
Original file line number Diff line number Diff line change
Expand Up @@ -283,6 +283,16 @@ class ToolContext:
# _call_model_sync assembly → byte-identical wire prefix.
rendered_system_prompt: "str | list[dict[str, Any]] | None" = None

# The explicit reasoning-effort level of the query running on this
# context (``QueryParams.thinking_effort`` — the session's ``/effort``,
# headless ``--effort``); ``None`` = unset, so the wire boundary falls
# back to the persisted ``settings.effort``. POPULATED by query() at turn
# entry beside ``rendered_system_prompt``, so a subagent spawned during
# the turn inherits the level its parent actually runs at — TS
# runAgent.ts:514-518 reads the session ``effortValue`` from app state
# for the same purpose (see ``run_agent.resolve_subagent_effort``).
thinking_effort: str | None = None

def __post_init__(self) -> None:
self.workspace_root = Path(self.workspace_root).resolve()
if self.cwd is None:
Expand Down
73 changes: 70 additions & 3 deletions tests/test_ch08_subagents_round4.py
Original file line number Diff line number Diff line change
Expand Up @@ -129,6 +129,7 @@ def setUp(self):
"ANTHROPIC_DEFAULT_SONNET_MODEL",
"ANTHROPIC_DEFAULT_HAIKU_MODEL",
"ANTHROPIC_BASE_URL",
"OPENAI_BASE_URL",
):
os.environ.pop(var, None)
# Hermetic: the resolver consults the user's real config.json for
Expand All @@ -153,6 +154,72 @@ def _deepseek(model="deepseek-flash"):

return DeepSeekProvider(api_key="test-key", model=model)

@staticmethod
def _openai(model="gpt-6-astra", **kwargs):
from src.providers.openai_provider import OpenAIProvider

return OpenAIProvider(api_key="test-key", model=model, **kwargs)

def _openai_subscription(
self, model="gpt-6-astra", catalog=("gpt-6-astra", "gpt-6-sol", "gpt-6-luna"),
):
# A ChatGPT-subscription session without reading the developer's real
# login: the route flag plus the account catalog the gate consults.
p = self._openai(model)
p._subscription_active = True
p.get_available_models = lambda: list(catalog)
return p

def test_openai_subscription_haiku_tier_uses_luna(self):
# Explore pins 'haiku'. With no openai row the alias went through the
# global Claude-id table, which OpenAI does not serve, and inherited
# the session model — every Explore ran on gpt-6-astra (~57 s vs
# ~33 s on luna for the same repo survey, 2026-09-24).
self.assertEqual(
get_agent_model(None, "haiku", self._openai_subscription()), "gpt-6-luna",
)

def test_openai_subscription_catalog_without_the_tier_inherits(self):
# The live gate: a login whose Codex catalog lacks luna must not be
# sent a model its account cannot use.
p = self._openai_subscription(catalog=("gpt-6-astra", "gpt-5.5"))
self.assertEqual(get_agent_model(None, "haiku", p), "gpt-6-astra")

def test_openai_unspecified_model_still_inherits(self):
# Deliberately no openai ``subagent_model``: general-purpose and
# custom agents with no model keep the session model.
self.assertEqual(
get_agent_model(None, None, self._openai_subscription()), "gpt-6-astra",
)

def test_openai_explicit_inherit_beats_the_explore_tier(self):
# The slow session passed model: "inherit" on its Explore call; the
# tool param still wins over the definition's haiku.
self.assertEqual(
get_agent_model("inherit", "haiku", self._openai_subscription()),
"gpt-6-astra",
)

def test_openai_api_key_inherits(self):
# On an API key the gate is the static registry list, which passes
# gpt-6-luna whether or not the key's project may use it — so the
# table stays off and Explore keeps the session model it ran before.
self.assertEqual(get_agent_model(None, "haiku", self._openai()), "gpt-6-astra")

def test_openai_custom_base_url_inherits(self):
# LiteLLM / vLLM / Azure proxies serve none of the GPT-6 ids.
p = self._openai("my-litellm-model", base_url="http://localhost:4000/v1")
self.assertEqual(get_agent_model(None, "haiku", p), "my-litellm-model")

def test_openai_env_base_url_inherits(self):
# The other base-URL channel, with no key (the case the provider's
# own OAuth guard exists for): still no table.
from src.providers.openai_provider import OpenAIProvider

os.environ["OPENAI_BASE_URL"] = "https://api.deepseek.com/v1"
p = OpenAIProvider(api_key="", model="deepseek-v4-pro")
self.assertEqual(get_agent_model(None, "haiku", p), "deepseek-v4-pro")

def test_anthropic_unspecified_uses_default_subagent_model(self):
# Goal ask #1: on the anthropic provider the default subagent model
# is claude-haiku-4-5 (the cheapest current-gen tier), NOT an
Expand Down Expand Up @@ -321,9 +388,9 @@ def test_registry_subagent_targets_are_available(self):
f"{name} does not list — the runtime gate would "
"silently ignore it",
)
# Today: anthropic + deepseek. If this drops to zero the tables were
# deleted and this test should go with them.
self.assertGreaterEqual(checked, 2)
# Today: anthropic + deepseek + openai. If this drops to zero the
# tables were deleted and this test should go with them.
self.assertGreaterEqual(checked, 3)

def test_tier_env_pin_beats_table(self):
os.environ["ANTHROPIC_DEFAULT_HAIKU_MODEL"] = "my-bedrock-haiku"
Expand Down
Loading
Loading