fix(agents): subagents run at their definition's or the parent's effort; Explore caps at low - #955
Merged
Merged
Conversation
…rt; Explore caps at low Delegating to a subagent took minutes at /effort max. A "what is this repo" question on the ChatGPT subscription (gpt-6-astra) spent 122 s inside one Explore subagent that read a handful of files. Two causes, both measured live: 1. run_agent built the subagent's QueryParams with no thinking_effort, so the wire boundary fell back to the saved settings.effort for EVERY subagent. The built-in Explore agent, whose job is fast read-only search, reasoned at max on every turn. An effort: in agent frontmatter was parsed and dropped. A session-only level (headless --effort, an unsaved /effort) never reached subagents. 2. The OpenAI provider had no subagent tier table, so Explore's `haiku` inherited the session model, the slowest GPT-6 tier. Fix: - query() captures its level on ToolContext.thinking_effort, beside rendered_system_prompt. - run_agent.resolve_subagent_effort: an authored definition's level is used as written, otherwise the parent's level (TS runAgent.ts:514-518). - A BUILT-IN definition's level is a ceiling: min(own, level in force). It adds nothing when nothing is configured, because the field alone is a 400 on Groq's Llama models and older Grok. - EXPLORE_AGENT.effort = "low". TS disables thinking for every non-fork subagent; that is deliberately not copied, because custom agents doing long proofs rely on the session level. - OpenAI subagent_tier_models (haiku->gpt-6-luna, sonnet->gpt-6-sol, opus/fable->gpt-6-astra) apply ONLY on the ChatGPT subscription, whose availability gate reads the account's own catalog. API-key and custom-base-URL sessions inherit, as TS agent.ts:97-112 does. Explore on the same task, 3 live runs each: 118-124 s before; 48-64 s after with model "inherit"; 31-35 s after with the model left to the definition (gpt-6-luna). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Delegating to a subagent was slow at
/effort max. On the ChatGPT subscription (gpt-6-astra), asking "what is this repo about?" took about 3 minutes in the TUI. 122 s of that was oneExploresubagent that read a handful of files.Root causes, all measured live:
run_agentsent nothinking_effort, so every subagent used the savedsettings.effort. As a result:maxon every turn.effort:set in agent frontmatter was parsed and then dropped.--effort, or/efforton a client that doesn't save preferences) never reached subagents.haikutier. On OpenAI that fell back to the session model, which was the slowest GPT-6 tier.Fix
query()records its effort level on theToolContext(thinking_effort), next to the existingrendered_system_promptcapture.run_agent.resolve_subagent_effortpicks the subagent's level:effort:is used as written.runAgent.ts:514-518).reasoning_effortis present at all.EXPLORE_AGENT.effort = "low". TS instead disables thinking for every non-fork subagent (runAgent.ts:720-725). That isn't copied here, because custom agents doing long proofs rely on the session's level.subagent_tier_models(haiku→gpt-6-luna,sonnet→gpt-6-sol,opus/fable→gpt-6-astra) apply only on the ChatGPT subscription. That's the one route whose model check reads the account's own catalog. API-key sessions and custom endpoints (base_urlor$OPENAI_BASE_URL: LiteLLM, vLLM, Azure) keep inheriting the session model, as TSagent.ts:97-112does. There is deliberately no OpenAIsubagent_model, so agents that name no model keep the session model.Measurements
All runs are live on the ChatGPT subscription with
settings.effort=max. They use the real Explore loop, real tools and the same task.model: "inherit"(the slow session's exact call)The headless CLI, from prompt to answer (the model omitted the
modelparam in all 4 runs), took 126 s and 89 s before the fix, and 44 s and 63 s after.Test plan
tests/test_subagent_effort.py(new, 18 tests) covers:effort:wins, and a built-in's level acts as a ceilingquery()records its level, including clearing a stale valuetests/test_ch08_subagents_round4.py:inheritbase_urland$OPENAI_BASE_URLall inherittests/test_team_runtime_e2e.pyandtests/workflow/test_runner_integration.py: teammate turns and workflow agents carry the parent's level.run_agent→ 8; removing the ceiling, the subscription-only scope, or thequery()capture → 3 each).test_system_prompt_full::…damped_wording…, fails identically onmain.Not in this PR (follow-ups)
extended_thinkingsetting (existing gap).enhanceSystemPromptWithEnvDetails).--effortnow reaches subagents, so Terminal-Bench baselines that involve subagents will shift.🤖 Generated with Claude Code