Skip to content

fix(agents): subagents run at their definition's or the parent's effort; Explore caps at low - #955

Merged
ericleepi314 merged 1 commit into
mainfrom
fix/subagent-effort-latency
Sep 25, 2026
Merged

ericleepi314 merged 1 commit into
mainfrom
fix/subagent-effort-latency

Conversation

@ericleepi314

Copy link
Copy Markdown
Collaborator

Summary

Delegating to a subagent was slow at /effort max. On the ChatGPT subscription (gpt-6-astra), asking "what is this repo about?" took about 3 minutes in the TUI. 122 s of that was one Explore subagent that read a handful of files.

Root causes, all measured live:

  1. Subagents ignored their definition's and their parent's effort. run_agent sent no thinking_effort, so every subagent used the saved settings.effort. As a result:
    • Explore, the fast read-only agent, reasoned at max on every turn.
    • An effort: set in agent frontmatter was parsed and then dropped.
    • A session-only effort (headless --effort, or /effort on a client that doesn't save preferences) never reached subagents.
  2. The OpenAI provider had no tier table. Explore asks for the haiku tier. On OpenAI that fell back to the session model, which was the slowest GPT-6 tier.

Fix

  • query() records its effort level on the ToolContext (thinking_effort), next to the existing rendered_system_prompt capture.
  • run_agent.resolve_subagent_effort picks the subagent's level:
    • A user, project or plugin agent's own effort: is used as written.
    • Otherwise the subagent uses the parent's level (TS runAgent.ts:514-518).
    • A built-in agent's level is a ceiling. It is the lower of its own level and the level already in force. If the session configured no effort, it adds none, because Groq's Llama models and older Grok return 400 when reasoning_effort is present at all.
  • EXPLORE_AGENT.effort = "low". TS instead disables thinking for every non-fork subagent (runAgent.ts:720-725). That isn't copied here, because custom agents doing long proofs rely on the session's level.
  • OpenAI subagent_tier_models (haiku→gpt-6-luna, sonnet→gpt-6-sol, opus/fable→gpt-6-astra) apply only on the ChatGPT subscription. That's the one route whose model check reads the account's own catalog. API-key sessions and custom endpoints (base_url or $OPENAI_BASE_URL: LiteLLM, vLLM, Azure) keep inheriting the session model, as TS agent.ts:97-112 does. There is deliberately no OpenAI subagent_model, so agents that name no model keep the session model.

Measurements

All runs are live on the ChatGPT subscription with settings.effort=max. They use the real Explore loop, real tools and the same task.

model / effort sent Explore time (3 runs)
before gpt-6-astra / max 118, 124, 120 s
after, model: "inherit" (the slow session's exact call) gpt-6-astra / low 48, 60, 64 s
after, model left to the agent definition gpt-6-luna / low 31, 35, 32 s

The headless CLI, from prompt to answer (the model omitted the model param in all 4 runs), took 126 s and 89 s before the fix, and 44 s and 63 s after.

Test plan

  • tests/test_subagent_effort.py (new, 18 tests) covers:
    • precedence: an agent's own effort: wins, and a built-in's level acts as a ceiling
    • a session with no effort configured sends no effort field
    • query() records its level, including clearing a stale value
    • the effort in the subagent's request body
    • the full main → Agent tool → subagent → main chain, with the agent list pinned so the test is hermetic
  • tests/test_ch08_subagents_round4.py:
    • tier resolution on the subscription, and the catalog check
    • inherit
    • an API-key session, a custom base_url and $OPENAI_BASE_URL all inherit
  • tests/test_team_runtime_e2e.py and tests/workflow/test_runner_integration.py: teammate turns and workflow agents carry the parent's level.
  • Mutation-checked: each part of the fix, removed on its own, turns 3 to 8 tests red (dropping the effort in run_agent → 8; removing the ceiling, the subscription-only scope, or the query() capture → 3 each).
  • Full suite: 10,811 passed. The one failure, test_system_prompt_full::…damped_wording…, fails identically on main.
  • Critic review: approved in round 2. Round 1's two blocking findings were the custom-endpoint tier table and the effort field on unconfigured sessions; both are fixed above.

Not in this PR (follow-ups)

  • General-purpose and Plan subagents still run at the session's level (intended).
  • Subagents never receive the session's extended_thinking setting (existing gap).
  • Agent results don't report the effort they ran at.
  • Subagent system prompts lack environment details (TS enhanceSystemPromptWithEnvDetails).
  • Headless --effort now reaches subagents, so Terminal-Bench baselines that involve subagents will shift.

🤖 Generated with Claude Code

…rt; Explore caps at low

Delegating to a subagent took minutes at /effort max. A "what is this repo"
question on the ChatGPT subscription (gpt-6-astra) spent 122 s inside one
Explore subagent that read a handful of files.

Two causes, both measured live:

1. run_agent built the subagent's QueryParams with no thinking_effort, so
   the wire boundary fell back to the saved settings.effort for EVERY
   subagent. The built-in Explore agent, whose job is fast read-only
   search, reasoned at max on every turn. An effort: in agent frontmatter
   was parsed and dropped. A session-only level (headless --effort, an
   unsaved /effort) never reached subagents.
2. The OpenAI provider had no subagent tier table, so Explore's `haiku`
   inherited the session model, the slowest GPT-6 tier.

Fix:
- query() captures its level on ToolContext.thinking_effort, beside
  rendered_system_prompt.
- run_agent.resolve_subagent_effort: an authored definition's level is
  used as written, otherwise the parent's level (TS runAgent.ts:514-518).
- A BUILT-IN definition's level is a ceiling: min(own, level in force).
  It adds nothing when nothing is configured, because the field alone is
  a 400 on Groq's Llama models and older Grok.
- EXPLORE_AGENT.effort = "low". TS disables thinking for every non-fork
  subagent; that is deliberately not copied, because custom agents doing
  long proofs rely on the session level.
- OpenAI subagent_tier_models (haiku->gpt-6-luna, sonnet->gpt-6-sol,
  opus/fable->gpt-6-astra) apply ONLY on the ChatGPT subscription, whose
  availability gate reads the account's own catalog. API-key and
  custom-base-URL sessions inherit, as TS agent.ts:97-112 does.

Explore on the same task, 3 live runs each: 118-124 s before; 48-64 s
after with model "inherit"; 31-35 s after with the model left to the
definition (gpt-6-luna).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ericleepi314
ericleepi314 merged commit 9cc6808 into main Sep 25, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant