fix(sdk): give an LLM call 300 s and retry a timeout only once - #39
Merged
Merged
Conversation
A provider request is not streamed, so its timeout bounds the whole response. At ~80 output tokens/s the old 120 s default capped one call at ~9.5k output tokens. An agent writing several large files in one turn hit that cap, and every retry sent the identical request and timed out the same way — five 120 s attempts, then the agent errored. - Default provider request timeout: 120 s -> 300 s, shared by the OpenRouter and Anthropic providers. - A timeout gets two attempts in total. Other retryable errors (rate_limit, server_error, network_error) keep five. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MmH33W8tV2jhB25w5GQ4E4
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Provider requests are not streamed, so the request timeout bounds the whole response. With the 120 s default and ~80 output tokens/s, one call can produce at most ~9.5k output tokens.
In production, an agent that wrote several large files in one turn hit this cap. Each retry sent the identical request and timed out the same way: 5 × 120 s, then the agent errored. A replacement agent failed the same way.
Change
DEFAULT_PROVIDER_REQUEST_TIMEOUT_MS = 300_000inprovider-request.ts, used by the OpenRouter and Anthropic providers. An explicittimeoutin config still overrides it.withRetryaccepts an optionalgetMaxAttempts(error), which can only lower the attempt cap.withLLMRetryuses it to captimeouterrors at 2 attempts in total.rate_limit,server_errorandnetwork_errorkeep 5.Worst case for a stuck call stays ~10 min (2 × 300 s instead of 5 × 120 s), but each attempt now has 2.5× longer to finish.
Tests
retry.test.ts: a timeout gives up after 2 attempts;server_errorkeeps the full 5.maxAttemptstest now usesserver_error, because a timeout is capped at 2.bun test src/core/agents src/core/llm,tsc --noEmitandbiome lintpass.🤖 Generated with Claude Code
https://claude.ai/code/session_01MmH33W8tV2jhB25w5GQ4E4