Skip to content

fix(sdk): give an LLM call 300 s and retry a timeout only once - #39

Merged
matej21 merged 1 commit into
mainfrom
fix/llm-timeout-retries
Sep 23, 2026
Merged

matej21 merged 1 commit into
mainfrom
fix/llm-timeout-retries

Conversation

@matej21

@matej21 matej21 commented Sep 23, 2026

Copy link
Copy Markdown
Member

Problem

Provider requests are not streamed, so the request timeout bounds the whole response. With the 120 s default and ~80 output tokens/s, one call can produce at most ~9.5k output tokens.

In production, an agent that wrote several large files in one turn hit this cap. Each retry sent the identical request and timed out the same way: 5 × 120 s, then the agent errored. A replacement agent failed the same way.

Change

  • DEFAULT_PROVIDER_REQUEST_TIMEOUT_MS = 300_000 in provider-request.ts, used by the OpenRouter and Anthropic providers. An explicit timeout in config still overrides it.
  • withRetry accepts an optional getMaxAttempts(error), which can only lower the attempt cap. withLLMRetry uses it to cap timeout errors at 2 attempts in total. rate_limit, server_error and network_error keep 5.

Worst case for a stuck call stays ~10 min (2 × 300 s instead of 5 × 120 s), but each attempt now has 2.5× longer to finish.

Tests

  • retry.test.ts: a timeout gives up after 2 attempts; server_error keeps the full 5.
  • The existing maxAttempts test now uses server_error, because a timeout is capped at 2.
  • bun test src/core/agents src/core/llm, tsc --noEmit and biome lint pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MmH33W8tV2jhB25w5GQ4E4

A provider request is not streamed, so its timeout bounds the whole
response. At ~80 output tokens/s the old 120 s default capped one call at
~9.5k output tokens. An agent writing several large files in one turn hit
that cap, and every retry sent the identical request and timed out the
same way — five 120 s attempts, then the agent errored.

- Default provider request timeout: 120 s -> 300 s, shared by the
  OpenRouter and Anthropic providers.
- A timeout gets two attempts in total. Other retryable errors
  (rate_limit, server_error, network_error) keep five.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmH33W8tV2jhB25w5GQ4E4
@matej21
matej21 merged commit 1a24fdf into main Sep 23, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant