Skip to content

Length recovery ignores configured output limit and retries at 64000 tokens #958

Description

@cenab

Describe the bug

An OpenAI-compatible model configured with a 4,096 output limit can retry a truncated response at 64,000 tokens. At ded163b, query.py uses ESCALATED_MAX_TOKENS = 64000; the override bypasses modelLimits. The ordinary OpenAI request omits an output field, relying on the endpoint default.

Notify authors

@agentforce314, @ericleepi314

To Reproduce

On ClawCodex 1.7.0/Python 3.13.15, configure an OpenAI-compatible loopback provider, model tsubasa-fast, and settings.modelLimits.tsubasa-fast with contextWindow: 32768 and maxOutputTokens: 4096 in the user config.json. Run clawcodex --nano --provider openai --model tsubasa-fast --allowed-tools Read --max-turns 3 -p "Greet briefly".

Return a valid streamed text completion with finish_reason: "length". The captured first request has no max-token field; the actual retry has max_tokens: 64000, despite the settings loader correctly returning 32768/4096. Tsubasa's real shared request preflight rejects that retry because it exceeds the model's output ceiling. The CLI exits 1 after the controlled 400.

Expected behavior

Bound recovery output by the selected provider/model limit and available context, or return an explicit truncation error if a larger retry is impossible. Preserve current behavior for models whose limits support 64,000. Add a regression at the actual query/provider boundary, including known-model and configured unknown-model limits.

Screenshots

Not applicable: the evidence is a captured HTTP request from the print-mode CLI.

Desktop

macOS arm64; Python 3.13.15; ClawCodex 1.7.0.

Additional context

137 existing provider/config/context tests passed. Eight native nano CLI requests across both Tsubasa aliases passed streamed text, actual fixture-file read/tool-history round trips and 401 checks before this separate length-response probe. HOME remained unchanged; native configuration roots isolated the fixture. This is a controlled client failure reproduction, not a live model claim. No ready-for-full-agent claim or custom transport is proposed. The current main head 07567fad73f825353cfb9e5c55d11d5423bd9c5f has identical query.py and models/context.py contents to the tested pin. Related merged PR #830 introduced compatible-provider length recovery; this report isolates its unconditional output escalation for a configured smaller model.

Provider-affiliated investigation for Tsubasa, prepared with AI assistance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions