Describe the bug
An OpenAI-compatible model configured with a 4,096 output limit can retry a truncated response at 64,000 tokens. At ded163b, query.py uses ESCALATED_MAX_TOKENS = 64000; the override bypasses modelLimits. The ordinary OpenAI request omits an output field, relying on the endpoint default.
Notify authors
@agentforce314, @ericleepi314
To Reproduce
On ClawCodex 1.7.0/Python 3.13.15, configure an OpenAI-compatible loopback provider, model tsubasa-fast, and settings.modelLimits.tsubasa-fast with contextWindow: 32768 and maxOutputTokens: 4096 in the user config.json. Run clawcodex --nano --provider openai --model tsubasa-fast --allowed-tools Read --max-turns 3 -p "Greet briefly".
Return a valid streamed text completion with finish_reason: "length". The captured first request has no max-token field; the actual retry has max_tokens: 64000, despite the settings loader correctly returning 32768/4096. Tsubasa's real shared request preflight rejects that retry because it exceeds the model's output ceiling. The CLI exits 1 after the controlled 400.
Expected behavior
Bound recovery output by the selected provider/model limit and available context, or return an explicit truncation error if a larger retry is impossible. Preserve current behavior for models whose limits support 64,000. Add a regression at the actual query/provider boundary, including known-model and configured unknown-model limits.
Screenshots
Not applicable: the evidence is a captured HTTP request from the print-mode CLI.
Desktop
macOS arm64; Python 3.13.15; ClawCodex 1.7.0.
Additional context
137 existing provider/config/context tests passed. Eight native nano CLI requests across both Tsubasa aliases passed streamed text, actual fixture-file read/tool-history round trips and 401 checks before this separate length-response probe. HOME remained unchanged; native configuration roots isolated the fixture. This is a controlled client failure reproduction, not a live model claim. No ready-for-full-agent claim or custom transport is proposed. The current main head 07567fad73f825353cfb9e5c55d11d5423bd9c5f has identical query.py and models/context.py contents to the tested pin. Related merged PR #830 introduced compatible-provider length recovery; this report isolates its unconditional output escalation for a configured smaller model.
Provider-affiliated investigation for Tsubasa, prepared with AI assistance.
Describe the bug
An OpenAI-compatible model configured with a 4,096 output limit can retry a truncated response at 64,000 tokens. At ded163b, query.py uses
ESCALATED_MAX_TOKENS = 64000; the override bypassesmodelLimits. The ordinary OpenAI request omits an output field, relying on the endpoint default.Notify authors
@agentforce314, @ericleepi314
To Reproduce
On ClawCodex 1.7.0/Python 3.13.15, configure an OpenAI-compatible loopback provider, model
tsubasa-fast, andsettings.modelLimits.tsubasa-fastwithcontextWindow: 32768andmaxOutputTokens: 4096in the userconfig.json. Runclawcodex --nano --provider openai --model tsubasa-fast --allowed-tools Read --max-turns 3 -p "Greet briefly".Return a valid streamed text completion with
finish_reason: "length". The captured first request has no max-token field; the actual retry hasmax_tokens: 64000, despite the settings loader correctly returning 32768/4096. Tsubasa's real shared request preflight rejects that retry because it exceeds the model's output ceiling. The CLI exits 1 after the controlled 400.Expected behavior
Bound recovery output by the selected provider/model limit and available context, or return an explicit truncation error if a larger retry is impossible. Preserve current behavior for models whose limits support 64,000. Add a regression at the actual query/provider boundary, including known-model and configured unknown-model limits.
Screenshots
Not applicable: the evidence is a captured HTTP request from the print-mode CLI.
Desktop
macOS arm64; Python 3.13.15; ClawCodex 1.7.0.
Additional context
137 existing provider/config/context tests passed. Eight native nano CLI requests across both Tsubasa aliases passed streamed text, actual fixture-file read/tool-history round trips and 401 checks before this separate length-response probe. HOME remained unchanged; native configuration roots isolated the fixture. This is a controlled client failure reproduction, not a live model claim. No ready-for-full-agent claim or custom transport is proposed. The current main head
07567fad73f825353cfb9e5c55d11d5423bd9c5fhas identical query.py and models/context.py contents to the tested pin. Related merged PR #830 introduced compatible-provider length recovery; this report isolates its unconditional output escalation for a configured smaller model.Provider-affiliated investigation for Tsubasa, prepared with AI assistance.