fix(server): forward echoed reasoning under reasoning_content too - #2094
Merged
Merged
Conversation
inureyes
added a commit
that referenced
this pull request
Oct 2, 2026
3 tasks done
AI21 Jamba-Reasoning-3B appeared to keep every reply in reasoning_content and missed the prompt cache on every turn. Measured on the real checkpoint, the empty content is finish_reason=length only: at max_tokens 64 every turn stops inside the primed <think> block, while at 2048 the model closes </think> and content is filled. Close-marker recognition works. The cache miss had a server-side cause. The raw-JSON render forwarded an echoed trace only as `reasoning`, but Qwen3-style and Jamba templates read `message.reasoning_content`. Jamba keeps its thinking instruction on an earlier user turn only when the next assistant message carries that field, so the turn re-rendered shorter than it was generated and the history-boundary snapshot never prefixed the next prompt. The trace now goes out under both spellings; templates that accept both read them as alternatives, so nothing renders twice. A client that echoes only content still gets the rewritten user turn and misses the cache; that is the template's own rule and is pinned by a test and documented. Real server, fresh process, 3 turns echoing reasoning_content: content non-empty each turn, cached_tokens 0/41/88 (was 0/0/0). Closes #2089
inureyes
force-pushed
the
fix/issue-2089-jamba-reasoning-content
branch
from
October 2, 2026 09:13
ea79329 to
dcaaaf9
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
jamba-v0.1-4bit, actually AI21-Jamba-Reasoning-3B-4bit): emptycontentisfinish_reason=lengthonly. Atmax_tokens64 all three turns stop inside the primed<think>block; at 2048 turns close</think>and fillcontent.</think>(id 542, not special) is recognized; no budget or marker change is warranted, matching other primed-thinking families.build_raw_json_messages_with_thinkingforwarded an echoed trace only asreasoning, while Qwen3-style and Jamba templates readmessage.reasoning_content. Jamba keeps its thinking instruction on an earlier user turn only when the following assistant message carries that field, so the turn re-rendered shorter and the history-boundary snapshot never prefixed the next prompt. The trace is now forwarded under both keys. Templates that accept both (Gemma 4, Laguna) read them as alternatives, so nothing renders twice.contentstill gets the earlier user turn without the instruction. That is the template's rule; it is pinned by a test, noted in the prefix-stability module doc, and documented indocs/supported-models.md.Test plan
server::chat_request::reasoning_content_tests(5 tests: both-key forwarding, Jamba boundary prefixes turn 2 with echoed reasoning, content-only rewrite pinned, no double render, primed-think split and truncation pinned);server::chat_requestfilter 126 passedcargo clippy --release --features cuda --lib --tests -- -D warnings,cargo fmt --checkreasoning_contentat max_tokens 2048: finish stop/stop/stop, contentParis./4/Louvre, cached_tokens 0/41/88 (before: prompt 66 ignored the echo, cached 0 on turns 2 and 3)Closes #2089