Skip to content

fix(server): forward echoed reasoning under reasoning_content too - #2094

Merged
inureyes merged 2 commits into
mainfrom
fix/issue-2089-jamba-reasoning-content
Oct 2, 2026
Merged

inureyes merged 2 commits into
mainfrom
fix/issue-2089-jamba-reasoning-content

Conversation

@inureyes

@inureyes inureyes commented Oct 2, 2026

Copy link
Copy Markdown
Member

Summary

  • Investigation on the real checkpoint (jamba-v0.1-4bit, actually AI21-Jamba-Reasoning-3B-4bit): empty content is finish_reason=length only. At max_tokens 64 all three turns stop inside the primed <think> block; at 2048 turns close </think> and fill content. </think> (id 542, not special) is recognized; no budget or marker change is warranted, matching other primed-thinking families.
  • The cache miss is a server bug: build_raw_json_messages_with_thinking forwarded an echoed trace only as reasoning, while Qwen3-style and Jamba templates read message.reasoning_content. Jamba keeps its thinking instruction on an earlier user turn only when the following assistant message carries that field, so the turn re-rendered shorter and the history-boundary snapshot never prefixed the next prompt. The trace is now forwarded under both keys. Templates that accept both (Gemma 4, Laguna) read them as alternatives, so nothing renders twice.
  • A client that echoes only content still gets the earlier user turn without the instruction. That is the template's rule; it is pinned by a test, noted in the prefix-stability module doc, and documented in docs/supported-models.md.

Test plan

  • New server::chat_request::reasoning_content_tests (5 tests: both-key forwarding, Jamba boundary prefixes turn 2 with echoed reasoning, content-only rewrite pinned, no double render, primed-think split and truncation pinned); server::chat_request filter 126 passed
  • cargo clippy --release --features cuda --lib --tests -- -D warnings, cargo fmt --check
  • Real server on GB10, fresh process, 3 turns echoing reasoning_content at max_tokens 2048: finish stop/stop/stop, content Paris./4/Louvre, cached_tokens 0/41/88 (before: prompt 66 ignored the echo, cached 0 on turns 2 and 3)
  • Same server, content-only echo: content non-empty, cached 0/0/0 (documented template behavior)

Closes #2089

inureyes added a commit that referenced this pull request Oct 2, 2026
@inureyes inureyes added status:done Completed type:bug Bug fixes, error corrections, or issue resolutions priority:low Low priority labels Oct 2, 2026
AI21 Jamba-Reasoning-3B appeared to keep every reply in reasoning_content and missed the prompt cache on every turn. Measured on the real checkpoint, the empty content is finish_reason=length only: at max_tokens 64 every turn stops inside the primed <think> block, while at 2048 the model closes </think> and content is filled. Close-marker recognition works.

The cache miss had a server-side cause. The raw-JSON render forwarded an echoed trace only as `reasoning`, but Qwen3-style and Jamba templates read `message.reasoning_content`. Jamba keeps its thinking instruction on an earlier user turn only when the next assistant message carries that field, so the turn re-rendered shorter than it was generated and the history-boundary snapshot never prefixed the next prompt. The trace now goes out under both spellings; templates that accept both read them as alternatives, so nothing renders twice.

A client that echoes only content still gets the rewritten user turn and misses the cache; that is the template's own rule and is pinned by a test and documented.

Real server, fresh process, 3 turns echoing reasoning_content: content non-empty each turn, cached_tokens 0/41/88 (was 0/0/0).

Closes #2089
@inureyes
inureyes force-pushed the fix/issue-2089-jamba-reasoning-content branch from ea79329 to dcaaaf9 Compare October 2, 2026 09:13
@inureyes
inureyes merged commit 9a0be04 into main Oct 2, 2026
140 of 145 checks passed
@inureyes
inureyes deleted the fix/issue-2089-jamba-reasoning-content branch October 2, 2026 09:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority:low Low priority status:done Completed type:bug Bug fixes, error corrections, or issue resolutions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: Jamba-Reasoning-3B chat output stays in reasoning_content

1 participant