Problem
On GB10, the checkpoint at /home/inureyes/models/mlx/jamba-v0.1-4bit served through /v1/chat/completions logs WARN mlxcel::thinking: generation produced tokens but content is empty and all output routed to reasoning_content (src/server/routes/chat.rs:2333-2368), returns empty content, and reports cached=0 every turn (history-boundary snapshot stored at 3314 tokens, next turn diverges at 3277; unchanged after #2082).
That directory name is misleading. Its README says it is mlx-community/AI21-Jamba-Reasoning-3B-4bit (hidden 2560, 28 layers). The tokenizer declares <think>/</think> (ids 541/542), and chat_template.jinja:98-100 always appends <|im_start|>assistant\n<think>\n (no enable_thinking gate). Routing the output to the reasoning channel is therefore expected. The bug is that </think> is never seen or never reached.
Investigation (do first)
Read finish_reason and completion_tokens from the warn line on a default-launch server.
finish_reason=length: the model never closed </think> within max_tokens. Fix the budget handling so the client gets usable content, matching how other primed-thinking families behave.
finish_reason=stop: </think> was emitted and not recognized. Check how LlamaTokenizerFast decodes id 542, infer_thinking_markers (src/tokenizer/mod.rs:495), is_prompt_primed_open_thinking (chat.rs:2220), prompt_primed_open_thinking (src/reasoning_stream.rs:406), and extract_reasoning_content (chat.rs:2476).
Cache miss
Before blaming rendering, rule out an entry larger than the whole prompt-cache store. Then check rendering: chat_template.jinja:46-62 drops reasoning from earlier assistant turns and keeps <think> only when </think> is in content or reasoning_content is sent, so an empty re-sent turn renders shorter than what was generated. In scope only if it persists once content is non-empty. Otherwise file it separately as a general reasoning-model issue.
Acceptance Criteria
Implement in one PR with one verification pass. Run the server under gpu-lock run, building without the lock.
Problem
On GB10, the checkpoint at
/home/inureyes/models/mlx/jamba-v0.1-4bitserved through/v1/chat/completionslogsWARN mlxcel::thinking: generation produced tokens but content is empty and all output routed to reasoning_content(src/server/routes/chat.rs:2333-2368), returns emptycontent, and reportscached=0every turn (history-boundary snapshot stored at 3314 tokens, next turn diverges at 3277; unchanged after #2082).That directory name is misleading. Its README says it is
mlx-community/AI21-Jamba-Reasoning-3B-4bit(hidden 2560, 28 layers). The tokenizer declares<think>/</think>(ids 541/542), andchat_template.jinja:98-100always appends<|im_start|>assistant\n<think>\n(noenable_thinkinggate). Routing the output to the reasoning channel is therefore expected. The bug is that</think>is never seen or never reached.Investigation (do first)
Read
finish_reasonandcompletion_tokensfrom the warn line on a default-launch server.finish_reason=length: the model never closed</think>withinmax_tokens. Fix the budget handling so the client gets usablecontent, matching how other primed-thinking families behave.finish_reason=stop:</think>was emitted and not recognized. Check howLlamaTokenizerFastdecodes id 542,infer_thinking_markers(src/tokenizer/mod.rs:495),is_prompt_primed_open_thinking(chat.rs:2220),prompt_primed_open_thinking(src/reasoning_stream.rs:406), andextract_reasoning_content(chat.rs:2476).Cache miss
Before blaming rendering, rule out an entry larger than the whole prompt-cache store. Then check rendering:
chat_template.jinja:46-62drops reasoning from earlier assistant turns and keeps<think>only when</think>is incontentorreasoning_contentis sent, so an empty re-sent turn renders shorter than what was generated. In scope only if it persists oncecontentis non-empty. Otherwise file it separately as a general reasoning-model issue.Acceptance Criteria
lengthcovers the budget behavior;stopcovers close-marker recognition with this tokenizer's markers.contenton every turn, and turns 2 and 3 reportcached_tokens > 0.Implement in one PR with one verification pass. Run the server under
gpu-lock run, building without the lock.contentwasfinish_reason=lengthonly (pinned by a test, documented); the cache miss with non-empty content came from echoed reasoning being forwarded to templates only asreasoning, notreasoning_content. Fixed; real-server 3 turns echoingreasoning_contentreturn non-empty content with cached_tokens 0/41/88. Content-only echo still misses by the template's own rule (documented).