diff --git a/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md new file mode 100644 index 000000000..034e1f61c --- /dev/null +++ b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md @@ -0,0 +1,43 @@ +# Technical Report: PR #2094 - Forward echoed reasoning under `reasoning_content` too + +**Date**: 2026-10-02 + +**Status**: Implemented and validated on GB10, awaiting merge. + +**Language**: Rust (chat request rendering, tests), Markdown + +**Risk**: Low (the raw-JSON message passed to chat templates gains a `reasoning_content` key carrying the same value as the existing `reasoning` key, only when a client echoed a trace) + +## Summary + +Issue #2089 reported that `AI21-Jamba-Reasoning-3B-4bit` (stored as `jamba-v0.1-4bit`) returned empty `content` and missed the prompt cache on every turn. The investigation split this into two findings. The empty `content` is `finish_reason=length`: the template always primes ``, and a 64-token budget ends inside the reasoning block. With a 2048-token budget the model closes `` and `content` is filled, so marker recognition is correct and no budget change was made. The cache miss had a server-side cause: an echoed trace reached chat templates only as `reasoning`, while this template reads `reasoning_content`. + +## 1. Measurements (before the fix) + +| max_tokens | turn 1 | turn 2 | turn 3 | +|---|---|---|---| +| 64 | length, content empty | length, content empty | length, content empty | +| 2048 | stop, `Paris.` | length (looped on a population question) | stop, `The Louvre.` | + +Turn 2 and 3 reported `cached_tokens=0` even when turn 1 ended with non-empty content. Echoing `reasoning_content` did not change the prompt length (66 tokens in both cases), which showed the server dropped the echo before rendering. + +## 2. Root cause + +The Jamba template adds a thinking instruction to the last user turn. On an earlier user turn it keeps the instruction only when the next assistant message starts with `` or defines `reasoning_content`. `build_raw_json_messages_with_thinking` folded both wire spellings into one field and emitted it as `reasoning` (the Gemma 4 spelling). The template therefore saw no reasoning, rendered turn 1's user message without the instruction on turn 2, and the turn-1 history-boundary snapshot (which contains the instruction) no longer prefixed turn 2's prompt. + +## 3. Change + +The forwarded trace is now written under both `reasoning` and `reasoning_content`, with the same gating as before (dropped for stripped turns and when the content already carries an inline `` block). Templates on hand that accept both (Gemma 4, Laguna) read them as alternatives, so the trace is never rendered twice; a test pins this. + +A client that echoes only `content` still gets the earlier user turn without the instruction. That is the template's own rule. It is pinned by a test, recorded in the prefix-stability notes of `chat_request.rs`, and documented in `docs/supported-models.md` together with the `max_tokens` guidance. + +## 4. Validation + +- `server::chat_request` tests: 126 passed, including 5 new tests in `chat_request_reasoning_content_tests.rs`. +- Clippy with `-D warnings` and rustfmt clean. +- Real server on GB10, fresh process, 3 turns at max_tokens 2048 echoing `reasoning_content`: finish stop/stop/stop, content `Paris.` / `4` / `Louvre`, prompt 46/93/135, cached_tokens 0/41/88. +- Same server, content-only echo: content non-empty, cached_tokens 0/0/0, as documented. + +## 5. Not covered + +Recovering cache reuse for content-only clients would need a snapshot at the point where the template stops rewriting history (before the last user turn) plus a warm-up that can start from it. That is a general change to the boundary snapshot design and is left for a separate issue. diff --git a/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md new file mode 100644 index 000000000..4b1860bc7 --- /dev/null +++ b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md @@ -0,0 +1,43 @@ +# 기술 보고서: PR #2094 - 되돌려 보낸 reasoning을 `reasoning_content`로도 전달 + +**날짜**: 2026-10-02 + +**상태**: GB10에서 구현 및 검증 완료, 머지 대기 중. + +**언어**: Rust (chat request 렌더링, 테스트), Markdown + +**위험도**: 낮음 (클라이언트가 reasoning을 되돌려 보낸 경우에만, chat template에 전달되는 raw JSON 메시지에 기존 `reasoning` 키와 같은 값을 가진 `reasoning_content` 키가 추가됩니다) + +## 요약 + +이슈 #2089는 `AI21-Jamba-Reasoning-3B-4bit`(`jamba-v0.1-4bit`로 저장됨)가 빈 `content`를 반환하고 매 turn prompt cache를 놓친다고 보고했습니다. 조사 결과 두 가지로 나뉩니다. 빈 `content`는 `finish_reason=length`입니다. template이 항상 ``를 미리 열어 두므로 64 토큰 예산은 reasoning 블록 안에서 끝납니다. 2048 토큰에서는 모델이 ``를 닫고 `content`가 채워지므로 marker 인식은 정상이며 예산 처리는 바꾸지 않았습니다. cache miss는 서버 쪽 원인이 있었습니다. 되돌려 보낸 reasoning이 chat template에 `reasoning`으로만 전달되었는데, 이 template은 `reasoning_content`를 읽습니다. + +## 1. 측정 (수정 전) + +| max_tokens | turn 1 | turn 2 | turn 3 | +|---|---|---|---| +| 64 | length, content 비어 있음 | length, content 비어 있음 | length, content 비어 있음 | +| 2048 | stop, `Paris.` | length (인구 질문에서 반복) | stop, `The Louvre.` | + +turn 1이 비어 있지 않은 content로 끝나도 turn 2와 3은 `cached_tokens=0`이었습니다. `reasoning_content`를 되돌려 보내도 prompt 길이가 같았고(두 경우 모두 66 토큰), 이는 서버가 렌더링 전에 그 값을 버렸다는 뜻입니다. + +## 2. 근본 원인 + +Jamba template은 마지막 user turn에 thinking 지시문을 붙입니다. 이전 user turn에는 다음 assistant 메시지가 ``로 시작하거나 `reasoning_content`를 정의할 때만 지시문을 유지합니다. `build_raw_json_messages_with_thinking`은 두 wire 표기를 하나의 필드로 합친 뒤 `reasoning`(Gemma 4 표기)으로만 내보냈습니다. 그래서 template은 reasoning을 보지 못했고, turn 2에서 turn 1의 user 메시지를 지시문 없이 렌더링했으며, 지시문을 포함한 turn 1의 history-boundary snapshot은 더 이상 turn 2 prompt의 prefix가 아니었습니다. + +## 3. 변경 내용 + +전달되는 reasoning을 이제 `reasoning`과 `reasoning_content` 두 키에 모두 씁니다. gating은 이전과 같습니다(strip 대상 turn이거나 content에 이미 inline `` 블록이 있으면 생략). 두 키를 모두 받는 기존 template(Gemma 4, Laguna)은 둘을 대안으로 읽으므로 reasoning이 두 번 렌더링되지 않으며, 테스트로 고정했습니다. + +`content`만 되돌려 보내는 클라이언트는 여전히 지시문이 빠진 이전 user turn을 받습니다. 이는 template 자체의 규칙입니다. 테스트로 고정했고, `chat_request.rs`의 prefix 안정성 설명과 `docs/supported-models.md`에 `max_tokens` 안내와 함께 기록했습니다. + +## 4. 검증 + +- `server::chat_request` 테스트 126개 통과, `chat_request_reasoning_content_tests.rs`의 새 테스트 5개 포함. +- `-D warnings` clippy와 rustfmt 통과. +- GB10 실서버, 새 프로세스, max_tokens 2048에서 `reasoning_content`를 되돌려 보내는 3 turn: finish stop/stop/stop, content `Paris.` / `4` / `Louvre`, prompt 46/93/135, cached_tokens 0/41/88. +- 같은 서버에서 content만 되돌려 보낸 경우: content는 비어 있지 않고 cached_tokens 0/0/0, 문서화된 동작과 일치. + +## 5. 다루지 않은 부분 + +content만 보내는 클라이언트의 cache 재사용을 되살리려면 template이 history를 다시 쓰기 시작하는 지점(마지막 user turn 앞)의 snapshot과 그 지점에서 시작할 수 있는 warm-up이 필요합니다. 이는 boundary snapshot 설계 전반의 변경이므로 별도 이슈로 남깁니다. diff --git a/docs/supported-models.md b/docs/supported-models.md index d165675d2..2371ac01b 100644 --- a/docs/supported-models.md +++ b/docs/supported-models.md @@ -524,6 +524,12 @@ answers. The scratchpad is no longer dropped: Chat Completions surfaces it as `reasoning_content` and, by default, an identical OpenRouter-compatible `reasoning` alias on both streaming deltas and non-streaming assistant messages. Both fields are omitted when the model produces no reasoning. Set `--reasoning-alias-field none` (or `MLXCEL_REASONING_ALIAS_FIELD=none`) to retain only `reasoning_content` when response bytes matter. This applies to every thinking family, including Qwen-style `` models. To turn thinking off rather than only suppress its alias, pass `chat_template_kwargs={"enable_thinking": false}` per request, or set the server default via `--chat-template-kwargs` or `LLAMA_ARG_CHAT_TEMPLATE_KWARGS`. A per-request value always wins over the server default. +#### AI21 Jamba-Reasoning + +`AI21-Jamba-Reasoning-3B` (model type `jamba`) always primes `` in its chat template, so every reply starts in `reasoning_content`. The answer reaches `content` once the model writes ``; a `max_tokens` budget that ends the decode first returns an empty `content` with `finish_reason` of `length`, as for the other primed-thinking families. Short factual answers measured 150 to 450 completion tokens at temperature 0 on this checkpoint, so set `max_tokens` to at least 1024. + +For multi-turn chat, echo each assistant turn's `reasoning_content` (or `reasoning`) back with its `content`. The template keeps its thinking instruction on an earlier user turn only when the reply that follows carries `reasoning_content`, and the server forwards an echoed trace under both field names, so the turn re-renders as it was generated and the follow-up reuses the prompt cache. A client that echoes only `content` gets the earlier user turn rendered without the instruction, which is the template's own rule, and that follow-up re-prefills from the start. + ### CLI reasoning display (`--show-reasoning`) `mlxcel generate` and `mlxcel run` decode with special tokens intact, so a diff --git a/src/server/chat_request.rs b/src/server/chat_request.rs index 6f12b0d11..8e63cbf04 100644 --- a/src/server/chat_request.rs +++ b/src/server/chat_request.rs @@ -47,6 +47,14 @@ //! bucket that naturally ages out. Documented, not silently masked. //! * `` block stripping across turns — invalidated by the //! `preserve_thinking=true` default. +//! * Templates that rewrite an earlier turn depending on the reply's +//! reasoning (AI21 Jamba-Reasoning keeps its thinking instruction on an +//! earlier user turn only when the following assistant message carries +//! `reasoning_content`, issue #2089). An echoed trace is forwarded under +//! both `reasoning` and `reasoning_content`, so such a turn re-renders +//! exactly as it was generated. A client that echoes only `content` +//! still gets the rewritten (shorter) turn; that is the template's own +//! rule and the follow-up misses the cache by design. //! * Tool-schema hashing: [`super::prompt_cache::key::tools_digest`] is //! order-preserving, so reordering tools invalidates the cache. This //! is intentional: HuggingFace templates iterate tools in order and @@ -1871,12 +1879,25 @@ fn build_raw_json_messages_with_thinking( // carries an inline `` block. Forwarding it on top of an // inline block would double-inject the same reasoning into // templates that render both channels. + // + // The value goes out under both spellings the wire accepts (issue + // #2089). The wire layer folds `reasoning` and `reasoning_content` + // into one field, but a template reads one of them: Gemma 4 reads + // `reasoning`, while Qwen3-style and AI21 Jamba-Reasoning templates + // read `message.reasoning_content`. Emitting only `reasoning` hid + // an echoed trace from the second group, and Jamba's template keys + // the thinking instruction on an earlier user turn to that field, + // so the turn re-rendered shorter than it had been generated and + // every follow-up missed the prompt cache. Templates that accept + // both read them as alternatives (`reasoning or + // reasoning_content`), so this never renders the trace twice. if !stripped && let Some(reasoning) = m.reasoning.as_ref() && !reasoning.is_empty() && !raw_content.contains("") { msg["reasoning"] = serde_json::Value::String(reasoning.clone()); + msg["reasoning_content"] = serde_json::Value::String(reasoning.clone()); } msg @@ -2113,3 +2134,7 @@ fn build_chat_messages_with_thinking( #[cfg(test)] #[path = "chat_request_tests.rs"] mod tests; + +#[cfg(test)] +#[path = "chat_request_reasoning_content_tests.rs"] +mod reasoning_content_tests; diff --git a/src/server/chat_request_reasoning_content_tests.rs b/src/server/chat_request_reasoning_content_tests.rs new file mode 100644 index 000000000..e3d52af02 --- /dev/null +++ b/src/server/chat_request_reasoning_content_tests.rs @@ -0,0 +1,177 @@ +// Copyright 2025-2026 Lablup Inc. +// +// Licensed under the Apache License, Version 2.0 (the "License"); +// you may not use this file except in compliance with the License. +// You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, software +// distributed under the License is distributed on an "AS IS" BASIS, +// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +// See the License for the specific language governing permissions and +// limitations under the License. + +//! Echoed `reasoning_content` and the AI21 Jamba-Reasoning template (#2089). +//! +//! Jamba-Reasoning-3B's template adds a thinking instruction to the last user +//! turn, and keeps it on an earlier user turn only when the assistant reply +//! that follows carries `reasoning_content` (or opens with ``). The +//! server used to forward an echoed trace under `reasoning` alone, so this +//! template never saw it: the earlier user turn re-rendered without the +//! instruction, the turn-1 history boundary stopped prefixing turn 2, and every +//! follow-up missed the prompt cache. + +use super::{build_raw_json_messages, prepare_chat_request_with_cache}; +use crate::server::chat_template::ChatTemplateProcessor; +use crate::server::types::ChatCompletionRequest; +use crate::tokenizer::ThinkingMarkers; + +/// The user/assistant branches of the shipped Jamba-Reasoning template +/// (`chat_template.jinja` of `AI21-Jamba-Reasoning-3B`), with the instruction +/// shortened to `THINK_PREFIX`. The prefix rule is copied verbatim. +fn jamba_reasoning_template() -> String { + r#"{%- set thinking_prefix = 'THINK_PREFIX\n' -%} +{%- for message in messages %} +{%- if message.role == "user" %} +{%- set prefix = '' %} +{%- if '' not in message.content %} +{%- if loop.last %}{%- set prefix = thinking_prefix %}{%- endif %} +{%- if not loop.last %} +{%- if loop.nextitem.role == 'assistant' and loop.nextitem.content.startswith('') or loop.nextitem.reasoning_content is defined and loop.nextitem.reasoning_content is not none %} +{%- set prefix = thinking_prefix %} +{%- endif %} +{%- endif %} +{%- endif %} +{{- '<|im_start|>user\n' + prefix + message.content + '<|im_end|>\n' }} +{%- elif message.role == "assistant" %} +{{- '<|im_start|>assistant\n' + message.content + '<|im_end|>\n' }} +{%- endif %} +{%- endfor %} +{%- if add_generation_prompt %}{{- '<|im_start|>assistant\n\n' }}{%- endif -%}"# + .to_string() +} + +fn request(messages: serde_json::Value) -> ChatCompletionRequest { + serde_json::from_value(serde_json::json!({ "model": "jamba", "messages": messages })) + .expect("valid chat request") +} + +async fn render( + processor: &ChatTemplateProcessor, + req: &ChatCompletionRequest, +) -> (String, Option) { + let prepared = prepare_chat_request_with_cache( + processor, + req, + None, + true, + true, + false, + &ThinkingMarkers::default(), + ) + .await + .expect("render succeeds"); + (prepared.prompt, prepared.history_prompt) +} + +#[test] +fn echoed_reasoning_content_is_forwarded_under_both_spellings() { + let req = request(serde_json::json!([ + {"role": "user", "content": "q1"}, + {"role": "assistant", "content": "Paris.", "reasoning_content": "trace"}, + {"role": "user", "content": "q2"}, + ])); + let raw = build_raw_json_messages(&req); + let assistant = &raw.as_array().expect("array")[1]; + assert_eq!(assistant["reasoning_content"], "trace"); + assert_eq!(assistant["reasoning"], "trace"); +} + +#[tokio::test] +async fn jamba_history_boundary_prefixes_the_next_turn_when_reasoning_is_echoed() { + let processor = ChatTemplateProcessor::with_template(jamba_reasoning_template()); + let turn1 = request(serde_json::json!([{"role": "user", "content": "q1"}])); + let (_, history1) = render(&processor, &turn1).await; + let history1 = history1.expect("prompt cache on: history boundary rendered"); + assert_eq!(history1, "<|im_start|>user\nTHINK_PREFIX\nq1<|im_end|>\n"); + + let turn2 = request(serde_json::json!([ + {"role": "user", "content": "q1"}, + {"role": "assistant", "content": "Paris.", "reasoning_content": "trace"}, + {"role": "user", "content": "q2"}, + ])); + let (prompt2, _) = render(&processor, &turn2).await; + assert!( + prompt2.starts_with(&history1), + "turn-1 boundary must prefix turn 2 when the client echoes reasoning_content;\n\ + boundary: {history1:?}\nturn 2: {prompt2:?}" + ); +} + +/// Pins the documented limit: a client that echoes only `content` gets an +/// earlier user turn without the instruction. That rewrite is the template's +/// own rule, not a server render choice, so the boundary cannot prefix turn 2 +/// and the follow-up misses the prompt cache by design. +#[tokio::test] +async fn jamba_content_only_echo_rewrites_the_earlier_user_turn() { + let processor = ChatTemplateProcessor::with_template(jamba_reasoning_template()); + let turn1 = request(serde_json::json!([{"role": "user", "content": "q1"}])); + let (_, history1) = render(&processor, &turn1).await; + let history1 = history1.expect("history boundary rendered"); + + let turn2 = request(serde_json::json!([ + {"role": "user", "content": "q1"}, + {"role": "assistant", "content": "Paris."}, + {"role": "user", "content": "q2"}, + ])); + let (prompt2, _) = render(&processor, &turn2).await; + assert!(prompt2.starts_with("<|im_start|>user\nq1<|im_end|>\n")); + assert!(!prompt2.starts_with(&history1)); +} + +/// Templates that accept both spellings read them as alternatives, so the +/// second key never renders the trace twice. +#[tokio::test] +async fn template_reading_either_spelling_renders_the_trace_once() { + let processor = ChatTemplateProcessor::with_template( + "{%- for m in messages %}{{ m.role }}:{{ m.get('reasoning') or m.get('reasoning_content') or '' }}|{{ m.content }}\n{%- endfor %}" + .to_string(), + ); + let req = request(serde_json::json!([ + {"role": "user", "content": "q1"}, + {"role": "assistant", "content": "a1", "reasoning_content": "TRACE"}, + {"role": "user", "content": "q2"}, + ])); + let (prompt, _) = render(&processor, &req).await; + assert_eq!(prompt.matches("TRACE").count(), 1, "{prompt:?}"); +} + +/// Pins the `finish_reason=length` branch the investigation measured: when +/// `max_tokens` ends the decode inside the primed `` block, the output is +/// all reasoning and `content` is empty, the same as every primed-thinking +/// family. When `` is reached, the answer lands in `content`. +#[test] +fn jamba_primed_think_splits_on_close_and_stays_reasoning_when_truncated() { + use crate::server::routes::chat::{extract_reasoning_content, is_prompt_primed_open_thinking}; + let markers = ThinkingMarkers { + think_start: Some("".to_string()), + think_end: Some("".to_string()), + think_start_tokens: Some(vec![541]), + think_end_tokens: Some(vec![542]), + ..ThinkingMarkers::default() + }; + let prompt = "<|im_start|>user\nTHINK_PREFIX\nq1<|im_end|>\n<|im_start|>assistant\n\n"; + assert!(is_prompt_primed_open_thinking(&markers, prompt)); + + let truncated = "The user asks for the capital. It is Paris. So we"; + assert_eq!( + extract_reasoning_content(truncated, true).as_deref(), + Some(truncated) + ); + + let closed = "The capital is Paris.\n\n\nParis."; + let reasoning = extract_reasoning_content(closed, true).expect("reasoning captured"); + assert!(reasoning.contains("The capital is Paris.")); + assert!(!reasoning.contains("")); +}