diff --git a/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md
new file mode 100644
index 000000000..034e1f61c
--- /dev/null
+++ b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md
@@ -0,0 +1,43 @@
+# Technical Report: PR #2094 - Forward echoed reasoning under `reasoning_content` too
+
+**Date**: 2026-10-02
+
+**Status**: Implemented and validated on GB10, awaiting merge.
+
+**Language**: Rust (chat request rendering, tests), Markdown
+
+**Risk**: Low (the raw-JSON message passed to chat templates gains a `reasoning_content` key carrying the same value as the existing `reasoning` key, only when a client echoed a trace)
+
+## Summary
+
+Issue #2089 reported that `AI21-Jamba-Reasoning-3B-4bit` (stored as `jamba-v0.1-4bit`) returned empty `content` and missed the prompt cache on every turn. The investigation split this into two findings. The empty `content` is `finish_reason=length`: the template always primes ``, and a 64-token budget ends inside the reasoning block. With a 2048-token budget the model closes `` and `content` is filled, so marker recognition is correct and no budget change was made. The cache miss had a server-side cause: an echoed trace reached chat templates only as `reasoning`, while this template reads `reasoning_content`.
+
+## 1. Measurements (before the fix)
+
+| max_tokens | turn 1 | turn 2 | turn 3 |
+|---|---|---|---|
+| 64 | length, content empty | length, content empty | length, content empty |
+| 2048 | stop, `Paris.` | length (looped on a population question) | stop, `The Louvre.` |
+
+Turn 2 and 3 reported `cached_tokens=0` even when turn 1 ended with non-empty content. Echoing `reasoning_content` did not change the prompt length (66 tokens in both cases), which showed the server dropped the echo before rendering.
+
+## 2. Root cause
+
+The Jamba template adds a thinking instruction to the last user turn. On an earlier user turn it keeps the instruction only when the next assistant message starts with `` or defines `reasoning_content`. `build_raw_json_messages_with_thinking` folded both wire spellings into one field and emitted it as `reasoning` (the Gemma 4 spelling). The template therefore saw no reasoning, rendered turn 1's user message without the instruction on turn 2, and the turn-1 history-boundary snapshot (which contains the instruction) no longer prefixed turn 2's prompt.
+
+## 3. Change
+
+The forwarded trace is now written under both `reasoning` and `reasoning_content`, with the same gating as before (dropped for stripped turns and when the content already carries an inline `` block). Templates on hand that accept both (Gemma 4, Laguna) read them as alternatives, so the trace is never rendered twice; a test pins this.
+
+A client that echoes only `content` still gets the earlier user turn without the instruction. That is the template's own rule. It is pinned by a test, recorded in the prefix-stability notes of `chat_request.rs`, and documented in `docs/supported-models.md` together with the `max_tokens` guidance.
+
+## 4. Validation
+
+- `server::chat_request` tests: 126 passed, including 5 new tests in `chat_request_reasoning_content_tests.rs`.
+- Clippy with `-D warnings` and rustfmt clean.
+- Real server on GB10, fresh process, 3 turns at max_tokens 2048 echoing `reasoning_content`: finish stop/stop/stop, content `Paris.` / `4` / `Louvre`, prompt 46/93/135, cached_tokens 0/41/88.
+- Same server, content-only echo: content non-empty, cached_tokens 0/0/0, as documented.
+
+## 5. Not covered
+
+Recovering cache reuse for content-only clients would need a snapshot at the point where the template stops rewriting history (before the last user turn) plus a warm-up that can start from it. That is a general change to the boundary snapshot design and is left for a separate issue.
diff --git a/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md
new file mode 100644
index 000000000..4b1860bc7
--- /dev/null
+++ b/TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md
@@ -0,0 +1,43 @@
+# 기술 보고서: PR #2094 - 되돌려 보낸 reasoning을 `reasoning_content`로도 전달
+
+**날짜**: 2026-10-02
+
+**상태**: GB10에서 구현 및 검증 완료, 머지 대기 중.
+
+**언어**: Rust (chat request 렌더링, 테스트), Markdown
+
+**위험도**: 낮음 (클라이언트가 reasoning을 되돌려 보낸 경우에만, chat template에 전달되는 raw JSON 메시지에 기존 `reasoning` 키와 같은 값을 가진 `reasoning_content` 키가 추가됩니다)
+
+## 요약
+
+이슈 #2089는 `AI21-Jamba-Reasoning-3B-4bit`(`jamba-v0.1-4bit`로 저장됨)가 빈 `content`를 반환하고 매 turn prompt cache를 놓친다고 보고했습니다. 조사 결과 두 가지로 나뉩니다. 빈 `content`는 `finish_reason=length`입니다. template이 항상 ``를 미리 열어 두므로 64 토큰 예산은 reasoning 블록 안에서 끝납니다. 2048 토큰에서는 모델이 ``를 닫고 `content`가 채워지므로 marker 인식은 정상이며 예산 처리는 바꾸지 않았습니다. cache miss는 서버 쪽 원인이 있었습니다. 되돌려 보낸 reasoning이 chat template에 `reasoning`으로만 전달되었는데, 이 template은 `reasoning_content`를 읽습니다.
+
+## 1. 측정 (수정 전)
+
+| max_tokens | turn 1 | turn 2 | turn 3 |
+|---|---|---|---|
+| 64 | length, content 비어 있음 | length, content 비어 있음 | length, content 비어 있음 |
+| 2048 | stop, `Paris.` | length (인구 질문에서 반복) | stop, `The Louvre.` |
+
+turn 1이 비어 있지 않은 content로 끝나도 turn 2와 3은 `cached_tokens=0`이었습니다. `reasoning_content`를 되돌려 보내도 prompt 길이가 같았고(두 경우 모두 66 토큰), 이는 서버가 렌더링 전에 그 값을 버렸다는 뜻입니다.
+
+## 2. 근본 원인
+
+Jamba template은 마지막 user turn에 thinking 지시문을 붙입니다. 이전 user turn에는 다음 assistant 메시지가 ``로 시작하거나 `reasoning_content`를 정의할 때만 지시문을 유지합니다. `build_raw_json_messages_with_thinking`은 두 wire 표기를 하나의 필드로 합친 뒤 `reasoning`(Gemma 4 표기)으로만 내보냈습니다. 그래서 template은 reasoning을 보지 못했고, turn 2에서 turn 1의 user 메시지를 지시문 없이 렌더링했으며, 지시문을 포함한 turn 1의 history-boundary snapshot은 더 이상 turn 2 prompt의 prefix가 아니었습니다.
+
+## 3. 변경 내용
+
+전달되는 reasoning을 이제 `reasoning`과 `reasoning_content` 두 키에 모두 씁니다. gating은 이전과 같습니다(strip 대상 turn이거나 content에 이미 inline `` 블록이 있으면 생략). 두 키를 모두 받는 기존 template(Gemma 4, Laguna)은 둘을 대안으로 읽으므로 reasoning이 두 번 렌더링되지 않으며, 테스트로 고정했습니다.
+
+`content`만 되돌려 보내는 클라이언트는 여전히 지시문이 빠진 이전 user turn을 받습니다. 이는 template 자체의 규칙입니다. 테스트로 고정했고, `chat_request.rs`의 prefix 안정성 설명과 `docs/supported-models.md`에 `max_tokens` 안내와 함께 기록했습니다.
+
+## 4. 검증
+
+- `server::chat_request` 테스트 126개 통과, `chat_request_reasoning_content_tests.rs`의 새 테스트 5개 포함.
+- `-D warnings` clippy와 rustfmt 통과.
+- GB10 실서버, 새 프로세스, max_tokens 2048에서 `reasoning_content`를 되돌려 보내는 3 turn: finish stop/stop/stop, content `Paris.` / `4` / `Louvre`, prompt 46/93/135, cached_tokens 0/41/88.
+- 같은 서버에서 content만 되돌려 보낸 경우: content는 비어 있지 않고 cached_tokens 0/0/0, 문서화된 동작과 일치.
+
+## 5. 다루지 않은 부분
+
+content만 보내는 클라이언트의 cache 재사용을 되살리려면 template이 history를 다시 쓰기 시작하는 지점(마지막 user turn 앞)의 snapshot과 그 지점에서 시작할 수 있는 warm-up이 필요합니다. 이는 boundary snapshot 설계 전반의 변경이므로 별도 이슈로 남깁니다.
diff --git a/docs/supported-models.md b/docs/supported-models.md
index d165675d2..2371ac01b 100644
--- a/docs/supported-models.md
+++ b/docs/supported-models.md
@@ -524,6 +524,12 @@ answers.
The scratchpad is no longer dropped: Chat Completions surfaces it as `reasoning_content` and, by default, an identical OpenRouter-compatible `reasoning` alias on both streaming deltas and non-streaming assistant messages. Both fields are omitted when the model produces no reasoning. Set `--reasoning-alias-field none` (or `MLXCEL_REASONING_ALIAS_FIELD=none`) to retain only `reasoning_content` when response bytes matter. This applies to every thinking family, including Qwen-style `` models. To turn thinking off rather than only suppress its alias, pass `chat_template_kwargs={"enable_thinking": false}` per request, or set the server default via `--chat-template-kwargs` or `LLAMA_ARG_CHAT_TEMPLATE_KWARGS`. A per-request value always wins over the server default.
+#### AI21 Jamba-Reasoning
+
+`AI21-Jamba-Reasoning-3B` (model type `jamba`) always primes `` in its chat template, so every reply starts in `reasoning_content`. The answer reaches `content` once the model writes ``; a `max_tokens` budget that ends the decode first returns an empty `content` with `finish_reason` of `length`, as for the other primed-thinking families. Short factual answers measured 150 to 450 completion tokens at temperature 0 on this checkpoint, so set `max_tokens` to at least 1024.
+
+For multi-turn chat, echo each assistant turn's `reasoning_content` (or `reasoning`) back with its `content`. The template keeps its thinking instruction on an earlier user turn only when the reply that follows carries `reasoning_content`, and the server forwards an echoed trace under both field names, so the turn re-renders as it was generated and the follow-up reuses the prompt cache. A client that echoes only `content` gets the earlier user turn rendered without the instruction, which is the template's own rule, and that follow-up re-prefills from the start.
+
### CLI reasoning display (`--show-reasoning`)
`mlxcel generate` and `mlxcel run` decode with special tokens intact, so a
diff --git a/src/server/chat_request.rs b/src/server/chat_request.rs
index 6f12b0d11..8e63cbf04 100644
--- a/src/server/chat_request.rs
+++ b/src/server/chat_request.rs
@@ -47,6 +47,14 @@
//! bucket that naturally ages out. Documented, not silently masked.
//! * `` block stripping across turns — invalidated by the
//! `preserve_thinking=true` default.
+//! * Templates that rewrite an earlier turn depending on the reply's
+//! reasoning (AI21 Jamba-Reasoning keeps its thinking instruction on an
+//! earlier user turn only when the following assistant message carries
+//! `reasoning_content`, issue #2089). An echoed trace is forwarded under
+//! both `reasoning` and `reasoning_content`, so such a turn re-renders
+//! exactly as it was generated. A client that echoes only `content`
+//! still gets the rewritten (shorter) turn; that is the template's own
+//! rule and the follow-up misses the cache by design.
//! * Tool-schema hashing: [`super::prompt_cache::key::tools_digest`] is
//! order-preserving, so reordering tools invalidates the cache. This
//! is intentional: HuggingFace templates iterate tools in order and
@@ -1871,12 +1879,25 @@ fn build_raw_json_messages_with_thinking(
// carries an inline `` block. Forwarding it on top of an
// inline block would double-inject the same reasoning into
// templates that render both channels.
+ //
+ // The value goes out under both spellings the wire accepts (issue
+ // #2089). The wire layer folds `reasoning` and `reasoning_content`
+ // into one field, but a template reads one of them: Gemma 4 reads
+ // `reasoning`, while Qwen3-style and AI21 Jamba-Reasoning templates
+ // read `message.reasoning_content`. Emitting only `reasoning` hid
+ // an echoed trace from the second group, and Jamba's template keys
+ // the thinking instruction on an earlier user turn to that field,
+ // so the turn re-rendered shorter than it had been generated and
+ // every follow-up missed the prompt cache. Templates that accept
+ // both read them as alternatives (`reasoning or
+ // reasoning_content`), so this never renders the trace twice.
if !stripped
&& let Some(reasoning) = m.reasoning.as_ref()
&& !reasoning.is_empty()
&& !raw_content.contains("")
{
msg["reasoning"] = serde_json::Value::String(reasoning.clone());
+ msg["reasoning_content"] = serde_json::Value::String(reasoning.clone());
}
msg
@@ -2113,3 +2134,7 @@ fn build_chat_messages_with_thinking(
#[cfg(test)]
#[path = "chat_request_tests.rs"]
mod tests;
+
+#[cfg(test)]
+#[path = "chat_request_reasoning_content_tests.rs"]
+mod reasoning_content_tests;
diff --git a/src/server/chat_request_reasoning_content_tests.rs b/src/server/chat_request_reasoning_content_tests.rs
new file mode 100644
index 000000000..e3d52af02
--- /dev/null
+++ b/src/server/chat_request_reasoning_content_tests.rs
@@ -0,0 +1,177 @@
+// Copyright 2025-2026 Lablup Inc.
+//
+// Licensed under the Apache License, Version 2.0 (the "License");
+// you may not use this file except in compliance with the License.
+// You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing, software
+// distributed under the License is distributed on an "AS IS" BASIS,
+// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+// See the License for the specific language governing permissions and
+// limitations under the License.
+
+//! Echoed `reasoning_content` and the AI21 Jamba-Reasoning template (#2089).
+//!
+//! Jamba-Reasoning-3B's template adds a thinking instruction to the last user
+//! turn, and keeps it on an earlier user turn only when the assistant reply
+//! that follows carries `reasoning_content` (or opens with ``). The
+//! server used to forward an echoed trace under `reasoning` alone, so this
+//! template never saw it: the earlier user turn re-rendered without the
+//! instruction, the turn-1 history boundary stopped prefixing turn 2, and every
+//! follow-up missed the prompt cache.
+
+use super::{build_raw_json_messages, prepare_chat_request_with_cache};
+use crate::server::chat_template::ChatTemplateProcessor;
+use crate::server::types::ChatCompletionRequest;
+use crate::tokenizer::ThinkingMarkers;
+
+/// The user/assistant branches of the shipped Jamba-Reasoning template
+/// (`chat_template.jinja` of `AI21-Jamba-Reasoning-3B`), with the instruction
+/// shortened to `THINK_PREFIX`. The prefix rule is copied verbatim.
+fn jamba_reasoning_template() -> String {
+ r#"{%- set thinking_prefix = 'THINK_PREFIX\n' -%}
+{%- for message in messages %}
+{%- if message.role == "user" %}
+{%- set prefix = '' %}
+{%- if '' not in message.content %}
+{%- if loop.last %}{%- set prefix = thinking_prefix %}{%- endif %}
+{%- if not loop.last %}
+{%- if loop.nextitem.role == 'assistant' and loop.nextitem.content.startswith('') or loop.nextitem.reasoning_content is defined and loop.nextitem.reasoning_content is not none %}
+{%- set prefix = thinking_prefix %}
+{%- endif %}
+{%- endif %}
+{%- endif %}
+{{- '<|im_start|>user\n' + prefix + message.content + '<|im_end|>\n' }}
+{%- elif message.role == "assistant" %}
+{{- '<|im_start|>assistant\n' + message.content + '<|im_end|>\n' }}
+{%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}{{- '<|im_start|>assistant\n\n' }}{%- endif -%}"#
+ .to_string()
+}
+
+fn request(messages: serde_json::Value) -> ChatCompletionRequest {
+ serde_json::from_value(serde_json::json!({ "model": "jamba", "messages": messages }))
+ .expect("valid chat request")
+}
+
+async fn render(
+ processor: &ChatTemplateProcessor,
+ req: &ChatCompletionRequest,
+) -> (String, Option) {
+ let prepared = prepare_chat_request_with_cache(
+ processor,
+ req,
+ None,
+ true,
+ true,
+ false,
+ &ThinkingMarkers::default(),
+ )
+ .await
+ .expect("render succeeds");
+ (prepared.prompt, prepared.history_prompt)
+}
+
+#[test]
+fn echoed_reasoning_content_is_forwarded_under_both_spellings() {
+ let req = request(serde_json::json!([
+ {"role": "user", "content": "q1"},
+ {"role": "assistant", "content": "Paris.", "reasoning_content": "trace"},
+ {"role": "user", "content": "q2"},
+ ]));
+ let raw = build_raw_json_messages(&req);
+ let assistant = &raw.as_array().expect("array")[1];
+ assert_eq!(assistant["reasoning_content"], "trace");
+ assert_eq!(assistant["reasoning"], "trace");
+}
+
+#[tokio::test]
+async fn jamba_history_boundary_prefixes_the_next_turn_when_reasoning_is_echoed() {
+ let processor = ChatTemplateProcessor::with_template(jamba_reasoning_template());
+ let turn1 = request(serde_json::json!([{"role": "user", "content": "q1"}]));
+ let (_, history1) = render(&processor, &turn1).await;
+ let history1 = history1.expect("prompt cache on: history boundary rendered");
+ assert_eq!(history1, "<|im_start|>user\nTHINK_PREFIX\nq1<|im_end|>\n");
+
+ let turn2 = request(serde_json::json!([
+ {"role": "user", "content": "q1"},
+ {"role": "assistant", "content": "Paris.", "reasoning_content": "trace"},
+ {"role": "user", "content": "q2"},
+ ]));
+ let (prompt2, _) = render(&processor, &turn2).await;
+ assert!(
+ prompt2.starts_with(&history1),
+ "turn-1 boundary must prefix turn 2 when the client echoes reasoning_content;\n\
+ boundary: {history1:?}\nturn 2: {prompt2:?}"
+ );
+}
+
+/// Pins the documented limit: a client that echoes only `content` gets an
+/// earlier user turn without the instruction. That rewrite is the template's
+/// own rule, not a server render choice, so the boundary cannot prefix turn 2
+/// and the follow-up misses the prompt cache by design.
+#[tokio::test]
+async fn jamba_content_only_echo_rewrites_the_earlier_user_turn() {
+ let processor = ChatTemplateProcessor::with_template(jamba_reasoning_template());
+ let turn1 = request(serde_json::json!([{"role": "user", "content": "q1"}]));
+ let (_, history1) = render(&processor, &turn1).await;
+ let history1 = history1.expect("history boundary rendered");
+
+ let turn2 = request(serde_json::json!([
+ {"role": "user", "content": "q1"},
+ {"role": "assistant", "content": "Paris."},
+ {"role": "user", "content": "q2"},
+ ]));
+ let (prompt2, _) = render(&processor, &turn2).await;
+ assert!(prompt2.starts_with("<|im_start|>user\nq1<|im_end|>\n"));
+ assert!(!prompt2.starts_with(&history1));
+}
+
+/// Templates that accept both spellings read them as alternatives, so the
+/// second key never renders the trace twice.
+#[tokio::test]
+async fn template_reading_either_spelling_renders_the_trace_once() {
+ let processor = ChatTemplateProcessor::with_template(
+ "{%- for m in messages %}{{ m.role }}:{{ m.get('reasoning') or m.get('reasoning_content') or '' }}|{{ m.content }}\n{%- endfor %}"
+ .to_string(),
+ );
+ let req = request(serde_json::json!([
+ {"role": "user", "content": "q1"},
+ {"role": "assistant", "content": "a1", "reasoning_content": "TRACE"},
+ {"role": "user", "content": "q2"},
+ ]));
+ let (prompt, _) = render(&processor, &req).await;
+ assert_eq!(prompt.matches("TRACE").count(), 1, "{prompt:?}");
+}
+
+/// Pins the `finish_reason=length` branch the investigation measured: when
+/// `max_tokens` ends the decode inside the primed `` block, the output is
+/// all reasoning and `content` is empty, the same as every primed-thinking
+/// family. When `` is reached, the answer lands in `content`.
+#[test]
+fn jamba_primed_think_splits_on_close_and_stays_reasoning_when_truncated() {
+ use crate::server::routes::chat::{extract_reasoning_content, is_prompt_primed_open_thinking};
+ let markers = ThinkingMarkers {
+ think_start: Some("".to_string()),
+ think_end: Some("".to_string()),
+ think_start_tokens: Some(vec![541]),
+ think_end_tokens: Some(vec![542]),
+ ..ThinkingMarkers::default()
+ };
+ let prompt = "<|im_start|>user\nTHINK_PREFIX\nq1<|im_end|>\n<|im_start|>assistant\n\n";
+ assert!(is_prompt_primed_open_thinking(&markers, prompt));
+
+ let truncated = "The user asks for the capital. It is Paris. So we";
+ assert_eq!(
+ extract_reasoning_content(truncated, true).as_deref(),
+ Some(truncated)
+ );
+
+ let closed = "The capital is Paris.\n\n\nParis.";
+ let reasoning = extract_reasoning_content(closed, true).expect("reasoning captured");
+ assert!(reasoning.contains("The capital is Paris."));
+ assert!(!reasoning.contains(""));
+}