Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 43 additions & 0 deletions TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.en.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
# Technical Report: PR #2094 - Forward echoed reasoning under `reasoning_content` too

**Date**: 2026-10-02

**Status**: Implemented and validated on GB10, awaiting merge.

**Language**: Rust (chat request rendering, tests), Markdown

**Risk**: Low (the raw-JSON message passed to chat templates gains a `reasoning_content` key carrying the same value as the existing `reasoning` key, only when a client echoed a trace)

## Summary

Issue #2089 reported that `AI21-Jamba-Reasoning-3B-4bit` (stored as `jamba-v0.1-4bit`) returned empty `content` and missed the prompt cache on every turn. The investigation split this into two findings. The empty `content` is `finish_reason=length`: the template always primes `<think>`, and a 64-token budget ends inside the reasoning block. With a 2048-token budget the model closes `</think>` and `content` is filled, so marker recognition is correct and no budget change was made. The cache miss had a server-side cause: an echoed trace reached chat templates only as `reasoning`, while this template reads `reasoning_content`.

## 1. Measurements (before the fix)

| max_tokens | turn 1 | turn 2 | turn 3 |
|---|---|---|---|
| 64 | length, content empty | length, content empty | length, content empty |
| 2048 | stop, `Paris.` | length (looped on a population question) | stop, `The Louvre.` |

Turn 2 and 3 reported `cached_tokens=0` even when turn 1 ended with non-empty content. Echoing `reasoning_content` did not change the prompt length (66 tokens in both cases), which showed the server dropped the echo before rendering.

## 2. Root cause

The Jamba template adds a thinking instruction to the last user turn. On an earlier user turn it keeps the instruction only when the next assistant message starts with `<think>` or defines `reasoning_content`. `build_raw_json_messages_with_thinking` folded both wire spellings into one field and emitted it as `reasoning` (the Gemma 4 spelling). The template therefore saw no reasoning, rendered turn 1's user message without the instruction on turn 2, and the turn-1 history-boundary snapshot (which contains the instruction) no longer prefixed turn 2's prompt.

## 3. Change

The forwarded trace is now written under both `reasoning` and `reasoning_content`, with the same gating as before (dropped for stripped turns and when the content already carries an inline `<think>` block). Templates on hand that accept both (Gemma 4, Laguna) read them as alternatives, so the trace is never rendered twice; a test pins this.

A client that echoes only `content` still gets the earlier user turn without the instruction. That is the template's own rule. It is pinned by a test, recorded in the prefix-stability notes of `chat_request.rs`, and documented in `docs/supported-models.md` together with the `max_tokens` guidance.

## 4. Validation

- `server::chat_request` tests: 126 passed, including 5 new tests in `chat_request_reasoning_content_tests.rs`.
- Clippy with `-D warnings` and rustfmt clean.
- Real server on GB10, fresh process, 3 turns at max_tokens 2048 echoing `reasoning_content`: finish stop/stop/stop, content `Paris.` / `4` / `Louvre`, prompt 46/93/135, cached_tokens 0/41/88.
- Same server, content-only echo: content non-empty, cached_tokens 0/0/0, as documented.

## 5. Not covered

Recovering cache reuse for content-only clients would need a snapshot at the point where the template stops rewriting history (before the last user turn) plus a warm-up that can start from it. That is a general change to the boundary snapshot design and is left for a separate issue.
43 changes: 43 additions & 0 deletions TECHNICAL_REPORTS/2094-jamba-reasoning-content-echo-20261002.ko.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
# 기술 보고서: PR #2094 - 되돌려 보낸 reasoning을 `reasoning_content`로도 전달

**날짜**: 2026-10-02

**상태**: GB10에서 구현 및 검증 완료, 머지 대기 중.

**언어**: Rust (chat request 렌더링, 테스트), Markdown

**위험도**: 낮음 (클라이언트가 reasoning을 되돌려 보낸 경우에만, chat template에 전달되는 raw JSON 메시지에 기존 `reasoning` 키와 같은 값을 가진 `reasoning_content` 키가 추가됩니다)

## 요약

이슈 #2089는 `AI21-Jamba-Reasoning-3B-4bit`(`jamba-v0.1-4bit`로 저장됨)가 빈 `content`를 반환하고 매 turn prompt cache를 놓친다고 보고했습니다. 조사 결과 두 가지로 나뉩니다. 빈 `content`는 `finish_reason=length`입니다. template이 항상 `<think>`를 미리 열어 두므로 64 토큰 예산은 reasoning 블록 안에서 끝납니다. 2048 토큰에서는 모델이 `</think>`를 닫고 `content`가 채워지므로 marker 인식은 정상이며 예산 처리는 바꾸지 않았습니다. cache miss는 서버 쪽 원인이 있었습니다. 되돌려 보낸 reasoning이 chat template에 `reasoning`으로만 전달되었는데, 이 template은 `reasoning_content`를 읽습니다.

## 1. 측정 (수정 전)

| max_tokens | turn 1 | turn 2 | turn 3 |
|---|---|---|---|
| 64 | length, content 비어 있음 | length, content 비어 있음 | length, content 비어 있음 |
| 2048 | stop, `Paris.` | length (인구 질문에서 반복) | stop, `The Louvre.` |

turn 1이 비어 있지 않은 content로 끝나도 turn 2와 3은 `cached_tokens=0`이었습니다. `reasoning_content`를 되돌려 보내도 prompt 길이가 같았고(두 경우 모두 66 토큰), 이는 서버가 렌더링 전에 그 값을 버렸다는 뜻입니다.

## 2. 근본 원인

Jamba template은 마지막 user turn에 thinking 지시문을 붙입니다. 이전 user turn에는 다음 assistant 메시지가 `<think>`로 시작하거나 `reasoning_content`를 정의할 때만 지시문을 유지합니다. `build_raw_json_messages_with_thinking`은 두 wire 표기를 하나의 필드로 합친 뒤 `reasoning`(Gemma 4 표기)으로만 내보냈습니다. 그래서 template은 reasoning을 보지 못했고, turn 2에서 turn 1의 user 메시지를 지시문 없이 렌더링했으며, 지시문을 포함한 turn 1의 history-boundary snapshot은 더 이상 turn 2 prompt의 prefix가 아니었습니다.

## 3. 변경 내용

전달되는 reasoning을 이제 `reasoning`과 `reasoning_content` 두 키에 모두 씁니다. gating은 이전과 같습니다(strip 대상 turn이거나 content에 이미 inline `<think>` 블록이 있으면 생략). 두 키를 모두 받는 기존 template(Gemma 4, Laguna)은 둘을 대안으로 읽으므로 reasoning이 두 번 렌더링되지 않으며, 테스트로 고정했습니다.

`content`만 되돌려 보내는 클라이언트는 여전히 지시문이 빠진 이전 user turn을 받습니다. 이는 template 자체의 규칙입니다. 테스트로 고정했고, `chat_request.rs`의 prefix 안정성 설명과 `docs/supported-models.md`에 `max_tokens` 안내와 함께 기록했습니다.

## 4. 검증

- `server::chat_request` 테스트 126개 통과, `chat_request_reasoning_content_tests.rs`의 새 테스트 5개 포함.
- `-D warnings` clippy와 rustfmt 통과.
- GB10 실서버, 새 프로세스, max_tokens 2048에서 `reasoning_content`를 되돌려 보내는 3 turn: finish stop/stop/stop, content `Paris.` / `4` / `Louvre`, prompt 46/93/135, cached_tokens 0/41/88.
- 같은 서버에서 content만 되돌려 보낸 경우: content는 비어 있지 않고 cached_tokens 0/0/0, 문서화된 동작과 일치.

## 5. 다루지 않은 부분

content만 보내는 클라이언트의 cache 재사용을 되살리려면 template이 history를 다시 쓰기 시작하는 지점(마지막 user turn 앞)의 snapshot과 그 지점에서 시작할 수 있는 warm-up이 필요합니다. 이는 boundary snapshot 설계 전반의 변경이므로 별도 이슈로 남깁니다.
6 changes: 6 additions & 0 deletions docs/supported-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -524,6 +524,12 @@ answers.

The scratchpad is no longer dropped: Chat Completions surfaces it as `reasoning_content` and, by default, an identical OpenRouter-compatible `reasoning` alias on both streaming deltas and non-streaming assistant messages. Both fields are omitted when the model produces no reasoning. Set `--reasoning-alias-field none` (or `MLXCEL_REASONING_ALIAS_FIELD=none`) to retain only `reasoning_content` when response bytes matter. This applies to every thinking family, including Qwen-style `<think>` models. To turn thinking off rather than only suppress its alias, pass `chat_template_kwargs={"enable_thinking": false}` per request, or set the server default via `--chat-template-kwargs` or `LLAMA_ARG_CHAT_TEMPLATE_KWARGS`. A per-request value always wins over the server default.

#### AI21 Jamba-Reasoning

`AI21-Jamba-Reasoning-3B` (model type `jamba`) always primes `<think>` in its chat template, so every reply starts in `reasoning_content`. The answer reaches `content` once the model writes `</think>`; a `max_tokens` budget that ends the decode first returns an empty `content` with `finish_reason` of `length`, as for the other primed-thinking families. Short factual answers measured 150 to 450 completion tokens at temperature 0 on this checkpoint, so set `max_tokens` to at least 1024.

For multi-turn chat, echo each assistant turn's `reasoning_content` (or `reasoning`) back with its `content`. The template keeps its thinking instruction on an earlier user turn only when the reply that follows carries `reasoning_content`, and the server forwards an echoed trace under both field names, so the turn re-renders as it was generated and the follow-up reuses the prompt cache. A client that echoes only `content` gets the earlier user turn rendered without the instruction, which is the template's own rule, and that follow-up re-prefills from the start.

### CLI reasoning display (`--show-reasoning`)

`mlxcel generate` and `mlxcel run` decode with special tokens intact, so a
Expand Down
25 changes: 25 additions & 0 deletions src/server/chat_request.rs
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,14 @@
//! bucket that naturally ages out. Documented, not silently masked.
//! * `<think>` block stripping across turns — invalidated by the
//! `preserve_thinking=true` default.
//! * Templates that rewrite an earlier turn depending on the reply's
//! reasoning (AI21 Jamba-Reasoning keeps its thinking instruction on an
//! earlier user turn only when the following assistant message carries
//! `reasoning_content`, issue #2089). An echoed trace is forwarded under
//! both `reasoning` and `reasoning_content`, so such a turn re-renders
//! exactly as it was generated. A client that echoes only `content`
//! still gets the rewritten (shorter) turn; that is the template's own
//! rule and the follow-up misses the cache by design.
//! * Tool-schema hashing: [`super::prompt_cache::key::tools_digest`] is
//! order-preserving, so reordering tools invalidates the cache. This
//! is intentional: HuggingFace templates iterate tools in order and
Expand Down Expand Up @@ -1871,12 +1879,25 @@ fn build_raw_json_messages_with_thinking(
// carries an inline `<think>` block. Forwarding it on top of an
// inline block would double-inject the same reasoning into
// templates that render both channels.
//
// The value goes out under both spellings the wire accepts (issue
// #2089). The wire layer folds `reasoning` and `reasoning_content`
// into one field, but a template reads one of them: Gemma 4 reads
// `reasoning`, while Qwen3-style and AI21 Jamba-Reasoning templates
// read `message.reasoning_content`. Emitting only `reasoning` hid
// an echoed trace from the second group, and Jamba's template keys
// the thinking instruction on an earlier user turn to that field,
// so the turn re-rendered shorter than it had been generated and
// every follow-up missed the prompt cache. Templates that accept
// both read them as alternatives (`reasoning or
// reasoning_content`), so this never renders the trace twice.
if !stripped
&& let Some(reasoning) = m.reasoning.as_ref()
&& !reasoning.is_empty()
&& !raw_content.contains("<think>")
{
msg["reasoning"] = serde_json::Value::String(reasoning.clone());
msg["reasoning_content"] = serde_json::Value::String(reasoning.clone());
}

msg
Expand Down Expand Up @@ -2113,3 +2134,7 @@ fn build_chat_messages_with_thinking(
#[cfg(test)]
#[path = "chat_request_tests.rs"]
mod tests;

#[cfg(test)]
#[path = "chat_request_reasoning_content_tests.rs"]
mod reasoning_content_tests;
Loading