feat(anthropic): 补齐停止原因与 Fable 5.1 / Opus 5.5 兼容性 - #58
Merged
Merged
Conversation
安全分类器拒答是 HTTP 200 + 空 content + stop_reason="refusal",引擎看到的
形状与"模型选择不说话"完全一致:无文本、无 tool call → 走 no_tool 退出 →
整跑结束且 trace 上看起来干净。Anthropic 的 stop_details(type / category /
explanation)此前从未被读取,response_metadata 只有 {"id": ...}。
2026-10-05 gdpval-220 取证正是卡在这里:20 个 run 以 llm_error 收场、另有 25 个
以 no_tool + 空交付物收场,而 trace / job.log / trial.log / breaker last_error
四处都没有记录哪些是拒答,只能逐个 case 手工解剖。
本次补齐三段:
1. providers/anthropic.py —— 新增 _anthropic_stop_details(),同时支持 SDK
model 对象与 dict;_to_llm_response 把 RAW stop_reason 与 stop_details 放进
response_metadata。raw 值与归一化后的 finish_reason 并存:max_tokens 会被
归一成 length,消费者不该为了判断"provider 是否拒答"而去记住哪些标记被重写。
category 是开放集合(cyber / bio / reasoning_extraction / frontier_llm /
general_harms / …,新模型持续新增),原样透传不做白名单。
2. 流式路径 —— StreamDelta 新增 stop_details 字段(沿用 provider 字段"终端
元数据折进 response_metadata"的既有模式),anthropic 的 message_delta 捕获
它,_streaming 的组装器折进 response_metadata。否则一次流式拒答仍然隐形。
3. TurnContext 新增 finish_reason / stop_details 字段,agent_loop 填充,
TrajectoryFileObserver 写进 JSONL。finish_reason 此前只在 metadata 里,
观察者拿不到一等字段。end_turn 不写(无信息量且几乎每行都有),其余
(length / tool_use / refusal / …)记录。
两个字段都有默认值,未迁移的 TurnContext 构造点保持有效 —— 不设置的生产者
与改动前一样看不见,但不会出错。
测试:新增 test_trajectory_stop_reason.py(5 条)+ test_provider_native_clients.py
的 refusal 组(7 条,含一条走真实 SDK 解析的 MockTransport 端到端)+
test_llm_runtime_stream_metadata.py 的流式组(2 条)。全量 1779 passed。
ruff 通过。
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to 20b288f (whose body was lost to truncation). Summary of that commit: forward effort without thinking; keep tool_use blocks in provider order among signed thinking (canonical tool_calls stay authoritative); retry once without historical thinking when a thinking signature is rejected, and report thinking_history_reset so the loop drops stale signatures before storing the new turn; record finish_reason/stop_details in JSON trajectories. This commit: - Signature retry now matches "signature" + "thinking" instead of the exact "bound to a different conversation" phrase. The wording is not a documented contract, and omitting historical thinking is always a valid request, so a narrower match would silently disable recovery on any rewording. - Trajectories also omit finish_reason "stop" (OpenAI-compatible normal end), matching the existing end_turn omission. - Docs/changelog: note the per-request double cost for direct client callers that ignore thinking_history_reset, and the effort-without-thinking change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Anthropic 的 classifier refusal 是 HTTP 200,可能没有内容,也可能留下部分文本。此前 trace 无法可靠区分拒答、正常结束和输出截断。本次将停止原因完整传递到响应与两种 trajectory 格式,并补齐当前 Claude 的签名历史兼容性。
response_metadata保留原始stop_reason与结构化stop_details;finish_reason继续将max_tokens归一为length。未知 category 和 SDK 扩展字段透传。TurnContext向观察者提供finish_reason、stop_details,JSON 与 JSONL 均记录;新增字段有默认值。claude-fable-5-1和claude-opus-5-5:未配置 thinking 时仍转发 effort;空 thinking 文本的签名与 interleaved tool-use 顺序完整保留,canonical tool calls 在过滤/修复后仍为权威来源。changes/58.feature.md,修复version-bump检查;项目版本保持 0.12.3。验证:
模型兼容性依据:Fable 5.1 迁移指南、Opus 5.5 迁移指南、preserved thinking。
本次保持现有 loop 终止策略;拒答后的 fallback 模型选择仍由 host 决定。直接使用 client 的调用方应消费
thinking_history_reset,同步清除自己的旧 thinking 历史。