diff --git a/CHANGELOG.md b/CHANGELOG.md index 6116804e..e099851a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,6 +23,28 @@ User-visible changes to Imp are recorded here. unknown charge, not a free one; a host that wants the old number for calls with no reported charge reads `estimated_cost` for them, knowing it is an estimate. +- Breaking: `Imp.Predict.ReActV2`'s `:last_request_note` reaches the model + as a user message with no assistant reply after it, as documented, in the + last request and whenever the returned history is passed back. It was + stored as a history entry with inputs and no outputs. The chat adapter + renders that as a finished exchange, so the note was followed by an + assistant message the model never gave, each field reading "Not supplied + for this conversation history message." That filler is Imp's; DSPy 3.2.1 + renders a missing history output as `None`. The note, and the entry for + inputs no step spent that comes before it when the first step failed, now + carry `tool_calls: %Imp.Adapter.Types.ToolCalls{tool_calls: []}` and + `tool_call_results: []`, like every other step the loop records. Migration: a + host that recognises the note in `metadata.history` by its shape (only the + first input's key) matches it by that key with an empty `tool_calls` list, + or by its text. A note recorded by an earlier version keeps the old + shape and still renders with the filler. `History` entries a host writes + keep Imp's rendering. +- In written tool mode (an LM that cannot call tools natively, where earlier + steps replay as text), a stored step that recorded no call and no other + output replays as its user message alone, as native replay already did. A + ReActV2 step that answered with nothing no longer replays as an assistant + message of filler. A turn that recorded any output, an answer included, + keeps its assistant message. ### Fixed diff --git a/decisions.md b/decisions.md index 601dfce1..ffaad55b 100644 --- a/decisions.md +++ b/decisions.md @@ -8,6 +8,7 @@ not necessarily when it was made. | Date | Decision | Source and reason | Status | Retires when | | --- | --- | --- | --- | --- | +| 2026-09-28 | A turn ReActV2 records with only a user side (its `:last_request_note`, inputs no step spent) is a step event that called nothing, and the chat adapter renders a step with no calls and no other output as its user message alone, in native and written tool modes. A stored step that recorded any other output keeps its assistant message. A `History` entry with inputs and no outputs keeps Imp's rendering, a finished exchange whose missing outputs read "Not supplied for this conversation history message." (DSPy 3.2.1 renders them as `None`; that divergence is separate.) | deepfates/imp#259. `lib/imp/predict/react_v2.ex` `append_user_turn/2`, `lib/imp/adapter/chat.ex` `render_written_tool_history_turn/3`, `test/react_v2_last_request_note_test.exs`, `test/adapter_chat_written_history_test.exs`. Recorded as inputs alone, the note rendered as a user message followed by an assistant reply the model never gave, in the request and on replay. | In force. | Does not retire. | | 2026-09-28 | A model call's `cost` (`Imp.Core.LMResponse`, the `:model_response` event) is only a charge the provider reported, and `nil` when it reported none; ReqLLM's catalog price is `estimated_cost`, never `cost`. | deepfates/imp#252. `lib/imp/core.ex` `LMResponse` moduledoc, `test/model_response_cost_test.exs`. Hosts sum `cost` as money spent (Dwell's daily cap); the catalog price differs from the charge when prices change, when routing picks another endpoint, and by ReqLLM rounding each line item to a millionth of a dollar. | In force. | Does not retire. | | 2026-09-28 | Imp follows semantic versioning. Before 1.0, a release that changes what a caller receives or can rely on bumps the minor version (`0.5` to `0.6`), and a release of fixes that change nothing a caller relies on bumps the patch; `{:imp, "~> 0.x"}` then never takes a breaking release unasked. | Owner, 2026-09-28. `CHANGELOG.md` marks each breaking change "Breaking:" with its migration, and `RELEASE_NOTES.md` names them. | In force. | Does not retire. | | 2026-09-26 | `Imp.Optimizer.GEPA` with no `:execution_profile` runs DSPy's GEPA (`:gepa_v0_1_4_merge`, or `:gepa_v0_1_4` with `use_merge: false`), and its reflection records are DSPy's; Imp's own search is `execution_profile: :beam_native`, chosen explicitly. | Owner, 2026-09-26. An upstream name promises upstream semantics; the published GEPA results (`research/RESULTS.md` R3, R5) used the pinned profile, as `:gepa_v0_1_4`; and no reason for the BEAM-native default was ever recorded. `lib/imp/optimizer/gepa.ex` moduledoc, `test/gepa_agent_reflection_test.exs`. | In force. | Does not retire. | diff --git a/lib/imp/adapter/chat.ex b/lib/imp/adapter/chat.ex index 7a737484..57265c07 100644 --- a/lib/imp/adapter/chat.ex +++ b/lib/imp/adapter/chat.ex @@ -1257,7 +1257,31 @@ defmodule Imp.Adapter.Chat do [%{role: :user, content: renderers.input_section.(field, format_value(value))}] end - Enum.reject([user, assistant | result_messages], &blank_message?/1) + # A turn with no calls and no other output has no assistant message, as in + # native replay. That is a user turn the loop recorded on its own (a note, + # inputs no step spent) or a model reply that said nothing; a filler reply + # would put words in the model's mouth. A turn that recorded any output, + # an answer included, keeps its assistant message. + assistant = + if calls == [] and not recorded_output?(signature, turn), do: nil, else: assistant + + Enum.reject([user, assistant | result_messages], &(is_nil(&1) or blank_message?(&1))) + end + + defp recorded_output?(signature, turn) do + calls_field = + Map.get(signature.metadata, :tool_calls_field) || + Map.get(signature.metadata, "tool_calls_field") + + signature.outputs + |> Enum.reject(&Imp.FieldMap.same_name?(&1.name, calls_field)) + |> Enum.any?(fn field -> + case fetch_field(turn, field.name) do + nil -> false + value when is_binary(value) -> String.trim(value) != "" + _value -> true + end + end) end # A loop whose guidance names no finish tool answers in plain text, and has diff --git a/lib/imp/predict/react_v2.ex b/lib/imp/predict/react_v2.ex index 885f1334..137f9f03 100644 --- a/lib/imp/predict/react_v2.ex +++ b/lib/imp/predict/react_v2.ex @@ -115,8 +115,10 @@ defmodule Imp.Predict.ReActV2 do The last request says nothing about why it is being made unless `:last_request_note` is given: one line of host text put in front of it as a - user message and kept in the returned history like any other turn. Imp - writes no sentence of its own. + user message and kept in the returned history like any other turn. It is + recorded as a step that said nothing and called nothing, so it renders as + that user message alone, with no assistant reply after it, in the request and + whenever the history is passed back. Imp writes no sentence of its own. A request refused because the context window is full is not an interruption of this kind: a further request would be refused the same way, @@ -700,7 +702,7 @@ defmodule Imp.Predict.ReActV2 do # the history first, as the user turn they are. defp note_after_inputs(history, pending, %{last_request_note: note} = react) when is_binary(note) and note != "" and map_size(pending) > 0 do - history = history |> append_history(pending) |> append_note(react.signature, note) + history = history |> append_user_turn(pending) |> append_note(react.signature, note) {history, %{}} end @@ -768,13 +770,26 @@ defmodule Imp.Predict.ReActV2 do # prompt renders it as the last user message before the last request. defp append_note(history, signature, text) when is_binary(text) and text != "" do case Imp.Signature.input_names(signature) do - [first | _rest] -> append_history(history, %{first => text}) + [first | _rest] -> append_user_turn(history, %{first => text}) [] -> history end end defp append_note(history, _signature, _none), do: history + # A turn the loop records with only a user side (the note, or inputs no step + # spent) is a step event that called nothing and said nothing. The chat + # adapter renders such a step as its user message alone, in every tool mode, + # and a host that hands the history back gets the same rendering. Stored as + # inputs alone, it would be a `History` entry, which the adapter renders as a + # finished exchange with a filler reply the model never gave. + defp append_user_turn(history, fields) do + append_history( + history, + Map.merge(fields, %{tool_calls: %ToolCalls{tool_calls: []}, tool_call_results: []}) + ) + end + defp forced_submit_prediction(react, history, pending) do forced = forced_submit_program(react, %{type: "tool", name: "submit"}) diff --git a/test/adapter_chat_written_history_test.exs b/test/adapter_chat_written_history_test.exs new file mode 100644 index 00000000..e8fc97ed --- /dev/null +++ b/test/adapter_chat_written_history_test.exs @@ -0,0 +1,77 @@ +defmodule Imp.Adapter.ChatWrittenHistoryTest do + use ExUnit.Case, async: true + + # A request whose signature describes a tool-calls output replays its stored + # turns as text: the turn's outputs as the assistant's message, then the + # results. A turn that recorded no output and no call has no assistant + # message; any recorded output, an answer included, keeps it. + + defmodule TextOnlyLM do + @behaviour Imp.LM + defstruct [:handler] + + @impl true + def generate(%__MODULE__{handler: handler}, messages, opts), + do: {:ok, handler.(messages, opts)} + + def tool_calling_capability(%__MODULE__{}), do: false + end + + @filler "Not supplied for this conversation history message." + + defp text_only_lm(replies) do + owner = self() + counter = :counters.new(1, []) + + %TextOnlyLM{ + handler: fn messages, _opts -> + :counters.add(counter, 1, 1) + send(owner, {:request, messages}) + Enum.at(replies, :counters.get(counter, 1) - 1) + end + } + end + + test "a Predict step with a tool-calls field keeps an answered turn's assistant message" do + signature = + Imp.Signature.ensure("question, history, tools: array -> answer, tool_calls: array") + + signature = %{ + signature + | metadata: Map.put(signature.metadata, :tool_calls_field, :tool_calls) + } + + history = Imp.History.new([%{question: "q1", answer: "42", tool_calls: []}]) + + program = + Imp.Predict.new(signature, + lm: text_only_lm(["[[ ## answer ## ]]\n43\n\n[[ ## tool_calls ## ]]\n[]"]) + ) + + assert {:ok, _prediction} = + Imp.Predict.call(program, %{question: "q2", history: history, tools: []}) + + assert_received {:request, messages} + index = Enum.find_index(messages, &(&1.role == :user and &1.content =~ "q1")) + assert %{role: :assistant, content: answered} = Enum.at(messages, index + 1) + assert answered =~ "42" + end + + test "a ReActV2 step that answered with nothing replays with no assistant message" do + lm = text_only_lm(["", "[[ ## next_thought ## ]]\nsecond\n\n[[ ## tool_calls ## ]]\n[]"]) + look = Imp.tool(:look, "Look at a thing", fn _ -> %{"seen" => true} end) + program = Imp.react("intent -> answer", [look], lm: lm) + + assert {:ok, first} = Imp.call(program, %{intent: "hello"}) + assert first.metadata.termination_reason == :answered + assert_received {:request, _first} + + assert {:ok, _second} = + Imp.call(program, %{intent: "again", history: first.metadata.history}) + + assert_received {:request, replayed} + refute Enum.any?(replayed, &(to_string(&1.content) =~ @filler)) + index = Enum.find_index(replayed, &(&1.role == :user and &1.content =~ "hello")) + refute Enum.at(replayed, index + 1).role == :assistant + end +end diff --git a/test/react_v2_last_request_note_test.exs b/test/react_v2_last_request_note_test.exs index 22e914e1..56be8821 100644 --- a/test/react_v2_last_request_note_test.exs +++ b/test/react_v2_last_request_note_test.exs @@ -71,6 +71,136 @@ defmodule ReActV2LastRequestNoteTest do refute Enum.any?(forced, &(&1[:role] == :user and String.trim(&1[:content] || "") == "")) end + # An LM that writes its tool calls as text: the step lists the tools and + # replays earlier steps as text, never as native tool messages. + defmodule TextOnlyLM do + @behaviour Imp.LM + defstruct [:handler] + + @impl true + def generate(%__MODULE__{handler: handler}, messages, opts), + do: {:ok, handler.(messages, opts)} + + def tool_calling_capability(%__MODULE__{}), do: false + end + + @filler "Not supplied for this conversation history message." + + # The first request looks; every later one submits. The replies are native + # tool calls, or the same calls written as text for an LM that cannot call + # tools. + defp moded_lm(owner, native?) do + counter = :counters.new(1, []) + + handler = fn messages, opts -> + :counters.add(counter, 1, 1) + n = :counters.get(counter, 1) + send(owner, {:request, n, messages, opts}) + step_reply(native?, if(n == 1, do: :look, else: :submit)) + end + + if native?, do: Imp.LM.Static.new(handler: handler), else: %TextOnlyLM{handler: handler} + end + + defp step_reply(true, :look), + do: %{next_thought: "look first", tool_calls: [%{id: "c", name: "look", arguments: %{}}]} + + defp step_reply(true, :submit), + do: %{tool_calls: [%{id: "s", name: "submit", arguments: %{answer: "ok", confidence: 1.0}}]} + + defp step_reply(false, :look), + do: + ~s([[ ## next_thought ## ]]\nlook first\n\n[[ ## tool_calls ## ]]\n[{"name": "look", "arguments": {}}]) + + defp step_reply(false, :submit), + do: + ~s([[ ## next_thought ## ]]\nsubmitting\n\n[[ ## tool_calls ## ]]\n[{"name": "submit", "arguments": {"answer": "ok", "confidence": 1.0}}]) + + defp note_index(messages, note), + do: Enum.find_index(messages, &(&1.role == :user and to_string(&1.content) =~ note)) + + for {mode, native?} <- [native: true, prompt: false] do + @native native? + + test "#{mode} tools: the note ends the forced request as a user message, with no invented reply" do + owner = self() + note = "You have used every turn. Submit the answer you have now." + + program = + Imp.react(@signature, [look()], + lm: moded_lm(owner, @native), + max_iters: 1, + last_request_note: note + ) + + assert {:ok, prediction} = Imp.call(program, %{intent: "hello"}) + assert prediction.metadata.termination_reason == :forced_submit + + [_first, {forced, _opts}] = requests(2) + refute Enum.any?(forced, &(to_string(&1.content) =~ @filler)) + + index = note_index(forced, note) + assert index + after_note = Enum.drop(forced, index + 1) + + # Natively the note is the last message. A step that lists its tools as + # text ends every request on that listing, a user message too. + if @native, + do: assert(after_note == []), + else: + assert([%{role: :user, content: listing}] = after_note) && assert(listing =~ "tools") + end + + test "#{mode} tools: a returned history replays the note as the user turn the model answered" do + owner = self() + note = "You have used every turn. Submit the answer you have now." + + program = + Imp.react(@signature, [look()], + lm: moded_lm(owner, @native), + max_iters: 1, + last_request_note: note + ) + + assert {:ok, prediction} = Imp.call(program, %{intent: "hello"}) + _ = requests(2) + + # A host keeps the history as data and hands it back on the next turn. + history = prediction.metadata.history |> Imp.History.dump() |> Imp.History.load!() + assert {:ok, _prediction} = Imp.call(program, %{intent: "again", history: history}) + + {replayed, _opts} = receive(do: ({:request, 3, m, o} -> {m, o})) + refute Enum.any?(replayed, &(to_string(&1.content) =~ @filler)) + + # What follows the note is the reply the model actually gave: the submit. + index = note_index(replayed, note) + assert index + reply = Enum.at(replayed, index + 1) + assert reply.role == :assistant + + if @native, + do: assert(Enum.map(reply.tool_calls, & &1.function.name) == ["submit"]), + else: assert(reply.content =~ "submit") + end + end + + # A `History` entry a host wrote with inputs and no outputs is a finished + # exchange whose outputs were not recorded, and keeps Imp's rendering, with + # the filler (DSPy 3.2.1 renders `None`). Only the turns the loop records + # itself say they have no reply. + test "a host's input-only history entry keeps Imp's rendering" do + owner = self() + program = Imp.react(@signature, [look()], lm: moded_lm(owner, true), max_iters: 1) + history = Imp.History.new([%{intent: "earlier"}]) + + assert {:ok, _prediction} = Imp.call(program, %{intent: "hello", history: history}) + [{first, _opts}, _forced] = requests(2) + + index = note_index(first, "earlier") + assert %{role: :assistant, content: filler} = Enum.at(first, index + 1) + assert filler =~ @filler + end + test "no note leaves the forced request exactly as it was" do owner = self() diff --git a/test/react_v2_last_text_test.exs b/test/react_v2_last_text_test.exs index e83f4fd5..e4436576 100644 --- a/test/react_v2_last_text_test.exs +++ b/test/react_v2_last_text_test.exs @@ -186,6 +186,9 @@ defmodule ReActV2LastTextTest do [inputs, note] = Enum.take(user_contents(last), -2) assert inputs =~ "hello" assert note =~ "Last one." + + # No step answered either of them, so no assistant turn follows them. + assert Enum.map(Enum.take(last, -2), & &1.role) == [:user, :user] end # A turn whose step and last request both got no response has no answer and