feat(telemetry): one agent_turn span per speech handle - #2500
Conversation
9a563af to
91cea4d
Compare
🦋 Changeset detectedLatest commit: 2cc06f9 The changes in this PR will be included in the next version bump. This PR includes changesets to release 39 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
91cea4d to
dd4e334
Compare
241cd41 to
3c94ca8
Compare
3c94ca8 to
47d445d
Compare
47d445d to
c5d72a6
Compare
c5d72a6 to
2ee50d2
Compare
2ee50d2 to
3c49f49
Compare
548c6e5 to
308e57f
Compare
308e57f to
e95ba38
Compare
3d4201a to
fb87ab8
Compare
0aca960 to
c150336
Compare
c150336 to
b50cc56
Compare
b50cc56 to
31a8d6d
Compare
31a8d6d to
bb31be9
Compare
bb31be9 to
7ae16ae
Compare
7ae16ae to
291af61
Compare
291af61 to
6ebf417
Compare
6ebf417 to
31677aa
Compare
There was a problem hiding this comment.
Devin Review found 2 new potential issues.
2 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
| this._agentTurnSpan = span; | ||
| this._agentTurnContext = trace.setSpan(otelContext.active(), span); | ||
| this._agentTurnStartedAt = startedAt; | ||
| this._agentTurnAgentName = agentName; | ||
| this._agentTurnGenerations = carry.generations; |
There was a problem hiding this comment.
🟡 Discarded preemptive turn loses initial queue wait
After a preemptive handoff, _queueWaitRecorded remains false on the successor. Its authorization overwrites the first generation's queue wait on the shared span.
Learn more
The queue-wait attribute is meant to describe the first generation of a turn. recordQueueWait stamps it once per handle, guarded by _queueWaitRecorded. A discarded preemptive handle can already have recorded that attribute, but _continueAgentTurn copies only the span and generation count. When the replacement handle is authorized, its fresh flag permits another stamp on the same span.
Example: An initial preemptive attempt waits 300 ms before authorization; a replacement waits 10 ms. The exported turn reports 10 ms rather than the first generation's 300 ms.
Recommended fix: Carry _queueWaitRecorded with the turn and restore it on the successor, or guard the attribute at the span/turn level rather than per handle.
Was this helpful? React with 👍 or 👎 to provide feedback.
31677aa to
81529d0
Compare
Port of livekit/agents#7143. A reply that calls a tool runs two generations (LLM steps) in two tasks; they were two agent_turn spans linked only by lk.parent_generation_id, so one response rendered as two turns. A speech handle is now exactly one agent_turn. - The speech handle owns the span: the first reply task (pipeline, realtime, or say) opens agent_turn under agent_session; the follow-up generation after a tool call continues the open span. It ends with the speech in SpeechHandle._markDone, recording the speech's error redaction-aware and the gen_ai.invoke_agent.duration histogram for the whole turn (new in JS: otel_metrics.recordInvokeAgentDuration and trace_types.METRIC_GEN_AI_INVOKE_AGENT_DURATION; the Python metric already existed). - Each generation is a `generation` event with lk.generation_id and lk.parent_generation_id (SpeechHandle._generationId / _parentGenerationId, `<speech id>_<step>` like Python); the span carries the latest generation id and the new lk.generation_count. - A preemptive generation discarded for a successor answering the same user turn hands its open agent_turn over (preemptive_generation_discarded event, lk.speech_id follows the speech that answered), both on a newer attempt and on the real reply after onUserTurnCompleted invalidated it. The queue-wait and interruption helpers tolerate an ended span. - `say` gets an agent_turn too, as in Python's _tts_task; JS had none. Tests: agent_turn_span.test.ts (tool call is one turn with two generation events and every step nested; plain reply is one generation; discarded preemptive hand-off; LLM failure fails the turn; duration metric when sampled out; sampled-out hand-off). The preemptive-guard stand-in handle gained _takeAgentTurn; the PII key test skips METRIC_* names, which are not attribute keys. The handoff runs inside generateReply, before the reply task is created: Task starts its body synchronously and the task opens the turn in its first statements, so a handoff performed after generateReply() returned came too late, leaving the successor's own agent_turn unended and its llm_node, tts_node and function_tool dangling from a span that was never exported (seen in a cloud export as bare nodes and turns whose generation count did not match their events). A regression test drives the real reply task. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… the turn's speech id - endAgentTurn records a thrown string or object as an exception, so the turn does not read as a success while the handle reports a failure - the reply steps stamp the turn's speech id (the root speech's when a realtime tool reply continues its parent's turn on a new handle) instead of overwriting it with the step's own handle id Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With queue waits measured per generation, a tool reply on the same handle would replace the turn's lk.speech_queue_wait with its own; the span keeps the wait before the reply started. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rd; chained reply ids - a realtime model that answers its tool calls itself (autoToolReplyGeneration) opens the reply generation after the tool call's speech is done: the turn is now held open for it (AutoToolReplyTurnHold) and adopted by the model's next generation, ended after 5 s or at activity close when none comes - the pipeline task stamps lk.interrupted only while its speech still owns the turn, so a discarded preemptive attempt unwinding cannot write the successor's verdict - a handle that adopted a turn but never emitted a generation passes the numbering on from where it stood, so the next id follows an existing one; chained replies tested to three generations Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
81529d0 to
2cc06f9
Compare
Port of livekit/agents#7143. Stacked on #2499.
Description
A reply that calls a tool runs two generations (LLM steps) in two tasks. They were two
agent_turnspans linked only bylk.parent_generation_id, so one response rendered as two turns. A speech handle is now exactly oneagent_turn.say) opensagent_turnunderagent_session; the follow-up generation after a tool call continues the open span instead of opening a second one. It ends with the speech inSpeechHandle._markDone, recording the speech's error (redaction-aware) and thegen_ai.invoke_agent.durationmetric for the whole turn.generationevent on the span withlk.generation_idandlk.parent_generation_id; the span'slk.generation_idnames the latest generation and the newlk.generation_counthow many there were.lk.speech_idis set at creation.llm_node,function_tool,tts_node,realtime_inferenceandagent_speakingnest under the one turn.agent_turnto that successor (preemptive_generation_discardedevent;lk.speech_idfollows the speech that answered). An attempt cancelled with no successor still ends as its own turn. The queue-wait and interruption helpers tolerate a span already ended with the speech.Changes Made
voice/speech_handle.ts: span ownership,_generationId/_parentGenerationId,_takeAgentTurn/_continueAgentTurn, end in_markDone.voice/agent_activity.ts:withAgentTurnandcontinueDiscardedTurnwrap the pipeline, realtime andsaytasks; per-step span creation removed.telemetry/otel_metrics.ts:recordInvokeAgentDuration(gen_ai.invoke_agent.duration, units, job attribution);telemetry/trace_types.ts:lk.generation_count, the metric name.Adaptations from the Python source
saypreviously had noagent_turnspan in JS; it now gets one like Python's TTS task.lk.generation_id/lk.parent_generation_idwere never stamped onagent_turnin JS before (the constants existed unused); they are now, as<speech_id>_<step>.Testing
voice/agent_turn_span.test.ts(6 tests): a tool-calling reply is oneagent_turnwith twogenerationevents and bothllm_nodes, the tool,tts_nodeandagent_speakinginside it; a plain reply is one generation; the discarded preemptive hand-off; an LLM failure fails the turn; the duration metric when sampled out.agentssuite green; build, typecheck, lint, API report updated.Review follow-ups
agent_turninstead of opening a second one:continueToolReplyTurnhands the open span over, the turn keeps the tool call's speech id, and the reply's generation is numbered after and parented to the tool call's (lk.generation_id/lk.parent_generation_idcontinue the original handle's numbering). The exported trace matches Python's: one turn, two generations, parent link intact.exception(): the first non-cancellation rejection of an owned task is stored on the handle before it is marked done. The pipeline also stores a genuine LLM node failure the way Python's_on_llm_task_donedoes, since nothing awaits that task's rejection.lk.generation_countcounts the generation events on the turn, carried across handoffs, instead of the current handle's own step count: a turn continued from a discarded preemptive attempt or a tool call reports every generation it shows. Python's_continue_agent_turnneeds the same carry.🤖 Generated with Claude Code