feat(telemetry): interruption detail, handoff span, fallback events, text input - #2499
Conversation
🦋 Changeset detectedLatest commit: b9709d2 The changes in this PR will be included in the next version bump. This PR includes changesets to release 39 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
b21aa74 to
f3e8092
Compare
8b62e64 to
a6e6bd6
Compare
a6e6bd6 to
fc1883c
Compare
b9fc426 to
b4ecdb0
Compare
a7c5d40 to
cd08174
Compare
3f7c958 to
e0ee56e
Compare
e0ee56e to
a94afcd
Compare
a94afcd to
56df935
Compare
56df935 to
0c3efe3
Compare
There was a problem hiding this comment.
Note
Newer findings are available below. Devin Review posted a newer report on this PR, in addition to the findings presented here.
Devin Review found 1 new potential issue.
3 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
| override get model(): string { | ||
| return this.nextInstance().model; | ||
| } | ||
|
|
||
| /** The provider of the instance that serves next (see {@link model}). */ | ||
| override get provider(): string { | ||
| return this.nextInstance().provider; |
There was a problem hiding this comment.
🟡 Recovered primary misattributes fallback usage
When primary recovery finishes before metrics finalize, model and provider switch back despite the secondary serving. The secondary's usage is attributed to the primary.
Learn more
Fallback streams emit wrapper-level metrics after output collection. Those metrics read the adapter's dynamic model and provider through monitorMetrics. A failed primary starts recovery immediately, and recovery can mark it available while the secondary request is still running. The getters then select the primary even though servedLlm correctly records the secondary on spans. The same timing can affect wrapper-level TTS metrics because its fallback adapter uses equivalent dynamic getters.
Example: The primary fails, then its probe succeeds while the secondary is generating. The secondary returns 100 tokens. By metric finalization, nextInstance() selects the recovered primary, so those 100 tokens carry the primary model and provider.
Recommended fix: Store the serving child per fallback stream and use that child for wrapper metrics. Extend the base LLM/TTS metric construction with protected model/provider getters, analogous to responseModel, then override them in fallback streams. Keep adapter-level getters for pre-request attribution.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Codex found the same issue:
| Step | What happens |
|---|---|
| 1 | Primary fails. The adapter sends the request to secondary. |
| 2 | Secondary starts generating. A background probe checks whether primary has recovered. |
| 3 | The probe succeeds. Primary becomes the preferred provider for the next request. |
| 4 | Secondary finishes with 100 output tokens. |
| 5 | The adapter emits usage metrics. Its model and provider getters now return primary. |
82c8dc2 to
1e87395
Compare
a1a3905 to
d4824c7
Compare
chenghao-mou
left a comment
There was a problem hiding this comment.
Can we also port livekit/agents#7373 and livekit/agents#7374 here?
| override get model(): string { | ||
| return this.nextInstance().model; | ||
| } | ||
|
|
||
| /** The provider of the instance that serves next (see {@link model}). */ | ||
| override get provider(): string { | ||
| return this.nextInstance().provider; |
There was a problem hiding this comment.
Codex found the same issue:
| Step | What happens |
|---|---|
| 1 | Primary fails. The adapter sends the request to secondary. |
| 2 | Secondary starts generating. A background probe checks whether primary has recovered. |
| 3 | The probe succeeds. Primary becomes the preferred provider for the next request. |
| 4 | Secondary finishes with 100 output tokens. |
| 5 | The adapter emits usage metrics. Its model and provider getters now return primary. |
chenghao-mou
left a comment
There was a problem hiding this comment.
Can we also port livekit/agents#7373 and livekit/agents#7374 here?
d4824c7 to
c8d87a5
Compare
c8d87a5 to
3272149
Compare
3272149 to
efde259
Compare
There was a problem hiding this comment.
Devin Review found 3 new potential issues.
4 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
| // Keep child timestamps anchored to the parent stream's current retry attempt. | ||
| child.startTimeOffset = this.startTimeOffset + (Date.now() - startTime) / 1000; | ||
| mainRef.current = child; | ||
| this.fallbackAdapter._servedStt = sttInstance; |
There was a problem hiding this comment.
🟡 Concurrent streams misattribute STT turns
When two streams share an adapter, each overwrites or clears _servedStt for the other. Their user_turn spans can name a provider that never transcribed that turn.
Learn more
Each FallbackSpeechStream assigns its chosen child to a single field on the adapter. stampSttIdentity reads that shared field through the adapter's getters at turn creation and completion. A second stream can change it while the first stream still transcribes; either stream's cleanup also clears the field regardless of who set it.
Example: Stream A uses fallback while stream B starts on a recovered primary. A's turn finishes during B's stream, so A's span says primary. A closing can then erase B's attribution.
Recommended fix: Attribute the serving child per stream or per recognition event/turn rather than storing stream-local identity on the shared adapter. Avoid unconditional clearing by one stream while another is active.
Was this helpful? React with 👍 or 👎 to provide feedback.
efde259 to
554dece
Compare
…text input Port of livekit/agents#7137. Interruptions: `agent_turn` carries `lk.interruption.source`, set by the caller that knows the cause (`audio_activity` for a VAD/STT barge-in and the realtime server's own speech detection, `user_turn` for a committed turn or a final transcript ending a pause, `programmatic` for session.interrupt(), tools and teardown); the first interruption's cause stands. The pipeline and `say` paths stamp `lk.playout.position`, the seconds actually played when the user cut in. The say and realtime paths now also stamp `lk.interrupted`, which only the pipeline reply did before. Agent handoff: `updateAgent()` opens an `update_agent` span under `agent_session` with `lk.previous_agent_label` / `lk.agent_label`; the old agent's `drain_agent_activity` (with `on_exit`) and the new agent's `start_agent_activity` / `resume_agent_activity` nest under it. The initial start stays under `session_start`. Fallback adapters: LLM, TTS and STT `model` / `provider` follow the instance that serves next (first available, else the primary), so `llm_node`, `tts_node` and `start_agent_activity` name a real model. The LLM and TTS attempt span carries `lk.fallback.label` / `lk.fallback.index` and the serving instance's request model and provider; the adapter's request span and the caller's node span get `gen_ai.response.model` / provider of the instance that answered, per request. The STT adapter's last-served tracking is replaced by the same next-instance rule. Text input: the keyterm-detection LLM pass runs in its own `keyterm_detection` span under the `agent_turn` whose reply added the user message, else under `agent_session`, with counts only (`lk.keyterms.count/added/removed`) plus model and provider. Adaptations: agents carry `id` where Python has `label`; JS `SpeechHandle` and `AgentActivity.interrupt` take the source as a second positional parameter / option instead of a keyword. The LLM and TTS stream base classes expose their request span to subclasses (`llmRequestSpan`, `ttsRequestSpan`), the JS counterpart of Python's `_llm_request_span`. A cancelled preemptive attempt names why it was dropped. SpeechHandle._cancel takes the cause like interrupt does: an attempt superseded by more of the user's turn (a later preemptive trigger, or the transcript changing at commit) is user_turn, one dropped by a barge-in audio_activity through interrupt(); a cancel with no cause (teardown, a pause) still reads as programmatic. A cloud export showed such attempts, cancelled when the user kept talking, labelled programmatic. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ack attribution; port livekit/agents#7373 and #7374 - user_turn names an STT that leaves the base getters at `unknown` by its label (prefix as provider) and normalizes the provider to the GenAI registry spelling - the pipeline reply task stamps its interruption verdict once it is over, so an interruption that lands while the tools run is recorded - an LLM fallback stream's usage metrics name the instance that served, not the one the adapter would pick next after a recovery - #7373: the fallback adapter's request span carries no operation name; the nested provider request is the `chat` - #7374: llm_node records its configured model and provider when the nested request is created, so a failover's serving provider is kept rather than overwritten when the node completes Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ed audio; live STT identity on user_turn - a TTS fallback stream's usage metrics name the instance that served, not the one the adapter would pick next once a failed instance recovered - chunked synthesis that fails after audio reached the caller still names the instance that produced it on the request and caller spans, as the streaming path does - user_turn reads the STT identity when it stamps the turn, at its start and again at its end, so a fallback that failed over names the instance that transcribed rather than the one snapshotted at activity start Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rving it model/provider name the instance serving the open stream (or the one that answered the last recognize()), and only fall back to the next-in-line instance between streams. A recovery probe finding the primary back no longer relabels a turn the fallback is transcribing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- only the electing stream clears it, so a stream ending cannot erase a newer stream's attribution; the one-slot limit is documented - recognize() no longer pins the slot: between requests the getters follow availability again, as a completed recognize() has no turn in flight Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
554dece to
b9709d2
Compare
Port of livekit/agents#7137. Stacked on #2498.
Description
Smaller coverage gaps on existing spans, each of which came up when reading a trace and not being able to answer a question from it.
Interruptions.
agent_turncarried a singlelk.interruptedboolean. It now also carrieslk.interruption.source, set by the caller that knows the cause:audio_activity(barge-in from VAD / STT activity),user_turn(a committed user turn preempting the reply),programmatic(session.interrupt(), a tool, teardown). The first cause wins. The pipeline andsaypaths also stamplk.playout.position, how many seconds had actually played when the user cut in.A cancelled preemptive attempt names why it was dropped too:
SpeechHandle._canceltakes the cause likeinterrupt. An attempt superseded by more of the user's turn (a later preemptive trigger, or the transcript changing at commit) isuser_turn, one dropped by a barge-in isaudio_activity; a cancel with no cause (teardown, a pause) still reads asprogrammatic. A cloud export showed such attempts, cancelled while the user kept talking, labelledprogrammatic.Agent handoff.
updateAgent()spans anupdate_agentbar (parentagent_session) withlk.previous_agent_labelandlk.agent_label. The old agent'sdrain_agent_activity(withon_exitinside) and the new agent'sstart_agent_activitynest under it. The initial start stays undersession_start.Fallback adapters (LLM, TTS, STT).
FallbackAdapter.model/.providerfollow the instance that serves next, sollm_node,tts_nodeandstart_agent_activityname a real model instead of the adapter. The attempt span carrieslk.fallback.label/lk.fallback.index, and a failover mid-request is recorded on the response side:gen_ai.response.modelandgen_ai.provider.nameof the serving instance on the adapter's request span and onllm_node/tts_node. Usage metrics needed no change.Text input. A
keyterm_detectionspan around the keyterm-detection LLM pass, nested under theagent_turnthat answers the user message (falling back toagent_session), so itsllm_requestno longer looks like a second inference step. Attributes are counts only (lk.keyterms.count/added/removed) plus model and provider; the terms stay in the session report as PII.Changes Made
voice/speech_handle.ts:InterruptionSource,interrupt(force, source);voice/agent_activity.ts:recordInterruption, source threaded through every interrupt path,lk.playout.position;voice/agent_session.ts:update_agentspan.llm|tts|stt/fallback_adapter.ts:nextInstance(),model/provideroverrides, served-instance attribution;llm/llm.ts,tts/tts.ts: protected request-span getters.voice/keyterm_detection.ts:keyterm_detectionspan.telemetry/trace_types.ts: eight new attributes (none PII).Adaptations from the Python source
lk.agent_label/lk.previous_agent_labeluseagent.id.recordInterruptionalso stampslk.interrupted=true: the JSsayand realtime paths never set it before (only the pipeline reply did)._activeSttlast-served tracking is removed with the getters that read it.llm_fallback_adapterspan name has no JS counterpart; the adapter's request span keeps the fixedllm_requestname and is found through itsllm_request_runchild.Testing
voice/coverage_spans.test.ts(barge-in source and playout position through a fake session,update_agentnesting, LLM fallback serving provider, first-cause-wins) andvoice/keyterm_detection_span.test.ts; fallback adapter tests updated for next-instance identity;agent_session_handoff.test.tsaccepts the threaded trace context.agentssuite green; build, typecheck, lint, API report updated.Review follow-ups
user_turnas their interruption source instead of theprogrammaticdefault (agent_activity.ts, test: "a committed user turn interrupts the queued replies for the same reason").LLMStreamstamps the request span's response model through a new protectedresponseModelgetter, which the fallback stream overrides with the served instance: the adapter's ownmodelnames the instance it would pick next, which after a partial failure is one the caller never heard from (llm.ts,llm/fallback_adapter.ts; test: "names the instance whose partial response the caller received before it failed"). Python's base stream has the same overwrite (model=self._llm.model), so the Python fallback needs the same two changes.user_turncarries the STT'smodelandprovider, as Python passes them, instead of the STT's label and a label-derived provider guess (agent_activity.ts; test: "names the STT model and provider on user_turn, as python does"). For a fallback adapter that is the instance expected to serve next, at activity start; a mid-turn failover is not reflected, matching Python.🤖 Generated with Claude Code