Skip to content

feat(telemetry): interruption detail, handoff span, fallback events, text input - #2499

Merged
davidzhao merged 5 commits into
dz/telemetry-rpcfrom
dz/telemetry-coverage
Sep 26, 2026
Merged

davidzhao merged 5 commits into
dz/telemetry-rpcfrom
dz/telemetry-coverage

Conversation

@davidzhao

@davidzhao davidzhao commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

Port of livekit/agents#7137. Stacked on #2498.

Description

Smaller coverage gaps on existing spans, each of which came up when reading a trace and not being able to answer a question from it.

Interruptions. agent_turn carried a single lk.interrupted boolean. It now also carries lk.interruption.source, set by the caller that knows the cause: audio_activity (barge-in from VAD / STT activity), user_turn (a committed user turn preempting the reply), programmatic (session.interrupt(), a tool, teardown). The first cause wins. The pipeline and say paths also stamp lk.playout.position, how many seconds had actually played when the user cut in.

A cancelled preemptive attempt names why it was dropped too: SpeechHandle._cancel takes the cause like interrupt. An attempt superseded by more of the user's turn (a later preemptive trigger, or the transcript changing at commit) is user_turn, one dropped by a barge-in is audio_activity; a cancel with no cause (teardown, a pause) still reads as programmatic. A cloud export showed such attempts, cancelled while the user kept talking, labelled programmatic.

Agent handoff. updateAgent() spans an update_agent bar (parent agent_session) with lk.previous_agent_label and lk.agent_label. The old agent's drain_agent_activity (with on_exit inside) and the new agent's start_agent_activity nest under it. The initial start stays under session_start.

Fallback adapters (LLM, TTS, STT). FallbackAdapter.model / .provider follow the instance that serves next, so llm_node, tts_node and start_agent_activity name a real model instead of the adapter. The attempt span carries lk.fallback.label / lk.fallback.index, and a failover mid-request is recorded on the response side: gen_ai.response.model and gen_ai.provider.name of the serving instance on the adapter's request span and on llm_node / tts_node. Usage metrics needed no change.

Text input. A keyterm_detection span around the keyterm-detection LLM pass, nested under the agent_turn that answers the user message (falling back to agent_session), so its llm_request no longer looks like a second inference step. Attributes are counts only (lk.keyterms.count/added/removed) plus model and provider; the terms stay in the session report as PII.

Changes Made

  • voice/speech_handle.ts: InterruptionSource, interrupt(force, source); voice/agent_activity.ts: recordInterruption, source threaded through every interrupt path, lk.playout.position; voice/agent_session.ts: update_agent span.
  • llm|tts|stt/fallback_adapter.ts: nextInstance(), model / provider overrides, served-instance attribution; llm/llm.ts, tts/tts.ts: protected request-span getters.
  • voice/keyterm_detection.ts: keyterm_detection span.
  • telemetry/trace_types.ts: eight new attributes (none PII).

Adaptations from the Python source

  • lk.agent_label / lk.previous_agent_label use agent.id.
  • recordInterruption also stamps lk.interrupted=true: the JS say and realtime paths never set it before (only the pipeline reply did).
  • STT: the now-unused _activeStt last-served tracking is removed with the getters that read it.
  • Python's llm_fallback_adapter span name has no JS counterpart; the adapter's request span keeps the fixed llm_request name and is found through its llm_request_run child.

Testing

  • New voice/coverage_spans.test.ts (barge-in source and playout position through a fake session, update_agent nesting, LLM fallback serving provider, first-cause-wins) and voice/keyterm_detection_span.test.ts; fallback adapter tests updated for next-instance identity; agent_session_handoff.test.ts accepts the threaded trace context.
  • Full agents suite green; build, typecheck, lint, API report updated.

Review follow-ups

  • A committed user turn also sweeps the replies queued behind the current one; they now record user_turn as their interruption source instead of the programmatic default (agent_activity.ts, test: "a committed user turn interrupts the queued replies for the same reason").
  • The LLM fallback records the instance whose partial response reached the caller before rethrowing (retry-on-chunk off), as the TTS fallback does for audio already heard. The base LLMStream stamps the request span's response model through a new protected responseModel getter, which the fallback stream overrides with the served instance: the adapter's own model names the instance it would pick next, which after a partial failure is one the caller never heard from (llm.ts, llm/fallback_adapter.ts; test: "names the instance whose partial response the caller received before it failed"). Python's base stream has the same overwrite (model=self._llm.model), so the Python fallback needs the same two changes.
  • user_turn carries the STT's model and provider, as Python passes them, instead of the STT's label and a label-derived provider guess (agent_activity.ts; test: "names the STT model and provider on user_turn, as python does"). For a fallback adapter that is the instance expected to serve next, at activity start; a mid-turn failover is not reflected, matching Python.

🤖 Generated with Claude Code

@changeset-bot

changeset-bot Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: b9709d2

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 39 packages
Name Type
@livekit/agents Patch
@livekit/agents-plugin-anam Patch
@livekit/agents-plugin-anthropic Patch
@livekit/agents-plugin-assemblyai Patch
@livekit/agents-plugin-azure Patch
@livekit/agents-plugin-baseten Patch
@livekit/agents-plugin-bey Patch
@livekit/agents-plugin-cartesia Patch
@livekit/agents-plugin-cerebras Patch
@livekit/agents-plugin-deepgram Patch
@livekit/agents-plugin-did Patch
@livekit/agents-plugin-elevenlabs Patch
@livekit/agents-plugin-fishaudio Patch
@livekit/agents-plugin-google Patch
@livekit/agents-plugin-hume Patch
@livekit/agents-plugin-inworld Patch
@livekit/agents-plugin-krisp Patch
@livekit/agents-plugin-lemonslice Patch
@livekit/agents-plugin-liveavatar Patch
@livekit/agents-plugin-livekit Patch
@livekit/agents-plugin-meta Patch
@livekit/agents-plugin-minimax Patch
@livekit/agents-plugin-mistral Patch
@livekit/agents-plugin-mistralai Patch
@livekit/agents-plugin-neuphonic Patch
@livekit/agents-plugin-openai Patch
@livekit/agents-plugin-perplexity Patch
@livekit/agents-plugin-phonic Patch
@livekit/agents-plugin-protoface Patch
@livekit/agents-plugin-resemble Patch
@livekit/agents-plugin-rime Patch
@livekit/agents-plugin-runway Patch
@livekit/agents-plugin-sarvam Patch
@livekit/agents-plugin-silero Patch
@livekit/agents-plugin-soniox Patch
@livekit/agents-plugin-tavus Patch
@livekit/agents-plugins-test Patch
@livekit/agents-plugin-trugen Patch
@livekit/agents-plugin-xai Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@davidzhao
davidzhao added this pull request to stack #2502 September 15, 2026 06:32
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from b21aa74 to f3e8092 Compare September 15, 2026 06:46
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch 2 times, most recently from 8b62e64 to a6e6bd6 Compare September 16, 2026 04:34
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from a6e6bd6 to fc1883c Compare September 20, 2026 05:02
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch 2 times, most recently from b9fc426 to b4ecdb0 Compare September 20, 2026 06:11
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch 3 times, most recently from a7c5d40 to cd08174 Compare September 20, 2026 07:21
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch 2 times, most recently from 3f7c958 to e0ee56e Compare September 20, 2026 17:43
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from e0ee56e to a94afcd Compare September 20, 2026 18:02
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from a94afcd to 56df935 Compare September 20, 2026 23:39
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from 56df935 to 0c3efe3 Compare September 21, 2026 03:28
@davidzhao
davidzhao marked this pull request as ready for review September 21, 2026 04:32

@devin-ai-integration devin-ai-integration Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Newer findings are available below. Devin Review posted a newer report on this PR, in addition to the findings presented here.

Devin Review found 1 new potential issue.

3 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

Comment on lines +124 to +130
override get model(): string {
return this.nextInstance().model;
}

/** The provider of the instance that serves next (see {@link model}). */
override get provider(): string {
return this.nextInstance().provider;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Recovered primary misattributes fallback usage

When primary recovery finishes before metrics finalize, model and provider switch back despite the secondary serving. The secondary's usage is attributed to the primary.

Learn more

Fallback streams emit wrapper-level metrics after output collection. Those metrics read the adapter's dynamic model and provider through monitorMetrics. A failed primary starts recovery immediately, and recovery can mark it available while the secondary request is still running. The getters then select the primary even though servedLlm correctly records the secondary on spans. The same timing can affect wrapper-level TTS metrics because its fallback adapter uses equivalent dynamic getters.

Example: The primary fails, then its probe succeeds while the secondary is generating. The secondary returns 100 tokens. By metric finalization, nextInstance() selects the recovered primary, so those 100 tokens carry the primary model and provider.

Recommended fix: Store the serving child per fallback stream and use that child for wrapper metrics. Extend the base LLM/TTS metric construction with protected model/provider getters, analogous to responseModel, then override them in fallback streams. Keep adapter-level getters for pre-request attribution.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex found the same issue:

Step What happens
1 Primary fails. The adapter sends the request to secondary.
2 Secondary starts generating. A background probe checks whether primary has recovered.
3 The probe succeeds. Primary becomes the preferred provider for the next request.
4 Secondary finishes with 100 output tokens.
5 The adapter emits usage metrics. Its model and provider getters now return primary.

@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from 82c8dc2 to 1e87395 Compare September 21, 2026 15:35
@davidzhao davidzhao added the review_effort:medium Focused changes that need domain context or careful validation label Sep 21, 2026
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch 2 times, most recently from a1a3905 to d4824c7 Compare September 23, 2026 15:10
devin-ai-integration[bot]

This comment was marked as resolved.

@chenghao-mou chenghao-mou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we also port livekit/agents#7373 and livekit/agents#7374 here?

Comment thread agents/src/voice/agent_activity.ts
Comment on lines +124 to +130
override get model(): string {
return this.nextInstance().model;
}

/** The provider of the instance that serves next (see {@link model}). */
override get provider(): string {
return this.nextInstance().provider;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex found the same issue:

Step What happens
1 Primary fails. The adapter sends the request to secondary.
2 Secondary starts generating. A background probe checks whether primary has recovered.
3 The probe succeeds. Primary becomes the preferred provider for the next request.
4 Secondary finishes with 100 output tokens.
5 The adapter emits usage metrics. Its model and provider getters now return primary.

@chenghao-mou chenghao-mou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we also port livekit/agents#7373 and livekit/agents#7374 here?

@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from d4824c7 to c8d87a5 Compare September 26, 2026 01:56
devin-ai-integration[bot]

This comment was marked as resolved.

@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from c8d87a5 to 3272149 Compare September 26, 2026 05:02
devin-ai-integration[bot]

This comment was marked as resolved.

@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from 3272149 to efde259 Compare September 26, 2026 05:26

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 3 new potential issues.

4 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

Comment thread agents/src/voice/agent_session.ts
Comment thread agents/src/stt/fallback_adapter.ts Outdated
Comment thread agents/src/stt/fallback_adapter.ts Outdated
// Keep child timestamps anchored to the parent stream's current retry attempt.
child.startTimeOffset = this.startTimeOffset + (Date.now() - startTime) / 1000;
mainRef.current = child;
this.fallbackAdapter._servedStt = sttInstance;

@devin-ai-integration devin-ai-integration Bot Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Concurrent streams misattribute STT turns

When two streams share an adapter, each overwrites or clears _servedStt for the other. Their user_turn spans can name a provider that never transcribed that turn.

Learn more

Each FallbackSpeechStream assigns its chosen child to a single field on the adapter. stampSttIdentity reads that shared field through the adapter's getters at turn creation and completion. A second stream can change it while the first stream still transcribes; either stream's cleanup also clears the field regardless of who set it.

Example: Stream A uses fallback while stream B starts on a recovered primary. A's turn finishes during B's stream, so A's span says primary. A closing can then erase B's attribution.

Recommended fix: Attribute the serving child per stream or per recognition event/turn rather than storing stream-local identity on the shared adapter. Avoid unconditional clearing by one stream while another is active.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from efde259 to 554dece Compare September 26, 2026 05:40
davidzhao and others added 5 commits September 25, 2026 22:51
…text input

Port of livekit/agents#7137.

Interruptions: `agent_turn` carries `lk.interruption.source`, set by the
caller that knows the cause (`audio_activity` for a VAD/STT barge-in and the
realtime server's own speech detection, `user_turn` for a committed turn or a
final transcript ending a pause, `programmatic` for session.interrupt(),
tools and teardown); the first interruption's cause stands. The pipeline and
`say` paths stamp `lk.playout.position`, the seconds actually played when the
user cut in. The say and realtime paths now also stamp `lk.interrupted`,
which only the pipeline reply did before.

Agent handoff: `updateAgent()` opens an `update_agent` span under
`agent_session` with `lk.previous_agent_label` / `lk.agent_label`; the old
agent's `drain_agent_activity` (with `on_exit`) and the new agent's
`start_agent_activity` / `resume_agent_activity` nest under it. The initial
start stays under `session_start`.

Fallback adapters: LLM, TTS and STT `model` / `provider` follow the instance
that serves next (first available, else the primary), so `llm_node`,
`tts_node` and `start_agent_activity` name a real model. The LLM and TTS
attempt span carries `lk.fallback.label` / `lk.fallback.index` and the
serving instance's request model and provider; the adapter's request span
and the caller's node span get `gen_ai.response.model` / provider of the
instance that answered, per request. The STT adapter's last-served tracking
is replaced by the same next-instance rule.

Text input: the keyterm-detection LLM pass runs in its own
`keyterm_detection` span under the `agent_turn` whose reply added the user
message, else under `agent_session`, with counts only
(`lk.keyterms.count/added/removed`) plus model and provider.

Adaptations: agents carry `id` where Python has `label`; JS `SpeechHandle`
and `AgentActivity.interrupt` take the source as a second positional
parameter / option instead of a keyword. The LLM and TTS stream base classes
expose their request span to subclasses (`llmRequestSpan`,
`ttsRequestSpan`), the JS counterpart of Python's `_llm_request_span`.

A cancelled preemptive attempt names why it was dropped. SpeechHandle._cancel
takes the cause like interrupt does: an attempt superseded by more of the
user's turn (a later preemptive trigger, or the transcript changing at
commit) is user_turn, one dropped by a barge-in audio_activity through
interrupt(); a cancel with no cause (teardown, a pause) still reads as
programmatic. A cloud export showed such attempts, cancelled when the user
kept talking, labelled programmatic.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ack attribution; port livekit/agents#7373 and #7374

- user_turn names an STT that leaves the base getters at `unknown` by its
  label (prefix as provider) and normalizes the provider to the GenAI
  registry spelling
- the pipeline reply task stamps its interruption verdict once it is over,
  so an interruption that lands while the tools run is recorded
- an LLM fallback stream's usage metrics name the instance that served, not
  the one the adapter would pick next after a recovery
- #7373: the fallback adapter's request span carries no operation name; the
  nested provider request is the `chat`
- #7374: llm_node records its configured model and provider when the nested
  request is created, so a failover's serving provider is kept rather than
  overwritten when the node completes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ed audio; live STT identity on user_turn

- a TTS fallback stream's usage metrics name the instance that served, not
  the one the adapter would pick next once a failed instance recovered
- chunked synthesis that fails after audio reached the caller still names
  the instance that produced it on the request and caller spans, as the
  streaming path does
- user_turn reads the STT identity when it stamps the turn, at its start
  and again at its end, so a fallback that failed over names the instance
  that transcribed rather than the one snapshotted at activity start

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rving it

model/provider name the instance serving the open stream (or the one that
answered the last recognize()), and only fall back to the next-in-line
instance between streams. A recovery probe finding the primary back no
longer relabels a turn the fallback is transcribing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- only the electing stream clears it, so a stream ending cannot erase a
  newer stream's attribution; the one-slot limit is documented
- recognize() no longer pins the slot: between requests the getters follow
  availability again, as a completed recognize() has no turn in flight

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@davidzhao
davidzhao force-pushed the dz/telemetry-coverage branch from 554dece to b9709d2 Compare September 26, 2026 05:54
@davidzhao
davidzhao merged commit 7e23bc5 into main Sep 26, 2026
12 of 14 checks passed
@davidzhao
davidzhao deleted the dz/telemetry-coverage branch September 26, 2026 06:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review_effort:medium Focused changes that need domain context or careful validation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants