Skip to content

feat(telemetry): complete the GenAI conventions for LLM-observability backends - #7396

Open
SanjuMLGeek wants to merge 1 commit into
livekit:mainfrom
SanjuMLGeek:skukadiya/genai-otel-main
Open

SanjuMLGeek wants to merge 1 commit into
livekit:mainfrom
SanjuMLGeek:skukadiya/genai-otel-main

Conversation

@SanjuMLGeek

@SanjuMLGeek SanjuMLGeek commented Sep 22, 2026 •

Copy link
Copy Markdown

Summary

Builds on the GenAI semantic conventions, from driving a voice session's trace into an LLM-observability backend (Datadog AI/LLM Observability) and fixing what did not render or double counted. track_inference_span already keeps llm_node and llm_request from both reporting the operation, and gen_ai.conversation.id is already set on the GenAI spans; this closes the gaps around them.

Every change is attribute-only or opt-in. Nothing alters today's default export except where a span was previously empty or counted twice.

Supersedes #7104, which was opened against an older base. Rebased onto main, including #7373 — the fallback-adapter work there replaces what this PR originally carried, leaving only the token half — and #7374, which now carries the node span's request identity, so the model/provider attribution this PR originally proposed is dropped in favour of it.

What it fixes

A delegating stream still double counts its tokens. #7373 stopped llm_fallback_adapter claiming gen_ai.operation.name, so a backend no longer files it as an LLM call — but the response side is ungated: gen_ai.usage.* still lands on both the adapter span and the llm_request beneath it, and a backend summing tokens across spans counts every call twice. The adapter now leaves the counts to the provider call, keyed off the same _genai_operation_name #7373 introduced rather than a second flag. This is a three-line change on top of that fix.

gen_ai.conversation.id on every span, not only the GenAI ones. A backend that groups a session by the attribute drops any span without it, so a tool or lifecycle span falls out of the session view. The id is stamped centrally on span creation, detached_span included.

Content where spans were blank. agent_turn carried no gen_ai.input.messages or output.messages, and the STT and TTS turns carried no content and no gen_ai.operation.name, so a chat renderer showed them empty and a backend filed them as generic spans. Since a turn span only ever holds its own messages, the conversation is recorded once end to end on agent_session at close, instructions excluded. Both turns now report their provider through set_request_attributes like every other span, so a value such as api.openai.com reaches the registry spelling openai instead of being written raw — previously an STT span and the LLM span beneath it could name the same provider two different ways.

Bounded content per inference span. Every inference span repeats the whole prompt and the whole history, so exported bytes grow with the square of the turn count: one 73-call session measured 2.3 MB of span content, 74% of it the same 28–37 KB prompt copied onto 104 spans. Oversized batches time out, are retried and then dropped, which is how a parent span goes missing. Two opt-in variables in the convention's own namespace, OTEL_INSTRUMENTATION_GENAI_CAPTURE_SYSTEM_INSTRUCTIONS and OTEL_INSTRUMENTATION_GENAI_MAX_INPUT_MESSAGES, both defaulting to today's untruncated behaviour, with a span that drops messages reporting how many in lk.gen_ai.input.messages_dropped.

Tool results a backend can read. The convention's tool_call_response part carries only id and response, so a backend resolving a result's label from the part alone rendered every result as "unknown" — the part's schema is open, the request part already requires name, and FunctionCallOutput.name is populated at every construction site, so it is reported when set and omitted when not.

Points I'd like maintainer input on

  • GenAIOperationName.TRANSCRIBE / SYNTHESIZE are not in the OTel registry, which has no speech operations yet. The attribute is an open enum, and without a value the STT and TTS steps of a voice turn are not GenAI operations to a backend and drop out of the session tree. Happy to move these to a separate class, or drop them, if you'd rather not imply they are registry values.
  • generation.py contains the one non-telemetry runtime change: the TTS input tee's first branch used to stop after one chunk and now drains, so the synthesized text can be reported as the span's input. This changes tee buffering slightly.
  • Opt-in truncation is the largest single piece here and is independent of the rest — happy to split it into its own PR if you'd prefer the tree fixes on their own.

Behaviour notes

  • A whitespace-only provider string was previously written verbatim on the STT/TTS spans and is now dropped, since gen_ai_provider_name strips and returns None. Falsy model/provider are skipped exactly as before.

Testing

Run against af1604ce (rebased onto main, including #7373, #7374 and #7376):

  • ruff format --check and ruff check — pass
  • mypy via scripts/check_types.py — 682 source files; the one error is a pre-existing missing boto3 stub in the AWS plugin, untouched here
  • uv run pytest --unit --audio_eot — 3615 passed, 6 failed, 9 errors. Every failure is in test_room.py, test_tokenizer.py, test_chat_ctx.py or test_audio_decoder.py — none in a file this PR touches. They need a live server or provider credentials (TimeoutError waiting for participants, APIError in test_summarize); test_audio_decoder.py passes 16/16 in isolation but takes 5m21s and times out under full-suite load.
  • New tests cover the parent/child shape a backend receives, not just the attributes — the span tree is what actually broke.

@CLAassistant

CLAassistant commented Sep 22, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@SanjuMLGeek
SanjuMLGeek force-pushed the skukadiya/genai-otel-main branch 5 times, most recently from ef602fc to 143a459 Compare September 22, 2026 17:42
@SanjuMLGeek
SanjuMLGeek marked this pull request as ready for review September 22, 2026 17:53
@SanjuMLGeek
SanjuMLGeek requested a review from a team as a code owner September 22, 2026 17:53
devin-ai-integration[bot]

This comment was marked as resolved.

@SanjuMLGeek
SanjuMLGeek force-pushed the skukadiya/genai-otel-main branch from 143a459 to ea4d186 Compare September 23, 2026 16:12
devin-ai-integration[bot]

This comment was marked as resolved.

@SanjuMLGeek
SanjuMLGeek force-pushed the skukadiya/genai-otel-main branch from ea4d186 to 7f3f0cb Compare September 23, 2026 16:52
… backends

Builds on the GenAI semantic conventions, from driving a voice session's trace
into Datadog AI/LLM Observability and fixing what did not render or double
counted. `track_inference_span` already keeps `llm_node` and `llm_request` from
both reporting the operation, and `gen_ai.conversation.id` is already set on the
GenAI spans; this closes the gaps around them. Every change is attribute-only or
opt-in, so nothing alters today's default export except where a span was
previously empty or counted twice.

A delegating stream still double counts its tokens. livekit#7373 stopped
`llm_fallback_adapter` claiming `gen_ai.operation.name`, so a backend no longer
files it as an LLM call, but the response side is ungated: `gen_ai.usage.*`
still lands on both the adapter span and the `llm_request` beneath it, and a
backend summing tokens across spans counts every call twice. The adapter now
leaves the counts to the provider call, keyed off the same
`_genai_operation_name` livekit#7373 introduced rather than a second flag.

The served model rather than the adapter's name. `.model` and `.provider` are a
component's stable identity, so through a `FallbackAdapter` every span reported
`model="FallbackAdapter"` and per-model cost and latency could not be broken
out; `metrics_metadata` names the instance that actually ran. The adapters' own
`.model` and `.provider` are untouched.

`gen_ai.conversation.id` on every span, not only the GenAI ones. A backend that
groups a session by the attribute drops any span without it, so a tool or
lifecycle span falls out of the session view. The id is stamped centrally on
span creation, `detached_span` included.

Content where spans were blank. `agent_turn` carried no `gen_ai.input.messages`
or `output.messages`, and the STT and TTS turns carried no content and no
`gen_ai.operation.name`, so a chat renderer showed them empty and a backend
filed them as generic spans. Since a turn span only ever holds its own messages,
the conversation is recorded once end to end on `agent_session` at close,
instructions excluded — they are the largest payload in a session and already on
the inference spans. Both turns now report their provider through the same
helper as every other span, so a provider reaches the registry spelling rather
than being written raw.

Bounded content per inference span. Every inference span repeats the whole
prompt and the whole history, so exported bytes grow with the square of the turn
count: one 73-call session measured 2.3 MB of span content, 74% of it the same
28-37 KB prompt copied onto 104 spans. Oversized batches time out, are retried
and then dropped, which is how a parent span goes missing. Two opt-in variables
in the convention's own namespace,
`OTEL_INSTRUMENTATION_GENAI_CAPTURE_SYSTEM_INSTRUCTIONS` and
`OTEL_INSTRUMENTATION_GENAI_MAX_INPUT_MESSAGES`, both defaulting to today's
untruncated behaviour, with a span that drops messages reporting how many in
`lk.gen_ai.input.messages_dropped`.

Tool results a backend can read. The convention's `tool_call_response` part
carries only `id` and `response`, so a backend resolving a result's label from
the part alone rendered every result as "unknown" — the part's schema is open,
the request part already requires `name`, and `FunctionCallOutput.name` is
populated at every construction site, so it is reported when set and omitted
when not.

Tests cover the parent/child shape a backend receives, not just the attributes —
the span tree is what actually broke.
@SanjuMLGeek
SanjuMLGeek force-pushed the skukadiya/genai-otel-main branch from 7f3f0cb to af1604c Compare September 23, 2026 16:56

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 new potential issue.

3 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Realtime fallback spans lose served model

With RealtimeModelFallbackAdapter, realtime spans report the adapter's model and provider instead of the active provider. Its metrics_metadata exposes the serving instance, while these properties stay RealtimeModelFallbackAdapter and livekit. Backends therefore attribute every fallback response to the adapter.

(Refers to this code)

Learn more

A realtime fallback adapter delegates each live session to one underlying realtime model. The adapter keeps that serving model in _active_instance, and metrics_metadata exposes its model and provider. By contrast, the adapter's model and provider are constant adapter labels. Reading those constants records neither the primary model nor a model selected after failover.

Example: An adapter backed by OpenAI and Google starts on OpenAI. Its realtime inference span records model RealtimeModelFallbackAdapter and provider livekit, rather than OpenAI's model and provider. After a swap to Google, the span records the same adapter labels again.

Recommended fix: Read the model and provider from self.llm.metrics_metadata when creating the realtime inference span, or provide an equivalent API that exposes the adapter's active serving instance.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants