feat(telemetry): complete the GenAI conventions for LLM-observability backends - #7396
SanjuMLGeek wants to merge 1 commit into
Conversation
ef602fc to
143a459
Compare
143a459 to
ea4d186
Compare
ea4d186 to
7f3f0cb
Compare
… backends Builds on the GenAI semantic conventions, from driving a voice session's trace into Datadog AI/LLM Observability and fixing what did not render or double counted. `track_inference_span` already keeps `llm_node` and `llm_request` from both reporting the operation, and `gen_ai.conversation.id` is already set on the GenAI spans; this closes the gaps around them. Every change is attribute-only or opt-in, so nothing alters today's default export except where a span was previously empty or counted twice. A delegating stream still double counts its tokens. livekit#7373 stopped `llm_fallback_adapter` claiming `gen_ai.operation.name`, so a backend no longer files it as an LLM call, but the response side is ungated: `gen_ai.usage.*` still lands on both the adapter span and the `llm_request` beneath it, and a backend summing tokens across spans counts every call twice. The adapter now leaves the counts to the provider call, keyed off the same `_genai_operation_name` livekit#7373 introduced rather than a second flag. The served model rather than the adapter's name. `.model` and `.provider` are a component's stable identity, so through a `FallbackAdapter` every span reported `model="FallbackAdapter"` and per-model cost and latency could not be broken out; `metrics_metadata` names the instance that actually ran. The adapters' own `.model` and `.provider` are untouched. `gen_ai.conversation.id` on every span, not only the GenAI ones. A backend that groups a session by the attribute drops any span without it, so a tool or lifecycle span falls out of the session view. The id is stamped centrally on span creation, `detached_span` included. Content where spans were blank. `agent_turn` carried no `gen_ai.input.messages` or `output.messages`, and the STT and TTS turns carried no content and no `gen_ai.operation.name`, so a chat renderer showed them empty and a backend filed them as generic spans. Since a turn span only ever holds its own messages, the conversation is recorded once end to end on `agent_session` at close, instructions excluded — they are the largest payload in a session and already on the inference spans. Both turns now report their provider through the same helper as every other span, so a provider reaches the registry spelling rather than being written raw. Bounded content per inference span. Every inference span repeats the whole prompt and the whole history, so exported bytes grow with the square of the turn count: one 73-call session measured 2.3 MB of span content, 74% of it the same 28-37 KB prompt copied onto 104 spans. Oversized batches time out, are retried and then dropped, which is how a parent span goes missing. Two opt-in variables in the convention's own namespace, `OTEL_INSTRUMENTATION_GENAI_CAPTURE_SYSTEM_INSTRUCTIONS` and `OTEL_INSTRUMENTATION_GENAI_MAX_INPUT_MESSAGES`, both defaulting to today's untruncated behaviour, with a span that drops messages reporting how many in `lk.gen_ai.input.messages_dropped`. Tool results a backend can read. The convention's `tool_call_response` part carries only `id` and `response`, so a backend resolving a result's label from the part alone rendered every result as "unknown" — the part's schema is open, the request part already requires `name`, and `FunctionCallOutput.name` is populated at every construction site, so it is reported when set and omitted when not. Tests cover the parent/child shape a backend receives, not just the attributes — the span tree is what actually broke.
7f3f0cb to
af1604c
Compare
There was a problem hiding this comment.
Devin Review found 1 new potential issue.
3 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
There was a problem hiding this comment.
🟡 Realtime fallback spans lose served model
With RealtimeModelFallbackAdapter, realtime spans report the adapter's model and provider instead of the active provider. Its metrics_metadata exposes the serving instance, while these properties stay RealtimeModelFallbackAdapter and livekit. Backends therefore attribute every fallback response to the adapter.
(Refers to this code)
Learn more
A realtime fallback adapter delegates each live session to one underlying realtime model. The adapter keeps that serving model in _active_instance, and metrics_metadata exposes its model and provider. By contrast, the adapter's model and provider are constant adapter labels. Reading those constants records neither the primary model nor a model selected after failover.
Example: An adapter backed by OpenAI and Google starts on OpenAI. Its realtime inference span records model RealtimeModelFallbackAdapter and provider livekit, rather than OpenAI's model and provider. After a swap to Google, the span records the same adapter labels again.
Recommended fix: Read the model and provider from self.llm.metrics_metadata when creating the realtime inference span, or provide an equivalent API that exposes the adapter's active serving instance.
Was this helpful? React with 👍 or 👎 to provide feedback.
Summary
Builds on the GenAI semantic conventions, from driving a voice session's trace into an LLM-observability backend (Datadog AI/LLM Observability) and fixing what did not render or double counted.
track_inference_spanalready keepsllm_nodeandllm_requestfrom both reporting the operation, andgen_ai.conversation.idis already set on the GenAI spans; this closes the gaps around them.Every change is attribute-only or opt-in. Nothing alters today's default export except where a span was previously empty or counted twice.
Supersedes #7104, which was opened against an older base. Rebased onto
main, including #7373 — the fallback-adapter work there replaces what this PR originally carried, leaving only the token half — and #7374, which now carries the node span's request identity, so the model/provider attribution this PR originally proposed is dropped in favour of it.What it fixes
A delegating stream still double counts its tokens. #7373 stopped
llm_fallback_adapterclaiminggen_ai.operation.name, so a backend no longer files it as an LLM call — but the response side is ungated:gen_ai.usage.*still lands on both the adapter span and thellm_requestbeneath it, and a backend summing tokens across spans counts every call twice. The adapter now leaves the counts to the provider call, keyed off the same_genai_operation_name#7373 introduced rather than a second flag. This is a three-line change on top of that fix.gen_ai.conversation.idon every span, not only the GenAI ones. A backend that groups a session by the attribute drops any span without it, so a tool or lifecycle span falls out of the session view. The id is stamped centrally on span creation,detached_spanincluded.Content where spans were blank.
agent_turncarried nogen_ai.input.messagesoroutput.messages, and the STT and TTS turns carried no content and nogen_ai.operation.name, so a chat renderer showed them empty and a backend filed them as generic spans. Since a turn span only ever holds its own messages, the conversation is recorded once end to end onagent_sessionat close, instructions excluded. Both turns now report their provider throughset_request_attributeslike every other span, so a value such asapi.openai.comreaches the registry spellingopenaiinstead of being written raw — previously an STT span and the LLM span beneath it could name the same provider two different ways.Bounded content per inference span. Every inference span repeats the whole prompt and the whole history, so exported bytes grow with the square of the turn count: one 73-call session measured 2.3 MB of span content, 74% of it the same 28–37 KB prompt copied onto 104 spans. Oversized batches time out, are retried and then dropped, which is how a parent span goes missing. Two opt-in variables in the convention's own namespace,
OTEL_INSTRUMENTATION_GENAI_CAPTURE_SYSTEM_INSTRUCTIONSandOTEL_INSTRUMENTATION_GENAI_MAX_INPUT_MESSAGES, both defaulting to today's untruncated behaviour, with a span that drops messages reporting how many inlk.gen_ai.input.messages_dropped.Tool results a backend can read. The convention's
tool_call_responsepart carries onlyidandresponse, so a backend resolving a result's label from the part alone rendered every result as "unknown" — the part's schema is open, the request part already requiresname, andFunctionCallOutput.nameis populated at every construction site, so it is reported when set and omitted when not.Points I'd like maintainer input on
GenAIOperationName.TRANSCRIBE/SYNTHESIZEare not in the OTel registry, which has no speech operations yet. The attribute is an open enum, and without a value the STT and TTS steps of a voice turn are not GenAI operations to a backend and drop out of the session tree. Happy to move these to a separate class, or drop them, if you'd rather not imply they are registry values.generation.pycontains the one non-telemetry runtime change: the TTS input tee's first branch used to stop after one chunk and now drains, so the synthesized text can be reported as the span's input. This changes tee buffering slightly.Behaviour notes
providerstring was previously written verbatim on the STT/TTS spans and is now dropped, sincegen_ai_provider_namestrips and returnsNone. Falsymodel/providerare skipped exactly as before.Testing
Run against
af1604ce(rebased ontomain, including #7373, #7374 and #7376):ruff format --checkandruff check— passmypyviascripts/check_types.py— 682 source files; the one error is a pre-existing missingboto3stub in the AWS plugin, untouched hereuv run pytest --unit --audio_eot— 3615 passed, 6 failed, 9 errors. Every failure is intest_room.py,test_tokenizer.py,test_chat_ctx.pyortest_audio_decoder.py— none in a file this PR touches. They need a live server or provider credentials (TimeoutErrorwaiting for participants,APIErrorintest_summarize);test_audio_decoder.pypasses 16/16 in isolation but takes 5m21s and times out under full-suite load.