Skip to content

fix: avoid duplicate token usage on llm fallback spans - #7532

Draft
chenghao-mou wants to merge 1 commit into
mainfrom
chenghao/fix/AGT-3602-fallback-token-usage
Draft

chenghao-mou wants to merge 1 commit into
mainfrom
chenghao/fix/AGT-3602-fallback-token-usage

Conversation

@chenghao-mou

Copy link
Copy Markdown
Member

Fallback wrappers repeat token usage from the serving provider request. Record GenAI usage only on the provider span, while preserving LLM metrics.

Adopts @SanjuMLGeek’s usage guard from #7396.

Addresses AGT-3602

Request summary and agent context

Model: GPT-6-Astra

Split the proposed telemetry improvements into focused PRs that are easy to review and revert. Use general problem summaries and omit customer details.

Co-authored-by: Sanjay Kukadiya <43816307+SanjuMLGeek@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant