Skip to content

fix: preserve fallback provider attribution on llm node spans - #7374

Merged
chenghao-mou merged 1 commit into
mainfrom
chenghao/fix/AGT-3551-fallback-provider-attribution
Sep 23, 2026
Merged

chenghao-mou merged 1 commit into
mainfrom
chenghao/fix/AGT-3551-fallback-provider-attribution

Conversation

@chenghao-mou

Copy link
Copy Markdown
Member

After failover, llm_node could report the fallback model with the primary provider. Preserve the serving-provider update from #7137 by recording the node's request identity before nested inference starts.

Addresses AGT-3551: issue.

The full local type check reports existing plugin dependency errors also present on main; the core package passes.

Initial prompt and agent context

Model: GPT-6

can you create draft PR for the fallback issue?

The requested issue is the llm_node provider overwrite found during the preceding telemetry review. The fallback stream records the serving provider, then node completion overwrites it with the original provider. This PR preserves the per-request attribution.

Record the node's request identity when its first nested LLM stream is
created. The serving-provider update can then survive completion instead
of being overwritten by the provider captured before failover.
@chenghao-mou
chenghao-mou marked this pull request as ready for review September 21, 2026 10:15
@chenghao-mou
chenghao-mou requested a review from a team as a code owner September 21, 2026 10:15

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Devin Review

@longcw longcw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from claude code:

What it misses. On the request where the failover happens, FallbackAdapter.model still returns the primary: _next_instance() marks the primary unavailable only after it fails. FallbackLLMStream._run then writes gen_ai.provider.name and gen_ai.response.model to _caller_span, but not gen_ai.request.model. The node span ends with the primary's request model beside the fallback's provider — the same mismatched pair, reversed. The

@chenghao-mou

chenghao-mou commented Sep 22, 2026 •

Copy link
Copy Markdown
Member Author

from claude code:

What it misses. On the request where the failover happens, FallbackAdapter.model still returns the primary: _next_instance() marks the primary unavailable only after it fails. FallbackLLMStream._run then writes gen_ai.provider.name and gen_ai.response.model to _caller_span, but not gen_ai.request.model. The node span ends with the primary's request model beside the fallback's provider — the same mismatched pair, reversed. The

Yeah, this is expected from this change:
started:

  • gen_ai.request.model: model A
  • gen_ai.provider.name: provider A

after fallback:

previously:

  • gen_ai.request.model: model A
  • gen_ai.response.model: model B
  • gen_ai.provider.name: provider A

now:

  • gen_ai.request.model: model A
  • gen_ai.response.model: model B
  • gen_ai.provider.name: provider B

so the fallback (requested A but B responded) is clear.

@chenghao-mou
chenghao-mou merged commit c20d272 into main Sep 23, 2026
28 checks passed
@chenghao-mou
chenghao-mou deleted the chenghao/fix/AGT-3551-fallback-provider-attribution branch September 23, 2026 13:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants