Retain provider timing and reasoning usage for timeout diagnosis - #118
Merged
Merged
Conversation
|
Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Provider timeouts currently collapse into a generic 504: operators cannot distinguish waiting for response headers from an incomplete response body, and successful OpenAI-compatible calls discard upstream generation IDs, provider names and reasoning-token usage. Preserve bounded HTTP transport timings and available upstream metadata per attempt in the response trace and the existing call ledger. Expose reported reasoning counts in
usage.completion_tokens_details.This supports investigation of intermittent DeepSeek timeouts. The same saved request has both timed out at 40.3 seconds and completed through the public router in 18.0 seconds buffered / 18.5 seconds streamed; these observations do not establish the cause. This PR adds evidence, not a latency fix. Provider request bodies, routing, retries, output limits and deadlines stay unchanged. No production deployment has been performed.
Validation: 60 distinct focused tests pass locally (one existing PostgreSQL-dependent skip). New tests use real TCP responses to distinguish stalled headers from a trickling body; verify HTTP success/error and ledger propagation, concurrent-request isolation, bounded metadata, absent/zero reasoning usage, and exclusion of headers/prompts/reasoning text. Rebuilding the saved failed request before and after the change produces the identical outbound URL/body hash. Full CI passes on head b863b34: Python suite, policy core assertions, router/AntSeed image build and boot.
See
docs/provider-diagnostics.mdfor field semantics and limitations, including missing metadata on incomplete buffered JSON and outer cancellation before an adapter result is returned.Summary by CodeRabbit