Skip to content

service: bound concurrent rate-limit requests - #1241

Draft
NVGreg wants to merge 2 commits into
envoyproxy:mainfrom
NVGreg:feat/redis-work-admission
Draft

NVGreg wants to merge 2 commits into
envoyproxy:mainfrom
NVGreg:feat/redis-work-admission

Conversation

@NVGreg

@NVGreg NVGreg commented Sep 15, 2026 •

Copy link
Copy Markdown

Problem

Slow cache calls can retain ShouldRateLimit handlers after callers cancel. Without configured bounds, more gRPC streams can keep arriving while that work is stuck.

Change

Add optional per-instance request admission, a cache-work deadline, and a per-connection gRPC stream cap. Calls above the admission limit fail immediately with ResourceExhausted, separate from an OVER_LIMIT quota decision. Add admission metrics, preserve HTTP caller contexts, and map service errors to HTTP status codes.

Runtime contract

MAX_CONCURRENT_REQUESTS, REQUEST_TIMEOUT, and GRPC_MAX_CONCURRENT_STREAMS all default to 0 (disabled); operators must set bounds explicitly. The gRPC cap applies per connection, not per instance or to HTTP /json. A conforming client waits for stream capacity or its own deadline; it does not get an immediate application overload response from the stream cap.

Admission capacity remains occupied until synchronous work and completion-metric recording return, even after caller cancellation. The deadline is a cancellation signal, not a guarantee that Radix promptly drains a response or that Redis stops a command already sent. Shadow mode does not override service failures; rollout must account for the caller's failure policy.

Validation

Focused server, settings, and service tests pass. The gRPC tests verify HTTP/2 SETTINGS and one-connection deadline/recovery behavior. The stalled-Redis test uses the real client against a loopback RESP fixture and verifies that a cancelled call retains its admission slot until backend work returns.

Add optional admission and deadline settings for ShouldRateLimit. Reject excess calls immediately and retain capacity until synchronous processing and completion metric recording finish. Preserve HTTP caller contexts and distinguish overload from quota decisions.

Signed-off-by: Gregory Giecold <ggiecold@nvidia.com>
Signed-off-by: Gregory Giecold <ggiecold@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant