Context
Spun out of the adversarial review of PR LeanerCloud/cloud-commitments-cli#1231 (SEC-04, hardened HTTP client for JWKS warmup and fetch). The PR keeps Warmup non-fatal on failure with the rationale "Google's CDN sometimes hiccups and we prefer the validator come up with a stale-on-error fetch on the first real request rather than crashloop the container." Filed for explicit team discussion because the global no-silent-fallbacks rule (feedback_no_silent_fallbacks.md) prefers fail-loud at the earliest checkpoint where the consumer can react.
Current behavior
internal/server/scheduledauth/validator.go (PR LeanerCloud/cloud-commitments-cli#1231):
resp, err := v.httpClient.Do(req)
if err != nil {
log.Printf("scheduledauth: WARN — JWKS warmup fetch failed for %s: %v "+
"(validator will retry on first request)", v.jwksURL, err)
return
}
A misconfigured SCHEDULED_TASK_OIDC_JWKS_URL (typo, blocked egress, expired DNS, scheme mismatch) starts the validator cleanly. Every request to /api/scheduled/* then fails at JWT verification time, returning 401 Unauthorized to the legitimate Cloud Scheduler call. The misconfiguration only surfaces in serving logs, not in the deploy pipeline.
Decision needed
Pick one explicitly and document it in the validator comment:
- A: keep current behavior (warmup never fatal), make sure the WARN line gets a metric/alert so a CDN-vs-misconfig hiccup is at least observable from the dashboard.
- B: turn warmup into a fail-closed startup probe and pair with a small retry budget (e.g. 3 attempts, exponential backoff, total cap of 60s) before exiting non-zero. The retry budget absorbs CDN hiccups; persistent failure stops the deploy.
Option B aligns with the no-silent-fallbacks rule and the global "fail-closed at the boundary" pattern; option A is the choice the PR has documented. Whichever the team picks, lock the comment to that decision rather than leaving the behavior implicit.
Closing-issue trailer for the fix PR
Closes #<this issue number>
Context
Spun out of the adversarial review of PR LeanerCloud/cloud-commitments-cli#1231 (SEC-04, hardened HTTP client for JWKS warmup and fetch). The PR keeps
Warmupnon-fatal on failure with the rationale "Google's CDN sometimes hiccups and we prefer the validator come up with a stale-on-error fetch on the first real request rather than crashloop the container." Filed for explicit team discussion because the global no-silent-fallbacks rule (feedback_no_silent_fallbacks.md) prefers fail-loud at the earliest checkpoint where the consumer can react.Current behavior
internal/server/scheduledauth/validator.go(PR LeanerCloud/cloud-commitments-cli#1231):A misconfigured
SCHEDULED_TASK_OIDC_JWKS_URL(typo, blocked egress, expired DNS, scheme mismatch) starts the validator cleanly. Every request to/api/scheduled/*then fails at JWT verification time, returning401 Unauthorizedto the legitimate Cloud Scheduler call. The misconfiguration only surfaces in serving logs, not in the deploy pipeline.Decision needed
Pick one explicitly and document it in the validator comment:
Option B aligns with the no-silent-fallbacks rule and the global "fail-closed at the boundary" pattern; option A is the choice the PR has documented. Whichever the team picks, lock the comment to that decision rather than leaving the behavior implicit.
Closing-issue trailer for the fix PR