Symptom
Third recurrence of #667. QA reproduced again:
Failed to approve: execution 9336a111-85c5-4b74-9991-367350de9bac could not be approved:
some purchases failed: [t4g.nano: purchase failed: failed to find offering:
failed to describe offerings: operation error EC2: DescribeReservedInstancesOfferings,
https response error StatusCode: 0, RequestID: , canceled, context deadline exceeded]
Same shape as the prior #667 reproductions. StatusCode: 0, RequestID: "" = the AWS SDK request was cancelled at the network layer; the 60s Lambda budget got blown.
Evidence pulled from prod
What we still don't have
The smoking gun for #667's root cause (Lambda budget vs SDK retry storm) is the elapsed-time on the failed PurchaseCommitment call. Without #668's logging firing on a fresh failure, we can't confirm the timing hypothesis from CloudWatch — but the wall-clock fits: SDK default retry config is 3 attempts × 30s per call = up to ~90s, overrunning the 60s Lambda budget.
What to ship now
Implement the 2-part fix from #667 (still in PROPOSED state there):
- Bump
lambda_timeout 60 -> 300s in terraform/environments/aws/github-{dev,staging,prod}.tfvars for the api Lambda. Trades a 5x cost ceiling for resilience; Lambda Function URL has no upstream ceiling that would also need raising.
- Tighten AWS SDK retry config on every purchase-path client (EC2, RDS, ElastiCache, MemoryDB, OpenSearch, Redshift, SavingsPlans) — pass
aws.WithRetryMaxAttempts(2) + per-request http.Client{Timeout: 15 * time.Second} so transient slow API calls fail fast (15s × 2 = 30s) and surface a retriable error to the user instead of stranding the execution in approved.
After both ship + redeploy, re-test against the same SKU (t4g.nano in us-east-1) and confirm the next-cycle CloudWatch trace shows purchase[<id>]: PurchaseCommitment succeeded in <ms> OR failed after <ms> (the latter now bounded by 30s).
Acceptance criteria
Cross-references
Severity
P0 / critical. Third recurrence. Every Approve click on a real t4g.nano (or similar latency-sensitive AWS API path) is at risk; the only protection today is #681's 10-min reaper, which cleans up but doesn't prevent.
Symptom
Third recurrence of #667. QA reproduced again:
Same shape as the prior #667 reproductions.
StatusCode: 0, RequestID: ""= the AWS SDK request was cancelled at the network layer; the 60s Lambda budget got blown.Evidence pulled from prod
plan_id: "",cloud_account_id: None, rec has nodetails/engine/cloud_account_id(i.e. a stripped-by-old-frontend payload, NOT a fresh post-fix(frontend/api): preserve details + engine on purchase POST (closes #597) #600 rec)approved(the bug(purchases): approve strands execution in 'approved' (purchase never runs, no error) on interrupted sync execution #632 stranding pattern that fix(purchase): reap purchase_executions stuck in approved/running >10m (closes #678) #681's reaper handles after 10 min)purchase[9336a111-...]:log lines anywhere in CloudWatch. PR fix(purchases): add execution-tagged logging to executeSinglePurchase (refs #667) #668's exec-tagged logging merged at 12:56 UTC but the Lambda was last deployed at 16:23 UTC; this execution evidently fired BEFORE the deploy, so the logging didn't capture it.What we still don't have
The smoking gun for #667's root cause (Lambda budget vs SDK retry storm) is the elapsed-time on the failed
PurchaseCommitmentcall. Without #668's logging firing on a fresh failure, we can't confirm the timing hypothesis from CloudWatch — but the wall-clock fits: SDK default retry config is 3 attempts × 30s per call = up to ~90s, overrunning the 60s Lambda budget.What to ship now
Implement the 2-part fix from #667 (still in PROPOSED state there):
lambda_timeout60 -> 300s interraform/environments/aws/github-{dev,staging,prod}.tfvarsfor the api Lambda. Trades a 5x cost ceiling for resilience; Lambda Function URL has no upstream ceiling that would also need raising.aws.WithRetryMaxAttempts(2)+ per-requesthttp.Client{Timeout: 15 * time.Second}so transient slow API calls fail fast (15s × 2 = 30s) and surface a retriable error to the user instead of stranding the execution inapproved.After both ship + redeploy, re-test against the same SKU (t4g.nano in us-east-1) and confirm the next-cycle CloudWatch trace shows
purchase[<id>]: PurchaseCommitment succeeded in <ms>ORfailed after <ms>(the latter now bounded by 30s).Acceptance criteria
lambda_timeout= 300 in all 3 env tfvarsaws.WithRetryMaxAttempts(2)+ per-request HTTP timeout ≤ 15sproviders/aws/recommendations/)purchase[<id>]: ... failed after <elapsed>log line in CloudWatchCross-references
details)Severity
P0 / critical. Third recurrence. Every Approve click on a real t4g.nano (or similar latency-sensitive AWS API path) is at risk; the only protection today is #681's 10-min reaper, which cleans up but doesn't prevent.