Skip to content

fix(purchases): #667 recurrent -- recommend implementing the Lambda timeout bump + SDK retry tightening NOW (P0) #683

Description

@cristim

Symptom

Third recurrence of #667. QA reproduced again:

Failed to approve: execution 9336a111-85c5-4b74-9991-367350de9bac could not be approved:
some purchases failed: [t4g.nano: purchase failed: failed to find offering:
failed to describe offerings: operation error EC2: DescribeReservedInstancesOfferings,
https response error StatusCode: 0, RequestID: , canceled, context deadline exceeded]

Same shape as the prior #667 reproductions. StatusCode: 0, RequestID: "" = the AWS SDK request was cancelled at the network layer; the 60s Lambda budget got blown.

Evidence pulled from prod

What we still don't have

The smoking gun for #667's root cause (Lambda budget vs SDK retry storm) is the elapsed-time on the failed PurchaseCommitment call. Without #668's logging firing on a fresh failure, we can't confirm the timing hypothesis from CloudWatch — but the wall-clock fits: SDK default retry config is 3 attempts × 30s per call = up to ~90s, overrunning the 60s Lambda budget.

What to ship now

Implement the 2-part fix from #667 (still in PROPOSED state there):

  1. Bump lambda_timeout 60 -> 300s in terraform/environments/aws/github-{dev,staging,prod}.tfvars for the api Lambda. Trades a 5x cost ceiling for resilience; Lambda Function URL has no upstream ceiling that would also need raising.
  2. Tighten AWS SDK retry config on every purchase-path client (EC2, RDS, ElastiCache, MemoryDB, OpenSearch, Redshift, SavingsPlans) — pass aws.WithRetryMaxAttempts(2) + per-request http.Client{Timeout: 15 * time.Second} so transient slow API calls fail fast (15s × 2 = 30s) and surface a retriable error to the user instead of stranding the execution in approved.

After both ship + redeploy, re-test against the same SKU (t4g.nano in us-east-1) and confirm the next-cycle CloudWatch trace shows purchase[<id>]: PurchaseCommitment succeeded in <ms> OR failed after <ms> (the latter now bounded by 30s).

Acceptance criteria

  • lambda_timeout = 300 in all 3 env tfvars
  • EC2/RDS/ElastiCache/MemoryDB/OpenSearch/Redshift/SavingsPlans AWS SDK clients constructed with aws.WithRetryMaxAttempts(2) + per-request HTTP timeout ≤ 15s
  • Regression test (Go): slow-API mock returns within ≤30s wall clock instead of SDK default 90s+
  • Recommendation collection path is NOT regressed (it needs the higher retry budget; verify by checking providers/aws/recommendations/)
  • Live verification: redeploy + a real t4g.nano purchase attempt either succeeds OR fails with a purchase[<id>]: ... failed after <elapsed> log line in CloudWatch

Cross-references

Severity

P0 / critical. Third recurrence. Every Approve click on a real t4g.nano (or similar latency-sensitive AWS API path) is at risk; the only protection today is #681's 10-min reaper, which cleans up but doesn't prevent.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions