You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
QA reproduced on PR #373's synchronous approve path:
Failed to approve: execution 3419bb8b-9785-4e8c-ae05-2193202d58af could not be approved:
some purchases failed: [t4g.nano: purchase failed: failed to find offering:
failed to describe offerings: operation error EC2: DescribeReservedInstancesOfferings,
https response error StatusCode: 0, RequestID: , canceled, context deadline exceeded]
StatusCode: 0, RequestID: "" = the HTTP request was cancelled at the network layer before AWS returned. Lambda function timeout = 60s.
Compounding symptom: execution status is left stuck at approved (not failed) — already tracked as #632 ("approve strands execution in 'approved' on interrupted sync execution"). Fix landing on this issue would unblock the time budget but #632 covers the state-machine repair separately.
Root cause
The AWS SDK v2 EC2 client has default behavior: 3 retries with exponential backoff, each up to a 30s request timeout. A single transient slow/failed DescribeReservedInstancesOfferings call can consume 30s × 3 + backoff = ~90-120s of wall clock. The synchronous approve handler is bounded by the Lambda invocation budget (60s on dev, see terraform/environments/aws/github-dev.tfvars:22 lambda_timeout = 60). Result: the SDK is still on its second retry when Lambda's deadline fires, the parent context cancels, the SDK reports "canceled, context deadline exceeded" on the in-flight HTTP call, and the purchase manager aggregates that as some purchases failed: [<rec>: ...].
This is NOT a pagination issue — buildOfferingFilters correctly narrows by instance-type + product-description + tenancy + scope + duration + offering-class (6 filters), so the result set is typically a single page.
Fix options (any one or combination)
Increase Lambda timeout for purchase endpoints — e.g. 300s. Cheapest fix; trades cost (longer invocations) for resilience. Update terraform/environments/aws/github-*.tfvarslambda_timeout. Note: API-Gateway-fronted paths have a 30s ceiling that we'd hit instead; Function URL has no such cap. Verify the deployment target.
Tighten SDK retry config on purchase-path clients — pass aws.Retryer with aws.WithMaxAttempts(2) and per-request http.Client{Timeout: 10*time.Second} so a transient slow API call fails fast (10s × 2 = 20s) and surfaces a retriable error to the user instead of stranding the execution. Concrete location: providers/aws/services/ec2/client.go NewClient (and the parity files for other services).
Ship (1) + (2) together as the immediate fix: bump dev/staging/prod lambda_timeout to 300s for the api function AND set EC2/RDS/ElastiCache/MemoryDB/OpenSearch/Redshift/SavingsPlans client retry config to MaxAttempts=2 + per-request 10s HTTP timeout. (3) is a separate architectural cleanup.
Acceptance criteria
lambda_timeout raised for the api Lambda in all envs (or per-route override if possible)
Regression test in providers/aws/services/ec2/client_test.go asserting that a slow-API mock fails within ≤30s wall clock instead of running to the SDK's default 90s+
No regression in the rec-collection flow (which can be slow legitimately and should keep its current retry budget — confirm by reading providers/aws/recommendations/)
Symptom
QA reproduced on PR #373's synchronous approve path:
StatusCode: 0, RequestID: ""= the HTTP request was cancelled at the network layer before AWS returned. Lambda function timeout = 60s.Compounding symptom: execution status is left stuck at
approved(notfailed) — already tracked as #632 ("approve strands execution in 'approved' on interrupted sync execution"). Fix landing on this issue would unblock the time budget but #632 covers the state-machine repair separately.Root cause
The AWS SDK v2 EC2 client has default behavior: 3 retries with exponential backoff, each up to a 30s request timeout. A single transient slow/failed
DescribeReservedInstancesOfferingscall can consume 30s × 3 + backoff = ~90-120s of wall clock. The synchronous approve handler is bounded by the Lambda invocation budget (60s on dev, seeterraform/environments/aws/github-dev.tfvars:22 lambda_timeout = 60). Result: the SDK is still on its second retry when Lambda's deadline fires, the parent context cancels, the SDK reports "canceled, context deadline exceeded" on the in-flight HTTP call, and the purchase manager aggregates that assome purchases failed: [<rec>: ...].This is NOT a pagination issue —
buildOfferingFilterscorrectly narrows byinstance-type+ product-description + tenancy + scope + duration + offering-class (6 filters), so the result set is typically a single page.Fix options (any one or combination)
terraform/environments/aws/github-*.tfvarslambda_timeout. Note: API-Gateway-fronted paths have a 30s ceiling that we'd hit instead; Function URL has no such cap. Verify the deployment target.aws.Retryerwithaws.WithMaxAttempts(2)and per-requesthttp.Client{Timeout: 10*time.Second}so a transient slow API call fails fast (10s × 2 = 20s) and surfaces a retriable error to the user instead of stranding the execution. Concrete location:providers/aws/services/ec2/client.goNewClient (and the parity files for other services).Recommendation
Ship (1) + (2) together as the immediate fix: bump dev/staging/prod
lambda_timeoutto 300s for the api function AND set EC2/RDS/ElastiCache/MemoryDB/OpenSearch/Redshift/SavingsPlans client retry config toMaxAttempts=2+ per-request 10s HTTP timeout. (3) is a separate architectural cleanup.Acceptance criteria
lambda_timeoutraised for theapiLambda in all envs (or per-route override if possible)aws.WithRetryMaxAttempts(2)+ per-request HTTP timeout ≤ 15sproviders/aws/services/ec2/client_test.goasserting that a slow-API mock fails within ≤30s wall clock instead of running to the SDK's default 90s+providers/aws/recommendations/)Cross-references
detailsfield — every dashboard purchase POSTs without per-rec details, silently defaults to Linux/Region/default #597 / PR fix(frontend/api): preserve details + engine on purchase POST (closes #597) #600 (frontend details preserve)