Symptom
Approving a purchase repeatedly strands the execution in approved with the purchase never running. Three executions on 2026-05-20 (19:32, 22:07, 22:56, all cudly-web / cristi@leanercloud.com) are stuck:
status = approved
error = (empty)
completed_at = NULL
recommendations[].purchased = false
The AWS side confirms no commitment was created (Savings Plans / EC2-RDS-ElastiCache RIs all 0 on account 909626172446).
Root cause
Manager.ApproveAndExecute (internal/purchase/approvals.go) flips the row to approved before running the purchase synchronously inside the HTTP/Lambda request, and only finalizeExecution (internal/purchase/manager.go) sets completed/failed after executePurchase returns. If that synchronous run is interrupted (Lambda timeout, cold-start eviction, panic), the row is left persisted as approved with no error and purchased=false — the purchase never executed and there is no automatic recovery. PR #623 (#621) made these rows visible in History but did not address the stranding itself.
This is the structural problem behind #621: a financial action silently no-ops and the row requires manual DB intervention to clear.
Fix options (pick per architecture)
- Roll back on interruption: wrap the approve+execute so that if execution does not finalize, the row returns to
pending (or a dedicated approval_interrupted) rather than staying approved, so the scheduler / a retry can pick it up safely. Guard against double-execution (idempotency on commitment creation).
- Durable queue: enqueue the approved execution to SQS and execute in a worker with visibility-timeout-based retry + a dead-letter queue, instead of synchronously in the request. The
approved row then has a real owner that always finalizes it.
- Recovery sweep: a scheduled job that finds
approved rows older than N minutes with purchased=false and re-drives or fails them, so a strand can never be permanent.
Whichever path: an approved row must never be a terminal state without an owner. Add a regression test simulating an interrupted execution and asserting the row does not remain permanently approved.
Immediate cleanup (done separately)
The 3 stranded rows are being cancelled (nothing was purchased on AWS, so it is safe). This issue is the durable fix so it stops recurring.
Surfaced while diagnosing the stuck-purchase incident (#621 follow-up).
Symptom
Approving a purchase repeatedly strands the execution in
approvedwith the purchase never running. Three executions on 2026-05-20 (19:32, 22:07, 22:56, allcudly-web/ cristi@leanercloud.com) are stuck:The AWS side confirms no commitment was created (Savings Plans / EC2-RDS-ElastiCache RIs all 0 on account 909626172446).
Root cause
Manager.ApproveAndExecute(internal/purchase/approvals.go) flips the row toapprovedbefore running the purchase synchronously inside the HTTP/Lambda request, and onlyfinalizeExecution(internal/purchase/manager.go) setscompleted/failedafterexecutePurchasereturns. If that synchronous run is interrupted (Lambda timeout, cold-start eviction, panic), the row is left persisted asapprovedwith no error andpurchased=false— the purchase never executed and there is no automatic recovery. PR #623 (#621) made these rows visible in History but did not address the stranding itself.This is the structural problem behind #621: a financial action silently no-ops and the row requires manual DB intervention to clear.
Fix options (pick per architecture)
pending(or a dedicatedapproval_interrupted) rather than stayingapproved, so the scheduler / a retry can pick it up safely. Guard against double-execution (idempotency on commitment creation).approvedrow then has a real owner that always finalizes it.approvedrows older than N minutes withpurchased=falseand re-drives or fails them, so a strand can never be permanent.Whichever path: an
approvedrow must never be a terminal state without an owner. Add a regression test simulating an interrupted execution and asserting the row does not remain permanentlyapproved.Immediate cleanup (done separately)
The 3 stranded rows are being cancelled (nothing was purchased on AWS, so it is safe). This issue is the durable fix so it stops recurring.
Surfaced while diagnosing the stuck-purchase incident (#621 follow-up).