Reviewed commit: be11bdcb5. Note: origin/main moved to 3e9660d06 during the review; re-verify against current main before changing code, since a finding may have been fixed or moved.
Where
.github/workflows/rollback.yml - no concurrency: key; writes state key github-<env>/terraform.tfstate at :230
.github/workflows/database-migration.yml - no concurrency: key
.github/workflows/cleanup-staging.yml - no concurrency: key
.github/workflows/destroy-fargate-dev.yml - no concurrency: key
- Contrast, all four deploy workflows declare one:
deploy-aws-lambda.yml:29-31, deploy-aws-fargate.yml:28-30, deploy-azure.yml:29, deploy-gcp.yml:27
deploy-aws-lambda.yml:150 writes the identical state key github-<environment>/terraform.tfstate
What
Four workflows that mutate the same Terraform state keys and the same production databases can run in unlimited parallel with each other, and in parallel with the deploy workflows. The deploy workflows' own concurrency groups do not cover them, and those groups are keyed on github.ref while the state key is ref-independent, so even between two deploy runs the group does not correspond to the resource being serialised.
Failure scenario
A rollback and a push-triggered deploy overlap. Both target github-<env>/terraform.tfstate. The deploy loses the lock race and fails; its if: failure() || cancelled() handler at deploy-aws-lambda.yml:190 then deletes the lock that the rollback is holding mid-apply, and a third run is free to acquire it and apply concurrently. The account in question holds the purchase-executing Lambda, its RDS instance and its secrets, so a split-brain state there is a production money-path outage, not a cosmetic one.
For database-migration.yml the equivalent is two concurrent migrate up runs against one production database, which produces the golang-migrate dirty state that requires manual CUDLY_FORCE_MIGRATION_VERSION recovery.
Fix direction
Add concurrency: {group: <state-key-or-db-identity>, cancel-in-progress: false} to each of the four workflows, keyed on the target (environment plus cloud), not on github.ref. While doing so, re-key the existing deploy-workflow groups the same way, since deploy-lambda-${{ github.ref }} does not serialise two runs that share a state key but differ in ref (a workflow_dispatch of environment=dev from a feature branch versus the push-to-main dev deploy are in different groups today and share one state key).
Related
Reviewed commit:
be11bdcb5. Note:origin/mainmoved to3e9660d06during the review; re-verify against currentmainbefore changing code, since a finding may have been fixed or moved.Where
.github/workflows/rollback.yml- noconcurrency:key; writes state keygithub-<env>/terraform.tfstateat:230.github/workflows/database-migration.yml- noconcurrency:key.github/workflows/cleanup-staging.yml- noconcurrency:key.github/workflows/destroy-fargate-dev.yml- noconcurrency:keydeploy-aws-lambda.yml:29-31,deploy-aws-fargate.yml:28-30,deploy-azure.yml:29,deploy-gcp.yml:27deploy-aws-lambda.yml:150writes the identical state keygithub-<environment>/terraform.tfstateWhat
Four workflows that mutate the same Terraform state keys and the same production databases can run in unlimited parallel with each other, and in parallel with the deploy workflows. The deploy workflows' own concurrency groups do not cover them, and those groups are keyed on
github.refwhile the state key is ref-independent, so even between two deploy runs the group does not correspond to the resource being serialised.Failure scenario
A rollback and a push-triggered deploy overlap. Both target
github-<env>/terraform.tfstate. The deploy loses the lock race and fails; itsif: failure() || cancelled()handler atdeploy-aws-lambda.yml:190then deletes the lock that the rollback is holding mid-apply, and a third run is free to acquire it and apply concurrently. The account in question holds the purchase-executing Lambda, its RDS instance and its secrets, so a split-brain state there is a production money-path outage, not a cosmetic one.For
database-migration.ymlthe equivalent is two concurrentmigrate upruns against one production database, which produces the golang-migratedirtystate that requires manualCUDLY_FORCE_MIGRATION_VERSIONrecovery.Fix direction
Add
concurrency: {group: <state-key-or-db-identity>, cancel-in-progress: false}to each of the four workflows, keyed on the target (environment plus cloud), not ongithub.ref. While doing so, re-key the existing deploy-workflow groups the same way, sincedeploy-lambda-${{ github.ref }}does not serialise two runs that share a state key but differ in ref (aworkflow_dispatchofenvironment=devfrom a feature branch versus the push-to-maindev deploy are in different groups today and share one state key).Related
deploy-aws-lambda.ymlfailure-path lock delete; the concurrency gap is what turns that delete from "questionable" into "actively deletes another run's lock", so the two fixes belong together.