Skip to content

fix(ci): no concurrency group on rollback, database-migration, cleanup-staging or destroy-fargate-dev #1593

Description

@cristim

Reviewed commit: be11bdcb5. Note: origin/main moved to 3e9660d06 during the review; re-verify against current main before changing code, since a finding may have been fixed or moved.

Where

  • .github/workflows/rollback.yml - no concurrency: key; writes state key github-<env>/terraform.tfstate at :230
  • .github/workflows/database-migration.yml - no concurrency: key
  • .github/workflows/cleanup-staging.yml - no concurrency: key
  • .github/workflows/destroy-fargate-dev.yml - no concurrency: key
  • Contrast, all four deploy workflows declare one: deploy-aws-lambda.yml:29-31, deploy-aws-fargate.yml:28-30, deploy-azure.yml:29, deploy-gcp.yml:27
  • deploy-aws-lambda.yml:150 writes the identical state key github-<environment>/terraform.tfstate

What

Four workflows that mutate the same Terraform state keys and the same production databases can run in unlimited parallel with each other, and in parallel with the deploy workflows. The deploy workflows' own concurrency groups do not cover them, and those groups are keyed on github.ref while the state key is ref-independent, so even between two deploy runs the group does not correspond to the resource being serialised.

Failure scenario

A rollback and a push-triggered deploy overlap. Both target github-<env>/terraform.tfstate. The deploy loses the lock race and fails; its if: failure() || cancelled() handler at deploy-aws-lambda.yml:190 then deletes the lock that the rollback is holding mid-apply, and a third run is free to acquire it and apply concurrently. The account in question holds the purchase-executing Lambda, its RDS instance and its secrets, so a split-brain state there is a production money-path outage, not a cosmetic one.

For database-migration.yml the equivalent is two concurrent migrate up runs against one production database, which produces the golang-migrate dirty state that requires manual CUDLY_FORCE_MIGRATION_VERSION recovery.

Fix direction

Add concurrency: {group: <state-key-or-db-identity>, cancel-in-progress: false} to each of the four workflows, keyed on the target (environment plus cloud), not on github.ref. While doing so, re-key the existing deploy-workflow groups the same way, since deploy-lambda-${{ github.ref }} does not serialise two runs that share a state key but differ in ref (a workflow_dispatch of environment=dev from a feature branch versus the push-to-main dev deploy are in different groups today and share one state key).

Related

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions