Skip to content

sec(ci): destroy and rollback workflows have no environment binding, so no reviewer gate applies #1591

Description

@cristim

Reviewed commit: be11bdcb5. Note: origin/main moved to 3e9660d06 during the review; re-verify against current main before changing code, since a finding may have been fixed or moved.

Where

  • .github/workflows/cleanup-staging.yml:47, :117, :187, :252 - four terraform destroy -auto-approve jobs, none declaring environment:
  • .github/workflows/destroy-fargate-dev.yml:38 - no environment:
  • .github/workflows/rollback.yml:45 (validate) and :381 (summary) - no environment:
  • Contrast: .github/workflows/deploy-aws-lambda.yml:124 and database-migration.yml:170,246,313 do declare environment:
  • terraform/environments/aws/ci-cd-permissions/role.tf:28-33 - trust policy accepts repo:<org/repo>:ref:refs/heads/main
  • Typed-string guards: cleanup-staging.yml:41, destroy-fargate-dev.yml:33
  • cleanup-staging.yml:228-232 - the Azure job breaks the state lease before destroying

What

The repo's destructive workflows have no GitHub Environment binding, so none of the protection that Environments provide - required reviewers, wait timers, branch restrictions - ever applies to them. The only guard is a typed-string confirmation in a preceding job in the same run, which the same person who dispatched the workflow types. A second human is never involved.

This is not mitigated by the OIDC trust policy, because the AWS deploy role's trust policy accepts ref:refs/heads/main independently of the environment:{dev,staging,prod} subjects. A job with no environment: binding running on main can still mint an OIDC token and assume cudly-terraform-deploy, the full deploy role. The environment approval gate is therefore architecturally bypassable, not merely unconfigured on these workflows.

Failure scenario

Any repository collaborator with write access dispatches Cleanup Staging Resources from main, types destroy into the confirmation input, and four terraform destroy -auto-approve runs delete the entire staging estate across AWS, Azure and GCP. The Azure job first breaks the state lease (cleanup-staging.yml:228-232), so a concurrent apply loses its lock as well. No second human is ever consulted, no reviewer is notified, and there is no wait timer in which to cancel. The same shape applies to destroy-fargate-dev.yml and to rollback.yml's unbound validate / summary jobs, which additionally hold id-token: write (see #1542).

Fix direction

Two changes, and both are needed - either alone leaves the gate bypassable:

  1. Bind every destroy / rollback job to a protected GitHub Environment with required reviewers.
  2. Remove repo:...:ref:refs/heads/main from the AWS role's sub allow-list in terraform/environments/aws/ci-cd-permissions/role.tf so deploy credentials are obtainable only through an environment-gated job. Note the bootstrap/runtime split: ci-cd-permissions/ is applied manually by a privileged human, so this change needs a bootstrap re-apply, not a workflow apply.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions