Reviewed commit: be11bdcb5. Note: origin/main moved to 3e9660d06 during the review; re-verify against current main before changing code, since a finding may have been fixed or moved.
Where
.github/workflows/cleanup-staging.yml:47, :117, :187, :252 - four terraform destroy -auto-approve jobs, none declaring environment:
.github/workflows/destroy-fargate-dev.yml:38 - no environment:
.github/workflows/rollback.yml:45 (validate) and :381 (summary) - no environment:
- Contrast:
.github/workflows/deploy-aws-lambda.yml:124 and database-migration.yml:170,246,313 do declare environment:
terraform/environments/aws/ci-cd-permissions/role.tf:28-33 - trust policy accepts repo:<org/repo>:ref:refs/heads/main
- Typed-string guards:
cleanup-staging.yml:41, destroy-fargate-dev.yml:33
cleanup-staging.yml:228-232 - the Azure job breaks the state lease before destroying
What
The repo's destructive workflows have no GitHub Environment binding, so none of the protection that Environments provide - required reviewers, wait timers, branch restrictions - ever applies to them. The only guard is a typed-string confirmation in a preceding job in the same run, which the same person who dispatched the workflow types. A second human is never involved.
This is not mitigated by the OIDC trust policy, because the AWS deploy role's trust policy accepts ref:refs/heads/main independently of the environment:{dev,staging,prod} subjects. A job with no environment: binding running on main can still mint an OIDC token and assume cudly-terraform-deploy, the full deploy role. The environment approval gate is therefore architecturally bypassable, not merely unconfigured on these workflows.
Failure scenario
Any repository collaborator with write access dispatches Cleanup Staging Resources from main, types destroy into the confirmation input, and four terraform destroy -auto-approve runs delete the entire staging estate across AWS, Azure and GCP. The Azure job first breaks the state lease (cleanup-staging.yml:228-232), so a concurrent apply loses its lock as well. No second human is ever consulted, no reviewer is notified, and there is no wait timer in which to cancel. The same shape applies to destroy-fargate-dev.yml and to rollback.yml's unbound validate / summary jobs, which additionally hold id-token: write (see #1542).
Fix direction
Two changes, and both are needed - either alone leaves the gate bypassable:
- Bind every destroy / rollback job to a protected GitHub Environment with required reviewers.
- Remove
repo:...:ref:refs/heads/main from the AWS role's sub allow-list in terraform/environments/aws/ci-cd-permissions/role.tf so deploy credentials are obtainable only through an environment-gated job. Note the bootstrap/runtime split: ci-cd-permissions/ is applied manually by a privileged human, so this change needs a bootstrap re-apply, not a workflow apply.
Related
Reviewed commit:
be11bdcb5. Note:origin/mainmoved to3e9660d06during the review; re-verify against currentmainbefore changing code, since a finding may have been fixed or moved.Where
.github/workflows/cleanup-staging.yml:47,:117,:187,:252- fourterraform destroy -auto-approvejobs, none declaringenvironment:.github/workflows/destroy-fargate-dev.yml:38- noenvironment:.github/workflows/rollback.yml:45(validate) and:381(summary) - noenvironment:.github/workflows/deploy-aws-lambda.yml:124anddatabase-migration.yml:170,246,313do declareenvironment:terraform/environments/aws/ci-cd-permissions/role.tf:28-33- trust policy acceptsrepo:<org/repo>:ref:refs/heads/maincleanup-staging.yml:41,destroy-fargate-dev.yml:33cleanup-staging.yml:228-232- the Azure job breaks the state lease before destroyingWhat
The repo's destructive workflows have no GitHub Environment binding, so none of the protection that Environments provide - required reviewers, wait timers, branch restrictions - ever applies to them. The only guard is a typed-string confirmation in a preceding job in the same run, which the same person who dispatched the workflow types. A second human is never involved.
This is not mitigated by the OIDC trust policy, because the AWS deploy role's trust policy accepts
ref:refs/heads/mainindependently of theenvironment:{dev,staging,prod}subjects. A job with noenvironment:binding running onmaincan still mint an OIDC token and assumecudly-terraform-deploy, the full deploy role. The environment approval gate is therefore architecturally bypassable, not merely unconfigured on these workflows.Failure scenario
Any repository collaborator with write access dispatches
Cleanup Staging Resourcesfrommain, typesdestroyinto the confirmation input, and fourterraform destroy -auto-approveruns delete the entire staging estate across AWS, Azure and GCP. The Azure job first breaks the state lease (cleanup-staging.yml:228-232), so a concurrent apply loses its lock as well. No second human is ever consulted, no reviewer is notified, and there is no wait timer in which to cancel. The same shape applies todestroy-fargate-dev.ymland torollback.yml's unboundvalidate/summaryjobs, which additionally holdid-token: write(see #1542).Fix direction
Two changes, and both are needed - either alone leaves the gate bypassable:
repo:...:ref:refs/heads/mainfrom the AWS role'ssuballow-list interraform/environments/aws/ci-cd-permissions/role.tfso deploy credentials are obtainable only through an environment-gated job. Note the bootstrap/runtime split:ci-cd-permissions/is applied manually by a privileged human, so this change needs a bootstrap re-apply, not a workflow apply.Related
rollback.ymlreasoninjection whose escalation path depends on exactly this missing binding on thevalidatejob.ci: prod down migration with steps>=total bypasses confirmation gate) - the same "typed-string guard is the only gate" pattern indatabase-migration.yml.