Azure deploy state has silently lost resource entries, leaving resources live in Azure but untracked by Terraform. Every subsequent apply then tries to create them and Azure rejects the duplicate, so the deploy is permanently red until someone imports by hand.
Currently biting: cost_management_reader and subscription_reader in module.compute_container_apps[0] both 409 with RoleAssignmentExists. Neither sets an explicit name, so azurerm mints a fresh UUID per apply while Azure dedupes on (scope, role, principal). That makes the failure permanent rather than transient.
Root cause: two defects that are only dangerous together
1. The concurrency group is keyed on branch, but the state file is keyed on environment.
concurrency:
group: deploy-azure-${{ github.ref }}
prepare maps every non-dispatch event to dev (*) ENVIRONMENT=dev ;;), and the state key is github-${ENVIRONMENT}.terraform.tfstate. So two runs on different branches land in different concurrency groups, run simultaneously, and apply against one state file.
2. Break stale state lock (if any) breaks the lease unconditionally.
No if:, no lease-age check, no check that a lease is even held, and no --lease-break-period, so it defaults to breaking an active lease immediately.
Alone, defect 1 is survivable: the second run would fail loudly with Error acquiring the state lock, which is the blob lease doing its job. Defect 2 is what converts a safe, noisy collision into a silent one. The second run breaks the first run's live lease, both proceed, and the later writer's stale read drops entries the earlier one had recorded.
The step is named "stale" but never establishes staleness.
Evidence
At the exact creation instant of cost_management_reader (2026-06-09T15:42:51Z), two deploy runs were in flight on different refs against the same dev state:
| run |
branch |
window |
| 27217251153 |
feat/multicloud-web-frontend |
15:32:24 -> 15:58:15 |
| 27217600951 |
main |
15:38:01 -> 16:03:01 |
June 9 had 14 deploy runs with repeated overlaps.
Not fully proven: subscription_reader (14:58:51Z) had only one run in flight (27214016369, 14:40:44 -> 15:07:32). The same mechanism explains it through a later lost write, but no specific pair has been pinned to it. Recording that as unproven rather than asserting it.
Alternative hypothesis, ruled out
A past state-key rename was considered, since dev.terraform.tfstate sits beside the active github-dev.terraform.tfstate. It is not the cause: dev.terraform.tfstate holds 0 resources, serial 284, lineage 4a6d936f-..., which is a different lineage from the active state (c2ca9703-...), and was last modified 2026-03-25, ten weeks before the June 9 creations. It never held these resources. github-staging.terraform.tfstate is likewise a stub.
Blast radius
Four workflows write this state: deploy-azure.yml, deploy-all.yml, cleanup-staging.yml, rollback.yml. Any pair overlapping can collide, so a fix confined to deploy-azure.yml is incomplete.
Suggested fix
Key the concurrency group on the thing that actually identifies the state file, not the branch, so all runs targeting one environment serialize regardless of ref. Then either delete the lease-break step or make it genuinely conditional: check lease state, require a minimum age, and never break an active lease. A loud Error acquiring the state lock is the correct outcome of a real collision and should not be engineered away.
Whatever lands must cover all four workflows.
Verification
Reproduce first: trigger two overlapping runs on different branches and confirm they currently run concurrently against one state file. After the fix, confirm the second serializes behind the first rather than failing or proceeding in parallel. A fix that merely makes collisions rarer is not a fix, since the failure mode is silent.
Related
Azure deploy state has silently lost resource entries, leaving resources live in Azure but untracked by Terraform. Every subsequent apply then tries to create them and Azure rejects the duplicate, so the deploy is permanently red until someone imports by hand.
Currently biting:
cost_management_readerandsubscription_readerinmodule.compute_container_apps[0]both 409 withRoleAssignmentExists. Neither sets an explicitname, so azurerm mints a fresh UUID per apply while Azure dedupes on (scope, role, principal). That makes the failure permanent rather than transient.Root cause: two defects that are only dangerous together
1. The concurrency group is keyed on branch, but the state file is keyed on environment.
preparemaps every non-dispatch event todev(*) ENVIRONMENT=dev ;;), and the state key isgithub-${ENVIRONMENT}.terraform.tfstate. So two runs on different branches land in different concurrency groups, run simultaneously, and apply against one state file.2.
Break stale state lock (if any)breaks the lease unconditionally.No
if:, no lease-age check, no check that a lease is even held, and no--lease-break-period, so it defaults to breaking an active lease immediately.Alone, defect 1 is survivable: the second run would fail loudly with
Error acquiring the state lock, which is the blob lease doing its job. Defect 2 is what converts a safe, noisy collision into a silent one. The second run breaks the first run's live lease, both proceed, and the later writer's stale read drops entries the earlier one had recorded.The step is named "stale" but never establishes staleness.
Evidence
At the exact creation instant of
cost_management_reader(2026-06-09T15:42:51Z), two deploy runs were in flight on different refs against the samedevstate:feat/multicloud-web-frontendmainJune 9 had 14 deploy runs with repeated overlaps.
Not fully proven:
subscription_reader(14:58:51Z) had only one run in flight (27214016369, 14:40:44 -> 15:07:32). The same mechanism explains it through a later lost write, but no specific pair has been pinned to it. Recording that as unproven rather than asserting it.Alternative hypothesis, ruled out
A past state-key rename was considered, since
dev.terraform.tfstatesits beside the activegithub-dev.terraform.tfstate. It is not the cause:dev.terraform.tfstateholds 0 resources, serial 284, lineage4a6d936f-..., which is a different lineage from the active state (c2ca9703-...), and was last modified 2026-03-25, ten weeks before the June 9 creations. It never held these resources.github-staging.terraform.tfstateis likewise a stub.Blast radius
Four workflows write this state:
deploy-azure.yml,deploy-all.yml,cleanup-staging.yml,rollback.yml. Any pair overlapping can collide, so a fix confined todeploy-azure.ymlis incomplete.Suggested fix
Key the concurrency group on the thing that actually identifies the state file, not the branch, so all runs targeting one environment serialize regardless of ref. Then either delete the lease-break step or make it genuinely conditional: check lease state, require a minimum age, and never break an active lease. A loud
Error acquiring the state lockis the correct outcome of a real collision and should not be engineered away.Whatever lands must cover all four workflows.
Verification
Reproduce first: trigger two overlapping runs on different branches and confirm they currently run concurrently against one state file. After the fix, confirm the second serializes behind the first rather than failing or proceeding in parallel. A fix that merely makes collisions rarer is not a fix, since the failure mode is silent.
Related