You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
fix(iac/aws): deploy role blocked on rds:DescribeDBInstances; full audit of cudly-terraform-deploy grants #1698
Deploy to AWS Lambda on main fails at Terraform Plan with an AccessDenied on rds:DescribeDBInstances — an action the deploy role already has. Deploys are blocked.
More importantly, this is the fifth time the cudly-terraform-deploy grant has been patched by adding whatever action 403'd next (#1496, #1514, #1671, and now RDS). Each patch was correct and each left the next gap in place, because nobody ever derived the full set of API calls the Terraform configuration actually makes. This issue tracks both the immediate unblock and the audit that stops the pattern.
Current behaviour
Run 30809235060, Terraform Plan step:
Error: reading RDS DB Instance (cudly-dev-426fc8af-postgres): operation error RDS: DescribeDBInstances,
StatusCode: 403, AccessDenied: User: arn:aws:sts::909626172446:assumed-role/cudly-terraform-deploy/GitHubActions
is not authorized to perform: rds:DescribeDBInstances on resource: arn:aws:rds:us-east-1:909626172446:db:*
because no identity-based policy allows the rds:DescribeDBInstances action
rds:DescribeDBInstances is granted in terraform/environments/aws/ci-cd-permissions/policy_data.tf (Sid RDSResourceScoped), scoped to arn:aws:rds:*:*:db:cudly-*. The DB instance is cudly-dev-426fc8af-postgres, which matches that prefix. The denial nevertheless names arn:aws:rds:us-east-1:909626172446:db:*.
Root cause
The Terraform id of aws_db_instance is the DbiResourceId (db-<opaque>), not the DB identifier. In terraform-provider-aws v5.100.0, findDBInstanceByID (internal/service/rds/instance.go) branches on the shape of the id: when it looks like a DbiResourceId it issues
and leaves DBInstanceIdentifier nil. RDS authorizes an identifier-less DescribeDBInstances against the wildcard ARN arn:aws:rds:<region>:<account>:db:*. No name-scoped grant can ever match that request, so the refresh 403s.
This is unconditional on every read, not a not-found fallback — the provider's retry with the plain identifier only fires on NotFound, and AccessDenied is not NotFound, so the plan hard-fails. (The repeated # (left over from a partially-failed replacement of this instance) lines in the plan output are deposed objects left by earlier failed applies. They are a consequence of this, not the trigger.)
This is a third distinct failure mode beyond the two already seen:
Granted but unmatchable resource — present with an ARN pattern the request can never match, either because the call carries no identifier (this bug) or because the real ARN uses a different segment / a server-assigned opaque id.
All three look completely fine when you read the policy. Only class 1 is discoverable by eye.
Steps to reproduce
Re-run Deploy to AWS Lambda on main; it fails at Terraform Plan.
Expected behaviour
terraform plan and terraform apply for terraform/environments/aws/ complete without AccessDenied, and the policies contain no grants that are dead on arrival.
Audit findings
Full enumeration of every resource/data block under terraform/environments/aws/ and its modules, diffed against the union of policy_compute.tf, policy_compute_b.tf, policy_data.tf, policy_networking.tf. Deploy-path status resolved against github-{dev,staging,prod}.tfvars and the two deploy workflows.
Blocking / near-blocking (fixed here)
Gap
Class
Effect
rds:DescribeDBInstances scoped to db:cudly-*
unmatchable resource
plan fails today, all envs
iam:PassRole missing ec2.amazonaws.com
unsatisfiable condition
apply fails on fck-nat launch template / ASG replacement (enable_nat_gateway is hardcoded true, all envs)
their only resource types are the two dead patterns above
The RDS proxy path is gated off (enable_rds_proxy = false, hardcoded in terraform/environments/aws/database.tf), so these are latent, not blocking. Tracked as a follow-up rather than widened to db-proxy:* in this change, because scoping an opaque id needs a tag condition whose satisfiability has to be verified rather than assumed — which is precisely the mistake #1671 made.
Latent gaps behind opt-in flags (not fixed here — follow-up)
module "monitoring" (SNS, CloudWatch dashboards, Logs Insights query definitions, X-Ray, GuardDuty, Security Hub) and modules/compute/aws/cleanup-lambda are not instantiated by any environment, and enable_cdn = false disables CloudFront on all three tfvars. None of their permissions exist in the four policies, and several would be mis-scoped if added naively:
logs:PutMetricFilter / DescribeMetricFilters / DeleteMetricFilter — missing. Note the comment at terraform/modules/compute/aws/lambda/migration-alarm.tf:25-29 claims the bootstrap already grants logs:PutMetricFilter; it does not. Flipping enable_migration_alarm today 403s.
logs:PutQueryDefinition / DescribeQueryDefinitions / DeleteQueryDefinition — no resource type, must be Resource = "*".
xray:GetSamplingRules, guardduty:CreateDetector, guardduty:ListDetectors — no resource type, must be "*".
GuardDuty detector ids are opaque 32-char hex; a detector/cudly-* scope would be dead on arrival.
The Security Hub hub ARN is literally hub/default; a cudly-* scope would be dead on arrival.
CloudWatch dashboard ARNs are arn:aws:cloudwatch::<account>:dashboard/<name> — empty region segment and a / separator, so folding them into the existing alarm:cudly-* statement would be dead on arrival.
SNS topic ARNs have no resource-type segment at all (arn:aws:sns:<region>:<account>:<name>), so a topic/cudly-* pattern would be dead on arrival.
logs:DeleteRetentionPolicy — missing; only reachable if a retention variable is set to 0.
Proposed fix
Move rds:DescribeDBInstances (and the DB proxy reads, which have the same unmatchable-ARN problem) into a Resource = "*" read-only statement; keep every mutating RDS action ARN-scoped. Return rds:DescribeDBSubnetGroups to the scoped statement — it does carry a name and does not need widening.
Extend the iam:PassedToService allowlist with the three service principals above.
Add the missing rds:ModifyDBSubnetGroup, iam:Update* and secretsmanager:*VersionStage/ListSecretVersionIds actions.
Drop the dead secretsmanager:ListSecrets grant.
Because adding iam:UpdateAssumeRolePolicy on role/cudly-* also matches cudly-terraform-deploy itself, add an explicit Deny (IAMDenyModifyDeployRole) on the role-mutating actions against the deploy role, mirroring the existing IAMDenyPassDeployRole. This also closes a pre-existing self-escalation: iam:AttachRolePolicy / iam:PutRolePolicy were already grantable against the deploy role.
Document each Resource = "*" with which of the three reasons applies (no resource type / identifier-less request / opaque id), so the next reviewer can tell a justified wildcard from a lazy one.
All four policies stay well under the 6144-character managed-policy limit.
Note on applying this
terraform/environments/aws/ci-cd-permissions/ is a bootstrap-only root. It is applied manually by a privileged human and never by a deploy workflow, so merging the fix does not by itself unblock deploys — the bootstrap apply has to follow.
References
Failing run: 30809235060 (Deploy to AWS Lambda, main)
Summary
Deploy to AWS Lambdaonmainfails at Terraform Plan with anAccessDeniedonrds:DescribeDBInstances— an action the deploy role already has. Deploys are blocked.More importantly, this is the fifth time the
cudly-terraform-deploygrant has been patched by adding whatever action 403'd next (#1496, #1514, #1671, and now RDS). Each patch was correct and each left the next gap in place, because nobody ever derived the full set of API calls the Terraform configuration actually makes. This issue tracks both the immediate unblock and the audit that stops the pattern.Current behaviour
Run 30809235060, Terraform Plan step:
rds:DescribeDBInstancesis granted interraform/environments/aws/ci-cd-permissions/policy_data.tf(SidRDSResourceScoped), scoped toarn:aws:rds:*:*:db:cudly-*. The DB instance iscudly-dev-426fc8af-postgres, which matches that prefix. The denial nevertheless namesarn:aws:rds:us-east-1:909626172446:db:*.Root cause
The Terraform id of
aws_db_instanceis the DbiResourceId (db-<opaque>), not the DB identifier. In terraform-provider-aws v5.100.0,findDBInstanceByID(internal/service/rds/instance.go) branches on the shape of the id: when it looks like a DbiResourceId it issuesand leaves
DBInstanceIdentifiernil. RDS authorizes an identifier-lessDescribeDBInstancesagainst the wildcard ARNarn:aws:rds:<region>:<account>:db:*. No name-scoped grant can ever match that request, so the refresh 403s.This is unconditional on every read, not a not-found fallback — the provider's retry with the plain identifier only fires on
NotFound, andAccessDeniedis notNotFound, so the plan hard-fails. (The repeated# (left over from a partially-failed replacement of this instance)lines in the plan output are deposed objects left by earlier failed applies. They are a consequence of this, not the trigger.)This is a third distinct failure mode beyond the two already seen:
kms:TagResourcegated onaws:ResourceTagduringCreateKey(fix(iac/aws): grant kms:TagResource for KMS CreateKey tag-on-create #1671), or a case-sensitive tag match against the wrong case (fix(iac/aws): case-insensitive tag match on kms:GetKeyPolicy deploy grant #1514).All three look completely fine when you read the policy. Only class 1 is discoverable by eye.
Steps to reproduce
Re-run
Deploy to AWS Lambdaonmain; it fails at Terraform Plan.Expected behaviour
terraform planandterraform applyforterraform/environments/aws/complete withoutAccessDenied, and the policies contain no grants that are dead on arrival.Audit findings
Full enumeration of every
resource/datablock underterraform/environments/aws/and its modules, diffed against the union ofpolicy_compute.tf,policy_compute_b.tf,policy_data.tf,policy_networking.tf. Deploy-path status resolved againstgithub-{dev,staging,prod}.tfvarsand the two deploy workflows.Blocking / near-blocking (fixed here)
rds:DescribeDBInstancesscoped todb:cudly-*iam:PassRolemissingec2.amazonaws.comenable_nat_gatewayis hardcoded true, all envs)iam:PassRolemissingvpc-flow-logs.amazonaws.comaws_flow_log(staging + prod)iam:PassRolemissingevents.amazonaws.comevents:PutTargetswithrole_arn(Fargate path,enable_scheduled_taskstrue everywhere)rds:ModifyDBSubnetGroupaz_countor the private subnet set changesiam:UpdateAssumeRolePolicy,iam:UpdateRole,iam:UpdateRoleDescriptioncudly-*rolesecretsmanager:UpdateSecretVersionStage,secretsmanager:ListSecretVersionIdsAWSCURRENTDead grants found (present in the policy, can never be satisfied)
secretsmanager:ListSecretsinside a statement scoped tosecret:cudly-*arn:aws:rds:*:*:proxy:cudly-*(SidRDSResourceScoped)db-proxy:segment and a server-assignedprx-<opaque>idarn:aws:rds:*:*:target-group:cudly-*(same Sid)prx-tg-<opaque>, never the resource namerds:DeleteDBProxy,rds:ModifyDBProxy,rds:RegisterDBProxyTargets,rds:DeregisterDBProxyTargetsThe RDS proxy path is gated off (
enable_rds_proxy = false, hardcoded interraform/environments/aws/database.tf), so these are latent, not blocking. Tracked as a follow-up rather than widened todb-proxy:*in this change, because scoping an opaque id needs a tag condition whose satisfiability has to be verified rather than assumed — which is precisely the mistake #1671 made.Latent gaps behind opt-in flags (not fixed here — follow-up)
module "monitoring"(SNS, CloudWatch dashboards, Logs Insights query definitions, X-Ray, GuardDuty, Security Hub) andmodules/compute/aws/cleanup-lambdaare not instantiated by any environment, andenable_cdn = falsedisables CloudFront on all three tfvars. None of their permissions exist in the four policies, and several would be mis-scoped if added naively:logs:PutMetricFilter/DescribeMetricFilters/DeleteMetricFilter— missing. Note the comment atterraform/modules/compute/aws/lambda/migration-alarm.tf:25-29claims the bootstrap already grantslogs:PutMetricFilter; it does not. Flippingenable_migration_alarmtoday 403s.logs:PutQueryDefinition/DescribeQueryDefinitions/DeleteQueryDefinition— no resource type, must beResource = "*".xray:GetSamplingRules,guardduty:CreateDetector,guardduty:ListDetectors— no resource type, must be"*".detector/cudly-*scope would be dead on arrival.hub/default; acudly-*scope would be dead on arrival.arn:aws:cloudwatch::<account>:dashboard/<name>— empty region segment and a/separator, so folding them into the existingalarm:cudly-*statement would be dead on arrival.arn:aws:sns:<region>:<account>:<name>), so atopic/cudly-*pattern would be dead on arrival.logs:DeleteRetentionPolicy— missing; only reachable if a retention variable is set to 0.Proposed fix
rds:DescribeDBInstances(and the DB proxy reads, which have the same unmatchable-ARN problem) into aResource = "*"read-only statement; keep every mutating RDS action ARN-scoped. Returnrds:DescribeDBSubnetGroupsto the scoped statement — it does carry a name and does not need widening.iam:PassedToServiceallowlist with the three service principals above.rds:ModifyDBSubnetGroup,iam:Update*andsecretsmanager:*VersionStage/ListSecretVersionIdsactions.secretsmanager:ListSecretsgrant.iam:UpdateAssumeRolePolicyonrole/cudly-*also matchescudly-terraform-deployitself, add an explicitDeny(IAMDenyModifyDeployRole) on the role-mutating actions against the deploy role, mirroring the existingIAMDenyPassDeployRole. This also closes a pre-existing self-escalation:iam:AttachRolePolicy/iam:PutRolePolicywere already grantable against the deploy role.Resource = "*"with which of the three reasons applies (no resource type / identifier-less request / opaque id), so the next reviewer can tell a justified wildcard from a lazy one.All four policies stay well under the 6144-character managed-policy limit.
Note on applying this
terraform/environments/aws/ci-cd-permissions/is a bootstrap-only root. It is applied manually by a privileged human and never by a deploy workflow, so merging the fix does not by itself unblock deploys — the bootstrap apply has to follow.References
Deploy to AWS Lambda,main)kms:GetKeyPolicy), fix(iac/aws): grant kms:TagResource for KMS CreateKey tag-on-create #1671 (kms:TagResourcetag-on-create)terraform/environments/aws/ci-cd-permissions/policy_data.tfinternal/service/rds/instance.goSeverity
Critical — production deploys are blocked, and the same class of defect has recurred four times.