Skip to content

fix(iac/aws): deploy role blocked on rds:DescribeDBInstances; full audit of cudly-terraform-deploy grants #1698

Description

@cristim

Summary

Deploy to AWS Lambda on main fails at Terraform Plan with an AccessDenied on rds:DescribeDBInstances — an action the deploy role already has. Deploys are blocked.

More importantly, this is the fifth time the cudly-terraform-deploy grant has been patched by adding whatever action 403'd next (#1496, #1514, #1671, and now RDS). Each patch was correct and each left the next gap in place, because nobody ever derived the full set of API calls the Terraform configuration actually makes. This issue tracks both the immediate unblock and the audit that stops the pattern.

Current behaviour

Run 30809235060, Terraform Plan step:

Error: reading RDS DB Instance (cudly-dev-426fc8af-postgres): operation error RDS: DescribeDBInstances,
StatusCode: 403, AccessDenied: User: arn:aws:sts::909626172446:assumed-role/cudly-terraform-deploy/GitHubActions
is not authorized to perform: rds:DescribeDBInstances on resource: arn:aws:rds:us-east-1:909626172446:db:*
because no identity-based policy allows the rds:DescribeDBInstances action

rds:DescribeDBInstances is granted in terraform/environments/aws/ci-cd-permissions/policy_data.tf (Sid RDSResourceScoped), scoped to arn:aws:rds:*:*:db:cudly-*. The DB instance is cudly-dev-426fc8af-postgres, which matches that prefix. The denial nevertheless names arn:aws:rds:us-east-1:909626172446:db:*.

Root cause

The Terraform id of aws_db_instance is the DbiResourceId (db-<opaque>), not the DB identifier. In terraform-provider-aws v5.100.0, findDBInstanceByID (internal/service/rds/instance.go) branches on the shape of the id: when it looks like a DbiResourceId it issues

input.Filters = []types.Filter{{Name: aws.String("dbi-resource-id"), Values: []string{id}}}

and leaves DBInstanceIdentifier nil. RDS authorizes an identifier-less DescribeDBInstances against the wildcard ARN arn:aws:rds:<region>:<account>:db:*. No name-scoped grant can ever match that request, so the refresh 403s.

This is unconditional on every read, not a not-found fallback — the provider's retry with the plain identifier only fires on NotFound, and AccessDenied is not NotFound, so the plan hard-fails. (The repeated # (left over from a partially-failed replacement of this instance) lines in the plan output are deposed objects left by earlier failed applies. They are a consequence of this, not the trigger.)

This is a third distinct failure mode beyond the two already seen:

  1. Missing action — plainly absent from the policy (fix(iac/aws): grant deploy role KMS key-lifecycle for OIDC signing-key replace #1496).
  2. Granted but unsatisfiable condition — present but gated on a condition that cannot hold at call time, e.g. kms:TagResource gated on aws:ResourceTag during CreateKey (fix(iac/aws): grant kms:TagResource for KMS CreateKey tag-on-create #1671), or a case-sensitive tag match against the wrong case (fix(iac/aws): case-insensitive tag match on kms:GetKeyPolicy deploy grant #1514).
  3. Granted but unmatchable resource — present with an ARN pattern the request can never match, either because the call carries no identifier (this bug) or because the real ARN uses a different segment / a server-assigned opaque id.

All three look completely fine when you read the policy. Only class 1 is discoverable by eye.

Steps to reproduce

Re-run Deploy to AWS Lambda on main; it fails at Terraform Plan.

Expected behaviour

terraform plan and terraform apply for terraform/environments/aws/ complete without AccessDenied, and the policies contain no grants that are dead on arrival.

Audit findings

Full enumeration of every resource/data block under terraform/environments/aws/ and its modules, diffed against the union of policy_compute.tf, policy_compute_b.tf, policy_data.tf, policy_networking.tf. Deploy-path status resolved against github-{dev,staging,prod}.tfvars and the two deploy workflows.

Blocking / near-blocking (fixed here)

Gap Class Effect
rds:DescribeDBInstances scoped to db:cudly-* unmatchable resource plan fails today, all envs
iam:PassRole missing ec2.amazonaws.com unsatisfiable condition apply fails on fck-nat launch template / ASG replacement (enable_nat_gateway is hardcoded true, all envs)
iam:PassRole missing vpc-flow-logs.amazonaws.com unsatisfiable condition apply fails creating/replacing aws_flow_log (staging + prod)
iam:PassRole missing events.amazonaws.com unsatisfiable condition apply fails on events:PutTargets with role_arn (Fargate path, enable_scheduled_tasks true everywhere)
rds:ModifyDBSubnetGroup missing apply fails if az_count or the private subnet set changes
iam:UpdateAssumeRolePolicy, iam:UpdateRole, iam:UpdateRoleDescription missing apply fails on any trust-policy / description drift on a cudly-* role
secretsmanager:UpdateSecretVersionStage, secretsmanager:ListSecretVersionIds missing apply fails on any secret version carrying a stage other than a lone AWSCURRENT

Dead grants found (present in the policy, can never be satisfied)

Grant Why dead
secretsmanager:ListSecrets inside a statement scoped to secret:cudly-* action has no resource type; an ARN-scoped grant is unsatisfiable
arn:aws:rds:*:*:proxy:cudly-* (Sid RDSResourceScoped) RDS proxy ARNs use the db-proxy: segment and a server-assigned prx-<opaque> id
arn:aws:rds:*:*:target-group:cudly-* (same Sid) target-group ids are prx-tg-<opaque>, never the resource name
rds:DeleteDBProxy, rds:ModifyDBProxy, rds:RegisterDBProxyTargets, rds:DeregisterDBProxyTargets their only resource types are the two dead patterns above

The RDS proxy path is gated off (enable_rds_proxy = false, hardcoded in terraform/environments/aws/database.tf), so these are latent, not blocking. Tracked as a follow-up rather than widened to db-proxy:* in this change, because scoping an opaque id needs a tag condition whose satisfiability has to be verified rather than assumed — which is precisely the mistake #1671 made.

Latent gaps behind opt-in flags (not fixed here — follow-up)

module "monitoring" (SNS, CloudWatch dashboards, Logs Insights query definitions, X-Ray, GuardDuty, Security Hub) and modules/compute/aws/cleanup-lambda are not instantiated by any environment, and enable_cdn = false disables CloudFront on all three tfvars. None of their permissions exist in the four policies, and several would be mis-scoped if added naively:

  • logs:PutMetricFilter / DescribeMetricFilters / DeleteMetricFilter — missing. Note the comment at terraform/modules/compute/aws/lambda/migration-alarm.tf:25-29 claims the bootstrap already grants logs:PutMetricFilter; it does not. Flipping enable_migration_alarm today 403s.
  • logs:PutQueryDefinition / DescribeQueryDefinitions / DeleteQueryDefinition — no resource type, must be Resource = "*".
  • xray:GetSamplingRules, guardduty:CreateDetector, guardduty:ListDetectors — no resource type, must be "*".
  • GuardDuty detector ids are opaque 32-char hex; a detector/cudly-* scope would be dead on arrival.
  • The Security Hub hub ARN is literally hub/default; a cudly-* scope would be dead on arrival.
  • CloudWatch dashboard ARNs are arn:aws:cloudwatch::<account>:dashboard/<name> — empty region segment and a / separator, so folding them into the existing alarm:cudly-* statement would be dead on arrival.
  • SNS topic ARNs have no resource-type segment at all (arn:aws:sns:<region>:<account>:<name>), so a topic/cudly-* pattern would be dead on arrival.
  • logs:DeleteRetentionPolicy — missing; only reachable if a retention variable is set to 0.

Proposed fix

  1. Move rds:DescribeDBInstances (and the DB proxy reads, which have the same unmatchable-ARN problem) into a Resource = "*" read-only statement; keep every mutating RDS action ARN-scoped. Return rds:DescribeDBSubnetGroups to the scoped statement — it does carry a name and does not need widening.
  2. Extend the iam:PassedToService allowlist with the three service principals above.
  3. Add the missing rds:ModifyDBSubnetGroup, iam:Update* and secretsmanager:*VersionStage/ListSecretVersionIds actions.
  4. Drop the dead secretsmanager:ListSecrets grant.
  5. Because adding iam:UpdateAssumeRolePolicy on role/cudly-* also matches cudly-terraform-deploy itself, add an explicit Deny (IAMDenyModifyDeployRole) on the role-mutating actions against the deploy role, mirroring the existing IAMDenyPassDeployRole. This also closes a pre-existing self-escalation: iam:AttachRolePolicy / iam:PutRolePolicy were already grantable against the deploy role.
  6. Document each Resource = "*" with which of the three reasons applies (no resource type / identifier-less request / opaque id), so the next reviewer can tell a justified wildcard from a lazy one.

All four policies stay well under the 6144-character managed-policy limit.

Note on applying this

terraform/environments/aws/ci-cd-permissions/ is a bootstrap-only root. It is applied manually by a privileged human and never by a deploy workflow, so merging the fix does not by itself unblock deploys — the bootstrap apply has to follow.

References

Severity

Critical — production deploys are blocked, and the same class of defect has recurred four times.

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions