Skip to content

fix(iac/aws): grant rds:DescribeDBSnapshots for the prod final-snapshot destroy path #160

Description

@cristim

Summary

rds:DescribeDBSnapshots is granted by no statement in any of the four cudly-terraform-deploy policies. On a prod database destroy the provider takes a final snapshot and then waits on it, which requires that action. Destroy-path only; it does not block the deploy that LeanerCloud/cloud-commitments-cli#1699 fixes.

Current behaviour

Grep across terraform/environments/aws/ci-cd-permissions/ returns no match for DescribeDBSnapshots in policy_compute.tf, policy_compute_b.tf, policy_data.tf or policy_networking.tf. rds:CreateDBSnapshot is granted (Sid RDSResourceScoped, scoped to arn:aws:rds:*:*:snapshot:cudly-*), so the snapshot can be created but its completion cannot be polled.

Why it matters

terraform/environments/aws/github-prod.tfvars:50 sets:

database_skip_final_snapshot   = false

so a prod aws_db_instance destroy or replacement takes a final snapshot. The AWS provider then polls DescribeDBSnapshots until the snapshot reaches available. Without the action the destroy fails partway through, after the snapshot has been requested.

dev and staging set database_skip_final_snapshot = true (github-dev.tfvars:52, github-staging.tfvars:49), so neither exercises this path.

Steps to verify the gap

  1. grep -rn DescribeDBSnapshots terraform/environments/aws/ci-cd-permissions/ — no results.
  2. Confirm database_skip_final_snapshot = false in github-prod.tfvars:50.
  3. A prod destroy or a replacement of aws_db_instance.main (terraform/modules/database/aws/main.tf:120) would 403 on rds:DescribeDBSnapshots while waiting for the final snapshot.

Expected behaviour

A prod DB destroy or replacement completes, including the final-snapshot wait.

Proposed fix

Add rds:DescribeDBSnapshots to policy_data.tf. Decide the scoping deliberately rather than by analogy — this is the exact judgement call LeanerCloud/cloud-commitments-cli#1698 was about:

  • The snapshot resource type exists and the final snapshot name is name-predictable (cudly-<env>-<hex>-final-snapshot), so arn:aws:rds:*:*:snapshot:cudly-* is syntactically valid, and Sid RDSResourceScoped already carries that ARN.
  • But the same enumerate-vs-target question that produced the rds:DescribeDBInstances outage applies here: if the provider polls with an identifier-less DescribeDBSnapshots (e.g. filtered by DB instance rather than by snapshot identifier), a scoped grant cannot authorize it and it belongs in Sid RDSDescribeAccountWide alongside the other reads.

Check the actual request shape in terraform-provider-aws v5.100.0 (internal/service/rds/, the final-snapshot wait path) before choosing. Do not assume either answer.

cudly-deploy-data is at 4588 of 6144 characters, so there is ample headroom either way.

Note this also requires a manual bootstrap apply by a privileged human, since terraform/environments/aws/ci-cd-permissions/ is applied outside any deploy workflow.

References

Severity

Medium — destroy/replacement path only, and only in prod, but it fails after the snapshot has been requested, leaving the operation half-done at the worst possible moment.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions