From 8a7a4207a65f34c804c925380df7bd6cde34a253 Mon Sep 17 00:00:00 2001 From: Cristian Magherusan-Stanciu Date: Mon, 3 Aug 2026 13:36:20 +0200 Subject: [PATCH 1/3] fix(iac/aws): unblock deploy on rds:DescribeDBInstances and close audited grant gaps terraform plan on main fails with AccessDenied on rds:DescribeDBInstances against arn:aws:rds:us-east-1:*:db:*, even though the action was already granted on arn:aws:rds:*:*:db:cudly-* and the instance name matches that prefix. The Terraform id of aws_db_instance is the DbiResourceId (db-), not the DB identifier. provider v5.100.0's findDBInstanceByID branches on the id shape and, for a DbiResourceId, sends DescribeDBInstances with Filters=[dbi-resource-id] and no DBInstanceIdentifier. RDS authorizes an identifier-less DescribeDBInstances against the wildcard ARN db:*, which no name-scoped grant can match. It is unconditional on every read: the provider's retry with the plain identifier only fires on NotFound, and AccessDenied is not NotFound, so the plan hard-fails. Rather than add the one action that 403'd, this audits every resource and data block under terraform/environments/aws/ and its modules against the union of the four deploy policies, and fixes the gaps that block the deploy path. It also names a third failure mode the previous four patches did not account for: a grant can be present and still unusable because the request carries no identifier, or because the real ARN uses a different segment or a server-assigned opaque id. Fixed here: - Move rds:DescribeDBInstances and the DB proxy reads to a read-only Resource = "*" statement; every mutating RDS action stays ARN-scoped. Return rds:DescribeDBSubnetGroups to the scoped statement, since that call does carry a name and needs no widening. - Add rds:ModifyDBSubnetGroup, which is required whenever az_count or the private subnet set changes. - Extend the iam:PassedToService allowlist with ec2.amazonaws.com (fck-nat instance profile, on the deploy path in every environment), vpc-flow-logs.amazonaws.com (staging and prod) and events.amazonaws.com (scheduled ECS tasks). Each was a granted-but-unsatisfiable condition of the same shape as the kms:TagResource bug. - Add iam:UpdateAssumeRolePolicy, iam:UpdateRole and iam:UpdateRoleDescription for trust-policy and description drift on cudly-* roles. - Add secretsmanager:UpdateSecretVersionStage and secretsmanager:ListSecretVersionIds for secret versions carrying a stage other than a lone AWSCURRENT. - Drop secretsmanager:ListSecrets, which has no resource type and so could never be satisfied inside an ARN-scoped statement. Nothing on the deploy path calls it. - Add IAMDenyModifyDeployRoleAndPolicies. iam:UpdateAssumeRolePolicy on role/cudly-* also matches cudly-terraform-deploy, so granting it would let the deploy role rewrite its own trust policy. The Deny closes that along with two pre-existing doors into the same escalation: AttachRolePolicy / PutRolePolicy on the deploy role, and CreatePolicyVersion on the cudly-deploy-* policies. Denying only the role side would have looked complete without being complete. - Document, for every Resource = "*", which of the three reasons applies (no resource type, identifier-less request, or opaque id) so a reviewer can tell a justified wildcard from a lazy one, and mark the known-dead RDS proxy ARN scopes in place so nothing new is added to them. Deliberately not fixed here, enumerated in the issue and tracked separately: the RDS proxy ARN scopes (proxy:/target-group: are unmatchable but the path is gated off by enable_rds_proxy = false, and scoping an opaque id needs a tag condition whose satisfiability has to be verified rather than assumed), and the permissions for the monitoring module, cleanup-lambda and CloudFront, none of which are instantiated on any deploy path today. All four managed policies stay under the 6144-character limit; cudly-deploy-data goes from 3548 to 4219 characters. This root is bootstrap-only and is applied manually by a privileged human, so merging does not by itself unblock deploys. Closes #1698 --- .../aws/ci-cd-permissions/policy_data.tf | 165 ++++++++++++++++-- 1 file changed, 150 insertions(+), 15 deletions(-) diff --git a/terraform/environments/aws/ci-cd-permissions/policy_data.tf b/terraform/environments/aws/ci-cd-permissions/policy_data.tf index eca60627e..c8c0f4017 100644 --- a/terraform/environments/aws/ci-cd-permissions/policy_data.tf +++ b/terraform/environments/aws/ci-cd-permissions/policy_data.tf @@ -39,6 +39,17 @@ resource "aws_iam_policy" "data" { "iam:UntagInstanceProfile", "iam:UntagPolicy", "iam:UntagRole", + # iam:UpdateAssumeRolePolicy / UpdateRole / UpdateRoleDescription are + # what the provider calls when a role's assume_role_policy, + # max_session_duration or description drifts (resourceRoleUpdate in + # internal/service/iam/role.go). Without them any edit to a trust + # policy or role description in terraform/modules/** fails the apply + # rather than the plan, which is why this gap stayed invisible. + # IAMDenyModifyDeployRoleAndPolicies below stops these from being + # turned on the deploy role itself. + "iam:UpdateAssumeRolePolicy", + "iam:UpdateRole", + "iam:UpdateRoleDescription", ] Resource = [ "arn:aws:iam::*:role/cudly-*", @@ -65,6 +76,20 @@ resource "aws_iam_policy" "data" { "lambda.amazonaws.com", "ecs-tasks.amazonaws.com", "rds.amazonaws.com", + # ec2: the fck-nat instance profile referenced by the NAT launch + # template and Auto Scaling group + # (terraform/modules/networking/aws, enable_nat_gateway is + # hardcoded true in terraform/environments/aws/networking.tf, so + # this is on the deploy path in every environment). + "ec2.amazonaws.com", + # vpc-flow-logs: aws_flow_log passes its delivery role via + # DeliverLogsPermissionArn (enable_flow_logs is true for staging + # and prod). + "vpc-flow-logs.amazonaws.com", + # events: EventBridge targets that carry a role_arn, used by the + # scheduled ECS tasks on the Fargate deploy path + # (enable_scheduled_tasks is true in every tfvars). + "events.amazonaws.com", ] } } @@ -85,6 +110,54 @@ resource "aws_iam_policy" "data" { Action = ["iam:PassRole"] Resource = ["arn:aws:iam::*:role/cudly-terraform-deploy"] }, + { + # IAMDenyPassDeployRole above closes the pass-the-deploy-role door. + # This closes the other two doors into the same escalation, both of + # which IAMRolesAndPolicies leaves open because cudly-terraform-deploy + # is itself a cudly-* role and cudly-deploy-* are cudly-* policies: + # + # 1. Role side: iam:AttachRolePolicy / iam:PutRolePolicy let the + # deploy role grant itself AdministratorAccess, and + # iam:UpdateAssumeRolePolicy (added to IAMRolesAndPolicies above) + # lets it rewrite its own trust policy so an arbitrary principal + # can assume it. + # 2. Policy side: iam:CreatePolicyVersion on cudly-deploy-data (or + # any sibling) rewrites the deploy role's own permissions in + # place, reaching the same admin without ever touching the role. + # Denying only the role side would look complete and would not be. + # + # Either turns a leaked GitHub Actions token into persistent + # account-wide admin. An explicit Deny always beats the Allow, so this + # closes both loops while leaving the workload roles and policies + # (cudly---*) fully manageable. + # + # This costs the deploy path nothing: cudly-terraform-deploy and the + # four cudly-deploy-* policies are all declared in this bootstrap + # root, which is applied manually by a privileged human and never by a + # deploy workflow. Action and Resource are matched as a cross product, + # so the combinations that do not apply (a role action against a + # policy ARN and vice versa) are simply inert. + Sid = "IAMDenyModifyDeployRoleAndPolicies" + Effect = "Deny" + Action = [ + "iam:AttachRolePolicy", + "iam:CreatePolicyVersion", + "iam:DeletePolicy", + "iam:DeletePolicyVersion", + "iam:DeleteRole", + "iam:DeleteRolePolicy", + "iam:DetachRolePolicy", + "iam:PutRolePolicy", + "iam:SetDefaultPolicyVersion", + "iam:UpdateAssumeRolePolicy", + "iam:UpdateRole", + "iam:UpdateRoleDescription", + ] + Resource = [ + "arn:aws:iam::*:role/cudly-terraform-deploy", + "arn:aws:iam::*:policy/cudly-deploy-*", + ] + }, { Sid = "IAMReadForPassRole" Effect = "Allow" @@ -112,11 +185,29 @@ resource "aws_iam_policy" "data" { } }, { - # RDS actions that take a specific resource ARN are scoped to - # cudly-* DB instances, subnet groups, and proxies. This prevents - # the deploy SA from deleting or modifying unrelated RDS instances - # in the same account. DescribeDBEngineVersions does not accept a - # resource ARN (account-wide catalogue lookup) and is split below. + # RDS actions whose request carries a name we control are scoped to + # cudly-* DB instances, subnet groups and snapshots. This prevents the + # deploy SA from deleting or modifying unrelated RDS resources in the + # same account. + # + # rds:DescribeDBInstances deliberately does NOT live here; see + # RDSDescribeAccountWide below. + # + # KNOWN DEAD SCOPES: the proxy actions below (DeleteDBProxy, + # ModifyDBProxy, RegisterDBProxyTargets, DeregisterDBProxyTargets) + # can never be authorized, because their only resource types are + # `proxy` and `target-group`, whose real ARNs are + # arn:aws:rds:::db-proxy:prx- and + # arn:aws:rds:::target-group:prx-tg-: a different ARN + # segment (`db-proxy`, not `proxy`) and a server-assigned opaque id + # that never contains the resource name. `proxy:cudly-*` and + # `target-group:cudly-*` therefore match nothing. Nothing exercises + # this today (modules/database/aws is instantiated with + # enable_rds_proxy = false, terraform/environments/aws/database.tf), + # so it is left as-is rather than widened to db-proxy:* here; fixing + # it needs a scoping mechanism that actually constrains an opaque id + # (tag condition), tracked separately. Do NOT add new proxy actions + # to this statement; they would be dead on arrival. Sid = "RDSResourceScoped" Effect = "Allow" Action = [ @@ -126,16 +217,13 @@ resource "aws_iam_policy" "data" { "rds:CreateDBSubnetGroup", "rds:DeleteDBInstance", "rds:DeleteDBSubnetGroup", - "rds:DescribeDBInstances", "rds:DescribeDBSubnetGroups", "rds:ListTagsForResource", "rds:DeleteDBProxy", "rds:DeregisterDBProxyTargets", - "rds:DescribeDBProxies", - "rds:DescribeDBProxyTargetGroups", - "rds:DescribeDBProxyTargets", "rds:ModifyDBInstance", "rds:ModifyDBProxy", + "rds:ModifyDBSubnetGroup", "rds:RegisterDBProxyTargets", "rds:RemoveTagsFromResource", ] @@ -148,11 +236,43 @@ resource "aws_iam_policy" "data" { ] }, { - # DescribeDBEngineVersions is a catalogue lookup that doesn't - # accept a resource ARN at the API level. Read-only, no state change. - Sid = "RDSDescribeEngineVersions" - Effect = "Allow" - Action = ["rds:DescribeDBEngineVersions"] + # RDS reads that an ARN-scoped grant can never authorize. + # + # rds:DescribeDBEngineVersions is a catalogue lookup with no resource + # type at all, so it can only be granted on "*". + # + # rds:DescribeDBInstances is here because of the request shape the AWS + # provider uses, not because the action lacks a resource type. The + # Terraform id of aws_db_instance IS the DbiResourceId (`db-`, + # set at internal/service/rds/instance.go in provider v5.100.0), and + # findDBInstanceByID branches on that: when the id looks like a + # DbiResourceId it sends DescribeDBInstances with + # Filters=[dbi-resource-id] and leaves DBInstanceIdentifier nil. RDS + # authorizes an identifier-less DescribeDBInstances against the + # wildcard ARN arn:aws:rds:::db:*, which + # `db:cudly-*` cannot match, so every refresh of the instance 403'd: + # AccessDenied: ... not authorized to perform: rds:DescribeDBInstances + # on resource: arn:aws:rds:us-east-1:...:db:* + # This is unconditional on every read, not a not-found fallback: the + # provider's retry with the plain identifier only fires on NotFound, + # and AccessDenied is not NotFound, so the plan hard-fails. + # + # The DB proxy reads are here for a third reason: their resource ARNs + # use the `db-proxy:`/`target-group:` segments with server-assigned + # opaque ids (see the RDSResourceScoped note above), so no name-based + # ARN scope can ever match them. + # + # All five are read-only and expose only resource metadata; every + # mutating RDS action stays ARN-scoped above. + Sid = "RDSDescribeAccountWide" + Effect = "Allow" + Action = [ + "rds:DescribeDBEngineVersions", + "rds:DescribeDBInstances", + "rds:DescribeDBProxies", + "rds:DescribeDBProxyTargetGroups", + "rds:DescribeDBProxyTargets", + ] Resource = "*" }, { @@ -171,6 +291,20 @@ resource "aws_iam_policy" "data" { Resource = "*" }, { + # secretsmanager:ListSecrets used to be in this list. It has no + # resource type per the AWS Service Authorization Reference, so + # inside an ARN-scoped statement it could never be satisfied: a + # silently dead grant. Nothing in the provider calls it (secret + # lookups go through DescribeSecret with an explicit id), so it is + # dropped rather than moved to a "*" statement. + # + # ListSecretVersionIds and UpdateSecretVersionStage are what + # aws_secretsmanager_secret_version calls when a version carries a + # stage other than a lone AWSCURRENT, and on the delete path when it + # has to move stages off the version before removing it + # (resourceSecretVersionDelete/Update in + # internal/service/secretsmanager/secret_version.go). Both take the + # secret ARN, so the cudly-* scope below applies. Sid = "SecretsManager" Effect = "Allow" Action = [ @@ -179,13 +313,14 @@ resource "aws_iam_policy" "data" { "secretsmanager:DescribeSecret", "secretsmanager:GetResourcePolicy", "secretsmanager:GetSecretValue", - "secretsmanager:ListSecrets", + "secretsmanager:ListSecretVersionIds", "secretsmanager:PutSecretValue", "secretsmanager:RestoreSecret", "secretsmanager:RotateSecret", "secretsmanager:TagResource", "secretsmanager:UntagResource", "secretsmanager:UpdateSecret", + "secretsmanager:UpdateSecretVersionStage", ] Resource = "arn:aws:secretsmanager:*:*:secret:cudly-*" }, From b3edcab58f8e0746e8d5c8f8420bc2a2f4c0c0ae Mon Sep 17 00:00:00 2001 From: Cristian Magherusan-Stanciu Date: Mon, 3 Aug 2026 15:18:01 +0200 Subject: [PATCH 2/3] fix(iac/aws): close audited deploy-role gaps in compute and networking Second slice of the #1698 audit, covering the compute, container, edge and networking resources. Same three defect classes as the RDS fix: an action that is missing, an action gated on a condition that cannot hold, and an action whose grant can never match the request. Blocking on the deploy path: - ec2:CreateLaunchTemplateVersion, ec2:ModifyLaunchTemplate and ec2:DeleteLaunchTemplateVersions. The fck-nat launch template sources its AMI from a most_recent = true lookup, so every upstream AMI republish takes the launch template's update path rather than create. Only Create/DeleteLaunchTemplate were granted, so the first AMI rotation after the ASG exists would have failed the apply. - EC2RunInstancesFckNAT used StringEquals on ec2:InstanceType. AWS authorizes one RunInstances call against every resource type it touches, and ec2:InstanceType is only in the request context for the instance leg, so the condition evaluated false on the volume, network-interface, security-group, subnet, image and launch-template legs and denied the call as a whole. Switched to StringEqualsIfExists, which keeps the t4g.nano restriction exactly where the key exists. This is the operator the IAM condition-operators reference documents for this call. - iam:CreateServiceLinkedRole for autoscaling.amazonaws.com, elasticloadbalancing.amazonaws.com and ecs.amazonaws.com. Service-linked roles are created lazily on first use of a service in an account/region, so these are invisible where the role already exists and a hard failure in a fresh region. Note ecs.application-autoscaling was already granted and is a different principal that does not cover plain ecs. Latent, fixed while the file is open: - ec2:DeleteTags, the other half of ec2:CreateTags. The provider removes dropped keys before adding new ones, so dropping any key from common_tags or default_tags on an existing EC2 resource failed the apply. - ec2:ReplaceRouteTableAssociation, which the association's update path calls instead of Disassociate plus Associate. - ecr:PutImageScanningConfiguration, whose sibling PutImageTagMutability was already granted. - The elasticloadbalancing Set* family plus ModifyListenerAttributes. Changing a load balancer's subnets or security groups, which is what bumping az_count does, goes through Set* rather than Modify*. - CloudFrontMutateTaggedOnly was gated on a case-sensitive StringEquals against "CUDly" while the distribution's Project tag resolves to the lowercase "cudly" via local.common_tags, so all four of its actions were silently denied. This is the same defect #1496 fixed on the KMS statements, left unfixed here because enable_cdn is false in every tfvars so nothing exercised it. It would also have killed the destroy path, since deleting a distribution requires UpdateDistribution to disable it first. The ECR and ELB statements go in policy_compute_b.tf rather than policy_compute.tf, which is the file's stated purpose: policy_compute is at 5933 of 6144 characters and would not hold them. Sizes after this change: compute 5933, compute-b 1290, data 4588, networking 2831. Refs #1698 --- .../aws/ci-cd-permissions/policy_compute.tf | 15 ++++++- .../aws/ci-cd-permissions/policy_compute_b.tf | 32 ++++++++++++++ .../aws/ci-cd-permissions/policy_data.tf | 15 +++++++ .../ci-cd-permissions/policy_networking.tf | 44 ++++++++++++++++--- 4 files changed, 100 insertions(+), 6 deletions(-) diff --git a/terraform/environments/aws/ci-cd-permissions/policy_compute.tf b/terraform/environments/aws/ci-cd-permissions/policy_compute.tf index e7d2b793d..55d1554c0 100644 --- a/terraform/environments/aws/ci-cd-permissions/policy_compute.tf +++ b/terraform/environments/aws/ci-cd-permissions/policy_compute.tf @@ -152,6 +152,19 @@ resource "aws_iam_policy" "compute" { # Function mutations live in policy_compute_b.tf — the function # resource type does not support aws:ResourceTag per the AWS # Service Authorization Reference, so it needs ARN scoping. + # + # StringEqualsIgnoreCase for the same reason as KMSMutateTaggedOnly + # below (see PR #1496): the distribution's Project tag comes from + # local.common_tags (terraform/environments/aws/main.tf), which sets + # Project = var.project_name, and project_name is the lowercase + # "cudly". That resource-level tag overrides the provider's + # default_tags value of "CUDly" for this distribution, so a + # case-sensitive match against "CUDly" never matched and every action + # in this statement was silently denied. This is the same defect + # #1496 fixed on the KMS statements; it was left unfixed here because + # enable_cdn is false in every tfvars, so nothing exercised it. Note + # it would also have killed the destroy path: deleting a distribution + # requires UpdateDistribution to disable it first. Sid = "CloudFrontMutateTaggedOnly" Effect = "Allow" Action = [ @@ -162,7 +175,7 @@ resource "aws_iam_policy" "compute" { ] Resource = "*" Condition = { - StringEquals = { + StringEqualsIgnoreCase = { "aws:ResourceTag/Project" = "CUDly" } } diff --git a/terraform/environments/aws/ci-cd-permissions/policy_compute_b.tf b/terraform/environments/aws/ci-cd-permissions/policy_compute_b.tf index eb213010a..278533932 100644 --- a/terraform/environments/aws/ci-cd-permissions/policy_compute_b.tf +++ b/terraform/environments/aws/ci-cd-permissions/policy_compute_b.tf @@ -97,6 +97,38 @@ resource "aws_iam_policy" "compute_b" { } } }, + { + # ecr:PutImageScanningConfiguration is the writer for + # aws_ecr_repository's image_scanning_configuration block. Its sibling + # ecr:PutImageTagMutability is granted in policy_compute.tf's + # ECRRepositoryScoped but this one was not, so any change to + # scan_on_push (or importing a repository whose setting differs from + # the config) fails the apply. Same repository ARN scope as the + # statement it belongs with; it lives here only because + # policy_compute.tf is at the 6144-char ceiling. + Sid = "ECRImageScanningConfig" + Effect = "Allow" + Action = ["ecr:PutImageScanningConfiguration"] + Resource = "arn:aws:ecr:*:*:repository/cudly-*" + }, + { + # The ELB update path that policy_compute.tf's ELBFargate misses. + # ELBFargate grants Create/Delete plus the Modify* family, but changing + # a load balancer's subnets or security groups (which is what bumping + # az_count does) goes through the Set* family instead, and + # ModifyListenerAttributes is the writer whose reader + # (DescribeListenerAttributes) is already granted. Resource = "*" + # matches ELBFargate, whose ARNs are not name-scopeable. + Sid = "ELBFargateSetAttributes" + Effect = "Allow" + Action = [ + "elasticloadbalancing:ModifyListenerAttributes", + "elasticloadbalancing:SetIpAddressType", + "elasticloadbalancing:SetSecurityGroups", + "elasticloadbalancing:SetSubnets", + ] + Resource = "*" + }, { Sid = "KMSAliasMutate" Effect = "Allow" diff --git a/terraform/environments/aws/ci-cd-permissions/policy_data.tf b/terraform/environments/aws/ci-cd-permissions/policy_data.tf index c8c0f4017..721a788a5 100644 --- a/terraform/environments/aws/ci-cd-permissions/policy_data.tf +++ b/terraform/environments/aws/ci-cd-permissions/policy_data.tf @@ -168,18 +168,33 @@ resource "aws_iam_policy" "data" { Resource = "*" }, { + # Service-linked roles are created lazily by AWS on the FIRST use of a + # service in an account/region, and the caller needs + # iam:CreateServiceLinkedRole for it. That makes these gaps invisible + # in an account that already has the role and a hard failure in a + # fresh region, which is exactly the kind of latent break this audit + # was meant to surface. autoscaling (fck-nat ASG, on every deploy + # path), elasticloadbalancing and ecs (Fargate ALB and cluster) were + # all missing; note that ecs.application-autoscaling is a different + # principal from plain ecs and does not cover it. Sid = "IAMServiceLinkedRole" Effect = "Allow" Action = ["iam:CreateServiceLinkedRole"] Resource = [ "arn:aws:iam::*:role/aws-service-role/rds.amazonaws.com/AWSServiceRoleForRDS", "arn:aws:iam::*:role/aws-service-role/ecs.application-autoscaling.amazonaws.com/AWSServiceRoleForApplicationAutoScaling_ECSService", + "arn:aws:iam::*:role/aws-service-role/autoscaling.amazonaws.com/AWSServiceRoleForAutoScaling", + "arn:aws:iam::*:role/aws-service-role/elasticloadbalancing.amazonaws.com/AWSServiceRoleForElasticLoadBalancing", + "arn:aws:iam::*:role/aws-service-role/ecs.amazonaws.com/AWSServiceRoleForECS", ] Condition = { StringLike = { "iam:AWSServiceName" = [ "rds.amazonaws.com", "ecs.application-autoscaling.amazonaws.com", + "autoscaling.amazonaws.com", + "elasticloadbalancing.amazonaws.com", + "ecs.amazonaws.com", ] } } diff --git a/terraform/environments/aws/ci-cd-permissions/policy_networking.tf b/terraform/environments/aws/ci-cd-permissions/policy_networking.tf index 3336d451c..0ff3ec119 100644 --- a/terraform/environments/aws/ci-cd-permissions/policy_networking.tf +++ b/terraform/environments/aws/ci-cd-permissions/policy_networking.tf @@ -22,6 +22,17 @@ resource "aws_iam_policy" "networking" { "ec2:CreateEgressOnlyInternetGateway", "ec2:CreateInternetGateway", "ec2:CreateLaunchTemplate", + # The fck-nat launch template sources its AMI from a + # most_recent = true data.aws_ami lookup, so every upstream AMI + # republish takes the launch template's UPDATE path rather than + # create: the provider calls CreateLaunchTemplateVersion for the new + # image_id and ModifyLaunchTemplate to move the default version + # (resourceLaunchTemplateUpdate). Only Create/DeleteLaunchTemplate + # were granted, so the first AMI rotation after the ASG exists would + # have failed the apply. + "ec2:CreateLaunchTemplateVersion", + "ec2:DeleteLaunchTemplateVersions", + "ec2:ModifyLaunchTemplate", "ec2:CreateRoute", "ec2:CreateRouteTable", "ec2:CreateSecurityGroup", @@ -39,6 +50,12 @@ resource "aws_iam_policy" "networking" { "ec2:DeleteRouteTable", "ec2:DeleteSecurityGroup", "ec2:DeleteSubnet", + # ec2:DeleteTags is the other half of ec2:CreateTags: the provider's + # tag update path removes dropped keys with DeleteTags before adding + # new ones with CreateTags, so dropping any key from common_tags / + # default_tags on an existing VPC, subnet, route table, gateway, + # security group or launch template fails the apply without it. + "ec2:DeleteTags", "ec2:DeleteVpc", "ec2:DeleteVpcEndpoints", "ec2:DescribeAccountAttributes", @@ -66,6 +83,11 @@ resource "aws_iam_policy" "networking" { "ec2:ModifyVpcAttribute", "ec2:ModifyVpcEndpoint", "ec2:ReplaceRoute", + # aws_route_table_association's update path calls + # ReplaceRouteTableAssociation rather than + # Disassociate + Associate, so moving a subnet between route tables + # needs it even though both halves of that pair are granted. + "ec2:ReplaceRouteTableAssociation", "ec2:RevokeSecurityGroupEgress", "ec2:RevokeSecurityGroupIngress", ] @@ -73,16 +95,28 @@ resource "aws_iam_policy" "networking" { }, { # ec2:RunInstances is needed for the fck-nat AutoScaling group - # (one t4g.nano per AZ). Scoping to the fck-nat launch template - # and restricting allowed instance types prevents the deploy SA - # from launching arbitrary large instances and attaching CUDly IAM - # roles to exfiltrate credentials or run compute at account cost. + # (one t4g.nano per AZ). Restricting allowed instance types prevents + # the deploy SA from launching arbitrary large instances and attaching + # CUDly IAM roles to exfiltrate credentials or run compute at account + # cost. + # + # The operator MUST be StringEqualsIfExists, not StringEquals. AWS + # authorizes a single RunInstances call against every resource type it + # touches (instance, volume, network-interface, security-group, + # subnet, image, key-pair, launch-template), and ec2:InstanceType is + # only in the request context for the INSTANCE leg. With a plain + # StringEquals the key is absent on every other leg, the condition + # evaluates false there, and the call is denied as a whole even though + # the instance type is correct. StringEqualsIfExists keeps the t4g.nano + # restriction exactly where the key exists and lets the supporting legs + # through, which is the pattern the IAM condition-operators reference + # documents for precisely this call. Sid = "EC2RunInstancesFckNAT" Effect = "Allow" Action = ["ec2:RunInstances"] Resource = "*" Condition = { - StringEquals = { + StringEqualsIfExists = { "ec2:InstanceType" = ["t4g.nano"] } } From 8ac4a6dc52c60a1cdb0db3c1986638fcac39f9ac Mon Sep 17 00:00:00 2001 From: Cristian Magherusan-Stanciu Date: Mon, 3 Aug 2026 16:12:42 +0200 Subject: [PATCH 3/3] fix(iac/aws): move rds:DescribeDBSubnetGroups to the account-wide read statement rds:DescribeDBSubnetGroups was the only rds:Describe* left in the ARN-scoped RDSResourceScoped statement, and it is in exactly the same position as the rds:DescribeDBInstances read that has been failing the deploy: aws_db_subnet_group.main (terraform/modules/database/aws/main.tf:73) is created in every environment, and a resource-scoped grant on a Describe authorizes only the form that targets one named resource, while the enumerate form carries no identifier and is authorized against the wildcard ARN. Worth recording why this is not the simpler claim that ARN-scoped Describes are always dead, because that claim is false and would have justified the wrong fix elsewhere. AWS's Service Authorization Reference confirms both actions DO support resource types (db and subgrp), and the real names DO match the patterns (cudly---postgres against db:cudly-*, cudly---db-subnet against subgrp:cudly-*). The scoped grants were syntactically valid; only the request shape differs. The production 403 on DescribeDBInstances is the empirical proof that the scoped form did not suffice. Also records two facts established during review, so they are not rediscovered: - enable_rds_proxy is a literal false at terraform/environments/aws/database.tf:28, not a variable, so no tfvars can override it. All three proxy resources are count = 0 in every environment, which makes the proxy Describes here belt-and-braces rather than load-bearing. - The RDSResourceScoped comment now says no rds:Describe* lives there at all, rather than naming only DescribeDBInstances. Policy size is unchanged at 4588 of 6144 characters: the action moved between statements rather than being added. Refs #1698 --- .../aws/ci-cd-permissions/policy_data.tf | 32 ++++++++++++++++--- 1 file changed, 27 insertions(+), 5 deletions(-) diff --git a/terraform/environments/aws/ci-cd-permissions/policy_data.tf b/terraform/environments/aws/ci-cd-permissions/policy_data.tf index 721a788a5..22996efb3 100644 --- a/terraform/environments/aws/ci-cd-permissions/policy_data.tf +++ b/terraform/environments/aws/ci-cd-permissions/policy_data.tf @@ -205,8 +205,8 @@ resource "aws_iam_policy" "data" { # deploy SA from deleting or modifying unrelated RDS resources in the # same account. # - # rds:DescribeDBInstances deliberately does NOT live here; see - # RDSDescribeAccountWide below. + # NO rds:Describe* action lives here; they are all in + # RDSDescribeAccountWide below. See that statement for why. # # KNOWN DEAD SCOPES: the proxy actions below (DeleteDBProxy, # ModifyDBProxy, RegisterDBProxyTargets, DeregisterDBProxyTargets) @@ -232,7 +232,6 @@ resource "aws_iam_policy" "data" { "rds:CreateDBSubnetGroup", "rds:DeleteDBInstance", "rds:DeleteDBSubnetGroup", - "rds:DescribeDBSubnetGroups", "rds:ListTagsForResource", "rds:DeleteDBProxy", "rds:DeregisterDBProxyTargets", @@ -272,12 +271,34 @@ resource "aws_iam_policy" "data" { # provider's retry with the plain identifier only fires on NotFound, # and AccessDenied is not NotFound, so the plan hard-fails. # + # rds:DescribeDBSubnetGroups is here for the same reason as + # DescribeDBInstances, and it is worth being explicit about why, + # because the tempting simplification is wrong in both directions. + # Neither action lacks a resource type: per AWS's Service + # Authorization Reference both DO support one (`db` and `subgrp` + # respectively), and the real names DO match the patterns + # (cudly---postgres against db:cudly-*, + # cudly---db-subnet against subgrp:cudly-*). So this is not + # "ARN-scoped Describes are always dead" and it is not a name-prefix + # bug; the scoped grants were syntactically valid. What differs is the + # request shape: a resource-scoped grant authorizes only the form that + # targets one named resource, while the enumerate form (no identifier + # in the request) is authorized against the wildcard ARN. The + # production 403 on DescribeDBInstances is the empirical proof that + # the scoped form did not suffice, and DescribeDBSubnetGroups sits in + # exactly the same position: aws_db_subnet_group.main + # (terraform/modules/database/aws/main.tf) is created in every + # environment, so leaving it scoped just primes the next 403. + # # The DB proxy reads are here for a third reason: their resource ARNs # use the `db-proxy:`/`target-group:` segments with server-assigned # opaque ids (see the RDSResourceScoped note above), so no name-based - # ARN scope can ever match them. + # ARN scope can ever match them. Those are belt-and-braces rather than + # load-bearing: enable_rds_proxy is a literal `false` at + # terraform/environments/aws/database.tf:28, not a variable, so no + # tfvars can turn the proxy resources on. # - # All five are read-only and expose only resource metadata; every + # All six are read-only and expose only resource metadata; every # mutating RDS action stays ARN-scoped above. Sid = "RDSDescribeAccountWide" Effect = "Allow" @@ -287,6 +308,7 @@ resource "aws_iam_policy" "data" { "rds:DescribeDBProxies", "rds:DescribeDBProxyTargetGroups", "rds:DescribeDBProxyTargets", + "rds:DescribeDBSubnetGroups", ] Resource = "*" },