Skip to content

sec(iac/aws): deploy boundary still permits escaping to another principal via lambda/ssm/ecs at service:* #1723

Description

@cristim

cudly-deploy-boundary (added by #1722) caps the deploy role's IAM writes, which closes the #1705 escalation. But the ceiling itself still grants several services at service:* on Resource: "*", and at least three of those let a principal execute code under a different principal, escaping the boundary entirely.

A permissions boundary constrains the principal it is attached to. It does not follow a different principal. So any action that runs code as something else is an escape hatch, regardless of what the boundary says — and critically, these need no iam:PassRole, so PassRoleCeiling's cudly-* scoping does not apply.

Service Escape Runs as
lambda:* UpdateFunctionCode on any function, then invoke that function's execution role
ssm:* SendCommand / StartSession to any SSM-managed instance that instance's profile
ecs:* UpdateService onto an existing task-def revision, or ExecuteCommand into a running task that task's role

This is the same criterion #1722 already applies, correctly, to organizations and sts. policy_boundary.tf:100-116 states those are "the two services that are NOT safe at service:* granularity". The criterion is right; the enumeration is incomplete.

Precondition: a non-boundaried compute principal in the account holding more than the ceiling. That is account-state-dependent — the same class of precondition as the OrganizationAccountAccessRole reasoning that was (correctly) treated as decisive for sts.

The gap is wide relative to need

The modules actually grant only:

  • lambda:InvokeFunction, lambda:InvokeFunctionUrl, lambda:GetFunctionUrlConfig
  • ecs:RunTask
  • three kms: verbs
  • a handful of ec2: RI/ENI calls
  • zero s3: actions

So the ceiling is far broader than anything the deploy path uses.

Why this was deliberately not fixed in #1722

Correct call, and it should stay that way. A too-narrow boundary fails at runtime, in production, long after a green apply — whereas everything #1722 can get wrong fails at apply time in CI. Landing the escalation fix first and narrowing separately is the right risk ordering.

Suggested approach

Narrow per service to the verbs the modules actually declare, using the same derive-from-source method the guard tests already use, and extend policy_guard_test.go so a module granting a verb absent from the ceiling fails CI rather than 403ing at runtime. That guard already exists and caught this class in review — mutation testing showed it fails when a module grants a service the ceiling lacks.

Do this as its own change with its own verification, and verify against live account state first: whether any non-boundaried Lambda, SSM-managed instance or ECS service actually exists determines the real severity here. That check was out of scope for #1722's review (no credentials).

Also worth recording

iam:AddRoleToInstanceProfile supports no condition keys and remains unconditioned. Not exploitable today, because it needs PassRole and that is scoped — but its precondition is identical to the hand-made-cudly-*-role residual #1722 does admit, so it belongs in the same list.

Found during the adversarial review of #1722.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions