Skip to content

sec(ci): lambda deploy has same unsafe failure-path state-lock delete as #438 #1269

Description

@cristim

Problem

.github/workflows/deploy-aws-lambda.yml (around lines 190-199) still has an unconditional failure-path state-lock release:

- name: Release state lock on failure
  if: failure() || cancelled()
  run: |
    ...
    aws s3 rm "s3://${BUCKET}/${LOCK_KEY}"

This is the same state-corruption race that #438 closed for deploy-aws-fargate.yml: a run that fails because it could not acquire the lock will execute this cleanup and delete the lock held by another run that is still actively applying, enabling parallel state writes.

The Fargate workflow now relies solely on the operator-triggered clear_stale_lock input (PR #818). The Lambda workflow should adopt the same approach.

Fix

Remove the automatic Release state lock on failure step from deploy-aws-lambda.yml and add an operator-triggered clear_stale_lock workflow_dispatch boolean input (mirroring the Fargate workflow). Terraform already releases its own lock on a clean apply error; a surviving lock means the run died abnormally and needs operator confirmation before clearing. Reference runbooks/terraform-stuck-lock.md.

Out of scope for #818

PR #818 is scoped to the Fargate workflow only; this is the parallel fix for the Lambda deploy path.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions