Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 17 additions & 6 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -753,15 +753,26 @@ jobs:
- name: Run guard script self-tests
run: bash scripts/test-gcp-secret-scope.sh

# Assert that the ECR repository selector used by destroy-fargate-dev.yml
# picks the repository that state owns and nothing else. The consumer
# force-deletes what the selector prints, so both directions are asserted:
# over-matching deletes images the workflow does not own, and matching
# nothing leaves the dev repository behind. Fast (shell only), so it always
# runs.
# Assert that the ECR repository selector used by destroy-fargate-dev.yml and
# cleanup-staging.yml picks the repository each state owns and nothing else.
# The consumers force-delete what the selector prints, so both directions are
# asserted: over-matching deletes images the workflow does not own, and
# matching nothing leaves that state's repository behind. The suite also
# asserts the wiring, which is what #1592 and #1820 each escaped: all three
# destroy steps still call scripts/force-delete-owned-ecr-repo.sh, that script
# still deletes only what the selector yields, and nothing else under
# .github/workflows or scripts/ runs `aws ecr delete-repository` unguarded.
# Fast (shell only), so it always runs.
ecr-delete-selection:
name: ECR delete selection scope
runs-on: ubuntu-latest
# ci.yml declares no workflow-level `permissions`, so a job without its own
# block gets the repository default, which is read/write on this repo. This
# job checks out the tree and runs a shell script against it; `contents:
# read` is all of that needs, and it is the same shape security-scan above
# uses (which adds `security-events: write` only because it uploads SARIF).
permissions:
contents: read

steps:
- name: Checkout code
Expand Down
62 changes: 24 additions & 38 deletions .github/workflows/cleanup-staging.yml
Original file line number Diff line number Diff line change
Expand Up @@ -138,25 +138,18 @@ jobs:
cd terraform/environments/aws
terraform init -backend-config=/tmp/backend.tfbackend

- name: Force-delete all staging ECR repos
run: |
for REPO in $(aws ecr describe-repositories \
--query "repositories[?starts_with(repositoryName,'cudly-staging')].repositoryName" \
--output text 2>/dev/null); do
# Defense in depth: refuse anything outside the staging naming
# pattern (in particular the bare 'cudly' prod-adjacent repo)
# even if the query filter above is ever loosened.
case "$REPO" in
cudly-staging*) ;;
*)
echo "Refusing to delete non-staging ECR repo '$REPO'; staging cleanup only deletes cudly-staging* repos"
exit 1
;;
esac
echo "Force-deleting ECR repo $REPO..."
aws ecr delete-repository --repository-name "$REPO" --force 2>/dev/null \
|| echo " Failed to delete $REPO (may already be gone)"
done
# Runs before `terraform destroy`: the repository is created with
# force_delete = false, so the destroy fails while images remain. Deletes
# only the repository THIS state owns, by exact name. The `cudly-staging*`
# prefix this used to select by also matched `cudly-staging-prod-mirror`,
# `cudly-staging-<hex>-backup` and the sibling staging state's repository,
# and force-deleted every image in them (#1820). Rationale, the
# `output -json` handling and why nothing here is swallowed: the script's
# header. Both staging jobs and destroy-fargate-dev.yml call the same
# script, so the guard cannot land in one workflow and not its sibling --
# which is how #1592 became #1820.
- name: Force-delete the ECR repo this state owns
run: ./scripts/force-delete-owned-ecr-repo.sh terraform/environments/aws

- name: Disable RDS deletion protection before destroy
run: |
Expand Down Expand Up @@ -223,25 +216,18 @@ jobs:
cd terraform/environments/aws
terraform init -backend-config=/tmp/backend.tfbackend

- name: Force-delete all staging ECR repos
run: |
for REPO in $(aws ecr describe-repositories \
--query "repositories[?starts_with(repositoryName,'cudly-staging')].repositoryName" \
--output text 2>/dev/null); do
# Defense in depth: refuse anything outside the staging naming
# pattern (in particular the bare 'cudly' prod-adjacent repo)
# even if the query filter above is ever loosened.
case "$REPO" in
cudly-staging*) ;;
*)
echo "Refusing to delete non-staging ECR repo '$REPO'; staging cleanup only deletes cudly-staging* repos"
exit 1
;;
esac
echo "Force-deleting ECR repo $REPO..."
aws ecr delete-repository --repository-name "$REPO" --force 2>/dev/null \
|| echo " Failed to delete $REPO (may already be gone)"
done
# Runs before `terraform destroy`: the repository is created with
# force_delete = false, so the destroy fails while images remain. Deletes
# only the repository THIS state owns, by exact name. The `cudly-staging*`
# prefix this used to select by also matched `cudly-staging-prod-mirror`,
# `cudly-staging-<hex>-backup` and the sibling staging state's repository,
# and force-deleted every image in them (#1820). Rationale, the
# `output -json` handling and why nothing here is swallowed: the script's
# header. Both staging jobs and destroy-fargate-dev.yml call the same
# script, so the guard cannot land in one workflow and not its sibling --
# which is how #1592 became #1820.
- name: Force-delete the ECR repo this state owns
run: ./scripts/force-delete-owned-ecr-repo.sh terraform/environments/aws

- name: Disable RDS deletion protection before destroy
run: |
Expand Down
47 changes: 11 additions & 36 deletions .github/workflows/destroy-fargate-dev.yml
Original file line number Diff line number Diff line change
Expand Up @@ -128,43 +128,18 @@ jobs:
cd terraform/environments/aws
terraform init -backend-config=/tmp/backend.tfbackend

# Runs before `terraform destroy` because the repository is created with
# force_delete = false (terraform/modules/registry/aws/main.tf), so the
# destroy fails while images remain.
#
# Deletes only the repository this state owns, read from `terraform
# output` and compared by exact equality against every repository in the
# account. The `contains(repositoryName,'cudly-dev')` filter this step
# used to run also force-deleted `backup-cudly-dev` and
# `cudly-dev-prod-mirror`, with every image in them (#1592).
# scripts/select-ecr-repos-to-delete.sh holds the comparison and the name
# table pinning it, including why a `cudly-dev*` prefix is not enough.
# Reads the name through `output -json`, not `output -raw`: on a state
# with no outputs at all, `output -raw <name>` exits 0 and writes its "No
# outputs found" warning to STDOUT, so the name became that warning text
# and re-dispatching the destroy after a completed one failed here before
# `terraform destroy` ever ran. `output -json` returns `{}` for that state
# and the two cases separate cleanly:
# no outputs at all -> already destroyed, skip the ECR cleanup
# outputs but not this one -> `jq -e` exits 1, the step fails loudly
# `exit 0` ends this step only; the steps after it still run.
# Runs before `terraform destroy`: the repository is created with
# force_delete = false, so the destroy fails while images remain. Deletes
# only the repository THIS state owns, by exact name. The
# `contains(repositoryName,'cudly-dev')` filter this step used to run also
# force-deleted `backup-cudly-dev` and `cudly-dev-prod-mirror`, with every
# image in them (#1592). Rationale, the `output -json` handling and why
# nothing here is swallowed: the script's header. This job and both
# cleanup-staging.yml staging jobs call the same script, so the guard
# cannot land in one workflow and not its sibling -- which is how #1592
# became #1820.
- name: Force-delete ECR repo
run: |
set -euo pipefail
OUTPUTS_JSON="$(terraform -chdir=terraform/environments/aws output -json)"
if [ "$(jq -r 'length' <<<"$OUTPUTS_JSON")" -eq 0 ]; then
echo "State has no outputs; the stack is already destroyed and there is no ECR repository to clean up."
exit 0
fi
OWNED_REPO="$(jq -er '.ecr_repository_name.value' <<<"$OUTPUTS_JSON")"
echo "This state owns ECR repository '$OWNED_REPO'"
aws ecr describe-repositories --query 'repositories[].repositoryName' --output text \
| tr '\t' '\n' \
| ./scripts/select-ecr-repos-to-delete.sh "$OWNED_REPO" \
| while IFS= read -r REPO; do
echo "Force-deleting ECR repo $REPO..."
aws ecr delete-repository --repository-name "$REPO" --force
done
run: ./scripts/force-delete-owned-ecr-repo.sh terraform/environments/aws

- name: Disable RDS deletion protection
run: |
Expand Down
84 changes: 84 additions & 0 deletions scripts/force-delete-owned-ecr-repo.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
#!/usr/bin/env bash
# force-delete-owned-ecr-repo.sh
#
# Force-deletes the one ECR repository a Terraform state owns, and nothing else.
#
# Usage: force-delete-owned-ecr-repo.sh TERRAFORM_STATE_DIR
#
# Run before `terraform destroy` on a state that created an ECR repository: the
# repository is created with force_delete = false
# (terraform/modules/registry/aws/main.tf), so the destroy fails while images
# remain in it.
#
# The state directory is a required argument rather than a constant because it
# is the identity of what gets deleted. Every caller happens to pass
# terraform/environments/aws today, but which state the name is read from is
# the whole safety property here, so it stays visible at each call site.
#
# The owned name is read from `terraform output` on the state the destroy is
# about to tear down and compared by exact equality against every repository in
# the account, by scripts/select-ecr-repos-to-delete.sh. The callers used to
# select by the `cudly-dev*` / `cudly-staging*` prefix, which also matches
# `cudly-staging-prod-mirror` and `cudly-staging-<hex>-backup` and force-deleted
# every image in them (#1592, #1820). The prefix also spans both staging states:
# cleanup-staging.yml's lambda and fargate jobs each create their own
# `cudly-staging-<random_id.suffix.hex>` repository (main.tf:55), so either job
# deleted the other's. The selector script holds the comparison and the name
# table pinning it, including why no prefix describes the owned repository
# uniquely.
#
# Reads the name through `output -json`, not `output -raw`: on a state with no
# outputs at all, `output -raw <name>` exits 0 and writes its "No outputs found"
# warning to STDOUT, so the name would become that warning text and re-running
# the cleanup after a completed one would fail here before `terraform destroy`
# ever ran. `output -json` returns `{}` for that state and the two cases
# separate cleanly:
# no outputs at all -> already destroyed, skip the ECR cleanup
# outputs but not this one -> `jq -e` exits 1, this script fails loudly
#
# Nothing is swallowed. The `2>/dev/null || echo "may already be gone"` the
# callers used to carry reported success after a failed listing or a failed
# delete, and a failed listing is indistinguishable from an empty account, so
# the cleanup did nothing and `terraform destroy` then failed on the images
# still in the repository. "Already gone" needs no swallowing: the repository is
# simply absent from the listing, the selector prints nothing and exits 0, and
# the loop body never runs.
#
# Exit codes:
# 0 cleanup completed, including the "state already destroyed" and "nothing
# to delete" cases, which are normal outcomes and not errors
# 2 usage error (wrong arity, or a state directory that does not exist)
# * anything the AWS CLI, terraform, jq or the selector fails with, unmasked

set -euo pipefail

SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"

if [[ $# -ne 1 ]]; then
echo "usage: $(basename "$0") TERRAFORM_STATE_DIR" >&2
exit 2
fi

STATE_DIR="$1"

if [[ ! -d "$STATE_DIR" ]]; then
echo "error: terraform state directory '${STATE_DIR}' does not exist" >&2
exit 2
fi

OUTPUTS_JSON="$(terraform -chdir="$STATE_DIR" output -json)"
if [[ "$(jq -r 'length' <<<"$OUTPUTS_JSON")" -eq 0 ]]; then
echo "State has no outputs; the stack is already destroyed and there is no ECR repository to clean up."
exit 0
fi

OWNED_REPO="$(jq -er '.ecr_repository_name.value' <<<"$OUTPUTS_JSON")"
echo "This state owns ECR repository '$OWNED_REPO'"

aws ecr describe-repositories --query 'repositories[].repositoryName' --output text \
| tr '\t' '\n' \
| "${SCRIPT_DIR}/select-ecr-repos-to-delete.sh" "$OWNED_REPO" \
| while IFS= read -r REPO; do
echo "Force-deleting ECR repo $REPO..."
aws ecr delete-repository --repository-name "$REPO" --force
done
Loading
Loading