From 3b73caa29a22c7b42aedb7596865c40cc9e79124 Mon Sep 17 00:00:00 2001 From: Cristian Magherusan-Stanciu Date: Wed, 12 Aug 2026 23:16:27 +0200 Subject: [PATCH] fix(ci): resolve deploy environment from inputs, not github.event_name Inside a reusable workflow `github.event_name` is the caller's event and is never the literal `workflow_call`, so the `workflow_dispatch|workflow_call` arm in the called deploy workflows is dead code. A release-triggered deploy-all.yml computes environment=prod and passes it to every cloud; each callee then falls through to the `*)` arm and returns dev. Azure and GCP would apply Terraform against dev while the run summary reported prod. Nothing has been mis-deployed yet: deploy-all.yml has never run. The workflow_dispatch path survives only by coincidence, the caller's event there genuinely being workflow_dispatch. Branch on whether an environment was requested instead. `inputs` is populated on workflow_dispatch and workflow_call, the only two triggers that carry an environment, so a non-empty input identifies a workflow_call invocation that no event name can. Where an empty input has no defensible default the step now fails rather than guessing a deployment target. deploy-aws-fargate.yml declares no trigger besides workflow_dispatch and workflow_call, so it loses its fallback outright. deploy-aws-lambda.yml already produced the right answer but hardcoded it in a `release) prod` arm instead of honouring the input; the input now wins and that arm serves only its own direct release trigger. deploy-all.yml is always the entry point of its own run, so its use of github.event_name is sound and stays. It gains the dev|staging|prod allowlist the callees already had, so an empty or unknown environment stops the fan-out rather than reaching four deployments, and it no longer interpolates inputs.environment straight into a shell command. Because that allowlist makes determine-deployment the first job that can deliberately fail, the notify job now treats any non-success there as a failure: the four deploy jobs are skipped rather than failed in that case, so the summary would otherwise announce that all deployments completed successfully for a run that deployed nothing. The comment on the deploy-azure job, which documented that the callee's prepare output need not agree with the environment computed above, was true only because of this bug and is corrected here. Closes #1805 --- .github/workflows/deploy-all.yml | 62 +++++++++++++++++++++--- .github/workflows/deploy-aws-fargate.yml | 17 ++++--- .github/workflows/deploy-aws-lambda.yml | 38 ++++++++++----- .github/workflows/deploy-azure.yml | 30 ++++++++++-- .github/workflows/deploy-gcp.yml | 30 ++++++++++-- 5 files changed, 142 insertions(+), 35 deletions(-) diff --git a/.github/workflows/deploy-all.yml b/.github/workflows/deploy-all.yml index bf79df9ac..6f9a8a8a6 100644 --- a/.github/workflows/deploy-all.yml +++ b/.github/workflows/deploy-all.yml @@ -15,6 +15,15 @@ name: Deploy to All Clouds +# Adding a trigger here, or a second caller of the deploy-* workflows, means +# revisiting the `prepare` comments in deploy-azure.yml, deploy-gcp.yml and +# deploy-aws-lambda.yml. Since #1805 those three resolve their environment from +# `inputs.environment` and fall back on `github.event_name`, which inside a +# reusable workflow is whatever event started THIS run. They are correct today +# only because this workflow is their only caller, has no `push` trigger, and +# allowlists the environment it passes. A `push:` trigger here would make +# Azure and GCP silently resolve dev for a caller that asked for something +# else, which is #1805 restored. on: workflow_dispatch: inputs: @@ -56,13 +65,36 @@ jobs: steps: - name: Set environment id: set-env + env: + # This workflow is always the entry point of its own run, never a + # reusable one, so `github.event_name` here really is the trigger that + # started it. The called workflows cannot rely on that and read + # `inputs.environment` instead (issue #1805). + EVENT_NAME: ${{ github.event_name }} + INPUT_ENVIRONMENT: ${{ inputs.environment }} run: | - if [[ "${{ github.event_name }}" == "release" ]]; then - echo "environment=prod" >> $GITHUB_OUTPUT + set -euo pipefail + INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}" + + if [ "$EVENT_NAME" = release ]; then + ENVIRONMENT=prod else - echo "environment=${{ inputs.environment }}" >> $GITHUB_OUTPUT + ENVIRONMENT="$INPUT_ENVIRONMENT" fi + # Every called workflow takes this value verbatim, so an empty or + # unknown value has to stop the fan-out here rather than reach four + # deployments. + case "$ENVIRONMENT" in + dev|staging|prod) ;; + *) + echo "::error::Refusing unknown environment: '$ENVIRONMENT'" + exit 1 + ;; + esac + + echo "environment=$ENVIRONMENT" >> "$GITHUB_OUTPUT" + - name: Set cloud providers id: set-clouds run: | @@ -166,8 +198,9 @@ jobs: # Deliberately carries no `concurrency` (#1801). The Terraform state guard is # the group on the called workflow's own build-and-deploy job, which runs as a # real job of this run and is serialized there. That group is derived from the - # called workflow's own `prepare` output, which is not necessarily the - # `environment` computed above, so do not assume the two agree. + # called workflow's own `prepare` output, which since #1805 is the + # `environment` computed above verbatim: `prepare` takes the non-empty input + # it is passed and fails the run rather than substituting a default. # # Putting that same group on this caller job would deadlock: this job would # hold the group while waiting on the inner job queued behind it. GitHub @@ -263,7 +296,22 @@ jobs: run: | FAILED=false - if [[ "${{ needs.deploy-aws-lambda.result }}" == "failure" ]] || \ + # determine-deployment is checked for anything other than success, not + # just "failure". When it stops the run the four deploy jobs are + # SKIPPED rather than failed, so testing only those would report "all + # deployments completed successfully" for a run that deployed nothing. + # + # The four deploy jobs are deliberately still tested for "failure" + # only. A CANCELLED run leaves them `cancelled` and reaches the + # success line, which is the same false report. That predates this + # change, so it is tracked in #1810 rather than fixed here. + # + # Note this step is the only success-by-default check in the repo. All + # four per-cloud `summary` jobs allowlist on `== "success"` instead, + # so an unenumerated result falls to their failure branch and errs + # safe. This denylist is the outlier, not the pattern. + if [[ "${{ needs.determine-deployment.result }}" != "success" ]] || \ + [[ "${{ needs.deploy-aws-lambda.result }}" == "failure" ]] || \ [[ "${{ needs.deploy-aws-fargate.result }}" == "failure" ]] || \ [[ "${{ needs.deploy-gcp.result }}" == "failure" ]] || \ [[ "${{ needs.deploy-azure.result }}" == "failure" ]]; then @@ -272,7 +320,7 @@ jobs: if [ "$FAILED" = true ]; then echo "" >> $GITHUB_STEP_SUMMARY - echo "❌ **One or more deployments failed. Check individual job logs.**" >> $GITHUB_STEP_SUMMARY + echo "❌ **Deployment did not complete successfully. Check individual job logs.**" >> $GITHUB_STEP_SUMMARY exit 1 else echo "" >> $GITHUB_STEP_SUMMARY diff --git a/.github/workflows/deploy-aws-fargate.yml b/.github/workflows/deploy-aws-fargate.yml index d4c6343c3..bcee839ea 100644 --- a/.github/workflows/deploy-aws-fargate.yml +++ b/.github/workflows/deploy-aws-fargate.yml @@ -74,16 +74,21 @@ jobs: - name: Determine environment id: set-env env: - EVENT_NAME: ${{ github.event_name }} INPUT_ENVIRONMENT: ${{ inputs.environment }} run: | set -euo pipefail INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}" - case "$EVENT_NAME" in - workflow_dispatch|workflow_call) ENVIRONMENT="$INPUT_ENVIRONMENT" ;; - *) ENVIRONMENT=dev ;; - esac + # workflow_dispatch and workflow_call are this workflow's only + # triggers and both declare `environment` as `required: true`, so the + # input is the single source of truth and there is nothing to fall + # back to. Branching on `github.event_name` here was issue #1805: + # inside a reusable workflow it reports the CALLER's event and is + # never "workflow_call", so a release-triggered caller passing + # environment=prod would have mis-resolved to a `dev` default arm. + # deploy-all.yml hardcodes this workflow off, so the case was latent + # here and live only for Azure and GCP. + ENVIRONMENT="$INPUT_ENVIRONMENT" # workflow_dispatch constrains this to the declared `choice` options, # but the workflow_call input is typed as a free-form string, so the @@ -91,7 +96,7 @@ jobs: case "$ENVIRONMENT" in dev|staging|prod) ;; *) - echo "::error::Refusing unknown environment: $ENVIRONMENT" + echo "::error::Refusing unknown environment: '$ENVIRONMENT'" exit 1 ;; esac diff --git a/.github/workflows/deploy-aws-lambda.yml b/.github/workflows/deploy-aws-lambda.yml index f52ac1c5b..bccf4423f 100644 --- a/.github/workflows/deploy-aws-lambda.yml +++ b/.github/workflows/deploy-aws-lambda.yml @@ -107,17 +107,31 @@ jobs: set -euo pipefail INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}" - # Inside a reusable workflow `github.event_name` is the CALLER's - # event, so it is never "workflow_call" — a caller's free-form - # `inputs.environment` arrives through the workflow_dispatch arm. - # Caveat: the `*` arm forces dev, so a push-triggered caller would - # have its requested environment silently ignored. deploy-all.yml - # triggers only on workflow_dispatch and release today. - case "$EVENT_NAME" in - release) TARGET_ENVIRONMENT=prod ;; - workflow_dispatch) TARGET_ENVIRONMENT="$INPUT_ENVIRONMENT" ;; - *) TARGET_ENVIRONMENT=dev ;; - esac + # Branch on whether an environment was REQUESTED, not on the event + # name. Inside a reusable workflow `github.event_name` is the CALLER's + # event and is never "workflow_call", so no event name can identify a + # workflow_call invocation (issue #1805). A non-empty input is always + # an explicit request and wins. + # + # An empty input means none was supplied, which this workflow's own + # `release` and `push` triggers are the legitimate causes of; every + # other event stops the run. Note that `release` is inherited by a + # workflow_call from a release-triggered caller, so a caller passing + # an EMPTY environment would resolve to prod here rather than + # failing. deploy-all.yml is the only caller and allowlists the value + # it passes, so that cannot happen. Telling the two apart would need + # `github.job_workflow_ref`, which is not worth the machinery for an + # unreachable path. + if [ -n "$INPUT_ENVIRONMENT" ]; then + TARGET_ENVIRONMENT="$INPUT_ENVIRONMENT" + elif [ "$EVENT_NAME" = release ]; then + TARGET_ENVIRONMENT=prod + elif [ "$EVENT_NAME" = push ]; then + TARGET_ENVIRONMENT=dev + else + echo "::error::No environment supplied on a '$EVENT_NAME'-triggered run; refusing to guess a deployment target." + exit 1 + fi # workflow_dispatch constrains this to the declared `choice` options, # but the workflow_call input is typed as a free-form string, so the @@ -125,7 +139,7 @@ jobs: case "$TARGET_ENVIRONMENT" in dev|staging|prod) ;; *) - echo "::error::Refusing unknown environment: $TARGET_ENVIRONMENT" + echo "::error::Refusing unknown environment: '$TARGET_ENVIRONMENT'" exit 1 ;; esac diff --git a/.github/workflows/deploy-azure.yml b/.github/workflows/deploy-azure.yml index 2d169d772..7c204e2d9 100644 --- a/.github/workflows/deploy-azure.yml +++ b/.github/workflows/deploy-azure.yml @@ -85,10 +85,30 @@ jobs: set -euo pipefail INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}" - case "$EVENT_NAME" in - workflow_dispatch|workflow_call) ENVIRONMENT="$INPUT_ENVIRONMENT" ;; - *) ENVIRONMENT=dev ;; - esac + # Branch on whether an environment was REQUESTED, not on the event + # name. Inside a reusable workflow `github.event_name` is the CALLER's + # event and is never "workflow_call", so no event name can identify a + # workflow_call invocation: a release-triggered deploy-all.yml passing + # environment=prod would have landed on a `dev` default arm and + # resolved to dev while the summary said prod (issue #1805). + # + # `inputs` is populated on workflow_dispatch and workflow_call, the + # only two triggers that carry an environment, so a non-empty input is + # always an explicit request and wins. An empty input means none was + # supplied: under `push` that is this workflow's own main-branch + # deploy and means dev, and every other event stops the run rather + # than guessing. A workflow_call from a push-triggered caller would + # read as `push` here too, since `github.event_name` is inherited, but + # deploy-all.yml is the only caller, has no push trigger, and + # allowlists the value it passes. + if [ -n "$INPUT_ENVIRONMENT" ]; then + ENVIRONMENT="$INPUT_ENVIRONMENT" + elif [ "$EVENT_NAME" = push ]; then + ENVIRONMENT=dev + else + echo "::error::No environment supplied on a '$EVENT_NAME'-triggered run; refusing to guess a deployment target." + exit 1 + fi # workflow_dispatch constrains this to the declared `choice` options, # but the workflow_call input is typed as a free-form string, so the @@ -96,7 +116,7 @@ jobs: case "$ENVIRONMENT" in dev|staging|prod) ;; *) - echo "::error::Refusing unknown environment: $ENVIRONMENT" + echo "::error::Refusing unknown environment: '$ENVIRONMENT'" exit 1 ;; esac diff --git a/.github/workflows/deploy-gcp.yml b/.github/workflows/deploy-gcp.yml index b5357a029..0a327f5b5 100644 --- a/.github/workflows/deploy-gcp.yml +++ b/.github/workflows/deploy-gcp.yml @@ -72,10 +72,30 @@ jobs: set -euo pipefail INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}" - case "$EVENT_NAME" in - workflow_dispatch|workflow_call) ENVIRONMENT="$INPUT_ENVIRONMENT" ;; - *) ENVIRONMENT=dev ;; - esac + # Branch on whether an environment was REQUESTED, not on the event + # name. Inside a reusable workflow `github.event_name` is the CALLER's + # event and is never "workflow_call", so no event name can identify a + # workflow_call invocation: a release-triggered deploy-all.yml passing + # environment=prod would have landed on a `dev` default arm and + # resolved to dev while the summary said prod (issue #1805). + # + # `inputs` is populated on workflow_dispatch and workflow_call, the + # only two triggers that carry an environment, so a non-empty input is + # always an explicit request and wins. An empty input means none was + # supplied: under `push` that is this workflow's own main-branch + # deploy and means dev, and every other event stops the run rather + # than guessing. A workflow_call from a push-triggered caller would + # read as `push` here too, since `github.event_name` is inherited, but + # deploy-all.yml is the only caller, has no push trigger, and + # allowlists the value it passes. + if [ -n "$INPUT_ENVIRONMENT" ]; then + ENVIRONMENT="$INPUT_ENVIRONMENT" + elif [ "$EVENT_NAME" = push ]; then + ENVIRONMENT=dev + else + echo "::error::No environment supplied on a '$EVENT_NAME'-triggered run; refusing to guess a deployment target." + exit 1 + fi # workflow_dispatch constrains this to the declared `choice` options, # but the workflow_call input is typed as a free-form string, so the @@ -83,7 +103,7 @@ jobs: case "$ENVIRONMENT" in dev|staging|prod) ;; *) - echo "::error::Refusing unknown environment: $ENVIRONMENT" + echo "::error::Refusing unknown environment: '$ENVIRONMENT'" exit 1 ;; esac