Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
62 changes: 55 additions & 7 deletions .github/workflows/deploy-all.yml
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,15 @@

name: Deploy to All Clouds

# Adding a trigger here, or a second caller of the deploy-* workflows, means
# revisiting the `prepare` comments in deploy-azure.yml, deploy-gcp.yml and
# deploy-aws-lambda.yml. Since #1805 those three resolve their environment from
# `inputs.environment` and fall back on `github.event_name`, which inside a
# reusable workflow is whatever event started THIS run. They are correct today
# only because this workflow is their only caller, has no `push` trigger, and
# allowlists the environment it passes. A `push:` trigger here would make
# Azure and GCP silently resolve dev for a caller that asked for something
# else, which is #1805 restored.
on:
workflow_dispatch:
inputs:
Expand Down Expand Up @@ -56,13 +65,36 @@ jobs:
steps:
- name: Set environment
id: set-env
env:
# This workflow is always the entry point of its own run, never a
# reusable one, so `github.event_name` here really is the trigger that
# started it. The called workflows cannot rely on that and read
# `inputs.environment` instead (issue #1805).
EVENT_NAME: ${{ github.event_name }}
INPUT_ENVIRONMENT: ${{ inputs.environment }}
run: |
if [[ "${{ github.event_name }}" == "release" ]]; then
echo "environment=prod" >> $GITHUB_OUTPUT
set -euo pipefail
INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}"

if [ "$EVENT_NAME" = release ]; then
ENVIRONMENT=prod
else
echo "environment=${{ inputs.environment }}" >> $GITHUB_OUTPUT
ENVIRONMENT="$INPUT_ENVIRONMENT"
fi

# Every called workflow takes this value verbatim, so an empty or
# unknown value has to stop the fan-out here rather than reach four
# deployments.
case "$ENVIRONMENT" in
dev|staging|prod) ;;
*)
echo "::error::Refusing unknown environment: '$ENVIRONMENT'"
exit 1
;;
Comment thread
coderabbitai[bot] marked this conversation as resolved.
esac

echo "environment=$ENVIRONMENT" >> "$GITHUB_OUTPUT"

- name: Set cloud providers
id: set-clouds
run: |
Expand Down Expand Up @@ -166,8 +198,9 @@ jobs:
# Deliberately carries no `concurrency` (#1801). The Terraform state guard is
# the group on the called workflow's own build-and-deploy job, which runs as a
# real job of this run and is serialized there. That group is derived from the
# called workflow's own `prepare` output, which is not necessarily the
# `environment` computed above, so do not assume the two agree.
# called workflow's own `prepare` output, which since #1805 is the
# `environment` computed above verbatim: `prepare` takes the non-empty input
# it is passed and fails the run rather than substituting a default.
#
# Putting that same group on this caller job would deadlock: this job would
# hold the group while waiting on the inner job queued behind it. GitHub
Expand Down Expand Up @@ -263,7 +296,22 @@ jobs:
run: |
FAILED=false

if [[ "${{ needs.deploy-aws-lambda.result }}" == "failure" ]] || \
# determine-deployment is checked for anything other than success, not
# just "failure". When it stops the run the four deploy jobs are
# SKIPPED rather than failed, so testing only those would report "all
# deployments completed successfully" for a run that deployed nothing.
#
# The four deploy jobs are deliberately still tested for "failure"
# only. A CANCELLED run leaves them `cancelled` and reaches the
# success line, which is the same false report. That predates this
# change, so it is tracked in #1810 rather than fixed here.
#
# Note this step is the only success-by-default check in the repo. All
# four per-cloud `summary` jobs allowlist on `== "success"` instead,
# so an unenumerated result falls to their failure branch and errs
# safe. This denylist is the outlier, not the pattern.
if [[ "${{ needs.determine-deployment.result }}" != "success" ]] || \
[[ "${{ needs.deploy-aws-lambda.result }}" == "failure" ]] || \
[[ "${{ needs.deploy-aws-fargate.result }}" == "failure" ]] || \
[[ "${{ needs.deploy-gcp.result }}" == "failure" ]] || \
[[ "${{ needs.deploy-azure.result }}" == "failure" ]]; then
Comment on lines +313 to 317

@coderabbitai coderabbitai Bot Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

sed -n '286,309p' .github/workflows/deploy-all.yml

# Validate in GitHub Actions with a selected reusable deployment job that exits
# as cancelled. Expected: the `notify` job fails and does not emit the success
# summary.

Repository: LeanerCloud/CUDly

Length of output: 1342


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- deployment job declarations and result aggregation ---'
sed -n '1,120p' .github/workflows/deploy-all.yml
sed -n '120,220p' .github/workflows/deploy-all.yml
sed -n '250,315p' .github/workflows/deploy-all.yml

printf '%s\n' '--- all references to deployment selectors and job conditions ---'
rg -n -C 3 'deploy-(aws-lambda|aws-fargate|gcp|azure)|determine-deployment|needs\..*\.result|if:' .github/workflows/deploy-all.yml

Repository: LeanerCloud/CUDly

Length of output: 20448


🏁 Script executed:

#!/bin/bash
set -euo pipefail

python3 - <<'PY'
from itertools import product

selectors = {
    "aws-lambda": {"all", "aws-only", "aws-gcp", "aws-azure"},
    "aws-fargate": set(),
    "gcp": {"all", "gcp-only", "aws-gcp", "gcp-azure"},
    "azure": {"all", "azure-only", "aws-azure", "gcp-azure"},
}
strategies = ["all", "aws-only", "gcp-only", "azure-only", "aws-gcp", "aws-azure", "gcp-azure"]
results = ["success", "failure", "cancelled", "skipped"]

for strategy in strategies:
    selected = [name for name, enabled in selectors.items() if strategy in enabled]
    false_positive = []
    for statuses in product(results, repeat=len(selected)):
        current = all(status != "failure" for status in statuses)
        proposed = all(status == "success" for status in statuses)
        if current and not proposed:
            false_positive.append(dict(zip(selected, statuses)))
    print(strategy, "selected=", selected, "current-success-but-not-proposed=", false_positive[:8],
          "count=", len(false_positive))
PY

printf '%s\n' '--- reusable workflow job-level conditions and cancellation-sensitive constructs ---'
for file in $(fd -i 'deploy-(aws-lambda|aws-fargate|gcp|azure)\.yml$' .github/workflows); do
  echo "### $file"
  rg -n -C 2 '^[[:space:]]+(if:|needs:|concurrency:)|cancel-in-progress|always\(\)|failure\(\)|cancelled\(\)' "$file" || true
done

Repository: LeanerCloud/CUDly

Length of output: 7179


🌐 Web query:

GitHub Actions documentation needs context result values cancelled skipped reusable workflow caller job result

💡 Result:

In GitHub Actions, you can determine the result of a job in a reusable workflow using the jobs context, which is available within the reusable workflow itself [1][2]. The jobs.<job_id>.result property provides the status of a specific job and can return one of the following string values: success, failure, cancelled, or skipped [1][2]. Regarding the context in a caller workflow: 1. Caller Workflow Limitations: When a reusable workflow is triggered, the caller workflow treats the entire job that calls the reusable workflow as a single unit [3]. You can access the outputs of the reusable workflow, but the individual internal job statuses (like whether a specific internal job was cancelled or skipped) are not automatically exposed as properties in the caller workflow's jobs context [1][4]. 2. Handling Skipped or Cancelled States: If you need to trigger logic in a caller workflow based on the specific outcome of an internal job within a reusable workflow, you must explicitly define outputs in the reusable workflow [3][4]. You can set these outputs to reflect the result of internal jobs using expressions like ${{ jobs.<job_id>.result }} [1][4]. 3. Concurrency and Cancellation: If you use concurrency with cancel-in-progress: true in both the caller and the called reusable workflow, ensure you are not using the same concurrency group name, as this may lead to the caller workflow being prematurely cancelled [5][6]. To cancel sub-workflows, you can apply concurrency settings directly within the reusable workflow definition [7]. For debugging purposes, you can inspect the available information by printing the context to the logs using ${{ toJSON(jobs) }} within the reusable workflow [4].

Citations:


Treat cancelled selected deployments as unsuccessful.

The current check flags only failure, so a selected job with result cancelled can produce a success summary. Compare each selected deployment result with success; keep unselected jobs as skipped.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/deploy-all.yml around lines 294 - 298, Update the
deployment-result condition in the summary job to compare each selected
deployment’s result against success rather than checking only for failure.
Ensure cancelled selected jobs are treated as unsuccessful while skipped,
unselected jobs do not trigger the condition; apply this to deploy-aws-lambda,
deploy-aws-fargate, deploy-gcp, and deploy-azure alongside determine-deployment.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid finding, and already tracked as #1810. Deliberately out of scope for this PR.

You are right about the substance: the Check for failures step tests each deploy job for == "failure", and a cancelled job is not failure, so a run killed mid-apply prints "All deployments completed successfully!". Converting to an allowlist on == "success" is the correct shape.

It is excluded here for one reason: this PR did not cause it. The four == "failure" tests are byte-identical to origin/main. The related change that was folded into this PR is the determine-deployment case, and that one qualified precisely because the new environment allowlist made that job able to fail for the first time, turning previously-impossible skipped states into reachable ones.

Some context recorded on #1810 that is worth having before anyone picks it up, since two of these were wrong when first written:

  • It is one site, not a class. == "failure" appears 4 times in deploy-all.yml and zero times in any other workflow. Enumerated across all 16 workflow files, every other .result comparison allowlists on success, including all four per-cloud summary jobs. deploy-all.yml is the outlier, not the pattern.
  • In-repo precedent exists for the fix shape: ci.yml:740-744 already tests contains(needs.*.result, 'failure') and contains(needs.*.result, 'cancelled') as separate states.
  • The four per-cloud summary jobs do print "Deployment failed" on a cancelled run. Wrong wording, safe direction, tracked separately on #1810 so it does not get conflated with this fix.

Leaving the code unchanged here; the gap is documented in-place with a pointer to #1810.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@cristim, thanks for the detailed clarification.

The cancellation gap is valid, but it predates this PR and is tracked in #1810. The added determine-deployment.result != "success" check addresses the new failure path introduced by this PR. Keeping the four existing deployment-result checks unchanged is appropriate for this PR’s scope.


✏️ Learnings added
Learnt from: cristim
Repo: LeanerCloud/CUDly PR: 1809
File: .github/workflows/deploy-all.yml:313-317
Timestamp: 2026-08-12T21:59:41.967Z
Learning: In `.github/workflows/deploy-all.yml`, the `notify` job checks the four deployment job results with `== "failure"`. A selected deployment with result `cancelled` can therefore reach the success summary. This pre-existing gap is tracked in GitHub issue `#1810` and is deliberately out of scope for PR `#1809`. The PR-specific fix is the `needs.determine-deployment.result != "success"` check because PR `#1809` introduces environment validation that can cause `determine-deployment` to fail and downstream jobs to be skipped.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.


🧠 Learnings used
Learnt from: cristim
Repo: LeanerCloud/CUDly PR: 1809
File: .github/workflows/deploy-all.yml:79-93
Timestamp: 2026-08-12T21:59:15.382Z
Learning: In `.github/workflows/deploy-all.yml`, only `workflow_dispatch` and `release` trigger the workflow. `release` resolves to `prod`, and `workflow_dispatch` requires an allowlisted `dev`, `staging`, or `prod` environment. The `Set environment` allowlist must fail for an empty or unsupported environment rather than map a potential future `push` event to `dev`, so that a new trigger cannot silently choose a deployment environment.

You are interacting with an AI system.

Expand All @@ -272,7 +320,7 @@ jobs:

if [ "$FAILED" = true ]; then
echo "" >> $GITHUB_STEP_SUMMARY
echo "❌ **One or more deployments failed. Check individual job logs.**" >> $GITHUB_STEP_SUMMARY
echo "❌ **Deployment did not complete successfully. Check individual job logs.**" >> $GITHUB_STEP_SUMMARY
exit 1
else
echo "" >> $GITHUB_STEP_SUMMARY
Expand Down
17 changes: 11 additions & 6 deletions .github/workflows/deploy-aws-fargate.yml
Original file line number Diff line number Diff line change
Expand Up @@ -74,24 +74,29 @@ jobs:
- name: Determine environment
id: set-env
env:
EVENT_NAME: ${{ github.event_name }}
INPUT_ENVIRONMENT: ${{ inputs.environment }}
run: |
set -euo pipefail
INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}"

case "$EVENT_NAME" in
workflow_dispatch|workflow_call) ENVIRONMENT="$INPUT_ENVIRONMENT" ;;
*) ENVIRONMENT=dev ;;
esac
# workflow_dispatch and workflow_call are this workflow's only
# triggers and both declare `environment` as `required: true`, so the
# input is the single source of truth and there is nothing to fall
# back to. Branching on `github.event_name` here was issue #1805:
# inside a reusable workflow it reports the CALLER's event and is
# never "workflow_call", so a release-triggered caller passing
# environment=prod would have mis-resolved to a `dev` default arm.
# deploy-all.yml hardcodes this workflow off, so the case was latent
# here and live only for Azure and GCP.
ENVIRONMENT="$INPUT_ENVIRONMENT"

# workflow_dispatch constrains this to the declared `choice` options,
# but the workflow_call input is typed as a free-form string, so the
# allowlist is enforced here rather than assumed.
case "$ENVIRONMENT" in
dev|staging|prod) ;;
*)
echo "::error::Refusing unknown environment: $ENVIRONMENT"
echo "::error::Refusing unknown environment: '$ENVIRONMENT'"
exit 1
;;
esac
Expand Down
38 changes: 26 additions & 12 deletions .github/workflows/deploy-aws-lambda.yml
Original file line number Diff line number Diff line change
Expand Up @@ -107,25 +107,39 @@ jobs:
set -euo pipefail
INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}"

# Inside a reusable workflow `github.event_name` is the CALLER's
# event, so it is never "workflow_call" — a caller's free-form
# `inputs.environment` arrives through the workflow_dispatch arm.
# Caveat: the `*` arm forces dev, so a push-triggered caller would
# have its requested environment silently ignored. deploy-all.yml
# triggers only on workflow_dispatch and release today.
case "$EVENT_NAME" in
release) TARGET_ENVIRONMENT=prod ;;
workflow_dispatch) TARGET_ENVIRONMENT="$INPUT_ENVIRONMENT" ;;
*) TARGET_ENVIRONMENT=dev ;;
esac
# Branch on whether an environment was REQUESTED, not on the event
# name. Inside a reusable workflow `github.event_name` is the CALLER's
# event and is never "workflow_call", so no event name can identify a
# workflow_call invocation (issue #1805). A non-empty input is always
# an explicit request and wins.
#
# An empty input means none was supplied, which this workflow's own
# `release` and `push` triggers are the legitimate causes of; every
# other event stops the run. Note that `release` is inherited by a
# workflow_call from a release-triggered caller, so a caller passing
# an EMPTY environment would resolve to prod here rather than
# failing. deploy-all.yml is the only caller and allowlists the value
# it passes, so that cannot happen. Telling the two apart would need
# `github.job_workflow_ref`, which is not worth the machinery for an
# unreachable path.
if [ -n "$INPUT_ENVIRONMENT" ]; then
TARGET_ENVIRONMENT="$INPUT_ENVIRONMENT"
elif [ "$EVENT_NAME" = release ]; then
TARGET_ENVIRONMENT=prod
elif [ "$EVENT_NAME" = push ]; then
TARGET_ENVIRONMENT=dev
else
echo "::error::No environment supplied on a '$EVENT_NAME'-triggered run; refusing to guess a deployment target."
exit 1
fi

# workflow_dispatch constrains this to the declared `choice` options,
# but the workflow_call input is typed as a free-form string, so the
# allowlist is enforced here rather than assumed.
case "$TARGET_ENVIRONMENT" in
dev|staging|prod) ;;
*)
echo "::error::Refusing unknown environment: $TARGET_ENVIRONMENT"
echo "::error::Refusing unknown environment: '$TARGET_ENVIRONMENT'"
exit 1
;;
esac
Expand Down
30 changes: 25 additions & 5 deletions .github/workflows/deploy-azure.yml
Original file line number Diff line number Diff line change
Expand Up @@ -85,18 +85,38 @@ jobs:
set -euo pipefail
INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}"

case "$EVENT_NAME" in
workflow_dispatch|workflow_call) ENVIRONMENT="$INPUT_ENVIRONMENT" ;;
*) ENVIRONMENT=dev ;;
esac
# Branch on whether an environment was REQUESTED, not on the event
# name. Inside a reusable workflow `github.event_name` is the CALLER's
# event and is never "workflow_call", so no event name can identify a
# workflow_call invocation: a release-triggered deploy-all.yml passing
# environment=prod would have landed on a `dev` default arm and
# resolved to dev while the summary said prod (issue #1805).
#
# `inputs` is populated on workflow_dispatch and workflow_call, the
# only two triggers that carry an environment, so a non-empty input is
# always an explicit request and wins. An empty input means none was
# supplied: under `push` that is this workflow's own main-branch
# deploy and means dev, and every other event stops the run rather
# than guessing. A workflow_call from a push-triggered caller would
# read as `push` here too, since `github.event_name` is inherited, but
# deploy-all.yml is the only caller, has no push trigger, and
# allowlists the value it passes.
if [ -n "$INPUT_ENVIRONMENT" ]; then
ENVIRONMENT="$INPUT_ENVIRONMENT"
elif [ "$EVENT_NAME" = push ]; then
ENVIRONMENT=dev
else
echo "::error::No environment supplied on a '$EVENT_NAME'-triggered run; refusing to guess a deployment target."
exit 1
fi

# workflow_dispatch constrains this to the declared `choice` options,
# but the workflow_call input is typed as a free-form string, so the
# allowlist is enforced here rather than assumed.
case "$ENVIRONMENT" in
dev|staging|prod) ;;
*)
echo "::error::Refusing unknown environment: $ENVIRONMENT"
echo "::error::Refusing unknown environment: '$ENVIRONMENT'"
exit 1
;;
esac
Expand Down
30 changes: 25 additions & 5 deletions .github/workflows/deploy-gcp.yml
Original file line number Diff line number Diff line change
Expand Up @@ -72,18 +72,38 @@ jobs:
set -euo pipefail
INPUT_ENVIRONMENT="${INPUT_ENVIRONMENT:-}"

case "$EVENT_NAME" in
workflow_dispatch|workflow_call) ENVIRONMENT="$INPUT_ENVIRONMENT" ;;
*) ENVIRONMENT=dev ;;
esac
# Branch on whether an environment was REQUESTED, not on the event
# name. Inside a reusable workflow `github.event_name` is the CALLER's
# event and is never "workflow_call", so no event name can identify a
# workflow_call invocation: a release-triggered deploy-all.yml passing
# environment=prod would have landed on a `dev` default arm and
# resolved to dev while the summary said prod (issue #1805).
#
# `inputs` is populated on workflow_dispatch and workflow_call, the
# only two triggers that carry an environment, so a non-empty input is
# always an explicit request and wins. An empty input means none was
# supplied: under `push` that is this workflow's own main-branch
# deploy and means dev, and every other event stops the run rather
# than guessing. A workflow_call from a push-triggered caller would
# read as `push` here too, since `github.event_name` is inherited, but
# deploy-all.yml is the only caller, has no push trigger, and
# allowlists the value it passes.
if [ -n "$INPUT_ENVIRONMENT" ]; then
ENVIRONMENT="$INPUT_ENVIRONMENT"
elif [ "$EVENT_NAME" = push ]; then
ENVIRONMENT=dev
else
echo "::error::No environment supplied on a '$EVENT_NAME'-triggered run; refusing to guess a deployment target."
exit 1
fi

# workflow_dispatch constrains this to the declared `choice` options,
# but the workflow_call input is typed as a free-form string, so the
# allowlist is enforced here rather than assumed.
case "$ENVIRONMENT" in
dev|staging|prod) ;;
*)
echo "::error::Refusing unknown environment: $ENVIRONMENT"
echo "::error::Refusing unknown environment: '$ENVIRONMENT'"
exit 1
;;
esac
Expand Down
Loading