Skip to content

ops(gcp): take Cloud Run secure-ingress default into effect (per-env LB enablement + override removal) #78

Description

@cristim

Summary

PR #73 flipped the cloud_run_ingress variable default in terraform/environments/gcp/variables.tf to INGRESS_TRAFFIC_INTERNAL_LOAD_BALANCER, but every shipped tfvars file currently overrides it back to INGRESS_TRAFFIC_ALL because the supporting Cloud Armor + Load Balancer stack isn't enabled in any environment yet (enable_cdn = false everywhere). That's why the PR used Refs #51 rather than Closes #51.

This issue tracks the per-env migration so the security gap doesn't sit invisible in the override list forever.

Current behaviour

$ grep -n "cloud_run_ingress" terraform/environments/gcp/*.tfvars*
terraform/environments/gcp/dev.tfvars.example:...:cloud_run_ingress = "INGRESS_TRAFFIC_ALL"
terraform/environments/gcp/github-dev.tfvars:...:cloud_run_ingress = "INGRESS_TRAFFIC_ALL"
terraform/environments/gcp/github-staging.tfvars:...:cloud_run_ingress = "INGRESS_TRAFFIC_ALL"
terraform/environments/gcp/github-prod.tfvars:...:cloud_run_ingress = "INGRESS_TRAFFIC_ALL"

All 4 envs override the new secure default. Cloud Run is currently behind IAM-level auth only (cloud_run_allow_unauthenticated = false, per the first half of #51) — defence-in-depth via LB + Cloud Armor is not yet in effect.

Steps to reproduce / verify the gap

# 1. Confirm every tfvars overrides the secure default
grep -n "cloud_run_ingress" terraform/environments/gcp/*.tfvars*

# 2. Confirm the deployed Cloud Run ingress annotation reflects the override
gcloud run services describe <svc> --region <region> \
  --format='value(spec.template.metadata.annotations[run.googleapis.com/ingress])'
# Returns: all
# Expected after migration: internal-and-cloud-load-balancing

Expected behaviour

Each env's tfvar should drop the cloud_run_ingress override once its LB + DNS + cert are provisioned, so the secure variable default takes effect and Cloud Run is reachable only via the LB (with Cloud Armor in front).

Proposed fix — per-env migration steps

Run this 4-step checklist for each env, in this order: dev → github-dev → github-staging → github-prod.

dev (terraform/environments/gcp/dev.tfvars.example)

  • Flip enable_cdn = true to instantiate the LB stack at terraform/modules/frontend/gcp/.
  • Provision DNS + TLS cert via frontend_domain_names + dns_managed_zone_name.
  • Drop the cloud_run_ingress = "INGRESS_TRAFFIC_ALL" line from the tfvar.
  • Verify with gcloud run services describe ... (annotation flips to internal-and-cloud-load-balancing) and curl against the direct Cloud Run URL (expect 403) vs the LB domain (expect 200).

github-dev (terraform/environments/gcp/github-dev.tfvars)

  • Flip enable_cdn = true to instantiate the LB stack.
  • Provision DNS + TLS cert via frontend_domain_names + dns_managed_zone_name.
  • Drop the cloud_run_ingress = "INGRESS_TRAFFIC_ALL" line.
  • Verify with gcloud run services describe ... and curl (direct URL → 403, LB domain → 200).

github-staging (terraform/environments/gcp/github-staging.tfvars)

  • Flip enable_cdn = true to instantiate the LB stack.
  • Provision DNS + TLS cert via frontend_domain_names + dns_managed_zone_name.
  • Drop the cloud_run_ingress = "INGRESS_TRAFFIC_ALL" line.
  • Verify with gcloud run services describe ... and curl (direct URL → 403, LB domain → 200).

github-prod (terraform/environments/gcp/github-prod.tfvars)

  • Flip enable_cdn = true to instantiate the LB stack.
  • Provision DNS + TLS cert via frontend_domain_names + dns_managed_zone_name.
  • Drop the cloud_run_ingress = "INGRESS_TRAFFIC_ALL" line.
  • Verify with gcloud run services describe ... and curl (direct URL → 403, LB domain → 200).

References

Severity

Medium — defence-in-depth gap. Cloud Run is currently behind IAM-level auth only; a misconfigured IAM binding (e.g. an accidental allUsers / allAuthenticatedUsers grant) would expose the service directly to the public internet without a Load Balancer / Cloud Armor in front to absorb abuse.

Effort per env

Medium — depends on whether DNS + cert are already provisioned. dev is the quickest (no production traffic, easiest to validate); github-prod requires the most coordination (DNS cutover, monitoring, rollback plan).

Why not a single PR

Each env's migration is operator-coordinated (DNS records, cert provisioning, traffic cutover validation) and breaks differently if it goes wrong. Better tracked as a checklist with one item closed per env than as a monolithic PR that has to land everywhere at once.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions