Skip to content

bug(api/executions): migration 000058 not applied on deployed RDS -> 500 on all execution reads/writes (plus admin-upsert SQL still uses dropped users.role) #945

Description

@cristim

Symptom

A user encountered a 500 on plan-creation in the deployed environment (Lambda cudly-dev-426fc8af-api, AWS account 909626172446, region us-east-1). All purchase-execution reads and writes are broken.

CloudWatch Evidence

[ERROR] API error: failed to get planned executions: failed to scan execution: number of field descriptions must equal number of destinations, got 22 and 25
[ERROR] API error: failed to get pending executions: failed to query executions: ERROR: column "executed_by_user_id" does not exist (SQLSTATE 42703)
[ERROR] API error: failed to save execution (row 1/3): failed to save purchase execution: ERROR: column "executed_by_user_id" of relation "purchase_executions" does not exist (SQLSTATE 42703)

And on every cold start:

[ERROR] Prepare ... ERROR: column "role" of relation "users" does not exist (SQLSTATE 42703) ... sql: INSERT INTO users ...
⚠️ Migration failed — app continuing with existing schema: failed to create admin user: failed to upsert admin user: ERROR: column "role" of relation "users" does not exist (SQLSTATE 42703)

Root Cause: Two Distinct Sub-Bugs

Sub-bug 1 — Migration 000058 not applied to deployed RDS

Migration 000058_purchase_executions_direct_execute_audit.up.sql adds three audit columns to purchase_executions: executed_by_user_id, executed_at, pre_approval_skip_reason. The deployed Lambda image was built before this migration was merged, so the image's /app/migrations directory does not contain the 000058 SQL files. As a result m.Up() in RunMigrations has never applied them, and the Go scanner (scanExecutionRows in internal/config/store_postgres.go) expects 25 columns while the table only has 22. This breaks ALL purchase-execution reads and writes with P1 impact.

Sub-bug 2 — Admin-upsert SQL references dropped users.role column (every cold start)

Migration 000057_drop_user_role_to_groups.up.sql (PR #912) dropped the users.role column and sessions.role column. However, internal/database/postgres/migrations/migrate.go was not updated: both ensureAdminUser and ensureAdminUserWithPassword still include role in their INSERT column lists, and assignAdminGroupAndWarn still has WHERE role = 'admin' in two UPDATE/SELECT queries. On every cold start after migration 057 was applied, these SQL statements fail with column "role" of relation "users" does not exist, causing RunMigrations to return an error.

Why the cold-start admin-upsert failure may be masking migration 058

RunMigrations runs m.Up() first (which applies pending SQL migrations in order), then calls ensureAdminUser only if an admin email is configured. If the deployed image predates migration 058, m.Up() simply has no 058 file to apply. The admin-upsert failure then causes RunMigrations to return an error, but app.go soft-continues. The net result is: the app boots, 058 is never applied (its SQL file is absent from the image), and every request that touches purchase_executions gets a column-not-found error.

The two bugs are independent in cause but compound in effect: even if the image were rebuilt with 058 included, the admin-upsert bug would still fire on every cold start as long as migrate.go references the dropped role column.

Fix Shape

Fix 1 — Remove role from all admin-upsert SQL in migrate.go

  • Remove role from the INSERT INTO users column list and VALUES clause in both ensureAdminUser and ensureAdminUserWithPassword.
  • Replace WHERE role = 'admin' in assignAdminGroupAndWarn with a group-membership check: WHERE $1::UUID = ANY(group_ids) (identify admin users by Administrators group membership, which is the post-057 invariant).

Fix 2 — Decouple admin-bootstrap failure from migration-step error

Restructure RunMigrations so that ensureAdminUser failure is logged but does NOT cause RunMigrations to return an error. The migration SQL step (m.Up()) and the admin bootstrap are independent concerns; conflating them makes cold-start logs ambiguous and means a stale migrate.go can make the health endpoint report migration failure even when the schema is fully up to date.

Fix 3 — Apply migration 058 to deployed RDS (ops action)

The deployed image must be rebuilt from the current branch (which includes 058) and redeployed. After the image swap, m.Up() will apply 058 on the next cold start. Operators should monitor CloudWatch for Database migrations completed successfully (version: 58) or higher.

If the deployed schema_migrations table is in a dirty state (unlikely given the soft-continue path, but possible), the operator can set CUDLY_FORCE_MIGRATION_VERSION=57 and redeploy once to clear the dirty flag, then remove the env var.

Acceptance Criteria

  • CloudWatch shows zero column "role" of relation "users" errors on cold start after deploy
  • CloudWatch shows zero column "executed_by_user_id" does not exist errors after image rebuild + redeploy
  • RunMigrations returns nil when m.Up() succeeds, even if the admin-bootstrap step logs a warning
  • Integration tests pass: go test -tags integration ./internal/database/postgres/migrations/...

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions