Symptom
A user encountered a 500 on plan-creation in the deployed environment (Lambda cudly-dev-426fc8af-api, AWS account 909626172446, region us-east-1). All purchase-execution reads and writes are broken.
CloudWatch Evidence
[ERROR] API error: failed to get planned executions: failed to scan execution: number of field descriptions must equal number of destinations, got 22 and 25
[ERROR] API error: failed to get pending executions: failed to query executions: ERROR: column "executed_by_user_id" does not exist (SQLSTATE 42703)
[ERROR] API error: failed to save execution (row 1/3): failed to save purchase execution: ERROR: column "executed_by_user_id" of relation "purchase_executions" does not exist (SQLSTATE 42703)
And on every cold start:
[ERROR] Prepare ... ERROR: column "role" of relation "users" does not exist (SQLSTATE 42703) ... sql: INSERT INTO users ...
⚠️ Migration failed — app continuing with existing schema: failed to create admin user: failed to upsert admin user: ERROR: column "role" of relation "users" does not exist (SQLSTATE 42703)
Root Cause: Two Distinct Sub-Bugs
Sub-bug 1 — Migration 000058 not applied to deployed RDS
Migration 000058_purchase_executions_direct_execute_audit.up.sql adds three audit columns to purchase_executions: executed_by_user_id, executed_at, pre_approval_skip_reason. The deployed Lambda image was built before this migration was merged, so the image's /app/migrations directory does not contain the 000058 SQL files. As a result m.Up() in RunMigrations has never applied them, and the Go scanner (scanExecutionRows in internal/config/store_postgres.go) expects 25 columns while the table only has 22. This breaks ALL purchase-execution reads and writes with P1 impact.
Sub-bug 2 — Admin-upsert SQL references dropped users.role column (every cold start)
Migration 000057_drop_user_role_to_groups.up.sql (PR #912) dropped the users.role column and sessions.role column. However, internal/database/postgres/migrations/migrate.go was not updated: both ensureAdminUser and ensureAdminUserWithPassword still include role in their INSERT column lists, and assignAdminGroupAndWarn still has WHERE role = 'admin' in two UPDATE/SELECT queries. On every cold start after migration 057 was applied, these SQL statements fail with column "role" of relation "users" does not exist, causing RunMigrations to return an error.
Why the cold-start admin-upsert failure may be masking migration 058
RunMigrations runs m.Up() first (which applies pending SQL migrations in order), then calls ensureAdminUser only if an admin email is configured. If the deployed image predates migration 058, m.Up() simply has no 058 file to apply. The admin-upsert failure then causes RunMigrations to return an error, but app.go soft-continues. The net result is: the app boots, 058 is never applied (its SQL file is absent from the image), and every request that touches purchase_executions gets a column-not-found error.
The two bugs are independent in cause but compound in effect: even if the image were rebuilt with 058 included, the admin-upsert bug would still fire on every cold start as long as migrate.go references the dropped role column.
Fix Shape
Fix 1 — Remove role from all admin-upsert SQL in migrate.go
- Remove
role from the INSERT INTO users column list and VALUES clause in both ensureAdminUser and ensureAdminUserWithPassword.
- Replace
WHERE role = 'admin' in assignAdminGroupAndWarn with a group-membership check: WHERE $1::UUID = ANY(group_ids) (identify admin users by Administrators group membership, which is the post-057 invariant).
Fix 2 — Decouple admin-bootstrap failure from migration-step error
Restructure RunMigrations so that ensureAdminUser failure is logged but does NOT cause RunMigrations to return an error. The migration SQL step (m.Up()) and the admin bootstrap are independent concerns; conflating them makes cold-start logs ambiguous and means a stale migrate.go can make the health endpoint report migration failure even when the schema is fully up to date.
Fix 3 — Apply migration 058 to deployed RDS (ops action)
The deployed image must be rebuilt from the current branch (which includes 058) and redeployed. After the image swap, m.Up() will apply 058 on the next cold start. Operators should monitor CloudWatch for Database migrations completed successfully (version: 58) or higher.
If the deployed schema_migrations table is in a dirty state (unlikely given the soft-continue path, but possible), the operator can set CUDLY_FORCE_MIGRATION_VERSION=57 and redeploy once to clear the dirty flag, then remove the env var.
Acceptance Criteria
- CloudWatch shows zero
column "role" of relation "users" errors on cold start after deploy
- CloudWatch shows zero
column "executed_by_user_id" does not exist errors after image rebuild + redeploy
RunMigrations returns nil when m.Up() succeeds, even if the admin-bootstrap step logs a warning
- Integration tests pass:
go test -tags integration ./internal/database/postgres/migrations/...
Symptom
A user encountered a 500 on plan-creation in the deployed environment (Lambda
cudly-dev-426fc8af-api, AWS account909626172446, regionus-east-1). All purchase-execution reads and writes are broken.CloudWatch Evidence
And on every cold start:
Root Cause: Two Distinct Sub-Bugs
Sub-bug 1 — Migration 000058 not applied to deployed RDS
Migration
000058_purchase_executions_direct_execute_audit.up.sqladds three audit columns topurchase_executions:executed_by_user_id,executed_at,pre_approval_skip_reason. The deployed Lambda image was built before this migration was merged, so the image's/app/migrationsdirectory does not contain the 000058 SQL files. As a resultm.Up()inRunMigrationshas never applied them, and the Go scanner (scanExecutionRowsininternal/config/store_postgres.go) expects 25 columns while the table only has 22. This breaks ALL purchase-execution reads and writes with P1 impact.Sub-bug 2 — Admin-upsert SQL references dropped
users.rolecolumn (every cold start)Migration
000057_drop_user_role_to_groups.up.sql(PR #912) dropped theusers.rolecolumn andsessions.rolecolumn. However,internal/database/postgres/migrations/migrate.gowas not updated: bothensureAdminUserandensureAdminUserWithPasswordstill includerolein their INSERT column lists, andassignAdminGroupAndWarnstill hasWHERE role = 'admin'in two UPDATE/SELECT queries. On every cold start after migration 057 was applied, these SQL statements fail withcolumn "role" of relation "users" does not exist, causingRunMigrationsto return an error.Why the cold-start admin-upsert failure may be masking migration 058
RunMigrationsrunsm.Up()first (which applies pending SQL migrations in order), then callsensureAdminUseronly if an admin email is configured. If the deployed image predates migration 058,m.Up()simply has no 058 file to apply. The admin-upsert failure then causesRunMigrationsto return an error, butapp.gosoft-continues. The net result is: the app boots, 058 is never applied (its SQL file is absent from the image), and every request that touchespurchase_executionsgets a column-not-found error.The two bugs are independent in cause but compound in effect: even if the image were rebuilt with 058 included, the admin-upsert bug would still fire on every cold start as long as
migrate.goreferences the droppedrolecolumn.Fix Shape
Fix 1 — Remove
rolefrom all admin-upsert SQL inmigrate.gorolefrom theINSERT INTO userscolumn list andVALUESclause in bothensureAdminUserandensureAdminUserWithPassword.WHERE role = 'admin'inassignAdminGroupAndWarnwith a group-membership check:WHERE $1::UUID = ANY(group_ids)(identify admin users by Administrators group membership, which is the post-057 invariant).Fix 2 — Decouple admin-bootstrap failure from migration-step error
Restructure
RunMigrationsso thatensureAdminUserfailure is logged but does NOT causeRunMigrationsto return an error. The migration SQL step (m.Up()) and the admin bootstrap are independent concerns; conflating them makes cold-start logs ambiguous and means a stalemigrate.gocan make the health endpoint report migration failure even when the schema is fully up to date.Fix 3 — Apply migration 058 to deployed RDS (ops action)
The deployed image must be rebuilt from the current branch (which includes 058) and redeployed. After the image swap,
m.Up()will apply 058 on the next cold start. Operators should monitor CloudWatch forDatabase migrations completed successfully (version: 58)or higher.If the deployed
schema_migrationstable is in a dirty state (unlikely given the soft-continue path, but possible), the operator can setCUDLY_FORCE_MIGRATION_VERSION=57and redeploy once to clear the dirty flag, then remove the env var.Acceptance Criteria
column "role" of relation "users"errors on cold start after deploycolumn "executed_by_user_id" does not existerrors after image rebuild + redeployRunMigrationsreturns nil whenm.Up()succeeds, even if the admin-bootstrap step logs a warninggo test -tags integration ./internal/database/postgres/migrations/...