Skip to content

Move DB migrations off the request/connect path into a one-shot deploy-time runner #31

Description

@cristim

Problem

Database migrations currently run lazily on DB connect, inside the
per-instance ensureDB path (internal/server/app.go), bounded by a
timeout (CUDLY_MIGRATION_TIMEOUT, default now 120s). This couples schema
evolution to request/cold-start traffic and has several structural failure
modes:

  • Per-instance race. Every Lambda cold start (and every Fargate task)
    independently tries to run migrations. Concurrent cold starts race on the
    same schema_migrations row; a migration interrupted mid-run (Lambda
    timeout, ENI drop) leaves dirty = true, which is terminal — every
    later boot then fail-opens and serves 500s on any query needing the
    unapplied columns.
  • No fail-the-deploy signal. Because the app intentionally fail-opens on
    migration failure (so liveness probes pass), a broken migration is invisible
    to the deploy pipeline and to the AWS/Lambda Errors metric. The bad
    build ships and the failure only surfaces as user-facing 500s.
  • Timeout is a band-aid. Raising the bound (PR fix(migrations): opt-in dirty auto-heal, raise timeout default, failure alarm cloud-commitments-cli#1124 took it 20s -> 120s)
    buys headroom but the migration still competes with request latency budgets
    and Lambda's hard 300s ceiling. A genuinely long DDL on a large table is
    still at risk.

This was the root cause of the recent prod outage (migrations 070-073 never
applied; purchase_delay_hours / revocation_window_closes_at /
executed_by_user_id missing). PR LeanerCloud/cloud-commitments-cli#1124 hardened the in-place path
(opt-in dirty auto-heal, higher timeout, a CloudWatch failure alarm) but did
not change where/when migrations run.

Proposed structural fix

Move migrations off the request/connect path into a one-shot, deploy-time
runner
:

  • A single dedicated runner (CI deploy step, init container, or a one-shot
    Lambda/Job invoked by the deploy pipeline) runs RunMigrations exactly once
    per deploy, before the new app version starts taking traffic.
  • Long timeout / no request-latency coupling — the runner can take minutes
    for a large index build without competing with request budgets or the
    serving Lambda's 300s ceiling.
  • Fail the deploy on migration failure — a dirty/failed migration aborts
    the rollout instead of fail-opening into a half-broken serving fleet.
  • No per-instance race — the serving Lambda / Fargate task no longer runs
    migrations at all (DB_AUTO_MIGRATE=false on the workload), so concurrent
    cold starts can't race or dirty the row.

Scope / out of scope

Related

Activity

  1. cristim commented on Jul 28, 2026

    @cristim
    MemberAuthor

    Reviewed commit: be11bdcb5. Note: origin/main moved to 3e9660d06 during the review; re-verify against current main before changing code, since a finding may have been fixed or moved.

    The 2026-07-28 full-repo review re-derived this and confirmed it is still present on be11bdcb5. Adding two mechanisms this issue does not currently name, both of which widen the blast radius beyond "migrations compete with request latency".

    Where

    • internal/server/app.go:577-597 (ensureDB holds app.dbMu for its full duration, taken at :583)
    • internal/server/app.go:621 (runMigrationsBounded)
    • internal/server/app.go:199 (runMigrationsBoundedWith builds its own context.WithTimeout(context.Background(), timeout))
    • internal/server/app.go:140 (migrationsTimeout, default 120s)
    • Callers: internal/server/http.go:162, internal/server/http.go:228, internal/server/lambda.go:27

    What the review adds

    1. ensureDB holds dbMu for the entire migration, so migrations do not merely compete with request latency, they serialise every other request behind them. This issue describes the per-instance race and the fail-open behaviour, but not the mutex. Every concurrent request on that instance, including /api/scheduled/*, blocks on dbMu for up to the full migrationsTimeout.

    2. The migration runner deliberately ignores the caller's deadline. runMigrationsBoundedWith builds a fresh context.Background() bounded by migrationsTimeout (default 120s, app.go:140), so the caller's 30s HTTP deadline has no effect on the DDL. The caller therefore cannot abandon it; it can only time out itself while the migration continues.

    Failure scenario

    A deploy adds an index build that takes 90s.

    The first request after the cold start enters ensureDB, takes dbMu, and blocks in runMigrationsBounded. Every concurrent request, API traffic and /api/scheduled/* (all of which call ensureDB first), blocks on dbMu, hits its own 30s ceiling, and returns 503 Service temporarily unavailable.

    Two consequences beyond the latency:

    • On Lambda, the invocation is killed at the function timeout while the migration goroutine is still mid-DDL in a container that is then frozen. That is precisely the state that leaves schema_migrations.dirty = true, the terminal failure mode the 120s default was raised to avoid. So the band-aid can itself produce the outcome it was meant to prevent, because the two timeouts (function and migration) are set independently and the longer one wins the DDL.
    • The scheduled tick that fires during this window returns 503 and is simply lost. There is no retry on the /api/scheduled/* path, so a 90s migration window silently drops whatever collection or analytics tick landed in it, with no record that it was skipped.

    Fix direction

    Consistent with the one-shot deploy-time runner already proposed here. If lazy migration must stay as a transitional safety net, one cheap intermediate improvement: do the connect under dbMu and the migration outside it, so a long DDL degrades throughput rather than blocking every request behind a mutex, and report pending / failed through /health as it already does.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions