Open-source execution layer that protects the billable outcome of image and video generation calls.
面向付费图像、视频生成调用的计费安全执行层。
FailoverAI is a reproducible reference implementation of one question most AI gateways skip: after you pay a provider to generate something, what exactly happened to the money when the call fails ambiguously?
A 503 is easy. The provider told you it cannot serve the request, so the gateway records the attempt and calls the backup.
A timeout is not easy. The provider may have accepted the request and started a billed generation before the connection dropped. If the client retries, it may pay for a second render of the same prompt. If it gives up, the user gets nothing even though the provider may deliver a finished asset. Every retry budget, fallback chain, and circuit breaker in a typical gateway treats this as a generic failure. It is not generic: it is a billing decision being made without evidence.
The same class of problem exists on the client side. A user double-clicking submit, or a client SDK retrying on a flaky network, sends the same request twice. Without a durable idempotency record, both copies run and both may be billed.
This matters most where each call is expensive: image and video generation, where a single generation can cost real money. For cheap text completions most users would rather just retry, and that is a reasonable trade-off this project does not try to talk you out of.
- Distinguish "the provider could not serve this" from "we do not know what the provider did".
- Never issue a second potentially billable call while the first one's outcome is unknown.
- Make replays of the same request return the original job, and reject reused keys carrying different input.
- Survive worker crashes without letting a stale worker overwrite a result.
- Give operators a way to re-check the provider and resolve unknown outcomes, instead of leaving them as permanent dead ends.
- Record everything needed to audit the bill: who was called, what came back, what was confirmed, what was written.
PostgreSQL is the single source of truth for jobs, idempotency records, provider attempts, leases, events, and the final result.
flowchart LR
UI[React console] --> API[FastAPI API]
API --> DB[(PostgreSQL authoritative state)]
W[Python worker<br/>lease + heartbeat + reconcile sweep] --> DB
subgraph R[Provider registry from PROVIDERS_JSON]
P[primary adapter]
B[backup adapter]
end
W --> P
W --> B
sequenceDiagram
participant C as Client
participant A as API
participant D as PostgreSQL
participant W as Worker
participant P as primary
participant B as backup
C->>A: POST /v1/jobs + Idempotency-Key
A->>D: insert queued job or replay existing job
W->>D: atomic claim (SKIP LOCKED)
W->>P: one business attempt
P-->>W: 503 availability error
W->>B: one fallback attempt
B-->>W: confirmed result
W->>D: fenced final write + events
C->>A: replay same request
A-->>C: same job, no new generation
stateDiagram-v2
[*] --> queued
queued --> running: claimed
running --> succeeded: confirmed result
running --> failed: user or adapter error, providers exhausted
running --> manual_review: timeout or disconnect after submit
manual_review --> succeeded: reconcile confirms provider result
manual_review --> queued: operator resolves retry
manual_review --> failed: operator resolves fail
Failure classification, from provider.py upward:
| Observed outcome | Classification | Next action |
|---|---|---|
429 or 503 |
Provider availability error | Call the backup once |
| Empty or malformed result | Provider availability error | Call the backup once |
| User request error | User error | Fail without fallback |
Adapter TypeError |
Internal adapter error | Fail without fallback |
| Timeout or interrupted connection | Result unconfirmed | Stop in manual_review |
The guarantees behind the table:
- Each provider receives at most one business attempt per job, enforced by a
UNIQUE(job_id, provider)constraint. - The final write requires the current lease token and
result IS NULL, so an expired worker cannot overwrite a live job. Idempotency-Keyis stored as an HMAC digest and compared against a canonical request hash under a row lock. Same key and same input returns the original job; same key and different input returns409. The raw key never leaves the API.- Timeout or disconnect after submit marks the job
result_unconfirmedand parks it inmanual_review. The worker's reconcile sweep re-checks the provider's status (RECONCILE_INTERVAL_SECONDS, bounded byRECONCILE_MAX_AGE_SECONDS), andmanual_reviewjobs can also be inspected and resolved by an operator.
Providers come from a registry configured with PROVIDERS_JSON. The reference stack ships two local fake providers so the demo runs without keys, plus adapters for real services:
| Adapter | Service |
|---|---|
fake |
Local deterministic provider used by the demo stack |
openai_images |
OpenAI Images API (OPENAI_API_KEY, OPENAI_BASE_URL) |
fal |
fal.ai (FAL_KEY) |
replicate |
Replicate (REPLICATE_API_TOKEN) |
Each adapter implements submit(), poll_status(), and cancel() and maps provider responses onto one outcome: confirmed, provider_availability, user_error, adapter_error, or result_unconfirmed. poll_status() is what makes reconcile possible: a job that ended ambiguously can be re-checked against the provider instead of guessed.
| Route | Purpose |
|---|---|
GET /v1/review-queue |
List jobs parked in manual_review |
POST /v1/jobs/{id}/reconcile |
Re-check the provider status for one job now |
POST /v1/jobs/{id}/resolve |
Operator decision: retry, fail, or confirm_result |
cp .env.example .env
docker compose up --buildOpen http://localhost:3001, select HTTP 503, click Generate image, watch primary → backup, then click Idempotent replay. The job ID and provider call count stay unchanged.
To see the billing-safe path, select Connection interrupted and submit again: the job parks in manual_review instead of silently retrying a possibly billed call. The console and the job trace read PostgreSQL records directly; the demo image is a deterministic local SVG, and no real provider is called.
Windows PowerShell:
Copy-Item .env.example .env
docker compose up --build| Confirmed failover and idempotent replay | Ambiguous outcome isolated for review |
|---|---|
![]() |
![]() |
curl -X POST http://localhost:8000/v1/jobs \
-H 'Content-Type: application/json' -H 'Idempotency-Key: demo-1' \
-d '{"prompt":"A lighthouse over a cloud ocean","model":"fake-image-v1","timeout":3,"provider_policy":{"primary":"primary","backup":"backup"}}'
curl http://localhost:8000/v1/jobs/<job_id>OpenAI-compatible image entry: POST /v1/images/generations. It waits at most OPENAI_WAIT_SECONDS, then returns a queryable job_id with HTTP 202.
- A timeout or disconnect no longer becomes either a blind retry or a silent loss. It becomes a stored, auditable
manual_reviewstate that a reconcile pass or an operator can close. - A client retry of the same request returns the original job instead of a second generation.
- A dead worker's job is reclaimed by a new worker, and the old worker's late write is rejected by the lease fence.
- Every provider call lands in
provider_attemptswith HTTP status, provider request ID, latency, error class, and a billable flag, so a finance-facing audit does not have to scrape logs.
- Image and video pipelines where a duplicate generation is a real charge, not just a retry.
- Internal AI gateways that want one job API while keeping per-provider attempt records for operators.
- Batch processing that must survive worker restarts without double-writing results.
These are reference scenarios, not claims about production deployments.
- Architecture
- Failure semantics
- Idempotency
- Provider failover
- Demo runbook
- Security
- Limitations
- Deep dive: where the money goes after a timeout (Chinese)
This is a reference implementation, not a production gateway. The default stack runs local fake providers; real adapters need their own keys and contracts to be verified per vendor. There is no authentication or multi-tenant isolation, and PostgreSQL is a single-node dependency. See limitations and the security notes.
Video job adapters beyond the current set, an authenticated multi-tenant boundary, richer review-queue workflows, and database high-availability guidance after the reference semantics remain stable.
MIT. See LICENSE.

