Skip to content

Repository files navigation

FailoverAI | Billing-safe execution layer for paid AI generation

English | 简体中文

Open-source execution layer that protects the billable outcome of image and video generation calls.

面向付费图像、视频生成调用的计费安全执行层。

FailoverAI is a reproducible reference implementation of one question most AI gateways skip: after you pay a provider to generate something, what exactly happened to the money when the call fails ambiguously?

The problem

A 503 is easy. The provider told you it cannot serve the request, so the gateway records the attempt and calls the backup.

A timeout is not easy. The provider may have accepted the request and started a billed generation before the connection dropped. If the client retries, it may pay for a second render of the same prompt. If it gives up, the user gets nothing even though the provider may deliver a finished asset. Every retry budget, fallback chain, and circuit breaker in a typical gateway treats this as a generic failure. It is not generic: it is a billing decision being made without evidence.

The same class of problem exists on the client side. A user double-clicking submit, or a client SDK retrying on a flaky network, sends the same request twice. Without a durable idempotency record, both copies run and both may be billed.

This matters most where each call is expensive: image and video generation, where a single generation can cost real money. For cheap text completions most users would rather just retry, and that is a reasonable trade-off this project does not try to talk you out of.

What is required

  • Distinguish "the provider could not serve this" from "we do not know what the provider did".
  • Never issue a second potentially billable call while the first one's outcome is unknown.
  • Make replays of the same request return the original job, and reject reused keys carrying different input.
  • Survive worker crashes without letting a stale worker overwrite a result.
  • Give operators a way to re-check the provider and resolve unknown outcomes, instead of leaving them as permanent dead ends.
  • Record everything needed to audit the bill: who was called, what came back, what was confirmed, what was written.

How it works

PostgreSQL is the single source of truth for jobs, idempotency records, provider attempts, leases, events, and the final result.

flowchart LR
  UI[React console] --> API[FastAPI API]
  API --> DB[(PostgreSQL authoritative state)]
  W[Python worker<br/>lease + heartbeat + reconcile sweep] --> DB
  subgraph R[Provider registry from PROVIDERS_JSON]
    P[primary adapter]
    B[backup adapter]
  end
  W --> P
  W --> B
Loading
sequenceDiagram
  participant C as Client
  participant A as API
  participant D as PostgreSQL
  participant W as Worker
  participant P as primary
  participant B as backup
  C->>A: POST /v1/jobs + Idempotency-Key
  A->>D: insert queued job or replay existing job
  W->>D: atomic claim (SKIP LOCKED)
  W->>P: one business attempt
  P-->>W: 503 availability error
  W->>B: one fallback attempt
  B-->>W: confirmed result
  W->>D: fenced final write + events
  C->>A: replay same request
  A-->>C: same job, no new generation
Loading
stateDiagram-v2
  [*] --> queued
  queued --> running: claimed
  running --> succeeded: confirmed result
  running --> failed: user or adapter error, providers exhausted
  running --> manual_review: timeout or disconnect after submit
  manual_review --> succeeded: reconcile confirms provider result
  manual_review --> queued: operator resolves retry
  manual_review --> failed: operator resolves fail
Loading

Failure classification, from provider.py upward:

Observed outcome Classification Next action
429 or 503 Provider availability error Call the backup once
Empty or malformed result Provider availability error Call the backup once
User request error User error Fail without fallback
Adapter TypeError Internal adapter error Fail without fallback
Timeout or interrupted connection Result unconfirmed Stop in manual_review

The guarantees behind the table:

  • Each provider receives at most one business attempt per job, enforced by a UNIQUE(job_id, provider) constraint.
  • The final write requires the current lease token and result IS NULL, so an expired worker cannot overwrite a live job.
  • Idempotency-Key is stored as an HMAC digest and compared against a canonical request hash under a row lock. Same key and same input returns the original job; same key and different input returns 409. The raw key never leaves the API.
  • Timeout or disconnect after submit marks the job result_unconfirmed and parks it in manual_review. The worker's reconcile sweep re-checks the provider's status (RECONCILE_INTERVAL_SECONDS, bounded by RECONCILE_MAX_AGE_SECONDS), and manual_review jobs can also be inspected and resolved by an operator.

Provider adapters

Providers come from a registry configured with PROVIDERS_JSON. The reference stack ships two local fake providers so the demo runs without keys, plus adapters for real services:

Adapter Service
fake Local deterministic provider used by the demo stack
openai_images OpenAI Images API (OPENAI_API_KEY, OPENAI_BASE_URL)
fal fal.ai (FAL_KEY)
replicate Replicate (REPLICATE_API_TOKEN)

Each adapter implements submit(), poll_status(), and cancel() and maps provider responses onto one outcome: confirmed, provider_availability, user_error, adapter_error, or result_unconfirmed. poll_status() is what makes reconcile possible: a job that ended ambiguously can be re-checked against the provider instead of guessed.

Operator API for unconfirmed jobs

Route Purpose
GET /v1/review-queue List jobs parked in manual_review
POST /v1/jobs/{id}/reconcile Re-check the provider status for one job now
POST /v1/jobs/{id}/resolve Operator decision: retry, fail, or confirm_result

30-second demo

cp .env.example .env
docker compose up --build

Open http://localhost:3001, select HTTP 503, click Generate image, watch primary → backup, then click Idempotent replay. The job ID and provider call count stay unchanged.

To see the billing-safe path, select Connection interrupted and submit again: the job parks in manual_review instead of silently retrying a possibly billed call. The console and the job trace read PostgreSQL records directly; the demo image is a deterministic local SVG, and no real provider is called.

Windows PowerShell:

Copy-Item .env.example .env
docker compose up --build

Screenshots

Confirmed failover and idempotent replay Ambiguous outcome isolated for review
Console trace: primary 503, backup confirmed, replay returned the same job Console trace: connection interrupted, job parked in manual review

API example

curl -X POST http://localhost:8000/v1/jobs \
  -H 'Content-Type: application/json' -H 'Idempotency-Key: demo-1' \
  -d '{"prompt":"A lighthouse over a cloud ocean","model":"fake-image-v1","timeout":3,"provider_policy":{"primary":"primary","backup":"backup"}}'
curl http://localhost:8000/v1/jobs/<job_id>

OpenAI-compatible image entry: POST /v1/images/generations. It waits at most OPENAI_WAIT_SECONDS, then returns a queryable job_id with HTTP 202.

What you get

  • A timeout or disconnect no longer becomes either a blind retry or a silent loss. It becomes a stored, auditable manual_review state that a reconcile pass or an operator can close.
  • A client retry of the same request returns the original job instead of a second generation.
  • A dead worker's job is reclaimed by a new worker, and the old worker's late write is rejected by the lease fence.
  • Every provider call lands in provider_attempts with HTTP status, provider request ID, latency, error class, and a billable flag, so a finance-facing audit does not have to scrape logs.

Where it fits

  • Image and video pipelines where a duplicate generation is a real charge, not just a retry.
  • Internal AI gateways that want one job API while keeping per-provider attempt records for operators.
  • Batch processing that must survive worker restarts without double-writing results.

These are reference scenarios, not claims about production deployments.

Documentation

Current boundaries

This is a reference implementation, not a production gateway. The default stack runs local fake providers; real adapters need their own keys and contracts to be verified per vendor. There is no authentication or multi-tenant isolation, and PostgreSQL is a single-node dependency. See limitations and the security notes.

Roadmap

Video job adapters beyond the current set, an authenticated multi-tenant boundary, richer review-queue workflows, and database high-availability guidance after the reference semantics remain stable.

License

MIT. See LICENSE.

About

Open-source gateway for reliable image, video and LLM jobs. | 面向可靠图片、视频和 LLM 任务的开源网关。

Resources

Security policy

Stars

138 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages