Skip to content

libsql-server: add an operation-owned namespace fence with positive drain - #35

Closed
tszymczyszyn-shopify wants to merge 33 commits into
v0.9.30-shopify-patchesfrom
dozer/task-928
Closed

tszymczyszyn-shopify wants to merge 33 commits into
v0.9.30-shopify-patchesfrom
dozer/task-928

Conversation

@tszymczyszyn-shopify

Copy link
Copy Markdown

Summary

This adds a durable namespace fence to libsql-server, owned by one operation at a time. It is meant to be the data-plane authority boundary when a database moves between servers. The fence lets one operation:

  1. stop new writes on the source and positively prove that no transaction admitted before the fence can still commit;
  2. create a target namespace that is quarantined from the first moment anything can see it, and writable only through an import capability the server issues to that operation;
  3. make a validated target readable while its writes stay closed;
  4. stop source SQL, dump and replication reads before target writes are published; and
  5. publish target writes with an idempotent compare-and-swap transition whose result can be inspected after a restart or a lost response.

The existing block_writes / block_reads config flags are process configuration, so they are not used as the authority. They are snapshotted once per program, no operation owns them, and they don't cover lifecycle, dump, admin-shell or replication paths.

Design

The full contract is in docs/NAMESPACE_FENCE.md. Main points:

  • States. Source: SOURCE_DRAINING → SOURCE_WRITE_FENCED → SOURCE_READ_DRAINING → SOURCE_READ_FENCED, with explicit CAS-guarded precommit rollback to SOURCE_WRITE_FENCED and RELEASED. Target: TARGET_QUARANTINED → TARGET_IMPORT_DRAINING → TARGET_VALIDATING → TARGET_WRITE_FENCED → TARGET_WRITABLE. TARGET_WRITABLE has no reverse transition, and TARGET_ABORTED handles abandonment. Any control state that is unknown, corrupt or indeterminate becomes UNKNOWN_UNAVAILABLE, which denies everything.
  • Durable CAS. Fence state lives in two new metastore tables (a versioned protobuf record plus command receipts keyed by (operation_id, command_id), each with a request fingerprint). The revision survives restart. Fence transitions and config writes serialise through BEGIN IMMEDIATE on the metastore. Replay lookup runs before any owner or revision check.
  • Authoritative write gate. The check sits in ManagedConnectionWalWrapper::begin_write_txn, before the writer slot is acquired, and uses admission generations. Every program (CoreConnection::run, each with_raw call, vacuum) records the generation it was admitted under, and every read transaction records the one it was opened under; a write transaction starts only when both equal the gate's current generation. So a transaction or program admitted under an older generation can never upgrade to a writer, even after the gate reopens (MIGRATION_WRITE_FENCED / stale_transaction). Refusals return SQLITE_AUTH, which SQLite's busy handler does not retry, and the typed reason travels out of band to the program layer, which reports it as the fence error. Write and DDL statements are also refused early, per statement, against the live gate.
  • Writer queue. Queue entries and the write slot carry an operation class. A queued writer re-checks its admission every time it takes the connection manager's lock, and the fence controller wakes every write queue of the namespace after each change of the write generation, so a writer that queued before the fence leaves the queue with MIGRATION_WRITE_FENCED rather than waiting for the slot and writing. Checkpoints are maintenance and are never refused. VACUUM is not maintenance and is skipped while writes are fenced.
  • Positive drain. AcquireSourceWriteFence closes write admission in memory (an INSTALLING gate, which also wakes queued writers) before it persists SOURCE_DRAINING, then waits on the connection manager's release notification for the writer that held the slot when admission closed. Elapsed time and the transaction timeout are never taken as evidence. Once no connection holds the slot for a write, it reads the frozen boundary (replication log id and last committed frame) under the slot lock and commits SOURCE_WRITE_FENCED with it. At the deadline it answers DRAINING and admission stays closed; with force_rollback it rolls the writer back through its registered rollback handle and still waits for the actual release. Replaying the same command resumes the drain. The whole command runs on its own task under the namespace's transition lock, so a lost response interrupts nothing. The default deadline is --namespace-fence-default-write-drain-ms (30 s).
  • Drain replay safety. A non-final DRAINING receipt resumes only while the durable record is still in that command’s draining state at the receipt’s exact revision. Release, read-fence clear, target abort, and a later drain cycle supersede the old receipt; replay then returns the historical result without restarting cancellation, rollback, or completion. A replay that does resume and finish remains classified as Resumed, and the admin response reports it as replayed.
  • Read fence. Read work holds a read lease on the namespace's fence controller for as long as it runs, and the lease is taken only after checking the gate under the controller's lease lock, so once SetSourceReadFence has closed read admission (an in-memory read-closing gate, published before SOURCE_READ_DRAINING is persisted) the set of leases can only shrink. Each SQL program holds one, including a Hrana cursor still producing rows, as does each describe and each admin-shell query. Attaching a namespace is a read of that namespace: the connection keeps a lease on every namespace it has attached for each later program until it detaches it. A connection idle inside a transaction holds none; its next program is refused with MIGRATION_READ_FENCED and its transaction rolled back. At the deadline (--namespace-fence-default-read-drain-ms, 30 s) running programs are cancelled through the connection's existing progress-handler cancel flag and report the read fence, and the drain still waits for the actual releases; if one does not come it answers DRAINING and a replay resumes. /beta/listen is refused where reads are denied and its stream ends when reads are fenced. /dump is checked against the gate before a connection is created and its export holds a lease: a running dump may finish, and at the deadline it is cancelled before its next row or in the middle of a write blocked on a peer that stopped reading, and its response body is aborted rather than completed, so a partial dump never ends with COMMIT;. Every replication call is refused at its start with FAILED_PRECONDITION and the x-libsql-fence-code metadata while streams are denied (counted, and logged at most once a minute per namespace). log_entries, snapshot and the frames of batch_log_entries hold a lease through a stream wrapper whose watcher ends the stream as soon as the gate closes, dropping the inner stream and the lease itself, so the lease is released even if the peer never reads again; the next poll yields the typed status. With the fence enabled, the RPC server and the user HTTP server send HTTP/2 keepalive pings (--namespace-fence-keepalive-interval-s, 30 s; 20 s timeout) so dead peers are detected.
  • Target creation. NamespaceStore::create_target_quarantined (also what the admin command runs) creates a target already in TARGET_QUARANTINED. For a name with no fence state, an in-memory target-creation gate is published before the metastore transaction that writes the marker, config row, record and receipt; it refuses every class but maintenance and observability and keeps create, fork and namespace loading from storing or setting anything up under the name. The committed quarantine gate replaces it, and only then is the config made visible (from the durable row) and the namespace loaded, so its first connection is created behind the quarantine. A name the server already knows is refused without touching its gate. The command runs on its own task; a replay completes a creation whose commit was not acknowledged, or one interrupted between its marker and its commit. AbortQuarantinedTarget leaves every normal class denied.
  • Import and seal. NamespaceStore::open_import_session issues a MigrationCapability (server-created, fields private, bound to the namespace, the owning operation and the current fence revision) and opens a connection that carries it. The WAL admits that connection's write transactions only while the fence is still TARGET_QUARANTINED, owned by that operation, at that revision, and the capability is live; every other connection, including the admin shell's raw access, is refused. ImportSession::with_raw runs a closure on that connection, counted as an import writer while it runs, and ImportSession::load_dump runs the server's existing dump loader under the capability. SealTargetImport closes import admission in memory, persists TARGET_IMPORT_DRAINING (the revision moves, so every issued import capability is invalidated), waits on release notifications for running import calls and for any import transaction still holding the write slot (with force_rollback, rolls it back at the deadline), and persists TARGET_VALIDATING. At the deadline it answers DRAINING and the target stays closed; only the owning operation's seal resumes it. Import never reopens.
  • Validation and publication. NamespaceStore::open_validation_session issues a Validate capability and a query_only connection for the owner at the current revision. Every call re-checks the live capability; the WAL independently refuses every validation write even if raw code disables query_only. A new RecordTargetValidation adds the server's current replication log id/frame and SQLite page count to the durable result, while a durable command-key preflight keeps exact replay and command-id conflict independent of a fresh snapshot. A successful result is required before PublishTargetReadableWriteFenced opens normal reads and clears the legacy read mirror. EnableTargetWrites commits TARGET_WRITABLE, publishes a new admission generation before responding and restores the original legacy blocks. It is replayable, returns ALREADY_APPLIED for a new same-goal command, and cannot be reversed; transactions opened before publication cannot upgrade after it.
  • Admin API. libsql-server/src/http/admin/fence.rs serves the contract of docs/NAMESPACE_FENCE.md §4 on the admin listener. GET /v1/fence/capabilities is always served (protocol version, whether fences are enabled, served commands, states, proxy stable_code support, server build and instance id, active fence count, and whether the metastore was restored from its backup at startup, with the generation) so preflight can reject a server that cannot take part. With --enable-namespace-fence, GET /v1/namespaces/:ns/fence (InspectFence: record, the owning operation's receipts or ?receipts=all, live admission and drain counters, the live replication log id) and one POST route per command are served; every command goes through the same store entry point as the internal API, so replay precedes every other check and responses carry the stored receipt with replayed: true. Command routes and validation-query (one read-only program under the operation's validation capability, bounded to 10 000 rows, refused rather than truncated) refuse to run without an admin auth key (admin_auth_required) and on a replica (not_primary). Bodies are strict (invalid_argument for unknown or malformed fields); a target creation refuses restore options and dump URLs (restore_not_allowed). Statuses follow the outcome: 200 for APPLIED/ALREADY_APPLIED, 202 for DRAINING, 409/412/403/423 for refusals, each with the current fence view.
  • Lifecycle interlocks. Config changes, delete, reset, fork (as source or destination), create over an existing record (with or without a dump_url), linking to a shared schema and schema migration are refused while a namespace's fence denies lifecycle work, with the fence code and 423 on the admin API. One check reads the fence registry, so it sees in-memory gates (a closing transition, a target being created, an indeterminate commit) and never loads the namespace; paths that persist through the metastore are refused again inside its transaction. A fork reads the source without a read lease, so it holds the source's transition lock for its whole run: a write or read fence command either completes first (and the fork is refused) or waits for the fork. Reset, which writes nothing to the metastore, is checked under the same lock. A fork destination is checked before anything is stored and again with its namespace entry held; before this, a fork onto an existing namespace that was not loaded would publish its config in memory and remove its directory. A dump URL is not fetched for a refused create. Shared-schema databases and namespaces linked to one cannot be fenced and a fenced namespace cannot be linked; registering a migration job also checks the schema and every linked namespace, so a link written by a binary that does not know fences cannot lead to a migration step refused halfway through a job.
  • Typed outcomes. Stable codes (MIGRATION_WRITE_FENCED, MIGRATION_READ_FENCED, MIGRATION_TARGET_QUARANTINED, FENCE_REVISION_MISMATCH, …) are mapped consistently across admin HTTP, user HTTP (423 plus an additive code field, and detail where there is one), Hrana, gRPC (FAILED_PRECONDITION plus metadata) and the replica write proxy, through an additive stable_code field on proxy.Error: the primary sets it (with code = SQL_ERROR) on a refused step or program, and answers a namespace it refuses before a program runs with FAILED_PRECONDITION and the code, never UNAVAILABLE; a replica maps either back to the same fence error and answers its client exactly as the primary would, without retrying. An older primary never sets the field and its errors keep their old mapping. Capability discovery reports proxy_stable_code: true. On Hrana (/v1, /v2, /v3, cursors, WebSocket) a refused step and a refused request are Hrana errors whose code is the stable code, and the stream stays usable; the legacy / API, which has no per-step codes, answers a batch with a refused step as a whole with 423; docs/NAMESPACE_FENCE.md §6.0 has the table. 401 and 404 are unchanged and carry no fence code. Fence denials never map to 500, 503 or UNAVAILABLE.
  • Replicated fence. metadata.DatabaseConfig gains an additive fence field (ReplicatedFence { state, revision }). The primary's replication hello fills it from the namespace's published gate while a fence is active and leaves it empty otherwise; it is never stored, and the rest of the configuration hello sends is unchanged. Both new fields are skipped by older peers, and a message without them encodes exactly as before.
  • Replica servers. A replica server whose primary refuses replication with a fence code (a refused hello, log_entries or snapshot, or a stream the primary ended) refuses local reads of its copy with the same code over every user protocol, and cancels reads it had admitted. Instead of reconnecting every second it retries at 1 s doubling to 15 s, counting each refused call (libsql_server_replica_fence_refusals_total{code}), and serves reads again once the primary answers hello. A replica that has never held a namespace and is asked for a name the primary refuses (for example a quarantined target) answers with the primary's code at once instead of retrying the handshake, and creates nothing: the namespace directory the attempt made is removed, and neither the in-memory config entry nor a fence controller is left behind, so the name is created normally once the primary admits it.
  • Recovery. Gates are installed from a registry that sits outside the namespace cache, before the first connection is created, so eviction and reload keep them. An integration test drives a one-entry cache under capacity pressure, reloads the fenced namespace through the user protocol, and observes the exact same revision, generation and write denial. Metastore recovery paths (handle() defaults, undecodable rows, filesystem recovery, destroy_on_error, backup restore) fail closed, helped by a per-namespace marker file.
  • Restart. A restart at any persistence boundary recovers the state before the command or the state it committed, and a namespace is open after a restart only if an opening transition committed. A source that restarts while draining completes the drain when the same command is replayed. A same-path integration restart reconstructs SOURCE_WRITE_FENCED, SOURCE_READ_FENCED, TARGET_QUARANTINED and TARGET_WRITE_FENCED before traffic, preserves typed user denials and durable exact-command receipts, and leaves TARGET_WRITABLE readable and writable. Crash recovery rebuilds a namespace's replication log under a new log id, so the frozen boundary names the log that is live when the drain is proven; the record keeps the log id the fence was acquired against, and a caller can see the rebuild by comparing the two: InspectFence and every admin response report the live log as incarnation.current_log_id.
  • Metastore restore provenance. The metastore's bottomless restore used to be run and its outcome discarded. The server now keeps whether it recovered the metastore from the backup and the generation it restored from, and reports it in the capability endpoint (metastore), in every fence view (provenance.metastore_restored_from_backup, metastore_restored_generation), in the libsql_server_metastore_restored_from_backup gauge, and in a startup warning. A restored record is still never trusted over a newer marker (the marker comparison makes such a namespace UNKNOWN_UNAVAILABLE); the provenance tells an operator why. When destroy_on_error rebuilds the metastore, the restore of the rebuilt one is what is reported.
  • Incident adoption. POST /v1/namespaces/:ns/fence/adopt (AdoptFence) hands an unfinished operation's fence to a new operation when the original owner's control record is lost. It needs the admin credential and a separate secret, --namespace-fence-adoption-key, in the x-libsql-fence-adoption-key header (kept as a SHA-256 digest and compared without an early exit; without it adoption is disabled), plus the current owner, two distinct approvers, an incident reference and a reason, which are stored in the record and receipt and written to one audit event under the libsql_server::fence::audit tracing target. It moves the owner and the revision and nothing else: every admission stays as it was, the old owner's capabilities are revoked and its commands are refused, and RELEASED, TARGET_WRITABLE and TARGET_ABORTED cannot be adopted. After a metastore rollback (the marker ahead of the metastore) it re-establishes the record the marker holds under the new owner, which makes the namespace available again behind the same gate. If the metastore has no config row for the name at all, adoption is refused with namespace_config_missing rather than recreating a configuration (JWT key, size limit, durability) from defaults; the namespace stays unavailable until an operator restores its config row.
  • Observability. Every fence command the server answers (committed, replayed or refused) emits one structured event under the libsql_server::fence::audit tracing target, with namespace, operation and command id, command, outcome, revisions, state before and after, drain kind and duration, forced actions, replay/conflict and server instance (and approvers, incident reference and reason for an adoption). Metrics use bounded labels only, never a namespace, operation id, command id, revision or caller: libsql_server_fence_transitions_total{command, outcome}, libsql_server_fence_drain_duration_seconds{kind} (from the start of a drain to its proof), libsql_server_fence_forced_total{kind = rollback | sql_cancel | dump_cancel | stream_termination}, libsql_server_fence_replays_total{result = replay | conflict}, libsql_server_fence_denials_total{code, surface} counted at each surface that refuses (http, hrana, rpc, proxy, dump, replication, admin_shell, lifecycle), libsql_server_fence_adoptions_total, and two gauges computed from the fence registry when /metrics is read: libsql_server_fence_namespaces{role, state} and libsql_server_fence_oldest_active_age_seconds.
  • Deployment. The feature is off by default behind --enable-namespace-fence. While a fence is active, the legacy block_* fields of the stored config are mirrored from the fence state as a best-effort guard against older binaries (a config write cannot overwrite them while the fence denies lifecycle work, and the fence row's foreign key makes an older binary's delete fail), and a capability endpoint lets preflight checks reject unsupported servers. Once the operation has released the namespace or enabled target writes, the stored row holds the namespace's own values again, and a restart keeps any config written since rather than the values saved when the fence was acquired.
  • Reuse. Target import and validation go through a reusable internal API (create_target_quarantined, open_import_session with ImportSession::{with_raw, load_dump}, and open_validation_session with ValidationSession::with_raw), with typed FenceError refusals, so bulk import can build on it.

Code-path audit

docs/NAMESPACE_FENCE.md §14 maps every path that can reach namespace data or lifecycle to how the fence covers it. That includes connection core and manager, the HTTP, Hrana and RPC paths, the write proxy, metastore recovery, namespace store lifecycle, admin endpoints, dump, the replication services, the admin shell, the schema scheduler, dump load, ATTACH and internal raw connections.

Every write, from whichever path, meets the one authoritative gate in begin_write_txn, before the writer slot is taken. SQL reads, dumps and replication streams hold a lease on the namespace's fence controller for as long as they run. Lifecycle operations go through FenceRegistry::check_lifecycle, and where they persist they are checked again inside the metastore transaction. Internal maintenance connections (checkpoints, the storage monitor, the replication logger, bottomless) are classed as maintenance and are never blocked.

Acceptance tests

Each acceptance requirement and the tests that cover it (docs/NAMESPACE_FENCE.md §17 names every test; all pass). Race tests park the task at named #[cfg(test)] hook points or use simulated time; none relies on a sleep.

# Requirement Covered by
1 Two operations race to acquire; one owner, typed conflict for the loser namespace::fence::drain::tests::acquire_race_single_owner, namespace::fence::transition::tests::acquire_race_single_owner, tests::fence::admin::concurrent_acquire_one_owner
2 Active writer commits or is rolled back before the acknowledgement; nothing commits after namespace::fence::drain::tests::{active_writer_commits_before_ack, forced_rollback_before_ack, no_commit_after_ack, installing_gate_closes_writes_before_persisting}
3 Autocommit, explicit transactions, queued writers, batches, DDL, schema jobs, old WebSockets, read-to-write upgrades cannot bypass connection::connection_manager::fence_tests::*, schema::scheduler::test::fence::*, tests::fence::protocol::{old_ws_session_cannot_write, batch_denied_mid_batch}
4 A program that captured config before the fence is refused at the WAL connection::connection_manager::fence_tests::wal_gate_rejects_program_admitted_before_fence
5 Pre-fence transactions cannot write after release or publication connection::connection_manager::fence_tests::stale_generation_cannot_write_after_release, namespace::fence::target::tests::stale_generation_cannot_write_after_enable_writes
6 Acquisition timeout returns DRAINING and admission stays closed namespace::fence::drain::tests::{deadline_returns_draining_and_stays_closed, replay_of_draining_resumes_and_completes}
7 Restart at every persistence boundary; indeterminate persistence keeps the gate closed namespace::fence::tests::{restart_at_each_boundary, restart_in_draining_waits_for_the_same_command, indeterminate_commit_keeps_gate_closed, read_and_target_boundaries::restart_at_each_read_and_target_boundary}, namespace::fence::controller::tests::indeterminate_*, tests::fence::lifecycle::{restart_keeps_fence, restart_after_enable_writes_stays_writable}
8 Evict and lazily reload a fenced namespace; identical admission tests::fence::lifecycle::evicted_namespace_reloads_same_gate, namespace::store::fence_tests::evicted_namespace_reloads_with_the_same_controller
9 Filesystem recovery, destroy_on_error, undecodable records, missing target quarantine, metastore rollback fail closed with provenance namespace::meta_store::fence_tests::recovery::*, namespace::meta_store::fence_tests::provenance::* (including a real bottomless restore), http::admin::fence::tests::restore_provenance_is_reported
10 Wrong owner, stale revision, invalid role/state, replay, command-id reuse; replay before revision check namespace::fence::transition::tests::{exhaustive_owner_commands, wrong_owner_is_refused, stale_revision_is_refused, role_mismatch, exact_replay_after_revision_advanced, command_id_reuse_with_different_fingerprint_conflicts}, namespace::meta_store::fence_tests::concurrent_cas_has_exactly_one_winner
11 Target creation raced with SQL, dump, replication and lifecycle is never observable namespace::fence::target::tests::{create_race_never_observable, creating_gate_refuses_before_commit, create_replay_completes_interrupted_creation, indeterminate_create_is_completed_by_replay}
12 Only the matching import capability writes a quarantined target; admin credentials and admin shell cannot namespace::fence::import::tests::import_requires_matching_capability, admin_shell::fence_tests::admin_shell_cannot_write_quarantined, tests::fence::admin::target_walk_over_http
13 Seal drains import writers, reaches TARGET_VALIDATING, cannot resume import; publication needs a validation receipt and is idempotent namespace::fence::import::tests::{seal_waits_for_import_writers, seal_deadline_leaves_import_draining_until_replayed, sealed_target_rejects_import}, namespace::fence::target::tests::{validation_session_is_read_only, publish_requires_validation_receipt, publish_is_idempotent}
14 Enable writes is idempotent, survives restart and response loss, irreversible namespace::fence::target::tests::{enable_writes_idempotent_and_irreversible, enable_writes_survives_restart}, namespace::fence::transition::tests::target_writable_is_irreversible
15 A lost EnableTargetWrites response is resolved from the receipt/state namespace::fence::target::tests::enable_writes_response_loss_resolved; classifying an unanswerable response as COMMIT_UNKNOWN is the caller's (see Limits)
16 Read fence drains SQL, dump, log_entries, snapshot, incl. dead peers and forced termination namespace::fence::read::tests::*, namespace::fence::stream::tests::*, admin_shell::fence_tests::admin_shell_read_denied, tests::fence::protocol::dump_codes
17 Delete, reset, fork, restore, config and schema mutation refused tests::fence::lifecycle::lifecycle_rejected_while_fenced, namespace::store::fence_tests::{reset_refused_while_fenced, lifecycle_refused_while_fenced}, schema::scheduler::test::fence::*
18 Codes through HTTP, Hrana, RPC, dump, replication and replica write proxy; distinct from auth/not-found; old peers compatible; no retry loops tests::fence::protocol::{http_codes, hrana_http_codes, hrana_ws_codes, dump_codes, auth_and_not_found_distinct, replica_proxy_preserves_code, denial_not_retried, replication_codes, replica_backs_off_on_fence_code, replica_lazy_creation_refused_by_fence}, rpc::proxy::fence_tests::*, libsql-replication rpc::test::{proxy_error_stable_code_is_additive, replicated_fence_is_additive}
19 Corrupt or unknown durable fence state fails closed namespace::fence::store::tests::corrupt_rows_are_unavailable, namespace::fence::record::tests::{unknown_format_version_is_rejected, garbage_is_rejected, revision_column_must_match}, namespace::meta_store::fence_tests::corrupt_fence_row_fails_closed
20 Metrics and structured audit logs, with bounded labels namespace::fence::audit::tests::{audit_event_fields, adoption_event_fields, metrics_and_bounded_labels}, tests::fence::observability::metrics_and_labels
21 Capability discovery and mixed-version protection tests::fence::admin::{capabilities, capabilities_when_disabled}, namespace::fence::tests::legacy_mirror::legacy_mirror_and_fk_guard; an older binary is not run (see Limits)
22 Adoption is two-person/audited, keeps admission closed, cannot reverse publication namespace::fence::tests::adoption::*, tests::fence::admin::{adopt_over_http, adopt_disabled_without_key}
— Import API usable by bulk import namespace::fence::import::tests::import_session_loads_dump_into_quarantined_target (tables, keys, a foreign key, indexes, autoincrement, a trigger, a view and FTS5)

Limits

docs/NAMESPACE_FENCE.md §18 lists what this change doesn't prove:

  • mixed-version behaviour is tested at the data and wire level (legacy mirror, foreign-key guard, protobuf unknown-field handling, capability endpoint), not by running an older binary against the same metastore;
  • test data is synthetic: multi-table schemas with keys, indexes, triggers, views and an FTS5 table, under the server's default settings plus small caches and short drain deadlines in the fence tests; real production schemas and settings are not part of the suite;
  • client SDK retry policy was not audited; the server only guarantees that fence denials use codes that are not conventionally retried;
  • whether a caller treats a response that neither replay nor inspection can answer as commit-unknown is the caller's policy; the server guarantees that replay and inspection answer while it is reachable;
  • two-person adoption is enforced as a separate key plus two recorded approvers, not as verified identities;
  • data that has already been delivered can't be recalled, and an older replica server stops replicating under a read fence but keeps serving local reads of what it already holds;
  • a newer replica server installs its local read denial when its ended stream reaches it, a moment after the primary's read drain completes, so the drain does not prove that replicas have stopped serving; one partitioned from the primary keeps serving what it holds;
  • rollback detection depends on the namespace directory's marker surviving; if the metastore and the namespace directory are both lost or restored together, the server cannot tell that a newer fence existed;
  • releasing and pinning a server build that contains the fence is outside this change.

Commits

The series is 31 commits on main at e4beacaa266f, in dependency order so they can be reviewed one at a time. 75 files change (+28,748 / −293), most of it the new namespace/fence module, its tests and the contract document.

  1. docs: the contract (docs/NAMESPACE_FENCE.md).
  2. Durable state: fence types and pure transition logic; metastore persistence with CAS and receipts; fail-closed metastore recovery.
  3. Write side: registry and controller installed before the first connection; the WAL write gate by admission generation; classed writer queue; positive source write drain; indeterminate-commit reconciliation and restart at every boundary.
  4. Read side: SQL read leases; dump and replication stream leases.
  5. Targets: quarantined creation; import capabilities and the seal drain; validation, publication and write enable.
  6. Surfaces: admin API and capability discovery; lifecycle interlocks; additive libsql-replication protocol fields; typed outcomes over HTTP, Hrana, dump, RPC and the replica write proxy; replica servers honouring the replicated fence and refusing lazy creation of a fenced name; restore provenance and the live log id.
  7. Operations: incident adoption; legacy-mirror and crash-boundary tests; restart and eviction integration tests; metrics and audit log.
  8. Finish: acceptance tests completed; transitional dead-code allowances removed; replay/retry edge cases fixed; restart-test expectations updated for resumed drain attribution.

Testing

Run locally with Rust 1.85.0, cargo-nextest and RUSTFLAGS="-D warnings --cfg tokio_unstable", as upstream CI does. At the final commit (43500cece3bf):

  • cargo fmt --all -- --check: clean.
  • cargo build -p libsql-server --tests: clean. The only warning is the existing clang warning from the vendored, generated SQLite C source.
  • cargo nextest run -p libsql-server --no-fail-fast: 418 passed, 3 skipped (existing #[ignore] tests), 0 retries. The base commit has 199 tests, so this series adds 219.
  • cargo nextest run -p libsql_replication: 15 passed, including the bootstrap check that the generated protobuf code matches the .proto files, and rpc::test::{proxy_error_stable_code_is_additive, replicated_fence_is_additive}.
  • cargo nextest run -p libsql-server --lib --retries 0 'namespace::fence': 142 passed in each of 10 consecutive runs.
  • cargo nextest run -p libsql-server --test tests --retries 0 'fence::': 28 passed. These are the admin, protocol, lifecycle, replica and observability integration tests.

On c3f2e47bca0c, before the last two commits (which change only libsql-server and the doc):

  • cargo check -p libsql-server -p libsql_replication --all-targets: clean.
  • cargo check --all-targets --all-features, workspace-wide: clean.
  • cargo tree -p libsql-server -i openssl: exits 101, so there is no OpenSSL in the tree.

The patch series was also applied with git am to a fresh clone at e4beacaa266f. It reproduced the final tree (9a52dc698ca6) and every per-commit tree. In that checkout, with its own target directory, formatting, the test build, the 142 fence unit tests, the 28 fence integration tests and the full libsql-server suite (418 passed, 0 retries) all passed again.

How the race and restart tests work:

  • Race tests park a task at named #[cfg(test)] hook points (FenceTestHooks, doc §16), or run under turmoil's simulated time. Admission against persistence, drain against commit, and response loss are ordered by those hooks, never by sleeps.
  • Crash-restart tests run each server lifetime on its own runtime and drop it with no shutdown code while a command is parked. The restart then takes the real dirty-recovery path on the same directory, at 17 write-fence boundaries and 29 read-fence, seal and write-enable boundaries.
  • Each guard was checked by removing it and seeing its tests fail. Examples: the WAL capability check, the drain's wait for the active writer and for read leases, the queued-writer wake, the stream watcher, the fork source check, the reset check, the per-command audit call, and the replica-side code mapping.

Pre-existing timing flakes sometimes needed a retry under full-suite load on intermediate commits. The final commit did not need any. One is embedded_replica::replica_no_resync_on_restart, a wall-clock ratio assertion that also fails on the base commit. The others are cluster::schema_dbs::schema_migration_basics and connection::connection_core::test::release_before_timeout.

penberg and others added 30 commits March 25, 2026 10:15
Add docs/NAMESPACE_FENCE.md, the contract and design for a durable,
operation-owned namespace fence: source and target state machines and
permission matrix, admin API, stable outcome codes and their mapping to
HTTP, Hrana, gRPC, the write proxy and replication, metastore schema and
CAS semantics, write admission generations and positive drain, read
fence, quarantined target lifecycle and the reusable import capability,
restart and eviction rules, fail-closed metastore recovery, deployment
flag and legacy-binary protection, code-path coverage, test strategy and
the acceptance-test map.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Add the I/O-free core of the namespace fence described in
docs/NAMESPACE_FENCE.md:

- state.rs: roles, fence states (including the non-durable UNFENCED,
  ABSENT and derived UNKNOWN_UNAVAILABLE), operation classes and the
  permission matrix.
- outcome.rs: stable outcome codes, their admin HTTP, user HTTP, Hrana,
  gRPC and proxy mappings (fence denials are never 5xx, 429 or gRPC
  UNAVAILABLE), bounded detail reasons and FenceError.
- command.rs: fence commands and requests, and the canonical SHA-256
  request fingerprint over everything except command_id.
- record.rs: fence records, command receipts and the on-disk marker,
  with a strict protobuf encoding (proto/namespace_fence.proto) whose
  decoder rejects unknown versions, enum values, malformed ids and
  self-contradictory records instead of defaulting.
- transition.rs: the pure apply() and complete_drain() functions.
  Replay and fingerprint conflicts are decided before owner, role,
  state and revision; TARGET_WRITABLE has no reverse transition;
  adoption changes only the owner.

Nothing outside the module uses it yet; the metastore, controller and
protocol layers follow.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Add the additive `namespace_fences` and `namespace_fence_receipts`
tables (created with --enable-namespace-fence, loaded and enforced
whenever they exist) and run every fence command as a compare-and-swap
in one BEGIN IMMEDIATE metastore transaction: read the record, the
command's receipt and the config row, decide with the pure transition
function, then commit the record, the receipt and the legacy block_*
mirror together. The revision is stored, so it survives restart.

Stored state is read strictly: an unknown format version, an
undecodable payload, a revision column that disagrees with its payload,
or a per-namespace `.fence` marker that is ahead of the metastore makes
the namespace UNKNOWN_UNAVAILABLE instead of guessing. The marker is
written after each commit (before it, for target creation) and
rewritten on load when it fell behind.

Ordinary config writes and deletes now take BEGIN IMMEDIATE and refuse
while the fence denies lifecycle operations, and a config write that
failed to persist is no longer published to the in-memory config.
Receipts of finished operations are pruned after a retention period.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
When namespace fences are in use (the flag is on, the fence tables exist,
or a namespace directory holds a fence marker), a namespace whose state
startup cannot establish is registered UNKNOWN_UNAVAILABLE instead of
being skipped or given a default config:

- an undecodable config row, fence row or marker;
- a directory with a marker that the metastore has no trustworthy record
  for: filesystem recovery, a metastore rebuilt by destroy_on_error, a
  metastore restored from an older backup or without fence tables, and a
  target whose creation was interrupted.

Such a namespace is refused by lookups, config writes and deletes with
FENCE_STATE_UNAVAILABLE, is reported by InspectFence, and is settled only
by a fence command that commits for it.

MetaStore::handle() no longer default-creates on read paths: the new
non-creating MetaStore::lookup serves NamespaceStore::with, fork's source
and ATTACH authorisation. handle(), used only by creating paths, refuses
unavailable names and names without a config whose directory holds a
marker. Filesystem recovery skips marked directories, destroy_on_error
moves the broken metastore aside instead of deleting it, and a fence row
for a namespace name that cannot be decoded stops startup.

Servers that never used fences recover as before.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…t connection

Add the in-memory authority for namespace fences:

- `FenceRegistry`, held by `NamespaceStore` outside the namespace cache and
  seeded from `MetaStore::load_fences()` before anything is served, so an
  evicted and reloaded namespace gets the controller it had. Namespaces
  without fence state get an UNFENCED controller on first load; deleting a
  namespace drops its controller.
- `FenceController`: a per-namespace transition lock and a `watch` gate
  (`GateSnapshot`: the durable fence, a write generation and an
  indeterminate flag). Commands commit in the metastore, are published to
  the gate and only then answered, on their own task so a lost response
  does not lose the publication. An error before COMMIT leaves the gate
  unchanged; a failed COMMIT closes the gate and refuses other commands
  with FENCE_COMMIT_INDETERMINATE until the same command is replayed.
- `FenceConnState`, bound to the controller for every connection a
  `MakeLegacyConnection` opens, starting with its held connection, and
  shared with the connection's WAL wrapper. The checks that use it land in
  the next commit.
- `cfg(test)` `FenceTestHooks` with the named hook points of the design.

`NamespaceStore::with` and `make_namespace` refuse a namespace whose fence
state is unavailable before any setup.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…tion

Make ManagedConnectionWalWrapper::begin_write_txn the authoritative
namespace fence check. Before queueing for the write slot it requires the
live gate to admit the connection's operation class and the program and
its read transaction to have been admitted under the gate's current write
generation. A refusal returns SQLITE_AUTH (not BUSY, so SQLite does not
retry it, and before acquire(), so no slot is released that was never
held) and leaves the typed outcome in the connection's FenceConnState.

- FenceConnState gains begin_program, begin_read_txn and admit_write.
  CoreConnection::run, every with_raw call and vacuum_if_needed start a
  program; the WAL wrapper records the generation of each new read
  transaction before its snapshot is taken.
- The Vm refuses Write and DDL statements early against the live gate and
  reports a WAL refusal as Error::NamespaceFence instead of SQLITE_AUTH.
  A plain SQLITE_AUTH from an authorizer is left unchanged.
- New detail stale_transaction on MIGRATION_WRITE_FENCED for a
  transaction or program that began under an earlier generation.
- Namespaces without a fence stay at generation 0 and behave as before.

Tests cover a program parked between admission and its write while the
fence is acquired and released, read-to-write upgrades, DDL, a header
pragma, BEGIN IMMEDIATE, VACUUM and raw writes, stale transactions after
release, and unchanged behaviour of unfenced namespaces.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Queue entries and the write slot now carry the operation class. A write
transaction waiting for the slot re-checks the connection's fence
admission every time it takes the manager's lock, and the fence
controller wakes every registered write queue after each change of the
write generation, so a writer queued before a fence leaves the queue
with MIGRATION_WRITE_FENCED instead of waiting for the slot and then
writing. Checkpoints ask for the slot as maintenance: they are never
refused and queue again when woken.

The manager exposes what the positive write drain needs: the active
writer and its class, a notification on every release, and
abort_active(), which uses the registered rollback handle. Abort no
longer panics when the connection has already closed.

VACUUM is skipped, and reported as skipped rather than failed, while
the fence denies normal writes, including when the fence closes between
the check and the statement. TRUNCATE checkpoints run in every state.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
AcquireSourceWriteFence now runs the drain of docs/NAMESPACE_FENCE.md
section 8.3 end to end, under the namespace's transition lock and on a
task of its own:

- an in-memory INSTALLING gate closes write admission (and moves the
  write generation, waking queued writers) before SOURCE_DRAINING is
  persisted; a command proven not to have committed removes it again;
- the drain waits on the connection manager's release notification for
  the writer that held the slot when admission closed, never on elapsed
  time or the transaction timeout; at the deadline it answers DRAINING
  (admission stays closed) or, with force_rollback, rolls the writer
  back and waits for the actual release;
- the frozen boundary (log id, last committed frame) is read under the
  write-slot lock once no writer holds it, and SOURCE_WRITE_FENCED is
  committed with it;
- replaying a DRAINING command resumes the same drain.

Primary connection makers register a write-drain source (their
connection manager, held weakly, and their replication log) with the
namespace's controller; NamespaceStore::execute_fence_command loads the
namespace before an acquisition so that the source exists. The default
drain deadline is --namespace-fence-default-write-drain-ms (30 s).
FrozenBoundary.frame_no becomes optional, for a log without frames.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…ach boundary

Add crash-restart tests for the namespace fence: each server lifetime runs on
its own runtime and is ended without any shutdown code while a fence command
is parked at a hook point, so the next start takes the real dirty-recovery
path on the same directory. They cover every persistence boundary of
AcquireSourceWriteFence and ReleaseSourceWriteFence (including a marker that
lags the metastore commit), a restart in SOURCE_DRAINING with a writer active
at the crash, indeterminate commits through the drain path (applied and not
applied), and lost acquisition responses resolved by replay and inspection.

The tests exposed that a source restarted while draining could never finish
its drain: dirty recovery rebuilds the replication log under a new log id,
and completing the drain refused a boundary on a log other than the one the
fence was acquired against, leaving the namespace in SOURCE_DRAINING for
good. The frozen boundary now names the log that is live when the drain is
proven, the record's identity keeps the acquisition log id, and the server
warns when the two differ. Write admission was durably closed throughout, so
the data at the boundary is unchanged.

The BeforeMetastoreCommit test hook can now report a commit as indeterminate
without running it. The contract document describes the restart and log
rebuild semantics.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
SetSourceReadFence now closes read admission in memory, persists
SOURCE_READ_DRAINING, and waits for every read lease already held before
it persists SOURCE_READ_FENCED. Leases are taken only after checking the
gate under the controller's lease lock, so once admission is closed the
set of leases can only shrink.

Each SQL program holds a lease for as long as it runs (including a Hrana
cursor still producing rows), as does describe and each admin shell
query. ATTACH of a namespace is a read of that namespace: the attaching
connection keeps a lease on it for every later program until it is
detached. A connection idle inside a transaction holds no lease; its next
program is refused with MIGRATION_READ_FENCED and rolled back.
/beta/listen is refused where reads are denied and ends when reads are
fenced.

At the deadline (--namespace-fence-default-read-drain-ms, 30s) running
programs are cancelled through the connection's progress-handler cancel
flag and report the read fence; the drain still waits for the actual
releases and answers DRAINING if they do not come, and a replay of the
same command resumes it. ClearSourceReadFence reopens reads with writes
still fenced. Dump and replication stream leases follow.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Dump and replication now hold read leases on the namespace fence, so the
source read fence drains them positively:

- /dump is admitted by the fence gate before a connection is created (a
  failure to create one is an error, not a panic) and the export holds a
  dump lease. At the read drain's deadline the export is cancelled before
  its next row, and a write blocked on a peer that stopped reading fails at
  once, so the lease is released without the peer. A cancelled dump ends
  its body with the fence error and never reaches its final COMMIT;.
- Replication hello, log_entries, batch_log_entries and snapshot are
  refused at their start with FAILED_PRECONDITION and x-libsql-fence-code
  while streams are denied; refusals are counted and logged at most once a
  minute per namespace. Streams (and the frames of a batch) are served
  through FencedStream, whose watcher ends the stream as soon as the gate
  closes or the drain cancels it, dropping the inner stream and the lease
  itself; the next poll yields the typed terminal status.
- With --enable-namespace-fence, the RPC server and the user HTTP server
  send HTTP/2 keepalive pings (--namespace-fence-keepalive-interval-s,
  default 30 s, 20 s timeout) so dead peers are detected.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Add `NamespaceStore::create_target_quarantined` (and route the admin
`CreateTargetQuarantined` command through it), which creates a namespace
as a quarantined migration target atomically with namespace creation:

- A name the server already knows (config in memory, or a namespace
  cache entry) is refused with FENCE_PRECONDITION_FAILED/namespace_exists
  without touching its gate.
- Otherwise the controller publishes an in-memory target-creation gate
  before the metastore transaction writes the marker, config row, record
  and receipt. The gate refuses every class but maintenance and
  observability, and `check_available` refuses the name, so create, fork
  and `with()` neither store nor set anything up for it.
- The committed record's quarantine gate replaces it; only then is the
  config published into the in-memory map (from the durable row, with
  the record's own block values) and the namespace loaded, so its first
  connection maker is created behind the quarantine gate.
- The command runs on its own task under the transition lock; a replay
  returns the stored result and completes the publication and load of a
  commit that was not acknowledged, or of a creation interrupted between
  its marker and its commit.

`CreateTargetRequest` is the typed entry point for the admin route and
bulk import. The target's log id is not written back into the record, to
keep the marker rule for same-revision records intact (documented).

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Add the operation-owned import path into a quarantined migration target and
the seal that ends it (docs/NAMESPACE_FENCE.md sections 7, 10.2 and 11).

- MigrationCapability (server-issued, fields private) and CapabilityPurpose.
  The fence controller keeps the live capability set and a count of running
  import calls; every published transition drops capabilities whose state,
  owner or revision no longer match.
- FenceConnState::with_capability: the WAL admits a capability connection's
  write transaction only while its capability matches the fence and is live;
  a validation connection never writes.
- NamespaceStore::open_import_session and ImportSession::{with_raw,
  load_dump}: the only way to write into TARGET_QUARANTINED. The dump loader
  is split into load_dump_sql so that the loader used for namespaces created
  from a dump also runs under an import capability.
- SealTargetImport closes import admission in memory, persists
  TARGET_IMPORT_DRAINING (invalidating every import capability), waits on
  release notifications for running import calls and for any import
  transaction holding the write slot (force_rollback rolls it back at the
  deadline), then persists TARGET_VALIDATING. A deadline leaves
  TARGET_IMPORT_DRAINING durable and closed until the owner's seal resumes it.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…mpaction timing

`snapshot_stream_ends_typed` polled `get_snapshot_file(1)` and unwrapped
its result while the compactor was still writing the first snapshot and
the snapshot merger could be replacing merged files. The lookup lists
the snapshot directory and then opens the chosen file, so under load it
could fail with `NotFound` (directory not created yet, or a file removed
by a merge between the listing and the open) and the test panicked.

Open the stream through the `snapshot` RPC itself in a bounded loop that
treats only "snapshot not found" and the vanished-file error as "not
yet", and asserts that no read lease is held after a failed attempt. An
opened stream holds its file open, so a later merge cannot affect it.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Serve the fence contract of docs/NAMESPACE_FENCE.md section 4 on the
admin listener (new http/admin/fence.rs):

- GET /v1/fence/capabilities, always served: protocol version, whether
  fences are enabled, served commands, states, proxy stable_code support,
  server build and instance id, and the number of active fences counted
  from the fence registry.
- GET /v1/namespaces/:ns/fence (InspectFence): the durable record, the
  owning operation's receipts (or ?receipts=all), live admission, drain
  counters and the live replication log id; read-only, and also served
  while fence tables exist with the flag off.
- One POST route per source and target command, all through
  NamespaceStore::execute_fence_command, and validation-query, which runs
  one read-only program under the operation's validation capability with
  a 10 000-row bound.

Command routes answer 404 unless --enable-namespace-fence is on, refuse
to run without an admin auth key (admin_auth_required) or on a replica
(not_primary), parse bodies strictly (invalid_argument), and refuse
restore options on target creation (restore_not_allowed). Responses carry
the outcome, replay flag, fence view, receipt and drain counters with the
outcome's admin status.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Refuse config changes, delete, reset, fork (as source or destination),
create over an existing record (with or without a dump URL), linking to a
shared schema and schema migration while a namespace's fence denies
lifecycle work.

A single check reads the fence registry, so it also sees in-memory gates
(a closing transition, a target being created, an indeterminate commit)
and never loads the namespace. Paths that persist through the metastore
are refused again inside its transaction.

A fork reads the source's log without a read lease, so it now holds the
source's transition lock for its whole run and checks the gate under it:
a write or read fence command either completes first and the fork is
refused, or waits for the fork. Reset writes nothing to the metastore and
is checked under the same lock. The fork destination is checked before
anything is stored and again with its namespace entry held; previously a
fork onto an existing, unloaded namespace would publish its config in
memory and remove its directory. A dump URL is no longer fetched for a
create that is refused.

Registering a schema migration job checks the schema and every linked
namespace, so a link written by a binary that does not know fences cannot
lead to a migration step being refused halfway through a job.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…tocols

Two additive proto3 fields for the namespace fence:

- `proxy.Error.stable_code` (tag 4): a stable machine-readable outcome
  such as `MIGRATION_WRITE_FENCED`, so a replica can return the same
  typed outcome the primary would have. Absent means "no typed
  outcome"; older peers skip it.
- `metadata.DatabaseConfig.fence` (tag 14, `ReplicatedFence { state,
  revision }`): the live fence as the primary's replication `hello`
  sees it. It is filled only by `hello` and only while a fence is
  active, and is never part of a stored configuration.

The primary now fills `DatabaseConfig.fence` in `hello` from the
namespace's fence gate. Filling `stable_code` on fence denials and
mapping it on the replica side come in a later change; until then no
server sets it, and capability discovery keeps reporting
`proxy_stable_code: false`.

Tests cover wire compatibility in both directions (an absent field
encodes exactly as before) and that `hello` carries the fence only while
one is active.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
A fence denial reaching the user-facing protocols is now a typed
answer everywhere instead of a generic or fatal error:

- HTTP error bodies for fence errors gain an additive `code` field
  (and `detail` when there is one), with `423` for data-plane
  denials, through every wrapper the error arrives in. The legacy
  `/` API answers a batch with a fenced step as a whole with `423`.
- Hrana gains `StmtError::Fence` and `BatchError::Fence`, whose code
  is the stable fence code, so step denials and whole-request
  denials (a read under a read fence, a quarantined target) are
  Hrana errors on `/v1`, `/v2`, `/v3`, cursors and WebSockets, and
  the stream stays usable.
- `/v1/execute` and `/v1/batch` answer whole-request denials with
  `423` and the code.

Integration tests cover each protocol, an old WebSocket transaction,
a batch denied mid-way, `/dump`, and that `401`/`404` stay distinct;
a lib test shows a dump cancelled by the read drain fails its
response body.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
… proxy

The primary's proxy service now fills the additive `stable_code` field
(with `code = SQL_ERROR`) for fence denials, on step errors and program
errors, streamed and unary. A fence denial before a program runs (the
namespace or JWT-key lookup, connection creation, a unary program
refused as a whole) is the typed `FAILED_PRECONDITION` status with the
stable code in `x-libsql-fence-code`, never `UNAVAILABLE`, which the
write proxy retries without bound.

On a replica, the write proxy maps a proxied error or status carrying a
fence code back to `Error::NamespaceFence`, so the replica answers its
client exactly as the primary would: `423` with the `code` field on the
HTTP APIs and the stable code as the Hrana error code. Errors from an
older primary, which never sets the field, keep their old mapping.
Capability discovery now reports `proxy_stable_code: true`.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
A replica server whose primary refuses replication of a namespace with a
fence code (a refused `hello`, `log_entries` or `snapshot`, or a stream
the primary ended with the typed status) now refuses local reads and
streams of its copy with the same code, and asks the reads it had already
admitted to stop. The denial is published on the namespace's fence
controller on the replica, so every user protocol reports it as it does
on the primary; writes keep going to the primary, which refuses them
itself. A `hello` the primary answers lifts it.

The refusal is no longer retried every second by the handshake loop: the
replica's replication loop waits 1 s, doubling to 15 s, until the primary
answers again, and counts each refused call in
`libsql_server_replica_fence_refusals_total{code}`.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…refuses

A replica server creating a namespace lazily for a name the primary's fence
refuses (a quarantined or aborted target, a read-fenced source, a fence state
the primary cannot establish) now fails the request at once with the primary's
code (Error::NamespaceFence: 423 and the stable code) instead of retrying the
handshake, and leaves no local namespace behind: the namespace directory the
setup created is removed (a directory that already existed is kept), and
NamespaceStore::with forgets the metastore entry handle() added and the
controller the attempt created, when neither holds anything durable or is in
use by another attempt. A later request, once the primary admits the name,
creates it normally.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Keep what the metastore's bottomless restore reports at startup (whether
it recovered the database and the generation it restored from) instead of
discarding it, record it on the MetaStore, and report it in the fence
capability endpoint (`metastore`), in every fence view (`provenance`), in
the `libsql_server_metastore_restored_from_backup` gauge and in a startup
warning. When destroy_on_error rebuilds the metastore, the restore of the
rebuilt metastore is what is reported.

Tests cover a real bottomless restore against a local S3 endpoint, the
admin rendering, and that `incarnation.current_log_id` names the rebuilt
replication log after a dirty restart while the stored record keeps the
log the source was acquired on.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Serve AdoptFence on the admin API (POST /v1/namespaces/:ns/fence/adopt).
It needs the admin credential and the separate adoption key configured
with --namespace-fence-adoption-key (SQLD_NAMESPACE_FENCE_ADOPTION_KEY),
presented in the x-libsql-fence-adoption-key header. The key is kept as
a SHA-256 digest and compared without an early exit; without a key,
adoption is refused with adoption_not_authorised. The request names the
current owner, two distinct approvers, an incident reference and a
reason, which are stored in the record and receipt and written to one
audit event (target libsql_server::fence::audit).

Adoption moves the owner and the revision and nothing else: every
admission stays as it was, capabilities of the old owner are revoked,
and finished operations cannot be adopted. After a metastore rollback
it re-establishes the record the marker holds under the new owner,
settles the unavailable name and restores the namespace's own block_*
values in memory. When the metastore holds no config row for the name,
adoption is refused with namespace_config_missing instead of inventing a
configuration.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Walk a source and two targets through every stored fence state and check
the protection an older binary gets (docs/NAMESPACE_FENCE.md section
13.2): the config row's block_* fields hold the fence's mirror and
nothing else in the row changes, a config write through the metastore is
refused and changes nothing, and an older binary's delete of the config
row fails on the fence row's foreign key. Release and write enable put
the namespace's own values back.

The test found that a restart overwrote the in-memory config of a
released source or a writable target with the block_* values saved when
the fence was acquired, although config writes after the operation had
stored the namespace's own values in the row since: a namespace blocked
after its release could come back unblocked. Loading the metastore, the
target config publication and adoption now take the namespace's own
config through fence_store::own_config, which uses the row as it is once
the record no longer mirrors the fence.

Crash the server at each persistence boundary of SetSourceReadFence,
SealTargetImport and EnableTargetWrites, and while a read or seal drain
waits for a reader or an import call, then restart it on the same
directory: the prior or the committed state is recovered with no
in-memory gate, reader or import call left, reads and import stay
closed once their draining state committed, target writes open only if
TARGET_WRITABLE committed, the marker is repaired, and a replay of the
same command completes an interrupted drain at once.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Every fence command the server answers (committed, replayed or refused)
now emits one structured event under the libsql_server::fence::audit
tracing target, with the namespace, operation and command id, outcome,
revisions, state before and after, drain kind and duration, forced
actions, replay/conflict and server instance; a committed adoption keeps
its approvers, incident reference and reason.

Metrics, all with bounded labels (docs/NAMESPACE_FENCE.md section 15):
transitions by command and outcome, drain duration by kind, forced
rollbacks and cancellations by kind, replays and conflicts, denials by
code and surface (http, hrana, rpc, proxy, dump, replication,
admin_shell, lifecycle), adoptions, and two gauges computed from the
fence registry when /metrics is read: namespaces by role and state, and
the age of the oldest active fence.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Add a test that the admin shell can neither read nor write a
quarantined migration target, through its own entry point and with a
raw write that skips its read admission, while the operation's import
session keeps working.

Map every acceptance requirement in docs/NAMESPACE_FENCE.md section 17
to the tests that cover it, correct the names of the corrupt-state
tests, list the transition and CAS tests that cover ownership, revision
and replay outcomes, and note the replica connection path in the
code-path coverage table.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
shopify-river and others added 3 commits September 30, 2026 12:47
…dule

The fence module carried a module-wide `allow(dead_code)` while its
consumers landed across the series. Everything is wired now, so remove
it and deal with the five items it was hiding:

- delete the `InBeginWriteTxnAfterCheck` and `AfterManagerRelease` hook
  points, which no code path reaches and no test arms;
- delete the unused `AuditReport::drain` accessor and the
  `FenceController::cancel_read_leases` wrapper (its one test now uses
  `cancel_read_leases_by_kind`);
- compile `DenialSurface::ALL`, `AuditReport::forced_kinds` and the
  `Fail`/`Indeterminate` hook outcomes only in the library's test build,
  which is the only place they are used or produced.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Keep stale DRAINING receipts historical once their exact drain state and revision have been superseded, and preserve resumed attribution when a live drain completes. This prevents old commands from cancelling or completing work in newer states while keeping replay responses and audit classification accurate.

Count fence errors delivered as terminal replication-stream statuses in the consecutive-refusal backoff, so reconnect pacing doubles from the first refusal as documented.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
@tszymczyszyn-shopify

Copy link
Copy Markdown
Author

Superseded by a stack of nine smaller draft PRs, split for review:

Together they carry exactly this PR's diff (same git patch-id --stable: 4bc63326; 75 files, +28,748/-293), rebased onto v0.9.30-shopify-patches. Fix-up commits are folded into the pieces they fix, and each piece builds and passes CI on its own. Please review and land bottom-up, starting with #41. Closing this one in favour of the stack.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants