Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "switchboard"
version = "0.2.0"
version = "0.3.0"
edition = "2024"
rust-version = "1.85"
license = "MIT"
Expand Down
36 changes: 20 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -472,27 +472,31 @@ database and `<key>.toml` override on disk, every health check green. This is th
class of fault behind override counts drifting between nodes and orphaned
overrides lingering as 404s.

The reconcile pass is the **level-triggered** backstop. On its own interval,
independent of `gen`, it reads the KV desired state *directly* and converges this
node's on-disk previews to it. Pruning classifies each preview directory from its
**own** `switchboard:preview:<label>` key (a plain `set`, immune to both the
`gen` propagation race and the index's lost-update race), so it never trusts the
index and never false-positives a live preview into an orphan.
The reconcile pass is the **level-triggered** backstop, and its authority is
**GitHub PR state** — needing no KV and no ePHPm-side change (the preview nodes
keep ePHPm's RESP listener off). On its own interval, independent of `gen`, it
lists every preview directory, parses each site key `<owner>-<repo>-pr-<N>` back
to its pull request, and asks GitHub whether that PR is still open. **Open (or an
unrecognised state) ⇒ keep; merged or closed ⇒ prune.** GitHub is more
authoritative than the KV index (which has a lost-update window), and a preview
that has fallen off every other signal is still correctly classified by the one
fact that actually decides whether it should exist. Re-*deploying* a missing but
still-open preview is deliberately left to the webhook path; this pass only
removes what GitHub says is gone.

| Flag | Env | Default | Meaning |
|---|---|---|---|
| `--reconcile-interval-secs` | `SWITCHBOARD_RECONCILE_INTERVAL_SECS` | `0` | Seconds between reconcile passes. **`0` disables it.** Requires `--kv-secret-file` and a resolvable API site key; a backstop, not the hot path, so a cadence well above the drain tick (e.g. 30–60s) is right. |
| `--reconcile-prune` | `SWITCHBOARD_RECONCILE_PRUNE` | `false` | Actually remove orphaned previews. **Off by default**: until set, the pass is observability-only and logs each orphan it *would* prune at `WARN`. |
| `--reconcile-deploy-missing` | `SWITCHBOARD_RECONCILE_DEPLOY_MISSING` | `false` | Enqueue a deploy for a desired preview whose directory is absent on this node (recovers a deploy a node never saw). Off by default — re-provisioning runs composer and a checkout. |
| `--reconcile-keep-sites` | `SWITCHBOARD_RECONCILE_KEEP_SITES` | *(empty)* | Comma-separated vhost directory names never pruned — the node's infra/test sites (e.g. `switchboard,site-a,site-b`). The API's own site key is always protected in addition. |
| `--reconcile-max-prunes-per-cycle` | `SWITCHBOARD_RECONCILE_MAX_PRUNES_PER_CYCLE` | `8` | Blast-radius guard: at most this many orphans are removed per pass, the excess deferred to later passes so a misconfiguration cannot wipe the fleet in one tick. |
| `--reconcile-api-site` | `SWITCHBOARD_RECONCILE_API_SITE` | *(derived from `--drain-host`)* | The switchboard-api vhost's canonical site key, whose keyspace holds the `switchboard:*` desired-state keys. Set only if it differs from the drain host's site key. |
| `--reconcile-interval-secs` | `SWITCHBOARD_RECONCILE_INTERVAL_SECS` | `0` | Seconds between reconcile passes. **`0` disables it.** Requires `--app-id`/`--app-key` (the pass queries GitHub); a backstop, not the hot path, so a cadence well above the drain tick (e.g. 30–60s) is right. |
| `--reconcile-prune` | `SWITCHBOARD_RECONCILE_PRUNE` | `false` | Actually remove orphaned previews (PRs merged/closed). **Off by default**: until set, the pass is observability-only and logs each orphan it *would* prune at `WARN`. |
| `--reconcile-keep-sites` | `SWITCHBOARD_RECONCILE_KEEP_SITES` | *(empty)* | Comma-separated vhost directory names never pruned — the node's infra/test sites (e.g. `site-a,site-b,preview.ephpm.dev`). The API's own site key is always protected in addition. (Belt-and-braces: a non-`<owner>-<repo>-pr-<N>` directory never parses as a preview, so it is never a prune candidate regardless.) |
| `--reconcile-owner` | `SWITCHBOARD_RECONCILE_OWNER` | `ephpm` | The GitHub owner/org previews belong to — the leading segment of `<owner>-<repo>-pr-<N>`, and the owner PR state is queried under. |

Rollout is fail-safe: enable the interval first and read the dry-run `WOULD
prune` logs; turn on `--reconcile-prune` once they look right; add
`--reconcile-deploy-missing` last. A KV read failure aborts the whole pass — the
reconcile prunes **nothing** rather than act against an authority it could not
read.
prune` logs; turn on `--reconcile-prune` once they look right. Uncertainty always
resolves to **keep** — a site key that does not parse as `<owner>-<repo>-pr-<N>`
(an infra vhost, or a hashed/overflow label) is skipped, a per-PR GitHub read
error keeps that preview, and a failure to mint the installation token aborts the
whole pass so a total GitHub outage prunes **nothing**.

### GitHub reporting (optional)

Expand Down
123 changes: 43 additions & 80 deletions src/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -279,15 +279,15 @@ pub struct Config {
/// that is lost or not-yet-gossiped strands a teardown (the preview stays
/// served, its database and override on disk) with every health check green
/// (switchboard#24). This pass is the level-triggered safety net: every
/// interval it reads the KV desired state *directly* (independent of `gen`)
/// and converges this node's on-disk previews to it — pruning orphans and,
/// with `--reconcile-deploy-missing`, enqueuing deploys it never saw.
/// interval it asks **GitHub** whether each on-disk preview's PR is still
/// open, and prunes the ones whose PR has merged or closed. It needs no KV
/// (the preview nodes keep ePHPm's RESP listener off) and no ePHPm-side
/// change — only the switchboard App credentials it already reports with.
///
/// Requires `--kv-secret-file` (the pass reads desired state over RESP) and a
/// resolvable API site key (`--reconcile-api-site`, else derived from
/// `--drain-host`); the daemon refuses to start with a non-zero interval and
/// neither. A sane cadence is well above the two-second drain tick — it is a
/// backstop, not the hot path — e.g. 30–60s.
/// Requires `--app-id`/`--app-key`; the daemon refuses to start with a
/// non-zero interval and no App credentials (the pass would query nothing).
/// A sane cadence is well above the two-second drain tick — it is a backstop,
/// not the hot path — e.g. 30–60s.
#[arg(long, default_value_t = 0, env = "SWITCHBOARD_RECONCILE_INTERVAL_SECS")]
pub reconcile_interval_secs: u64,

Expand All @@ -299,47 +299,21 @@ pub struct Config {
#[arg(long, default_value_t = false, env = "SWITCHBOARD_RECONCILE_PRUNE")]
pub reconcile_prune: bool,

/// Enqueue a deploy for a desired preview whose directory is absent on this
/// node. **Off by default**: re-provisioning runs composer and a checkout,
/// so recovering a missed *deploy* is a heavier, separate opt-in from
/// removing a stranded *teardown*. When on, the reconcile writes the exact
/// job document into the queue and the normal deploy path runs it.
#[arg(
long,
default_value_t = false,
env = "SWITCHBOARD_RECONCILE_DEPLOY_MISSING"
)]
pub reconcile_deploy_missing: bool,

/// Comma-separated vhost directory names the reconcile must never prune —
/// the node's infra/test sites that are not PR previews (e.g.
/// `switchboard,site-a,site-b`). The API's own site key is always protected
/// in addition to this list. A preview directory is only ever a prune
/// candidate when it is absent from this set *and* its own KV preview key is
/// gone or torn down.
/// `site-a,site-b,preview.ephpm.dev`). The API's own site key is always
/// protected in addition to this list. Belt-and-braces: a non-`<owner>-<repo>-pr-<N>`
/// directory never parses as a preview and so is never a prune candidate
/// regardless.
#[arg(long, default_value = "", env = "SWITCHBOARD_RECONCILE_KEEP_SITES")]
pub reconcile_keep_sites: String,

/// Upper bound on orphans pruned in a single pass — a blast-radius guard.
/// A misconfiguration (wrong keep-list, a flushed KV) cannot then wipe the
/// fleet in one tick; the excess is deferred to later passes, giving the
/// WARN/INFO logs time to be noticed. Floored at one.
#[arg(
long,
default_value_t = 8,
env = "SWITCHBOARD_RECONCILE_MAX_PRUNES_PER_CYCLE"
)]
pub reconcile_max_prunes_per_cycle: usize,

/// The switchboard-api vhost's canonical **site key**, whose gossip-replicated
/// keyspace holds the `switchboard:*` desired-state keys the reconcile reads.
///
/// Defaults to the site key derived from `--drain-host` (the same host the
/// drain kick addresses), which is correct whenever the API is reached at its
/// own vhost. Set it explicitly only if the API's KV site key differs from
/// its drain host.
#[arg(long, env = "SWITCHBOARD_RECONCILE_API_SITE")]
pub reconcile_api_site: Option<String>,
/// The GitHub owner/org the previews belong to — the leading segment of a
/// preview site key `<owner>-<repo>-pr-<N>`, and the owner the reconcile
/// queries PR state under. Defaults to `ephpm` (the fleet's org); set it if
/// previews are built for a different org.
#[arg(long, default_value = "ephpm", env = "SWITCHBOARD_RECONCILE_OWNER")]
pub reconcile_owner: String,

// ── pre-serve static-analysis gate ─────────────────────────────────
/// Path to the operator-controlled `ephpm analyze` policy file (YAML), used
Expand Down Expand Up @@ -521,43 +495,34 @@ impl Config {
self.reconcile_interval_secs > 0
}

/// The switchboard-api vhost's canonical site key, whose keyspace the
/// reconcile reads: the explicit `--reconcile-api-site` if set, otherwise
/// derived from `--drain-host` with this node's suffix rule.
///
/// `Ok(None)` when neither an override nor a drain host is available.
/// The switchboard-api vhost's canonical site key, derived from
/// `--drain-host` with this node's suffix rule — added to the reconcile
/// keep-list so the API's own vhost is never a prune candidate.
///
/// # Errors
///
/// Returns an error if a drain host is present but does not normalize to a
/// valid site key.
pub fn reconcile_api_site(&self) -> anyhow::Result<Option<String>> {
if let Some(explicit) = &self.reconcile_api_site {
return Ok(Some(explicit.clone()));
}
let Some(host) = &self.drain_host else {
return Ok(None);
};
/// `None` when no drain host is configured (single-node mode) or it does not
/// normalize to a valid site key; the keep-list simply omits it then (and the
/// `<owner>-<repo>-pr-<N>` parse rule already excludes an infra vhost anyway).
#[must_use]
pub fn drain_host_site_key(&self) -> Option<String> {
let host = self.drain_host.as_ref()?;
let suffix = self.effective_sites_domain_suffix();
crate::site_key::site_key(host, suffix.as_deref())
.map(Some)
.with_context(|| {
format!("cannot derive the switchboard-api site key from --drain-host {host:?}")
})
crate::site_key::site_key(host, suffix.as_deref()).ok()
}

/// The reconcile keep-list — vhost names never pruned — from
/// `--reconcile-keep-sites`, with the API site key always added.
/// `--reconcile-keep-sites`, with the API site key added when known.
#[must_use]
pub fn reconcile_keep_set(&self, api_site: &str) -> std::collections::HashSet<String> {
pub fn reconcile_keep_set(&self, api_site: Option<&str>) -> std::collections::HashSet<String> {
let mut set: std::collections::HashSet<String> = self
.reconcile_keep_sites
.split(',')
.map(str::trim)
.filter(|s| !s.is_empty())
.map(str::to_string)
.collect();
set.insert(api_site.to_string());
if let Some(api_site) = api_site {
set.insert(api_site.to_string());
}
set
}

Expand Down Expand Up @@ -599,22 +564,20 @@ impl Config {
"--fork-secrets has no effect without --allow-fork-deploy — set \
both to build forks with operator secrets, or neither"
);
// The reconcile reads desired state over RESP and addresses the API's
// KV keyspace by site key; without a secret or a resolvable site key it
// is silent dead configuration, so refuse to start rather than run a
// pass that can do nothing.
// The reconcile decides prune-vs-keep by asking GitHub for each preview's
// PR state, so it needs the App credentials. Without them every pass would
// abort (mint nothing, prune nothing) — silent dead configuration, so
// refuse to start rather than run a pass that can do nothing.
if self.reconcile_enabled() {
anyhow::ensure!(
self.kv_secret_file.is_some(),
"--reconcile-interval-secs is non-zero but --kv-secret-file is unset — \
the reconcile reads cluster desired state over RESP and cannot without \
ePHPm's [kv] secret (set it to 0 to disable the reconcile)"
self.github_reporting_enabled(),
"--reconcile-interval-secs is non-zero but --app-id/--app-key are unset — \
the reconcile queries GitHub for each preview's PR state and cannot without \
the switchboard App credentials (set the interval to 0 to disable it)"
);
anyhow::ensure!(
self.reconcile_api_site()?.is_some(),
"--reconcile-interval-secs is non-zero but the switchboard-api site key \
could not be resolved — set --reconcile-api-site, or --drain-host to \
derive it from"
!self.reconcile_owner.trim().is_empty(),
"--reconcile-owner must not be empty when the reconcile is enabled"
);
}
// A daemon that cannot remove a tenant's database when its PR closes is
Expand Down
Loading
Loading