The control-plane console for vmsync replication. Agents on each hypervisor enrol here and report what their host knows about its own replication state; operators read the availability page.
Its own module, with no dependencies outside the Go standard library. It
never touches libvirt, libnbd or any of vmsync's Go packages — it only speaks
JSON over HTTPS to agents — so it builds with a plain go build and none of
the native headers vmsync itself needs.
This is phase 4 of the control plane: agents report, the console sets the replication schedule they run, and failovers are driven from it.
What it still cannot do is act on a VM directly. It has no session to any hypervisor and issues no command; it publishes a schedule, and each agent decides what to run on its own host. So the console being unreachable does not stop replication — agents keep running the last schedule they cached.
A failover is issued the same way: as a one-shot operation the agent collects on its next poll and carries out itself. The console never opens a session to a hypervisor; it publishes an instruction, and the agent that owns the host decides whether to honour it.
The same split is why a warning here is only ever a warning. When the failover page says a replica is known not to match its source, it is repeating a verdict vmsync wrote and vmsync enforces — the engine refuses that promotion whatever this page shows. So the control stays offered rather than being withdrawn: taking it away would hide the explanation without removing the ability, since the same command remains available on the hypervisor itself.
The one thing to understand about the schedule is that it names when, not
whether. Removing an entry stops the timer; it does not stop anything else
invoking vmsync for that VM. To make a VM genuinely refuse to be
overwritten, set its replication role (vmsync -update-role=paused), which
vmsync enforces itself and which does not depend on this console at all.
It holds no hypervisor credentials and reaches into no host. Agents dial out to it, so hypervisors need no inbound port — which is what makes a DR site behind its own firewall workable.
It stores no inventory of replication pairs. The topology lives in each VM's own libvirt metadata and is rediscovered by agents on every report, so this console cannot drift out of step with reality about what replicates where.
go build -o vmsync-ui ./cmd/vmsync-ui
install -m 0755 vmsync-ui /usr/local/bin/
mkdir -p /etc/vmsync-ui
cp vmsync-ui.conf.example /etc/vmsync-ui/vmsync-ui.confGenerate a password hash per account — plaintext passwords are refused at startup, with a message naming this command:
vmsync-ui -hash-passwordThen edit /etc/vmsync-ui/vmsync-ui.conf:
| key | meaning |
|---|---|
listen |
Address to serve on. Both the agent API and the console share it. |
tls_cert / tls_key |
Required. Agents refuse a non-https:// UI. |
state_dir |
Agents, their token hashes, reports and the audit log. |
grafana_url |
Optional. Adds a "Trends" link; this console deliberately draws no time series. |
session_hours |
How long a sign-in lasts. Default 12. |
accounts |
username, password_hash, role (admin or readonly). |
insecure |
Serves plain HTTP for local development. Agents will refuse to talk to it, so this is browser-only. |
An unrecognised key is a startup error rather than a setting that silently does nothing — in a file carrying TLS paths and account roles, a quietly ignored typo is the worst outcome.
| capability | readonly | admin |
|---|---|---|
| View availability, agents, audit, failover state | yes | yes |
| Enrol and revoke agents | no | yes |
| Set the schedule | no | yes |
| Promote, invert, shut down, set a role, cancel | no | yes |
A reader sees a split brain and every failover the estate has performed; what they do not get is anything to click. Withholding the state would be the wrong half to withhold — the people watching the board are often exactly the ones who need to raise the alarm.
The split that matters most: -verify=compare and
-verify=fast suspend the source VM, and only -verify=online does not.
Read-only accounts get online alone — "just a verification" must not be a
production-impacting action for someone who was given the lesser role.
Attribution only means something if accounts are not shared. "Who failed over production at 3am" is a question that gets asked eventually.
On the Agents page, create a token for a hostname. It is shown once, is single-use, is bound to that hostname, and expires in 24 hours. Then on the hypervisor, follow agent enrolling procedure.
That exchanges the token for a long-lived credential and sends one report. Start the service afterwards.
If the UI has a publicly-trusted certificate, agents need nothing — no CA
file, and renewals are invisible to them. The agent's --ui-ca flag is only
for a private CA (distribute the CA once; it outlives many leaf renewals) or
a bare self-signed certificate (which does mean touching every host on every
renewal — prefer a small private CA).
Ordered worst-first, because the page is read to find what needs attention.
- Not replicated gets its own alarm-styled panel. A VM nobody configured replication for is otherwise absent from a green board and looks fine.
- Hosts with no agent are listed as blind spots. "I cannot see this" must never render as "this is healthy".
- Target missing gets its own alarm-styled panel. A source whose named replica is in no agent's report under that exact name — deleted, never created, misspelled, or a short name where the agent reports an FQDN — has nowhere for its syncs to land. Each row is shown verbatim, with a note saying whether any agent reports under that name at all. A reference is skipped when the source already has a pair row for a target under the same VM name: the pair proves the copy exists, so what remains is a spelling difference, not a missing copy.
- References match exactly (case-insensitive only):
hypervisor01in hand-typed vmsync metadata does not resolve to an agent reporting ashypervisor01.domain.tld. That surfaces as rows to fix — align the names and the rows clear — rather than being silently correlated away. - Behind is the replication lag — how much you would lose. A never-synced
target shows
—rather than0, which would read as "just synced". - promoted and paused get a distinct neutral colour. They are deliberate states; if a planned failover turned the board red, people would learn to ignore red.
- Verify failed is a second pill beside the status, never instead of it,
with the date the finding was recorded. The status answers how far behind
a copy is and goes on saying
okfor one that was compared against its source and did not match, so folding the two together would make the row choose between two facts that are both true — and dropping the marker would let a copy known to be wrong read as healthy. The target cell spells out what it means: vmsync refuses every sync into the replica while the finding stands, a plain full resync included, and the repair is a sync set to recopy once and re-verify, which clears it only if that second comparison passes. - Partial copy is a third pill, also beside the status, for a replica
whose rebuild was interrupted. It is the one marker on the board that
contradicts every other cell on its row, and the reason is worth
understanding: a full copy into an existing target renames the good disks to
<disk>.vmsync-replaced-<stamp>and writes new base images in their place without touching the domain's metadata, which is only updated when the run finishes. So a run that dies part way leaves a recent checkpoint, a small lag and a failure count of zero — all truthful, and all about the copy that was renamed aside rather than the disks now sitting there. Nothing else on the board moves. The dated explanation on the target cell says the replica is a partial copy, that the metadata beside it describes the copy the rebuild replaced, and that the complete copy may still be on the host under the.vmsync-replaced-<stamp>suffix; the source cell repeats the headline, because a resync runs from the source and the source's own metadata knows none of this. - Reasons appear on the row, not behind a click.
Freshness always comes from the target, because that is where vmsync
writes last_sync, last_checkpoint and failure_count. A source's own
metadata records where it replicates to, never when. The verification record
(verify_state, verify_failed_at) is read from the target and the same
report, but it is not a freshness fact at all: freshness says how far behind a
copy is, the record says it does not match, and a copy that is wrong does not
become right by being recent.
replica_incomplete comes from the target too, and it is the one field that
tells you the freshness figures beside it are about different disks. It is
carried raw and parsed only where it is rendered; nothing in this console
branches on it. The refusal it exists to justify lives in vmsync, on the host
holding the disks, because -promote runs there during a disaster with the
source host gone — so a value this console cannot parse still renders as a
warning, with the agent's own text shown verbatim, rather than disappearing.
POST /api/v1/agents/enrol -> {"agent_id","token"}
POST /api/v1/agents/{id}/report bearer; 204
GET /api/v1/agents/{id}/config bearer; long-poll, ETag-aware
internal/api/agents_test.go here and client_test.go in the agent are the
two halves of this contract, each driving a stub of the other. Editing one
without the other is how agents in the field stop working.
Notes for anyone changing it:
- 401 means revoked and the agent treats it as terminal, saying so loudly. Every enrolment failure returns a bare 401 with no detail — an unauthenticated caller learns only that it did not work.
- The config endpoint must genuinely hold the request for the
waitit is given before answering 304. That hold is what delivers a change to a hypervisor in seconds without any inbound connection. - Bearer tokens are stored as SHA-256 hashes. If
agents.jsonleaks, what leaks is hashes.
- No
WriteTimeouton the HTTP server, deliberately: agents long-poll for minutes, and a write deadline would sever those mid-hold.ReadHeaderTimeoutandIdleTimeoutcover slow-client abuse instead. - Sessions are in memory only. A restart signs everyone out and there is no session file to leak.
- Forms carry a per-session CSRF token; the session cookie is
HttpOnly,SecureandSameSite=Strict. - Passwords are PBKDF2-HMAC-SHA256 at 600k iterations, from the standard library as of Go 1.24.
The availability page raises one alarm above everything else: a pair where the target has been promoted and is running while its original source is running too. Two live copies of one VM, diverging from the moment the second started.
A promotion issued from the failover page can arm a fence against the displaced source, and that closes the common case: the old source's agent reads the token from the promoted domain's own libvirt and shuts its copy down cleanly. Detection still matters, because fencing is cooperative and therefore has three ways not to happen — the displaced host may have no agent, its agent may be unable to reach the promoted peer to read the token, or the guest may ignore the shutdown, which is latched rather than retried. None of those is rare during exactly the partition that motivated the failover, and a promotion performed without arming a fence at all (a drill, or from a shell) authorises nothing.
So this console does not claim to prevent split brain; it has no power control and never destroys a running guest. What it can always do is notice: neither agent can see the other, so the only place the condition is visible at all is here, where both reports arrive.
It is deliberately narrow — promoted and target running and source
running and the source's report current. That last condition is what keeps
the alarm worth reading. Reports are stored in place and never aged out, so a
host that died with its VM running keeps saying Active: true forever; without
a freshness gate the banner would fire on every correctly-executed failover of a
dead primary, and an alarm indistinguishable from success is one operators learn
to scroll past.
A promoted target whose source has simply gone quiet is reported separately and without alarm, showing how long ago the source was last heard from. That is the honest state: silence is not proof the old primary is down, and it is also not evidence that it is up. It becomes the real alarm only if that host starts reporting again with the VM still running.
Only source VMs are listed. A target is synced by the agent on its source host, so an entry against a target's own agent would be one that agent could never act on — the page does not offer it rather than accept a schedule that silently does nothing.
Each entry is an interval, a profile and an optional verification mode:
| Profile | For | Settings |
|---|---|---|
wan |
a link you pay for | zstd-5, network buffer, I/O depth 16 |
lan |
a switched site network | zstd-1, I/O depth 8 |
direct |
same host or a fast fabric | no compression, I/O depth 8 |
The profile is resolved to explicit vmsync flags here, and travels to
the agent as those fields — never as the name wan. The agent therefore
needs no opinion about what wan means, and what an operator reads on the
page is what will run.
-verify=compare and -verify=fast suspend the source VM for the
comparison. -verify=online does not. The page says so on the control, and
the server checks it again on save rather than trusting the form.
A saved change reaches a host on that agent's next poll, normally within seconds. Two estate-wide limits bound what that can cost: how many syncs one agent runs at once, editable under Estate defaults below, and how many may target the same host, which is stored and served but has no form yet.
Promotions, inversions, clean shutdowns and role changes are published as operations — one-shot instructions for one agent, alongside the standing schedule. They ride in the same document an agent long-polls, so issuing one is publishing it, and acknowledging one is removing it.
The lifecycle, and what each step protects against:
- An admin confirms. The UI writes an audit entry recording intent, then the operation, with a 15-minute deadline.
- It appears in that agent's config, moving the ETag, so a polling agent sees it within seconds.
- The agent executes it exactly once ever, against a durable ledger that records intent before acting.
- The result rides on every report the agent sends until this UI stops publishing the operation.
- The UI applies the consequences, records the result, and stops publishing — which is what tells the agent to forget it.
Only one operation per VM may be in flight. Two promotions of one domain, or a promote racing an invert, is not a state anyone should be able to create by clicking twice on a slow page.
Consequences are applied before the result is recorded. A crash between them leaves the operation still published and the agent still re-reporting it, so the whole step simply runs again — it is idempotent and keyed by operation ID. The other order would acknowledge the operation and lose the schedule changes forever.
Those consequences are the part whose absence silently breaks a pair:
- A successful promotion disables the old source's schedule entry.
Otherwise it keeps firing every interval against a domain that is now
live, is refused by vmsync each time, and climbs
failure_count. Disabled, not deleted, so the profile survives for the inversion. - A successful inversion moves the schedule entry to the new source's agent in a single write, flipping its target host. Two writes would let a crash land the entry under neither agent, and a missing schedule entry is indistinguishable from a VM nobody asked to replicate — the pair would stop silently, right after a failover.
The migrated entry also gets its target_disk_path re-aimed. That value
says where this direction's replicas go, so after an inversion it names the
new source's own disks — carried across unchanged, the reversed sync would
write to the wrong directory on the wrong host, and where that directory
happens to exist it would redefine the domain to match and orphan the
original disk. It is recomputed from where the new target's disks actually
are, read out of that host's own report, and left alone when they span more
than one directory (a single value cannot express that in either direction).
The migrated entry arrives disabled. The first sync in the reversed direction has no checkpoint chain and must be a full reinit, which the schedule cannot yet express; enabled, it would schedule a run that fails every interval.
Every operation carries an action_id, and it is this UI's own audit
entry id rather than a second identifier minted for the purpose. The agent
passes it to the engine as -action-id, and the engine stamps it into the
intent and outcome records it journals beside the disks it touched.
That is what closes a loop across three programs that keep separate records and share no storage: the audit entry says who asked and why, the operation says what was published to which agent, and the journal on the hypervisor says what the engine then did. The case it exists for is the one nobody plans for — somebody is holding a replica whose rebuild died, and needs to know what ran, under whose hand, and whether it was the first attempt or the third. Correlating three records by timestamp answers that until a run was retried, which is exactly the situation an interrupted rebuild tends to produce.
It is optional and omitempty in both directions. An operation issued before
the field existed carries none, an operation issued with no audit entry
carries none rather than a fabricated id pointing at an intent nobody wrote,
and an agent too old to know the field ignores it — the config an agent polls
is decoded leniently, unlike the reports it sends.
An operation can be cancelled while it has not reported. After that it fails loudly rather than pretending — the work has happened on a hypervisor and no UI state undoes it. An expired operation is still published on purpose: the agent must see it, refuse it and report the refusal, or the audit entry hangs open forever with nothing saying what became of it.
UI first, then agents. The report body is decoded with
DisallowUnknownFields, so an agent that sends a field this UI does not know
has its entire report rejected — domains, roles and sync results included —
and the symptom looks like every upgraded host going offline at once.
This has applied to operation_results, to the fence fields (fence_id,
fence_source, fence_armed_at_unix, fence_armed_by and the fenced
object), to the verification record (verify_state, verify_failed_at_unix),
and now to replica_incomplete. It applies to every future addition too,
which is why both halves of the contract are pinned by tests that name the
strings literally: TestAReportCarryingFenceStateIsAccepted,
TestAReportCarryingAVerificationFailureIsAccepted and
TestAReportCarryingAnInterruptedRebuildIsAccepted here, and
TestSendReportCarriesFenceStateUnderTheAgreedNames and
TestSendReportCarriesVerifyStateUnderTheAgreedNames in the agent. Changing
one without the other fails there rather than in the field.
The verification fields and replica_incomplete are all omitempty, which is
why the order still matters in spite of how rare each condition is: a report
from an agent ahead of its UI decodes cleanly for every healthy domain and is
rejected outright the first time a replica fails a verification or a rebuild
is interrupted — the symptom would arrive weeks after the upgrade that caused
it, on exactly the host with something wrong.
For replica_incomplete the order is UI first, then engines, then agents,
because a third program is involved: the engine is what arms and clears the
field, and -promote is what refuses on the strength of it. That refusal
rests on this field alone — nothing writes a failure_count alongside it as a
second carrier — so an engine too old to know the field will still accept a
half-written replica, on the DR host, during the incident. Every host that
drives syncs or promotions has to be upgraded before the refusal can be relied
on. Until then this console's warnings are the only thing standing between
an operator and a partial copy, which is an argument for reading them, not for
trusting them as an interlock.
The other direction is lenient, deliberately. An agent decodes the
config it polls without DisallowUnknownFields, so a newer UI sending a
field an older agent has never heard of — shutdown_timeout_sec, say — is
ignored by that agent rather than breaking it. That asymmetry is what makes
"UI first" a safe order rather than merely a preferred one: a UI ahead of
its agents degrades to the old behaviour, while agents ahead of their UI
stop reporting entirely.
(A --standalone file is parsed strictly, for the opposite reason: it was
typed by a person, and a silently ignored key there looks like the scheduler
not working.)
/failover is where a failover is actually driven. It lists every domain
that participates in replication, and beside each one only the actions its
current state allows.
That restriction is the page's whole safety design. A console that offers every action on every row is one that invites the wrong one during an incident, when the person reading it is under pressure and moving fast:
| state | what is offered | why not the others |
|---|---|---|
| an ordinary replica | Promote | there is nothing to invert until a failover has happened |
| a source | Shut down, and Invert once its target is promoted | promoting a source would ask vmsync to overwrite the original with its own copy |
| promoted and running | Shut down and Set role, both behind a typed confirmation, and Set role narrowed to source |
this is the copy serving production: stopping or relabelling it takes the workload down and discards the fence against its returning source, and neither is undoable from here. source stays because it is the one value still true of a running copy |
| promoted, never started | Finish the failover | vmsync writes the promotion record before booting, so this is what a crash or a refused start leaves behind — and it must be finishable from here |
| paused, including anything a fence stopped | Set role | the way back |
None of it replaces the checks underneath. vmsync refuses a promotion whose replica is not usable, and the agent refuses an operation whose peer does not match the VM's own libvirt metadata. This layer exists so the common case never reaches those refusals.
Rows that need a decision sort first — a split brain above all, then promoted domains, then the old sources of unresolved failovers, then anything paused. Nobody should have to scroll to find the VM that is running twice.
An agent cannot see another agent. A source's own metadata records where it replicates to, never that the target has since been promoted — so the Invert action, and the split-brain banner, exist only because the control plane hears from both hosts and cross-references them.
A -verify that found a replica's contents differing from its source leaves
the verdict on the domain, and the agent reports it. This page shows it four
times over, which is deliberate:
- a verify failed pill on the row, beside the role rather than folded into it — a replica that failed verification still carries an ordinary role and an ordinary age, and this is the only thing on the row saying its contents are wrong rather than old;
- a dated warning under the row's contents line, naming the peer it was compared against and saying that vmsync refuses every sync into it while the finding stands — the Full resync offered here included — so the repair is a sync set to recopy once, then re-verify;
- a warning on the promote cell outside the collapsed control, repeated as a pill on the control's own summary. Everything else in that cell collapses so a row stays one line at rest, but a caution that has to be opened to be read is one a hurried operator has already clicked past;
- and the same pill and date on the source's row, because that is where Full resync and Force clean resync are offered. Force clean is the one control vmsync lets past the finding, and it does not repair the copy — it drops the record without re-verifying — so its note says so where the operator is about to click it.
Open the control and the form says the rest: vmsync refuses the promotion on its own, only force gets past that refusal, and ticking it changes what vmsync allows, not what the copy contains.
The Promote button is not withdrawn. That is the point, not an oversight.
This console publishes instructions and enforces nothing (see What it does,
and what it deliberately cannot), so withdrawing the button would not prevent
the promotion — it would only hide it from the one place that explains what is
wrong, while vmsync -promote -force-promote on the hypervisor stays exactly
as available. The console reports and warns; the engine refuses.
The one place this console does withdraw controls, and the reasoning is the mirror image of the paragraph above: these controls would be refused by the engine, so offering them can only produce a failed operation the operator had no way to predict.
A VM that was failed over to carries last_promoted_at, and that field survives
every role change — which matters because every route out of promoted rewrites
the role. Shutting the copy down records paused. A fence records fenced.
Either way the promotion record goes in the same write, and an hour later the row
looks like an ordinary idle replica: same role, same checkpoint, same recent
sync, same zero failure count. That is the state in which this console used to
offer Roll back over it, and Force clean resync over it from the
source's row, with nothing between the click and a full overwrite.
So while the record is set and this copy is not the authoritative one, the row shows a has served live pill and a warning explaining what is refused, and:
- Roll back is withheld;
- Full resync and Force clean resync are withheld on the source's row, with the warning shown there too — that is where they are fired from, and the record is on the other end;
- Set role stays available, but
targetis dropped from its menu.sourceandpauseddestroy nothing, andsourceis how you keep this copy.
Separate from the above, and stricter. A row that is promoted and running is
the copy currently serving production after a failover. Stopping it takes the
workload down, and there is no undo from this page.
So on that row alone:
- Shut down cleanly and Set role both require the operator to type the VM's name, and both name the VM and the host in their button labels. A page that lists a whole estate puts those buttons next to each other on every row, and nothing else distinguishes the one that stops production.
- Set role offers
sourceonly. Everything else would record that something else is the live copy while this one is taking writes, and vmsync refuses those outright.
The fence is the reason Shut down carries a warning and Set role does not. A
promotion that armed one holds a token authorising the agent on the displaced
source's host to stop that source when it comes back. Shutting this copy down
records paused afterwards, which discards it; source keeps it, because
that transition rewrites this end only and the arrangement the token describes is
still in force. Where it is discarded, no role change puts it back — re-arming
means vmsync -update-role promoted and then vmsync -promote -fence-source. The
warning appears only when a token was actually armed; a drill that armed none is
not warned about losing one.
source is a step, not a resolution: the other end still records itself as a
source, so nothing replicates either way until an Invert from that end's row
finishes the pair. The row keeps reporting as unresolved until then.
The confirmation is enforced when the form is submitted, not by the input's
required attribute — a stale tab, a back button or a curl skips that, which is
the same hole the server-side re-check exists to close. Every other row on the
page stays one click, because every other row is undoable.
Two roles are excluded from the pill and the warning, and the second one matters:
promoted (the row already says it in the present tense) and source — the
primary an inversion produced. That pair is resolved, and the record is still
there only because the inversion deliberately keeps it, to refuse a later
-update-role target on the live primary. Marking it as undecided put a standing
"release this" instruction on the primary of every failed-over-and-inverted pair;
following it would have stripped the guard, and it could not be followed anyway —
the release refuses a running domain. The controls still read the record; only
the reporting excludes these two.
The way out is printed, not offered:
vmsync -target-uri qemu:///system -target-domain web01 -release-promotion
There is deliberately no button for it, no operation kind, and no field on an
operation that carries it. Everything else vmsync guards can be satisfied from
this console; this is the one that asks for a person at a shell on the host
holding the data. If the failover stands, Invert is the alternative that
keeps the data — and it is offered on the old source's row even after the
promoted copy has been shut down and demoted to paused, fenced or target,
because the same record is what tells vmsync that end really did serve. That
matters because demoting the copy is the normal precondition for reaching for
an inversion, so a console that only offered Invert against a promoted peer
recommended a remedy it then withheld. On a fan-out source the row is also about
the copy that served rather than whichever target happens to be listed first.
Every action is also re-checked server-side when it is submitted. The
predicates that decide what a row offers are evaluated again, against state read
at that moment, before an operation is created — so a tab left open across a
failover, a back button, a re-submitted POST or a hand-made curl cannot issue
an action the page would no longer draw. This does not make the console an
enforcement point (the engine is still the only one); it makes the console stop
promising things the engine will refuse.
A full copy into an existing target — -reinit, a force-clean, an ordinary
full sync — renames the replica's good disks to
<disk>.vmsync-replaced-<stamp> and writes new base images in their place,
with no overlay, while the target domain keeps its old metadata. Only a
run that finishes updates that metadata. So a run killed part way — a dropped
link, a restarted agent, a power loss — leaves a domain whose disks are a
truncated copy and whose last_checkpoint, last_sync and failure_count
all describe the complete copy it renamed aside.
Nothing else on this page moves. The role is unchanged, the status is
unchanged, and contents as of is computed from that stale record, so the
one figure a promoting operator reads as the data-loss window is not about the
disks at all. vmsync arms replica_incomplete on the domain before it starts
writing and clears it in the same metadata write that records success, and
this page shows what survives:
- a partial copy pill on the replica's row, beside the role like the verification one and for a harder reason — that marker at least moves the status word, and this one moves nothing;
- a dated explanation under the row naming the verb that armed it, when that
run started, which host it was writing to, and the
.vmsync-replaced-<stamp>suffix the complete copy may still be sitting under. That last part is the actionable half: nothing removes those files for you, and they are usually the last whole copy of the VM on that host; - the same pill and explanation on the source's row, because Full resync is offered there and is the only repair — the marker is cleared by the same write that records a rebuild succeeding, so nothing but a run that finishes clears it. The source's own metadata records none of this: it does not know the copy it wrote died half way;
- and a warning on the promote cell, outside the collapsed control and repeated as a pill on its summary. Open it and the form says vmsync will refuse the promotion, that only force gets past that refusal, and that forcing changes what vmsync allows rather than how much of the copy exists. It points at the displaced disks first: promoting the complete older copy by putting those files back loses a known amount of data, while promoting the partial one loses an unknown amount.
An unreadable value still warns. The value is a single comma-separated
k=v line, parsed here only for display, and a verb from a newer vmsync — or
a corrupted value — renders the warning anyway with the agent's own text shown
verbatim beside it. The absence of this marker is what reads as safe to
promote, so it may only ever fail towards the alarm. Unknown keys are ignored
rather than poisoning the parse, which is what stops the first key a newer
engine adds from turning every affected row into an unexplained blob.
The refusal is not this console's. It lives in vmsync, on the host holding
the disks, because -promote and -restore run locally on the target and
during a real disaster the source host is gone — anything a refusal depends on
has to be readable on the survivor. This page reports and explains; see
Upgrade order for why every engine that drives syncs must be upgraded before
that refusal can be relied on.
Each row shows what the domain's disks occupy and what is free on the
storage under them, because an inversion's -replaced-disk-action=rename
keeps the displaced copy and therefore needs room for both at once. Where
the figures say it will not fit, the row says so. Where nothing measured
them, it shows a dash rather than 0 B — a dash is missing data, and 0 B
reads as a measurement.
The promote form has a stop the old source checkbox. Ticking it arms a fence: the displaced source's agent reads the token from the promoted domain's own libvirt and shuts that VM down, so one VM does not end up serving in two places. It is off by default and opt-in the whole way down — a DR drill is a promotion too, and a drill that stopped production would be worse than the split brain it was rehearsing for. See the fencing section of the main README.
Note what does not travel: which source to fence. The agent has vmsync
resolve that from the promoted domain's own replica_source, so this
console can request a fence but can never choose its victim.
target, source and paused only. promoted is deliberately absent:
promotion has real preconditions — vmsync verifies a usable replica exists
and reports the data-loss window it accepts — and writing the role directly
would reach the same recorded state having checked none of them. The button
for that is three columns to the left.
This is also the way back from a fence: a fenced VM is paused, and setting
it to target lets it receive again.
Any operation that has not reported can be cancelled, which stops it being published and frees that VM for another. It does not undo one an agent has already run — the work happened on a hypervisor, and no console state changes that.
A VM a fence stopped and a VM an operator paused are both just paused
in libvirt. Nothing in the metadata tells them apart, and they call for
completely different responses. So agents report what their own fence ledger
says, and the row shows it:
fenced— stopped by vmsync because a named peer was promoted, with when, by whom, and which copy displaced it. Explicitly not an administrative pause.fence failed— a fence was attempted and the domain is still running. This one exists nowhere else: the attempt leaves no mark in libvirt, the fence is latched so nothing retries it, and the VM simply keeps running beside a promoted copy — indistinguishable, without this, from a failover nobody has got to yet. It sorts above everything, including the inferred split brain, because something already tried to resolve this and could not.
The alarm clears when the domain stops, not when the ledger changes: the ledger entry is latched forever by design, so a fence that failed in March and was then handled by hand would otherwise still be shouting in December.
A promoted row also shows whether its own promotion armed a fence, and against whom. Absence is the drill.
A clean shutdown that overruns is reported as a failure — and when a fence asked for it, fences latch, so nothing tries again and the console shows a live split brain that is really just a database taking its time. Setting this per VM is what avoids that.
Three steps, in order: the VM's own value on its schedule row, the estate default under Estate defaults on the schedule page, then vmsync's own 300 seconds. A blank per-VM box inherits, and shows the estate value as its placeholder.
The same order is implemented in the agent, because the two paths that shut a domain down are not the same code:
- a shutdown operation carries the value resolved when it was issued, so the instruction means the same thing whenever it runs — one that silently meant 300 seconds in March and 900 in April, because somebody edited a setting in between, is not one anybody can audit;
- a fence has no operation behind it and resolves from the config it last polled, because during the partition that usually causes one there is no control plane to ask.
The agent clamps whatever it is sent (30s–3600s) rather than trusting it: this UI is a separately-versioned program, and the number decides how long a production VM is given before its shutdown is called a failure.
On-demand "sync now". The per-host target budget is stored and served but
has no form (it is a row per host, unlike the two defaults above). reinit
exists as an operation kind and is not implemented on the agent side.