Skip to content

Latest commit

 

History

History
1428 lines (1200 loc) · 110 KB

File metadata and controls

1428 lines (1200 loc) · 110 KB

Runtime installation, update and diagnosis

POLICY.md owns repository rules and the operations contract owns the operational ones. This page describes how the runtime behind the MCP bridge, the session relay and the completion hook is installed, updated, rolled back and diagnosed, and it is written against that contract's clause numbers so a reader can check a claim against the rule it came from.

The runtime is one Go binary, crw, shipped in a release archive. The plugin package declares the server and the Stop hook and reaches the binary through the installer's pointer (plugin packaging); the package carries no runtime of its own.

Updating is the half that can lose something. A first install has nothing to destroy; a second one is standing on a runtime somebody is using and a database nobody can rebuild, so most of what follows is about what is read before anything moves and what is put back when it does not. Updating an installation is where that lives.

Command What it does Contract
crw install install, crw install update Verify a release archive, install it as a new runtime directory, exercise it and move the owned pointer to it OPS-2.4
crw install rollback [<dir>] Point the owned pointer back at the runtime the last promotion replaced, or at a runtime directory the host record lists OPS-2.4
crw install remove <dir> Delete one runtime directory nothing selects, points at or runs out of OPS-2.4
crw install register-mcp [--owner plugin] Write the bridge record the plugin's declared server reads OPS-2.2
crw install hook [--owner plugin] Write the Stop settings the plugin's declared hook reads OPS-6.3
crw install register-service [--remove] Write and enable the one systemd user unit that starts the relay service when the user manager starts; with --remove, disable and delete it OPS-4.1, OPS-6.1
crw install status, crw doctor Read the installation, classify it and report the six check results; write nothing (the relay readings behind them are the relay's doctor without --probe-write, which creates no file of its own and never opens its write gate; what it still does is named under Installing the runtime) OPS-2.1, OPS-2.2, OPS-6.1
crw-dev skills link --check or --apply Skill links into Codex, from a checkout OPS-2.3

Every crw install and crw doctor command prints one JSON document. Runtime installation is never folded into the skill links: crw-dev skills link belongs to the repository's development binary because it links a checkout, and a release archive has none. It stays idempotent, it refuses to replace an existing directory or a foreign link, and its LINKED, MISSING and CONFLICT words mean the same thing wherever this page uses them. A plugin installation has no skill links at all.

The Python installer, scripts/runtime_install.py, installed the Python fence release and was the development and rollback path until todo 44 removed the Python execution path; the Python fence installer is the one section of this page about it. Moving a host from the Python runtime to this one was the cutover, not an install alone: it moved the store's ownership. It ran at todos 42 and 43; the Python runtime left the repository in todo 44 (decision 48) and the takeover controller in refactor R1 (decision 54), so no order between crw install install, which moves the pointer, and a takeover is left to settle.

What an installation is

A release archive is crw_<version>_<os>_<arch>.tar.gz, published with a SHA256SUMS beside it (releases). It holds crw, the two compatibility names codex-session-relay and codex-thread-bridge as links to it, and the licences. crw dispatches on the name it was started under, so each name is the component it names. The completion hook is crw hook; its former crw-completion-hook link is retired (decision 66), and a runtime installed before that still carries one.

An installation of it is two things under the destination:

Path What it is
<destination>/bin-<version>-<digest12>/bin/ The runtime: crw and the two links, where <digest12> is the start of the archive's SHA-256
<destination>/current The owned pointer: a directory symlink naming the selected runtime directory

The destination is ~/.local/share/crw-runtime and nothing else. Both of the plugin's declared commands name $HOME/.local/share/crw-runtime/current/bin/ and nothing else, because HOME is the one variable a hook and an MCP server both receive (how hooks and MCP servers load). So crw install has no --dest, and neither it nor the wiring honours XDG_DATA_HOME: every crw install command acts on <home>/.local/share/crw-runtime, where <home> is HOME, or this user's passwd entry when HOME is not set. A host record whose pointer names another link is refused by every command before it acts, naming the repair, and crw install status reports that reading as destinationAgrees (decisions 11 and 38). For a temporary or isolated installation, run the commands under another HOME, and move with it everything HOME does not decide, as the isolated-home integration test (internal/runtime/integration) does: CODEX_HOME and XDG_STATE_HOME inside the same tree, CODEX_SESSION_RELAY_SCOPE_DIR set to a directory there, and CODEX_SESSION_RELAY_STATE and CODEX_SESSION_RELAY_MARKER_ROOT unset or pointed there too. Each is read on its own. A CODEX_HOME left naming another Codex home has the install judge that home's registrations and crw install hook write its Stop settings naming the temporary pointer; an XDG_STATE_HOME left naming another state home puts the temporary host record there, or is refused where the record there names another pointer; and without CODEX_SESSION_RELAY_SCOPE_DIR the relay finds its scope registry from this user's passwd entry, never from HOME. crw doctor still takes --dest, to read another destination, never to install into one.

A path the commands cannot use as given is refused rather than guessed at. HOME has to be absolute, hold no .. and not start with exactly two slashes (//home/..., which pathlib keeps as spelled and a lexical join folds to one; three or more fold to one in both). A relative XDG_STATE_HOME is a usage error (exit 2) to every crw install command not given --record: read against the working directory it would put the host record where nothing else looks. The doctor does not read against the working directory either, but it reports rather than refuses: crw doctor answers hostRecordState ACCESS_ERROR and exits 0. A path from HOME, CODEX_HOME, XDG_STATE_HOME or a path flag that holds a byte that is not UTF-8 is a usage error naming where it came from, because a record or settings document written with a replacement character names a file that does not exist. The execution policy path is the one exception, recorded as os.fsdecode spells it (the execution policy).

The host record, ${XDG_STATE_HOME:-~/.local/state}/codex-relay-workflow/host-record.json, says which runtime is selected and records every install, its measured points and who placed the pointer (the host record). The two settings records the plugin's commands read sit in the Codex home: crw-bridge-mcp.json for the server and crw-completion-hook.json for the Stop hook.

Nothing here manages PATH. The declared server and hook name the pointer by absolute path, but a skill command that runs codex-session-relay finds whatever PATH finds (how skill commands reach the relay).

Installing the runtime

crw install needs a crw to run it. The archive carries one, so unpack it anywhere temporary and run that copy against the archive itself:

tar -xzf crw_<version>_<os>_<arch>.tar.gz -C <scratch>
<scratch>/crw install install --from crw_<version>_<os>_<arch>.tar.gz \
    --socket <app-server-socket>   # SHA256SUMS beside the archive, or --sums <file>

--release <tag> fetches the archive for this host's target and its SHA256SUMS from that GitHub release instead of --from. Either way the archive has to be named for this host's operating system and architecture, be listed exactly once in SHA256SUMS and hash to the listed digest, and nothing under the destination or in the host record is created before all three hold.

--backup-state-to <dir> (on install, update and rollback) is the acknowledgement that the swap brings the additive DAG zone, or ordinary indexes on tables the store already holds, to a store that lacks them: the command copies the whole relay state directory to <dir> before promoting, and without it that swap refuses and names the flag (the route).

What follows is one run, in this order, and the result lists the steps it took:

  1. Claim the runtime directory with an exclusive mkdir and a claim file (the claim a run leaves behind).
  2. Unpack the archive into it and read the binary's digest.
  3. Record the install entries.
  4. Exercise the candidate through its own concrete executables, never through current, which still names the predecessor: the relay's doctor must report actorReachability.socketConnect as ok, and the bridge must answer an MCP session that lists its tools and calls get_capabilities. Both run against the App Server socket --socket names. The bridge falls back to <CODEX_HOME>/app-server-control/app-server-control.sock without it, but the relay has no default socket: its doctor answers socketConnect as not configured, so a run without --socket fails at exercise the candidate (exit 1) even with an App Server listening at that path. A run that cannot exercise the candidate records no point and promotes nothing.
  5. Under the host-wide promotion lock: read whether it is safe to swap, establish that the pointer is this command's, refuse a second owner of the bridge or the Stop hook, and refuse Stop settings that name through the pointer something a Go runtime does not serve (one Stop settings document); nothing rewrites them.
  6. Commit the selection, then replace the pointer, and read the pointer back.
  7. Settle the claim.

OPS-2.4 sequences an update as measure, install, measure again, and step 4 is the measurement that produces the point; promoting before it would select a runtime that unpacks cleanly and fails the moment it is used. A failure at any step up to the promotion leaves the previous runtime selected and the pointer where it was (what a failed update restores). Nothing here removes, moves or recreates the store: update failure and store loss are different accidents and the recovery for one must not cause the other.

The exercise and the swap gate read the store through the relay's doctor and service status and a catalog read that takes no lock. They run doctor with no option, so it only reads: it opens the database read-only, creates no .probe- file and no SQLite sidecar where SQLite allows, makes no read-write connection and never opens write-gate.lock; doctor --probe-write, which writes a temporary file and begins and rolls back a write transaction, is not what they run (decision 76). The ownership reading copies the database into the temporary directory, the worker-policy reading takes the daemon lock for an instant when a worker record exists, and a store left after an unclean shutdown is read the plain way, which may create SQLite's -shm index beside it; none of these writes the store's data. The runtime opens it afterwards, for the relay commands the skills run and for the Stop hook's guard whenever it has to read the store, in the relay's default state directory with no variable set. Until todo 43 the Go build refused that directory unless CRW_ALLOW_LIVE_STATE=1 was set, so an install left a runtime that could not serve a live host's store; the guard now refuses it only under test isolation (the live-state guard). A host whose store the Python runtime still owns moves through the cutover first.

The record is not the replacement

The staging claim is written last. It says this staging finished, and until the selection is committed and the owned pointer names the runtime there is nothing finished to say, so by the time writing it can fail the declared commands already reach the new runtime. The replacement has happened and only its record has not, and those are reported as two outcomes rather than folded into one.

Field Answers
promoted whether this run replaced a runtime
inService whether this runtime directory must be kept: true while the record selects it or the pointer names it, true once its claim has settled, and true when none of that could be read. False only when the readings say so
claimSettled whether the claim recording it was written
claim the claim's own outcomes (settled, released), its read-back and the selection snapshot that decided them
recoveryRequires what has to be done next, under the same key a refusal reports it

So crw install install has four exit statuses:

Status What this run changed The record What it means
0 it landed, or there was nothing to change written the run finished; alreadyInstalled says when this archive was already the selected runtime
3 it landed not written the pointer names the new runtime and the claim that records it did not settle
1 nothing not written refused, or failed and put back what it had changed; whatever the host selected and reached before, it still does
2 nothing not written a usage error, found before anything was read

Exit 3 is not a refusal and must not be read as one. Non-zero here means the opposite of what it means everywhere else in this command: the change landed, and a process may be running out of the runtime it changed. A wrapper that reads every non-zero status as "nothing changed" would report the old runtime as selected, or clean up a runtime that is in service. Key a cleanup decision on inService and never on the status alone: a competing install can supersede this runtime between the promotion and the result, and then status 3 is still correct about this run while inService is false.

Which accident happened, and what to do about it, is in recoveryRequires, derived from the claim as it reads back and from a selection snapshot taken under the promotion lock:

What the result says What to do
another run held the claim's lock wait for that run; this call wrote nothing
the claim could not be read back make it readable or remove it, then run the install again; the runtime is in service and must not be deleted
what this host selects could not be established read the host record before acting
the record selects this runtime and the pointer does not name it read the pointer before rerunning, because a rerun replaces that link first
the record selects this runtime clear what stopped the write and run the same install again: it finishes an interrupted promotion and rebuilds nothing
the record no longer selects it, and something may still reach it leave the directory alone
nothing selects it or points at it nothing; another run moved the selection on, so do not rerun to settle it

The snapshot is consistent, not durable. Nothing holds the promotion lock until an operator reads the result, so re-read it before acting if time has passed.

Updating an installation

crw install update is the same run as crw install install; the name says which one you meant. A new archive is a new runtime directory, because the directory is named for the archive's version and digest, and the predecessor is never removed by an update. It becomes the host record's outgoing selection, which crw install rollback returns to.

The pointer is what moves

<destination>/current is a directory symlink, and everything that starts the runtime names a path through it: the plugin's Stop command, its server launcher, and the three executables the settings records name. Those strings are stable across every update, so an update rewrites none of them and changes nothing Codex has trusted. It moves the link.

The pointer is a way to reach a runtime and never an identity. A process already started keeps the runtime it started in after the pointer moves: a running bridge or daemon goes on running its own directory's binary, which is why no update removes a predecessor. The next process started through the pointer is the new runtime.

A registered command and the runtime it reaches are separate claims, and crw doctor reports them separately: the pointer's state and target, which kind of runtime the target is (go-binary, or unknown), and whether the record selects what the pointer names (runtime.agrees). A current repointed by hand at another directory is caught by that comparison rather than passing because the command strings are unchanged.

One Stop settings document

crw install hook writes one plugin-owned Stop settings document, and every runtime the pointer names serves it: relayExecutable <destination>/current/bin/codex-session-relay names the runtime through the pointer, so no promotion and no rollback rewrites that document (decision 18). Every move of the pointer reads it first, under the promotion lock, and refuses, with nothing changed, when it names through the pointer a path a Go runtime does not serve (bin/crw and its compatibility links); settings a user owns are judged with the second owners. The settings are reported under settings with action none.

Until refactor R1 a move onto a Go runtime also replaced a Python-era document (one naming .../current/bin/python3), archiving it as crw-completion-hook.json.superseded-<time>, and a rollback could point at a Python env-* runtime that served the same document. The relay host's first Go install made that replacement, and no Python runtime is left on it, so both are retired (decision 61); a Python-era document found now is judged like any other and refused when it names a path a Go runtime does not serve.

The claim a run leaves behind

The runtime directory's name is deterministic and it is created with an exclusive mkdir, which is what proves a run owns it. A run killed outright would otherwise leave the directory behind and every retry of the same archive would refuse at the existence check for ever.

A run leaves two files in the directory, because they answer two questions. The lock, .crw-staging-lock, answers whether anybody is still building, and it is created once and never replaced: an advisory lock belongs to an inode rather than to a name, so a lock on a file later replaced by rename would sit on an unlinked inode while the next reader found the new one free. The claim, .crw-staging-claim.json, answers what that run said it was doing, and it is rewritten when the staging settles. Removing anything needs positive proof of ownership, so the claim has to carry this command's marker (crw install), its claim version and a state from the declared set. Readable JSON at that path is not proof, and neither is a claim the retired Python installer wrote (runtime_install.py), which is somebody else's file here (decision 63).

Observed Answer
No claim, and the directory holds files Somebody else's. Refused, nothing touched, even when the host record selects something inside it
No claim, and the directory is empty Taken over as it stands
A claim of this command's, the lock held Another run is building it. Refused, nothing touched
A claim of this command's, the lock free, nothing selecting or naming it An abandoned staging. Removed (through a tombstone, as remove removes) and built again, but only under remove's rules, read again under the promotion lock: kept, and the run refused, while a process may run out of it, a relay daemon record cannot be read, or a registration names it or cannot be read. One the record's outgoing names was in service, so it is kept and its claim settled. Where there is no process table (darwin) it is kept too, and the answer gives the recovery for a staging that was never promoted. A tombstone (.crw-removing-<name>) that holds such a claim is not reclaimed this way: an install reclaims the directory under a runtime's own name, never a tombstone for its own sake, and crw install remove finishes it (Removing a runtime)
A claim, and whether anyone holds it could not be established Kept, and reported with what recovery needs
A settled claim, every component selects it and the pointer names it Already installed. Reported, nothing rebuilt
A settled claim the record selects for one component and not another Refused, naming each component's selection; crw install rollback <dir> selects every component and swaps nothing
A selected runtime the host cannot launch as it stands (bin/crw not a regular file this user may execute, or a link that does not resolve to it) Refused, with repair: the commands that restore it in place (crw extracted from the archive the directory is named for, chmod 755, ln -sfn crw for each link), after which the same install answers already installed. Its directory is named for the archive, so it cannot be built again beside itself
A settled claim, and nothing selects it any more Kept. It is a runtime that was promoted once, and a process may still be running out of it
An unsettled claim for a runtime that IS selected An interrupted promotion. Finished rather than rebuilt
A lock held with no claim written A run between taking the lock and writing its claim. Refused, nothing touched

Finishing an interrupted promotion asks a narrower question about the link than a promotion does: not whether it agrees with the selection, which it cannot while the promotion is unfinished, but whether it still names a runtime this host record accounts for. It asks OPS-4.4 again as well, because the interrupted run recorded no gate verdict and every cell of the gate reads state that moves while nobody is looking.

Liveness is the lock and never a recorded process id: inside a container sharing a kernel the same process id under the same boot id is a different process. Where flock is unavailable the answer is that nobody could tell, and an owner nobody could establish is never read as an owner that is gone. Deciding and acting are one step, under the directory's own <env>.crw-lock and then the promotion lock (the lock order), so two retries that both find the same abandoned staging cannot both act on it.

Reading whether it is safe to swap

OPS-4.4 sequences an update around a daemon that is not running and open attempts that have been reconciled. Three readings answer that, each filling only its own cell:

Cell The reading that answers it
daemon the selected relay's service status, whose running is decided by the lock a supervisor holds
inFlight whether a store is there at all, then the selected relay's doctor, whose contents.openAttempts counts in-flight and held-uncertain attempts
storeSchema the store's own schema, read read-only from sqlite_master without opening the store for writing, against the schema the candidate binary declares

The in-flight cell reads twice, and the order is the point. The relay reports contents unavailable both for a store that is missing and for one it cannot read, and those are opposite answers here: an absent store has no open attempt, an unreadable one has an unknown number.

The swap proceeds only when the daemon is established stopped, the open attempts are established zero, and the schema is established compatible: the verdict is ALLOWED. A cell that answered no decides BLOCKED, and a cell that could not be read decides UNESTABLISHED; both keep the existing installation, and the refusal names the cells. This command never starts or stops a daemon. OPS-4.1 gives the service to the scope operator, so a running daemon is a refusal here and not something to resolve.

STOPPED describes a moment that has already passed: the relay's liveness answer releases its lock before returning. Taking the reading inside the promotion lock narrows the window to the promotion's own length; it cannot close it, because that lock excludes other runs of this command and says nothing to a supervisor.

Why the schema reading compares statements and not versions

The relay declares schema version 1, has never raised it, and grows its schema through separate CREATE ... IF NOT EXISTS statements. Every store therefore agrees with every candidate at version one, and comparing versions would detect neither a downgrade nor an upgrade while looking exactly like a check.

So the cell compares each object's CREATE statement in the store's sqlite_master with the statements the candidate declares, keyed by kind and name (index sync_ready, not sync_ready), over every object the catalog holds except the ones SQLite maintains for itself. Runs of whitespace outside quoted text are normalised away and nothing else is.

Answer Observed Decision
NO_STORE no store exists at the resolved selection allowed, and reported as absence rather than as agreement
AGREES the same schema objects, defined identically allowed
EXTENDS the candidate declares objects the store does not hold refused
DIFFERS a shared object is defined differently refused
NARROWS the store holds objects the candidate does not declare refused
EXTENDS_ZONE the only difference is the additive DAG zone arriving: every added object is the zone's refused, unless the command takes the OPS-4.5 backup (below)
NARROWS_ZONE the only difference is the additive DAG zone leaving: every object the candidate does not declare is the zone's allowed
EXTENDS_INDEX the only difference is ordinary indexes arriving on tables both sides declare (a few may leave beside them) refused, unless the command takes the OPS-4.5 backup (below)
NARROWS_INDEX the only difference is ordinary indexes leaving: every object the candidate does not declare is one, on a table both sides declare allowed

The relay applies its whole schema on every write-open, so a candidate whose schema is not the store's applies the difference the moment its daemon first starts. OPS-4.5 reserves that for its own decision, with a copied backup of the whole state directory taken first, so an update never waves it through. The Go and Python runtimes execute the same schema statements (decision 14), and no Go release changes them before the commit point (cutover).

The one release that changes the schema is the one that adds the DAG zone (DAG plans, decision 74): its dag_* tables are created by the first write-open and are declared by that build. That decision is made (D-01), so the gate has an answer of its own for the zone and, below, for ordinary indexes, and for nothing else, and the supported install command has a route through both.

The history indexes (CRW-301). A second change reaches the frozen v1 schema script itself: six CREATE INDEX IF NOT EXISTS statements for the tables that only grow (attempts, supervisor_attempts, acks, sync_outbox, supervisor_messages, managed_start_requests; the list and the reason for each are in contract/schema/relay-sqlite-history-indexes.json). They are not zone objects. Each is an ordinary index, non-unique and without a SQL function (the partial index on acks compares a column with a literal), on a table both sides declare, so the gate gives them the answers of the paragraph on ordinary indexes below (pull request #397). A candidate that declares them over a store that lacks them reads EXTENDS_INDEX, which refuses without --backup-state-to and passes with it; a store that holds them against a candidate that does not declare them reads NARROWS_INDEX, which is allowed. The ratchet test TestEveryNonUniqueIndexOfTheShippedSchemaIsOrdinaryButOne keeps the six ordinary. A store that lacks the zone as well, meeting a build that brings the zone and these indexes together, is the case that paragraph leaves out: the zone beside an index is a plain EXTENDS, and --backup-state-to does not release it. The build itself creates the indexes: the schema script runs on every write-open, so the first command to open a store of the previous version builds them, one index in a transaction of its own, and a command that declares itself read-only builds them too, since it tries the writable open first (a read that opens nothing for writing builds nothing). The swap gate tests TestTheHistoryIndexesArriveAsAnIndexExtendsAndLeaveAsAnIndexNarrows and TestTheHistoryIndexesAgainstRealStores pin this reading.

The zone arrives (EXTENDS_ZONE). Installing a build that declares the zone onto a store that has none (every store there is) refuses, and the refusal names the route: crw install update --from ... --backup-state-to DIR (the same flag is on install and rollback). The flag is the operator's acknowledgement, and under it the command itself takes the OPS-4.5 backup, so the backup is guaranteed by the route and not by anyone's memory. Inside the promotion lock, after the daemon-stopped and no-open-attempt cells have answered and before anything is promoted, it copies the whole state directory the gate read to DIR: copy only, byte for byte, the source opened read-only and nothing moved, recreated or deleted, here or on a failure. Directories and regular files are copied with their bytes and, once every byte is in place and verified, their permission bits, each synced after its mode is set so the mode and not only the bytes survives a power loss (the directory the backup is made in stays 0700, and a copy always keeps its owner able to open it: a source the installer could read only through its group gets the owner's read bit, and the manifest records both modes); a symbolic link to a regular file is copied as the file's bytes under the link's name (a link alone would back up nothing), and when relay.sqlite3 is such a link the real file's -wal and -shm are copied beside it, so a restore opens with the commits only the log held; a socket, a FIFO or a device is listed as skipped. Each file is hashed while it is read and synced; then the state directory is read again, and the listing, every size and every file's digest, and the digest of every file in the copy, must be what was copied. Any difference, in any file, refuses the swap ("the state directory changed under the copy"): the copy is of one moment or it is not made. The backup, its manifest and the directories made for them are synced in their parents after the manifest is written, so a power loss cannot keep the files and lose the names that reach them. DIR and its manifest must not exist, must not lie inside the state directory, and must not lie inside the runtime destination tree (a failed run removes its candidate runtime, and a backup there would go with it); the free space on its filesystem must cover the directory. The record is written last, beside the backup and not inside it (DIR.manifest.json: source, destination, time, issue, every entry with its size, mode and digest, what was skipped, an aggregate digest), and the command's result carries the same facts as swapGate.stateBackup. A backup that fails part-way stays where it is and is reported with partial: true; a swap that fails after the backup keeps it too, and a rerun needs a new destination.

The route carries the zone arriving alone. A running daemon, an open attempt, a cell that could not be read, another schema object arriving or leaving, or any object defined differently (a dag_* object included) still refuses with the acknowledgement as without it, and no backup is taken for a swap that does not happen. The zone's objects are the ones the build's own zone statements create, not whatever starts with dag_: an object this build does not know, or a later step of the zone ledger, is a plain EXTENDS or DIFFERS until a build that knows it is the one asking. The acknowledgement where nothing arrives is not a refusal and takes no backup. The route exists in a build that carries it: the installed current/bin/crw of an earlier build reads the same arrival as a plain EXTENDS, so the update is run with the new build's binary from the archive (the documented path).

The zone leaves (NARROWS_ZONE). Returning to a runtime that does not declare the zone, on an update, on a promotion an interrupted run left to finish, or on a rollback that moves the pointer, is not refused for that reason: the zone is additive, so an older runtime opens a store that has it (it validates only the frozen tables) and never reads or writes the zone, which stays in the store for a newer runtime to read again.

Ordinary indexes arrive (EXTENDS_INDEX) or leave (NARROWS_INDEX). A release may add an index to a table the schema already declares, and then the store lacks an object the candidate declares, which is a plain EXTENDS. The gate gives a difference made only of ordinary indexes on tables both sides declare an answer of its own. An index is ordinary when it is not unique and calls no SQL function, in its key or in the WHERE of a partial index. SQLite decides that and the gate does not match text: it builds every table both readings declare in memory from the store's own statement for it, executes the one CREATE INDEX statement (a statement with a semicolon, even inside a literal, is not classified), and asks SQLite for the table the index is on, whether it is unique and, from SQLite's own compiled program for the statement, whether it calls a function. A unique index can fail to build over existing rows and then fails the writes of a runtime that never heard of it, and so can a function: json_extract over text that is not JSON raises, which an older runtime's ordinary write would meet. Neither is ordinary, however few rows are involved. An index on a table that is new, that differs between the two readings, or that the statement does not name correctly is not ordinary either, and neither is a trigger, a view, a new table or a column change: each keeps the answer it had (EXTENDS, DIFFERS or NARROWS), alone or beside an ordinary index, and the backup route does not carry it. The zone beside an index is neither class (a store without the zone meeting a build that brings the zone and an index together is a plain EXTENDS, as is the zone and an index leaving together a plain NARROWS), so the route is taken for the zone first, with a build that carries it alone (refactor backlog).

An arrival refuses without --backup-state-to and passes with it, by the same route as the zone: the backup is taken after the daemon-stopped and no-open-attempt cells have answered and before anything is promoted, and the candidate builds each index on its first write-open. The schema script runs as separate statements, each in its own transaction, so the write lock is held for the build of one index at a time, never for the script. Measured on a synthetic store built by this build (the zone present, the six history indexes of pull request #384 absent, filled with the generator of that change's benchmark; page cache warm, on a host that was running other jobs; WAL mode, so readers were never held), with a second connection updating one row every 2 ms to see the lock:

Store Events / attempts The first write-open (six indexes) The largest index (attempts_sent_at) Each of the other five Longest write stall Growth of the file
89.8 MB (larger than the operating store's about 65 MB) 80,000 / 160,000 0.12 to 0.13 s 0.10 s 3 to 22 ms 0.08 to 0.105 s 8 MB
447.8 MB 400,000 / 800,000 0.50 to 0.56 s 0.48 to 0.54 s 2 to 45 ms 0.33 to 0.43 s 40 MB

The build is a small fraction of a second at the operating store's size and scales about linearly with the attempts table, which holds the one large index. A writer that meets the lock waits for it: the probe's longest wait can exceed the build by up to one busy-handler retry step, and no probe write failed. Three runs of each size are in the figures, and the readers' longest wait was 3 ms.

An index the candidate does not declare leaves with no acknowledgement and no backup, on an update, on a promotion an interrupted run left to finish, and on a rollback. An ordinary index is derived from its table's rows and is never needed for correctness, and SQLite maintains every index of a table on every write whichever runtime made it, with no function that could fail, so an older runtime opens such a store and keeps it up to date without reading it; the index stays for a newer runtime to find. The one effect is on speed and on the order of a query that has no ORDER BY, which may now follow a leftover index. A unique index, or one that calls a function, that the candidate does not declare is not that: an older runtime's write can fail on it, and the store reads a plain NARROWS. A new index meant for this route is therefore non-unique and function-free; the ratchet test names the one index of the shipped schema that is not (journal_managed_creation, whose WHERE guards json_extract with json_valid), and a new exception fails it. The test of what calls a function reads SQLite's compiled program, whose opcodes SQLite does not promise to keep, so the tests keep json_extract as a control that fails when a driver upgrade changes them.

The warning against write commands stays: opening a store for writing with this build creates the zone, and that happens outside this route (a relay command run by hand against a live state directory takes no backup), so the route makes the warning unnecessary only inside the install command. Run no write command of this build against a live state directory before the install has taken its backup.

The order a swap commits in

The selection in the host record and the pointer on disk are two truths, and both the order they are written in and the lock they are written under are the safety argument. They are written inside one critical section holding the promotion lock, and the selection is committed first.

The reverse order has a real failure: the link lands, the record write then fails, recovery reads a selection that does not name this runtime, concludes the candidate was never promoted, removes it, and leaves the declared commands pointing into a directory that no longer exists. OPS-4.4 requires every state transition to be committed before its side effect. Recovery also refuses to remove a runtime the pointer names, so neither truth alone can authorise deleting a runtime the other one is still using. Every judgement the promotion makes is decided on state read inside that lock.

The promotion lock is an advisory lock on one host-wide file beside the host record, host-record.json.promotion-lock, created once, never unlinked, and released by the operating system when its owner dies. It excludes runs of this command and of the Python installer, which take the same lock, and nothing else. A kill inside the window leaves a runtime that is selected and unsettled, which the next run recognises and finishes rather than rebuilds.

Ownership of the pointer is established from the record before it is replaced. Renaming over an existing symlink succeeds whoever created it, so a current this host record never recorded placing is left alone; a real directory at that path fails the rename outright, which is the safe direction. The pointer placed is then proved: it has to resolve, with no loop and nothing dangling, to the runtime directory itself by identity, and <pointer>/bin/crw has to be a regular file this user may execute.

Every crw install command takes its locks in one order (decision 33): a runtime directory's <env>.crw-lock first, then the host-wide promotion lock, then the host record's own .crw-lock inside both. The settings records' locks, a claim's and the launcher copy's are leaves: nothing else is taken while one is held. It is runtime_install.py's order too. Every wait before a command's first write ends when the command is interrupted (SIGINT, SIGTERM, SIGHUP): it stops waiting, writes nothing and says so (register-mcp and hook answer the outcome interrupted). A wait inside a sequence already under way, such as the restore after a failed promotion or the entry drop after a directory was set aside, runs to completion, because stopping there would leave the host half-written.

What a failed update restores

A failed update leaves the previous runtime selected, the previous pointer target in place, the Stop settings as it found them and the store exactly as it was. The result says which step failed rather than only that something did: failedStep names it, retriable says whether the same archive can be installed again, residualPaths names what the run left, and recoveryRequires what has to happen first.

Putting a selection back is narrower than it sounds, and deliberately. The restoration runs under the promotion's own lock and puts back only the entries that still name what this run wrote; an entry another run has promoted since is left alone. Putting a selection back includes putting it back to nothing, for a first install that had no previous selection, and putting the pointer back includes putting it back to absence. A restoration that cannot be read back reports a residual pointer and keeps the candidate rather than claiming the rollback completed.

A candidate that is selected, whose record could not be read, or that the pointer names, is kept. Otherwise the run drops its install entries and removes the directory it created, through a tombstone as remove does, and the result says whether that removal was verified: verified means the same archive can be installed again; not verified names the residual path and what recovery needs, with the original failure reported beside the cleanup failure.

Two failure points an update might be expected to have do not exist here. Registration is not part of an update: the declared commands name the pointer, so nothing is registered again. And nothing starts: OPS-4.1 gives the service to the scope operator, so this command refuses while a daemon runs and never starts one.

This is a POSIX path. The runtime layout, the directory symlink and the advisory locks are POSIX assumptions; Windows is out of scope rather than approximated.

Rolling back

crw install rollback points the owned pointer back at the host record's outgoing selection, the runtime the last promotion replaced. crw install rollback <dir> points it at a runtime directory the record lists exactly, as an install entry's environment, instead; an empty <dir> is a usage error. It takes the target's <env>.crw-lock and then the promotion lock (the lock order), finds the target again there and refuses with nothing written when the record moved it, and then applies a promotion's gate, ownership and second-owner rules, commits the selection before the link, and proves the pointer it placed. The runtime it leaves stays installed and becomes the outgoing selection, so a second rollback returns to it. Where there is no outgoing selection, a bare rollback refuses and writes nothing.

The target has to be a Go runtime (bin/crw) launchable as it stands, judged without running it, and its claim has to be settled, or unsettled with nobody holding it where the record shows a promotion committed it (an exit 3). A rollback never rewrites the Stop settings (one Stop settings document). A directory the record lists that is not a Go runtime - a Python env-* runtime the fence installer made - is refused with nothing changed (decision 61).

Install rollback is not takeover rollback. crw install rollback moves which runtime the pointer names, and it rewrites no settings. It does not move the store's ownership. After the cutover the store is owned by the Go runtime, and handing it back to the Python fence release was crw relay takeover rollback --to python --python-relay <path>, which named the Python relay by absolute path and never through the pointer (cutover rollback); since todo 44 no Python candidate is launched and that rollback is refused (decision 48), and refactor R1 deleted the takeover command (decision 54). One does not imply the other, and a return to Python needed both; with the store's half refused and crw install rollback taking only a Go runtime, this revision offers no return to Python. The one order fixed for the pointer is the plugin payload's (update and roll back): the payload goes back before the runtime does.

Removing a runtime

crw install remove <dir> deletes one bin-* runtime directory directly under the destination, and only when nothing may still be using it; a Python env-* directory the retired installer made is not one it takes (decision 63). The directory is judged by file identity, so naming it through a symlink, a bind mount or another spelling of the destination passes no check its own name fails. It refuses a directory the record selects, one the pointer names or might name, one with no claim this command wrote, one whose claim cannot be read, a staging another run still holds, any directory a live process runs out of (its /proc/<pid>/exe, its working directory, or what its arguments run, resolving inside it, and every relay daemon a daemon.json or scope claim records alive), and any directory a registration the host reads names a path inside: the Stop settings, the bridge record, the cached plugin declarations, hooks.json and config.toml. A process it cannot rule out or a CRW registration it cannot read or judge refuses too, and the answer names it; another program's entry in hooks.json or config.toml that it cannot judge does not (what remove reads). A process whose working directory cannot be read and that runs something by a relative path is one it cannot rule out, whatever user runs it, unless that user (not root) is provably shut out of the directory by the permissions of the directory or a parent. What counts as running something is the program and, for a shell (sh, bash, dash and the like), the script it reads; a runtime directory holds no Python, so an interpreter's script operand is not followed (decision 62). So on a host where a root agent runs a relative shell script remove refuses, names the process, and gives the manual recovery below; a root Python agent run relative (every Azure VM's WALinuxAgent) does not refuse it.

The process reading is this host's process table, in this command's PID namespace, and every answer that rests on it says so under processTable: a process in a container sharing the directory, or on another host sharing a network home, is not seen, so run crw install remove where the runtime's processes run. Where there is no process table it can read (darwin has no procfs), whether a relay or a bridge still runs out of the directory cannot be established, and every remove refuses, fail-closed, with the recovery by hand under recoveryRequires: stop the relay daemon started from it (<dir>/bin/codex-session-relay service stop) and end every Codex session whose bridge it started, delete the directory, then run crw install status to see that the host record and the pointer still name the runtime you meant. Its install entries stay in the record, where a rollback naming it is refused because the directory is gone.

A removal survives a kill. Under the directory's <env>.crw-lock and then the promotion lock, the directory is renamed in one step to its tombstone, <destination>/.crw-removing-<name>; its install entries leave the host record in one write, with an outgoing that names it; only then is the tombstone deleted. A drop that cannot be written renames the directory back and refuses, and a deletion that does not finish exits 3 naming the tombstone. A kill anywhere leaves either the whole directory or a tombstone. crw install status lists every tombstone under interruptedRemovals, and crw install remove of the tombstone, or of the runtime's name when only its tombstone is left, finishes it once no process runs out of it and no registration names it. Do not delete a tombstone by hand: finishing it also drops what the host record still lists under the runtime's name. crw doctor's residue lists a tombstone whose claim is an abandoned staging's (a reclaim, or a failed run's candidate, killed before its deletion finished) and names the same crw install remove for it as status does. A tombstone a removal of a settled runtime left, or an empty one, is not residue by the staging decision, so status alone reports it. A tombstone that holds files and carries no claim of this command's is somebody else's directory: status reports it ours: false, and nothing removes it.

What remove cannot see is a command fixed before the update that will start a process later: a turn holds the hook command it resolved when it started until it ends (the turn-command cache). Remove a runtime the pointer left only once the turns that started before the pointer moved have ended.

What remove reads

Every registration the host reads is read as the program that reads it runs it, and each path it names is resolved through the owned pointer and every link, as are the interpreters and scripts it is followed through. Each finding and each entry that could not be read or judged names its row, the source file and the field; the row numbers are those of the retention scan the command's readers came from, which crw doctor retention-scan ran until it was retired (decision 59):

# Registration What is read
4 <CODEX_HOME>/crw-*.json (crw-completion-hook.json, crw-bridge-mcp.json, any other), and every settings document a Stop command below names relayExecutable, bridgeExecutable, interpreterPath, command, each as written (the retired adapterEntryPoint and adapterInterpreter run nothing and are not read, decision 66) (the launchers run them with no shell and no expansion), and each args entry as an argument. A value that is not an absolute path, or not a string or a list of strings, is unreadable
5 <CODEX_HOME>/plugins/cache/crw/crw/*/wiring/hooks/*.json, wiring/mcp.json, .mcp.json in every version directory (a stray file there declares nothing) every hook command, under the grammar below, and every MCP server, as Codex starts it
9 <CODEX_HOME>/hooks.json every hook command, as row 5
10 <CODEX_HOME>/config.toml mcp_servers.* every MCP server, as row 5 (a relative cwd is unplaced)

hooks.json and config.toml hold other programs' registrations too, so what the reading cannot judge of one entry there (a hook command, an MCP server) holds a removal back only when the entry is CRW's: when a word it names or runs, a path it classifies, or where that path resolves is crw, codex-session-relay, codex-thread-bridge or crw-completion-hook, or lies in the destination (spelled from $HOME, ${HOME} or ~ too). Another tool's SessionStart hook running an interpreter's script, or a server started over ssh or by node, leaves a CRW runtime unused and no longer refuses a removal (decision 68); on the relay host three such entries did until todo 43 ran the removal against an edited copy of the Codex home. An entry with a field of the wrong type (a server whose args is 42, a hook whose command is a list, an event that is not a list, a server that is not a table) is judged the same way: every string it holds is read as a shell line and as one program, under the entry's own cwd, env and PATH, and searched for the names above and the destination, and the entry holds a removal back only when that finds CRW in it. Its server name counts (CRW registers the bridge as codex-thread-bridge), its map keys are matched by spelling and never read as paths, and a number or boolean names nothing. What any entry names inside the directory is still found, a file that cannot be read or parsed, or whose top-level structure is not what the host reads (hooks that is not an object, mcp_servers that is not a table), still refuses, and the CRW files (rows 4 and 5) refuse on anything they hold that cannot be judged. crw doctor reads the same two files as one document the way Codex accepts it, so there a server of another program with args = 42 makes the bridge's registration unreadable; the doctor asks whether the selected runtime is registered as the host reads it, and remove asks whether any registration can name the directory it would delete.

Row 8, the <CODEX_HOME>/crw-stop-hook.py launcher copy, is no longer read (decision 67): the relay host holds none, and the Python bootstrap that fell back to it left the plugin cache with the pre-native payloads. A Stop command's settings are the document the word after crw hook (or after a crw-completion-hook link an older runtime still carries) names, which a runtime installed before decision 66 reads, else <CODEX_HOME>/crw-completion-hook.json; CRW_COMPLETION_HOOK_CONFIG is not read.

A hook command, a shell script a reference reaches, and a program handed to sh -c are parsed with mvdan.cc/sh/v3/syntax (Bash grammar; decisions.md 37) and judged only as far as they are written in a grammar read completely: simple commands joined by ;, &, &&, ||, |, |& and newlines; words that are literal once a leading ~, $HOME, $CODEX_HOME and (row 5) ${PLUGIN_ROOT} are expanded; redirections to such words; and as commands exit, true, :, exec, an absolute path or a bare name found on PATH (a relative or empty PATH directory met first leaves the name unreadable), sh, bash and dash with -e, -u, -x, -f and -c, and env with -i and --. Every other construct is unreadable and listed, never interpreted. An MCP server is judged as Codex starts it: command and args exec'd with no shell and no expansion, in its declared cwd, with its declared env over HOME and PATH.

The relay records are every daemon.json in the relay state root, each directory under it (a link to one included), CODEX_SESSION_RELAY_STATE with ~ expanded, the state directory --state names, and each stateDir the relay's scope registry records, together with the registry's own claims. The registry is the one the relay resolves in this command's environment: CODEX_SESSION_RELAY_SCOPE_DIR alone when it is set, otherwise <passwd home>/.codex-session-relay/scopes. A pid counts only while its start time and boot still match the record; a pid or workerPid that is present but not a positive integer pid, a record naming a boot whose current id cannot be read, and an alive process whose executable or command line cannot be read are unreadable.

Reading an installation

Two read-only commands, each answering a different question:

  • crw install status answers what a replacement actually left: what the host record selects, what the owned pointer names and whether they agree, whether the record's pointer is the fixed destination's (destinationAgrees, with the repair when it is not), every runtime directory with its claim, every unfinished removal (interruptedRemovals), the outgoing selection, the promotion lock and the two settings records.
  • crw doctor is the host diagnosis: the host record reading, which runtime kind is selected, the Codex CLI version against the one the wiring was measured on, the App Server observed through the selected bridge, the pointer, the promotion lock, the settings records and every executable they name, the classification of each component, the residue on the destination (an abandoned staging a later install would reclaim, and the tombstone of an interrupted removal, which crw install remove finishes), the relay's own scope reading, and the six results. --temporary records the destination as a temporary one, so a proof taken there is never read as a claim about a host.

None of them takes a lock across the whole reading, so a host changing underneath is described in pieces.

How skill commands reach the relay

The skills run codex-session-relay ... as a shell command (relay usage), and that name is found on PATH. The installer does not manage PATH: it places the name in <destination>/current/bin/, and nothing puts that directory on anyone's PATH. So either put it there, ahead of any other copy, in the environment the Codex tasks run in, or run the relay by that absolute path. Check what a task actually reaches:

command -v codex-session-relay
readlink -f "$(command -v codex-session-relay)"
readlink -f <destination>/current                  # the runtime directory the pointer selects

The first answer has to lie inside the second. On a Go runtime (bin-<version>-<digest12>) it is that directory's bin/crw, which the name links to. On a host still on the Python fence release, before the cutover, the pointer selects an env-* directory and the answer is its bin/codex-session-relay, a console script of that environment.

A stale codex-session-relay earlier on PATH, such as a console script under ~/.local/bin whose shebang names a Python virtual environment in a development checkout, runs whatever that checkout holds against the same store, which can be code from before the fence release and so outside the cutover's fence. Neither crw install nor crw doctor reads PATH for it, so this reading is the one that finds it.

The one definition

internal/runtime/definition is the single compatibility definition OPS-1.1 requires: the two components, their console-script names (the compatibility names beside crw), version, licence and the tool that identifies the bridge. crw-dev ci contracts checks it against what it names outside itself: each licence is in the checkout and among the files the release archives carry, the identity tool is a tool the bridge's contract lists, and the links the installer places are the ones the release build makes. Release digests are not in it: they live in the release's SHA256SUMS and in the host record (decision 35). A host record still states definitionVersion 1 (decision 34).

Until todo 44 the definition was a file, scripts/crw_runtime/components.json, which kept the Python packages' trees and source digests for the Python installer to re-derive; it left with its last Python reader, and the Go package, which had carried the fields a Go install uses, became the only copy (decision 47). The ported bridge's upstream provenance, which the file recorded because the import brought source rather than history, is kept in its PROVENANCE.md beside its licence, the provenance narrative OPS-1.5 says is retained rather than replaced.

What the definition does not carry is as important. Installed locations, entry points, host names and measured points are host facts. OPS-3.2 makes a real record a private receipt, so they go to the host record outside this repository and this repository never commits one.

The host record

The host record is the other half of the definition and is never committed. It is recordVersion 1, which the Go and Python installers both read and write: Go install entries carry fields of their own beside the Python-era entries, which keep theirs. It holds, per component, one entry per install and the measured points, plus what is selected, the outgoing selection a rollback returns to, and the owned pointer with who placed it and when.

A Go install entry records location (the runtime's bin/), environment (the runtime directory), entryPoint (the component's compatibility name), binaryDigest and integrity (the binary's SHA-256), target, version, archiveDigest, reachedVia, and source, the repository commit and tree the binary was built from. An entry the Python installer wrote carries interpreter, interpreterPath and installMode instead of the binary fields. crw doctor reads both, and classifies only Go installs: a Python install keeps the Python installer's classification until the Python path is removed.

Under OPS-1.3 a point means the combination was exercised, so no amount of reading bytes produces one, and an install whose bytes match but whose combination nobody has run classifies unmeasured and is preserved rather than reused. The install's exercise is what produces a point (installing the runtime, step 4), and a connection, protocol or tool-call failure records no point and says why. A Go point records install, installDigest, codexCli, host and appServer, the dimensions it is matched on, beside date, measuredBy, method and archiveDigest. Points are appended, never replaced.

The appServer dimension is the bridge's whole get_capabilities payload, serialized and compared by exact string equality: the identity of the server a combination was exercised against is everything the round trip reported, not a version string someone chose to trust. So a change to that payload's shape changes the dimension, and a point measured before it does not cover a bridge built after it. The old point is not wrong and is not discarded; it remains evidence about the build it was taken against. The same property is why the payload carries no timestamp: a dimension that varied between two calls to the same build would never match itself.

Reading a record, and what happens when it cannot be read

Every record these commands read (the host record, the settings records, a claim, the Codex configuration, the hook file) is read at a narrow boundary that turns a failure into an answer rather than a crash. The answer is one of four states, decided by an ordered observation rather than by a convenience test:

State What was observed
ABSENT nothing exists at the path. The only state that may be read as a host with no history
PRESENT it was read. An existing record with nothing in it is PRESENT, not ABSENT
UNREADABLE something is there and its shape cannot be read: a directory or other non-regular file, a symlink whose target is established missing or looping, invalid UTF-8, unparseable JSON, or containers of the wrong type
ACCESS_ERROR nothing could be established: a permission or I/O failure reaching the path, a symlink whose target could not be resolved, or a parent directory that cannot be traversed

The distinction that matters most is the last row. Being unable to ask is not being told no, so a failure to establish existence is never reported as absence, and a permission problem is never reported as a malformed record. A refusal names what failed and where.

What this guarantees: the worst case for a record these commands read is a named refusal, not a crash. What it does not guarantee: that a record which could have been read is never refused. That direction is the safe one, and the refusal carries its reason.

  • "Nothing was written" is scoped to what can be guaranteed. Malformed input detected before the first mutating step refuses and the target file's bytes are unchanged. A read-back failure after a write reports the write instead of denying it (record_applied_unverified, config_applied_unverified) and exits non-zero, because reporting a landed write as a refusal that wrote nothing invites a retry over a file that now exists.
  • A read-only diagnosis reports rather than refuses. crw doctor names the failed reading in hostRecordState and hostRecordReading and continues with what it could still observe, and never reads an unreadable record as a clean host. Commands that write refuse outright.

One writer for the host record

Every change to the host record takes the record's lock, loads the record inside it, applies the caller's narrow delta and saves. A caller never hands back a record it loaded earlier: an install that loaded the record, spent minutes unpacking and exercising a runtime, and then saved what it had loaded would overwrite whatever another run committed in between, and holding a lock over that save would not help, because the staleness is already inside the value. So a caller says what it learned (this install, these points, this selection) and the merge happens against the record as it then stands. A failed install drops only the install entries keyed to the directory it created and leaves the selection exactly as found, because another run's successful promotion is not this run's to undo.

Installation ownership

crw doctor classifies each component of the runtime the pointer names from the OPS-2.1 signals: whether its entry point resolves into a recorded install, whether the binary's digest matches the recorded binaryDigest, whether a measured point covers the combination it runs under, and whether a registration conflicts with it or the owned pointer names something other than what the record selects. A Go install is an archive, not a checkout, so the checkout signals do not exist for it: a fork is bytes that differ from the recorded digest. A signal that cannot be read stops classification and is reported; it never counts as a signal that agreed. A runtime a host cannot launch, a bin/crw that is not an executable regular file or a compatibility name that does not resolve to it, is reported as such, because a digest says nothing about either.

The five OPS-2.2 classes are evaluated in their fixed order and the first match wins: conflict, fork, foreign, unmeasured, own. Only own is reused. An install whose bytes match and whose combination has no point classifies unmeasured, and that is the intended answer, not a gap to close by relaxing the rule: OPS-1.3 refuses to let a matching digest stand in for a run nobody performed. Exercising it, which is what an install does, is the way out.

Nothing outside a recorded path is ever overwritten, and no predecessor is removed by an install or an update.

MCP registration

The plugin package declares the server, and installing the package registers it. What the package cannot carry is where the runtime is and which execution policy the bridge runs under, so that goes into <CODEX_HOME>/crw-bridge-mcp.json, which crw install register-mcp --owner plugin writes and the launcher reads at every start (the native wiring). The record names the owner, the server name, the bridge executable (<destination>/current/bin/codex-thread-bridge unless --bridge-command says otherwise), its arguments (--bridge-arg, repeatable), and, from version 2, the execution policy. --dry-run reports what would be written and writes nothing.

Writing the record is not registering a server. On a host with no plugin installed the record is inert, and the command says so. Registered, recorded and a tool actually called stay three claims.

Who registers the server

One owner registers each surface. --owner plugin is the only owner crw install writes for: the user-owned registration, a [mcp_servers] table in config.toml, retired with the Python installer. The refusal still runs both ways, because a host can acquire the other half from elsewhere. A record is refused while the Codex configuration already starts this bridge, under this server's name or under any other, and the refusal names that entry: the plugin's declaration beside it would run a second bridge. The launcher stands down for a record that names another owner. A record or configuration that could not be read refuses rather than defaulting, because installing on an unanswered question is how the second bridge arrives. Every writer of the record takes the ownership lock beside it (crw-mcp-ownership) while it writes.

The execution policy the plugin bridge runs under

Codex starts a plugin-declared server with the App Server's own environment. Measured on Codex Desktop 0.154.0 for Linux, that is HOME LANG LOGNAME PATH SHELL USER and nothing else. A bridge started that way reads no execution policy: get_capabilities reports presence_only with no roles, and a child created through the Desktop tools is never asked whether it runs its role's pair. --execution-policy gives the plugin-owned record the one fact that closes that gap:

crw install register-mcp --owner plugin --execution-policy /path/to/execution-policy.json

The record becomes version 2 and gains executionPolicy, with two fields: path, the file as given (expanded and made absolute, not resolved, like the relay's own launch declaration), and digest, the SHA-256 of the bytes this run read. It never carries what the file says. Before anything is written, the file goes through the bridge's own parser in this binary, so a policy the bridge would refuse to start under is refused here instead, as execution_policy_unreadable. The output reports the mode and the declared role pairs, the same values get_capabilities discloses, and says which parser judged them: the installed runtime parses the file again every time it starts and decides for itself.

At every start the launcher reads the record and does one of two things. It refuses and exits 2, naming the record and the repair, when the file is missing, is not a regular file, cannot be read, or no longer hashes to digest, and also when its own environment already names a different policy file or digest. Otherwise it starts the bridge with CODEX_THREAD_BRIDGE_EXECUTION_POLICY set to path and CODEX_THREAD_BRIDGE_EXECUTION_POLICY_DIGEST set to digest, and the bridge refuses to start if the bytes it parses hash to anything else. Refusing to start is the visible failure. The declaration marks the server not required, so the session continues without the bridge's tools, and it never continues with a bridge that checks no role. A version-1 record names no policy and starts the bridge with the environment the launcher was given.

The policy is part of the registration's identity, so the only rerun that succeeds is an identical one, answered record_unchanged. Any other difference is refused as record_differs and nothing is written: another file, the same file with other contents, a rerun that drops the flag, or adding a policy to a version-1 record. The refusal names the repair: move the record aside by hand, then run crw install register-mcp again. Every edit to the policy file, including adding an exception, therefore has two consequences. The relay picks the edit up when its daemon restarts. The bridge record has to be moved aside and registered again, and a thread started in between has no bridge tools. The digest is what lets a changed file fail visibly instead of being enforced unregistered.

A created record is reported only after the policy file has been hashed again, following the write, and still matched. A mismatch immediately before the write writes nothing and answers record_policy_changed; a mismatch after it gets the same answer and the move-aside repair, and the record stays, because the launcher refuses its stale digest at every start. A removal by path cannot exclude a writer that does not take the ownership lock, such as an editor, so nothing here removes a record.

Codex starts the server once for each thread it loads. That was observed on Codex Desktop 0.154.0: one App Server process with a separate bridge child per thread, and the child's start time matching the thread's creation to the second. A record written now therefore takes effect for threads started afterwards, with no App Server restart, and a thread already running keeps the bridge it spawned. Neither is established on a host until a new thread's get_capabilities reports the digest the record names. A package replacement may or may not reach threads that are already running; what two replacements measured says what was seen, and updating safely gives the order and the checks.

Six results that never imply one another

crw doctor reports the six OPS-6.1 fields under checks.results, in the OPS-6.2 shape, each with its own evidence, command, acting process and measurement time. A field with no timed observation behind it reports unknown; no time is ever invented or copied from another field.

Field Established by Never established by
installed Every component classifies own The runtime being present
mcpExposed Tool names observed in a live Codex session A record or a declaration
connected The relay's doctor reporting actorReachability.socketConnect as ok A socket file on disk
deliveryAccepted An attempt that recorded a returned turn id A dispatch or an absent error
verificationComplete Every OPS-6.4 condition at once A completed turn or a green check
alwaysActive A supervised runtime surviving a host restart, observed after one Any of the five above, or a registered unit no restart has tested

A live session is the only thing that can list the tools a Codex session exposes, and the diagnosis's own session with the bridge is not Codex's, so mcpExposed stays not_verified from crw doctor; read it in a task. deliveryAccepted needs work created and delivered, which a diagnosis never does, so it is not_applicable unless a trial produced it. settingsPreserved is reported beside the six and never merged into them: it answers that the diagnosis wrote nothing.

One shared service and one store

OPS-3.1 puts one relay service and one durable store behind an entire operating scope, which is one host, one OS user and one App Server. A second repository or a second project installs into that same scope and reuses the same service and the same store; nothing here creates a daemon or a store per project, per repository or per parent. crw doctor reports the resolved scope, the store and the service by calling the relay's own doctor rather than by rediscovering any of it, and reports other stores beside the resolved one without adopting any of them. Equality of path strings is not proof under OPS-3.4; proof is doctor from each participating process reporting the same state directory together with assignment-find --issue returning the expected relationship.

The relay service unit

A relay service started by hand does not come back after a host restart: service start leaves a record that reads "recorded before a different boot", and nothing starts it again. crw install register-service --socket <app-server-socket> [--state <dir>] registers the one systemd user unit that does. It is a command of its own, like hook and register-mcp, because an install or an update never starts a daemon and gains no side effect here. Starting at boot, before anyone logs in, needs the user manager to start at boot, which a host with lingering off does only at login.

The unit is written to ${XDG_CONFIG_HOME:-~/.config}/systemd/user/crw-relay.service (--unit-dir and --unit-name change either) and enabled with systemctl --user enable. It is Type=oneshot with RemainAfterExit=yes, wanted by default.target; its ExecStart is the relay's own service start and its ExecStop is service stop. Both name <destination>/current/bin/codex-session-relay with the state directory and socket resolved at registration, so a pointer move never rewrites the unit, and the relay's own checks (intent, launch policy, scope) apply to the start unchanged. The unit has no restart policy: after the one start it acts only when an operator acts on it, and a refused start (service_disabled, already_running) is a failed unit with the relay's answer in the journal, never a retry that could start the relay in the middle of an update. It unsets CODEX_THREAD_BRIDGE_EXECUTION_POLICY and CODEX_SESSION_RELAY_SCOPE_DIR, which the user manager would otherwise pass in. --scope-dir <dir> registers an isolated target instead (a temporary state directory and socket): the unit sets that scope variable and the start carries --allow-isolated-scope.

One owner registers this surface, and the command asks the manager as well as the file. It refuses, writing nothing:

Outcome When
unit_foreign something at the path is not the installer's unit (it lacks X-CRW-Owner=crw-install), or is not a regular file
unit_differs the installer's unit says something else; run --remove, then register again
unit_name_taken the manager loads the name from another file, or it is masked
unit_modified a <name>.d directory, drop-ins (also read once the unit is enabled and loaded, when the unit stays enabled and the exit status is 3), or a manager definition older than the disk (NeedDaemonReload), which disable would re-read
unit_second_owner another unit in the directory runs the relay's service start or run
unit_dir_in_runtime the unit directory lies inside the installer's destination, where removing a runtime would delete it
unit_unreadable no runtime is installed, systemd could not be asked, or the relay cannot read its launch declaration

It also reads what the boot start will find, through the pointer's relay in the unit's environment (service status), and reports serviceEnabled, launchPolicySource and scopeAuthority. A disabled service intent, an undeclared execution policy and an XDG_STATE_HOME the unit does not carry are warnings: the boot start would refuse service_disabled, withhold role-bound deliveries, or write its host record elsewhere. It never changes the intent. Exit status 3 means a change may have landed and a later step did not or could not be confirmed (written and not enabled: run it again; an enable or disable that failed, possibly after changing some links: read systemctl --user is-enabled; deleted and daemon-reload failed: run that); 1 is a refusal with nothing changed, 2 a usage error. --remove disables and deletes only a unit it wrote, only when the file is exactly what the command writes (disable follows Also= into other units), the manager resolves the name to that file and it is not running; it never stops the relay.

The maintenance order is the relay's own: service stop (systemctl --user stop crw-relay.service runs the same stop while the unit is active, but a unit that never started has none to run), crw install update, which refuses while a daemon runs, then systemctl --user restart crw-relay.service or service start. After a hand service stop the unit still reads active (exited), so systemctl --user start does nothing and restart is the command. Measured with the unit's own commands: a start with no App Server socket present succeeds and its worker stays up, so a start before the App Server is listening is not refused; a start not ready within the relay's 20 seconds on a heavily loaded host is ended by the relay and leaves a partial store (write-gate.lock without a database) that the next start refuses as store_owned_by_other, which is the relay's behaviour and is in the refactor backlog.

crw doctor keeps alwaysActive at not_verified and names the default unit file it found or did not (a unit under another name or directory is not read). A unit file is a registration; surviving a host restart is a measurement of one.

The completion hook

The plugin package declares the Stop hook that catches a managed turn ending without the records a completion needs, and crw install hook --owner plugin writes the settings it reads, <CODEX_HOME>/crw-completion-hook.json:

Setting Written as
owner, configVersion, event plugin, 1, Stop
relayExecutable <destination>/current/bin/codex-session-relay, or --relay-command
mode observe (--mode), which classifies and records and never holds a turn
timeoutSeconds the adapter's own budget (--guard-timeout, default 5), under the registered timeout (--timeout, default 10, at most 10)
journalRoot, journalPolicy <CODEX_HOME>/crw-completion-hook/journal (--journal-root) and every_invocation
markerRoot, dbPath, socketPath the relay's own resolution, or --marker-root, --db-path and --socket
isolationAssertedBy required with --mode hold, and nothing else

The decision is not made in the hook. The hook contract fixes the rules and the relay's guard implements them (in the hook's own process, or in the store's owner, which crw hook asks over control.sock); crw hook is the piece between the host and that guard, and it exits 0 on every path, because exit 2 is the host's blocking code.

The settings are written before anything could read them, and every precondition is checked before any write: the budget against the registered timeout, a hook file that already registers this adapter for Stop (the user-owned registration, refused by name, because two registrations run twice on every Stop), and settings already there that this command cannot act on. Settings that already say something else are refused rather than overwritten, because they carry the mode: they answer config_differs, naming the differing fields; move the document aside by hand to write these flags' settings instead (one Stop settings document).

The one exception is a document written before decision 66. Those settings also recorded adapterInterpreter /usr/bin/env and adapterEntryPoint <destination>/current/bin/crw-completion-hook for the Python launchers the plugin package no longer carries. Nothing reads either key now: the Go hook, the doctor and every move of the pointer accept a document that still carries them, so a host keeps working on its current settings. A document that differs from the one this command would write only by those two keys and installedBy is rewritten without them, answering config_replaced with the keys it dropped under retiredFields (a --dry-run answers config_would_create and says so). On the relay host, run crw install hook --owner plugin with the flags the settings were written with, after the runtime carrying this change is installed, to rewrite the document once. Any other difference is still config_differs. The run registers nothing and leaves the hook file untouched. Written, registered and observed to have fired stay three separate claims.

--socket is recorded as socketPath. It is what lets the guard tell that the state directory it resolved holds another App Server's store: a store records the socket it serves, and a Stop hook that inherits CODEX_SESSION_RELAY_STATE from a second installation would otherwise read that store, find no relationship for its assignment, and hold a child that has finished. Configure it wherever more than one installation shares a machine.

The marker root follows the relay's own resolution, including CODEX_SESSION_RELAY_MARKER_ROOT. A default that skipped it would be a disagreement: the coordinator would publish its intents under one tree while this hook looked under another, and every managed turn would read as unmanaged.

Trust, and the command that is fixed when a turn starts

Codex runs a declared hook only once it has been trusted, and nothing in this repository grants trust. Trust is recorded against the declaration's content, so a plugin update that changes the hook command's text needs the hook trusted again, and Codex asks for it; until then nothing fires. An update of the runtime changes no declaration, because the command names the pointer, and so needs no new trust (plugin packaging).

A hook command is fixed when a turn starts, and every Stop of that turn reuses it; the next turn resolves the declaration installed then (the turn-command cache). A turn that started while the package still declared the Python bootstrap goes on running it: ${PLUGIN_ROOT}/wiring/crw_stop_hook.py first, and <CODEX_HOME>/crw-stop-hook.py once that version directory is gone. crw install places no such launcher and leaves an existing copy exactly as it is; the Python installer placed it. Either launcher reads the same settings and runs the adapter they name, which is the Go hook once the settings are Go-era. The packaged launchers left the package in todo 43, after the cutover commit: the version directory such a turn names is the pre-native one, which the install that brought the native wiring removed, so the copy is what answers it. The copy and the Python runtime stay while any live or resumable task can still run such a command, and the operator removes them once the turns that could still run one have ended (retention).

Who registers the hook

The plugin package declares this hook, and a host holding a second registration of it would run both on every Stop: each leaves a row, and only the one that claims the event's accepted record asks the guard (one accepted record per Stop event). crw install hook writes for the plugin owner only; the user-owned registration, an entry in <CODEX_HOME>/hooks.json, retired with the Python installer, and one already there is refused by name. The hook repeats the check at run time, because installing the plugin is not a command this repository runs: crw hook --plugin-launch stands down in silence unless the settings name the plugin as owner.

Only a verdict that agrees with itself is acted on. A verdict whose own decision releases while its hook_output holds did not come from the guard, and any disagreement reads as guard_verdict_incomplete and releases. Failures keep their own names: a runtime that could not be run carries its errno, and an exit of 2 carrying the relay's own error record is the relay declining a request it understood, while an exit of 2 carrying nothing is its argument parser refusing before any command ran. Every one of these releases the turn and is recorded.

One accepted record per Stop event

A turn can end more than once. When any Stop hook holds, the host appends a continuation to the same turn and fires Stop again; when a message was waiting, it appends that and does the same. Each of those is its own Stop event, and each one gets its own decision. Two registrations answering one Stop, or one Stop delivered twice, are a different thing: one event handled twice. The adapter keeps exactly one accepted record per event, asks the guard once for it, and still leaves a row for every invocation, so the two cases stay apart instead of being counted as one number per turn.

What identifies an event

The host hands a Stop hook nine fields and no per-invocation identifier. None of them separates two events of one turn: turn_id is kept across a continuation chain, stop_hook_active is false on a turn's first Stop and true on every later one, and last_assistant_message can repeat word for word. What differs is the transcript. Every sampling that ends in a Stop leaves one final answer, and before running Stop hooks the host records it in the file named by transcript_path with the turn id, the thread id and an item id of its own. So an event is

(session_id, turn_id, stop_hook_active, answer item id)

where the answer item is the one Stop of the turn, as the transcript shows it, that reported the payload's text under the payload's stop_hook_active. It is established only when the latest sampling's answer is recorded, exactly one Stop of the turn matches, and the matching answer's thread is the delivered session. The key is a SHA-256 over the four values, so no host value becomes a path component and nothing is minted per invocation.

The transcript is read backwards from its end to the turn's start, bounded at 64 MiB and 0.75 seconds, inside the margin the launcher keeps over the guard budget. A missing or unreadable transcript, a scan that hits either bound, an unfinished last line and any failed condition leave the identity unestablished, with the reason on the row. An unestablished invocation is asked about exactly as before and is never deduplicated.

Accepted records and attempt rows

A claim is two create-once files. The first is the host's, <CODEX_HOME>/crw-completion-hook/stop-events/<key>.json: every registration the host starts for one Stop inherits that Codex home, while the settings each one reads, and so its journal root, may differ, so this is the file two registrations of one host always meet. Only the invocation that creates it owns the event, and only the owner asks the guard. The owner creates the accepted record, <journalRoot>/accepted/<key>.json; asks the guard; writes its row; and then writes accepted/<key>.outcome.json naming the session, the turn, the outcome and the row. An invocation that finds either file already there asks nothing, prints nothing and writes a row whose adapterOutcome is duplicate_invocation. When the host's file can be neither created nor found, nobody can own the event, so nobody asks: the invocation writes an arbitration_failed row and releases the Stop, the adapter's ordinary failure direction.

Rows are <journalRoot>/<YYYYMMDD>/<32 hex>.json, one JSON object on one line with sorted keys, and every count of them counts invocations. A row carries sessionId, turnId, adapterOutcome, eventKey, eventIdentity, acceptance (accepted, duplicate, unestablished, unclaimable, claim_failed or unarbitrated), acceptedAs and guardInvoked. faults_only leaves out the rows of answered and duplicate invocations of identified events; no_journal keeps no rows at all.

The outcome record says what the adapter answered, not what the host received: it is written before the answer is printed, exactly as the row is. A claimant killed between its claim and its outcome leaves a claim with no outcome, and the event still has its owner, so a later delivery of it is a duplicate and the Stop was released.

Reading it back

crw-dev stop-events --journal-root <root>, in the repository's development binary, reads the rows, the accepted records and the host ledgers the claims name, and answers one verdict (the same reading as scripts/stop_events.py, deleted in todo 44, gave). FALSE (exit 1) means an event was accepted more than once; UNREADABLE (exit 3) means the reading cannot vouch for what it read, and names why; TRUE (exit 0) otherwise. --session and --turn choose the events of one turn, --since and --until a window in the records' own YYYY-MM-DDTHH:MM:SSZ format, --journal-root repeats for every root the host's registrations write to, and --codex-home adds a host whose ledger no claim names yet. The invocations it cannot judge, whose identity was not established or that had no owner, are counted by reason and never read as answered.

When the window contains readable version-2 rows with no event key, the optional excludedInvocations list names each row, timestamp, session and turn, and its unestablished:<reason> or no_event:<outcome> reason. evidence holds the recorded acceptance, event key, accepted path, adapter outcome, guard invocation/decision, hold and event identity. excludedFrom: per_event_acceptance_count means only that the row cannot join an event's acceptance count; it does not remove uncertainty from the window. preventsTrue is true except for the existing native pre-scan unreachable exemption. An unestablished guard call still makes the verdict UNREADABLE, even if it wrote no claim. Keyed unclaimable or failed claims remain in unjudgedInvocations; malformed rows remain in rowsUnreadable.

For no_event:stdin_unreadable, evidence.stdinRead adds the recorded detail and elapsedMs and a case naming what the row establishes. A row written since CRW-504 carries its own account of the failed read in stdinRead (below) and its case is the cause that account names; a row without it is read as before.

Case What it establishes
input_late The 100 ms input allocation ran out while the hook waited for the host to supply the payload.
work_ended The invocation's own work context ended while the hook waited (a settings budget shorter than the allocation, or a cancelled run).
read_error The stdin descriptor, or the reader standing in for it, failed; error holds the Go error text.
invalid_utf8 The adapter read bytes but could not decode them as UTF-8. A row without stdinRead is recognised by its detail.
read_failed_cause_unrecorded A row without stdinRead: reading stdin failed and the adapter discarded the underlying cause.
detail_unrecognized A row without stdinRead whose detail matches no known branch; retain it without assigning a cause.

stdinRead is an optional object on stdin_unreadable rows, added beside the existing keys the way runtime was; rows written before it stay valid. It has five keys: cause (input_late, work_ended, read_error or invalid_utf8), error (the Go error text of the failed read, an OS or context message about the stdin descriptor with no payload content, and for a recovered panic only its Go type), bytesRead (the bytes the completed reads reported before the failure was recorded, which counts what arrived and not what the host meant to write, and is a lower bound when the reader was still running), and waitStartedMs and waitEndedMs (when the wait began and ended, in milliseconds from process entry, the origin of elapsedMs). The judge reports error, bytesRead, waitStartedMs and waitEndedMs beside the case. A row whose stdinRead is not that closed shape, whose cause disagrees with its detail, or that carries it on another outcome is unreadable.

These diagnostics change no count, reason, exit code or verdict. Other outcomes omit stdinRead; malformed rows never receive it. A row without stdinRead cannot distinguish a host that supplies bytes or EOF late from a descriptor error or the work deadline. In particular, elapsed time near 100 ms does not prove a timeout, a restart, or which session was involved. Correlate such a row with host stdin write/EOF/error logs and rollout evidence; when those are absent, report the cause as undetermined rather than attaching the nearest turn by timestamp. A row with stdinRead says which of the four happened and how many bytes had arrived, but not why the host was late; the session and turn are unknown before the payload parses and are not recorded.

Host input that remains incomplete when the adapter must wait beyond its 100 ms allocation is the intended late-input release case in decisions 24 and 32. Already-readable input is still taken. There is no separate stdin byte cap. Empty input whose writer closes normally reaches JSON parsing and is stdin_not_json, not stdin_unreadable: a clean end of input is not a read failure and never carries stdinRead.

Limits

The guarantee holds among the registrations of one Codex home, which is every registration one host starts for a Stop. The claim files are never removed, like the rows. That the host records the answer before running Stop hooks was observed in every isolated run and is consistent with every record in the live journal, but it is not a documented host contract. A turn whose Stops repeat the same text trades deduplication for safety: those Stops are asked about by every registration, and a reading of a window that holds them is UNREADABLE.

Registration is not firing

Written settings, a declared hook, a trusted hook and a hook that fired are four different facts, and the only evidence of the last one is a record of this turn. End a turn in an ordinary task started after the change, take its session and turn ids, and look for them in the journal:

journal="${CODEX_HOME:-$HOME/.codex}/crw-completion-hook/journal"
grep -l '"sessionId": "<session>"' "$journal"/*/*.json | xargs -r grep -l '"turnId": "<turn>"'

A row naming that session and turn is a callback this procedure can attribute; its adapterOutcome says what the adapter did. No row is not a smaller number of callbacks: the turn did not reach this hook's journal, and the cause is one of these, read in order:

Cause How to tell
The hook is not trusted Codex has not been asked to trust this declaration since the command text last changed; nothing fires until it is
The pointer names no runtime that reads --plugin-launch crw doctor: runtime.state, runtime.kind (anything but go-binary has no bin/crw), and a Go build older than decision 26 (the native wiring)
No usable settings crw doctor: settings.crw-completion-hook.json.state; the hook releases in silence without settings, with settings it cannot read, and with settings another owner holds
Journalling is off journalPolicy is no_journal, or faults_only and nothing faulted
A different journal the settings name another journalRoot than the one you read

A subagent's turn is no signal: on the measured host subagent turns recorded no Stop at all. The daemon is not on this path. The guard reads the marker and a read-only database, so a stopped daemon is not observable from a Stop and is never inferred from one. runtime_install.py hook-status, the Python installer's cell-by-cell reading of the same question, has no crw counterpart (the Python fence installer); its Go port was deleted unused in wave R1 (decision 57 in the port decisions).

The composed acceptance run

Installing, updating and hooking each have their own rules above. What none of them states is the sequence a host actually lives through, with the state that has to survive it put there before the first install and read again after the last refusal. The Go installer's tests run that sequence against temporary homes: install, the same install again, an update that fails at a step and puts everything back, rollback and remove (internal/runtime/install, lifecycle_test.go, decisions_test.go and restore_test.go), and the native Stop command and the launchers through the pointer (wiring_test.go; the retired Python launchers from the pre-native testdata).

Seven questions, seven readings

The composed run asks seven questions and forbids one reading standing in for another. A reading that could not be made is unreadable, never false, and never the value of the reading beside it.

Question Answered by Read from Never established by it
Skill link crw-dev skills link --check, on a linked host its report that a linked skill is loaded or trusted by a host. A plugin host has no links; codex plugin list is its reading
Runtime reach crw doctor runtime.state, runtime.kind, runtime.agrees and each components.<name>.class that the runtime works for a task
MCP tool exposure a fresh Codex task the bridge tools it lists, and get_capabilities anything about a task started before the change
App Server connection crw doctor checks.results.connected that a socket that accepted a connection will accept delivery
Real hook callback the journal the rows naming the turn you ended that a firing was judged correctly
Model and permission preservation a byte comparison of config.toml the file before the run and after the crw commands that anything a Codex process writes later was preserved
Delivery acceptance a trial checks.results.deliveryAccepted from a run that created and delivered work that an accepted delivery was acted on

Running the combination against a real host

A real combination is an operator action, not a check. It needs a home, a Codex home, a host record and a state directory that are yours to change, and it establishes nothing until it is recorded. The destination is $HOME/.local/share/crw-runtime of the HOME the commands run under, and the live half reaches this runtime only through the plugin's declared commands, so run the block under the HOME the Codex process runs with. The block names the host record and the Codex home on every command, and the state directory wherever one is read, because none of them derives from another: the Codex home is where the settings the plugin reads are written, and the host record defaults to $XDG_STATE_HOME/codex-relay-workflow/host-record.json, whatever the Codex home. <codex-home> is the Codex home that process reads. The Go build serves a store in that HOME's default relay state directory, which the live-state guard refuses only under test isolation, so only a host that still runs the Python runtime waits for the cutover.

# Substitute every <...> below before running any of it. They are placeholders, not literals, and
# an unsubstituted one is a shell redirection rather than a value.

# A NEW receipt directory, so no earlier run's files are read as this one's. mkdir without -p is
# the check, and it ends the procedure rather than running the rest into a directory it refused.
mkdir <receipt> || exit 1

# Preservation is a comparison. No crw command below writes config.toml, so the whole file is
# compared rather than keys guessed out of it. A fresh Codex home has none; record the absence.
if [ -f <codex-home>/config.toml ]; then
    cp <codex-home>/config.toml <receipt>/config.before.toml
else
    printf 'no configuration existed before this run\n' > <receipt>/config.before.absent
fi

# Nothing puts crw on PATH. The install runs the copy unpacked from the release archive into
# <scratch> (see "Installing the runtime" above); every later command runs the one the pointer
# then selects.
# The exit status belongs IN the receipt: a receipt that kept the result and lost the status
# cannot say whether the install refused, or that a 3 means the change landed.
<scratch>/crw install install --release <tag> --record <record> \
    --codex-home <codex-home> --state <state> --socket <socket> > <receipt>/install.json
printf 'install exit=%s\n' "$?" > <receipt>/install.exit

crw="$HOME/.local/share/crw-runtime/current/bin/crw"

"$crw" install register-mcp --owner plugin --record <record> \
    --codex-home <codex-home> --execution-policy <policy-file> > <receipt>/register-mcp.json
printf 'register-mcp exit=%s\n' "$?" > <receipt>/register-mcp.exit

"$crw" install hook --owner plugin --record <record> \
    --codex-home <codex-home> --socket <socket> > <receipt>/hook.json
printf 'hook exit=%s\n' "$?" > <receipt>/hook.exit

"$crw" doctor --record <record> --codex-home <codex-home> \
    --state <state> --socket <socket> > <receipt>/doctor.json
printf 'doctor exit=%s\n' "$?" > <receipt>/doctor.exit

# The other half of the preservation reading, taken before anything else can write the file.
if [ -f <codex-home>/config.toml ]; then
    cp <codex-home>/config.toml <receipt>/config.after.toml
    cmp <receipt>/config.before.toml <receipt>/config.after.toml > <receipt>/config.cmp 2>&1
    printf 'cmp exit=%s\n' "$?" >> <receipt>/config.cmp
else
    printf 'no configuration existed after the crw commands\n' > <receipt>/config.after.absent
fi

#   ... then, in Codex: trust the hook if it asks, start a fresh task, read the bridge tools and
#   get_capabilities there, end a real turn, and note that turn's session and turn ids. Only then:

journal=<codex-home>/crw-completion-hook/journal
grep -l '"sessionId": "<session>"' "$journal"/*/*.json | xargs -r grep -l '"turnId": "<turn>"' \
    > <receipt>/rows-for-this-turn.txt
printf 'rows grep exit=%s\n' "$?" >> <receipt>/rows-for-this-turn.txt

# From a checkout, the exactly-once reading for that turn:
go run -tags dev ./cmd/crw-dev stop-events --journal-root "$journal" --codex-home <codex-home> \
    --session <session> --turn <turn> > <receipt>/stop-events.json
printf 'stop-events exit=%s\n' "$?" > <receipt>/stop-events.exit

What this block is, and what it is not. It installs, writes both records, takes the readings and reads the results back. It does not re-run the install, it does not present a second archive, and it does not fail an update; this page will not tell an operator to break a runtime their host is using in order to watch it come back. The tests above exercise those stages against temporary homes.

cmp exiting 0 is preservation. A difference means something wrote the file between the two copies, and the receipt holds both sides for reading which keys moved. rows-for-this-turn.txt naming no file is a turn that did not reach the hook (registration is not firing), not a smaller count. More than one row for the turn is not a duplicate by itself: a turn whose Stop was held ends again, and that is a second event; stop-events.json separates the two.

Three of the seven are readings of something live, and the block supplies none of it: an App Server accepting connections at <socket>, a session that actually listed the bridge tools, and a relay that can carry an assignment to a returned turn id. crw doctor reports deliveryAccepted as not_applicable because it creates no work; a live trial (live-trial.md) is how a host gets that reading. If the live half is absent, the honest receipt records the absence for that row and says the rest. It never carries a row forward as though the question had been put.

A real run records the exact release it installed, the host it ran on, the destination kind, and the answer to each of the seven with the command that produced it and the time it was produced. Those receipts are host facts: they belong in the private record outside this repository, not in a commit. This page is the procedure and the shape. It is not a record that anybody ran it.

The Python fence installer

scripts/runtime_install.py installed the Python runtime, the fence release that step 0 of the cutover deployed, and it was the rollback path until the cutover committed (todo 43). Todo 44 removed it with the rest of the Python execution path. It shared the destination, the owned pointer, the host record, the promotion lock and both settings files with crw install; what it installed was a Python virtual environment, <destination>/env-1-<digest>, built from the checkout's packages, and it took --dest, which is why crw install still refuses a host record whose pointer names another link. Its subcommands were install, diagnose (with --trial), register-mcp, hook (which also placed the fallback launcher <CODEX_HOME>/crw-stop-hook.py), hook-status and verify-definition; crw install and crw doctor carry what a Go install needs of them.

Two properties of a Python runtime outlive this installer, and the cutover's retention rule rests on them. pip writes an absolute shebang into every console script, so a process started through current reports and keeps its concrete env-* directory after the pointer moves; and a Stop command fixed before the native wiring names python3 and a .py launcher. Both kept env-* directories and the <CODEX_HOME>/crw-stop-hook.py copy in place until nothing could still resolve to them. The packaged launchers did not wait for that: such a command names the pre-native version directory, not whatever the package ships now, so todo 43 retired them from the package after the cutover commit.

Trial mode

runtime_install.py diagnose --trial was the only command that filled deliveryAccepted itself: it registered one relationship, emitted and delivered once, and recorded the returned turn id. It has no crw counterpart and left with the installer; a live trial is a different thing with a similar name.

The full reference this page carried for the Python installer, its design record and its acceptance procedure, is this page at the parent of the commit that rewrote it for crw, the oldest one that names todo 41 (git log --format=%H --grep='(todo 41)' -- docs/runtime-install.md | tail -1).

What none of this establishes

Running these commands against a temporary destination proves what they did there. It is not evidence about a host's real Codex home, its installed runtime, its bridge record or its operational database. installed, mcpExposed, connected, deliveryAccepted, verificationComplete and alwaysActive are six separate facts under OPS-6.1, and none of them is read from another.

Written settings are not a fired hook, and a fired hook is not a delivered hold. That a settings file is there says nothing about the host having run the hook, about the runtime it names being able to answer, or about any turn having been judged. Those claims need the host's own evidence.

A successful update is not one of them either. That the pointer moved, that the gate found the daemon stopped and no attempt open, and that the store's schema was compatible are readings taken at one moment, about one destination. They say a swap was permitted and performed; they do not say the new runtime works, and the point that would say so is measured before the swap rather than after it. Nor does a refused update establish that a store is healthy: the gate reads whether it is safe to replace a runtime, and reads nothing about whether the data in the store is correct.