POLICY.md owns repository rules and the operations contract owns the operational ones. This page describes how the runtime behind the MCP bridge, the session relay and the completion hook is installed, updated, rolled back and diagnosed, and it is written against that contract's clause numbers so a reader can check a claim against the rule it came from.
The runtime is one Go binary, crw, shipped in a release archive. The plugin package declares the
server and the Stop hook and reaches the binary through the installer's pointer
(plugin packaging); the package carries no runtime of its own.
Updating is the half that can lose something. A first install has nothing to destroy; a second one is standing on a runtime somebody is using and a database nobody can rebuild, so most of what follows is about what is read before anything moves and what is put back when it does not. Updating an installation is where that lives.
| Command | What it does | Contract |
|---|---|---|
crw install install, crw install update |
Verify a release archive, install it as a new runtime directory, exercise it and move the owned pointer to it | OPS-2.4 |
crw install rollback [<dir>] |
Point the owned pointer back at the runtime the last promotion replaced, or at a runtime directory the host record lists | OPS-2.4 |
crw install remove <dir> |
Delete one runtime directory nothing selects, points at or runs out of | OPS-2.4 |
crw install register-mcp [--owner plugin] |
Write the bridge record the plugin's declared server reads | OPS-2.2 |
crw install hook [--owner plugin] |
Write the Stop settings the plugin's declared hook reads | OPS-6.3 |
crw install register-service [--remove] |
Write and enable the one systemd user unit that starts the relay service when the user manager starts; with --remove, disable and delete it |
OPS-4.1, OPS-6.1 |
crw install status, crw doctor |
Read the installation, classify it and report the six check results; write nothing (the relay readings behind them are the relay's doctor without --probe-write, which creates no file of its own and never opens its write gate; what it still does is named under Installing the runtime) |
OPS-2.1, OPS-2.2, OPS-6.1 |
crw-dev skills link --check or --apply |
Skill links into Codex, from a checkout | OPS-2.3 |
Every crw install and crw doctor command prints one JSON document. Runtime installation is never
folded into the skill links: crw-dev skills link belongs to the repository's development binary
because it links a checkout, and a release archive has none. It stays idempotent, it refuses to
replace an existing directory or a foreign link, and its LINKED, MISSING and CONFLICT words
mean the same thing wherever this page uses them. A plugin installation has no skill links at all.
The Python installer, scripts/runtime_install.py, installed the Python fence release and was the
development and rollback path until todo 44 removed the Python execution path;
the Python fence installer is the one section of this page about it.
Moving a host from the Python runtime to this one was the cutover, not an install
alone: it moved the store's ownership. It ran at todos 42 and 43; the Python runtime left the
repository in todo 44 (decision 48) and the takeover controller in refactor R1 (decision 54), so no
order between crw install install, which moves the pointer, and a takeover is left to settle.
A release archive is crw_<version>_<os>_<arch>.tar.gz, published with a SHA256SUMS beside it
(releases). It holds crw, the two compatibility names
codex-session-relay and codex-thread-bridge as links to it, and the licences. crw dispatches
on the name it was started under, so each name is the component it names. The completion hook is
crw hook; its former crw-completion-hook link is retired (decision 66), and a runtime
installed before that still carries one.
An installation of it is two things under the destination:
| Path | What it is |
|---|---|
<destination>/bin-<version>-<digest12>/bin/ |
The runtime: crw and the two links, where <digest12> is the start of the archive's SHA-256 |
<destination>/current |
The owned pointer: a directory symlink naming the selected runtime directory |
The destination is ~/.local/share/crw-runtime and nothing else. Both of the plugin's declared
commands name $HOME/.local/share/crw-runtime/current/bin/ and nothing else, because HOME is the
one variable a hook and an MCP server both receive
(how hooks and MCP servers load). So
crw install has no --dest, and neither it nor the wiring honours XDG_DATA_HOME: every
crw install command acts on <home>/.local/share/crw-runtime, where <home> is HOME, or this
user's passwd entry when HOME is not set. A host record whose pointer names another link is
refused by every command before it acts, naming the repair, and crw install status reports that
reading as destinationAgrees (decisions 11 and 38). For a temporary or
isolated installation, run the commands under another HOME, and move with it everything HOME
does not decide, as the isolated-home integration test (internal/runtime/integration) does:
CODEX_HOME and XDG_STATE_HOME inside the same tree, CODEX_SESSION_RELAY_SCOPE_DIR set to a
directory there, and CODEX_SESSION_RELAY_STATE and CODEX_SESSION_RELAY_MARKER_ROOT unset or
pointed there too. Each is read on its own. A CODEX_HOME left naming another Codex home has the
install judge that home's registrations and crw install hook write its Stop settings naming the
temporary pointer; an XDG_STATE_HOME left naming
another state home puts the temporary host record there, or is refused where the record there names
another pointer; and without CODEX_SESSION_RELAY_SCOPE_DIR the relay finds its scope registry from
this user's passwd entry, never from HOME. crw doctor still takes --dest, to read another
destination, never to install into one.
A path the commands cannot use as given is refused rather than guessed at. HOME has to be
absolute, hold no .. and not start with exactly two slashes (//home/..., which pathlib keeps as
spelled and a lexical join folds to one; three or more fold to one in both). A relative
XDG_STATE_HOME is a usage error (exit 2) to every crw install command not given --record: read
against the working directory it would put the host record where nothing else looks. The doctor does
not read against the working directory either, but it reports rather than refuses: crw doctor
answers hostRecordState ACCESS_ERROR and exits 0. A path from HOME, CODEX_HOME, XDG_STATE_HOME or a path flag that holds a byte
that is not UTF-8 is a usage error naming where it came from, because a record or settings document
written with a replacement character names a file that does not exist. The execution policy path is
the one exception, recorded as os.fsdecode spells it
(the execution policy).
The host record, ${XDG_STATE_HOME:-~/.local/state}/codex-relay-workflow/host-record.json, says
which runtime is selected and records every install, its measured points and who placed the
pointer (the host record). The two settings records the plugin's commands read
sit in the Codex home: crw-bridge-mcp.json for the server and crw-completion-hook.json for the
Stop hook.
Nothing here manages PATH. The declared server and hook name the pointer by absolute path, but a
skill command that runs codex-session-relay finds whatever PATH finds
(how skill commands reach the relay).
crw install needs a crw to run it. The archive carries one, so unpack it anywhere temporary and
run that copy against the archive itself:
tar -xzf crw_<version>_<os>_<arch>.tar.gz -C <scratch>
<scratch>/crw install install --from crw_<version>_<os>_<arch>.tar.gz \
--socket <app-server-socket> # SHA256SUMS beside the archive, or --sums <file>--release <tag> fetches the archive for this host's target and its SHA256SUMS from that GitHub
release instead of --from. Either way the archive has to be named for this host's operating system
and architecture, be listed exactly once in SHA256SUMS and hash to the listed digest, and nothing
under the destination or in the host record is created before all three hold.
--backup-state-to <dir> (on install, update and rollback) is the acknowledgement that the swap brings the additive DAG zone, or ordinary
indexes on tables the store already holds, to a store that lacks them: the command copies the whole relay state directory to <dir> before
promoting, and without it that swap refuses and names the flag (the route).
What follows is one run, in this order, and the result lists the steps it took:
- Claim the runtime directory with an exclusive
mkdirand a claim file (the claim a run leaves behind). - Unpack the archive into it and read the binary's digest.
- Record the install entries.
- Exercise the candidate through its own concrete executables, never through
current, which still names the predecessor: the relay'sdoctormust reportactorReachability.socketConnectasok, and the bridge must answer an MCP session that lists its tools and callsget_capabilities. Both run against the App Server socket--socketnames. The bridge falls back to<CODEX_HOME>/app-server-control/app-server-control.sockwithout it, but the relay has no default socket: itsdoctoranswerssocketConnectasnot configured, so a run without--socketfails atexercise the candidate(exit 1) even with an App Server listening at that path. A run that cannot exercise the candidate records no point and promotes nothing. - Under the host-wide promotion lock: read whether it is safe to swap, establish that the pointer is this command's, refuse a second owner of the bridge or the Stop hook, and refuse Stop settings that name through the pointer something a Go runtime does not serve (one Stop settings document); nothing rewrites them.
- Commit the selection, then replace the pointer, and read the pointer back.
- Settle the claim.
OPS-2.4 sequences an update as measure, install, measure again, and step 4 is the measurement that produces the point; promoting before it would select a runtime that unpacks cleanly and fails the moment it is used. A failure at any step up to the promotion leaves the previous runtime selected and the pointer where it was (what a failed update restores). Nothing here removes, moves or recreates the store: update failure and store loss are different accidents and the recovery for one must not cause the other.
The exercise and the swap gate read the store through the relay's doctor and service status and a
catalog read that takes no lock. They run doctor with no option, so it only reads: it opens the
database read-only, creates no .probe- file and no SQLite sidecar where SQLite allows, makes no
read-write connection and never opens write-gate.lock; doctor --probe-write, which writes a
temporary file and begins and rolls back a write transaction, is not what they run
(decision 76).
The ownership reading copies the database into the temporary directory, the worker-policy reading
takes the daemon lock for an instant when a worker record exists, and a store left after an unclean
shutdown is read the plain way, which may create SQLite's -shm index beside it; none of these writes
the store's data. The
runtime opens it afterwards, for
the relay commands the skills run and for the Stop hook's guard whenever it has to read the store,
in the relay's default state directory with no variable set. Until todo 43 the Go build refused
that directory unless CRW_ALLOW_LIVE_STATE=1 was set, so an install left a runtime that could
not serve a live host's store; the guard now refuses it only under test isolation
(the live-state guard). A host whose
store the Python runtime still owns moves through the cutover first.
The staging claim is written last. It says this staging finished, and until the selection is committed and the owned pointer names the runtime there is nothing finished to say, so by the time writing it can fail the declared commands already reach the new runtime. The replacement has happened and only its record has not, and those are reported as two outcomes rather than folded into one.
| Field | Answers |
|---|---|
promoted |
whether this run replaced a runtime |
inService |
whether this runtime directory must be kept: true while the record selects it or the pointer names it, true once its claim has settled, and true when none of that could be read. False only when the readings say so |
claimSettled |
whether the claim recording it was written |
claim |
the claim's own outcomes (settled, released), its read-back and the selection snapshot that decided them |
recoveryRequires |
what has to be done next, under the same key a refusal reports it |
So crw install install has four exit statuses:
| Status | What this run changed | The record | What it means |
|---|---|---|---|
0 |
it landed, or there was nothing to change | written | the run finished; alreadyInstalled says when this archive was already the selected runtime |
3 |
it landed | not written | the pointer names the new runtime and the claim that records it did not settle |
1 |
nothing | not written | refused, or failed and put back what it had changed; whatever the host selected and reached before, it still does |
2 |
nothing | not written | a usage error, found before anything was read |
Exit 3 is not a refusal and must not be read as one. Non-zero here means the opposite of what it
means everywhere else in this command: the change landed, and a process may be running out of the
runtime it changed. A wrapper that reads every non-zero status as "nothing changed" would report the
old runtime as selected, or clean up a runtime that is in service. Key a cleanup decision on
inService and never on the status alone: a competing install can supersede this runtime between
the promotion and the result, and then status 3 is still correct about this run while inService
is false.
Which accident happened, and what to do about it, is in recoveryRequires, derived from the claim
as it reads back and from a selection snapshot taken under the promotion lock:
| What the result says | What to do |
|---|---|
| another run held the claim's lock | wait for that run; this call wrote nothing |
| the claim could not be read back | make it readable or remove it, then run the install again; the runtime is in service and must not be deleted |
| what this host selects could not be established | read the host record before acting |
| the record selects this runtime and the pointer does not name it | read the pointer before rerunning, because a rerun replaces that link first |
| the record selects this runtime | clear what stopped the write and run the same install again: it finishes an interrupted promotion and rebuilds nothing |
| the record no longer selects it, and something may still reach it | leave the directory alone |
| nothing selects it or points at it | nothing; another run moved the selection on, so do not rerun to settle it |
The snapshot is consistent, not durable. Nothing holds the promotion lock until an operator reads the result, so re-read it before acting if time has passed.
crw install update is the same run as crw install install; the name says which one you meant.
A new archive is a new runtime directory, because the directory is named for the archive's version
and digest, and the predecessor is never removed by an update. It becomes the host record's
outgoing selection, which crw install rollback returns to.
<destination>/current is a directory symlink, and everything that starts the runtime names a path
through it: the plugin's Stop command, its server launcher, and the three executables the settings
records name. Those strings are stable across every update, so an update rewrites none of them and
changes nothing Codex has trusted. It moves the link.
The pointer is a way to reach a runtime and never an identity. A process already started keeps the runtime it started in after the pointer moves: a running bridge or daemon goes on running its own directory's binary, which is why no update removes a predecessor. The next process started through the pointer is the new runtime.
A registered command and the runtime it reaches are separate claims, and crw doctor reports them
separately: the pointer's state and target, which kind of runtime the target is (go-binary,
or unknown), and whether the record selects what the pointer names (runtime.agrees). A
current repointed by hand at another directory is caught by that comparison rather than passing
because the command strings are unchanged.
crw install hook writes one plugin-owned Stop settings document, and every runtime the pointer
names serves it: relayExecutable <destination>/current/bin/codex-session-relay names the
runtime through the pointer, so no promotion and no rollback rewrites that document
(decision 18). Every move of the pointer reads it first, under the promotion
lock, and refuses, with nothing changed, when it names through the pointer a path a Go runtime does
not serve (bin/crw and its compatibility links); settings a user owns are judged with the second
owners. The settings are reported under settings with action none.
Until refactor R1 a move onto a Go runtime also replaced a Python-era document (one naming
.../current/bin/python3), archiving it as crw-completion-hook.json.superseded-<time>, and a
rollback could point at a Python env-* runtime that served the same document. The relay host's
first Go install made that replacement, and no Python runtime is left on it, so both are retired
(decision 61); a Python-era document found now is judged like any other and refused when it
names a path a Go runtime does not serve.
The runtime directory's name is deterministic and it is created with an exclusive mkdir, which is
what proves a run owns it. A run killed outright would otherwise leave the directory behind and
every retry of the same archive would refuse at the existence check for ever.
A run leaves two files in the directory, because they answer two questions. The lock,
.crw-staging-lock, answers whether anybody is still building, and it is created once and never
replaced: an advisory lock belongs to an inode rather than to a name, so a lock on a file later
replaced by rename would sit on an unlinked inode while the next reader found the new one free. The
claim, .crw-staging-claim.json, answers what that run said it was doing, and it is rewritten when
the staging settles. Removing anything needs positive proof of ownership, so the claim has to carry
this command's marker (crw install), its claim version and a state from the declared set.
Readable JSON at that path is not proof, and neither is a claim the retired Python installer wrote
(runtime_install.py), which is somebody else's file here (decision 63).
| Observed | Answer |
|---|---|
| No claim, and the directory holds files | Somebody else's. Refused, nothing touched, even when the host record selects something inside it |
| No claim, and the directory is empty | Taken over as it stands |
| A claim of this command's, the lock held | Another run is building it. Refused, nothing touched |
| A claim of this command's, the lock free, nothing selecting or naming it | An abandoned staging. Removed (through a tombstone, as remove removes) and built again, but only under remove's rules, read again under the promotion lock: kept, and the run refused, while a process may run out of it, a relay daemon record cannot be read, or a registration names it or cannot be read. One the record's outgoing names was in service, so it is kept and its claim settled. Where there is no process table (darwin) it is kept too, and the answer gives the recovery for a staging that was never promoted. A tombstone (.crw-removing-<name>) that holds such a claim is not reclaimed this way: an install reclaims the directory under a runtime's own name, never a tombstone for its own sake, and crw install remove finishes it (Removing a runtime) |
| A claim, and whether anyone holds it could not be established | Kept, and reported with what recovery needs |
| A settled claim, every component selects it and the pointer names it | Already installed. Reported, nothing rebuilt |
| A settled claim the record selects for one component and not another | Refused, naming each component's selection; crw install rollback <dir> selects every component and swaps nothing |
A selected runtime the host cannot launch as it stands (bin/crw not a regular file this user may execute, or a link that does not resolve to it) |
Refused, with repair: the commands that restore it in place (crw extracted from the archive the directory is named for, chmod 755, ln -sfn crw for each link), after which the same install answers already installed. Its directory is named for the archive, so it cannot be built again beside itself |
| A settled claim, and nothing selects it any more | Kept. It is a runtime that was promoted once, and a process may still be running out of it |
| An unsettled claim for a runtime that IS selected | An interrupted promotion. Finished rather than rebuilt |
| A lock held with no claim written | A run between taking the lock and writing its claim. Refused, nothing touched |
Finishing an interrupted promotion asks a narrower question about the link than a promotion does: not whether it agrees with the selection, which it cannot while the promotion is unfinished, but whether it still names a runtime this host record accounts for. It asks OPS-4.4 again as well, because the interrupted run recorded no gate verdict and every cell of the gate reads state that moves while nobody is looking.
Liveness is the lock and never a recorded process id: inside a container sharing a kernel the same
process id under the same boot id is a different process. Where flock is unavailable the answer is
that nobody could tell, and an owner nobody could establish is never read as an owner that is gone.
Deciding and acting are one step, under the directory's own <env>.crw-lock and then the promotion
lock (the lock order), so two retries that both find the same
abandoned staging cannot both act on it.
OPS-4.4 sequences an update around a daemon that is not running and open attempts that have been reconciled. Three readings answer that, each filling only its own cell:
| Cell | The reading that answers it |
|---|---|
daemon |
the selected relay's service status, whose running is decided by the lock a supervisor holds |
inFlight |
whether a store is there at all, then the selected relay's doctor, whose contents.openAttempts counts in-flight and held-uncertain attempts |
storeSchema |
the store's own schema, read read-only from sqlite_master without opening the store for writing, against the schema the candidate binary declares |
The in-flight cell reads twice, and the order is the point. The relay reports contents unavailable both for a store that is missing and for one it cannot read, and those are opposite answers here: an absent store has no open attempt, an unreadable one has an unknown number.
The swap proceeds only when the daemon is established stopped, the open attempts are established
zero, and the schema is established compatible: the verdict is ALLOWED. A cell that answered no
decides BLOCKED, and a cell that could not be read decides UNESTABLISHED; both keep the existing
installation, and the refusal names the cells. This command never starts or stops a daemon. OPS-4.1
gives the service to the scope operator, so a running daemon is a refusal here and not something to
resolve.
STOPPED describes a moment that has already passed: the relay's liveness answer releases its lock
before returning. Taking the reading inside the promotion lock narrows the window to the promotion's
own length; it cannot close it, because that lock excludes other runs of this command and says
nothing to a supervisor.
The relay declares schema version 1, has never raised it, and grows its schema through separate
CREATE ... IF NOT EXISTS statements. Every store therefore agrees with every candidate at version
one, and comparing versions would detect neither a downgrade nor an upgrade while looking exactly
like a check.
So the cell compares each object's CREATE statement in the store's sqlite_master with the
statements the candidate declares, keyed by kind and name (index sync_ready, not sync_ready),
over every object the catalog holds except the ones SQLite maintains for itself. Runs of whitespace
outside quoted text are normalised away and nothing else is.
| Answer | Observed | Decision |
|---|---|---|
NO_STORE |
no store exists at the resolved selection | allowed, and reported as absence rather than as agreement |
AGREES |
the same schema objects, defined identically | allowed |
EXTENDS |
the candidate declares objects the store does not hold | refused |
DIFFERS |
a shared object is defined differently | refused |
NARROWS |
the store holds objects the candidate does not declare | refused |
EXTENDS_ZONE |
the only difference is the additive DAG zone arriving: every added object is the zone's | refused, unless the command takes the OPS-4.5 backup (below) |
NARROWS_ZONE |
the only difference is the additive DAG zone leaving: every object the candidate does not declare is the zone's | allowed |
EXTENDS_INDEX |
the only difference is ordinary indexes arriving on tables both sides declare (a few may leave beside them) | refused, unless the command takes the OPS-4.5 backup (below) |
NARROWS_INDEX |
the only difference is ordinary indexes leaving: every object the candidate does not declare is one, on a table both sides declare | allowed |
The relay applies its whole schema on every write-open, so a candidate whose schema is not the store's applies the difference the moment its daemon first starts. OPS-4.5 reserves that for its own decision, with a copied backup of the whole state directory taken first, so an update never waves it through. The Go and Python runtimes execute the same schema statements (decision 14), and no Go release changes them before the commit point (cutover).
The one release that changes the schema is the one that adds the DAG zone (DAG plans, decision 74): its
dag_* tables are created by the first write-open and are declared by that build. That decision is made (D-01), so the gate has an answer of
its own for the zone and, below, for ordinary indexes, and for nothing else, and the supported install command has a route through both.
The history indexes (CRW-301). A second change reaches the frozen v1 schema script itself: six CREATE INDEX IF NOT EXISTS statements
for the tables that only grow (attempts, supervisor_attempts, acks, sync_outbox, supervisor_messages,
managed_start_requests; the list and the reason for each are in contract/schema/relay-sqlite-history-indexes.json). They are not zone
objects. Each is an ordinary index, non-unique and without a SQL function (the partial index on acks compares a column with a literal), on a table
both sides declare, so the gate gives them the answers of the paragraph on ordinary indexes below (pull request #397). A candidate that declares
them over a store that lacks them reads EXTENDS_INDEX, which refuses without --backup-state-to and passes with it; a store that holds them
against a candidate that does not declare them reads NARROWS_INDEX, which is allowed. The ratchet test
TestEveryNonUniqueIndexOfTheShippedSchemaIsOrdinaryButOne keeps the six ordinary. A store that lacks the zone as well, meeting a build that
brings the zone and these indexes together, is the case that paragraph leaves out: the zone beside an index is a plain EXTENDS, and
--backup-state-to does not release it. The build itself creates the indexes: the schema script runs on every write-open, so the first command to
open a store of the previous version builds them, one index in a transaction of its own, and a command that declares itself read-only builds them
too, since it tries the writable open first (a read that opens nothing for writing builds nothing). The swap gate tests
TestTheHistoryIndexesArriveAsAnIndexExtendsAndLeaveAsAnIndexNarrows and TestTheHistoryIndexesAgainstRealStores pin this reading.
The zone arrives (EXTENDS_ZONE). Installing a build that declares the zone onto a store that has none (every store there is) refuses,
and the refusal names the route: crw install update --from ... --backup-state-to DIR (the same flag is on install and rollback). The flag is the
operator's acknowledgement, and under it the command itself takes the OPS-4.5 backup, so the backup is guaranteed by the route and not by
anyone's memory. Inside the promotion lock, after the daemon-stopped and no-open-attempt cells have answered and before anything is promoted,
it copies the whole state directory the gate read to DIR: copy only, byte for byte, the source opened read-only and nothing moved, recreated
or deleted, here or on a failure. Directories and regular files are copied with their bytes and, once every byte is in place and verified, their permission bits, each synced after its mode is set so the mode and not only the bytes survives a power loss (the directory the backup is made in stays 0700, and a copy always keeps its owner able to open it: a source the installer could read only through its group gets the owner's read bit, and the manifest records both modes); a symbolic link to a regular file is copied as the
file's bytes under the link's name (a link alone would back up nothing), and when relay.sqlite3 is such a link the real file's -wal and -shm
are copied beside it, so a restore opens with the commits only the log held; a socket, a FIFO or a device is listed as skipped. Each file is
hashed while it is read and synced; then the state directory is read again, and the listing, every size and every file's digest, and the digest of
every file in the copy, must be what was copied. Any difference, in any file, refuses the swap ("the state directory changed under the copy"):
the copy is of one moment or it is not made. The backup, its manifest and the directories made for them are synced in their parents after the manifest is written, so a power loss cannot keep the files and lose the names that reach them. DIR and its manifest must not exist, must not lie inside the state directory, and must not lie
inside the runtime destination tree (a failed run removes its candidate runtime, and a backup there would go with it); the free space on its
filesystem must cover the directory. The record is written last, beside the backup and not inside it (DIR.manifest.json: source, destination,
time, issue, every entry with its size, mode and digest, what was skipped, an aggregate digest), and the command's result carries the same facts
as swapGate.stateBackup. A backup that fails part-way stays where it is and is reported with partial: true; a swap that fails after the backup keeps
it too, and a rerun needs a new destination.
The route carries the zone arriving alone. A running daemon, an open attempt, a cell that could not be read, another schema object arriving or
leaving, or any object defined differently (a dag_* object included) still refuses with the acknowledgement as without it, and no backup is
taken for a swap that does not happen. The zone's objects are the ones the build's own zone statements create, not whatever starts with dag_: an
object this build does not know, or a later step of the zone ledger, is a plain EXTENDS or DIFFERS until a build that knows it is the one asking.
The acknowledgement where nothing arrives is not a refusal and takes no backup. The route exists in a build that carries it: the installed
current/bin/crw of an earlier build reads the same arrival as a plain EXTENDS, so the update is run with the new build's binary from the archive
(the documented path).
The zone leaves (NARROWS_ZONE). Returning to a runtime that does not declare the zone, on an update, on a promotion an interrupted run left to
finish, or on a rollback that moves the pointer, is not refused for that reason: the zone is additive, so an older runtime opens a store that has
it (it validates only the frozen tables) and never reads or writes the zone, which stays in the store for a newer runtime to read again.
Ordinary indexes arrive (EXTENDS_INDEX) or leave (NARROWS_INDEX). A release may add an index to a table the schema already declares, and then
the store lacks an object the candidate declares, which is a plain EXTENDS. The gate gives a difference made only of ordinary indexes on tables
both sides declare an answer of its own. An index is ordinary when it is not unique and calls no SQL function, in its key or in the WHERE of a
partial index. SQLite decides that and the gate does not match text: it builds every table both readings declare in memory from the store's own
statement for it, executes the one CREATE INDEX statement (a statement with a semicolon, even inside a literal, is not classified), and asks
SQLite for the table the index is on, whether it is unique and, from SQLite's own compiled program for the statement, whether it calls a function.
A unique index can fail to build over existing rows and then fails the writes of a runtime that never heard of it, and so can a function:
json_extract over text that is not JSON raises, which an older runtime's ordinary write would meet. Neither is ordinary, however few rows are
involved. An index on a table that is new, that differs between the two readings, or that the statement does not name correctly is not ordinary
either, and neither is a trigger, a view, a new table or a column change: each keeps the answer it had (EXTENDS, DIFFERS or NARROWS),
alone or beside an ordinary index, and the backup route does not carry it. The zone beside an index is neither class (a store without the zone
meeting a build that brings the zone and an index together is a plain EXTENDS, as is the zone and an index leaving together a plain
NARROWS), so the route is taken for the zone first, with a build that carries it alone (refactor backlog).
An arrival refuses without --backup-state-to and passes with it, by the same route as the zone: the backup is taken after the daemon-stopped and
no-open-attempt cells have answered and before anything is promoted, and the candidate builds each index on its first write-open. The schema
script runs as separate statements, each in its own transaction, so the write lock is held for the build of one index at a time, never for the
script. Measured on a synthetic store built by this build (the zone present, the six history indexes of pull request #384 absent, filled with the
generator of that change's benchmark; page cache warm, on a host that was running other jobs; WAL mode, so readers were never held), with a second
connection updating one row every 2 ms to see the lock:
| Store | Events / attempts | The first write-open (six indexes) | The largest index (attempts_sent_at) |
Each of the other five | Longest write stall | Growth of the file |
|---|---|---|---|---|---|---|
| 89.8 MB (larger than the operating store's about 65 MB) | 80,000 / 160,000 | 0.12 to 0.13 s | 0.10 s | 3 to 22 ms | 0.08 to 0.105 s | 8 MB |
| 447.8 MB | 400,000 / 800,000 | 0.50 to 0.56 s | 0.48 to 0.54 s | 2 to 45 ms | 0.33 to 0.43 s | 40 MB |
The build is a small fraction of a second at the operating store's size and scales about linearly with the attempts table, which holds the one large index. A writer that meets the lock waits for it: the probe's longest wait can exceed the build by up to one busy-handler retry step, and no probe write failed. Three runs of each size are in the figures, and the readers' longest wait was 3 ms.
An index the candidate does not declare leaves with no acknowledgement and no backup, on an update, on a promotion an interrupted run left to finish,
and on a rollback. An ordinary index is derived from its table's rows and is never needed for correctness, and SQLite maintains every index of a
table on every write whichever runtime made it, with no function that could fail, so an older runtime opens such a store and keeps it up to date
without reading it; the index stays for a newer runtime to find. The one effect is on speed and on the order of a query that has no ORDER BY,
which may now follow a leftover index. A unique index, or one that calls a function, that the candidate does not declare is not that: an older
runtime's write can fail on it, and the store reads a plain NARROWS. A new index meant for this route is therefore non-unique and function-free;
the ratchet test names the one index of the shipped schema that is not (journal_managed_creation, whose WHERE guards json_extract with
json_valid), and a new exception fails it. The test of what calls a function reads SQLite's compiled program, whose opcodes SQLite does not promise
to keep, so the tests keep json_extract as a control that fails when a driver upgrade changes them.
The warning against write commands stays: opening a store for writing with this build creates the zone, and that happens outside this route (a relay command run by hand against a live state directory takes no backup), so the route makes the warning unnecessary only inside the install command. Run no write command of this build against a live state directory before the install has taken its backup.
The selection in the host record and the pointer on disk are two truths, and both the order they are written in and the lock they are written under are the safety argument. They are written inside one critical section holding the promotion lock, and the selection is committed first.
The reverse order has a real failure: the link lands, the record write then fails, recovery reads a selection that does not name this runtime, concludes the candidate was never promoted, removes it, and leaves the declared commands pointing into a directory that no longer exists. OPS-4.4 requires every state transition to be committed before its side effect. Recovery also refuses to remove a runtime the pointer names, so neither truth alone can authorise deleting a runtime the other one is still using. Every judgement the promotion makes is decided on state read inside that lock.
The promotion lock is an advisory lock on one host-wide file beside the host record,
host-record.json.promotion-lock, created once, never unlinked, and released by the operating
system when its owner dies. It excludes runs of this command and of the Python installer, which
take the same lock, and nothing else. A kill inside the window leaves a runtime that is selected
and unsettled, which the next run recognises and finishes rather than rebuilds.
Ownership of the pointer is established from the record before it is replaced. Renaming over an
existing symlink succeeds whoever created it, so a current this host record never recorded
placing is left alone; a real directory at that path fails the rename outright, which is the safe
direction. The pointer placed is then proved: it has to resolve, with no loop and nothing dangling,
to the runtime directory itself by identity, and <pointer>/bin/crw has to be a regular file this
user may execute.
Every crw install command takes its locks in one order (decision 33): a
runtime directory's <env>.crw-lock first, then the host-wide promotion lock, then the host
record's own .crw-lock inside both. The settings records' locks, a claim's and the launcher
copy's are leaves: nothing else is taken while one is held. It is runtime_install.py's order too.
Every wait before a command's first write ends when the command is interrupted (SIGINT, SIGTERM,
SIGHUP): it stops waiting, writes nothing and says so (register-mcp and hook answer the outcome
interrupted). A wait inside a sequence already under way, such as the restore after a failed
promotion or the entry drop after a directory was set aside, runs to completion, because stopping
there would leave the host half-written.
A failed update leaves the previous runtime selected, the previous pointer target in place, the
Stop settings as it found them and the store exactly as it was. The result says which step failed
rather than only that something did: failedStep names it, retriable says whether the same
archive can be installed again, residualPaths names what the run left, and recoveryRequires what
has to happen first.
Putting a selection back is narrower than it sounds, and deliberately. The restoration runs under the promotion's own lock and puts back only the entries that still name what this run wrote; an entry another run has promoted since is left alone. Putting a selection back includes putting it back to nothing, for a first install that had no previous selection, and putting the pointer back includes putting it back to absence. A restoration that cannot be read back reports a residual pointer and keeps the candidate rather than claiming the rollback completed.
A candidate that is selected, whose record could not be read, or that the pointer names, is kept. Otherwise the run drops its install entries and removes the directory it created, through a tombstone as remove does, and the result says whether that removal was verified: verified means the same archive can be installed again; not verified names the residual path and what recovery needs, with the original failure reported beside the cleanup failure.
Two failure points an update might be expected to have do not exist here. Registration is not part of an update: the declared commands name the pointer, so nothing is registered again. And nothing starts: OPS-4.1 gives the service to the scope operator, so this command refuses while a daemon runs and never starts one.
This is a POSIX path. The runtime layout, the directory symlink and the advisory locks are POSIX assumptions; Windows is out of scope rather than approximated.
crw install rollback points the owned pointer back at the host record's outgoing selection, the
runtime the last promotion replaced. crw install rollback <dir> points it at a runtime directory
the record lists exactly, as an install entry's environment, instead; an empty <dir> is a usage
error. It takes the target's <env>.crw-lock and then the promotion lock
(the lock order), finds the target again there and refuses with
nothing written when the record moved it, and then applies a promotion's gate, ownership and
second-owner rules, commits the selection before the link, and proves the pointer it placed. The
runtime it leaves stays installed and becomes the outgoing selection, so a second rollback returns
to it. Where there is no outgoing selection, a bare rollback refuses and writes nothing.
The target has to be a Go runtime (bin/crw) launchable as it stands, judged without running
it, and its claim has to be settled, or unsettled with nobody holding it where the record shows a
promotion committed it (an exit 3). A rollback never rewrites the Stop settings
(one Stop settings document). A directory the record lists that is
not a Go runtime - a Python env-* runtime the fence installer made - is refused with nothing
changed (decision 61).
Install rollback is not takeover rollback. crw install rollback moves which runtime the pointer
names, and it rewrites no settings. It does not move the store's ownership. After the cutover
the store is owned by the Go runtime, and handing it back to the Python fence release was
crw relay takeover rollback --to python --python-relay <path>, which named the Python relay by
absolute path and never through the pointer (cutover rollback); since
todo 44 no Python candidate is launched and that rollback is refused (decision 48), and refactor R1
deleted the takeover command (decision 54). One does not imply the other, and a return to
Python needed both; with the store's half refused and crw install rollback taking only a Go
runtime, this revision offers no return to Python. The one order fixed for the pointer is the
plugin payload's (update and roll back): the payload
goes back before the runtime does.
crw install remove <dir> deletes one bin-* runtime directory directly under the destination,
and only when nothing may still be using it; a Python env-* directory the retired installer made
is not one it takes (decision 63). The directory is judged by file identity,
so naming it through a symlink, a bind mount or another spelling of the destination passes no check
its own name fails. It refuses a directory the record selects, one the pointer names or might name,
one with no claim this command wrote, one whose claim cannot be read, a
staging another run still holds, any directory a live process runs out of (its /proc/<pid>/exe,
its working directory, or what its arguments run, resolving inside it, and every relay daemon a
daemon.json or scope claim records alive), and any directory a registration the host reads names
a path inside: the Stop settings, the bridge record, the cached plugin declarations, hooks.json
and config.toml. A process it cannot rule out or a
CRW registration it cannot read or judge refuses too, and the answer names it; another program's
entry in hooks.json or config.toml that it cannot judge does not (what remove reads). A process whose working
directory cannot be read and that runs something by a relative path is one it cannot rule out,
whatever user runs it, unless that user (not root) is provably shut out of the directory by the
permissions of the directory or a parent. What counts as running something is the program and,
for a shell (sh, bash, dash and the like), the script it reads; a runtime directory holds no
Python, so an interpreter's script operand is not followed (decision 62). So on a host where a
root agent runs a relative shell script remove refuses, names the process, and gives the manual
recovery below; a root Python agent run relative (every Azure VM's WALinuxAgent) does not
refuse it.
The process reading is this host's process table, in this command's PID namespace, and every answer
that rests on it says so under processTable: a process in a container sharing the directory, or on
another host sharing a network home, is not seen, so run crw install remove where the runtime's
processes run. Where there is no process table it can read (darwin has no procfs), whether a relay
or a bridge still runs out of the directory cannot be established, and every remove refuses,
fail-closed, with the recovery by hand under recoveryRequires: stop the relay daemon started from
it (<dir>/bin/codex-session-relay service stop) and end every Codex session whose bridge it
started, delete the directory, then run crw install status to see that the host record and the
pointer still name the runtime you meant. Its install entries stay in the record, where a rollback
naming it is refused because the directory is gone.
A removal survives a kill. Under the directory's <env>.crw-lock and then the promotion lock, the
directory is renamed in one step to its tombstone, <destination>/.crw-removing-<name>; its install
entries leave the host record in one write, with an outgoing that names it; only then is the
tombstone deleted. A drop that cannot be written renames the directory back and refuses, and a
deletion that does not finish exits 3 naming the tombstone. A kill anywhere leaves either the whole
directory or a tombstone. crw install status lists every tombstone under interruptedRemovals,
and crw install remove of the tombstone, or of the runtime's name when only its tombstone is left,
finishes it once no process runs out of it and no registration names it. Do not delete a tombstone
by hand: finishing it also drops what the host record still lists under the runtime's name.
crw doctor's residue lists a tombstone whose claim is an abandoned staging's (a reclaim, or a
failed run's candidate, killed before its deletion finished) and names the same crw install remove for it as status does. A tombstone a removal of a settled runtime left, or an empty one,
is not residue by the staging decision, so status alone reports it. A
tombstone that holds files and carries no claim of this command's is somebody else's directory:
status reports it ours: false, and nothing removes it.
What remove cannot see is a command fixed before the update that will start a process later: a turn holds the hook command it resolved when it started until it ends (the turn-command cache). Remove a runtime the pointer left only once the turns that started before the pointer moved have ended.
Every registration the host reads is read as the program that reads it runs it, and each path it
names is resolved through the owned pointer and every link, as are the interpreters and scripts it
is followed through. Each finding and each entry that could not be read or judged names its row,
the source file and the field; the row numbers are those of the retention scan the command's
readers came from, which crw doctor retention-scan ran until it was retired (decision 59):
| # | Registration | What is read |
|---|---|---|
| 4 | <CODEX_HOME>/crw-*.json (crw-completion-hook.json, crw-bridge-mcp.json, any other), and every settings document a Stop command below names |
relayExecutable, bridgeExecutable, interpreterPath, command, each as written (the retired adapterEntryPoint and adapterInterpreter run nothing and are not read, decision 66) (the launchers run them with no shell and no expansion), and each args entry as an argument. A value that is not an absolute path, or not a string or a list of strings, is unreadable |
| 5 | <CODEX_HOME>/plugins/cache/crw/crw/*/wiring/hooks/*.json, wiring/mcp.json, .mcp.json in every version directory (a stray file there declares nothing) |
every hook command, under the grammar below, and every MCP server, as Codex starts it |
| 9 | <CODEX_HOME>/hooks.json |
every hook command, as row 5 |
| 10 | <CODEX_HOME>/config.toml mcp_servers.* |
every MCP server, as row 5 (a relative cwd is unplaced) |
hooks.json and config.toml hold other programs' registrations too, so what the reading cannot
judge of one entry there (a hook command, an MCP server) holds a removal back only when the entry
is CRW's: when a word it names or runs, a path it classifies, or where that path resolves is
crw, codex-session-relay, codex-thread-bridge or crw-completion-hook, or lies in the
destination (spelled from $HOME, ${HOME} or ~ too). Another tool's SessionStart hook
running an interpreter's script, or a server started over ssh or by node, leaves a CRW runtime
unused and no longer refuses a removal (decision 68); on the relay host three such entries did
until todo 43 ran the removal against an edited copy of the Codex home. An entry with a field of the
wrong type (a server whose args is 42, a hook whose command is a list, an event that is not
a list, a server that is not a table) is judged the same way: every string it holds is read as a
shell line and as one program, under the entry's own cwd, env and PATH, and searched for
the names above and the destination, and the entry holds a removal back only when that finds CRW in
it. Its server name counts (CRW registers the bridge as codex-thread-bridge), its map keys are
matched by spelling and never read as paths, and a number or boolean names nothing. What any entry
names inside the directory is still found, a file that cannot be read or parsed, or whose top-level
structure is not what the host reads (hooks that is not an object, mcp_servers that is not a
table), still refuses, and the CRW files (rows 4 and 5) refuse on anything they hold that cannot be
judged. crw doctor reads the same two files as one document the way Codex accepts it, so there a
server of another program with args = 42 makes the bridge's registration unreadable; the doctor
asks whether the selected runtime is registered as the host reads it, and remove asks whether any
registration can name the directory it would delete.
Row 8, the <CODEX_HOME>/crw-stop-hook.py launcher copy, is no longer read (decision 67): the
relay host holds none, and the Python bootstrap that fell back to it left the plugin cache with the
pre-native payloads. A Stop command's settings are the document the word after crw hook (or after
a crw-completion-hook link an older runtime still carries) names, which a runtime installed
before decision 66 reads, else <CODEX_HOME>/crw-completion-hook.json; CRW_COMPLETION_HOOK_CONFIG
is not read.
A hook command, a shell script a reference reaches, and a program handed to sh -c are parsed with
mvdan.cc/sh/v3/syntax (Bash grammar; decisions.md 37) and judged only as far as they are written
in a grammar read completely: simple commands joined by ;, &, &&, ||, |, |& and
newlines; words that are literal once a leading ~, $HOME, $CODEX_HOME and (row 5)
${PLUGIN_ROOT} are expanded; redirections to such words; and as commands exit, true, :,
exec, an absolute path or a bare name found on PATH (a relative or empty PATH directory met first
leaves the name unreadable), sh, bash and dash with -e, -u, -x, -f and -c, and env
with -i and --. Every other construct is unreadable and listed, never interpreted. An MCP server
is judged as Codex starts it: command and args exec'd with no shell and no expansion, in its
declared cwd, with its declared env over HOME and PATH.
The relay records are every daemon.json in the relay state root, each directory under it (a link
to one included), CODEX_SESSION_RELAY_STATE with ~ expanded, the state directory --state names,
and each stateDir the relay's scope registry records, together with the registry's own claims.
The registry is the one the relay resolves in this command's environment:
CODEX_SESSION_RELAY_SCOPE_DIR alone when it is set, otherwise
<passwd home>/.codex-session-relay/scopes. A pid counts only while its start time and boot still
match the record; a pid or workerPid that is present but not a positive integer pid, a record
naming a boot whose current id cannot be read, and an alive process whose executable or command
line cannot be read are unreadable.
Two read-only commands, each answering a different question:
crw install statusanswers what a replacement actually left: what the host record selects, what the owned pointer names and whether they agree, whether the record's pointer is the fixed destination's (destinationAgrees, with the repair when it is not), every runtime directory with its claim, every unfinished removal (interruptedRemovals), the outgoing selection, the promotion lock and the two settings records.crw doctoris the host diagnosis: the host record reading, which runtime kind is selected, the Codex CLI version against the one the wiring was measured on, the App Server observed through the selected bridge, the pointer, the promotion lock, the settings records and every executable they name, the classification of each component, the residue on the destination (an abandoned staging a later install would reclaim, and the tombstone of an interrupted removal, whichcrw install removefinishes), the relay's own scope reading, and the six results.--temporaryrecords the destination as a temporary one, so a proof taken there is never read as a claim about a host.
None of them takes a lock across the whole reading, so a host changing underneath is described in pieces.
The skills run codex-session-relay ... as a shell command (relay usage),
and that name is found on PATH. The installer does not manage PATH: it places the name in
<destination>/current/bin/, and nothing puts that directory on anyone's PATH. So either put it
there, ahead of any other copy, in the environment the Codex tasks run in, or run the relay by that
absolute path. Check what a task actually reaches:
command -v codex-session-relay
readlink -f "$(command -v codex-session-relay)"
readlink -f <destination>/current # the runtime directory the pointer selectsThe first answer has to lie inside the second. On a Go runtime (bin-<version>-<digest12>) it is
that directory's bin/crw, which the name links to. On a host still on the Python fence release,
before the cutover, the pointer selects an env-* directory and the answer is its
bin/codex-session-relay, a console script of that environment.
A stale codex-session-relay earlier on PATH, such as a console script under ~/.local/bin whose
shebang names a Python virtual environment in a development checkout, runs whatever that checkout
holds against the same store, which can be code from before the fence release and so outside the
cutover's fence. Neither crw install nor crw doctor reads PATH for it, so this reading is
the one that finds it.
internal/runtime/definition is the single compatibility definition OPS-1.1 requires: the two
components, their console-script names (the compatibility names beside crw), version, licence
and the tool that identifies the bridge. crw-dev ci contracts checks it against what it names
outside itself: each licence is in the checkout and among the files the release archives carry,
the identity tool is a tool the bridge's contract lists, and the links the installer places are
the ones the release build makes. Release digests are not in it: they live in the release's
SHA256SUMS and in the host record (decision 35). A host record still
states definitionVersion 1 (decision 34).
Until todo 44 the definition was a file, scripts/crw_runtime/components.json, which kept the
Python packages' trees and source digests for the Python installer to re-derive; it left with its
last Python reader, and the Go package, which had carried the fields a Go install uses, became
the only copy (decision 47). The ported bridge's upstream provenance, which the file recorded
because the import brought source rather than history, is kept in
its PROVENANCE.md beside its licence, the
provenance narrative OPS-1.5 says is retained rather than replaced.
What the definition does not carry is as important. Installed locations, entry points, host names and measured points are host facts. OPS-3.2 makes a real record a private receipt, so they go to the host record outside this repository and this repository never commits one.
The host record is the other half of the definition and is never committed. It is recordVersion
1, which the Go and Python installers both read and write: Go install entries carry fields of their
own beside the Python-era entries, which keep theirs. It holds, per component, one entry per install and the measured points, plus what is
selected, the outgoing selection a rollback returns to, and the owned pointer with who placed
it and when.
A Go install entry records location (the runtime's bin/), environment (the runtime
directory), entryPoint (the component's compatibility name), binaryDigest and integrity (the
binary's SHA-256), target, version, archiveDigest, reachedVia, and source, the repository
commit and tree the binary was built from. An entry the Python installer wrote carries
interpreter, interpreterPath and installMode instead of the binary fields. crw doctor reads
both, and classifies only Go installs: a Python install keeps the Python installer's classification
until the Python path is removed.
Under OPS-1.3 a point means the combination was exercised, so no amount of reading bytes produces
one, and an install whose bytes match but whose combination nobody has run classifies unmeasured
and is preserved rather than reused. The install's exercise is what produces a point
(installing the runtime, step 4), and a connection, protocol or tool-call
failure records no point and says why. A Go point records install, installDigest, codexCli,
host and appServer, the dimensions it is matched on, beside date, measuredBy, method and
archiveDigest. Points are appended, never replaced.
The appServer dimension is the bridge's whole get_capabilities payload, serialized and compared
by exact string equality: the identity of the server a combination was exercised against is
everything the round trip reported, not a version string someone chose to trust. So a change to that
payload's shape changes the dimension, and a point measured before it does not cover a bridge built
after it. The old point is not wrong and is not discarded; it remains evidence about the build it
was taken against. The same property is why the payload carries no timestamp: a dimension that
varied between two calls to the same build would never match itself.
Every record these commands read (the host record, the settings records, a claim, the Codex configuration, the hook file) is read at a narrow boundary that turns a failure into an answer rather than a crash. The answer is one of four states, decided by an ordered observation rather than by a convenience test:
| State | What was observed |
|---|---|
ABSENT |
nothing exists at the path. The only state that may be read as a host with no history |
PRESENT |
it was read. An existing record with nothing in it is PRESENT, not ABSENT |
UNREADABLE |
something is there and its shape cannot be read: a directory or other non-regular file, a symlink whose target is established missing or looping, invalid UTF-8, unparseable JSON, or containers of the wrong type |
ACCESS_ERROR |
nothing could be established: a permission or I/O failure reaching the path, a symlink whose target could not be resolved, or a parent directory that cannot be traversed |
The distinction that matters most is the last row. Being unable to ask is not being told no, so a failure to establish existence is never reported as absence, and a permission problem is never reported as a malformed record. A refusal names what failed and where.
What this guarantees: the worst case for a record these commands read is a named refusal, not a crash. What it does not guarantee: that a record which could have been read is never refused. That direction is the safe one, and the refusal carries its reason.
- "Nothing was written" is scoped to what can be guaranteed. Malformed input detected before
the first mutating step refuses and the target file's bytes are unchanged. A read-back failure
after a write reports the write instead of denying it (
record_applied_unverified,config_applied_unverified) and exits non-zero, because reporting a landed write as a refusal that wrote nothing invites a retry over a file that now exists. - A read-only diagnosis reports rather than refuses.
crw doctornames the failed reading inhostRecordStateandhostRecordReadingand continues with what it could still observe, and never reads an unreadable record as a clean host. Commands that write refuse outright.
Every change to the host record takes the record's lock, loads the record inside it, applies the caller's narrow delta and saves. A caller never hands back a record it loaded earlier: an install that loaded the record, spent minutes unpacking and exercising a runtime, and then saved what it had loaded would overwrite whatever another run committed in between, and holding a lock over that save would not help, because the staleness is already inside the value. So a caller says what it learned (this install, these points, this selection) and the merge happens against the record as it then stands. A failed install drops only the install entries keyed to the directory it created and leaves the selection exactly as found, because another run's successful promotion is not this run's to undo.
crw doctor classifies each component of the runtime the pointer names from the OPS-2.1 signals:
whether its entry point resolves into a recorded install, whether the binary's digest matches the
recorded binaryDigest, whether a measured point covers the combination it runs under, and
whether a registration conflicts with it or the owned pointer names something other than what the
record selects. A Go install is an archive, not a checkout, so
the checkout signals do not exist for it: a fork is bytes that differ from the recorded digest. A
signal that cannot be read stops classification and is reported; it never counts as a signal that
agreed. A runtime a host cannot launch, a bin/crw that is not an executable regular file or a
compatibility name that does not resolve to it, is reported as such, because a digest says nothing
about either.
The five OPS-2.2 classes are evaluated in their fixed order and the first match wins: conflict,
fork, foreign, unmeasured, own. Only own is reused. An install whose bytes match and whose
combination has no point classifies unmeasured, and that is the intended answer, not a gap to
close by relaxing the rule: OPS-1.3 refuses to let a matching digest stand in for a run nobody
performed. Exercising it, which is what an install does, is the way out.
Nothing outside a recorded path is ever overwritten, and no predecessor is removed by an install or an update.
The plugin package declares the server, and installing the package registers it. What the package
cannot carry is where the runtime is and which execution policy the bridge runs under, so that goes
into <CODEX_HOME>/crw-bridge-mcp.json, which crw install register-mcp --owner plugin writes and
the launcher reads at every start (the native wiring). The
record names the owner, the server name, the bridge executable
(<destination>/current/bin/codex-thread-bridge unless --bridge-command says otherwise), its
arguments (--bridge-arg, repeatable), and, from version 2, the execution policy. --dry-run
reports what would be written and writes nothing.
Writing the record is not registering a server. On a host with no plugin installed the record is inert, and the command says so. Registered, recorded and a tool actually called stay three claims.
One owner registers each surface. --owner plugin is the only owner crw install writes for:
the user-owned registration, a [mcp_servers] table in config.toml, retired with the Python
installer. The refusal still runs both ways, because a host can acquire the other half from
elsewhere. A record is refused while the Codex configuration already starts this bridge, under this
server's name or under any other, and the refusal names that entry: the plugin's declaration beside
it would run a second bridge. The launcher stands down for a record that names another owner. A
record or configuration that could not be read refuses rather than defaulting, because installing on
an unanswered question is how the second bridge arrives. Every writer of the record takes the
ownership lock beside it (crw-mcp-ownership) while it writes.
Codex starts a plugin-declared server with the App Server's own environment. Measured on Codex
Desktop 0.154.0 for Linux, that is HOME LANG LOGNAME PATH SHELL USER and nothing else. A bridge
started that way reads no execution policy: get_capabilities reports presence_only with no
roles, and a child created through the Desktop tools is never asked whether it runs its role's pair.
--execution-policy gives the plugin-owned record the one fact that closes that gap:
crw install register-mcp --owner plugin --execution-policy /path/to/execution-policy.json
The record becomes version 2 and gains executionPolicy, with two fields: path, the file as given
(expanded and made absolute, not resolved, like the relay's own launch declaration), and digest,
the SHA-256 of the bytes this run read. It never carries what the file says. Before anything is
written, the file goes through the bridge's own parser in this binary, so a policy the bridge would
refuse to start under is refused here instead, as execution_policy_unreadable. The output reports
the mode and the declared role pairs, the same values get_capabilities discloses, and says which
parser judged them: the installed runtime parses the file again every time it starts and decides for
itself.
At every start the launcher reads the record and does one of two things. It refuses and exits 2,
naming the record and the repair, when the file is missing, is not a regular file, cannot be read,
or no longer hashes to digest, and also when its own environment already names a different policy
file or digest. Otherwise it starts the bridge with CODEX_THREAD_BRIDGE_EXECUTION_POLICY set to
path and CODEX_THREAD_BRIDGE_EXECUTION_POLICY_DIGEST set to digest, and the bridge refuses to
start if the bytes it parses hash to anything else. Refusing to start is the visible failure. The
declaration marks the server not required, so the session continues without the bridge's tools, and
it never continues with a bridge that checks no role. A version-1 record names no policy and starts
the bridge with the environment the launcher was given.
The policy is part of the registration's identity, so the only rerun that succeeds is an identical
one, answered record_unchanged. Any other difference is refused as record_differs and nothing is
written: another file, the same file with other contents, a rerun that drops the flag, or adding a
policy to a version-1 record. The refusal names the repair: move the record aside by hand, then run
crw install register-mcp again. Every edit to the policy file, including adding an exception,
therefore has two consequences. The relay picks the edit up when its daemon restarts. The bridge
record has to be moved aside and registered again, and a thread started in between has no bridge
tools. The digest is what lets a changed file fail visibly instead of being enforced unregistered.
A created record is reported only after the policy file has been hashed again, following the write,
and still matched. A mismatch immediately before the write writes nothing and answers
record_policy_changed; a mismatch after it gets the same answer and the move-aside repair, and the
record stays, because the launcher refuses its stale digest at every start. A removal by path cannot
exclude a writer that does not take the ownership lock, such as an editor, so nothing here removes a
record.
Codex starts the server once for each thread it loads. That was observed on Codex Desktop 0.154.0:
one App Server process with a separate bridge child per thread, and the child's start time matching
the thread's creation to the second. A record written now therefore takes effect for threads started
afterwards, with no App Server restart, and a thread already running keeps the bridge it spawned.
Neither is established on a host until a new thread's get_capabilities reports the digest the
record names. A package replacement may or may not reach threads that are already running;
what two replacements measured says what was
seen, and updating safely gives the order and the checks.
crw doctor reports the six OPS-6.1 fields under checks.results, in the OPS-6.2 shape, each with
its own evidence, command, acting process and measurement time. A field with no timed observation
behind it reports unknown; no time is ever invented or copied from another field.
| Field | Established by | Never established by |
|---|---|---|
installed |
Every component classifies own |
The runtime being present |
mcpExposed |
Tool names observed in a live Codex session | A record or a declaration |
connected |
The relay's doctor reporting actorReachability.socketConnect as ok |
A socket file on disk |
deliveryAccepted |
An attempt that recorded a returned turn id | A dispatch or an absent error |
verificationComplete |
Every OPS-6.4 condition at once | A completed turn or a green check |
alwaysActive |
A supervised runtime surviving a host restart, observed after one | Any of the five above, or a registered unit no restart has tested |
A live session is the only thing that can list the tools a Codex session exposes, and the
diagnosis's own session with the bridge is not Codex's, so mcpExposed stays not_verified from
crw doctor; read it in a task. deliveryAccepted needs work created and delivered, which a
diagnosis never does, so it is not_applicable unless a trial produced it. settingsPreserved is
reported beside the six and never merged into them: it answers that the diagnosis wrote nothing.
OPS-3.1 puts one relay service and one durable store behind an entire operating scope, which is one
host, one OS user and one App Server. A second repository or a second project installs into that
same scope and reuses the same service and the same store; nothing here creates a daemon or a store
per project, per repository or per parent. crw doctor reports the resolved scope, the store and
the service by calling the relay's own doctor rather than by rediscovering any of it, and reports
other stores beside the resolved one without adopting any of them. Equality of path strings is not
proof under OPS-3.4; proof is doctor from each participating process reporting the same state
directory together with assignment-find --issue returning the expected relationship.
A relay service started by hand does not come back after a host restart: service start leaves a record that reads
"recorded before a different boot", and nothing starts it again. crw install register-service --socket <app-server-socket> [--state <dir>] registers the one systemd user unit that does. It is a command of its own, like hook and register-mcp,
because an install or an update never starts a daemon and gains no side effect here. Starting at boot, before anyone logs in,
needs the user manager to start at boot, which a host with lingering off does only at login.
The unit is written to ${XDG_CONFIG_HOME:-~/.config}/systemd/user/crw-relay.service (--unit-dir and --unit-name change
either) and enabled with systemctl --user enable. It is Type=oneshot with RemainAfterExit=yes, wanted by
default.target; its ExecStart is the relay's own service start and its ExecStop is service stop. Both name
<destination>/current/bin/codex-session-relay with the state directory and socket resolved at registration, so a pointer
move never rewrites the unit, and the relay's own checks (intent, launch policy, scope) apply to the start unchanged. The unit
has no restart policy: after the one start it acts only when an operator acts on it, and a refused start (service_disabled,
already_running) is a failed unit with the relay's answer in the journal, never a retry that could start the relay in the
middle of an update. It unsets CODEX_THREAD_BRIDGE_EXECUTION_POLICY and CODEX_SESSION_RELAY_SCOPE_DIR, which the user manager
would otherwise pass in. --scope-dir <dir> registers an isolated target instead (a temporary state directory and socket):
the unit sets that scope variable and the start carries --allow-isolated-scope.
One owner registers this surface, and the command asks the manager as well as the file. It refuses, writing nothing:
| Outcome | When |
|---|---|
unit_foreign |
something at the path is not the installer's unit (it lacks X-CRW-Owner=crw-install), or is not a regular file |
unit_differs |
the installer's unit says something else; run --remove, then register again |
unit_name_taken |
the manager loads the name from another file, or it is masked |
unit_modified |
a <name>.d directory, drop-ins (also read once the unit is enabled and loaded, when the unit stays enabled and the exit status is 3), or a manager definition older than the disk (NeedDaemonReload), which disable would re-read |
unit_second_owner |
another unit in the directory runs the relay's service start or run |
unit_dir_in_runtime |
the unit directory lies inside the installer's destination, where removing a runtime would delete it |
unit_unreadable |
no runtime is installed, systemd could not be asked, or the relay cannot read its launch declaration |
It also reads what the boot start will find, through the pointer's relay in the unit's environment (service status), and
reports serviceEnabled, launchPolicySource and scopeAuthority. A disabled service intent, an undeclared execution policy
and an XDG_STATE_HOME the unit does not carry are warnings: the boot start would refuse service_disabled, withhold
role-bound deliveries, or write its host record elsewhere. It never changes the intent. Exit status 3 means a change may have
landed and a later step did not or could not be confirmed (written and not enabled: run it again; an enable or disable that
failed, possibly after changing some links: read systemctl --user is-enabled; deleted and daemon-reload failed: run that); 1 is a
refusal with nothing changed, 2 a usage error. --remove disables and deletes only a unit it wrote, only when the file is exactly
what the command writes (disable follows Also= into other units), the manager resolves the name to that file and it is not
running; it never stops the relay.
The maintenance order is the relay's own: service stop (systemctl --user stop crw-relay.service runs the same stop while the
unit is active, but a unit that never started has none to run), crw install update, which refuses while a daemon runs, then
systemctl --user restart crw-relay.service or service start. After a hand service stop the unit still reads
active (exited), so systemctl --user start does nothing and restart is the command. Measured with the unit's own commands: a
start with no App Server socket present succeeds and its worker stays up, so a start before the App Server is listening is not
refused; a start not ready within the relay's 20 seconds on a heavily loaded host is ended by the relay and leaves a partial store
(write-gate.lock without a database) that the next start refuses as store_owned_by_other, which is the relay's behaviour
and is in the refactor backlog.
crw doctor keeps alwaysActive at not_verified and names the default unit file it found or did not (a unit under another
name or directory is not read). A unit file is a registration; surviving a host restart is a measurement of one.
The plugin package declares the Stop hook that catches a managed turn ending without the records a
completion needs, and crw install hook --owner plugin writes the settings it reads,
<CODEX_HOME>/crw-completion-hook.json:
| Setting | Written as |
|---|---|
owner, configVersion, event |
plugin, 1, Stop |
relayExecutable |
<destination>/current/bin/codex-session-relay, or --relay-command |
mode |
observe (--mode), which classifies and records and never holds a turn |
timeoutSeconds |
the adapter's own budget (--guard-timeout, default 5), under the registered timeout (--timeout, default 10, at most 10) |
journalRoot, journalPolicy |
<CODEX_HOME>/crw-completion-hook/journal (--journal-root) and every_invocation |
markerRoot, dbPath, socketPath |
the relay's own resolution, or --marker-root, --db-path and --socket |
isolationAssertedBy |
required with --mode hold, and nothing else |
The decision is not made in the hook. The hook contract
fixes the rules and the relay's guard implements them (in the hook's own process, or in the store's
owner, which crw hook asks over control.sock); crw hook is the piece between the host and
that guard, and it exits 0 on every path, because exit 2 is the host's blocking code.
The settings are written before anything could read them, and every precondition is checked before
any write: the budget against the registered timeout, a hook file that already registers
this adapter for Stop (the user-owned registration, refused by name, because two registrations run
twice on every Stop), and settings already there that this command cannot act on. Settings that
already say something else are refused rather than overwritten, because they carry the mode: they
answer config_differs, naming the differing fields; move the document aside by hand to write these
flags' settings instead (one Stop settings document).
The one exception is a document written before decision 66. Those settings also recorded
adapterInterpreter /usr/bin/env and adapterEntryPoint <destination>/current/bin/crw-completion-hook
for the Python launchers the plugin package no longer carries. Nothing reads either key now: the Go
hook, the doctor and every move of the pointer accept a document that still carries them, so a host
keeps working on its current settings. A document that differs from the one this command would
write only by those two keys and installedBy is rewritten without them, answering
config_replaced with the keys it dropped under retiredFields (a --dry-run answers
config_would_create and says so). On the relay host, run crw install hook --owner plugin with
the flags the settings were written with, after the runtime carrying this change is installed, to
rewrite the document once. Any other difference is still config_differs.
The run registers nothing and leaves the hook file untouched. Written, registered and observed to
have fired stay three separate claims.
--socket is recorded as socketPath. It is what lets the guard tell that the state directory it
resolved holds another App Server's store: a store records the socket it serves, and a Stop hook that
inherits CODEX_SESSION_RELAY_STATE from a second installation would otherwise read that store, find
no relationship for its assignment, and hold a child that has finished. Configure it wherever more
than one installation shares a machine.
The marker root follows the relay's own resolution, including CODEX_SESSION_RELAY_MARKER_ROOT. A
default that skipped it would be a disagreement: the coordinator would publish its intents under one
tree while this hook looked under another, and every managed turn would read as unmanaged.
Codex runs a declared hook only once it has been trusted, and nothing in this repository grants trust. Trust is recorded against the declaration's content, so a plugin update that changes the hook command's text needs the hook trusted again, and Codex asks for it; until then nothing fires. An update of the runtime changes no declaration, because the command names the pointer, and so needs no new trust (plugin packaging).
A hook command is fixed when a turn starts, and every Stop of that turn reuses it; the next turn
resolves the declaration installed then (the turn-command cache).
A turn that started while the package still declared the Python bootstrap goes on running it:
${PLUGIN_ROOT}/wiring/crw_stop_hook.py first, and <CODEX_HOME>/crw-stop-hook.py once that version
directory is gone. crw install places no such launcher and leaves an existing copy exactly as it is;
the Python installer placed it. Either launcher reads the same settings and runs the adapter they
name, which is the Go hook once the settings are Go-era. The packaged launchers left the package in
todo 43, after the cutover commit: the version directory such a turn names is the pre-native one,
which the install that brought the native wiring removed, so the copy is what answers it. The copy
and the Python runtime stay while any live or resumable task can still run such a command, and the
operator removes them once the turns that could still run one have ended
(retention).
The plugin package declares this hook, and a host holding a second registration of it would run both
on every Stop: each leaves a row, and only the one that claims the event's accepted record asks the
guard (one accepted record per Stop event). crw install hook
writes for the plugin owner only; the user-owned registration, an entry in <CODEX_HOME>/hooks.json,
retired with the Python installer, and one already there is refused by name. The hook repeats the
check at run time, because installing the plugin is not a command this repository runs:
crw hook --plugin-launch stands down in silence unless the settings name the plugin as owner.
Only a verdict that agrees with itself is acted on. A verdict whose own decision releases while its
hook_output holds did not come from the guard, and any disagreement reads as
guard_verdict_incomplete and releases. Failures keep their own names: a runtime that could not be
run carries its errno, and an exit of 2 carrying the relay's own error record is the relay
declining a request it understood, while an exit of 2 carrying nothing is its argument parser
refusing before any command ran. Every one of these releases the turn and is recorded.
A turn can end more than once. When any Stop hook holds, the host appends a continuation to the same turn and fires Stop again; when a message was waiting, it appends that and does the same. Each of those is its own Stop event, and each one gets its own decision. Two registrations answering one Stop, or one Stop delivered twice, are a different thing: one event handled twice. The adapter keeps exactly one accepted record per event, asks the guard once for it, and still leaves a row for every invocation, so the two cases stay apart instead of being counted as one number per turn.
The host hands a Stop hook nine fields and no per-invocation identifier. None of them separates two
events of one turn: turn_id is kept across a continuation chain, stop_hook_active is false on a
turn's first Stop and true on every later one, and last_assistant_message can repeat word for
word. What differs is the transcript. Every sampling that ends in a Stop leaves one final answer,
and before running Stop hooks the host records it in the file named by transcript_path with the
turn id, the thread id and an item id of its own. So an event is
(session_id, turn_id, stop_hook_active, answer item id)
where the answer item is the one Stop of the turn, as the transcript shows it, that reported the
payload's text under the payload's stop_hook_active. It is established only when the latest
sampling's answer is recorded, exactly one Stop of the turn matches, and the matching answer's
thread is the delivered session. The key is a SHA-256 over the four values, so no host value becomes
a path component and nothing is minted per invocation.
The transcript is read backwards from its end to the turn's start, bounded at 64 MiB and 0.75 seconds, inside the margin the launcher keeps over the guard budget. A missing or unreadable transcript, a scan that hits either bound, an unfinished last line and any failed condition leave the identity unestablished, with the reason on the row. An unestablished invocation is asked about exactly as before and is never deduplicated.
A claim is two create-once files. The first is the host's,
<CODEX_HOME>/crw-completion-hook/stop-events/<key>.json: every registration the host starts for
one Stop inherits that Codex home, while the settings each one reads, and so its journal root, may
differ, so this is the file two registrations of one host always meet. Only the invocation that
creates it owns the event, and only the owner asks the guard. The owner creates the accepted record,
<journalRoot>/accepted/<key>.json; asks the guard; writes its row; and then writes
accepted/<key>.outcome.json naming the session, the turn, the outcome and the row. An invocation
that finds either file already there asks nothing, prints nothing and writes a row whose
adapterOutcome is duplicate_invocation. When the host's file can be neither created nor found,
nobody can own the event, so nobody asks: the invocation writes an arbitration_failed row and
releases the Stop, the adapter's ordinary failure direction.
Rows are <journalRoot>/<YYYYMMDD>/<32 hex>.json, one JSON object on one line with sorted keys, and
every count of them counts invocations. A row carries sessionId, turnId, adapterOutcome,
eventKey, eventIdentity, acceptance (accepted, duplicate, unestablished, unclaimable,
claim_failed or unarbitrated), acceptedAs and guardInvoked. faults_only leaves out the rows
of answered and duplicate invocations of identified events; no_journal keeps no rows at all.
The outcome record says what the adapter answered, not what the host received: it is written before the answer is printed, exactly as the row is. A claimant killed between its claim and its outcome leaves a claim with no outcome, and the event still has its owner, so a later delivery of it is a duplicate and the Stop was released.
crw-dev stop-events --journal-root <root>, in the repository's development binary, reads the rows,
the accepted records and the host ledgers the claims name, and answers one verdict (the same reading
as scripts/stop_events.py, deleted in todo 44, gave). FALSE (exit 1) means an event was
accepted more than once; UNREADABLE (exit 3) means the reading cannot vouch for what it read, and
names why; TRUE (exit 0) otherwise.
--session and --turn choose the events of one turn, --since and --until a window in the
records' own YYYY-MM-DDTHH:MM:SSZ format, --journal-root repeats for every root the host's
registrations write to, and --codex-home adds a host whose ledger no claim names yet. The
invocations it cannot judge, whose identity was not established or that had no owner, are counted by
reason and never read as answered.
When the window contains readable version-2 rows with no event key, the optional
excludedInvocations list names each row, timestamp, session and turn, and its
unestablished:<reason> or no_event:<outcome> reason. evidence holds the recorded
acceptance, event key, accepted path, adapter outcome, guard invocation/decision,
hold and event identity. excludedFrom: per_event_acceptance_count means only that
the row cannot join an event's acceptance count; it does not remove uncertainty
from the window. preventsTrue is true except for the existing native pre-scan
unreachable exemption. An unestablished guard call still makes the verdict
UNREADABLE, even if it wrote no claim. Keyed unclaimable or failed claims remain
in unjudgedInvocations; malformed rows remain in rowsUnreadable.
For no_event:stdin_unreadable, evidence.stdinRead adds the recorded detail and elapsedMs
and a case naming what the row establishes. A row written since CRW-504 carries its own account
of the failed read in stdinRead (below) and its case is the cause that account names; a row
without it is read as before.
| Case | What it establishes |
|---|---|
input_late |
The 100 ms input allocation ran out while the hook waited for the host to supply the payload. |
work_ended |
The invocation's own work context ended while the hook waited (a settings budget shorter than the allocation, or a cancelled run). |
read_error |
The stdin descriptor, or the reader standing in for it, failed; error holds the Go error text. |
invalid_utf8 |
The adapter read bytes but could not decode them as UTF-8. A row without stdinRead is recognised by its detail. |
read_failed_cause_unrecorded |
A row without stdinRead: reading stdin failed and the adapter discarded the underlying cause. |
detail_unrecognized |
A row without stdinRead whose detail matches no known branch; retain it without assigning a cause. |
stdinRead is an optional object on stdin_unreadable rows, added beside the existing keys the way
runtime was; rows written before it stay valid. It has five keys: cause (input_late,
work_ended, read_error or invalid_utf8), error (the Go error text of the failed read, an OS or
context message about the stdin descriptor with no payload content, and for a recovered panic only
its Go type), bytesRead (the bytes the completed reads reported before the failure was recorded,
which counts what arrived and not what the host meant to write, and is a lower bound when the
reader was still running), and waitStartedMs and waitEndedMs (when the wait began and ended, in
milliseconds from process entry, the origin of elapsedMs). The judge reports error, bytesRead,
waitStartedMs and waitEndedMs beside the case. A row whose stdinRead is not that closed shape,
whose cause disagrees with its detail, or that carries it on another outcome is unreadable.
These diagnostics change no count, reason, exit code or verdict. Other outcomes omit stdinRead;
malformed rows never receive it. A row without stdinRead cannot distinguish a host that supplies
bytes or EOF late from a descriptor error or the work deadline. In particular, elapsed time near
100 ms does not prove a timeout, a restart, or which session was involved. Correlate such a row with
host stdin write/EOF/error logs and rollout evidence; when those are absent, report the cause as
undetermined rather than attaching the nearest turn by timestamp. A row with stdinRead says which
of the four happened and how many bytes had arrived, but not why the host was late; the session and
turn are unknown before the payload parses and are not recorded.
Host input that remains incomplete when the adapter must wait beyond its 100 ms
allocation is the intended late-input release case in decisions
24
and 32.
Already-readable input is still taken. There is no separate stdin byte cap.
Empty input whose writer closes normally reaches JSON parsing and is
stdin_not_json, not stdin_unreadable: a clean end of input is not a read failure and never
carries stdinRead.
The guarantee holds among the registrations of one Codex home, which is every registration one host
starts for a Stop. The claim files are never removed, like the rows. That the host records the answer
before running Stop hooks was observed in every isolated run and is consistent with every record in
the live journal, but it is not a documented host contract. A turn whose Stops repeat the same text
trades deduplication for safety: those Stops are asked about by every registration, and a reading of
a window that holds them is UNREADABLE.
Written settings, a declared hook, a trusted hook and a hook that fired are four different facts, and the only evidence of the last one is a record of this turn. End a turn in an ordinary task started after the change, take its session and turn ids, and look for them in the journal:
journal="${CODEX_HOME:-$HOME/.codex}/crw-completion-hook/journal"
grep -l '"sessionId": "<session>"' "$journal"/*/*.json | xargs -r grep -l '"turnId": "<turn>"'A row naming that session and turn is a callback this procedure can attribute; its adapterOutcome
says what the adapter did. No row is not a smaller number of callbacks: the turn did not reach this
hook's journal, and the cause is one of these, read in order:
| Cause | How to tell |
|---|---|
| The hook is not trusted | Codex has not been asked to trust this declaration since the command text last changed; nothing fires until it is |
The pointer names no runtime that reads --plugin-launch |
crw doctor: runtime.state, runtime.kind (anything but go-binary has no bin/crw), and a Go build older than decision 26 (the native wiring) |
| No usable settings | crw doctor: settings.crw-completion-hook.json.state; the hook releases in silence without settings, with settings it cannot read, and with settings another owner holds |
| Journalling is off | journalPolicy is no_journal, or faults_only and nothing faulted |
| A different journal | the settings name another journalRoot than the one you read |
A subagent's turn is no signal: on the measured host subagent turns recorded no Stop at all. The
daemon is not on this path. The guard reads the marker and a read-only database, so a stopped daemon
is not observable from a Stop and is never inferred from one. runtime_install.py hook-status, the
Python installer's cell-by-cell reading of the same question, has no crw counterpart
(the Python fence installer); its Go port was deleted unused in
wave R1 (decision 57 in the port decisions).
Installing, updating and hooking each have their own rules above. What none of them states is the
sequence a host actually lives through, with the state that has to survive it put there before the
first install and read again after the last refusal. The Go installer's tests run that sequence
against temporary homes: install, the same install again, an update that fails at a step and puts
everything back, rollback and remove (internal/runtime/install, lifecycle_test.go,
decisions_test.go and restore_test.go), and the native Stop command and the launchers through
the pointer (wiring_test.go; the retired Python launchers from the pre-native testdata).
The composed run asks seven questions and forbids one reading standing in for another. A reading
that could not be made is unreadable, never false, and never the value of the reading beside it.
| Question | Answered by | Read from | Never established by it |
|---|---|---|---|
| Skill link | crw-dev skills link --check, on a linked host |
its report | that a linked skill is loaded or trusted by a host. A plugin host has no links; codex plugin list is its reading |
| Runtime reach | crw doctor |
runtime.state, runtime.kind, runtime.agrees and each components.<name>.class |
that the runtime works for a task |
| MCP tool exposure | a fresh Codex task | the bridge tools it lists, and get_capabilities |
anything about a task started before the change |
| App Server connection | crw doctor |
checks.results.connected |
that a socket that accepted a connection will accept delivery |
| Real hook callback | the journal | the rows naming the turn you ended | that a firing was judged correctly |
| Model and permission preservation | a byte comparison of config.toml |
the file before the run and after the crw commands |
that anything a Codex process writes later was preserved |
| Delivery acceptance | a trial | checks.results.deliveryAccepted from a run that created and delivered work |
that an accepted delivery was acted on |
A real combination is an operator action, not a check. It needs a home, a Codex home, a host record
and a state directory that are yours to change, and it establishes nothing until it is recorded. The
destination is $HOME/.local/share/crw-runtime of the HOME the commands run under, and the live
half reaches this runtime only through the plugin's declared commands, so run the block under the
HOME the Codex process runs with. The block names the host record and the Codex home on every
command, and the state directory wherever one is read, because none of them derives from another:
the Codex home is where the settings the plugin reads are written, and the host record defaults to
$XDG_STATE_HOME/codex-relay-workflow/host-record.json, whatever the Codex home. <codex-home> is
the Codex home that process reads. The Go build serves a store in that HOME's default relay
state directory, which the live-state guard
refuses only under test isolation, so only a host that still runs the Python runtime waits for
the cutover.
# Substitute every <...> below before running any of it. They are placeholders, not literals, and
# an unsubstituted one is a shell redirection rather than a value.
# A NEW receipt directory, so no earlier run's files are read as this one's. mkdir without -p is
# the check, and it ends the procedure rather than running the rest into a directory it refused.
mkdir <receipt> || exit 1
# Preservation is a comparison. No crw command below writes config.toml, so the whole file is
# compared rather than keys guessed out of it. A fresh Codex home has none; record the absence.
if [ -f <codex-home>/config.toml ]; then
cp <codex-home>/config.toml <receipt>/config.before.toml
else
printf 'no configuration existed before this run\n' > <receipt>/config.before.absent
fi
# Nothing puts crw on PATH. The install runs the copy unpacked from the release archive into
# <scratch> (see "Installing the runtime" above); every later command runs the one the pointer
# then selects.
# The exit status belongs IN the receipt: a receipt that kept the result and lost the status
# cannot say whether the install refused, or that a 3 means the change landed.
<scratch>/crw install install --release <tag> --record <record> \
--codex-home <codex-home> --state <state> --socket <socket> > <receipt>/install.json
printf 'install exit=%s\n' "$?" > <receipt>/install.exit
crw="$HOME/.local/share/crw-runtime/current/bin/crw"
"$crw" install register-mcp --owner plugin --record <record> \
--codex-home <codex-home> --execution-policy <policy-file> > <receipt>/register-mcp.json
printf 'register-mcp exit=%s\n' "$?" > <receipt>/register-mcp.exit
"$crw" install hook --owner plugin --record <record> \
--codex-home <codex-home> --socket <socket> > <receipt>/hook.json
printf 'hook exit=%s\n' "$?" > <receipt>/hook.exit
"$crw" doctor --record <record> --codex-home <codex-home> \
--state <state> --socket <socket> > <receipt>/doctor.json
printf 'doctor exit=%s\n' "$?" > <receipt>/doctor.exit
# The other half of the preservation reading, taken before anything else can write the file.
if [ -f <codex-home>/config.toml ]; then
cp <codex-home>/config.toml <receipt>/config.after.toml
cmp <receipt>/config.before.toml <receipt>/config.after.toml > <receipt>/config.cmp 2>&1
printf 'cmp exit=%s\n' "$?" >> <receipt>/config.cmp
else
printf 'no configuration existed after the crw commands\n' > <receipt>/config.after.absent
fi
# ... then, in Codex: trust the hook if it asks, start a fresh task, read the bridge tools and
# get_capabilities there, end a real turn, and note that turn's session and turn ids. Only then:
journal=<codex-home>/crw-completion-hook/journal
grep -l '"sessionId": "<session>"' "$journal"/*/*.json | xargs -r grep -l '"turnId": "<turn>"' \
> <receipt>/rows-for-this-turn.txt
printf 'rows grep exit=%s\n' "$?" >> <receipt>/rows-for-this-turn.txt
# From a checkout, the exactly-once reading for that turn:
go run -tags dev ./cmd/crw-dev stop-events --journal-root "$journal" --codex-home <codex-home> \
--session <session> --turn <turn> > <receipt>/stop-events.json
printf 'stop-events exit=%s\n' "$?" > <receipt>/stop-events.exitWhat this block is, and what it is not. It installs, writes both records, takes the readings and reads the results back. It does not re-run the install, it does not present a second archive, and it does not fail an update; this page will not tell an operator to break a runtime their host is using in order to watch it come back. The tests above exercise those stages against temporary homes.
cmp exiting 0 is preservation. A difference means something wrote the file between the two copies,
and the receipt holds both sides for reading which keys moved. rows-for-this-turn.txt naming no
file is a turn that did not reach the hook (registration is not firing),
not a smaller count. More than one row for the turn is not a duplicate by itself: a turn whose Stop
was held ends again, and that is a second event; stop-events.json separates the two.
Three of the seven are readings of something live, and the block supplies none of it: an App Server
accepting connections at <socket>, a session that actually listed the bridge tools, and a relay
that can carry an assignment to a returned turn id. crw doctor reports deliveryAccepted as
not_applicable because it creates no work; a live trial (live-trial.md) is how a
host gets that reading. If the live half is absent, the honest receipt records the absence for that
row and says the rest. It never carries a row forward as though the question had been put.
A real run records the exact release it installed, the host it ran on, the destination kind, and the answer to each of the seven with the command that produced it and the time it was produced. Those receipts are host facts: they belong in the private record outside this repository, not in a commit. This page is the procedure and the shape. It is not a record that anybody ran it.
scripts/runtime_install.py installed the Python runtime, the fence release that step 0 of
the cutover deployed, and it
was the rollback path until the cutover committed (todo 43). Todo 44 removed it with the rest of the
Python execution path. It shared the destination, the owned pointer, the host record, the promotion
lock and both settings files with crw install; what it installed was a Python virtual environment,
<destination>/env-1-<digest>, built from the checkout's packages, and it took --dest, which is
why crw install still refuses a host record whose pointer names another link. Its subcommands were
install, diagnose (with --trial), register-mcp, hook (which also placed the fallback
launcher <CODEX_HOME>/crw-stop-hook.py), hook-status and verify-definition; crw install and
crw doctor carry what a Go install needs of them.
Two properties of a Python runtime outlive this installer, and the cutover's retention rule rests on
them. pip writes an absolute shebang into every console script, so a process started through
current reports and keeps its concrete env-* directory after the pointer moves; and a Stop
command fixed before the native wiring names python3 and a .py launcher. Both kept env-*
directories and the <CODEX_HOME>/crw-stop-hook.py copy in place until nothing could still resolve
to them. The packaged launchers did not wait for that: such a command
names the pre-native version directory, not whatever the package ships now, so todo 43 retired
them from the package after the cutover commit.
runtime_install.py diagnose --trial was the only command that filled deliveryAccepted itself: it
registered one relationship, emitted and delivered once, and recorded the returned turn id. It has
no crw counterpart and left with the installer; a live trial is a different
thing with a similar name.
The full reference this page carried for the Python installer, its design record and its acceptance
procedure, is this page at the parent of the commit that rewrote it for crw, the oldest one that
names todo 41 (git log --format=%H --grep='(todo 41)' -- docs/runtime-install.md | tail -1).
Running these commands against a temporary destination proves what they did there. It is not
evidence about a host's real Codex home, its installed runtime, its bridge record or its operational
database. installed, mcpExposed, connected, deliveryAccepted, verificationComplete and
alwaysActive are six separate facts under OPS-6.1, and none of them is read from another.
Written settings are not a fired hook, and a fired hook is not a delivered hold. That a settings file is there says nothing about the host having run the hook, about the runtime it names being able to answer, or about any turn having been judged. Those claims need the host's own evidence.
A successful update is not one of them either. That the pointer moved, that the gate found the daemon stopped and no attempt open, and that the store's schema was compatible are readings taken at one moment, about one destination. They say a swap was permitted and performed; they do not say the new runtime works, and the point that would say so is measured before the swap rather than after it. Nor does a refused update establish that a store is healthy: the gate reads whether it is safe to replace a runtime, and reads nothing about whether the data in the store is correct.