Skip to content

feat(spawn): support config/claude-launcher and mirasim process classification - #1

Closed
Arobotmaster wants to merge 38 commits into
mainfrom
fm/fm-fork-mirasim-live-delivery
Closed

Arobotmaster wants to merge 38 commits into
mainfrom
fm/fm-fork-mirasim-live-delivery

Conversation

@Arobotmaster

Copy link
Copy Markdown
Owner

Summary

Support config/claude-launcher to allow wrapping Claude Code invocations with mirasim claude, enabling dynamic local proxy model execution without altering the core Claude harness identity, hooks, or control plane.

Key Changes

  1. bin/fm-spawn.sh: Read and validate config/claude-launcher (supporting claude and mirasim); wrap worker invocations with "$MIRASIM_BIN" claude for crewmate/scout tasks when set to mirasim.
  2. bin/fm-agent-process-lib.sh: Classify mirasim process as agent.
  3. AGENTS.md & docs/configuration.md: Document config/claude-launcher schema and semantics.
  4. tests/fm-claude-launcher.test.sh: 7 regression unit tests verifying launcher configuration, validation, process classification, and secondmate exclusion.

Live Verification Evidence

  • Verified end-to-end via an isolated Herdr lab session (fm-lab-fm-fork-mirasim-*) and clean test workspace.
  • Executed managed fm-spawn.sh with claude-opus-5-5 via Mirasim.
  • Confirmed pre-registered workspace trust, .claude/settings.local.json hooks execution, hook signals updating agent busy-state, and Mirasim traffic session recording HTTP 200 model invocations.

kunchenguid and others added 30 commits September 22, 2026 09:24
* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue kunchenguid#3793).
The original 0.25s window after confirmation was widened to 80 polls in
kunchenguid#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
…nchenguid#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
* Add verified Devin CLI worker adapter

* no-mistakes(review): Drop Devin resolver refusal and launch marker

* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs

* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts

* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals

* fix(control): never pair Devin interrupt presses on an idle agent

A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.

* fix(devin): disable Claude hook import and commit attribution for workers

The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.

* test(devin): extend live guard and record Herdr and revert-picker evidence

The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.

* no-mistakes(document): Correct Devin documentation links and lifecycle guidance

---------

Co-authored-by: Denis Beliaev <battler73@yandex.ru>
…id#5322)

fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.

Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.
…nguid#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning
…kunchenguid#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
…chenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
kunchenguid#5362 and kunchenguid#4878) (kunchenguid#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859.
…enguid#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing
…unchenguid#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg
kunchenguid#5391)

The listener resolves its server from that store before it polls. Without a session for this board, it exits in the gap after the build has already sampled a live claim.
…#5390)

* fix(bin): prune a torn-down task's wake rows at teardown

Prune pending durable wake rows (.wake-queue) for a task when it is torn
down, clearing stale wakes for its target window, signal wakes for its status
or turn-ended files, and task-specific check wakes.

Fixes kunchenguid#3419.
Adjacent to kunchenguid#5252.

- bin/fm-wake-lib.sh: add fm_wake_queue_prune_task
- bin/fm-teardown.sh: call fm_wake_queue_prune_task in cleanup_firstmate_home_children and main teardown
- tests/fm-wake-queue.test.sh: add test_wake_queue_prune_task

* no-mistakes(document): docs: note teardown prunes a task's wake rows

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
…ixture readiness (kunchenguid#5392)

* Make portable tests match resolved host paths

Summary:
- Match macOS full Node command paths by basename in the Gemini behavior test.
- Mirror symlink-resolved Nix PATH behavior and give the loaded-host race bounded headroom.

Testing:
- bin/fm-lint.sh
- bin/fm-test-run.sh tests/fm-on.test.sh tests/fm-gemini-harness.test.sh tests/fm-procevent.test.sh

Related:
- None

* no-mistakes(review): Mirror production PATH helper rules per directory group in tests

* no-mistakes(test): Wait for orphan runner start marker instead of fixed sleep

* no-mistakes(document): Clarify gemini ancestry test comment for versioned node comm

---------

Co-authored-by: Sandeep Salwan <salwansa@amazon.com>
…ns (kunchenguid#5389)

The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation.

Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427)

Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358)

* feat(bin): add an opt-in per-home worker account pin

A home that mixes work and personal accounts for one runner had no way to
say which account its workers launch on: Claude workers inherited whatever
CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient
root, and an ambient API key outranked both, with no signal at launch.

config/claude-account and config/pi-account now pin that choice per home.
With neither file every launch is unchanged. With one, every launch of that
runner from the home (ship, scout, local secondmate, raw Claude command, and
relaunch) runs under the declared root, and the spawn refuses before any
endpoint exists when the file is malformed or the runner's own check
(claude auth status, pi auth check with a model-listing fallback) says the
pinned account is not signed in. The check runs in a cleared environment so
an ambient credential cannot answer for an empty root. A pinned Claude launch
sheds the environment credentials Claude ranks above a stored login; a pinned
Pi launch needs an explicit <provider>/<id> model for a declared provider and
also carries --provider. The chosen account is printed on the spawned line and
recorded in the task record, and relaunch checks the pin before stopping the
running agent.

* test(secondmate): give the concurrent config-push wait room for a slow host

test_config_reread_serializes_concurrent_pushes waited about two seconds for
the first fm-config-push.sh to reach its first send-keys. On a slower host
that push takes four to five seconds, so the test failed on main before the
push ever got there. The loop still leaves as soon as the marker appears, so
the larger bound costs nothing where the push is fast.

* no-mistakes(review): Refuse raw Claude account overrides under a pin
yasuhito and others added 8 commits September 23, 2026 15:50
…ts (kunchenguid#5470)

* fix(bin): keep the Herdr lab session option before a -- delimiter

fm-herdr-lab.sh run appended --session <lab> after every argument, so a
command with a passthrough delimiter such as agent start ... -- <agent args>
handed the session flag to the agent and Herdr routed the call by the
caller's ambient socket instead of the lab.
The helper now inserts --session <lab> immediately before the first --
delimiter and keeps the trailing form otherwise.

* no-mistakes(document): Clarify Herdr lab session option placement
* feat(bin): guard the partition, harness pin, and bounded exec for a non-Pi supervision host

Lease liveness is now the pure record test in every calling context, so an
unmarked main honors a live branch lease held by a separate process, and a
lease file engages the guard's claim serialization for any caller; a home
with no lease files still takes no lock.

bin/fm-harness.sh honors FM_SUPERVISION_PRIMARY_HARNESS while
FM_SUPERVISION_ACTOR=branch, so a supervision branch running under another
harness resolves own, crew, and secondmate to the primary's harness.

fm_tasks_axi's watchdog moves into bin/fm-timeout-lib.sh as fm_exec_timed with
a separate grace: the perl watchdog is preferred, runs the command in its own
process group against wall-clock deadlines, forwards TERM/INT/HUP, and reaps
the group, so a descendant holding captured output can no longer keep the
caller waiting past the bound on a host without timeout.

The Claude Stop auto-arm header records that Claude drops the exit 2 of a hook
it terminated at the configured timeout, re-measured on Claude Code 2.1.281.

* fix(bin): state that fm_exec_timed cannot reach a descendant in its own process group

Live runs of real Claude and Pi engine turns under the bound showed both CLIs
start every tool command in a process group of its own, so those processes end
through the engine's own TERM handling rather than the group signal or reap.
Also clears the new timeout test's ShellCheck findings.

* no-mistakes(document): Clarify cross-harness lease documentation
…henguid#2648)

* feat(bin): make the ship-branch prefix configurable per project

fm-brief.sh hardcoded every generated ship branch to fm/<task-id>, which
leaks that firstmate produced the branch/PR - unwanted for a third-party
public repo that does not use this tooling.

Add an optional --branch-prefix flag to fm-brief.sh (default "fm/", so
existing installs are unaffected) and teach fm-project-mode.sh - the
registry's single-owner parser - to resolve a project's optional
"branch=<prefix>" data/projects.md annotation via a new --branch-prefix
query, order-independent with the existing mode/+yolo tokens. Firstmate
resolves the override at task intake and passes it explicitly, mirroring
how --mode already works; fm-brief.sh itself never reads the registry.

An empty override resolves to a bare "<task-id>" branch rather than a
leading slash. All five previously hardcoded fm/$ID sites (branch
creation, never-push rule text, definition-of-done text, and the status
message) now render the resolved prefix consistently.

* no-mistakes(review): Wire branch-prefix intake in AGENTS.md; fix fm-merge-local.sh hardcoded fm/ prefix

* no-mistakes(document): docs: document configurable ship-branch prefix in architecture.md

* no-mistakes(review): Persist immutable branch contracts

* no-mistakes(document): Document configurable ship branch prefixes

* no-mistakes(lint): Captain: fix ShellCheck test warnings

* fix(bin): map bearings PR rows to their recorded ship branch (kunchenguid#1887)

fm-bearings-snapshot.sh keyed a PR back to its task by string-matching
the headRefName against the fm/ prefix, so any project whose branch
prefix was overridden (e.g. via kunchenguid#2648's branch=<prefix> registry
annotation) had its PRs silently drop to task "-" in the bearings
view, exactly the third fm/-assumption issue kunchenguid#1887 named alongside
fm-merge-local.sh and fm-bearings-snapshot.sh itself.

fm-fleet-snapshot.sh now surfaces each task's recorded branch=
metadata field in its JSON task rows, and fm-bearings-snapshot.sh
cross-references a PR's headRefName against those recorded branches
before falling back to the legacy fm/ prefix heuristic, so a custom
branch prefix maps a PR back to its real task.

Adds a regression test proving a PR opened against a fix/<task-id>
branch resolves to that task instead of "-"; confirmed it fails on
the prior startswith("fm/") logic and passes with this change.
ShellCheck clean; full fm-bearings-snapshot.test.sh and
fm-fleet-snapshot-view.test.sh suites pass.

* fix(ci): align lint arithmetic-looking assignment and stale Bearings snapshot count

- Quote the --branch-prefix want_value assignment in fm-brief.sh, fm-promote.sh,
  and fm-spawn.sh so ShellCheck SC2100 no longer misreads the plain string
  'branch-prefix' as arithmetic shorthand.
- Bump the Stock macOS Bash snapshot job's hardcoded Bearings test-count
  assertion from 59 to 60: this PR added a Bearings test, so the count was
  stale, not the feature.

* no-mistakes(review): fix(bin): honor recorded ship branch in relaunch and review-diff

* no-mistakes(document): docs: complete branch-prefix flag in brief and promote headers

* fix(lint): quote branch-prefix parser token; drop unused BRANCH_Q after rebase

* no-mistakes(review): Restore %q branch escaping in promotion instructions with regression test

* no-mistakes(document): document recorded ship branch and prefix flag

fm-review-diff.sh's header is the owner of its branch-resolution
contract; it still described only the legacy local-branch behavior
after the change made review-diff honor state/<id>.meta's recorded
ship branch. README's feature bullet enumerates the registry's
optional flags and was missing the new branch=<prefix> override.

* no-mistakes(lint): Silence SC2016 on intentional single-quoted sed expression

* no-mistakes(review): address branch-prefix review findings in DoD and project-mode

* no-mistakes(test): branch-prefix suites pass under tasks-axi 0.2.6; environment-only failure

* no-mistakes(document): purge stale fm/ branch naming from docs and headers

* fix(test): assert the merged epoch status wording in the branch-prefix override test

The rebase resolution of tests/fm-brief.test.sh kept the branch's
pre-merge \`done: ready in branch ...\` assertion while the merged
fm-dod-lib.sh (carrying main's epoch-stamped status line) renders
\`done [at=<epoch>]: ready in branch ...\`. Align the assertion so the
override-consistency test matches the behavior it verifies.

* no-mistakes(review): Address remaining branch-prefix findings in four bin scripts

* no-mistakes(test): skip real-tasks-axi tests below the repo's 0.2.6 floor

* no-mistakes(document): document spawn's branch-prefix registry deviation notice
…d per-rule confidence floors (kunchenguid#5478)

* feat(bin): send dispatch resolver only the brief's task sections

* Sent Jev only the scaffolded Captain's intent and Firstmate spec
  sections, falling back to the whole brief when neither heading is
  present, so the identical setup, rules, and definition-of-done
  boilerplate no longer reads as a signal about the task
* Added an optional per-rule min_confidence that replaces the global 0.6
  floor for that rule; a picked rule below its own floor falls to the most
  probable other option that clears its floor, or returns ambiguous
* Kept files with no declared floor on the exact previous behavior and
  kept the model blind to the new field
* Recorded the live old-versus-new comparison over scaffolded fixtures

* no-mistakes(review): share brief heading parser, add kind line, fix floors

* no-mistakes(test): stop sending ship delivery mode to jev, keep scout tag

* no-mistakes(document): docs: list shared brief heading lib in scripts inventory
* feat(bin): supervision host core behind config/supervision-host

Add the supervision host (bin/fm-supervision-host.sh): beside a Claude
primary it owns the watcher cycle for the Stop auto-arm and, while the
away-posture record exists, hands each wake to a bounded headless Claude
engine session that runs the supervision branch's contract - the same
generated prompt, row eligibility, wake grant, per-actor drain, outcome
store, leases, and away relocation the Pi branch uses. Attended wakes pass
straight to main. Every path that cannot finish a wake hands it to main
with a supervision-host line; the park ends itself before the Stop hook
timeout with a cycle-boundary wake.

- bin/fm-supervision-engine-lib.sh: opt-in parse, verified engines
  (claude, default sonnet), one bounded engine turn, and a reap of engine
  tool processes that sit in their own process groups.
- bin/fm-branch-report.sh: the command twin of fm_branch_report, scoped
  to the tasks the current host turn claimed.
- bin/fm-branch-dispatch.mjs: command entry to the Pi dispatch module, so
  eligibility and the wake prompt have one owner.
- bin/fm-claude-stop-autoarm.sh runs the host in the arm's place when
  config/supervision-host exists; nothing changes without the file.
- bin/fm-watch-arm.sh --stop: home-scoped stop without a re-arm.
- bin/fm-lease-lib.sh: an opted-in home takes the lease-command lock for
  unmarked main too, closing the first-claim race; the refusal tells the
  caller to leave the lease alone and retry.
- /afk launches no away daemon on an opted-in Claude home; /quiet still
  does. Session start renders the host's main-side protocol there.

* fix(bin): relay a host turn's outcomes when the captain returns mid-turn, and log per-turn engine cost

Live validation found two supervision host gaps. A captain who returns while
an engine turn is running gets a return brief rendered before that turn's
outcomes exist, so the host now hands the close to main with those outcomes.
Claude reports a resumed conversation's running cost, so the engine lib now
derives each turn's cost from the total the host records, and the host log
records every close's destination.

* docs(verification): record the supervision host's live evidence

The dated live results behind docs/supervision-host.md: the Claude engine's
live guard, the away-wake cases against real workers, the engine's cost
reporting, and the flag-off before-and-after regression.

* docs: describe the supervision host ledger as covering every close

* no-mistakes(review): Harden supervision host ownership, boundary, ack, and late outcomes

* no-mistakes(review): Recheck park boundary just before starting an engine turn

* no-mistakes(review): Cap park boundary, deliver all host lines, reject incomplete results

* no-mistakes(document): Correct supervision host documentation and stale pointers
…t-in, and delivery (kunchenguid#5506)

Attestation MATCH; contract-class restore; CI/NM green. Squash-merged by Kun's firstmate.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.