Acknowledge node send before typing it; one zmx ls per presence pass - #582
Merged
Merged
Conversation
node send exited 75 ("may still have been applied") for messages that
landed: the daemon typed the whole message (a zmx ls gate over every
session, each chunk, the submit beat, Enter) before acknowledging, and on
a CPU-starved machine that outlasted the CLI's 10s wait.
- messageNode is acknowledged once the message is on the board and queued
for its session; typing runs afterwards on a per-target chain that keeps
send order, and follow-ups wait behind it. A failure after the ack is
staged to the loop's memory and logged, not broadcast, so no other
client takes it for its own verdict.
- One zmx ls per presence pass (ZmxSessionLauncher.SessionListing),
shared with sends and joined while in flight, instead of a full listing
per node per read. Starting or killing a session invalidates it, and a
"not alive" that would refuse a send or allow a resolution is confirmed
against a fresh listing.
- runZmx awaits exit through terminationHandler and reads on a GCD
thread instead of blocking the cooperative pool.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
The shared listing can be older than a task's end, and a session whose task ended is a husk whose wrapper shell would take the keystrokes. The send gate now always takes a fresh zmx ls (typing runs after the ack, so this costs the sender nothing) and types only into a task zmx reports alive. A row with exit_code= but no ended= counts as ended too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: scgopi <scgopireddy@gmail.com>
- A composite child's store lives for one command, so it types inline instead of on a chain no later send or parent could see (order kept). - Each post-ack typing is bounded by deliveryDeadline; a hung one is staged and logged, and the loop's follow-ups and wakes move on. - Every outcome is logged against the acknowledged request: send-typed, send-staged (delivery-failed, session-gone, deadline), send-dropped. - Lifecycle checks (start/terminate results, a resume that may have died, the first-pass kickoff, isSessionAlive) take fresh listings; a hung in-flight listing is joined for at most 30s. - A send whose drain changes nothing is broadcast once, not twice. - Tests: composite-child ordering with a slow transport, a hung typing, one broadcast per send, and a real zmx husk that send refuses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: scgopi <scgopireddy@gmail.com>
withDeadline cancels the typing it gives up on but cannot stop it: the cancelled send read as a failure, so the abandoned task respawned an unattended loop, typed again behind the chain's next message, and staged the message a second time. It now returns as soon as it sees it was cancelled; typeLogged has already staged it once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: scgopi <scgopireddy@gmail.com>
A deadline that fires during the respawn settle returned the sleep at once, and the retry then typed again inside the cancelled task. The check now follows the settle as well as each delivery attempt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: scgopi <scgopireddy@gmail.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Ends the false exit-75
node sendtimeouts: the daemon now acknowledges a send once it is recorded and queued, and types it afterwards. Presence passes also stop running one fullzmx lsper node.Why
DaemonTimeoutTriage measured it on 2026-10-01: a send landed and still exited 75. The daemon typed the whole message before acknowledging it: a full serial
zmx lsover all 27 sessions, then each chunk, the 400ms beat, then Enter. Every unreachable session costszmx lsa 1000ms probe timeout, so on a starved machine the request ran 10–30s against the CLI's 10s wait.Timing
Rig: an isolated graphcoded (private
GRAPHCODE_SUPPORT_DIRandZMX_DIR; the live daemon was never touched), with 20 goal loops whose sessions runcatbehind a fakeclaude, and one client held connected so presence passes run.Builds: main =
acdb2a62, whose daemon sources are identical to this PR's base. PR =0da20eaa.Method:
node sendruns per build.handle_mscomes from that daemon'sgraphcoded.log.zmx lscalls per pass were counted with a loggingzmxshim.Load: the Edge storm had ended by measurement time (load 4.5), so the load is synthetic: 60 busy processes, load average 54–72 on 10 cores.
A: load only, all 20 sessions reachable
node sendwall p50 / p90 / maxhandle_msp50 / p90 / maxzmx lsper presence passB: same load, plus 2 session daemons SIGSTOPped
This is the field's failure mode: their rows read
err=Timeout status=unreachable, and eachzmx lstakes 2.0 s.node sendwall p50 / p90 / maxhandle_msp50 / p90 / maxzmx lsper presence passSynthetic CPU load alone did not reproduce the field's multi-second
zmx ls; unreachable sessions do. On main, every send pays onelsof about 1 s per unreachable session, so four or more cross the CLI's 10 s deadline. The PR's acknowledgement does not depend on anyzmxcall. Its wall time excludes the typing by design. The typing still runs after the ack, gated by a listing of its own.Changes
GraphStore).messageNodebroadcasts its verdict once the message is mirrored to the Mailroom and queued, before the settle drain.--follow-upto the same loop waits behind a send that is still being typed.deliveryDeadline. A typing that hangs is staged to memory, and the loop's queue moves on.conn/seq), assend-typed,send-staged(reasondelivery-failed,session-goneordeadline) orsend-dropped(target deleted while queued)..errorOccurred, because that reaches every connection and the next waiting CLI would take it as its own verdict.accepted — typing it in nowinstead ofdelivered. The remote Python shim matches.zmx lsper pass (ZmxSessionLauncher.SessionListing).ProjectRegistry.pollPresencebrackets each tick as one pass.isSessionAlive.runZmxno longer blocks the pool. It awaits exit viaterminationHandlerand reads stdout on a GCD thread. Before, it calledreadDataToEndOfFile+waitUntilExiton the cooperative pool.Issue #215 is now enforced in
ZmxSessionLauncher.sendGate, which every local send passes before typing. It takes a freshzmx ls, never the shared listing, and types only whenparseSessionTaskStatefinds the session's row with noended=,exit_code=orerr=.Independent review
ReviewPR582SendAck reviewed
0da20eaaadversarially. It swapped main's files in under the tests and found no vacuous tests. Its verdict was "not ready, block on F1 and F2". It re-reviewed47957204, verified F1–F9 fixed, and found F10 in the F2 fix, then F10b in the F10 fix. Both are fixed in11ff1572, and its final verdict at that head is READY (113 of 114 of its probes and tests pass; the one failure is a round-1 probe asserting the shared listing sees a task end, which it does not by design, since every typing and lifecycle check takes a fresh listing):zmx sendblocked the loop's follow-ups and wakes foreverdeliveryDeadline, staged and loggedisSessionAlivealways freshconn/seqon every outcome;send-typedandsend-droppedaddedgraphChangedbroadcasts per sendlscould stall every joining reader; nothing provedsendcallssendGatetypeLoggedstages it onceTest plan
RED: xcodebuild test -only-testing:graphcodeTests/SendAcknowledgementTests with deliverAdHocMessage typing inline as on main -> aSendIsAcknowledgedBeforeItIsTyped failed, events were typed-then-acknowledged
RED: xcodebuild test -only-testing:graphcodeTests/HuskSendGateTests with parseSessionTaskState treating ended= and exit_code= rows as alive -> 4 issues: aHuskAtListingTimeIsNeverTypedInto failed for all 3 markers and aSessionThatEndsAfterTheSharedListingIsNeverTypedInto failed
RED: same suite with sendGate reading the shared listing instead of a fresh one -> aSessionThatEndsAfterTheSharedListingIsNeverTypedInto failed, typed into the husk
RED: MessageDeliveryTests with the sendGate line removed from send -> aSessionWhoseTaskEndedIsNeverTypedInto failed, a real zmx husk was typed into
RED: SubGraphAddressingTests with child stores chaining instead of typing inline -> aMessageAddressedToAChildLoopReachesItsTransport failed, order lost
RED: SendAcknowledgementTests without the typing deadline -> aTypingThatHangsIsStagedAtTheDeadlineAndFreesTheLoop exceeded its 60 s time limit
RED: SendAcknowledgementTests with the closing broadcast unconditional -> aSendIsBroadcastOnceWhenItsDrainChangesNothing failed, 2 broadcasts
RED: SendAcknowledgementTests without the cancellation checks in typeAdHocMessage -> aTypingAbandonedAtTheDeadlineIsNeitherRetriedNorStagedTwice failed: 2 attempts, 1 respawn, staged twice
RED: SendAcknowledgementTests without the check after the respawn settle -> aDeadlineDuringTheRespawnSettleDoesNotTypeAgain failed, 2 attempts
GREEN: xcodebuild test -only-testing for SendAcknowledgementTests, SessionListingTests, HuskSendGateTests, RunCollectingOutputTests, MessageDeliveryTests, SubGraphAddressingTests, GoalResolutionFollowUpTests, RespawnOnSendTests, MessageAndSpawnTests, CreatedByLoopTests -> 79 tests in 10 suites passed, exit 0
REGRESSION: xcodebuild -scheme graphcode test at 11ff157 with private DerivedData -> 2017 tests in 217 suites passed, exit 0; graphcoded and graphcode-cli builds exit 0; make check exit 0; swift build plus scripts/cli-smoke.sh -> smoke exit 0