Skip to content

Acknowledge node send before typing it; one zmx ls per presence pass - #582

Merged
scgopi merged 7 commits into
mainfrom
fix/send-ack-async-zmx
Oct 2, 2026
Merged

scgopi merged 7 commits into
mainfrom
fix/send-ack-async-zmx

Conversation

@scgopi

@scgopi scgopi commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Summary

Ends the false exit-75 node send timeouts: the daemon now acknowledges a send once it is recorded and queued, and types it afterwards. Presence passes also stop running one full zmx ls per node.

Why

DaemonTimeoutTriage measured it on 2026-10-01: a send landed and still exited 75. The daemon typed the whole message before acknowledging it: a full serial zmx ls over all 27 sessions, then each chunk, the 400ms beat, then Enter. Every unreachable session costs zmx ls a 1000ms probe timeout, so on a starved machine the request ran 10–30s against the CLI's 10s wait.

Timing

Rig: an isolated graphcoded (private GRAPHCODE_SUPPORT_DIR and ZMX_DIR; the live daemon was never touched), with 20 goal loops whose sessions run cat behind a fake claude, and one client held connected so presence passes run.

Builds: main = acdb2a62, whose daemon sources are identical to this PR's base. PR = 0da20eaa.

Method:

  • 10 node send runs per build.
  • Wall clock is the CLI process time.
  • handle_ms comes from that daemon's graphcoded.log.
  • zmx ls calls per pass were counted with a logging zmx shim.

Load: the Edge storm had ended by measurement time (load 4.5), so the load is synthetic: 60 busy processes, load average 54–72 on 10 cores.

A: load only, all 20 sessions reachable

Metric main PR
node send wall p50 / p90 / max 0.54 / 0.58 / 0.71 s 0.02 / 0.04 / 0.05 s
messageNode handle_ms p50 / p90 / max 522 / 551 / 628 2 / 3 / 4
zmx ls per presence pass 20 (8 passes) 1 (9 passes)
presence pass duration 0.7–2.4 s 0.3–0.4 s
exit 75 0 / 10 0 / 10

B: same load, plus 2 session daemons SIGSTOPped

This is the field's failure mode: their rows read err=Timeout status=unreachable, and each zmx ls takes 2.0 s.

Metric main PR
node send wall p50 / p90 / max 2.53 / 2.56 / 2.63 s 0.02 / 0.03 / 0.03 s
messageNode handle_ms p50 / p90 / max 2514 / 2533 / 2615 1 / 2 / 2
zmx ls per presence pass 20 (2 passes) 1 (8 passes)
presence pass duration 40.8–41.1 s 2.3 s
exit 75 0 / 10 0 / 10

Synthetic CPU load alone did not reproduce the field's multi-second zmx ls; unreachable sessions do. On main, every send pays one ls of about 1 s per unreachable session, so four or more cross the CLI's 10 s deadline. The PR's acknowledgement does not depend on any zmx call. Its wall time excludes the typing by design. The typing still runs after the ack, gated by a listing of its own.

Changes

  • Ack before typing (GraphStore). messageNode broadcasts its verdict once the message is mirrored to the Mailroom and queued, before the settle drain.
    • Typing runs afterwards on a per-target chain, so sends keep their order. A --follow-up to the same loop waits behind a send that is still being typed.
    • Each typing is bounded by deliveryDeadline. A typing that hangs is staged to memory, and the loop's queue moves on.
    • A composite's child is typed inline, because its store lives for one command.
    • Known limit, by design: a send to a loop inside a composite is still typed before its acknowledgement, so exit 75 remains possible for those loops on a starved machine.
  • Not-live targets still answer before the ack. A not-live target still gets the synchronous "staged to its memory" verdict. For a finished loop, liveness is checked against a fresh listing before the ack, so a child reporting to a resolved parent still hears "staged".
  • Failures after the ack are discoverable, not broadcast.
    • Every outcome is logged against the acknowledged request (conn/seq), as send-typed, send-staged (reason delivery-failed, session-gone or deadline) or send-dropped (target deleted while queued).
    • A staged failure also goes to the loop's memory.
    • None of this is sent as .errorOccurred, because that reaches every connection and the next waiting CLI would take it as its own verdict.
  • One broadcast per send. A send whose drain changes nothing is broadcast once.
  • CLI wording. The CLI prints accepted — typing it in now instead of delivered. The remote Python shim matches.
  • One zmx ls per pass (ZmxSessionLauncher.SessionListing).
    • Presence, activity and usage read a shared listing.
    • ProjectRegistry.pollPresence brackets each tick as one pass.
    • Outside a pass a listing is reused for 2 s. A listing in flight is joined unless it has run past 30 s.
    • Starting or killing a session invalidates the listing.
    • Lifecycle decisions take fresh listings: start and terminate results, a resume that may have died, the first-pass kickoff, and isSessionAlive.
  • runZmx no longer blocks the pool. It awaits exit via terminationHandler and reads stdout on a GCD thread. Before, it called readDataToEndOfFile + waitUntilExit on the cooperative pool.

Issue #215 is now enforced in ZmxSessionLauncher.sendGate, which every local send passes before typing. It takes a fresh zmx ls, never the shared listing, and types only when parseSessionTaskState finds the session's row with no ended=, exit_code= or err=.

Independent review

ReviewPR582SendAck reviewed 0da20eaa adversarially. It swapped main's files in under the tests and found no vacuous tests. Its verdict was "not ready, block on F1 and F2". It re-reviewed 47957204, verified F1–F9 fixed, and found F10 in the F2 fix, then F10b in the F10 fix. Both are fixed in 11ff1572, and its final verdict at that head is READY (113 of 114 of its probes and tests pass; the one failure is a round-1 probe asserting the shared listing sees a task end, which it does not by design, since every typing and lifecycle check takes a fresh listing):

# Finding Outcome
F1 Sends into a composite child lost order: a fresh child store per command meant a fresh chain per send ✅ child stores type inline
F2 The typing chain had no deadline, so a hung zmx send blocked the loop's follow-ups and wakes forever ✅ bounded by deliveryDeadline, staged and logged
F3 The sub-graph test passed only because its stub returned synchronously ✅ slow stub, two sends, order asserted
F4 Lifecycle checks trusted a shared listing up to 30 s old ✅ fresh listings
F5 The resolved-target verdict could accept a just-ended husk ✅ isSessionAlive always fresh
F6 Staged sends could not be tied to their request; a deleted target was dropped silently ✅ conn/seq on every outcome; send-typed and send-dropped added
F7 Duplicated #346 comment ✅ removed
F8 Two graphChanged broadcasts per send ✅ the second is skipped when nothing changed
F9 One hung ls could stall every joining reader; nothing proved send calls sendGate ✅ joins bounded at 30 s; real-zmx husk test
F10 A deadline-abandoned typing still respawned, retried behind the next message, and staged twice ✅ returns once cancelled; typeLogged stages it once
F10b A deadline landing in the respawn settle still retried inside the cancelled task ✅ checked after the settle too

Test plan

RED: xcodebuild test -only-testing:graphcodeTests/SendAcknowledgementTests with deliverAdHocMessage typing inline as on main -> aSendIsAcknowledgedBeforeItIsTyped failed, events were typed-then-acknowledged
RED: xcodebuild test -only-testing:graphcodeTests/HuskSendGateTests with parseSessionTaskState treating ended= and exit_code= rows as alive -> 4 issues: aHuskAtListingTimeIsNeverTypedInto failed for all 3 markers and aSessionThatEndsAfterTheSharedListingIsNeverTypedInto failed
RED: same suite with sendGate reading the shared listing instead of a fresh one -> aSessionThatEndsAfterTheSharedListingIsNeverTypedInto failed, typed into the husk
RED: MessageDeliveryTests with the sendGate line removed from send -> aSessionWhoseTaskEndedIsNeverTypedInto failed, a real zmx husk was typed into
RED: SubGraphAddressingTests with child stores chaining instead of typing inline -> aMessageAddressedToAChildLoopReachesItsTransport failed, order lost
RED: SendAcknowledgementTests without the typing deadline -> aTypingThatHangsIsStagedAtTheDeadlineAndFreesTheLoop exceeded its 60 s time limit
RED: SendAcknowledgementTests with the closing broadcast unconditional -> aSendIsBroadcastOnceWhenItsDrainChangesNothing failed, 2 broadcasts
RED: SendAcknowledgementTests without the cancellation checks in typeAdHocMessage -> aTypingAbandonedAtTheDeadlineIsNeitherRetriedNorStagedTwice failed: 2 attempts, 1 respawn, staged twice
RED: SendAcknowledgementTests without the check after the respawn settle -> aDeadlineDuringTheRespawnSettleDoesNotTypeAgain failed, 2 attempts
GREEN: xcodebuild test -only-testing for SendAcknowledgementTests, SessionListingTests, HuskSendGateTests, RunCollectingOutputTests, MessageDeliveryTests, SubGraphAddressingTests, GoalResolutionFollowUpTests, RespawnOnSendTests, MessageAndSpawnTests, CreatedByLoopTests -> 79 tests in 10 suites passed, exit 0
REGRESSION: xcodebuild -scheme graphcode test at 11ff157 with private DerivedData -> 2017 tests in 217 suites passed, exit 0; graphcoded and graphcode-cli builds exit 0; make check exit 0; swift build plus scripts/cli-smoke.sh -> smoke exit 0

scgopi and others added 7 commits October 1, 2026 19:57
node send exited 75 ("may still have been applied") for messages that
landed: the daemon typed the whole message (a zmx ls gate over every
session, each chunk, the submit beat, Enter) before acknowledging, and on
a CPU-starved machine that outlasted the CLI's 10s wait.

- messageNode is acknowledged once the message is on the board and queued
  for its session; typing runs afterwards on a per-target chain that keeps
  send order, and follow-ups wait behind it. A failure after the ack is
  staged to the loop's memory and logged, not broadcast, so no other
  client takes it for its own verdict.
- One zmx ls per presence pass (ZmxSessionLauncher.SessionListing),
  shared with sends and joined while in flight, instead of a full listing
  per node per read. Starting or killing a session invalidates it, and a
  "not alive" that would refuse a send or allow a resolution is confirmed
  against a fresh listing.
- runZmx awaits exit through terminationHandler and reads on a GCD
  thread instead of blocking the cooperative pool.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
The shared listing can be older than a task's end, and a session whose
task ended is a husk whose wrapper shell would take the keystrokes. The
send gate now always takes a fresh zmx ls (typing runs after the ack, so
this costs the sender nothing) and types only into a task zmx reports
alive. A row with exit_code= but no ended= counts as ended too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
- A composite child's store lives for one command, so it types inline
  instead of on a chain no later send or parent could see (order kept).
- Each post-ack typing is bounded by deliveryDeadline; a hung one is
  staged and logged, and the loop's follow-ups and wakes move on.
- Every outcome is logged against the acknowledged request: send-typed,
  send-staged (delivery-failed, session-gone, deadline), send-dropped.
- Lifecycle checks (start/terminate results, a resume that may have
  died, the first-pass kickoff, isSessionAlive) take fresh listings; a
  hung in-flight listing is joined for at most 30s.
- A send whose drain changes nothing is broadcast once, not twice.
- Tests: composite-child ordering with a slow transport, a hung typing,
  one broadcast per send, and a real zmx husk that send refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
withDeadline cancels the typing it gives up on but cannot stop it: the
cancelled send read as a failure, so the abandoned task respawned an
unattended loop, typed again behind the chain's next message, and staged
the message a second time. It now returns as soon as it sees it was
cancelled; typeLogged has already staged it once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
A deadline that fires during the respawn settle returned the sleep at
once, and the retry then typed again inside the cancelled task. The
check now follows the settle as well as each delivery attempt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
@scgopi
scgopi merged commit 6010b28 into main Oct 2, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant