Conversation
Keep legacy JSONL as a read-only prefix and publish immutable batches with atomic same-directory rename. Verify uncertain commits and fence the store until recovery when their outcome cannot be established. Treat derived metadata as repairable without hiding explicit update failures. Cache compact verified indexes rather than replaying history on hot operations. Existing logs remain readable, but older SDK versions cannot read new batch commits. This requires one writer and atomic rename; it does not guarantee power-loss durability.
Keep activation identity separate from runtime residency. Drain admitted steps and nested work without starting successors, preserve pending wakes and background starts, and reject unsafe handoffs while old resources remain outstanding. Serialize event projection updates, fence uncertain writes, and prevent revoked contexts from publishing late effects. Preserve runtime refusal errors in RPC envelopes. BREAKING CHANGE: reacquire a closed session from its manager before reopening it. Disposed Session references no longer revive or forward to replacement runtimes. Operation-scoped mutation helpers expire with their owning operation.
Persist delivery identity and the original-request fingerprint with each accepted input. Join concurrent matching deliveries, replay the original receipt after consumption or restart, and reject conflicting reuse without changing legacy calls.
Exercise retained wakes, completed and pending tool calls, replayed delivery receipts, commit-time ownership fencing, and uncertain filesystem commits across two managers. Document host obligations, storage compatibility, and the existing HTTP error envelope.
A replace hook that returned a new id made executeToolCall emit tool_started and tool_completed under it, so the reducer never cleared the original entry and the rewritten tool_exec case re-entered on the same head forever. The id pairs the result with the LLM's tool_use, so a hook may swap name and input but not identity. Guard the loop on head progress as well, so any future path that leaves the head pending ends the turn instead of spinning. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…ried A host that drains on a deadline could not use park: nothing bounded it and one settled failure refused it forever. Bound every wait a drain depends on. Restore the worker effect drain bound that the stop path still had, bound the park wait for workers that outlive it, and give parkSession a timeoutMs that bounds the caller's wait without cancelling the drain. Bound waiting for a turn in the session write queue too: appends are serialised, so one stalled writer otherwise blocked session_closed forever and close could not return. A turn that never comes definitely did not commit, and the store fences because nothing may be ordered behind an open outcome. Report drain failures once instead of retaining them. The store failure, the session and agent scheduler errors, and the service publication failures are settled history once surfaced; keeping them downgraded every later handoff of that runtime from park to revoke. Park now stays retryable after it reports one — the runtime never returns to ready, so it still admits no new work — while a park that found the runtime gone stays terminal. The fence classification is deliberately unchanged: an unclassified append failure may still have landed, so it fences. BREAKING CHANGE: SessionStore takes an options bag after applyEvent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
Fencing only worked for a host somebody had already revoked. Nobody can revoke a host that is merely unreachable, so a partitioned owner kept writing and kept retrying; the only signal it could ever get was its own refused append. Add SessionOwnershipLostError, the seam a store bound to a host's lease throws once that lease moved on. It is a definite noncommit that also stops the runtime: the store fences, the session revokes itself rather than parking — a replacement owns the log, so there is nothing to drain — and the manager drops the residency so a later access reloads and asks who owns it now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…pped A tenure keeps `entry` so re-admission can be refused while a disposed runtime still holds resources. Nothing ever cleared it, and tenures are never pruned, so every session id the process touched kept its store, agents and plugin contexts reachable for the life of the process — a pod cycling sessions leaked all of them. admission() creates a tenure on first access, so this needed no activateSession call to accumulate. Clear the reference once the runtime reports no unsafe resources, and retry when outstanding cleanup settles. The tombstone itself stays: forgetting it would let the next access silently re-own a session the host had parked. Report the guarded count in getRuntimeCacheStats so a host can watch it settle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
… listings reconcileMetadata discarded the events it was handed and ran a full recover(), so every session load re-read and re-digested every batch a second time — the caller had just loaded that log for exactly this reason. Trust the cached verification instead. A failed rename left its .pending-* file behind, and each retry mints a new name, so a persistently failing commit filled the directory with orphans that only a later recover() would collect. Recovery never reads pending files, so drop it where it was written. getAllSessionMetadata lost its per-session guard, so one unreadable meta.json failed the listing for every other session. Restore it and say which session was skipped; a direct getMetadata still surfaces the error. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
… notifies Scoping turned two things that could not fail into things that do. notify is documented as ephemeral and never persisted, and the callers that outlive their operation are exactly the ones that cannot handle a throw — an interval, a detached promise, a callback that outlived its method. A late notification is now dropped. RuntimeFileStore reads returned Result and threw a lease error before returning one. Only a write can land behind the runtime that replaces this one, so reads pass through; write and remove still take a resource lease. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
performDisposal released its teardown lease as a plain statement, so a rejecting agent drain skipped the release, the detach and markDisposed — leaving activeCount above zero, every later waitForIdle hanging, and the activity stuck in unloading. A worker's emit applied the reducer only while the runtime still ran, so an append that committed during teardown left emit() resolving on a state behind the durable log, and everything the worker decided next was computed from it. callPluginMethod awaited the session-wide scheduler tail, coupling one call's latency to unrelated wakes already queued. Await only what the call armed. The user-chat receipt map was copied whole on every accepted message and never pruned — quadratic over a replay, unbounded over a session. Copy only when there is a receipt to record, and keep the most recent ones. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
Record the ownership seam, the two drain bounds and the park retry contract, and say plainly that nothing enforces the single-writer requirement the storage protocol depends on: rename replaces, so two stores over one directory overwrite each other silently. Note the parked-service status gap rather than implying the projection is always truthful. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
Quiesce workers before draining runtime operations, retain initializing runtimes until cleanup settles, and retire failed acquisitions. Cache closed runtimes without accumulating historical handles. Preserve delivery receipts for the session lifetime, honor notification continuations, fence inference file writes, and remove partially written pending batches. Add regression coverage and correct the failing test types.
runCloseHooks memoized each plugin's hook promise and never dropped it on rejection, so a park retry re-awaited the dead rejection and the hook body ran exactly once, ever. cleanupFailed latched with ||= and had no reset, so it stayed an unconditional term of hasUnsafeResources() — forgetEntry never released the entry and activateSession refused that session id for the process lifetime. Both are reachable from the shipped services plugin, which rethrows an AggregateError when a status publication or a process-group stop fails. A rejected hook is no longer cached, and a later cleanup that succeeds clears the flag. This is what SESSION-LIFECYCLE.md already promised. Awaiting the per-agent scheduler drain with Promise.all threw on the first rejection and dropped the rest — but waitForScheduler splices its failures before throwing, so every other agent's failure was gone for good and the next park looked clean while a durable wake stayed un-armed. Settle them all and aggregate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…h it recover() excluded `.pending-` and treated every other name as a batch, so any foreign file in `batches/` shifted the sequence and raised a contiguity error that fenced the session. The condition is on disk, so a fresh store fails identically and the README's recovery advice does not help. The sources are mundane — .DS_Store, an editor swap file, and NFS silly-rename files, which the kernel creates in that directory when an open file is unlinked, which is exactly what this store does to pending files. Select batches positively instead: a name the store did not write is not a missing batch. listSessions() asserted every id unfenced, and the fence is never cleared, so one bad log permanently threw out of the listing — defeating readMetadataOrSkip 45 lines below it, whose whole purpose is that one unreadable session must not take the listing down for the others. getStats() awaits it bare, so GET /status, the endpoint an orchestrator polls, returned 500 for the process lifetime. The fenced session is now skipped and a direct getMetadata still surfaces the error. Also: clamp a negative `since` in loadRange (both stores sliced from the end, dropping events and returning negative cursors that made a client re-poll forever), keep the real rename failure instead of the readback's ENOENT, prune `fenced` alongside its three sibling maps, drop the getMetadata override that re-implements the base method, and export SessionOwnershipLostError from the barrel that listed the other five. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
beginPark's zero-entry release left tenure.entries and tenure.entry populated where the success branch clears both; the two release paths now leave identical bookkeeping. A host has to implement EventStore, so export it with LoadRangeOptions and LoadRangeResult. Tighten sessions.getEvents to `.min(-1)` so an out-of-range cursor is refused at the edge rather than clamped in two stores. A worker's terminal event lost to a swallowed append error produced no log line at all; warn on the branch that drops it. Comments: the close-reason map named "parking causes" for idle and disposed, which are eviction; the fence comment ran to four lines over the repo's three-line ceiling; and a worker comment inlined root-cause analysis that belongs in a commit message. Two tests wrote under an unrelated project's name in /tmp — use the /tmp/roj-* prefix every sibling in that file already uses. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…sumers SESSION-LIFECYCLE.md was not in the package `files`, and npm auto-includes only README/LICENSE, so every consumer of @roj-ai/sdk got README's "See SESSION-LIFECYCLE.md" pointing at nothing — the whole host contract for park and ownership loss was absent from the tarball. The docs told a host to keep the two timeouts inside its shutdown budget without naming a single default, and four of the waits a drain performs have no bound at all. Add the table: every wait in performPark order, its bound and its default, with the unbounded ones named as such. A host cannot make a park fit a budget by tuning knobs, and should size parkSession's timeoutMs and fall back to revoke. Say that a park retry re-runs both hook phases, so onSessionPark and onSessionClose must both be idempotent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
The guard refused a hook event before writing anything, but threw a plain EventStoreError. SessionStore fences on every error that is not an EventAppendError, so one refused hook event blocked every later write of the runtime, including the session_reopened that reopen() needs. The guard now throws ClosedSessionAppendError, an EventAppendError: a definite noncommit that does not fence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
… map admission() creates a tenure on first access to any valid id, and nothing removed it. A server asked about many ids that have no session kept one tenure per id for the life of the process. A tenure that first access created, and that no host holds a handle to, is now dropped when the load finds no session. A tenure a host activated stays: dropping it would let the next access re-admit an id the host revoked. getRuntimeCacheStats().tenureCount makes the map size watchable. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
…ime stops uploadAsync writes meta.json with pendingStart and commits attachment_uploaded before it takes a processing lease. When the runtime was unloading at that point, the method returned "Session runtime is unavailable", but the next load found pendingStart and processed the upload anyway. A client that retried got a duplicate. Any runtime that cannot start the upload now answers as parking already did: the upload is committed and the next load starts it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
matej21
force-pushed
the
fix/session-handoff-drain-seams
branch
from
September 28, 2026 13:34
f4273ef to
e4289bf
Compare
…ches The batch store wrote one file per append and never compacted them. Listing sessions replayed every session's full log, all at once, and deleted leftover files as a side effect. The two things the batches fixed need less than that: - A trailing line without its newline is an append that never completed. Reads ignore it, and the next append drops it by swapping in the log without it through a rename. - A failed append is rolled back to the last committed line, so it is a definite noncommit. When the rollback fails too, the outcome is unknown and the next load decides it. Listing reads only meta.json again. meta.json is derived and may lag the log: a failed metadata write no longer fails a committed append, and the counters are rebuilt from the log on the next append, range read or load. The on-disk format is main's again, so logs stay readable by older SDKs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
…over another writer The committed size was cached per store instance and survived an ownership round trip. When the interim owner died mid-append, the returning host kept its stale size because a load with a torn tail left the cache alone. Its next append then skipped the tail repair and glued onto the fragment, and a failed append rolled the log back to the stale size, deleting the interim owner's committed events while reporting a definite noncommit. A load now sets the cached size from the log it read, or drops it on a torn tail so the next append repairs it. A rollback that finds more bytes than the failed append could have written reports an unknown outcome and truncates nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
A failed rollback threw before the metadata step, so the session was not marked stale. The appended bytes may still be in the log, yet the next append advanced meta.json incrementally from the lagging counters, and loadRange then gave the last events wrong indexes. The session is now marked stale before the unknown outcome is reported, so the next append or range read rebuilds the counters from the log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…id not commit uploadAsync writes meta.json with pendingStart before it appends the processing event. When that append failed, the handler returned an error but left the metadata in place, so the next load started the upload and delivered a ready attachment the client had already retried. A definite noncommit now removes the metadata and the stored file. An unknown outcome keeps them: the event may have landed, and then the next load must start the upload. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
Worker emit appended the sub-event first and reduced it afterwards. A reducer that threw then left the event in the log, and every replay threw on it again, so the session could no longer load. The reducer now runs before the append, and a rejected event is never written. Local state still changes only after the append commits, reduced from the state current at that point so a concurrent emit is not lost. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
writeQueueTimeoutMs existed only on SessionManagerOptions, so createSystem, bootstrap and the shipped hosts always ran with the 30 second default. A host that drains on a deadline could not keep the bound under its own budget. createSystem now takes the option, and Config reads it from SESSION_WRITE_QUEUE_TIMEOUT_MS. Validation rejects zero, which would fence the first write that waits, and delays setTimeout cannot honour. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…fence The write queue deadline started when a write was queued, so it measured the whole queue ahead of it. A batch of healthy but slow appends tripped it, and writes that never reached the store were reported as unknown outcomes. The bound now starts when an append reaches the head. An append that does not settle in time has an unknown outcome and fences the store. The writes behind it get EventAppendError with the fence as cause, because they never reached the store. A fence also notified nobody unless ownership was lost. The runtime stayed resident and ready and refused every later write until park, dispose or idle eviction. Every fence now stops the runtime the way ownership loss does. The manager drops the residency, and the next access reloads from the log, which decides the uncertain write. The lookup goes by runtime, so a reopened runtime is dropped too. Breaking for custom event stores: an append that does not settle within writeQueueTimeoutMs now counts as an unknown outcome. The EventStore interface documents the error contract. SessionStore.onOwnershipLost is replaced by onFenced. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
Delivery receipts stay in the projection for the session lifetime, keyed by the caller's delivery ID, and the ID had no upper bound. A client could store arbitrarily long keys that bypass the message size limit. sendMessage now rejects an ID longer than 256 characters before anything is accepted. The event schema keeps accepting longer IDs so existing logs still replay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
forgetEntry retried clearing tenure.entry once, after waitForLocalCleanup. That waited only for hooks and local cleanup, while hasUnsafeResources also counts pending writes, the scheduler tail and tracked resources. A revoke during a stalled append therefore kept the store, agents and plugin contexts reachable for good, and retainedRuntimeCount never returned to zero. Session.whenSafe() now settles once all of those have settled, and forgetEntry waits for it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…voked id The services close hook threw whenever a status publication had failed, even a stale one, although it had already stopped the processes and released the ports. On revoke that set cleanupFailed, and nothing retries a revoke, so the runtime stayed unsafe for good: admission refused the id as unloading and activateSession refused it too. A throw from a close hook now means that resources may still be held. A publication failure is still reported by a park, which is retried, and is only logged on every other close. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…cceptance A fence now stops the runtime, so a retry on the same runtime after an unknown outcome is refused as session_runtime_unavailable before it writes anything. It used to reject through the fenced store. Either way the retry cannot duplicate the message, and a fresh load still replays the committed receipt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…vates After ownership loss the manager dropped the cache entry, but the tenure stayed active. Every later access on the stale host replayed the full log, ran the ready hooks (services started processes, workers resumed), stepped the agent, and fenced again on the first write. The host saw only a generic revoked runtime. The tenure now moves to a lost state. Access is refused cheaply with the domain error session_ownership_lost, and parkSession rejects with SessionOwnershipLostError, until the host calls activateSession again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…nure refuses The review fixes changed what the README described: the write queue timeout now bounds one append at the head of the queue and is settable through createSystem and the environment, a lost tenure refuses access until reactivation, and two file stores over one session can delete each other's events, not just interleave them. Store authors also need the error contract in one place. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
The four commits already on this branch make the agent runtime survive a host
handoff. This PR fixes what they broke, and finishes the seams a host driver
actually has to plug into — a pool of pods where one dies and another takes over,
and where SIGTERM means drain within a budget.
Scope is deliberately the SDK side. A Postgres-backed event store is not here.
The blocking bug
beforeToolCallreturning{ action: 'replace' }with a new id madeexecuteToolCallemittool_started/tool_completedunder it, so the reducernever cleared the original entry — and the rewritten
tool_execcase re-enterswhile anything stays pending. The agent re-ran the tool forever, appending two
events per iteration.
The id is the LLM's
tool_useid and the result has to pair with it, so a hookmay now swap name and input but not identity. The loop also guards on head
progress, so a future path that leaves the head pending ends the turn instead of
spinning.
Draining
Nothing bounded a drain, and one settled failure refused it forever. Neither is
usable from a host with a shutdown budget.
Bounds. The worker effect drain lost the bound the stop path still had, so a
stalled append held the runtime lease and
onSessionClose('parked')waited on itwithout limit. Restored, and the park wait is bounded too — stragglers are
stopped so their state parks and the next host resumes them.
parkSessiontakestimeoutMs, which bounds the caller's wait without cancelling the drain.The one that was not obvious: appends became serialised in
SessionStore, whichis right — state has to follow the log — but it made one stalled write block
every later one, including the
session_closedthatclose()has to writefirst.
close()could not return at all.writeQueueTimeoutMsnow bounds eachappend once it reaches the head of the queue, so a long queue of healthy writes
never trips it. An append that does not settle in time has an unknown outcome
and fences the store, because nothing may be ordered behind a write that may
still land. Writes queued behind it never reached the store, so they are refused
with
EventAppendError, the fence as itscause. Hosts set it throughSessionManagerOptions,createSystemorSESSION_WRITE_QUEUE_TIMEOUT_MS.A fenced runtime stops. A fenced store accepts no write, so a resident
runtime behind it was dead until something parked or disposed it — one transient
error from a custom store was enough. Every fence now revokes the runtime and the
manager drops it; the next access reloads from the log, which decides whether
the uncertain write landed. Until that write settles, the id refuses access as
unloading.Retries. Five accumulators recorded failures and never cleared them: the
store failure, the session and agent scheduler errors, the service publication
failures. One recoverable append error or one failed wake downgraded every later
handoff of that runtime from park to hard revoke, for as long as the pod kept it
resident. They are reported once now, and park stays retryable after reporting —
the runtime never returns to
ready, so it still admits no new work. A park thatfound the runtime already gone stays terminal.
The seam that was missing
Fencing only worked against a host somebody had already revoked. Nobody can
revoke a host that is merely unreachable — that is what a partition is — so a
stale owner kept writing and kept retrying. The only signal it can ever get is
its own refused append.
SessionOwnershipLostErroris that seam: a store bound to a host's lease throwsit once the lease moved on. It is a definite noncommit that also stops the
runtime — the store fences, the session revokes itself rather than parking
(a replacement owns the log, so there is nothing to drain), and the manager drops
the residency and marks the tenure lost. Reloading would replay the log and
restart services and workers only to fence again on the first write, so later
access on that host is refused with
session_ownership_lost, andparkSessionrejects with
SessionOwnershipLostError, until the host callsactivateSessionagain.
The epoch itself stays where it belongs: the host builds an
EventStorebound toits lease and passes it in. The SDK does not decide ownership.
Leak
tenure.entryexists to refuse re-admission while a disposed runtime still holdsresources, and nothing ever cleared it. Tenures are never pruned and
admission()creates one on first access, so every session id the processtouched kept its store, agents and plugin contexts reachable for the life of the
process. Cleared once the runtime is safe:
Session.whenSafe()settles afterpending writes, the scheduler tail, tracked resources and local cleanup, so a
revoke during a stalled append releases the runtime once the append settles.
getRuntimeCacheStats().retainedRuntimeCountmakesit watchable. The tombstone itself stays — forgetting it would let the next
access silently re-own a parked session.
Error channel
Scoping turned two things that could not fail into things that do.
notifyisephemeral and never persisted, and its callers are exactly the ones that cannot
handle a throw — an interval, a detached promise, a callback that outlived its
method; a late notification is dropped.
RuntimeFileStorereads returnResultand threw a lease error before returning one; only a write can land behind a
replacement runtime, so reads pass through.
Event store
The branch first moved the file store to one immutable batch file per append.
That fixed two real defects of the JSONL log on
main: a torn last line madethe whole session unloadable, and a partly written append was reported as a
definite noncommit, so a retry duplicated it. It also brought one file per
append with no compaction, and a listing that replayed every session's full log
at once. This PR returns to a single
events.jsonland fixes the two defectsdirectly:
without one is ignored by reads and dropped by the next append, which swaps in
the log without it through a rename.
did not commit. When the rollback fails too, the outcome is unknown, the
metadata is marked stale, and the next load decides it.
finds lines it did not write reports an unknown outcome and truncates nothing,
so a host that takes a session back cannot delete what another host committed.
Listing reads only
meta.jsonagain, never an event log.meta.jsonis derivedand may lag the log: a failed metadata write no longer fails a committed append,
and the counters are rebuilt from the log on the next append, range read or load.
The on-disk format is
main's, so older SDKs still read it.Review follow-ups
Only a tenure that first access created is dropped; one a host activated stays,
so a revoked id cannot be re-admitted.
getRuntimeCacheStats().tenureCountmakes the map watchable.
ClosedSessionAppendError, a definitenoncommit, so one refused hook event no longer fences every later write,
session_reopenedincluded.uploadAsyncreports an upload that is already committed asprocessingwhenthe runtime cannot start it, as parking already did. The next load starts it;
returning an error made a retrying client upload it twice.
main, which addedcontexttouser-chat.sendMessage. Thedelivery fingerprint covers it, so a retry with a different context conflicts.
Second review
uploadAsyncwrotependingStartbefore itsprocessingevent. When thatappend definitely did not commit, the next load started the upload anyway and
a retrying client got two attachments. The upload is now discarded.
failed every replay. The reducer now runs as a check before the append and
again from current state after it.
the hook had already released the processes. A close hook now throws only when
something is still held; the services hook re-throws a publication failure
only on park, where it is retried.
deliveryIdis limited to 256 characters.EventAppendErroris adefinite noncommit,
SessionOwnershipLostErrorstops the runtime, anythingelse is an unknown outcome and fences.
Smaller
performDisposalreleased its teardown lease as a plain statement, so arejecting agent drain skipped the release, the detach and
markDisposed—leaving
activeCountabove zero and every laterwaitForIdlehanging. Aworker's
emitapplied the reducer only while the runtime still ran, so anappend that committed during teardown left
emit()resolving on a state behindthe log.
callPluginMethodawaited the session-wide scheduler tail.User-chat keeps delivery receipts for the session lifetime, including in the
replayed projection, so a retry dedupes exactly however late it arrives. Memory
grows with the number of distinct delivery IDs.
Deliberately not done
The fence classification stays as it was, but it is breaking for store
authors. An append slower than
writeQueueTimeoutMsnow counts as an unknownoutcome. It was suggested that only
EventAppendOutcomeUnknownErrorandEventLogCorruptionErrorshould fence. Theyshould not: an unclassified failure from a custom store may still have landed, so
fencing unless the store said otherwise is the safe default, and
session-store.test.tsalready documents it.Terminal service status during park. A service that exits while the runtime
is parking cannot publish
service_status_changed(stopped), so a replaying hostsees it as
running. Publishing it without a lease does not fix this — theunscoped path is gated by the same
tryOperationand fails identically. It needseither publishing before writes stop or reconciling on load. Documented in the
README as a known gap.
A single writer is still not enforced. The store caches where the next
append starts, so two stores over one directory interleave or cut each other's
lines, and a torn-tail repair can delete what the other committed. Loading
refreshes the cache, so a host that takes a session back must load before
appending. The README says so; a host that runs more than one instance fences
externally through
SessionOwnershipLostError. The file store's per-sessioncaches are not pruned either, because the
EventStoreinterface has no releasehook.
An upload whose owner was lost mid-commit. A revoked runtime cannot write,
so an upload whose
processingevent was refused by ownership loss cannotdiscard itself, and the new owner may still start it.
A failed revoke cleanup. When a close hook reports resources still held after
a revoke, nothing retries it, and the id stays refused on that host until the
process restarts.
Snapshots. Taking over a session still replays its full log. That belongs
with the store that will serve a pool.
Stale service publication failures. A status publication that failed is
reported by the next
stopServiceor park, not by the call that caused it.An upload revoked mid-commit. If the runtime is revoked while
attachment_uploadedis being written, the event can land and the method stillreports an error. The runtime cannot tell whether the write landed before it was
detached.
Verification
bun run ts:build,bun run lint, andbun test packages/*/src packages/*/tests— 1747 pass, 0 fail.
Each fix has a test that fails without it; for the review follow-ups and the
JSONL rollback and tail repair this was checked by removing the fix.
🤖 Generated with Claude Code
https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB