Skip to content

Bound the session drain, fix the handoff regressions, and add the ownership seam - #31

Open
matej21 wants to merge 33 commits into
mainfrom
fix/session-handoff-drain-seams
Open

matej21 wants to merge 33 commits into
mainfrom
fix/session-handoff-drain-seams

Conversation

@matej21

@matej21 matej21 commented Sep 8, 2026 •

Copy link
Copy Markdown
Member

What this is

The four commits already on this branch make the agent runtime survive a host
handoff. This PR fixes what they broke, and finishes the seams a host driver
actually has to plug into — a pool of pods where one dies and another takes over,
and where SIGTERM means drain within a budget.

Scope is deliberately the SDK side. A Postgres-backed event store is not here.

The blocking bug

beforeToolCall returning { action: 'replace' } with a new id made
executeToolCall emit tool_started/tool_completed under it, so the reducer
never cleared the original entry — and the rewritten tool_exec case re-enters
while anything stays pending. The agent re-ran the tool forever, appending two
events per iteration.

The id is the LLM's tool_use id and the result has to pair with it, so a hook
may now swap name and input but not identity. The loop also guards on head
progress, so a future path that leaves the head pending ends the turn instead of
spinning.

Draining

Nothing bounded a drain, and one settled failure refused it forever. Neither is
usable from a host with a shutdown budget.

Bounds. The worker effect drain lost the bound the stop path still had, so a
stalled append held the runtime lease and onSessionClose('parked') waited on it
without limit. Restored, and the park wait is bounded too — stragglers are
stopped so their state parks and the next host resumes them. parkSession takes
timeoutMs, which bounds the caller's wait without cancelling the drain.

The one that was not obvious: appends became serialised in SessionStore, which
is right — state has to follow the log — but it made one stalled write block
every later one, including the session_closed that close() has to write
first. close() could not return at all. writeQueueTimeoutMs now bounds each
append once it reaches the head of the queue, so a long queue of healthy writes
never trips it. An append that does not settle in time has an unknown outcome
and fences the store, because nothing may be ordered behind a write that may
still land. Writes queued behind it never reached the store, so they are refused
with EventAppendError, the fence as its cause. Hosts set it through
SessionManagerOptions, createSystem or SESSION_WRITE_QUEUE_TIMEOUT_MS.

A fenced runtime stops. A fenced store accepts no write, so a resident
runtime behind it was dead until something parked or disposed it — one transient
error from a custom store was enough. Every fence now revokes the runtime and the
manager drops it; the next access reloads from the log, which decides whether
the uncertain write landed. Until that write settles, the id refuses access as
unloading.

Retries. Five accumulators recorded failures and never cleared them: the
store failure, the session and agent scheduler errors, the service publication
failures. One recoverable append error or one failed wake downgraded every later
handoff of that runtime from park to hard revoke, for as long as the pod kept it
resident. They are reported once now, and park stays retryable after reporting —
the runtime never returns to ready, so it still admits no new work. A park that
found the runtime already gone stays terminal.

The seam that was missing

Fencing only worked against a host somebody had already revoked. Nobody can
revoke a host that is merely unreachable — that is what a partition is — so a
stale owner kept writing and kept retrying. The only signal it can ever get is
its own refused append.

SessionOwnershipLostError is that seam: a store bound to a host's lease throws
it once the lease moved on. It is a definite noncommit that also stops the
runtime — the store fences, the session revokes itself rather than parking
(a replacement owns the log, so there is nothing to drain), and the manager drops
the residency and marks the tenure lost. Reloading would replay the log and
restart services and workers only to fence again on the first write, so later
access on that host is refused with session_ownership_lost, and parkSession
rejects with SessionOwnershipLostError, until the host calls activateSession
again.

The epoch itself stays where it belongs: the host builds an EventStore bound to
its lease and passes it in. The SDK does not decide ownership.

Leak

tenure.entry exists to refuse re-admission while a disposed runtime still holds
resources, and nothing ever cleared it. Tenures are never pruned and
admission() creates one on first access, so every session id the process
touched kept its store, agents and plugin contexts reachable for the life of the
process. Cleared once the runtime is safe: Session.whenSafe() settles after
pending writes, the scheduler tail, tracked resources and local cleanup, so a
revoke during a stalled append releases the runtime once the append settles. getRuntimeCacheStats().retainedRuntimeCount makes
it watchable. The tombstone itself stays — forgetting it would let the next
access silently re-own a parked session.

Error channel

Scoping turned two things that could not fail into things that do. notify is
ephemeral and never persisted, and its callers are exactly the ones that cannot
handle a throw — an interval, a detached promise, a callback that outlived its
method; a late notification is dropped. RuntimeFileStore reads return Result
and threw a lease error before returning one; only a write can land behind a
replacement runtime, so reads pass through.

Event store

The branch first moved the file store to one immutable batch file per append.
That fixed two real defects of the JSONL log on main: a torn last line made
the whole session unloadable, and a partly written append was reported as a
definite noncommit, so a retry duplicated it. It also brought one file per
append with no compaction, and a listing that replayed every session's full log
at once. This PR returns to a single events.jsonl and fixes the two defects
directly:

  • A line counts as committed once its newline is written. A trailing line
    without one is ignored by reads and dropped by the next append, which swaps in
    the log without it through a rename.
  • A failed append is rolled back to the last committed line, so it definitely
    did not commit. When the rollback fails too, the outcome is unknown, the
    metadata is marked stale, and the next load decides it.
  • Loading a session refreshes where the next append starts. A rollback that
    finds lines it did not write reports an unknown outcome and truncates nothing,
    so a host that takes a session back cannot delete what another host committed.

Listing reads only meta.json again, never an event log. meta.json is derived
and may lag the log: a failed metadata write no longer fails a committed append,
and the counters are rebuilt from the log on the next append, range read or load.
The on-disk format is main's, so older SDKs still read it.

Review follow-ups

  • A lookup of a session id that has no session no longer leaves a tenure behind.
    Only a tenure that first access created is dropped; one a host activated stays,
    so a revoked id cannot be re-admitted. getRuntimeCacheStats().tenureCount
    makes the map watchable.
  • The closed-session guard now throws ClosedSessionAppendError, a definite
    noncommit, so one refused hook event no longer fences every later write,
    session_reopened included.
  • uploadAsync reports an upload that is already committed as processing when
    the runtime cannot start it, as parking already did. The next load starts it;
    returning an error made a retrying client upload it twice.
  • Rebased onto main, which added context to user-chat.sendMessage. The
    delivery fingerprint covers it, so a retry with a different context conflicts.

Second review

  • uploadAsync wrote pendingStart before its processing event. When that
    append definitely did not commit, the next load started the upload anyway and
    a retrying client got two attachments. The upload is now discarded.
  • A worker reducer that threw after its sub-event committed left an event that
    failed every replay. The reducer now runs as a check before the append and
    again from current state after it.
  • A failed service status publication made a revoked id unsafe forever, although
    the hook had already released the processes. A close hook now throws only when
    something is still held; the services hook re-throws a publication failure
    only on park, where it is retried.
  • deliveryId is limited to 256 characters.
  • The README states the event store error contract: EventAppendError is a
    definite noncommit, SessionOwnershipLostError stops the runtime, anything
    else is an unknown outcome and fences.

Smaller

performDisposal released its teardown lease as a plain statement, so a
rejecting agent drain skipped the release, the detach and markDisposed —
leaving activeCount above zero and every later waitForIdle hanging. A
worker's emit applied the reducer only while the runtime still ran, so an
append that committed during teardown left emit() resolving on a state behind
the log. callPluginMethod awaited the session-wide scheduler tail.

User-chat keeps delivery receipts for the session lifetime, including in the
replayed projection, so a retry dedupes exactly however late it arrives. Memory
grows with the number of distinct delivery IDs.

Deliberately not done

The fence classification stays as it was, but it is breaking for store
authors.
An append slower than writeQueueTimeoutMs now counts as an unknown
outcome. It was suggested that only
EventAppendOutcomeUnknownError and EventLogCorruptionError should fence. They
should not: an unclassified failure from a custom store may still have landed, so
fencing unless the store said otherwise is the safe default, and
session-store.test.ts already documents it.

Terminal service status during park. A service that exits while the runtime
is parking cannot publish service_status_changed(stopped), so a replaying host
sees it as running. Publishing it without a lease does not fix this — the
unscoped path is gated by the same tryOperation and fails identically. It needs
either publishing before writes stop or reconciling on load. Documented in the
README as a known gap.

A single writer is still not enforced. The store caches where the next
append starts, so two stores over one directory interleave or cut each other's
lines, and a torn-tail repair can delete what the other committed. Loading
refreshes the cache, so a host that takes a session back must load before
appending. The README says so; a host that runs more than one instance fences
externally through SessionOwnershipLostError. The file store's per-session
caches are not pruned either, because the EventStore interface has no release
hook.

An upload whose owner was lost mid-commit. A revoked runtime cannot write,
so an upload whose processing event was refused by ownership loss cannot
discard itself, and the new owner may still start it.

A failed revoke cleanup. When a close hook reports resources still held after
a revoke, nothing retries it, and the id stays refused on that host until the
process restarts.

Snapshots. Taking over a session still replays its full log. That belongs
with the store that will serve a pool.

Stale service publication failures. A status publication that failed is
reported by the next stopService or park, not by the call that caused it.

An upload revoked mid-commit. If the runtime is revoked while
attachment_uploaded is being written, the event can land and the method still
reports an error. The runtime cannot tell whether the write landed before it was
detached.

Verification

bun run ts:build, bun run lint, and bun test packages/*/src packages/*/tests
— 1747 pass, 0 fail.

Each fix has a test that fails without it; for the review follow-ups and the
JSONL rollback and tail repair this was checked by removing the fix.

🤖 Generated with Claude Code

https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB

matej21 and others added 20 commits September 28, 2026 14:40
Keep legacy JSONL as a read-only prefix and publish immutable batches with atomic same-directory rename. Verify uncertain commits and fence the store until recovery when their outcome cannot be established.

Treat derived metadata as repairable without hiding explicit update failures. Cache compact verified indexes rather than replaying history on hot operations.

Existing logs remain readable, but older SDK versions cannot read new batch commits. This requires one writer and atomic rename; it does not guarantee power-loss durability.
Keep activation identity separate from runtime residency. Drain admitted steps and nested work without starting successors, preserve pending wakes and background starts, and reject unsafe handoffs while old resources remain outstanding.

Serialize event projection updates, fence uncertain writes, and prevent revoked contexts from publishing late effects. Preserve runtime refusal errors in RPC envelopes.

BREAKING CHANGE: reacquire a closed session from its manager before reopening it. Disposed Session references no longer revive or forward to replacement runtimes. Operation-scoped mutation helpers expire with their owning operation.
Persist delivery identity and the original-request fingerprint with each accepted input. Join concurrent matching deliveries, replay the original receipt after consumption or restart, and reject conflicting reuse without changing legacy calls.
Exercise retained wakes, completed and pending tool calls, replayed delivery receipts, commit-time ownership fencing, and uncertain filesystem commits across two managers. Document host obligations, storage compatibility, and the existing HTTP error envelope.
A replace hook that returned a new id made executeToolCall emit tool_started
and tool_completed under it, so the reducer never cleared the original entry
and the rewritten tool_exec case re-entered on the same head forever. The id
pairs the result with the LLM's tool_use, so a hook may swap name and input
but not identity.

Guard the loop on head progress as well, so any future path that leaves the
head pending ends the turn instead of spinning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…ried

A host that drains on a deadline could not use park: nothing bounded it and one
settled failure refused it forever.

Bound every wait a drain depends on. Restore the worker effect drain bound that
the stop path still had, bound the park wait for workers that outlive it, and
give parkSession a timeoutMs that bounds the caller's wait without cancelling
the drain. Bound waiting for a turn in the session write queue too: appends are
serialised, so one stalled writer otherwise blocked session_closed forever and
close could not return. A turn that never comes definitely did not commit, and
the store fences because nothing may be ordered behind an open outcome.

Report drain failures once instead of retaining them. The store failure, the
session and agent scheduler errors, and the service publication failures are
settled history once surfaced; keeping them downgraded every later handoff of
that runtime from park to revoke. Park now stays retryable after it reports one
— the runtime never returns to ready, so it still admits no new work — while a
park that found the runtime gone stays terminal.

The fence classification is deliberately unchanged: an unclassified append
failure may still have landed, so it fences.

BREAKING CHANGE: SessionStore takes an options bag after applyEvent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
Fencing only worked for a host somebody had already revoked. Nobody can revoke
a host that is merely unreachable, so a partitioned owner kept writing and kept
retrying; the only signal it could ever get was its own refused append.

Add SessionOwnershipLostError, the seam a store bound to a host's lease throws
once that lease moved on. It is a definite noncommit that also stops the
runtime: the store fences, the session revokes itself rather than parking —
a replacement owns the log, so there is nothing to drain — and the manager
drops the residency so a later access reloads and asks who owns it now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…pped

A tenure keeps `entry` so re-admission can be refused while a disposed runtime
still holds resources. Nothing ever cleared it, and tenures are never pruned, so
every session id the process touched kept its store, agents and plugin contexts
reachable for the life of the process — a pod cycling sessions leaked all of
them. admission() creates a tenure on first access, so this needed no
activateSession call to accumulate.

Clear the reference once the runtime reports no unsafe resources, and retry when
outstanding cleanup settles. The tombstone itself stays: forgetting it would let
the next access silently re-own a session the host had parked.

Report the guarded count in getRuntimeCacheStats so a host can watch it settle.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
… listings

reconcileMetadata discarded the events it was handed and ran a full recover(),
so every session load re-read and re-digested every batch a second time — the
caller had just loaded that log for exactly this reason. Trust the cached
verification instead.

A failed rename left its .pending-* file behind, and each retry mints a new
name, so a persistently failing commit filled the directory with orphans that
only a later recover() would collect. Recovery never reads pending files, so
drop it where it was written.

getAllSessionMetadata lost its per-session guard, so one unreadable meta.json
failed the listing for every other session. Restore it and say which session was
skipped; a direct getMetadata still surfaces the error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
… notifies

Scoping turned two things that could not fail into things that do.

notify is documented as ephemeral and never persisted, and the callers that
outlive their operation are exactly the ones that cannot handle a throw — an
interval, a detached promise, a callback that outlived its method. A late
notification is now dropped.

RuntimeFileStore reads returned Result and threw a lease error before returning
one. Only a write can land behind the runtime that replaces this one, so reads
pass through; write and remove still take a resource lease.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
performDisposal released its teardown lease as a plain statement, so a rejecting
agent drain skipped the release, the detach and markDisposed — leaving
activeCount above zero, every later waitForIdle hanging, and the activity stuck
in unloading.

A worker's emit applied the reducer only while the runtime still ran, so an
append that committed during teardown left emit() resolving on a state behind
the durable log, and everything the worker decided next was computed from it.

callPluginMethod awaited the session-wide scheduler tail, coupling one call's
latency to unrelated wakes already queued. Await only what the call armed.

The user-chat receipt map was copied whole on every accepted message and never
pruned — quadratic over a replay, unbounded over a session. Copy only when there
is a receipt to record, and keep the most recent ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
Record the ownership seam, the two drain bounds and the park retry contract, and
say plainly that nothing enforces the single-writer requirement the storage
protocol depends on: rename replaces, so two stores over one directory overwrite
each other silently. Note the parked-service status gap rather than implying the
projection is always truthful.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
Quiesce workers before draining runtime operations, retain initializing runtimes until cleanup settles, and retire failed acquisitions. Cache closed runtimes without accumulating historical handles.

Preserve delivery receipts for the session lifetime, honor notification continuations, fence inference file writes, and remove partially written pending batches. Add regression coverage and correct the failing test types.
runCloseHooks memoized each plugin's hook promise and never dropped it on
rejection, so a park retry re-awaited the dead rejection and the hook body ran
exactly once, ever. cleanupFailed latched with ||= and had no reset, so it stayed
an unconditional term of hasUnsafeResources() — forgetEntry never released the
entry and activateSession refused that session id for the process lifetime. Both
are reachable from the shipped services plugin, which rethrows an AggregateError
when a status publication or a process-group stop fails.

A rejected hook is no longer cached, and a later cleanup that succeeds clears the
flag. This is what SESSION-LIFECYCLE.md already promised.

Awaiting the per-agent scheduler drain with Promise.all threw on the first
rejection and dropped the rest — but waitForScheduler splices its failures before
throwing, so every other agent's failure was gone for good and the next park
looked clean while a durable wake stayed un-armed. Settle them all and aggregate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…h it

recover() excluded `.pending-` and treated every other name as a batch, so any
foreign file in `batches/` shifted the sequence and raised a contiguity error
that fenced the session. The condition is on disk, so a fresh store fails
identically and the README's recovery advice does not help. The sources are
mundane — .DS_Store, an editor swap file, and NFS silly-rename files, which the
kernel creates in that directory when an open file is unlinked, which is exactly
what this store does to pending files. Select batches positively instead: a name
the store did not write is not a missing batch.

listSessions() asserted every id unfenced, and the fence is never cleared, so one
bad log permanently threw out of the listing — defeating readMetadataOrSkip 45
lines below it, whose whole purpose is that one unreadable session must not take
the listing down for the others. getStats() awaits it bare, so GET /status, the
endpoint an orchestrator polls, returned 500 for the process lifetime. The fenced
session is now skipped and a direct getMetadata still surfaces the error.

Also: clamp a negative `since` in loadRange (both stores sliced from the end,
dropping events and returning negative cursors that made a client re-poll
forever), keep the real rename failure instead of the readback's ENOENT, prune
`fenced` alongside its three sibling maps, drop the getMetadata override that
re-implements the base method, and export SessionOwnershipLostError from the
barrel that listed the other five.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
beginPark's zero-entry release left tenure.entries and tenure.entry populated
where the success branch clears both; the two release paths now leave identical
bookkeeping.

A host has to implement EventStore, so export it with LoadRangeOptions and
LoadRangeResult. Tighten sessions.getEvents to `.min(-1)` so an out-of-range
cursor is refused at the edge rather than clamped in two stores.

A worker's terminal event lost to a swallowed append error produced no log line
at all; warn on the branch that drops it.

Comments: the close-reason map named "parking causes" for idle and disposed,
which are eviction; the fence comment ran to four lines over the repo's
three-line ceiling; and a worker comment inlined root-cause analysis that belongs
in a commit message. Two tests wrote under an unrelated project's name in /tmp —
use the /tmp/roj-* prefix every sibling in that file already uses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
…sumers

SESSION-LIFECYCLE.md was not in the package `files`, and npm auto-includes only
README/LICENSE, so every consumer of @roj-ai/sdk got README's "See
SESSION-LIFECYCLE.md" pointing at nothing — the whole host contract for park and
ownership loss was absent from the tarball.

The docs told a host to keep the two timeouts inside its shutdown budget without
naming a single default, and four of the waits a drain performs have no bound at
all. Add the table: every wait in performPark order, its bound and its default,
with the unbounded ones named as such. A host cannot make a park fit a budget by
tuning knobs, and should size parkSession's timeoutMs and fall back to revoke.

Say that a park retry re-runs both hook phases, so onSessionPark and
onSessionClose must both be idempotent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wys12r3nbiusCdxfCtcGBB
The guard refused a hook event before writing anything, but threw a plain
EventStoreError. SessionStore fences on every error that is not an
EventAppendError, so one refused hook event blocked every later write of the
runtime, including the session_reopened that reopen() needs.

The guard now throws ClosedSessionAppendError, an EventAppendError: a
definite noncommit that does not fence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
… map

admission() creates a tenure on first access to any valid id, and nothing
removed it. A server asked about many ids that have no session kept one
tenure per id for the life of the process.

A tenure that first access created, and that no host holds a handle to, is
now dropped when the load finds no session. A tenure a host activated stays:
dropping it would let the next access re-admit an id the host revoked.
getRuntimeCacheStats().tenureCount makes the map size watchable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
…ime stops

uploadAsync writes meta.json with pendingStart and commits
attachment_uploaded before it takes a processing lease. When the runtime
was unloading at that point, the method returned "Session runtime is
unavailable", but the next load found pendingStart and processed the
upload anyway. A client that retried got a duplicate.

Any runtime that cannot start the upload now answers as parking already
did: the upload is committed and the next load starts it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
@matej21
matej21 force-pushed the fix/session-handoff-drain-seams branch from f4273ef to e4289bf Compare September 28, 2026 13:34
matej21 and others added 8 commits September 28, 2026 15:52
…ches

The batch store wrote one file per append and never compacted them. Listing
sessions replayed every session's full log, all at once, and deleted leftover
files as a side effect. The two things the batches fixed need less than that:

- A trailing line without its newline is an append that never completed.
  Reads ignore it, and the next append drops it by swapping in the log without
  it through a rename.
- A failed append is rolled back to the last committed line, so it is a
  definite noncommit. When the rollback fails too, the outcome is unknown and
  the next load decides it.

Listing reads only meta.json again. meta.json is derived and may lag the log:
a failed metadata write no longer fails a committed append, and the counters
are rebuilt from the log on the next append, range read or load.

The on-disk format is main's again, so logs stay readable by older SDKs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015oMWoa8H4wY3Q1JLeqkkvs
…over another writer

The committed size was cached per store instance and survived an ownership
round trip. When the interim owner died mid-append, the returning host kept
its stale size because a load with a torn tail left the cache alone. Its next
append then skipped the tail repair and glued onto the fragment, and a failed
append rolled the log back to the stale size, deleting the interim owner's
committed events while reporting a definite noncommit.

A load now sets the cached size from the log it read, or drops it on a torn
tail so the next append repairs it. A rollback that finds more bytes than the
failed append could have written reports an unknown outcome and truncates
nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
A failed rollback threw before the metadata step, so the session was not
marked stale. The appended bytes may still be in the log, yet the next append
advanced meta.json incrementally from the lagging counters, and loadRange then
gave the last events wrong indexes. The session is now marked stale before the
unknown outcome is reported, so the next append or range read rebuilds the
counters from the log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…id not commit

uploadAsync writes meta.json with pendingStart before it appends the
processing event. When that append failed, the handler returned an error but
left the metadata in place, so the next load started the upload and delivered
a ready attachment the client had already retried. A definite noncommit now
removes the metadata and the stored file. An unknown outcome keeps them: the
event may have landed, and then the next load must start the upload.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
Worker emit appended the sub-event first and reduced it afterwards. A reducer
that threw then left the event in the log, and every replay threw on it again,
so the session could no longer load. The reducer now runs before the append,
and a rejected event is never written. Local state still changes only after
the append commits, reduced from the state current at that point so a
concurrent emit is not lost.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
writeQueueTimeoutMs existed only on SessionManagerOptions, so createSystem,
bootstrap and the shipped hosts always ran with the 30 second default. A host
that drains on a deadline could not keep the bound under its own budget.
createSystem now takes the option, and Config reads it from
SESSION_WRITE_QUEUE_TIMEOUT_MS. Validation rejects zero, which would fence
the first write that waits, and delays setTimeout cannot honour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…fence

The write queue deadline started when a write was queued, so it measured the
whole queue ahead of it. A batch of healthy but slow appends tripped it, and
writes that never reached the store were reported as unknown outcomes. The
bound now starts when an append reaches the head. An append that does not
settle in time has an unknown outcome and fences the store. The writes behind
it get EventAppendError with the fence as cause, because they never reached
the store.

A fence also notified nobody unless ownership was lost. The runtime stayed
resident and ready and refused every later write until park, dispose or idle
eviction. Every fence now stops the runtime the way ownership loss does. The
manager drops the residency, and the next access reloads from the log, which
decides the uncertain write. The lookup goes by runtime, so a reopened
runtime is dropped too.

Breaking for custom event stores: an append that does not settle within
writeQueueTimeoutMs now counts as an unknown outcome. The EventStore interface
documents the error contract. SessionStore.onOwnershipLost is replaced by
onFenced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
Delivery receipts stay in the projection for the session lifetime, keyed by
the caller's delivery ID, and the ID had no upper bound. A client could store
arbitrarily long keys that bypass the message size limit. sendMessage now
rejects an ID longer than 256 characters before anything is accepted. The
event schema keeps accepting longer IDs so existing logs still replay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
matej21 and others added 5 commits September 28, 2026 21:22
forgetEntry retried clearing tenure.entry once, after waitForLocalCleanup.
That waited only for hooks and local cleanup, while hasUnsafeResources also
counts pending writes, the scheduler tail and tracked resources. A revoke
during a stalled append therefore kept the store, agents and plugin contexts
reachable for good, and retainedRuntimeCount never returned to zero.
Session.whenSafe() now settles once all of those have settled, and
forgetEntry waits for it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…voked id

The services close hook threw whenever a status publication had failed, even
a stale one, although it had already stopped the processes and released the
ports. On revoke that set cleanupFailed, and nothing retries a revoke, so
the runtime stayed unsafe for good: admission refused the id as unloading and
activateSession refused it too. A throw from a close hook now means that
resources may still be held. A publication failure is still reported by a
park, which is retried, and is only logged on every other close.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…cceptance

A fence now stops the runtime, so a retry on the same runtime after an
unknown outcome is refused as session_runtime_unavailable before it writes
anything. It used to reject through the fenced store. Either way the retry
cannot duplicate the message, and a fresh load still replays the committed
receipt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…vates

After ownership loss the manager dropped the cache entry, but the tenure
stayed active. Every later access on the stale host replayed the full log,
ran the ready hooks (services started processes, workers resumed), stepped
the agent, and fenced again on the first write. The host saw only a generic
revoked runtime.

The tenure now moves to a lost state. Access is refused cheaply with the
domain error session_ownership_lost, and parkSession rejects with
SessionOwnershipLostError, until the host calls activateSession again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
…nure refuses

The review fixes changed what the README described: the write queue timeout now
bounds one append at the head of the queue and is settable through createSystem
and the environment, a lost tenure refuses access until reactivation, and two
file stores over one session can delete each other's events, not just
interleave them. Store authors also need the error contract in one place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YKJFaSJo72tboWbPzjECzB
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant