A workspace command can stop for five different reasons, but a client (and the model reading the result) cannot reliably tell them apart:
| Why it stopped |
What the client receives today |
| Never started: queue deadline passed |
503 WORKSPACE_QUEUE_TIMEOUT (distinct, good) |
| Execution timeout |
200 with timedOut: true, signal: 'SIGKILL'. The COMMAND_TIMEOUT → 504 mapping in service/src/workspace-tools/router.ts is never produced. |
| Worker's own assignment deadline hit |
EXECUTION_ABORTED, identical to a user cancel |
| User or caller cancelled |
EXECUTION_ABORTED |
| Worker lost mid-run |
The API waits out the full execution deadline, then 504 ASSIGNMENT_EXPIRED "Bridge assignment exceeded its deadline" |
The effective timeout is never echoed. The router clamps timeoutMs to JOB_TIMEOUT silently (Math.min in router.ts). The result protocol (exitCode, signal, stdout, stderr, truncated, timedOut) has no field for the budget that applied. deadlineBudgetMs goes only to the server log.
Observed on a LibreChat trusted-VM deployment: before the bridge advertised maxCommandTimeoutMs, agents requested 110-120 s and were killed at 90 s with no indication why. A lost worker costs the full deadline (up to ~125 s) before the caller learns anything.
Proposal
- Add
appliedTimeoutMs (after every clamp) and a stopReason (exited, timeout, deadline, cancelled, worker_lost) to the command result, and return it on every settlement path.
- Give the worker's own deadline abort a distinct code from a caller cancel.
- Detect worker loss while waiting for settlement (lease or heartbeat check), and fail with a distinct
WORKER_LOST code instead of waiting out the deadline.
- Either produce
COMMAND_TIMEOUT or remove the dead mapping.
- Never re-run a timed-out command automatically. Mutations are not idempotent.
Acceptance
- Each of the five cases above yields a distinct, documented code or
stopReason, with a test per case.
- A result from a clamped request reports the clamped
appliedTimeoutMs.
- Killing the worker process mid-command fails the request within a heartbeat interval, not at the execution deadline.
A workspace command can stop for five different reasons, but a client (and the model reading the result) cannot reliably tell them apart:
WORKSPACE_QUEUE_TIMEOUT(distinct, good)timedOut: true,signal: 'SIGKILL'. TheCOMMAND_TIMEOUT→ 504 mapping inservice/src/workspace-tools/router.tsis never produced.EXECUTION_ABORTED, identical to a user cancelEXECUTION_ABORTEDASSIGNMENT_EXPIRED"Bridge assignment exceeded its deadline"The effective timeout is never echoed. The router clamps
timeoutMstoJOB_TIMEOUTsilently (Math.mininrouter.ts). The result protocol (exitCode,signal,stdout,stderr,truncated,timedOut) has no field for the budget that applied.deadlineBudgetMsgoes only to the server log.Observed on a LibreChat trusted-VM deployment: before the bridge advertised
maxCommandTimeoutMs, agents requested 110-120 s and were killed at 90 s with no indication why. A lost worker costs the full deadline (up to ~125 s) before the caller learns anything.Proposal
appliedTimeoutMs(after every clamp) and astopReason(exited,timeout,deadline,cancelled,worker_lost) to the command result, and return it on every settlement path.WORKER_LOSTcode instead of waiting out the deadline.COMMAND_TIMEOUTor remove the dead mapping.Acceptance
stopReason, with a test per case.appliedTimeoutMs.