Skip to content

Rework router worker sync around job ID labels - #4361

Open
un-def wants to merge 1 commit into
masterfrom
pr_rework_smg_worker_sync
Open

un-def wants to merge 1 commit into
masterfrom
pr_rework_smg_worker_sync

Conversation

@un-def

@un-def un-def commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

The sync matched SMG router workers to jobs by URL and removed every worker that wasn't a ready dstack worker, so users couldn't register workers of their own. It also probed every worker on each sync and deregistered the ones that failed, so a flaky connection from the server made a worker flap even when the router, which reaches it within the cluster, had no problem with it. Probes ran one at a time, and only for the connection mode and runtime type of the workers already registered, which broke services that mix them or switch between them in a rolling deployment.

Now each worker that dstack registers is labeled with its job ID (dstack.ai/job-id), and the sync works in terms of jobs:

  • A worker is probed only while some router lacks it. Once it's registered, its health is left to the router, which checks it anyway. A worker that is truly unreachable has its job terminated, which deregisters it.
  • A worker is removed once its job is no longer running. Workers that dstack can't map to a job are left alone, so users can register their own.
  • Workers registered before the label existed are mapped to jobs by address. The ones whose job has already stopped are no longer removed: the router marks them unhealthy, and a rolling deployment of the router drops them.
  • Unregistered workers are probed for HTTP, SGLang gRPC and vLLM gRPC at once, with no hints from the registered ones.
  • Routers and workers are processed in parallel, up to 8 replicas at a time. An unexpected error in one of them is logged and no longer aborts the sync for the others.

Also:

  • gather_async() is added, and gather_map_async() is built on it. Both accept max_concurrency, and cancel the remaining awaitables when one fails unless return_exceptions is set.
  • Router responses are validated with pydantic.

The sync matched SMG router workers to jobs by URL and removed every
worker that wasn't a ready dstack worker, so users couldn't register
workers of their own. It also probed every worker on each sync and
deregistered the ones that failed, so a flaky connection from the
server made a worker flap even when the router, which reaches it
within the cluster, had no problem with it. Probes ran one at a time,
and only for the connection mode and runtime type of the workers
already registered, which broke services that mix them or switch
between them in a rolling deployment.

Now each worker that dstack registers is labeled with its job ID
(`dstack.ai/job-id`), and the sync works in terms of jobs:

* A worker is probed only while some router lacks it. Once it's
  registered, its health is left to the router, which checks it
  anyway. A worker that is truly unreachable has its job terminated,
  which deregisters it.
* A worker is removed once its job is no longer running. Workers that
  dstack can't map to a job are left alone, so users can register
  their own.
* Workers registered before the label existed are mapped to jobs by
  address. The ones whose job has already stopped are no longer
  removed: the router marks them unhealthy, and a rolling deployment
  of the router drops them.
* Unregistered workers are probed for HTTP, SGLang gRPC and vLLM gRPC
  at once, with no hints from the registered ones.
* Routers and workers are processed in parallel, up to 8 replicas at
  a time. An unexpected error in one of them is logged and no longer
  aborts the sync for the others.

Also:

* `gather_async()` is added, and `gather_map_async()` is built on it.
  Both accept `max_concurrency`, and cancel the remaining awaitables
  when one fails unless `return_exceptions` is set.
* Router responses are validated with pydantic.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant