Skip to content

roadmap: tensor-only Manager-Based runtime and three-backend scope reduction #1811

Description

@TATP-233

Summary

Goal: close the remaining tensor-runtime migration by making the Manager-Based environment the sole training entry point, remove the remaining generic NumPy Manager boundary and env.tensor_runtime switch, and reduce the active scope to mujoco, mjwarp, and genesis.

Recommended path: first extend the public unisim-core contracts and remove private/backend-name probing; then complete the tensor-native Manager; then absorb the scoped direct FlashSAC motion-tracking runtime into Manager-owned fused owner terms; only then remove tensor_runtime and shelve the out-of-scope backends.

Delivery boundary: no compatibility seam is intended in the final state. There is no NpEnv, no tensor_runtime=false, no direct task-owned env registration, and no runtime task code that dispatches on backend identity or backend-private attributes.

Estimated review scale: large. This is a structural migration requiring several independently reviewable PRs and at least one new ADR.

Durable responsibility: UniLab owns the Manager runtime and task-owner contracts; unisim-core owns public backend tensor capabilities; unilab_rl owns the injected off-policy runner contract.

Decisions required before implementation begins:

  1. Whether a fused Manager-owned owner term is considered part of the sole Manager-Based API even though it bypasses per-term generic execution.
  2. Whether backend scope reduction is a permanent support decision or a temporary runtime scope for this roadmap.
  3. Whether mujoco remains a supported production backend after the tensor-only cutover, or is retained only as the canonical host-bridge/reference backend.
  4. Whether removal of env.tensor_runtime is a breaking release boundary or can ride the next minor.
  5. What RNG owner and reproducibility contract replaces the env-owned NumPy generator.
  6. Whether selected-reset readiness is a strong public postcondition or an optional/diagnostic-only optimization.
  7. What performance regression budget is acceptable while absorbing the direct FlashSAC runtime.

Decisions confirmed by the roadmap owner

These choices are now the execution direction for this roadmap. They remain linked here rather than reopened during implementation unless repository evidence invalidates them.

Decision Choice Boundary and consequence
Fused owner execution Fused Manager-owned terms are part of the Manager-Based API Fused observation/reward/reset components may bypass per-term generic execution, but only when registered through Manager config, executed inside the Manager lifecycle, protected by an owner semantic fingerprint, and consuming public SimBackend APIs. They are not a direct environment entry point and do not receive backend-specific branches.
Backend scope Temporary runtime gate, not permanent removal Restrict this runtime to mujoco, mjwarp, and genesis. Keep out-of-scope adapters in unisim-core during the migration. Re-enabling each backend later requires its own capability, parity, and support-matrix evidence. Permanent adapter deletion is a separate maintainer decision.
MuJoCo role Remains a production backend MuJoCo continues as the canonical in-process HOST_BRIDGE implementation. Its packed state/control/reset boundaries remain first-class contract surfaces, not a test-only fallback.
Breaking boundary Remove env.tensor_runtime in the next intended release without an alias This is a deliberate breaking config change. Owner YAML, tests, docs, run-config schema, and changelog are updated together. No deprecated dual-mode field, hidden fallback, or compatibility alias is retained.
RNG contract One Manager-owned Torch generator, seeded by the owner seed The Manager owns the deterministic device RNG used by tensor observations, commands, events, and selected reset. Seed/reset state and selected-row reproducibility are explicit contract tests. Legacy NumPy-stream parity is not a compatibility requirement.
Selected-reset readiness Strong public postcondition Once set_state_tensor returns, later public state/sensor views are authoritative. MJWarp may publish lazily at the first view; Genesis may publish immediately; MuJoCo keeps its paired selected reset/read boundary. No Manager/task readiness physics step remains. Runtime diagnostics describe implementation details but never control dispatch.
Performance budget Migration allowance is explicitly pessimistic See Performance guardrails: architectural child PRs receive a bounded, non-cumulative migration allowance, while closeout and the mandatory optimization phase still require convergence toward the Phase 0/direct baseline. The original 5% estimate was too optimistic for absorbing a highly fused direct runtime; the updated thresholds reserve room for transient Manager dispatch, reset-transaction staging, and small-kernel overhead without treating that regression as acceptable final performance.

Current State

The prior NpEnv removal roadmap (#1701) removed the legacy environment class and moved the public lifecycle to TorchEnv/TorchEnvState. Four of its five slices are closed. The remaining open slice is #1703, specifically event/randomization tensor transactions, command term internals, explicit RNG provenance, and selected-row reset semantics.

The current runtime is still hybrid rather than fully tensor-native:

  • Generic Manager execution still has host/NumPy boundaries in event terms, reset composition, row handling, and selected-row reset.
  • env.tensor_runtime remains a per-owner switch.
  • FlashSAC g1_motion_tracking uses a task-owned direct runtime (TorchG1MotionTrackingFlashSACEnv) rather than the generic Manager entry point.
  • Task code still contains one backend-private _sensor_map probe and backend-name readiness branches (refactor(runtime): replace backend-name readiness hacks with declared lifecycle capabilities #1790).
  • Active backend breadth exceeds the scope that can be validated during this migration.

Recent profiling on the low-clock multi-GPU server showed why finishing this correctly matters. Reusing stable device-resident sensor views and eliminating redundant DLPack boundaries improved collector throughput by roughly 22–25% on G1 walk owners. The remaining cost is concentrated in Manager Python term dispatch, reset composition, and readiness barriers rather than MJWarp physics. A blanket conversion to more small Torch operations is not sufficient; the reset/update path needs fewer, larger public operations.

End State

  • make_manager_based_rl_env() is the sole production training environment entry point.
  • ManagerBasedRlEnv executes update, reset, observation, reward, termination, command, event, curriculum, metric, and recorder paths without a generic NumPy Manager carrier.
  • Public Manager state and row selectors are Torch tensors. NumPy may exist only inside a backend adapter that declares a host-bridge data plane.
  • env.tensor_runtime and env.tensor_runtime_device no longer exist. Tensor execution is invariant for Manager-Based training; only backend-declared data planes differ.
  • Scoped task owners may register fused Manager-owned observation/reward/reset terms for performance, but they must execute through the Manager lifecycle and consume only public SimBackend APIs.
  • The active backend scope is mujoco, mjwarp, and genesis. Other adapters are gated out of this runtime until separately re-enabled.
  • Single-GPU mjwarp training does not require setting CUDA_VISIBLE_DEVICES; explicit topology controls remain available for multi-GPU and debugging.
  • Selected reset leaves public state/sensor views authoritative without an extra control step for MJWarp and Genesis.
  • No task or runtime code probes backend-private attributes or dispatches capabilities by backend name.

Non-Goals

  • No new backend is added.
  • No claim that MuJoCo, MJWarp, and Genesis produce identical trajectories.
  • No removal of the full-CPU off-policy profile unless separately decided.
  • No productization of IsaacGym/IsaacSim/new external CUDA IPC owners here.
  • No broad deletion of every optional backend adapter before the new Manager contract is stable.
  • No acceptance of a hidden NumPy fallback or alias during migration.

Dependency And Related Work

Phase Plan

Phase 0 — decision, ADR, and server baseline

Freeze the public contract decisions before touching runtime code. The seven roadmap decisions above are already selected; Phase 0 converts them into the durable ADR and baseline evidence, not into a second design review.

  • Add an ADR covering the sole Manager-Based entry point, tensor-only runtime, direct-runtime absorption, backend scope reduction, and removal of env.tensor_runtime.
  • Decide the exact unisim-core public additions: selected-reset publication, sensor inventory, and public state widths.
  • Record the confirmed backend decision in the ADR: this is a temporary runtime gate; permanent adapter deletion requires a separate decision.
  • Capture a low-clock multi-GPU server baseline for:
    • FlashSAC / g1_motion_tracking / MJWarp
    • FlashSAC / g1_walk_flat / MJWarp
    • FastSAC / g1_walk_flat / MJWarp
    • Corresponding Genesis smoke/performance where available
  • Record collector phase metrics and a CPU profile so later PRs cannot claim success from an unstable baseline.

Exit criteria: approved ADR, frozen contract names/semantics, and reproducible baseline artifacts linked to this issue.

Phase 1 — public backend contracts and governance

This work belongs primarily in unisim-core, with UniLab consuming it only after the contract is released or pinned through the sibling workspace.

  • Add a public selected-reset publication contract to the tensor lifecycle:
    • after set_state_tensor returns, subsequent public state/sensor views are authoritative;
    • backends may implement immediate or lazy publication;
    • callers do not issue an extra physics step to obtain readiness.
  • Declare that contract for MJWarp.
  • Declare and implement that contract for Genesis, relying on its forced state publication rather than a diagnostic-only branch.
  • Preserve the paired selected reset/read boundary for MuJoCo host-bridge execution.
  • Add a public sensor inventory/descriptor API sufficient to replace task-side _sensor_map probing.
  • Add public qpos/qvel state-width access to replace undeclared backend nq/nv attribute probing.
  • Add an architecture gate that rejects backend-private attribute access and backend-name capability dispatch in UniLab task/runtime code.
  • Update refactor(runtime): replace backend-name readiness hacks with declared lifecycle capabilities #1790 so readiness is negotiated only through the declared contract.

Exit criteria: all three scoped backends pass the new contract tests; no UniLab task/runtime private backend probe remains.

Phase 2 — tensor-native Manager core

This completes and supersedes the remaining #1703 work under the narrower backend scope.

  • Convert Manager row APIs to Torch tensor/slice selectors.
  • Keep all Manager public input/output/state carriers as Torch tensors.
  • Convert remaining command term internals to device tensors.
  • Convert remaining event/randomization transactions to device tensors.
  • Convert curriculum, metric, and recorder reset/update carriers.
  • Replace host reset composition with a tensor reset transaction:
    • selected rows remain device tensors;
    • qpos/qvel/model-field writes are device tensors;
    • host-bridge conversion is owned by the MuJoCo adapter.
  • Implement the Manager-owned Torch RNG decision:
    • one device generator owned by the Manager runtime;
    • initialization from the resolved owner seed;
    • explicit seed/reset behavior;
    • deterministic selected-row reproducibility where the row selector is reproducible;
    • no requirement to preserve the legacy NumPy RNG stream bit-for-bit.
  • Make selected-row reset semantics explicit and parity-tested.
  • Batch all logging/telemetry boundaries so a step performs at most one intentional host flush, controlled by logging interval.
  • Remove generic NumPy input/output branches from Manager execution.

Exit criteria: canonical G1 walk owners run on mujoco, mjwarp, and genesis with no generic NumPy Manager carrier and no hidden host fallback.

Phase 3 — absorb the direct FlashSAC motion runtime

The direct runtime must become Manager-owned functionality rather than a second environment entry point.

  • Move motion command state and adaptive sampling into a Manager command term/component.
  • Move motion reset sampling and full-width qpos/qvel construction into a Manager reset event.
  • Move direct task-owned observation computation into Manager-owned fused observation terms.
  • Move direct task-owned reward/termination computation into Manager-owned fused terms.
  • Preserve the existing semantic-fingerprint protection as a Manager-owner contract.
  • Consume the Phase 1 selected-reset publication contract; do not add a readiness control step.
  • Remove registration of make_torch_g1_motion_tracking_flashsac_env.
  • Delete or internalize TorchG1MotionTrackingFlashSACEnv once no caller remains.
  • Keep preallocated output buffers and packed state reads inside Manager-owned components.

Exit criteria: G1MotionTrackingSAC is registered only through the Manager factory for all scoped backends, and the direct class is no longer a public training path.

Phase 4 — remove the tensor-runtime switch

This is a breaking configuration cleanup and must happen only after Manager execution is the only supported path.

  • Remove ManagerBasedRlEnvCfg.tensor_runtime.
  • Remove ManagerBasedRlEnvCfg.tensor_runtime_device.
  • Remove all tensor_runtime=false branches and cold-proxy opt-outs.
  • Remove runner-side manual tensor-runtime injection.
  • Derive public device placement from backend capability and rank topology.
  • Update owner YAMLs, tests, benchmarks, docs, and run-config schema.
  • Add a migration note describing the removal without reintroducing a compatibility alias.

Exit criteria: production source, configs, tests, scripts, and current docs no longer reference env.tensor_runtime.

Phase 5 — backend scope reduction

  • Gate motrix, newton, drake, isaacgym, isaacsim, and superdex out of the tensor-only Manager runtime with actionable errors.
  • Keep their adapters temporarily present in unisim-core; do not delete unrelated backend code in this roadmap.
  • Restrict UniLab registry/config discovery/CI/benchmarks/docs to mujoco, mjwarp, and genesis.
  • Update the generated support matrix so no shelved backend is presented as supported.
  • Add a separate follow-up decision for permanently deleting adapters and owners.

Exit criteria: CLI/CI/docs expose only the scoped backends, and out-of-scope backends fail closed rather than silently falling back.

Phase 6 — single-GPU default device behavior

  • Make single-GPU MJWarp resolve the current default CUDA device without requiring CUDA_VISIBLE_DEVICES.
  • Keep explicit device selection for multi-GPU training.
  • Fail closed when multiple GPUs are visible and no topology is selected.
  • Preserve explicit CUDA_VISIBLE_DEVICES for launchers, debugging, MPS, and GPU isolation.
  • Update train/eval/benchmark docs and tests.

Exit criteria: a single-GPU host can run the canonical MJWarp training command without setting CUDA_VISIBLE_DEVICES.

Phase 7 — parity, performance, and closeout

  • Add Manager parity tests for observations, rewards, terminations, selected reset, commands, RNG, and episode metrics across the three scoped backends.
  • Add FlashSAC motion-tracking semantic and rollout snapshot tests after direct-runtime absorption.
  • Run make check, make test, and make test-all on the final integration head.
  • Run focused remote CI on the final head when targeting main.
  • Re-run the Phase 0 server benchmarks.
  • Compare against the Phase 0 baseline and classify the result against the staged performance budget below.
    A result outside the migration gate blocks closeout unless the maintainers explicitly accept a documented exception and open the optimization follow-up before merge.
  • Confirm selected reset performs no redundant physics step on MJWarp or Genesis.
  • Confirm backend-private probing and backend-name dispatch gates pass.
  • Update support matrix, tensor-runtime docs, ADR status, and release notes.

Exit criteria: all acceptance criteria below pass and the roadmap can be closed without leaving a parallel direct entry point.

Acceptance Criteria

Architecture

  • make_manager_based_rl_env() is the sole production environment factory for training.
  • No task-owned direct environment class is registered as a training entry point.
  • Generic Manager execution exposes no NumPy public carrier.
  • env.tensor_runtime and env.tensor_runtime_device do not exist.
  • Active backend scope is limited to mujoco, mjwarp, and genesis.
  • No task/runtime code accesses backend-private attributes.
  • No task/runtime capability decision is based on backend name.
  • Public backend behavior is negotiated only through SimBackend capabilities/methods.

Contracts

  • MJWarp and Genesis public views are authoritative after selected reset.
  • MuJoCo host-bridge selected reset keeps its explicit paired read boundary.
  • Sensor inventory and state widths are available through public backend APIs.
  • Unsupported task/backend/term combinations fail closed.

Runtime

  • Canonical G1 walk owners run on all three scoped backends.
  • FlashSAC g1_motion_tracking runs through the sole Manager entry point.
  • Single-GPU MJWarp runs without CUDA_VISIBLE_DEVICES.
  • Multi-GPU topology remains explicit.
  • Logging host boundaries are batched and interval-controlled.

Validation

  • make check passes.
  • make test passes.
  • make test-all passes.
  • Applicable remote CI passes on the final head when the PR base is main.
  • Phase 0 server benchmark comparison is recorded.
  • Final integration head meets the closeout gate in Performance guardrails, or has a maintainer-approved documented exception plus an open optimization issue.
  • Support matrix and documentation match the implemented runtime.

Performance Guardrails

Performance evidence must be captured on the low-clock multi-GPU server, not only on the local high-clock workstation. Local workstation results are reference only and are not the acceptance authority.

Minimum benchmark matrix:

Case Backends
FlashSAC / g1_motion_tracking MJWarp, Genesis
FlashSAC / g1_walk_flat MJWarp, Genesis
FastSAC / g1_walk_flat MJWarp, MuJoCo

Record for every case:

  • collector steps/s;
  • update_state_ms;
  • reset_done_ms;
  • reset_done_reset_call_ms;
  • backend/physics step time;
  • replay/bookkeeping time;
  • CPU profile function-call count;
  • selected-reset redundant-step status;
  • backend tensor-runtime diagnostics where available.

Why the original 5% estimate is too optimistic

Absorbing TorchG1MotionTrackingFlashSACEnv is not a carrier rename. The direct owner currently fuses action transformation, state reads, relative transforms, reward, termination, adaptive command updates, selected reset, final observation construction, and preallocated output assembly. Moving that into Manager-owned components can reintroduce:

  • Manager term-dispatch Python overhead;
  • more small Torch kernels;
  • additional validation/readiness boundaries;
  • reset transaction overhead;
  • logging and state-object copies.

Previous experiments already showed that replacing a compact NumPy reset composer with dozens of small Torch operations can regress selected reset despite removing host transfers. A 5% bound is therefore not a realistic first migration gate for FlashSAC motion tracking.

Staged budget

These percentages are migration and closeout ceilings, not targets. They do not compose across PRs and they do not replace the mandatory optimization phase. A child PR may use its allowance only when parity is unchanged and the measured cause is temporary migration staging; persistent architectural overhead must be fixed before review or the PR must be restructured rather than accumulating debt.

Use two explicit gates rather than one unrealistic final number.

1. Per-phase migration gate

Applies immediately after each architectural PR, before it is merged into the roadmap integration branch:

Workload Maximum acceptable regression versus its immediate pre-PR baseline
FastSAC / g1_walk_flat / MJWarp 10% collector throughput
FlashSAC / g1_walk_flat / MJWarp 10% collector throughput
FlashSAC / g1_motion_tracking / MJWarp 20% collector throughput during Phases 2–3

Additional per-phase constraints:

  • no increase in reset_done_ms greater than 50%, or more than 2 ms absolute, on the benchmark server (whichever is smaller);
  • no new per-step device synchronization in the selected-reset path;
  • no regression in the semantic/rollout parity tests.

A child PR exceeding this gate is rolled back or restructured by default. It may be merged beyond the ceiling only with a maintainer-approved comment that identifies the measured cause, explains why the overhead is transient, and links an already-open optimization task with an owner; that exception applies to one integration slice and must be revisited on the next measured child PR.

Measurement boundary:

  • collect paired pre/post samples in the same server session, using the same GPU, process topology, environment count, policy/collector settings, asset cache, CUDA/MPS setting, and CPU governor state;
  • record the CPU governor/topology state even when it cannot be controlled;
  • report at least three short, interleaved pre/post samples plus the observed min/median/max, not only best runs; use bounded sampling windows and do not add warm-up sleeps or full training runs just to enlarge the sample;
  • report collector steps/s and the listed phase metrics together, so a throughput gain cannot hide a reset or update-state regression;
  • keep local high-clock workstation results reference-only and outside acceptance;
  • a result inside run-to-run noise passes only if its upper regression estimate is also inside the applicable ceiling.

2. Roadmap closeout gate

Applies to the final integration head after direct-runtime absorption:

Workload Closeout requirement
FastSAC / g1_walk_flat / MJWarp no worse than 10% below Phase 0 baseline throughput
FlashSAC / g1_walk_flat / MJWarp no worse than 10% below Phase 0 baseline throughput
FlashSAC / g1_motion_tracking / MJWarp no worse than 25% below Phase 0 direct-runtime baseline throughput
FlashSAC / g1_motion_tracking reset no more than 5 ms, or 50%, slower than Phase 0 direct runtime on the benchmark server (whichever is smaller)

The 25% motion-tracking budget is a deliberate one-time migration allowance, not a target and not a support claim. It exists so the sole-entry-point architecture can land without preserving a parallel direct mode when the residual overhead has a diagnosed owner and a funded optimization path. If the measured regression exceeds it, closeout is blocked and one of the following must happen before final merge:

  1. optimize the Manager-owned fused terms until the gate is met;
  2. narrow Phase 3 by moving more computation into a single fused Manager component;
  3. obtain an explicit maintainer-approved exception and open a time-boxed optimization issue with a performance owner.

The closeout table cannot be satisfied by weakening benchmark length, environment count, parity checks, or profiler instrumentation. If server time is insufficient to collect the full matrix, record the available paired subset and keep the roadmap open; do not extrapolate the missing cases.

At the end of Phase 2 and again after Phase 3, measure the integration head against the Phase 0/direct baseline as a cumulative-drift checkpoint. The closeout ceilings apply provisionally to those checkpoints. If a checkpoint exceeds its ceiling, stop further architectural child PRs on that workload until the regression is optimized away or the affected slice is restructured; per-child allowances cannot be chained to normalize cumulative loss.

Required optimization follow-up

The expected end state after migration is not a 25% regression. A separate follow-up issue must be opened after Phase 3 even if the closeout gate is met. Its target is:

  • FlashSAC g1_motion_tracking throughput at or above the Phase 0 direct-runtime baseline;
  • selected reset at or below the Phase 0 direct-runtime timing;
  • no increase in low-clock-server CPU function-call count versus Phase 0.

The optimization issue must be opened when Phase 3 merges, even if Phase 3 is faster than the closeout ceiling. It remains open until at least one repeated low-clock-server run meets those targets or a maintainer explicitly re-scopes the target with new evidence. The roadmap issue may close only with that follow-up linked and its acceptance state recorded.

Likely follow-up techniques are fixed-capacity reset graphs, fused observation/reward kernels, stronger state-read publication, and batched finite diagnostics. None of those optimizations may reintroduce a direct environment entry point or backend private probing.

Rollout And Branch Strategy

Use the repository roadmap workflow:

  1. Create dev/issue-<this-issue>-tensor-manager from the approved base.
  2. Create each child branch from the integration branch.
  3. Target child PRs at the integration branch.
  4. Validate each child against the current integration head before review.
  5. Run focused checks and make test-all on each child head.
  6. Revalidate the integration head after each merge.
  7. Open the final integration PR to the declared base.
  8. Run remote CI on the final head when the base is main.

No phase should delete the old switch before its replacement path is complete.

Explicit Decisions To Record

The roadmap owner has selected the following direction; the ADR records these as decisions rather than reopening them during implementation:

  • Fused Manager-owned terms are part of the Manager-Based API, not a direct-mode exception.
  • Backend scope reduction is a temporary runtime gate; permanent deletion requires a separate decision.
  • MuJoCo remains a production host-bridge backend.
  • Remove env.tensor_runtime as the intended breaking boundary, with no compatibility alias.
  • Use one Manager-owned Torch RNG seeded from the owner seed; legacy NumPy stream parity is not required.
  • Selected-reset readiness is a strong public postcondition, not diagnostic-controlled behavior.
  • Use the staged performance budget above; the per-child FlashSAC motion-tracking allowance is 20%, the one-time closeout allowance is 25%, with mandatory optimization follow-up.
  • No compatibility alias or hidden fallback will be retained.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions