Skip to content

fix(training): export ONNX at opset 18 and decouple play artifact failures - #1656

Merged
TATP-233 merged 1 commit into
mainfrom
fix/issue-1654-onnx-export-opset18
Sep 26, 2026
Merged

TATP-233 merged 1 commit into
mainfrom
fix/issue-1654-onnx-export-opset18

Conversation

@TATP-233

@TATP-233 TATP-233 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • export_policy_onnx default opset 17 → 18. With torch 2.14 the dynamo exporter emits the onnxscript function library (e.g. aten_isnan from torch.nan_to_num in the uni_rl SAC actor wrapper) at opset 18; requesting 17 made torch version-convert the model back while the function stayed at 18, and the optimizer InlinePass crashed with an opset mismatch — killing the play phase right after training finished.
  • Added nonfatal_play_step (unilab.training.run) and wrapped both post-training artifacts — ONNX/JIT policy export and video rendering — in it across all three play paths (train_offpolicy.py, train_appo.py, train_rsl_rl.py), so one artifact's failure no longer blocks the other or crashes the play command; failures print a warning plus traceback and the run continues.
  • Training behavior unchanged; only the play/export phase is affected.

Linked Work

Validation

  • make test-all passed on the final local head before this PR was created or updated
  • Additional task-specific validation listed below

Commands actually run:

make check                       # ruff + mypy + pyright: all pass
make test-all                    # exit 0 (incl. slow suite; benchmark module-mode 35/35, script-mode 36/36)
uv run pytest tests/training/test_onnx_export.py -q          # 6 passed (incl. new nan_to_num regression test)
uv run pytest tests/training/test_training_helpers.py -q -k nonfatal   # 2 passed
uv run pytest tests/scripts/test_train_scripts.py -m slow -q -k "nonfatal or onnx_export_failure or video_render_failure or skip_onnx"  # 3 passed
uv run train --algo sac --task g1_walk_flat --sim mujoco algo.max_iterations=10
# end-to-end: training -> playback -> policy.onnx exported -> ONNX Runtime verified (max_diff 8.94e-08) -> play_video.mp4 recorded

Remote CI route:

Impact

  • Backend impact: none (export/play orchestration only; verified on mujoco)
  • Platform impact: both (verified on macOS; change is platform-independent)
  • Training effect expected: no (playback-phase artifact handling only)

Artifacts

  • W&B:
  • benchmark result:
  • video / screenshot: logs/fast_sac/G1WalkFlat/2026-09-27_03-24-29_mujoco/play_video.mp4 (local)
  • ONNX / checkpoint: logs/fast_sac/G1WalkFlat/2026-09-27_03-24-29_mujoco/policy.onnx (opset 18, ORT-verified vs PyTorch, max_diff 8.94e-08)

Checklist

  • Added or updated tests where needed
  • Updated docs if behavior or workflow changed (no doc contract change; opset is not documented anywhere)
  • Linked the driving issue
  • Noted any follow-up work explicitly

…lures (#1654)

torch 2.14's dynamo exporter emits the onnxscript function library
(e.g. aten_isnan from torch.nan_to_num in the SAC actor) at opset 18;
requesting opset 17 crashed the optimizer's InlinePass with an opset
mismatch after version conversion. Bump the shared export default to 18.

Also wrap ONNX export and video rendering in a shared nonfatal_play_step
guard across the off-policy, APPO, and rsl-rl play paths so one artifact
failure no longer blocks the other.
@TATP-233
TATP-233 requested a review from caozx1110 as a code owner September 26, 2026 19:26
@TATP-233
TATP-233 merged commit 26c0ade into main Sep 26, 2026
8 checks passed
@TATP-233
TATP-233 deleted the fix/issue-1654-onnx-export-opset18 branch September 26, 2026 19:32
@TATP-233 TATP-233 mentioned this pull request Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant