Skip to content

chore(deps): bump torch to 2.14 and unify CUDA wheels on cu130 - #1647

Merged
TATP-233 merged 1 commit into
mainfrom
chore/torch-2.14-cu130
Sep 26, 2026
Merged

TATP-233 merged 1 commit into
mainfrom
chore/torch-2.14-cu130

Conversation

@TATP-233

@TATP-233 TATP-233 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • torch 依赖由两行平台拆分合并为一行 torch>=2.9,<2.15;tool.uv.sources 统一走 r2-cu130(cu128 索引最高只有 torch 2.11),删除 pytorch-cu128 索引
  • uv.lock:torch 2.8.0 / 2.8.0+cu128 / 2.9.0+cu130 → 2.14.0(darwin/win 之外 PyPI)+ 2.14.0+cu130(linux/win),triton 3.8.0,nvidia 依赖 cu12 → cu13
  • ROCm 同步:pyproject.rocm.toml torch==2.14.0 + triton-rocm==3.8.0(torch 2.14 官方配套 triton~=3.8);放宽 setuptools<70 遗留 pin(与 torch 2.14+rocm7.2 的 setuptools>=77 冲突);重新生成 uv.rocm.lock
  • 更新 torch CUDA source 契约测试(tests/scripts/test_torch_cuda_source.py)与文档中 cu128 表述(中英 installation/contributing、CONTRIBUTING.md)
  • 动机:torch 2.14 的 MPS 后端改进对 FlashSAC 有显著增益。实测(M5 Max, g1_walk_flat, mujoco, 40 iter):端到端 16.1k → 21.7k steps/s(+35%),learner 213 → 147 ms/iter;standalone learner 62.7 → 45.6 ms/iter(1.38x)。FastSAC 为 GEMM-bound 无变化(52.4 → 54.9 ms,噪声内)
  • scripts/tools/setup_isaacsim_env.sh 的 cu128 是 IsaacSim 独立 pin 的环境,不在本次范围

Linked Work

  • Issue: n/a
  • Parent roadmap (when applicable): n/a
  • Roadmap declared base (when applicable): n/a
  • Milestone: n/a
  • Base branch: main

Validation

  • make test-all passed on the final local head before this PR was created or updated
  • Additional task-specific validation listed below

Commands actually run:

make check                 # pyright 0 errors;ruff 全绿(注:check-tests 中 tests/utils/test_experiment_tracking.py 的 5 个 F821 来自本地无关未提交改动,HEAD 版本 lint 通过)
make test                  # 1527 passed, 29 skipped;仅有的 2 个失败为本次改动的 torch CUDA source 契约测试,修正后 2 passed
uv run --no-sync pytest tests/scripts/test_torch_cuda_source.py -q   # 2 passed
uv run --no-sync train --algo flashsac --task g1_walk_flat --sim mujoco algo.max_iterations=40 training.no_play=true   # torch 2.14 smoke 正常,Steps/s 21,736

make test-all(slow 套件)本地未跑(M 系列 Mac 无 CUDA,慢速训练矩阵不适用);CUDA cu130 / ROCm rocm7.2 实机验证依赖 CI 与 Linux 机器。

Remote CI route:

Impact

  • Backend impact: none
  • Platform impact: both(macOS/Linux torch 版本升级;Linux CUDA wheel cu128 → cu130,需要驱动支持 CUDA 13)
  • Training effect expected: yes(Apple Silicon 上 FlashSAC +35% 吞吐;CUDA 侧预期持平或有增益,待 CI 验证)

Artifacts

  • W&B: n/a
  • benchmark result: 见 Summary 实测数据

Rebase update

- pyproject: collapse the platform-split torch pins into torch>=2.9,<2.15
- uv.sources: route linux/win torch to the cu130 index (cu128 tops out at
  torch 2.11); drop the pytorch-cu128 index
- uv.lock: torch 2.8.0/2.8.0+cu128/2.9.0+cu130 -> 2.14.0/2.14.0+cu130,
  triton 3.8.0, nvidia deps cu12 -> cu13
- rocm: torch==2.14.0 + triton-rocm==3.8.0 (torch 2.14 pairs with
  triton~=3.8); relax the setuptools<70 typo-era pin which conflicts with
  torch 2.14+rocm7.2 (requires setuptools>=77); regenerate uv.rocm.lock
- update the torch CUDA source contract tests and cu128 doc references

Measured on Apple Silicon (M5 Max, g1_walk_flat, mujoco): FlashSAC
end-to-end 16.1k -> 21.7k steps/s (+35%), learner 213ms -> 147ms/iter;
FastSAC unaffected (GEMM-bound).
@TATP-233
TATP-233 force-pushed the chore/torch-2.14-cu130 branch from d6e4a9e to da1514e Compare September 26, 2026 17:59
@TATP-233
TATP-233 merged commit efaac64 into main Sep 26, 2026
8 checks passed
@TATP-233
TATP-233 deleted the chore/torch-2.14-cu130 branch September 26, 2026 18:09
@TATP-233 TATP-233 mentioned this pull request Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant