perf(training): training.log_interval 配置项 + rsl_rl PPO TensorBoard 批量写 patch - #1648
Merged
Merged
Conversation
- pyproject: collapse the platform-split torch pins into torch>=2.9,<2.15 - uv.sources: route linux/win torch to the cu130 index (cu128 tops out at torch 2.11); drop the pytorch-cu128 index - uv.lock: torch 2.8.0/2.8.0+cu128/2.9.0+cu130 -> 2.14.0/2.14.0+cu130, triton 3.8.0, nvidia deps cu12 -> cu13 - rocm: torch==2.14.0 + triton-rocm==3.8.0 (torch 2.14 pairs with triton~=3.8); relax the setuptools<70 typo-era pin which conflicts with torch 2.14+rocm7.2 (requires setuptools>=77); regenerate uv.rocm.lock - update the torch CUDA source contract tests and cu128 doc references Measured on Apple Silicon (M5 Max, g1_walk_flat, mujoco): FlashSAC end-to-end 16.1k -> 21.7k steps/s (+35%), learner 213ms -> 147ms/iter; FastSAC unaffected (GEMM-bound).
…ging for rsl_rl PPO - conf (sac/flashsac/appo/ppo): new training.log_interval key (default 1, no behavior change); off-policy builders in unilab-rl read it directly, train_appo forwards it to APPORunner. - experiment.patch_rsl_rl_tensorboard_logging: wrap the rsl_rl logger's SummaryWriter so each iteration's ~15-30 add_scalar calls become a single event record, and gate backend writes to every log_interval iterations (console output unaffected, final iteration always logged); wired up in train_rsl_rl alongside the existing rsl_rl patches. On network filesystems (FUSE/GFS) per-record writes in TensorBoard's writer thread saturate the async queue and block the training loop (~160 ms/iteration for fast_sac g1_walk_flat/mjwarp on RTX 5090). Requires unilab-rl with log_interval support (unilabsim/unilab_rl#44); bump the unilab-rl pin when that release lands. Closes #1646
TATP-233
force-pushed
the
perf/tb-batched-scalar-logging
branch
from
September 26, 2026 18:00
162dd45 to
4bed1f1
Compare
1 of 2 tasks
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
概要
UniLab 侧配套改动,落地 #1646 的优化方案:
training.log_interval配置项 + PPO(rsl_rl)TensorBoard 批量写 patch。uni_rl 侧的批量写/降频实现在 unilabsim/unilab_rl#44,并由
unilab-rl==1.4.0发布。改动
sac/flashsac/warp_sac/appo/ppo五个 ownerconfig.yaml):新增training.log_interval: 1。默认 1、行为不变;off-policy builder 在 unilab-rl 侧直接读取该键。src/unilab/scripts/train_appo.py:透传log_interval给APPORunner。src/unilab/training/experiment.py:新增patch_rsl_rl_tensorboard_logging——包装 rsl_rl logger 的 SummaryWriter,把每 iteration 约 15–30 次add_scalar合并为单条 event record,并按log_interval节流后端写;控制台输出与 episode 簿记不受影响,最后一个 iteration 始终记录。src/unilab/scripts/train_rsl_rl.py:在既有patch_rsl_rl_action_std_logging旁挂接。tests/config/test_config_system.py:断言五个算法 owner config 均暴露有效training.log_interval。依赖与合入顺序
origin/main@d72f7f07,并保持堆叠在 chore(deps): bump torch to 2.14 and unify CUDA wheels on cu130 #1647(da1514e6)之上;请先合入 chore(deps): bump torch to 2.14 and unify CUDA wheels on cu130 #1647,再合入本 PR。pyproject.toml已 pinunilab-rl==1.4.0,该版本包含 perf(logging): TensorBoard 标量按步合并为单条 event,新增 log_interval 降频开关 unilabsim/unilab_rl#44,因此不再有额外 release blocker。training.log_intervalstruct 缺键问题。验证
Rebase 后最终本地 head(
4bed1f18):drake_uni.runtimeimport 无法解析,非本 PR 引入。tests/scripts/test_check_docs.py::test_documentation_files_match_current_repo_contracts失败,已在干净origin/main@d72f7f07复现,属 main 既有问题;本 PR 不顺手修改以保持范围。g1_walk_flat/mjwarp,fast_sac 5000 iterations,日志目录在 GFS/FUSE 上):修复前 cycle 墙钟约 206 ms,其中约 165 ms 被 TensorBoard 写盘阻塞;修复后墙钟与埋点时间一致,collector/wait_for_learner_action从约 156 ms 降到 33 ms。Closes #1646