Skip to content

chore(deps): bump unilab-rl to 1.4.1 - #1655

Merged
TATP-233 merged 2 commits into
mainfrom
chore/unilab-rl-1.4.1
Sep 26, 2026
Merged

TATP-233 merged 2 commits into
mainfrom
chore/unilab-rl-1.4.1

Conversation

@TATP-233

Copy link
Copy Markdown
Collaborator

Summary

Bump the optional unilab-rl (uni_rl) runtime dependency from 1.4.0 to 1.4.1 (latest release on PyPI). No driving issue — routine dependency refresh.

unilab-rl 1.4.1 (release notes) is not a pure packaging bump: it replaces the timing/* / lowercase perf/* TensorBoard/wandb tags with a canonical metric schema (retired keys fail closed), removes the reward/reward_metrics logger interfaces, and adds log_interval throttling (UniLab configs already expose training.log_interval). This PR therefore also migrates UniLab's off-policy metric consumers:

  • pyproject.toml / pyproject.rocm.toml: uni_rl extra pin and dev-group pin (both must agree or uv lock has no solution)
  • uv.lock / uv.rocm.lock: surgical pin update to unilab-rl 1.4.1 (version, specifiers, sdist/wheel URLs+hashes from a real uv lock resolve; no resolver churn)
  • tests/utils/test_experiment_tracking.py, tests/algos/test_offpolicy_dp_sync.py: 12 tests updated to the new logger API — log_collector() without mean_reward, return_mean_ep100 instead of reward=, canonical Perf/* / Loss/* / Train/* keys, seconds-vs-milliseconds unit changes, retired-key negative assertions
  • scripts/benchmark/rl/extract_offpolicy_metrics.py, benchmark_offpolicy_dp_scaling.py: tag maps migrated with newest-first (tag, scale) candidates so historical (pre-1.4.1) event files still extract; retired perf/effective_samples_per_sec is derived from run config + Perf/iteration_time when absent (learner_throughput_source records the origin); two retired-without-replacement rows report NaN on new runs by design
  • docs/sphinx/source/{en,zh_CN}/2-user_guide/1-training/3-logging.md, 2-algorithms/2-appo.md, zh_CN 3-sac.md: bilingual docs migrated to the canonical schema with a 1.4.1 old-to-new tag table and unit exceptions; historical run evidence in 5-genesis.md / 7-newton.md / 5-support_matrix.md intentionally untouched

src/ needed no changes: UniLab runtime code does not call the logger APIs directly (verified by grep).

Linked Work

  • Issue: none — routine dependency refresh
  • Base branch: main

Validation

  • make test-all passed on the final local head before this PR was created or updated — see note below on the 2 pre-existing asset failures
  • Additional task-specific validation listed below

Commands actually run (final local head c3c0465):

uv lock --check                      # main lock; ROCm lock checked via pyproject swap
uv sync --extra mujoco --extra motrix --extra uni_rl
uv pip show unilab-rl                # 1.4.1
make check                           # PASS: ruff, mypy 134 files, pyright 0 errors
uv run pytest -m "not slow" -q       # 1557 passed, 27 skipped; 2 pre-existing failures
uv run pytest tests/benchmark/ -q    # 78 passed, 1 skipped
uv run pytest tests/scripts/test_check_docs.py -q   # 19 passed
UNILAB_DOCS_SKIP_AUTODOC=1 sphinx-build -b html -n  # no new warnings

The 2 remaining failures (tests/base/test_sim_backend_smoke.py: allegro_hand YCB meshes not downloaded from the asset hub locally) are the same pre-existing failures documented on main in #1640/#1643, unrelated to this change. Full make test-all locally is blocked by the same missing assets; relying on current-head remote CI as the gate per the contributing workflow.

Remote CI route:

  • Base main: current-head CI recorded below once the run starts.

Impact

  • Backend impact: none (packaging + logging-consumers only)
  • Platform impact: Linux (uni_rl wheels are Linux x86_64); macOS unaffected
  • Training effect expected: yes — training runs on 1.4.1 emit the canonical TensorBoard/wandb metric schema; dashboards/queries keyed on retired timing/* / perf/* tags must use the new tags (migration table in docs/sphinx/source/en/2-user_guide/1-training/3-logging.md). FastSAC/FlashSAC/WarpSAC on NVIDIA CUDA move to a single whole-update-cycle CUDA Graph per the upstream release notes.

Artifacts

  • N/A

Checklist

  • Added or updated tests where needed (12 tests migrated; 7 new benchmark-extraction tests)
  • Updated docs if behavior or workflow changed (bilingual logging/algorithm pages)
  • Linked the driving issue — none; routine refresh
  • Noted any follow-up work explicitly — derived learner throughput in benchmark_offpolicy_dp_scaling.py is an approximation (rank-0 Perf/iteration_time only), documented in the script docstring

@TATP-233
TATP-233 merged commit 9ff999f into main Sep 26, 2026
8 checks passed
@TATP-233
TATP-233 deleted the chore/unilab-rl-1.4.1 branch September 26, 2026 19:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant