chore(deps): bump unilab-rl to 1.4.1 - #1655
Merged
Merged
Conversation
This was referenced Sep 26, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bump the optional
unilab-rl(uni_rl) runtime dependency from 1.4.0 to 1.4.1 (latest release on PyPI). No driving issue — routine dependency refresh.unilab-rl 1.4.1 (release notes) is not a pure packaging bump: it replaces the
timing/*/ lowercaseperf/*TensorBoard/wandb tags with a canonical metric schema (retired keys fail closed), removes thereward/reward_metricslogger interfaces, and addslog_intervalthrottling (UniLab configs already exposetraining.log_interval). This PR therefore also migrates UniLab's off-policy metric consumers:pyproject.toml/pyproject.rocm.toml:uni_rlextra pin and dev-group pin (both must agree oruv lockhas no solution)uv.lock/uv.rocm.lock: surgical pin update tounilab-rl 1.4.1(version, specifiers, sdist/wheel URLs+hashes from a realuv lockresolve; no resolver churn)tests/utils/test_experiment_tracking.py,tests/algos/test_offpolicy_dp_sync.py: 12 tests updated to the new logger API —log_collector()withoutmean_reward,return_mean_ep100instead ofreward=, canonicalPerf/*/Loss/*/Train/*keys, seconds-vs-milliseconds unit changes, retired-key negative assertionsscripts/benchmark/rl/extract_offpolicy_metrics.py,benchmark_offpolicy_dp_scaling.py: tag maps migrated with newest-first(tag, scale)candidates so historical (pre-1.4.1) event files still extract; retiredperf/effective_samples_per_secis derived from run config +Perf/iteration_timewhen absent (learner_throughput_sourcerecords the origin); two retired-without-replacement rows report NaN on new runs by designdocs/sphinx/source/{en,zh_CN}/2-user_guide/1-training/3-logging.md,2-algorithms/2-appo.md, zh_CN3-sac.md: bilingual docs migrated to the canonical schema with a 1.4.1 old-to-new tag table and unit exceptions; historical run evidence in5-genesis.md/7-newton.md/5-support_matrix.mdintentionally untouchedsrc/needed no changes: UniLab runtime code does not call the logger APIs directly (verified by grep).Linked Work
mainValidation
make test-allpassed on the final local head before this PR was created or updated — see note below on the 2 pre-existing asset failuresCommands actually run (final local head c3c0465):
The 2 remaining failures (
tests/base/test_sim_backend_smoke.py: allegro_hand YCB meshes not downloaded from the asset hub locally) are the same pre-existing failures documented onmainin #1640/#1643, unrelated to this change. Fullmake test-alllocally is blocked by the same missing assets; relying on current-head remote CI as the gate per the contributing workflow.Remote CI route:
main: current-head CI recorded below once the run starts.Impact
timing/*/perf/*tags must use the new tags (migration table indocs/sphinx/source/en/2-user_guide/1-training/3-logging.md). FastSAC/FlashSAC/WarpSAC on NVIDIA CUDA move to a single whole-update-cycle CUDA Graph per the upstream release notes.Artifacts
Checklist
benchmark_offpolicy_dp_scaling.pyis an approximation (rank-0Perf/iteration_timeonly), documented in the script docstring