Skip to content

Null-space preference critic: dual-critic PPO for uwlab_rl, plus grad-clip and wandb-resume fixes - #43

Draft
daphnechen wants to merge 7 commits into
UW-Lab:mainfrom
daphnechen:fix/pref-critic-clip
Draft

daphnechen wants to merge 7 commits into
UW-Lab:mainfrom
daphnechen:fix/pref-critic-clip

Conversation

@daphnechen

Copy link
Copy Markdown

Description

Adds a dual-critic PPO variant to uwlab_rl in which a preference reward gets its own critic and
its gradient is projected into the null space of the task gradient:

g = g_task + beta * (I - g_task_hat g_task_hat^T) * g_pref

The intent is that preference pressure selects among success-optimal behaviours rather than trading
against task success. Two fixes found while running it are included.

No new dependencies. No linked issue — this was developed as a research branch rather than from a
filed proposal, and I'm opening it as a draft for that reason.

Empirical status, stated up front

The projection's central claim did not survive its control, and a reviewer should know that
before reading the code. Measured on peg-in-hole (OmniReset, UR5e + Robotiq 2F-85, 4,000 iterations,
8 seeds per dose, 16 at the operating point), the projection was compared against naive mixing —
adding beta * g_pref directly, no projection, available in this branch as
projection_mode="sum" — at matched compliance rather than matched beta, since the two modes
reach different compliance at the same beta:

compliance naive mixing projected Fisher (two-sided)
+0.129 / +0.130 0 / 8 lost 0 / 16 lost p = 1.0
+0.174 / +0.177 8 / 8 lost 8 / 8 lost p = 1.0
+0.190 / +0.183 8 / 8 lost 7 / 8 lost p = 1.0
+0.210 / +0.201 8 / 8 lost 5 / 8 lost p = 0.20

Both modes are safe at roughly +0.13 compliance and lose the task by roughly +0.175; the cliff sits
in the same place with the projection and without it. On this evidence the dose, not the
projection, selects the safe operating point.

To be precise about what that does and does not mean: this is no evidence of an advantage, and it
is not evidence of equivalence — two doses that both lose nothing cannot discriminate between
them, and the single pairing that favours the projection is p = 0.20 and not exactly
compliance-matched.

So the honest framing of this contribution is scale-free preference control with a null-space
option
, not a demonstrated improvement from projecting.

What's here

  • source/uwlab_rl/.../rsl_rl/nullspace/ — nullspace_ppo.py, dual_actor_critic.py,
    dual_storage.py, projection.py, reward_split.py, runner.py
  • source/uwlab_rl/uwlab_rl/rsl_rl/nullspace_cfg.py — config surface, including
    projection_mode in ("gradient", "advantage", "sum"); "sum" is the no-projection baseline
  • source/uwlab_tasks/.../omnireset/config/ur5e_robotiq_2f85/ — agent config for the above
  • source/uwlab_rl/test/nullspace/test_grad_clip_and_pref_norm.py — tests for the clipping and
    preference-normalisation behaviour
  • scripts/.../rsl_rl/train.py — resume the same wandb run across preemption restarts when
    WANDB_RUN_ID is set, instead of opening a new run per restart
  • source/uwlab_assets/uwlab_assets/__init__.py — make the asset cache dir overridable via
    UWLAB_ASSET_CACHE_DIR

A note on size

This is +1,897 lines across 14 files, which is larger than the guidance in the PR template asks
for. Two commits are independently useful and trivially splittable if you'd prefer them separately:

  • 62eb368 — UWLAB_ASSET_CACHE_DIR override (3 lines, unrelated to the rest)
  • 64768df — wandb resume across preemption restarts (17 lines, unrelated to the rest)

Happy to split those out, or to drop the research code entirely and land only the two fixes.

The branch is 7 commits ahead of main and 4 behind, but no file is touched on both sides of the
merge-base, so it merges cleanly without a rebase.

Type of change

  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue)

Checklist

  • I have read and understood the contribution guidelines
  • I have run the pre-commit checks with ./uwlab.sh --format
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • I have updated the changelog and the corresponding version in the extension's config/extension.toml file
  • I have added my name to the CONTRIBUTORS.md or my name already exists there

Unticked items are genuinely not done rather than overlooked: the pre-commit toolchain isn't
installed in this environment, so I've left the format check to CI; and the changelog/version bumps
and CONTRIBUTORS.md entry are held until you say whether you want this branch at all in its
current shape. All new files do carry the standard copyright + SPDX header, and all changed Python
files compile.

🤖 Generated with Claude Code

https://claude.ai/code/session_01BXF4A9BzFr7PX66G1x72yu

daphnechen and others added 7 commits August 28, 2026 22:11
Adds uwlab_rl/rsl_rl/nullspace/:
  projection.py      pure-torch null-space projection + diagnostics (unit-tested, no sim needed)
  dual_actor_critic  second value head; shared-trunk or separate-critic arch
  dual_storage       parallel reward/value/return/advantage buffers, two GAE passes,
                     each stream normalised over its own full batch
  nullspace_ppo      one importance ratio and one clip shared by both surrogates; the
                     preference gradient is projected orthogonally to the task gradient
                     before it reaches the actor; per-stream time-out bootstrapping
  reward_split       task/preference split read off RewardManager._step_reward, with a
                     pluggable preference source (zero / gaussian noise / named terms)
  runner             resolves the above by name (upstream _construct_algorithm uses eval()
                     in its own module namespace)

Why the reward split reads _step_reward instead of restructuring the RewardManager:
omnireset's `progress_context` term returns zeros but caches the goal distances that r_dist,
r_success, three termination terms, the reset-distribution monitor and the data-collection
configs all read back off it -- and RewardManager skips any term whose weight is 0.0 without
calling it. Splitting, reordering or reweighting the manager corrupts the reward silently.

Also ports the parent checkout's uncommitted UWLAB_TMP_DIR fix in get_temp_dir(); without it
a fresh clone dies on /tmp/uwlab permissions after ~2 min of Isaac boot on shared machines.

Registers OmniReset-Ur5eRobotiq2f85-RelCartesianOSC-State-Nullspace-v0 (same env as the
baseline task, so a beta=0 run is comparable line-for-line) and threads the null-space
options through train.py.
$HOME is ephemeral on Singularity, so the hard-coded ~/.cache/uwlab/assets would re-download
~7 GB of USD/HDR assets on every job. Pointing this at a persistent writable mount lets the
first job populate any cache miss and every later job hit it.
…leak fix)

Every scripted manner preference (mechanical power, EE speed, action smoothness) is
monotonically improved by shrinking action noise, while perturbing noise near a local optimum
of the mean policy is ~second order in task return. Large first-order preference gradient
against ~zero first-order task gradient means the projector does not merely permit exploration
collapse -- it selects for it. That is entropy collapse arriving through the exact channel
built to find task-neutral directions.

gSDE noise is an optimisation parameter, not a deployed behavioural property (the policy is
deployed on the mean action), so masking gives up no legitimate preference.

Layer 1 (pref_mask_noise, default ON): exclude the noise params from g_pref and run the
projection *restricted to* the mean-action subspace. Restriction rather than zero-padding
matters: the projector subtracts a multiple of the full g_task, which has components on the
excluded coords, so zero-padding feeds a correction back onto exactly the parameters being
protected. Covered by test_masking_is_not_equivalent_to_zeroing_gpref_in_full_space.

Layer 2 (pref_detach_noise_features, opt-in): Layer 1 does NOT close the leak on its own --
gSDE's variance is mm(actor[:-1](obs)**2, exp(log_std)**2), so the preference can still shrink
exploration by reshaping the trunk features feeding the noise head. Detaching both inputs to
the variance inside the preference surrogate closes it. The resulting ratio is value-identical
to the task ratio (detach changes no values), so "one ratio, one clip" still holds; a runtime
assertion enforces that.

Also:
- ActionRatePreference: the noise-bait probe. r_pref = -||a - a_prev||^2, maximally satisfiable
  by shrinking noise and barely satisfiable otherwise. Sanity run B (Gaussian) is blind to this
  failure -- an unsystematic gradient does not preferentially shrink noise -- so the Gaussian
  run is retained but is no longer the sharp test.
- Guard metrics in EVERY arm including beta=0 (guard_entropy, guard_noise_std, and the active
  scoping flags), since the collapse is only visible as drift relative to the beta=0 reference.
- Entropy bonus documented as staying on the task side: it is a regulariser on the optimisation,
  not a preference about behaviour, and must keep its unrestricted path to the noise params.
- projection_mode now has three distinct behaviours (`sum` and `advantage` previously shared
  one code path). `sum` is arm B': dual critic, separately normalised advantages, NO
  projection -- the arm that separates scale-invariance from the projection itself. Without it
  a C-beats-B result is consistent with either being the contribution.
- pref_removed_frac / cos_before logged as a TIME SERIES in every arm, including beta=0 and the
  unprojected ablations, via a diagnostic-only backward that applies nothing. Alignment moves
  over training and the late-training regime (g_task -> 0) is where the self-scheduling claim
  lives; a converged scalar is not evidence for it.
- EndEffectorHeightPreference: a deliberately high-conflict preference ("keep the EE low" fights
  the lift the task requires). The action-rate bait measures pref_removed_frac ~ 1e-3, i.e. the
  constraint is barely binding -- in high-dimensional parameter space near-orthogonality is the
  default, so on such preferences the projected and unprojected arms are near-identical BY
  CONSTRUCTION and Phase 2 would produce no result however well the code works. Screen candidate
  preferences on pref_removed_frac and span a range.
guard_noise_std is a product -- realised std ~ ||f(obs)|| * sigma -- and Layer 1 masks only
sigma. The aggregate therefore cannot distinguish 'the trunk reshaped its features' from
'sigma shrank', which are different findings with different remedies. Log both factors
(guard_feat_norm, guard_sigma) per iteration in every arm so the Layer 2 decision is
mechanical rather than a judgement call.

Also asserts Layer 1's claim at runtime: the combined gradient must equal g_task EXACTLY on
the excluded noise coordinates. If sigma moves in a masked arm that is an implementation bug,
not a leak, and the two must never be confused when reading the guard metrics.
…he preference reward

The beta=0 null was not a null. NullspacePPO inherited upstream rsl_rl's single global
clip_grad_norm_ over the whole policy, so the preference critic's gradient entered the same
norm as the actor's. The preference critic regresses onto unnormalised returns, and with
pref_source=action_rate its value loss reached 1e5-1e6 against max_grad_norm=1.0. That coupled
the preference critic to the actor even at beta=0 -- and because -||a - a_prev||^2 scales with
action noise, the coupling was strongest in exactly the arms that preserve exploration
(NOTES 29).

Two changes, both on by default:

- grad_clip_mode="per_group": clip the actor, the task critic and the preference critic each
  on its own budget. "global" keeps the old behaviour so pre-fix runs can be reproduced.
- normalize_pref_reward=True: scale the preference reward by a running std of its discounted
  sum (rsl_rl's EmpiricalDiscountedVariationNormalization) before it reaches the preference
  critic. Per-stream advantage normalisation already makes the actor's preference signal
  scale-free, so this fixes the critic's scale without changing what the actor sees. The task
  stream is not normalised. The normaliser is registered on the policy, so it is checkpointed
  and survives auto-resume.

Pre-clip gradient norms are now logged per group (grad_norm/{actor,critic,critic_pref}, or
grad_norm/global), so a cross-group throttle is directly visible rather than inferred.

Also corrects two docstrings that promised what the bug broke: critic_arch="separate" did not
by itself guarantee the beta=0 overlay, and beta=0 reproduced baseline PPO only for
pref_source="zero".

Validated:
- 10 new unit tests (test_grad_clip_and_pref_norm.py): the bug under a global clip; actor
  isolation, and invariance of the actor gradient to preference-critic scale, under per-group
  clipping; the critic_pref/critic prefix trap; normaliser behaviour on large and on zero
  rewards; updates under torch.inference_mode; a torch.save/load round trip.
- Local A/B smoke, beta=0 with pref_source=action_rate, 6 iterations at 256 envs. Legacy global
  clip: preference value loss 7.1e3 rising to 1.9e4, global grad norm 126-478 against
  max_grad_norm=1.0. Fix: preference value loss 0.17-6.5, actor grad norm 0.27-0.53 (never
  clipped), normaliser buffers present in the checkpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BXF4A9BzFr7PX66G1x72yu
Amulet specs set a stable WANDB_RUN_ID with WANDB_RESUME=allow so a job that is
preempted and restarted continues one wandb run instead of opening a new one per
start. On resume wandb loads the run's stored config, and rsl_rl's
WandbSummaryWriter then re-sends its own config containing values that differ on
every start (log_dir, env_cfg.log_dir, resume settings). wandb raises ConfigError
on a changed value unless allow_val_change=True, which killed the restart before
training began.

Patch Config.update in this process only, and only when resuming an explicitly
named run, so those config updates overwrite instead of raising.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BXF4A9BzFr7PX66G1x72yu
@github-actions github-actions Bot added bug Something isn't working asset labels Sep 26, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

asset bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant