Rename the predicatorv3 configs to empiric with rounds and subsets - #189
Merged
Merged
Conversation
scripts/configs/empiric/ holds only the paper's experiment: the seven arms (approaches.yaml) on the five benchmark settings (envs.yaml) with the shared flags and the principled joint belief (common.yaml), which benchmark.yaml includes. With --round joint_r1 it resolves to the same 35 runs as continual_principled_belief_r1.yaml, over seeds 0-4. A launch names its round with --round or a ROUND key, which suffixes every experiment id, and launch.py refuses a continual launch without one, since those runs auto-resume from their run folders. --envs, --approaches and --seeds pick a subset. EXTENDS, which gave arms round-specific ids, is gone. Deleted: the phased exp_*.yaml configs and their all.yaml menus, oracle.yaml, the dated continual launchers, the menu entries outside the benchmark, relaunch_on_timeout.py (the self-requeue trap covers timeouts) and scripts/domino_debug/. random_actions_pybullet.yaml moves next to the ExoPredicator configs it belongs with. Tests that loaded dated launchers now load the benchmark; the arms outside it (scene package, real-to-sim, from assets, scene only, zero shot) keep their tests with their flags defined there. Doc links to deleted launchers point at the iclr-empiric-submission tag.
This was referenced Sep 26, 2026
3 tasks done
yichao-liang
added a commit
that referenced
this pull request
Sep 30, 2026
Docs and comments reached the launchers and the Domino band probe that #189 deleted through the iclr-empiric-submission tag. The tag is moving to the refactored master, where those files no longer exist, so the 24 mentions now name commit 72d7258, the tag's current target, which stays in master's history.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
scripts/configs/predicatorv3/(51 files) becomesscripts/configs/empiric/, which holds only the paper's experiment:common.yaml: the shared flags, including the principled joint belief (belief_joint_draws: 16, no carried posterior).envs.yaml: the five benchmark settings (Balloons composition, four-span Bridge, two-jug Boil, ramp Fan, Domino high-friction turn).approaches.yaml: the seven arms (EMPIRIC, model-free, model-free + scene package, standalone, oracle dynamics, no fitting, no uncertainty).benchmark.yaml: the default launcher, which includes the three, over seeds 0-4.With
--round joint_r1,benchmark.yamlresolves to the same 35 runs ascontinual_principled_belief_r1.yamldid: same experiment ids, approaches, envs, args and flags. The only differences are seeds 0-4 instead of 0-1, and two top-level flags that equal their defaults.Launching:
--roundor aROUNDkey, which suffixes every experiment id (<env>-<arm>_<round>).launch.pyrefuses a continual launch without one, since those runs auto-resume from their run folders.--envs,--approachesand--seeds N-Mpick a subset.benchmark.yamland setsROUND.EXTENDS, which gave arms round-specific ids, is removed.Deleted:
exp_*.yamlconfigs, theirall.yamlmenus,common.yamlandoracle.yaml;continual_*launchers and the menu entries outside the benchmark (the scene-only, zero-shot, from-assets, real-to-sim, EMPIRIC + scene package and Sonnet entries; their classes stay);scripts/engaging/relaunch_on_timeout.py, whose job the self-requeue trap now does;scripts/domino_debug/.random_actions_pybullet.yamlmoves next to the ExoPredicator configs it belongs with, andgenerate_random_action_gifs.pynow finds it.Tests that loaded dated launchers now load the benchmark. The arms outside it keep their tests, with their flags defined in the tests. Doc links to deleted launchers point at the files under the
iclr-empiric-submissiontag, and code comments that pointed at the band probe name it at that tag too.Test plan
test_continual_comparison_approach.pypasses (41 tests) and the other shards' failures were only that helper.benchmark.yaml --round joint_r1againstcontinual_principled_belief_r1.yaml: identical runs apart from the seeds.🤖 Generated with Claude Code