Stop naming the removed submit tools in the agent's prompts - #192
Merged
Merged
Conversation
scripts/configs/empiric/ holds only the paper's experiment: the seven arms (approaches.yaml) on the five benchmark settings (envs.yaml) with the shared flags and the principled joint belief (common.yaml), which benchmark.yaml includes. With --round joint_r1 it resolves to the same 35 runs as continual_principled_belief_r1.yaml, over seeds 0-4. A launch names its round with --round or a ROUND key, which suffixes every experiment id, and launch.py refuses a continual launch without one, since those runs auto-resume from their run folders. --envs, --approaches and --seeds pick a subset. EXTENDS, which gave arms round-specific ids, is gone. Deleted: the phased exp_*.yaml configs and their all.yaml menus, oracle.yaml, the dated continual launchers, the menu entries outside the benchmark, relaunch_on_timeout.py (the self-requeue trap covers timeouts) and scripts/domino_debug/. random_actions_pybullet.yaml moves next to the ExoPredicator configs it belongs with. Tests that loaded dated launchers now load the benchmark; the arms outside it (scene package, real-to-sim, from assets, scene only, zero shot) keep their tests with their flags defined there. Doc links to deleted launchers point at the iclr-empiric-submission tag.
The EMPIRIC agent and its baselines all run under the continual protocol. This removes the phased experiment path (learn sessions between online learning cycles, solve sessions per test task) and the code only it reached: - AgentModelBasedApproach, the agent explorers, the subgoal-annotation execution monitor and the sampler learning mixin - the NL world model and GNN dynamics baselines - the Docker sandbox backend and the sampler synthesis tools - the learn-note templates, the oracle-simulator branches, the per-cycle fit diagnostics, and the settings flags nothing reads anymore AgentModelFreeApproach keeps the option model, tool-context sync and checkpointing that the continual arms inherit. Its task solver raises, since the arms play levels through ContinualPlayMixin.play_level. ExoPredicator-era code is untouched.
The phased solve sessions delivered plans through submit_plan and submit_policy, whose capture gate re-ran each plan (repeat rollouts, physics and rule-parameter margins, a necessity check, a robot clearance probe) before accepting it. The continual arms never had these tools. This removes them together with: - the solve and learn prompt templates, their goldens and tests - the attempt deadline and the capture fields on the tool context - the standalone arm's particle override scope - the journal's strategy document and test-phase rollback helpers - the skills' contact_objects hook, which only the clearance probe read - 14 settings flags that lost their last reader The empiric configs stop setting the three removed flags they named; none of them changed behavior. Tool descriptions and probe docstrings the agent can read still mention submit_plan and are left for a separate change, since editing them changes the agent's prompts.
The run_python descriptions, the sim probe's notices and its docstrings still told the agent about submit_plan, submit_policy and the capture gate, which no continual arm has. Plan text now points at skills_execute_plan, whose grammar it shares, and the probe says its runs never act in the environment. The adaptive info-seeking notice now sends the agent to sim.run(plan, physics_sweep=True), which is what arms it. This changes the text the agents read, so runs from here on see different tool descriptions than the paper's runs did.
This was referenced Sep 26, 2026
yichao-liang
force-pushed
the
empiric-submit-tools-cleanup
branch
from
September 28, 2026 06:00
a3b4c01 to
d7c9bac
Compare
yichao-liang
changed the base branch from
empiric-submit-tools-cleanup
to
master
September 28, 2026 07:01
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #191 (base
empiric-submit-tools-cleanup); review and merge after it.The continual arms never had
submit_plan,submit_policyor the capture gate, but text they read still named them.This rewrites that text:
run_pythondescriptions (tools/synthesis.py, and the shared probe description intools/exploration.py): plans use "the same grammar asskills_execute_plan", which shares the probe's parser; nothing the probe runs "acts in the environment"; and the evaluator check comes "before executing the plan".simprobe's notices: a physics sweep with failing points asks for design margin before the plan is executed; the no-widthphysics_sweeperror drops its paragraph aboutsubmit_plan's ensemble margin; the adaptive info-seeking notice sends the agent tosim.run(plan, physics_sweep=True), which is what arms it now.help(sim.run), and theevaluate_trajectoryhelper's docstring.This changes what the agents read: runs from here on see different tool descriptions than the paper's runs did.
A rendered before/after diff of the
run_pythondescriptions (synthesis and exploration tools, with and without the joint belief) shows only the phrases above changing.Tests that pinned the old wording now check for
skills_execute_planand that no description namessubmit_plan.Test plan
run_pythondescriptions before and after; only the intended phrases differ.🤖 Generated with Claude Code