Skip to content

Stop naming the removed submit tools in the agent's prompts - #192

Merged
yichao-liang merged 4 commits into
masterfrom
empiric-prompt-submit-refs
Sep 28, 2026
Merged

yichao-liang merged 4 commits into
masterfrom
empiric-prompt-submit-refs

Conversation

@yichao-liang

Copy link
Copy Markdown
Collaborator

Summary

Stacked on #191 (base empiric-submit-tools-cleanup); review and merge after it.

The continual arms never had submit_plan, submit_policy or the capture gate, but text they read still named them.
This rewrites that text:

  • The run_python descriptions (tools/synthesis.py, and the shared probe description in tools/exploration.py): plans use "the same grammar as skills_execute_plan", which shares the probe's parser; nothing the probe runs "acts in the environment"; and the evaluator check comes "before executing the plan".
  • The sim probe's notices: a physics sweep with failing points asks for design margin before the plan is executed; the no-width physics_sweep error drops its paragraph about submit_plan's ensemble margin; the adaptive info-seeking notice sends the agent to sim.run(plan, physics_sweep=True), which is what arms it now.
  • The probe's docstrings, which the agent can read through help(sim.run), and the evaluate_trajectory helper's docstring.
  • Internal comments and docstrings that still described the capture gate as live.

This changes what the agents read: runs from here on see different tool descriptions than the paper's runs did.
A rendered before/after diff of the run_python descriptions (synthesis and exploration tools, with and without the joint belief) shows only the phrases above changing.

Tests that pinned the old wording now check for skills_execute_plan and that no description names submit_plan.

Test plan

  • Rendered the run_python descriptions before and after; only the intended phrases differ.
  • yapf, isort 5.10.1, docformatter, mypy and pylint on this commit (sbatch replay of CI's static jobs).
  • The 8 CI shards in the CI container.

🤖 Generated with Claude Code

scripts/configs/empiric/ holds only the paper's experiment: the seven
arms (approaches.yaml) on the five benchmark settings (envs.yaml) with
the shared flags and the principled joint belief (common.yaml), which
benchmark.yaml includes. With --round joint_r1 it resolves to the same
35 runs as continual_principled_belief_r1.yaml, over seeds 0-4.

A launch names its round with --round or a ROUND key, which suffixes
every experiment id, and launch.py refuses a continual launch without
one, since those runs auto-resume from their run folders. --envs,
--approaches and --seeds pick a subset. EXTENDS, which gave arms
round-specific ids, is gone.

Deleted: the phased exp_*.yaml configs and their all.yaml menus,
oracle.yaml, the dated continual launchers, the menu entries outside
the benchmark, relaunch_on_timeout.py (the self-requeue trap covers
timeouts) and scripts/domino_debug/. random_actions_pybullet.yaml moves
next to the ExoPredicator configs it belongs with.

Tests that loaded dated launchers now load the benchmark; the arms
outside it (scene package, real-to-sim, from assets, scene only, zero
shot) keep their tests with their flags defined there. Doc links to
deleted launchers point at the iclr-empiric-submission tag.
The EMPIRIC agent and its baselines all run under the continual protocol.
This removes the phased experiment path (learn sessions between online
learning cycles, solve sessions per test task) and the code only it
reached:

- AgentModelBasedApproach, the agent explorers, the subgoal-annotation
  execution monitor and the sampler learning mixin
- the NL world model and GNN dynamics baselines
- the Docker sandbox backend and the sampler synthesis tools
- the learn-note templates, the oracle-simulator branches, the per-cycle
  fit diagnostics, and the settings flags nothing reads anymore

AgentModelFreeApproach keeps the option model, tool-context sync and
checkpointing that the continual arms inherit. Its task solver raises,
since the arms play levels through ContinualPlayMixin.play_level.
ExoPredicator-era code is untouched.
The phased solve sessions delivered plans through submit_plan and
submit_policy, whose capture gate re-ran each plan (repeat rollouts,
physics and rule-parameter margins, a necessity check, a robot
clearance probe) before accepting it. The continual arms never had
these tools. This removes them together with:

- the solve and learn prompt templates, their goldens and tests
- the attempt deadline and the capture fields on the tool context
- the standalone arm's particle override scope
- the journal's strategy document and test-phase rollback helpers
- the skills' contact_objects hook, which only the clearance probe read
- 14 settings flags that lost their last reader

The empiric configs stop setting the three removed flags they named;
none of them changed behavior. Tool descriptions and probe docstrings
the agent can read still mention submit_plan and are left for a
separate change, since editing them changes the agent's prompts.
The run_python descriptions, the sim probe's notices and its docstrings
still told the agent about submit_plan, submit_policy and the capture
gate, which no continual arm has. Plan text now points at
skills_execute_plan, whose grammar it shares, and the probe says its
runs never act in the environment. The adaptive info-seeking notice now
sends the agent to sim.run(plan, physics_sweep=True), which is what
arms it.

This changes the text the agents read, so runs from here on see
different tool descriptions than the paper's runs did.
@yichao-liang
yichao-liang force-pushed the empiric-submit-tools-cleanup branch from a3b4c01 to d7c9bac Compare September 28, 2026 06:00
@yichao-liang
yichao-liang changed the base branch from empiric-submit-tools-cleanup to master September 28, 2026 07:01
@yichao-liang
yichao-liang merged commit 372aba7 into master Sep 28, 2026
14 checks passed
@yichao-liang
yichao-liang deleted the empiric-prompt-submit-refs branch September 28, 2026 09:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant