Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/amps/empiric-from-assets.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ This is a reduction in supplied domain implementation, not reconstruction from r

## Pilot

The configuration is [continual_from_assets_pilot_r1.yaml](../../scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml).
The configuration is [continual_from_assets_pilot_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml).
It runs Opus 5 on seeds 0 and 1 of Fan + ramp, four-span Bridge, Domino, Balloons and Boil.
There are ten intended runs total, each with its training and test levels.
Fan uses the previously reviewed 3 mm ramp with the 10 cm landing extension and the repaired shared skills, not Fan maze.
Expand All @@ -49,7 +49,7 @@ Bridge seeds 0-1 are array `23407427`; Fan maze seeds 0-1 are array `23407428`.
The user immediately corrected Fan maze to Fan + ramp; both tasks of array `23407428` were cancelled with their logs preserved.
The abandoned maze runs are not part of the intended pilot and must not be resumed or counted as task failures.
Bridge array `23407427` remains unchanged.
The [expansion-only configuration](../../scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml) submits only the eight new tasks, with Bridge and Fan maze disabled to prevent duplicates.
The [expansion-only configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml) submits only the eight new tasks, with Bridge and Fan maze disabled to prevent duplicates.
Its runtime menus are pinned to the frozen implementation.
The resolved Fan + ramp flags match the repaired EMPIRIC ramp cohort, apart from making the default fitting flag explicit.
The eight replacement/additional tasks were submitted on account d as the following two-seed arrays:
Expand Down
6 changes: 3 additions & 3 deletions docs/amps/fan-development.md
Original file line number Diff line number Diff line change
Expand Up @@ -303,20 +303,20 @@ The backup account `dat` also passed but is not in the active pool.
Account `c` has an active limit marker until September 23 at 00:00 UTC and is excluded.
Usage percentages were unavailable from the service endpoint; successful probes establish current access, not a guarantee of sufficient remaining quota for full runs.
Each task requests 8 CPUs and 16 GB, with requeue enabled for preemption, time limits, and recognized account-limit exits.
The launch configuration is [continual_fan_ramp_skill_repair_r1.yaml](../../scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml).
The launch configuration is [continual_fan_ramp_skill_repair_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml).
The shared repair and its measured extra interaction cost are documented in [the switch investigation](fan-switch-seed4-investigation.md).

Seed 0's original job stopped after account `a` reported that its organization had disabled subscription access for Claude Code.
This was an infrastructure interruption after training succeeded and the test reached 207 steps, not a task failure.
With the user's approval, job `23395354` resumes only seed 0 on `dat`, from the same frozen runtime, experiment key, run directory, sandbox, and recorded state.
The resume configuration is [continual_fan_ramp_skill_repair_seed0_resume.yaml](../../scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml).
The resume configuration is [continual_fan_ramp_skill_repair_seed0_resume.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml).

### Six matched comparison arms

The user approved five seeds each for all six other paper agents on this repaired Fan ramp setup.
All 30 tasks were verified running on compute nodes after submission.
They use the same frozen runtime `ff11bc76f4652c6964e73beda7e41a6655d670d9`, reviewed geometry, observation noise, step budget, and preflight-off setting as the repaired EMPIRIC cohort.
The [six-arm configuration](../../scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml) resolves to exactly six five-seed arrays, with only the intended approach and ablation flag differences.
The [six-arm configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml) resolves to exactly six five-seed arrays, with only the intended approach and ablation flag differences.

- Oracle dynamics: array `23398858`, seeds 0-4, accounts b/d.
- Direct agent: array `23398859`, seeds 0-4, account dat.
Expand Down
2 changes: 1 addition & 1 deletion docs/amps/oracle-fan-ramp-cohort-investigation.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,7 @@ The driver is [audit_oracle_ramp_20260922.py](/home/ycliang/predicators/logs/aud

Seeds 9 and 10 were submitted as array `23463117` and started on compute nodes using accounts b and d.
The launch flags and arguments were checked against the previous extra-seed config; only the seed range changes.
The [launch configuration](/home/ycliang/predicators/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml) retains the frozen runtime and original cohort identifier.
The [launch configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml) retains the frozen runtime and original cohort identifier.
The Markdown benchmark tracks all eleven seeds, retaining all failures.
Replacement plot monitor `23463368` refreshes the report and figures every 60 seconds when results change.
The expanded-cohort report tests passed: 8 tests.
Expand Down
2 changes: 1 addition & 1 deletion docs/comparisons/empiric-validation-r2.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# EMPIRIC r2: prospective robustness checks

The requested cohort is two additional seeds per domain, seeds 3 and 4, across Boil, Domino, Balloons, Bridge, and Fan.
The launcher is [continual_empiric_benchmark_r2.yaml](../../scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml).
The launcher is [continual_empiric_benchmark_r2.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml).
Its experiment key is `<domain>-mb_opus_benchmark_r2`, under `agent_continual`.
The main figures retain historical EMPIRIC and show this cohort separately as **EMPIRIC r2**.
All finished outcomes count, including failures; an unfinished run is not a failed run.
Expand Down
2 changes: 1 addition & 1 deletion docs/comparisons/five-seed-launch-review.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ These checks do not establish future solve rates or resolve the previously docum

## Launch and reporting

The launcher is [continual_benchmark_five_seeds.yaml](../../scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml).
The launcher is [continual_benchmark_five_seeds.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml).
All 60 destination seed directories were checked absent before submission.
Jobs use `mit_preemptable`, checkpoint resume, automatic requeue, and account labels `a,b,c,d`, all confirmed usable by the user during this launch review.
The user's account e corresponds to the launcher's `dat` label and is reserved as backup, not included in the normal rotation.
Expand Down
16 changes: 8 additions & 8 deletions docs/comparisons/ten-agent-opus-benchmark.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,20 +46,20 @@ python ../../../scripts/plotting/plot_benchmark_arms.py benchmark-arms-opus

## Settings and cohort selection

The five settings are the menu defaults in [envs/continual.yaml](../../scripts/configs/predicatorv3/envs/continual.yaml): Boil two-jug test, Domino high-friction turn, Fan ramp test, Bridge four-span test row and Balloons composition test levels.
The five settings are the menu defaults in [envs/continual.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/envs/continual.yaml): Boil two-jug test, Domino high-friction turn, Fan ramp test, Bridge four-span test row and Balloons composition test levels.
EMPIRIC and the direct agent ran before the benchmark rounds under their own round keys; the four-span EMPIRIC runs are the preflight-off ones (r2 seed 0, r3 seeds 1-2).
The other arms ran as `<arm>_opus_benchmark_<round>` from frozen worktrees:

- Oracle dynamics, zero-shot and no harness fitting: [continual_five_ablations_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml) at `2982f5876` (/home/ycliang/predicators-five-arms-frozen-20260918).
- Standalone simulator and no explicit uncertainty: round r2 from [continual_standalone_no_uncertainty_r2.yaml](../../scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml) at `0ddc8f7f4` (/home/ycliang/predicators-standalone-nounc-frozen-20260918), after the September 18 revision of both surfaces.
- Scene only: [continual_scene_only_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml).
- Agentic real-to-sim: [continual_real_to_sim_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml) at `5ea35c91c` (/home/ycliang/predicators-real-to-sim-frozen-20260918).
- Oracle dynamics, zero-shot and no harness fitting: [continual_five_ablations_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml) at `2982f5876` (/home/ycliang/predicators-five-arms-frozen-20260918).
- Standalone simulator and no explicit uncertainty: round r2 from [continual_standalone_no_uncertainty_r2.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml) at `0ddc8f7f4` (/home/ycliang/predicators-standalone-nounc-frozen-20260918), after the September 18 revision of both surfaces.
- Scene only: [continual_scene_only_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml).
- Agentic real-to-sim: [continual_real_to_sim_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml) at `5ea35c91c` (/home/ycliang/predicators-real-to-sim-frozen-20260918).
Its prompt asks the agent to write its own `simulator.py` scene from the engine, the manifest and the assets and to model the mechanisms, the same workflow as EMPIRIC.
- EMPIRIC + scene package: [continual_empiric_scene_package_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml).
- EMPIRIC + scene package: [continual_empiric_scene_package_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml).
Fan and Balloons ran as r1 at `d575a8447` (/home/ycliang/predicators-empiric-pkg-frozen-20260918).
Boil, Bridge and Domino ran as r2 at `8051333e5` (/home/ycliang/predicators-empiric-pkg-split-frozen-20260919), after their environment files were split so the observable sim core can be shared without the hidden mechanisms; their r1 runs were cancelled and are excluded.
To resume the paused seeds, relaunch the same launcher and round key from the same worktree so auto-resume picks up the scorecard and checkpoint.
- Direct agent + scene assets: [continual_direct_scene_files_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml) at `f6f609636` (/home/ycliang/predicators-direct-scene-files-frozen-20260919).
- Direct agent + scene assets: [continual_direct_scene_files_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml) at `f6f609636` (/home/ycliang/predicators-direct-scene-files-frozen-20260919).
It is the direct agent plus the real-to-sim arm's engine wrapper, scene manifest and URDF and mesh files as read-only references; its prompt only asks it to solve the levels, with no simulator, model files, fitting or model gate.
The rendered prompt is in [the prompt review](../prompt-review/2026-09-19-direct-scene-files/direct_agent_scene_files.md).

Expand All @@ -80,7 +80,7 @@ This is not a matched preflight ablation.
The archived Fan transfer experiment is a separate pilot with two seeds each for EMPIRIC and the direct agent.
The other agents have not been launched on this variant; their empty rows are missing results, not failures.
Only finished seeds enter the bars and curves; pending runs are listed below and do not count as zero successes.
See the [illustrated task description](../amps/fan-exposed-transfer.md) and [launch configuration](../../scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml).
See the [illustrated task description](../amps/fan-exposed-transfer.md) and [launch configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml).
This superseded pilot is omitted from the figure; its tables and logs remain archived below.
It is also excluded from the paper figure and its data selection.

Expand Down
2 changes: 1 addition & 1 deletion docs/envs/bridge/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,5 +15,5 @@ Gray pads mark the two leg sites; standing blocks are legs, lying blocks are spa
## Oracle solve trajectories

`oracle_solve_<spec>.mp4`: `oracle_process_planning` solving one test task end-to-end (seed 0), recorded with `--make_test_videos`.
The launch flags mirror the `bridge` entry in `scripts/configs/predicatorv3/envs/all.yaml`, plus `--no_repeated_arguments_in_grounding True` (set globally by `common.yaml` for config-launched runs, required on a bare CLI for the full spec to plan).
The launch flags mirror the `bridge` entry of the retired phased menu `scripts/configs/predicatorv3/envs/all.yaml` (tag `iclr-empiric-submission`), plus `--no_repeated_arguments_in_grounding True` (set globally by that menu's `common.yaml` for config-launched runs, required on a bare CLI for the full spec to plan).
The full-spec video was recorded with the default `pybullet_birrt_path_subsample_ratio 1`.
5 changes: 2 additions & 3 deletions docs/envs/domino/continuous-perception.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,9 +89,8 @@ contiguous recorded motion. Re-time an episode after the merge.
`after_step` returns `obs` **unchanged**. Shipping is a pure write-only side
effect, so *when* it happens is unobservable to the rollout — deferring every
chunk to the end produces a bit-identical twin trajectory. Everything reading
state mid-episode (`subgoal_annotations` monitor,
`agent_bilevel_max_execution_replans`, `terminate_on_goal_reached`) reads the
twin's own deterministic simulation either way.
state mid-episode (`terminate_on_goal_reached`) reads the twin's own
deterministic simulation either way.

**No protocol change is needed.** `execute_chunks` already packs a list of chunks
into one `StepRequest` (`real_robot_bridge.py:176-185`), and `_split_actions` is
Expand Down
8 changes: 4 additions & 4 deletions docs/protocol/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ There is no human oracle for these envs, our levels are not ordered by difficult
| Scorecard | `scorecard.json` per run, aggregated across runs |
| Recording JSONL and replay viewer | Per-level recording plus the existing trajectory viewer |
| Swarm | One Slurm job per env and seed, plus an aggregator |
| Benchmarking harness (model configs x games, tags) | The launcher configs in `scripts/configs/predicatorv3/` |
| Benchmarking harness (model configs x games, tags) | The benchmark config in `scripts/configs/empiric/` |

## 4. Protocol definition

Expand Down Expand Up @@ -297,7 +297,7 @@ The prompt gives no schedule and no certification rule.
- `predicators/agent_sdk/prompts/play_system.md` and `play_query.md`: the continual prompts, with golden renders like the existing templates.
- `predicators/approaches/agent_continual_approach.py`: `AgentSessionMixin` plus the tool context, the session loop, session resume, and `learn.run` as a sub-session launcher over the existing synthesis code.
- `scripts/aggregate_scorecards.py`: scorecards to tables and curves, parameterised by the aggregation chosen later.
- `scripts/configs/predicatorv3/continual_common.yaml` plus the menus `envs/continual.yaml` and `approaches/continual.yaml`: a launcher includes them, un-parks one env and the arms it compares, and names each arm with `EXTENDS` (for example `continual_balloons_compose_r2.yaml`); one job per env and seed and arm.
- `scripts/configs/empiric/benchmark.yaml`: the seven arms (`approaches.yaml`) on the five settings (`envs.yaml`) with the shared flags (`common.yaml`); a launch names its round with `--round` or a `ROUND` key, which suffixes every experiment id, and `--envs`, `--approaches` and `--seeds` pick a subset; one job per env and arm, one array task per seed.

### 6.2 Entry point

Expand Down Expand Up @@ -423,7 +423,7 @@ Tests: `tests/run/test_continual.py` pins the counts, the recordings, the preemp
Launching:

```bash
python scripts/engaging/launch.py -c predicatorv3/continual_balloons_compose_r2.yaml --partition mit_preemptable
python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 --envs balloons --partition mit_preemptable
```

The launcher passes `--auto_resume`, so a requeue resumes from the scorecard and the level recording in the run directory it adopts (section 4.7).
Expand Down Expand Up @@ -460,7 +460,7 @@ Step 2, the agent arm, landed the same day:

Tests: `tests/agent_sdk/test_continual_tools.py` drives the tools over a real session on cover (a win through `skills_execute_plan`, divergences on positive and `NOT` expectations, parse errors, game over then reset, give up and run end, the cap hit inside a tool); `tests/approaches/test_agent_continual_approach.py` runs the play loop on boil with a scripted agent in place of the LLM (a round that acts and a round that gives up, the continuation of the conversation by its recorded id, the attempts record, the checkpoint, the resume of an in-flight round after a preemption, the idle guard).

Launching the agent arm: write a launcher that includes `continual_common.yaml` and the menus and un-parks `agent_continual` with `EXTENDS` (see the header of `continual_common.yaml`).
Launching the agent arm: `--approaches mb_opus` on `empiric/benchmark.yaml` (see the header of `benchmark.yaml`).

First agent result, boil seed 0, job 21964274 (2026-09-04): both levels won, 2127 steps and 5 resets on level 1 (one session, 129 turns, 85 skill invocations, 44 failed, one horizon game over), 374 steps and no reset on level 2, 55 min active, $24.66.
The agent asked for no learning session and ran no model rollout: it measured the dynamics by probing the real environment, wrote a recipe into its journal, and replayed it on level 2.
Expand Down
4 changes: 2 additions & 2 deletions docs/protocol/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,10 +133,10 @@ An audit of every transcript (all tool calls, including the Python the agents ra

## How to run and view

Launch (un-skip the arms you want in the yaml; each job requeues and resumes itself):
Launch (name the round; `--envs`, `--approaches` and `--seeds` pick a subset; each job requeues and resumes itself):

```bash
PYTHONPATH=. python scripts/engaging/launch.py -c predicatorv3/continual_balloons_compose_r2.yaml --partition mit_preemptable
PYTHONPATH=. python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 --envs balloons --partition mit_preemptable
```

Add `--accounts a,b` to spread the runs over several Claude accounts (one token file per account under `~/.claude-tokens/`, see `scripts/engaging/claude_accounts.py`); each seed is assigned round-robin and the scorecard records which account it used.
Expand Down
Loading
Loading