diff --git a/docs/amps/empiric-from-assets.md b/docs/amps/empiric-from-assets.md index 1b8d3062dd..2c935e1868 100644 --- a/docs/amps/empiric-from-assets.md +++ b/docs/amps/empiric-from-assets.md @@ -29,7 +29,7 @@ This is a reduction in supplied domain implementation, not reconstruction from r ## Pilot -The configuration is [continual_from_assets_pilot_r1.yaml](../../scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml). +The configuration is [continual_from_assets_pilot_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml). It runs Opus 5 on seeds 0 and 1 of Fan + ramp, four-span Bridge, Domino, Balloons and Boil. There are ten intended runs total, each with its training and test levels. Fan uses the previously reviewed 3 mm ramp with the 10 cm landing extension and the repaired shared skills, not Fan maze. @@ -49,7 +49,7 @@ Bridge seeds 0-1 are array `23407427`; Fan maze seeds 0-1 are array `23407428`. The user immediately corrected Fan maze to Fan + ramp; both tasks of array `23407428` were cancelled with their logs preserved. The abandoned maze runs are not part of the intended pilot and must not be resumed or counted as task failures. Bridge array `23407427` remains unchanged. -The [expansion-only configuration](../../scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml) submits only the eight new tasks, with Bridge and Fan maze disabled to prevent duplicates. +The [expansion-only configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml) submits only the eight new tasks, with Bridge and Fan maze disabled to prevent duplicates. Its runtime menus are pinned to the frozen implementation. The resolved Fan + ramp flags match the repaired EMPIRIC ramp cohort, apart from making the default fitting flag explicit. The eight replacement/additional tasks were submitted on account d as the following two-seed arrays: diff --git a/docs/amps/fan-development.md b/docs/amps/fan-development.md index 8a5dbc6ecd..583b7d28a7 100644 --- a/docs/amps/fan-development.md +++ b/docs/amps/fan-development.md @@ -303,20 +303,20 @@ The backup account `dat` also passed but is not in the active pool. Account `c` has an active limit marker until September 23 at 00:00 UTC and is excluded. Usage percentages were unavailable from the service endpoint; successful probes establish current access, not a guarantee of sufficient remaining quota for full runs. Each task requests 8 CPUs and 16 GB, with requeue enabled for preemption, time limits, and recognized account-limit exits. -The launch configuration is [continual_fan_ramp_skill_repair_r1.yaml](../../scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml). +The launch configuration is [continual_fan_ramp_skill_repair_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml). The shared repair and its measured extra interaction cost are documented in [the switch investigation](fan-switch-seed4-investigation.md). Seed 0's original job stopped after account `a` reported that its organization had disabled subscription access for Claude Code. This was an infrastructure interruption after training succeeded and the test reached 207 steps, not a task failure. With the user's approval, job `23395354` resumes only seed 0 on `dat`, from the same frozen runtime, experiment key, run directory, sandbox, and recorded state. -The resume configuration is [continual_fan_ramp_skill_repair_seed0_resume.yaml](../../scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml). +The resume configuration is [continual_fan_ramp_skill_repair_seed0_resume.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml). ### Six matched comparison arms The user approved five seeds each for all six other paper agents on this repaired Fan ramp setup. All 30 tasks were verified running on compute nodes after submission. They use the same frozen runtime `ff11bc76f4652c6964e73beda7e41a6655d670d9`, reviewed geometry, observation noise, step budget, and preflight-off setting as the repaired EMPIRIC cohort. -The [six-arm configuration](../../scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml) resolves to exactly six five-seed arrays, with only the intended approach and ablation flag differences. +The [six-arm configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml) resolves to exactly six five-seed arrays, with only the intended approach and ablation flag differences. - Oracle dynamics: array `23398858`, seeds 0-4, accounts b/d. - Direct agent: array `23398859`, seeds 0-4, account dat. diff --git a/docs/amps/oracle-fan-ramp-cohort-investigation.md b/docs/amps/oracle-fan-ramp-cohort-investigation.md index 33ad547a2d..8c91f2d588 100644 --- a/docs/amps/oracle-fan-ramp-cohort-investigation.md +++ b/docs/amps/oracle-fan-ramp-cohort-investigation.md @@ -90,7 +90,7 @@ The driver is [audit_oracle_ramp_20260922.py](/home/ycliang/predicators/logs/aud Seeds 9 and 10 were submitted as array `23463117` and started on compute nodes using accounts b and d. The launch flags and arguments were checked against the previous extra-seed config; only the seed range changes. -The [launch configuration](/home/ycliang/predicators/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml) retains the frozen runtime and original cohort identifier. +The [launch configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml) retains the frozen runtime and original cohort identifier. The Markdown benchmark tracks all eleven seeds, retaining all failures. Replacement plot monitor `23463368` refreshes the report and figures every 60 seconds when results change. The expanded-cohort report tests passed: 8 tests. diff --git a/docs/comparisons/empiric-validation-r2.md b/docs/comparisons/empiric-validation-r2.md index 47da3f16c8..c2384b783d 100644 --- a/docs/comparisons/empiric-validation-r2.md +++ b/docs/comparisons/empiric-validation-r2.md @@ -1,7 +1,7 @@ # EMPIRIC r2: prospective robustness checks The requested cohort is two additional seeds per domain, seeds 3 and 4, across Boil, Domino, Balloons, Bridge, and Fan. -The launcher is [continual_empiric_benchmark_r2.yaml](../../scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml). +The launcher is [continual_empiric_benchmark_r2.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml). Its experiment key is `-mb_opus_benchmark_r2`, under `agent_continual`. The main figures retain historical EMPIRIC and show this cohort separately as **EMPIRIC r2**. All finished outcomes count, including failures; an unfinished run is not a failed run. diff --git a/docs/comparisons/five-seed-launch-review.md b/docs/comparisons/five-seed-launch-review.md index 2a7155730e..e7f43ae094 100644 --- a/docs/comparisons/five-seed-launch-review.md +++ b/docs/comparisons/five-seed-launch-review.md @@ -35,7 +35,7 @@ These checks do not establish future solve rates or resolve the previously docum ## Launch and reporting -The launcher is [continual_benchmark_five_seeds.yaml](../../scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml). +The launcher is [continual_benchmark_five_seeds.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml). All 60 destination seed directories were checked absent before submission. Jobs use `mit_preemptable`, checkpoint resume, automatic requeue, and account labels `a,b,c,d`, all confirmed usable by the user during this launch review. The user's account e corresponds to the launcher's `dat` label and is reserved as backup, not included in the normal rotation. diff --git a/docs/comparisons/ten-agent-opus-benchmark.md b/docs/comparisons/ten-agent-opus-benchmark.md index e8a8b9a3c8..00d1b45d2b 100644 --- a/docs/comparisons/ten-agent-opus-benchmark.md +++ b/docs/comparisons/ten-agent-opus-benchmark.md @@ -46,20 +46,20 @@ python ../../../scripts/plotting/plot_benchmark_arms.py benchmark-arms-opus ## Settings and cohort selection -The five settings are the menu defaults in [envs/continual.yaml](../../scripts/configs/predicatorv3/envs/continual.yaml): Boil two-jug test, Domino high-friction turn, Fan ramp test, Bridge four-span test row and Balloons composition test levels. +The five settings are the menu defaults in [envs/continual.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/envs/continual.yaml): Boil two-jug test, Domino high-friction turn, Fan ramp test, Bridge four-span test row and Balloons composition test levels. EMPIRIC and the direct agent ran before the benchmark rounds under their own round keys; the four-span EMPIRIC runs are the preflight-off ones (r2 seed 0, r3 seeds 1-2). The other arms ran as `_opus_benchmark_` from frozen worktrees: -- Oracle dynamics, zero-shot and no harness fitting: [continual_five_ablations_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml) at `2982f5876` (/home/ycliang/predicators-five-arms-frozen-20260918). -- Standalone simulator and no explicit uncertainty: round r2 from [continual_standalone_no_uncertainty_r2.yaml](../../scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml) at `0ddc8f7f4` (/home/ycliang/predicators-standalone-nounc-frozen-20260918), after the September 18 revision of both surfaces. -- Scene only: [continual_scene_only_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml). -- Agentic real-to-sim: [continual_real_to_sim_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml) at `5ea35c91c` (/home/ycliang/predicators-real-to-sim-frozen-20260918). +- Oracle dynamics, zero-shot and no harness fitting: [continual_five_ablations_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml) at `2982f5876` (/home/ycliang/predicators-five-arms-frozen-20260918). +- Standalone simulator and no explicit uncertainty: round r2 from [continual_standalone_no_uncertainty_r2.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml) at `0ddc8f7f4` (/home/ycliang/predicators-standalone-nounc-frozen-20260918), after the September 18 revision of both surfaces. +- Scene only: [continual_scene_only_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml). +- Agentic real-to-sim: [continual_real_to_sim_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml) at `5ea35c91c` (/home/ycliang/predicators-real-to-sim-frozen-20260918). Its prompt asks the agent to write its own `simulator.py` scene from the engine, the manifest and the assets and to model the mechanisms, the same workflow as EMPIRIC. -- EMPIRIC + scene package: [continual_empiric_scene_package_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml). +- EMPIRIC + scene package: [continual_empiric_scene_package_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml). Fan and Balloons ran as r1 at `d575a8447` (/home/ycliang/predicators-empiric-pkg-frozen-20260918). Boil, Bridge and Domino ran as r2 at `8051333e5` (/home/ycliang/predicators-empiric-pkg-split-frozen-20260919), after their environment files were split so the observable sim core can be shared without the hidden mechanisms; their r1 runs were cancelled and are excluded. To resume the paused seeds, relaunch the same launcher and round key from the same worktree so auto-resume picks up the scorecard and checkpoint. -- Direct agent + scene assets: [continual_direct_scene_files_benchmark_r1.yaml](../../scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml) at `f6f609636` (/home/ycliang/predicators-direct-scene-files-frozen-20260919). +- Direct agent + scene assets: [continual_direct_scene_files_benchmark_r1.yaml](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml) at `f6f609636` (/home/ycliang/predicators-direct-scene-files-frozen-20260919). It is the direct agent plus the real-to-sim arm's engine wrapper, scene manifest and URDF and mesh files as read-only references; its prompt only asks it to solve the levels, with no simulator, model files, fitting or model gate. The rendered prompt is in [the prompt review](../prompt-review/2026-09-19-direct-scene-files/direct_agent_scene_files.md). @@ -80,7 +80,7 @@ This is not a matched preflight ablation. The archived Fan transfer experiment is a separate pilot with two seeds each for EMPIRIC and the direct agent. The other agents have not been launched on this variant; their empty rows are missing results, not failures. Only finished seeds enter the bars and curves; pending runs are listed below and do not count as zero successes. -See the [illustrated task description](../amps/fan-exposed-transfer.md) and [launch configuration](../../scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml). +See the [illustrated task description](../amps/fan-exposed-transfer.md) and [launch configuration](https://github.com/BasisResearch/predicators/blob/iclr-empiric-submission/scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml). This superseded pilot is omitted from the figure; its tables and logs remain archived below. It is also excluded from the paper figure and its data selection. diff --git a/docs/envs/bridge/README.md b/docs/envs/bridge/README.md index ba000619e2..9b3850a5a1 100644 --- a/docs/envs/bridge/README.md +++ b/docs/envs/bridge/README.md @@ -15,5 +15,5 @@ Gray pads mark the two leg sites; standing blocks are legs, lying blocks are spa ## Oracle solve trajectories `oracle_solve_.mp4`: `oracle_process_planning` solving one test task end-to-end (seed 0), recorded with `--make_test_videos`. -The launch flags mirror the `bridge` entry in `scripts/configs/predicatorv3/envs/all.yaml`, plus `--no_repeated_arguments_in_grounding True` (set globally by `common.yaml` for config-launched runs, required on a bare CLI for the full spec to plan). +The launch flags mirror the `bridge` entry of the retired phased menu `scripts/configs/predicatorv3/envs/all.yaml` (tag `iclr-empiric-submission`), plus `--no_repeated_arguments_in_grounding True` (set globally by that menu's `common.yaml` for config-launched runs, required on a bare CLI for the full spec to plan). The full-spec video was recorded with the default `pybullet_birrt_path_subsample_ratio 1`. diff --git a/docs/envs/domino/continuous-perception.md b/docs/envs/domino/continuous-perception.md index 5cd22e104a..815eb223c1 100644 --- a/docs/envs/domino/continuous-perception.md +++ b/docs/envs/domino/continuous-perception.md @@ -89,9 +89,8 @@ contiguous recorded motion. Re-time an episode after the merge. `after_step` returns `obs` **unchanged**. Shipping is a pure write-only side effect, so *when* it happens is unobservable to the rollout — deferring every chunk to the end produces a bit-identical twin trajectory. Everything reading -state mid-episode (`subgoal_annotations` monitor, -`agent_bilevel_max_execution_replans`, `terminate_on_goal_reached`) reads the -twin's own deterministic simulation either way. +state mid-episode (`terminate_on_goal_reached`) reads the twin's own +deterministic simulation either way. **No protocol change is needed.** `execute_chunks` already packs a list of chunks into one `StepRequest` (`real_robot_bridge.py:176-185`), and `_split_actions` is diff --git a/docs/protocol/design.md b/docs/protocol/design.md index bde2331281..8a36619c0a 100644 --- a/docs/protocol/design.md +++ b/docs/protocol/design.md @@ -56,7 +56,7 @@ There is no human oracle for these envs, our levels are not ordered by difficult | Scorecard | `scorecard.json` per run, aggregated across runs | | Recording JSONL and replay viewer | Per-level recording plus the existing trajectory viewer | | Swarm | One Slurm job per env and seed, plus an aggregator | -| Benchmarking harness (model configs x games, tags) | The launcher configs in `scripts/configs/predicatorv3/` | +| Benchmarking harness (model configs x games, tags) | The benchmark config in `scripts/configs/empiric/` | ## 4. Protocol definition @@ -297,7 +297,7 @@ The prompt gives no schedule and no certification rule. - `predicators/agent_sdk/prompts/play_system.md` and `play_query.md`: the continual prompts, with golden renders like the existing templates. - `predicators/approaches/agent_continual_approach.py`: `AgentSessionMixin` plus the tool context, the session loop, session resume, and `learn.run` as a sub-session launcher over the existing synthesis code. - `scripts/aggregate_scorecards.py`: scorecards to tables and curves, parameterised by the aggregation chosen later. -- `scripts/configs/predicatorv3/continual_common.yaml` plus the menus `envs/continual.yaml` and `approaches/continual.yaml`: a launcher includes them, un-parks one env and the arms it compares, and names each arm with `EXTENDS` (for example `continual_balloons_compose_r2.yaml`); one job per env and seed and arm. +- `scripts/configs/empiric/benchmark.yaml`: the seven arms (`approaches.yaml`) on the five settings (`envs.yaml`) with the shared flags (`common.yaml`); a launch names its round with `--round` or a `ROUND` key, which suffixes every experiment id, and `--envs`, `--approaches` and `--seeds` pick a subset; one job per env and arm, one array task per seed. ### 6.2 Entry point @@ -423,7 +423,7 @@ Tests: `tests/run/test_continual.py` pins the counts, the recordings, the preemp Launching: ```bash -python scripts/engaging/launch.py -c predicatorv3/continual_balloons_compose_r2.yaml --partition mit_preemptable +python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 --envs balloons --partition mit_preemptable ``` The launcher passes `--auto_resume`, so a requeue resumes from the scorecard and the level recording in the run directory it adopts (section 4.7). @@ -460,7 +460,7 @@ Step 2, the agent arm, landed the same day: Tests: `tests/agent_sdk/test_continual_tools.py` drives the tools over a real session on cover (a win through `skills_execute_plan`, divergences on positive and `NOT` expectations, parse errors, game over then reset, give up and run end, the cap hit inside a tool); `tests/approaches/test_agent_continual_approach.py` runs the play loop on boil with a scripted agent in place of the LLM (a round that acts and a round that gives up, the continuation of the conversation by its recorded id, the attempts record, the checkpoint, the resume of an in-flight round after a preemption, the idle guard). -Launching the agent arm: write a launcher that includes `continual_common.yaml` and the menus and un-parks `agent_continual` with `EXTENDS` (see the header of `continual_common.yaml`). +Launching the agent arm: `--approaches mb_opus` on `empiric/benchmark.yaml` (see the header of `benchmark.yaml`). First agent result, boil seed 0, job 21964274 (2026-09-04): both levels won, 2127 steps and 5 resets on level 1 (one session, 129 turns, 85 skill invocations, 44 failed, one horizon game over), 374 steps and no reset on level 2, 55 min active, $24.66. The agent asked for no learning session and ran no model rollout: it measured the dynamics by probing the real environment, wrote a recipe into its journal, and replayed it on level 2. diff --git a/docs/protocol/overview.md b/docs/protocol/overview.md index d9ea1a28c3..1f55741a57 100644 --- a/docs/protocol/overview.md +++ b/docs/protocol/overview.md @@ -133,10 +133,10 @@ An audit of every transcript (all tool calls, including the Python the agents ra ## How to run and view -Launch (un-skip the arms you want in the yaml; each job requeues and resumes itself): +Launch (name the round; `--envs`, `--approaches` and `--seeds` pick a subset; each job requeues and resumes itself): ```bash -PYTHONPATH=. python scripts/engaging/launch.py -c predicatorv3/continual_balloons_compose_r2.yaml --partition mit_preemptable +PYTHONPATH=. python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 --envs balloons --partition mit_preemptable ``` Add `--accounts a,b` to spread the runs over several Claude accounts (one token file per account under `~/.claude-tokens/`, see `scripts/engaging/claude_accounts.py`); each seed is assigned round-robin and the scorecard records which account it used. diff --git a/docs/uncertainty/principled-belief.md b/docs/uncertainty/principled-belief.md index aff9a15b6a..c5a6e7f3eb 100644 --- a/docs/uncertainty/principled-belief.md +++ b/docs/uncertainty/principled-belief.md @@ -106,7 +106,7 @@ Scripts and logs: `~/claude_sbatch/principled_belief_20260925/interpenetration_c ## Implementation status (2026-09-25) -Built on the `principled-belief` branch, behind `belief_joint_draws` (settings default 0; `continual_common.yaml` sets 16, and the No uncertainty arm keeps 0). +Built on the `principled-belief` branch, behind `belief_joint_draws` (settings default 0; `scripts/configs/empiric/common.yaml` sets 16, and the No uncertainty arm keeps 0). - Parameter factor: `ParameterBelief` with line posteriors, discrete parameters as categorical lines (`DiscretePosterior`), fresh draws (`ParameterBelief.sample`), and `prior_parameter_belief` before a fit; the fit tool reports the belief in place of the identifiability section. - State factor: `change_point_belief` and `feature_ranges` in `observation_belief.py`; `ContinualRun.belief()` uses it when the joint belief is on, for every arm with a state estimate. diff --git a/predicators/agent_sdk/belief_probe.py b/predicators/agent_sdk/belief_probe.py index 1a0defd3b3..41f7a00918 100644 --- a/predicators/agent_sdk/belief_probe.py +++ b/predicators/agent_sdk/belief_probe.py @@ -1,21 +1,19 @@ """Exploration probe API exposed to agents via ``run_python``. -``BeliefProbe`` is a thin facade over the machinery the curated tools -already use - ``parse_sketch_from_text`` (plan grammar), -``execute_plan_forward`` (forward executor over the option model), the -tools' state-modification and rendering helpers - so probe rollouts -behave identically to ``submit_plan`` rollouts. What it adds is -composability: the agent can set the sim to any task state (or a -modified copy), read full-precision features, run partial plans, render, -snapshot/restore, and write sweep loops in one ``run_python`` call -instead of one tool round-trip per experiment. - -By construction nothing the probe executes can be captured as the -answer - submission happens only through ``submit_plan`` on the -true initial state. The task evaluator is reachable, but only as a -read-only preview: ``run(trials>=2, solved=True)`` and -``refine(require_solved=True)`` score rollouts through the same gate the -capture path uses (see ``_require_solved_evaluator``). +``BeliefProbe`` is a thin facade over shared machinery - +``parse_sketch_from_text`` (plan grammar), ``execute_plan_forward`` +(forward executor over the option model), the tools' state-modification +and rendering helpers. What it adds is composability: the agent can set +the sim to any task state (or a modified copy), read full-precision +features, run partial plans, render, snapshot/restore, and write sweep +loops in one ``run_python`` call instead of one tool round-trip per +experiment. + +Nothing the probe executes acts in the environment: plans reach the +robot only through the play tools (``skills_execute_plan``). The task +evaluator is reachable, but only as a read-only preview: +``run(trials>=2, solved=True)`` and ``refine(require_solved=True)`` +score rollouts with it (see ``_require_solved_evaluator``). In synthesis sessions the same facade probes the CANDIDATE simulator (the ``simulator.py`` under edit, freshly fitted - see @@ -38,15 +36,15 @@ Sequence, Set, Tuple, Union from predicators import utils -from predicators.agent_sdk.config import RefinementConfig, ToolSurfaceConfig, \ - ValidationConfig +from predicators.agent_sdk.config import RefinementConfig, ToolSurfaceConfig from predicators.agent_sdk.parallel_rollouts import prefetch_parallel from predicators.agent_sdk.tools.context import absolute_rollout_seed, \ decorrelated_rollout_seed from predicators.agent_sdk.tools.scene import apply_state_modifications, \ draw_pybullet_annotation, render_pybullet_image, render_scene_image from predicators.agent_sdk.tools.verdicts import _EvalStateCollector, \ - evaluate_states_with, load_ground_sampler_fns, make_solved_check + _policy_source_path, evaluate_states_with, load_ground_sampler_fns, \ + make_solved_check from predicators.structs import State, Task, excluded_object_type_names if TYPE_CHECKING: @@ -62,29 +60,17 @@ class ProbeBudgetExceeded(Exception): """A probe call ran past a wall-clock budget. Raised cooperatively at probe checkpoints (every sim call) when the - run_python per-call limit or the solve attempt's wall clock has - expired. ``run_python`` catches it specially: the code's printed + run_python per-call limit has expired. ``run_python`` catches it + specially: the code's printed output so far is returned with the budget message appended, so a stopped sweep still hands the agent its partial results. """ def _check_time_budget(ctx: "ToolContext") -> None: - """Raise :class:`ProbeBudgetExceeded` when a wall-clock budget is up. - - Never fires during the final-submission nudge - (``ctx.capture_best_effort_plan``): with the budget spent, the one - thing left is submitting, and blocking that would forfeit the task. - """ - if ctx.capture_best_effort_plan: - return + """Raise :class:`ProbeBudgetExceeded` when the run_python call's time limit + is up.""" now = time.monotonic() - attempt_dl = ctx.attempt_deadline - if attempt_dl is not None and now > attempt_dl: - raise ProbeBudgetExceeded( - "the attempt's wall-clock exploration budget is exhausted. Stop " - "exploring NOW and submit your single best plan via " - "submit_plan on the current task (omit task_idx).") call_dl = ctx.python_call_deadline if call_dl is not None and now > call_dl: call_timeout = ToolSurfaceConfig.from_cfg().python_call_timeout @@ -236,16 +222,15 @@ def __getattr__(self, name: str) -> Any: class ProbeResult(_StrLikeResult): """Outcome of one ``BeliefProbe.run`` call. - Attributes mirror the ``submit_plan`` report: ``steps`` is - a list of per-step dicts (``option``, ``num_actions``, ``failure``, - ``added``, ``deleted``, ``subgoals_missing`` - the step's ``-> - {atoms}`` annotations that did NOT hold in the post-state (the - forward-pass divergence signal), ``image`` - the saved post-step - scene image path, if rendering is available; with ``contacts=True`` - also ``contacts`` - the step's contact-pair span lines), plus - ``goal_reached`` and ``final_atoms``. ``notes`` carries caveats - (ignored region annotations, horizon overruns). ``print(result)`` - renders the same step-by-step summary the tool prints. + ``steps`` is a list of per-step dicts (``option``, ``num_actions``, + ``failure``, ``added``, ``deleted``, ``subgoals_missing`` - the + step's ``-> {atoms}`` annotations that did NOT hold in the + post-state (the forward-pass divergence signal), ``image`` - the + saved post-step scene image path, if rendering is available; with + ``contacts=True`` also ``contacts`` - the step's contact-pair span + lines), plus ``goal_reached`` and ``final_atoms``. ``notes`` carries + caveats (ignored region annotations, horizon overruns). + ``print(result)`` renders the step-by-step summary. """ steps: List[Dict[str, Any]] goal_reached: bool @@ -668,9 +653,9 @@ def __repr__(self) -> str: " the plan fails INSIDE the identified-parameter " "uncertainty range. The real environment may sit at any " "of these points (success can be non-monotonic: passing " - "neighbors do NOT cover the points between them), and the " - "capture gate re-runs this same sweep - add design margin " - "until every point passes before submitting.") + "neighbors do NOT cover the points between them) - add " + "design margin until every point passes before executing " + "the plan.") lines.extend(f"NOTE: {n_}" for n_ in self.notes) return "\n".join(lines) @@ -776,7 +761,7 @@ class ProbeRefineResult(_StrLikeResult): as more than it means. ``plan_lines`` holds one line per sketch step with the refined params filled in (``[?]`` for steps the search never refined) - paste them into ``sim.run`` or - ``submit_plan``. ``near_miss`` is the deepest validation + ``skills_execute_plan``. ``near_miss`` is the deepest validation failure (step index, the exact params that got furthest, and why they failed), also populated on timeout/exhaustion. ``note`` carries caveats. @@ -872,15 +857,14 @@ class BeliefProbe: sim.state("domino_1") # full-precision features sim.render("after_push") sim.restore(sid) - # Search params for a suffix from here (nothing is captured): + # Search params for a suffix from here (simulation only): print(sim.refine( "Place(robot:robot)[0.46, 1.32, 0.55, -1.0] ~ [0.05, 0.05, " "0.0, 0.5] -> {SomeSubgoal(domino_1:domino)}")) The "current state" is just a ``State`` object; ``run`` executes from - it (the option model resets the sim env from that state, exactly as - ``submit_plan`` does from a task init) and advances it to - the rollout's final state. + it (the option model resets the sim env from that state) and advances + it to the rollout's final state. """ # Distinct deterministic rng streams per instance (see refine). @@ -930,7 +914,9 @@ def _option_model(self) -> Any: def _fresh_scope(self) -> Optional[Callable[..., Any]]: """Select isolation for the model this probe actually executes.""" - if not ValidationConfig.from_cfg().fresh_env: + # pylint: disable-next=import-outside-toplevel + from predicators.settings import CFG + if not CFG.agent_plan_validation_fresh_env: return None ctx = self._ctx if ctx.probe_option_model_provider is not None: @@ -1205,20 +1191,6 @@ def predicates(self, max_trajectories=max_trajectories, max_groundings_per_predicate=max_groundings_per_predicate) - def samplers(self) -> str: - """Reload ``samplers.py`` and install its per-skill samplers. - - Sampler-synthesis sessions only. Loads ``LEARNED_SAMPLERS`` - fresh from the file (snapshotting it into - ``samplers_versions/``), validates the option-name -> callable - map, installs it so ``refine`` draws from the draft samplers, - and reports a per-option sanity check (return shape, in-box - draws) on a representative train-task state. Call it after every - edit of ``samplers.py``. - """ - self._require_available("samplers") - return self._artifact_loader("samplers")() - def _artifact_loader(self, name: str) -> Callable[..., str]: ctx = self._ctx _check_time_budget(ctx) @@ -1739,8 +1711,8 @@ def _parse_sketch(self, plan_text: str) -> Any: the probe task starts at the current state and carries no evaluator, and ``notices`` lists parse caveats to surface (e.g. region annotations ignored because ground samplers are off). - Same grammar and parser as ``submit_plan`` (``~ [w]`` search - regions included). + Same grammar and parser as ``skills_execute_plan``, plus ``~ + [w]`` search regions. """ # pylint: disable=import-outside-toplevel from predicators.agent_sdk import bilevel_sketch @@ -1896,7 +1868,7 @@ def run( real steps are spent on it. Needs a declared observation-noise channel. - ``plan_text`` uses the same grammar as ``submit_plan``: + ``plan_text`` uses the same grammar as ``skills_execute_plan``: one option per line, ``Option(obj:type, ...)[params]`` with exact continuous params (``[]`` for none); ``-> {atoms}`` subgoal annotations are optional but CHECKED - each step's @@ -1907,10 +1879,10 @@ def run( diverges here means a rule is more permissive than the env). Advances the current state to the rollout's final state (``restore`` a snapshot to rewind). - Like ``submit_plan``, each step's post-state is - rendered to a saved image whose path lands in the step report; - pass ``render=False`` inside tight sweep loops to skip that. - Exploratory only: results are never captured. + Each step's post-state is rendered to a saved image whose path + lands in the step report; pass ``render=False`` inside tight + sweep loops to skip that. Exploratory only: nothing runs in the + environment. Without the joint belief, ``trials=N`` (N > 1) runs the SAME plan N times and returns a ``ProbeTrialsResult`` with the @@ -1929,8 +1901,7 @@ def run( EVALUATOR, reporting per-trial ``solved``/``reward``. Reaching the goal atoms is NOT the same as being scored a solve - the evaluator can reject a goal-reaching route - so check ``solved`` - counts here BEFORE submitting via ``submit_plan`` instead of - discovering rejections one submission at a time. + counts here BEFORE executing the plan in the environment. ``contacts=True`` (single-run mode only, ``trials=1``) records every physical contact during the rollout and reports, per step, @@ -1946,7 +1917,7 @@ def run( parameter at the ends of its 95% interval with the others at their most likely values; without it, a grid spanning the +-1-sigma uncertainty range of the identified physical - parameters (the SAME points the capture gate checks) - each on + parameters - each on a fresh env at the base motion-planner seed, plus once at the fitted values, and returns a ``ProbeSweepResult`` with per-point outcomes. The @@ -1957,31 +1928,23 @@ def run( it), so ``trials=`` at the fitted values CANNOT see this. A design near a feasibility boundary (e.g. the minimal block count) is exactly where such holes live: sweep it and add - margin until EVERY point passes before submitting, instead of - discovering PARAM-SENSITIVE rejections one capture at a time. + margin until EVERY point passes before executing it for real. Rollouts are deterministic per point, so each point costs one rollout and its outcome is a measurement, not a sample. The current state is NOT advanced and nothing is rendered. ``seed=S`` overrides the base motion-planner seed for this call. Trials report the planner seed each ran at (trial ``i`` runs at - ``S + i``; without ``seed=`` at ``base + i``), and - ``submit_plan``'s validation rollouts report theirs the - same way. A single run (``trials=1``) executes entirely at - ``S``; a physics sweep runs every point at ``S`` instead of the - base. Without ``seed=``, and from the task's unmodified initial - state (plain ``reset()``, no rollout since), ``trials=N`` runs - the IDENTICAL rollout set as ``submit_plan``'s N-rollout capture - gate (fresh env per rollout, planner seeds ``base..base+N-1``), - so a trials score here is exactly the gate's verdict on this - plan. + ``S + i``; without ``seed=`` at ``base + i``). A single run + (``trials=1``) executes entirely at ``S``; a physics sweep runs + every point at ``S`` instead of the base. SUBSTRATE: a default single run executes on the WARM shared session env from the probe's current state - a feature for mid-exploration probing (after ``reset(mods=...)`` or a partial - rollout), but a different physics substrate from trials, - validation rollouts, and the real episode, which all run on a - freshly constructed env. A warm pass at a seed where a fresh + rollout), but a different physics substrate from trials, physics + sweeps and the real episode, which all run on a freshly + constructed env. A warm pass at a seed where a fresh trial failed is evidence of shared-env optimism, not seed luck. ``fresh=True`` (single-run mode only) reproduces a failed trial's substrate exactly: ``run(plan, seed=, @@ -2078,8 +2041,8 @@ def run( probe_task, sketch_steps, all_predicates, notices = \ self._parse_sketch(plan_text) # Ground via the shared helper so an annotated Wait waits for - # its annotated atoms here exactly as in refine, submit_plan, - # and real execution (see submit_plan's grounding comment). + # its annotated atoms here exactly as in refine and real + # execution. grounded: List[Any] = [] for st in sketch_steps: params = (st.initial_params if st.initial_params is not None else @@ -2123,12 +2086,8 @@ def _horizon_note(total_actions: int) -> Optional[str]: "physics_sweep=True, but no identified physical " "parameters with nonzero posterior width are deployed " "this cycle, so there is no uncertainty range to " - "sweep. (submit_plan's rule-parameter ensemble margin " - "is a different, automatic gate over the learned rule " - "constants - it still runs at submission and is not " - "reachable through physics_sweep.) Use trials= to " - "measure execution reliability at the current physics " - "instead.") + "sweep. Use trials= to measure execution reliability " + "at the current physics instead.") point_dicts: List[Dict[str, Any]] = [] all_points: List[Optional[Dict[str, float]]] = \ [None] + sweep_points @@ -2140,8 +2099,7 @@ def _horizon_note(total_actions: int) -> Optional[str]: def _one_point( point: Optional[Dict[str, float]]) -> Dict[str, Any]: # Base planner seed at every point (no - # decorrelated_rollout_seed), matching the capture - # gate's margin rollouts: with the seed held fixed, an + # decorrelated_rollout_seed): with the seed held fixed, an # outcome flip between points is attributable to the # physics perturbation alone. with (sweep_scope() if point is None else sweep_scope( @@ -2287,7 +2245,7 @@ def _trial_on_step(i: int, outcome: Any) -> None: stop_on_failure=True) # Score INSIDE the scope: the evaluator's # certificate probes at the (fresh) env the - # rollout ran on, same as the capture path. + # rollout ran on. if collector is not None: coarse = collector.coarse if (collector is not None and not coarse @@ -2447,8 +2405,8 @@ def _on_step(i: int, outcome: Any) -> None: f"state exactly (features: {feats}) - a failure " "here may partly reflect start-state " "reconstruction error, not plan margin.") - # Same per-step audit image submit_plan saves; the - # env already sits at the post-step state here. + # Per-step audit image; the env already sits at the + # post-step state here. img = render_scene_image( ctx, f"probe_step_{i}_{outcome.option.name}") if render else None @@ -2554,18 +2512,16 @@ def run_policy( Policy-mode counterpart of ``run``: loads the sandbox's ``policy.py`` fresh (or executes ``source`` directly) and drives - its ``get_option(state, memory)`` through the belief model with - the same failure-surfacing semantics as ``submit_policy`` and - the real executor - option failures land in - ``memory['last_failure']`` and the policy is asked again; - get_option bugs end the episode. + its ``get_option(state, memory)`` through the belief model - + option failures land in ``memory['last_failure']`` and the + policy is asked again; get_option bugs end the episode. Starting from the CURRENT probe state is the point: perturb or advance the state first (``reset(mods=...)``, a partial ``run``) and check that the policy RECOVERS from off-nominal states, not just the initial one. ``trials=N`` repeats the rollout from the SAME current state on fresh envs (fresh policy memory per - trial). Never captures - deliver via ``submit_policy``. Also + trial). Nothing runs in the environment. Also available in learn sessions (probing a candidate simulator); there it runs against the candidate model. """ @@ -2575,7 +2531,6 @@ def run_policy( from predicators.agent_sdk.policy_execution import \ build_policy_option_fn, execute_policy_forward - from predicators.agent_sdk.tools.testing import _policy_source_path from predicators.settings import CFG # pylint: enable=import-outside-toplevel @@ -2871,17 +2826,18 @@ def suggest_probes(self, "alternatives cannot be ranked." ]) if not ctx.info_seeking_active(): - # Adaptive info-seeking: hold probing back until the capture - # gate has actually caught a fragile plan. Spending real steps + # Adaptive info-seeking: hold probing back until a physics + # sweep has actually found a fragile plan. Spending real steps # to reduce uncertainty before then is the step tax this mode # removes on easy levels. return ProbeSuggestResult([], [ - "adaptive info-seeking: no plan has been refused as " - "parameter-sensitive yet, so probing is not worth real " - "steps. Submit your best plan; if the capture gate refuses " - "it because a parameter's uncertainty threatens the goal, " - "it will name that parameter and this call will then rank " - "the probes that reduce it." + "adaptive info-seeking: no physics sweep has found a " + "parameter-sensitive plan yet, so probing is not worth " + "real steps. Stress-test your best plan with " + "sim.run(plan, physics_sweep=True); if a parameter's " + "uncertainty threatens the goal, the sweep names that " + "parameter and this call will then rank the probes that " + "reduce it." ]) probe_task, sketch_steps, all_predicates, notices = \ self._parse_sketch(sketch_text) @@ -2903,7 +2859,6 @@ def _on_rollout() -> None: rng=rng, max_draws=max(1, int(max_draws)), top_k=max(1, int(top_k)), - parameterized_samplers=ctx.parameterized_samplers or None, on_rollout=_on_rollout) return ProbeSuggestResult(suggestions, list(notices) + notes) @@ -2974,7 +2929,6 @@ def plan_scorer(plan: List[Any], rng=rng, max_draws=max(1, max_draws), top_k=max(1, int(top_k)), - parameterized_samplers=ctx.parameterized_samplers or None, on_rollout=lambda: _check_time_budget(ctx), plan_scorer=plan_scorer) notices.append( @@ -2994,7 +2948,7 @@ def refine(self, require_solved: bool = False) -> "ProbeRefineResult": """Backtracking parameter search for a sketch FROM THE CURRENT STATE. - Same grammar as ``submit_plan``, and composable: + Same grammar as ``skills_execute_plan``, and composable: refine a plan *suffix* from a snapshot where the prefix already executed, so the search budget goes to the step that matters instead of re-descending through the whole plan. @@ -3010,7 +2964,7 @@ def refine(self, current state. Returns best-found params (also on TIMEOUT - the refined prefix is reported as far as it got) plus per-step sample counts and the deepest near-miss. Exploratory only: - nothing is captured. + nothing runs in the environment. ``require_solved=True`` (implies ``require_goal``) additionally gates final-step acceptance on the TASK EVALUATOR's public @@ -3095,7 +3049,6 @@ def gated_solved_check(states: List[State], labels: List[Any], check_subgoals=True, check_final_goal=require_goal, run_id="probe", - parameterized_samplers=ctx.parameterized_samplers or None, strip_latent_wait_targets=not ctx.latent_tracking_available, solved_check=solved_check) refined_plan, success = outcome.plan, outcome.success @@ -3219,7 +3172,6 @@ def _select_on_joint_draws( check_subgoals=True, check_final_goal=require_goal, run_id="probe", - parameterized_samplers=ctx.parameterized_samplers or None, strip_latent_wait_targets=not ctx.latent_tracking_available, solved_check=solved_check) extra_samples += outcome.total_samples diff --git a/predicators/agent_sdk/bilevel_sketch.py b/predicators/agent_sdk/bilevel_sketch.py index ff66682b1c..c9ddee5535 100644 --- a/predicators/agent_sdk/bilevel_sketch.py +++ b/predicators/agent_sdk/bilevel_sketch.py @@ -8,9 +8,6 @@ - ``sketch_types``: shared dataclasses (``GroundSampler``, ``SketchStep``) that parsing constructs and refinement/execution consume. -- ``sketch_prompts``: ``build_solve_system_prompt`` and - ``build_solve_prompt``, the solve/explore system-prompt and query - builders (rendered from ``prompts/*.md``). - ``sketch_parsing``: the sketch-line grammar - step/plan formatters and the parsers for subgoal / ``~`` ground-sampler annotations and continuous params. @@ -27,8 +24,6 @@ parse_region_annotations, parse_sketch_from_text, \ parse_subgoal_annotations, strip_code_fences, strip_region_annotations, \ strip_subgoal_annotations -from predicators.agent_sdk.sketch_prompts import build_early_stop_note, \ - build_solve_prompt, build_solve_system_prompt from predicators.agent_sdk.sketch_refinement import DeepestFailure, \ InfoScorer, RefineOutcome, StepProbeSuggestion, ground_step, \ refine_and_validate_report, refine_sketch, resolve_refine_timeout, \ @@ -44,9 +39,6 @@ "SketchStep", "StepOutcome", "StepProbeSuggestion", - "build_early_stop_note", - "build_solve_prompt", - "build_solve_system_prompt", "execute_plan_forward", "format_plan_lines", "format_sketch_lines", diff --git a/predicators/agent_sdk/config.py b/predicators/agent_sdk/config.py index 9930b66cdd..4b20e3c85c 100644 --- a/predicators/agent_sdk/config.py +++ b/predicators/agent_sdk/config.py @@ -29,9 +29,7 @@ class SessionConfig: max_turns: int max_buffer_size: int agent_timeout: int - use_docker_sandbox: bool use_local_sandbox: bool - docker_image: str use_scratchpad: bool @classmethod @@ -44,27 +42,16 @@ def from_cfg(cls) -> "SessionConfig": max_turns=CFG.agent_sdk_max_agent_turns_per_iteration, max_buffer_size=CFG.agent_sdk_max_buffer_size, agent_timeout=CFG.agent_sdk_agent_timeout, - use_docker_sandbox=CFG.agent_sdk_use_docker_sandbox, use_local_sandbox=CFG.agent_sdk_use_local_sandbox, - docker_image=CFG.agent_sdk_docker_image, use_scratchpad=CFG.agent_planner_use_scratchpad, ) @dataclass(frozen=True) class RefinementConfig: - """Plan-sketch refinement: search budgets, gates, and ground samplers. - - Consumed at handler entry by ``submit_plan`` (tools/testing.py) and - by the probe's ``refine`` (belief_probe.py). - """ + """Plan-sketch refinement settings the probe's ``refine`` reads.""" ground_samplers: bool - refinement_timeout_per_step: float - refinement_timeout_min: float max_samples_per_step: int - check_subgoals: bool - log_state: bool - use_llm_initial_params: bool @classmethod def from_cfg(cls) -> "RefinementConfig": @@ -72,45 +59,7 @@ def from_cfg(cls) -> "RefinementConfig": # Flags keep their names for experiment-yaml compatibility. return cls( ground_samplers=CFG.agent_bilevel_ground_samplers, - refinement_timeout_per_step=( - CFG.agent_bilevel_refinement_timeout_per_step), - refinement_timeout_min=CFG.agent_bilevel_refinement_timeout_min, max_samples_per_step=CFG.agent_bilevel_max_samples_per_step, - check_subgoals=CFG.agent_bilevel_check_subgoals, - log_state=CFG.agent_bilevel_log_state, - use_llm_initial_params=CFG.agent_bilevel_use_llm_initial_params, - ) - - -@dataclass(frozen=True) -class ValidationConfig: - """Capture-validation rollouts and the cross-attempt journal. - - Consumed at handler entry by ``submit_plan`` (tools.py) and the - probe's ``run(trials=N)`` (belief_probe.py); ``use_journal`` gates - the journal / attempt-log channel. - """ - rollouts: int - rollouts_after_flaky: int - fresh_env: bool - physics_margin: bool - rule_param_margin: bool - necessity: bool - use_journal: bool - - @classmethod - def from_cfg(cls) -> "ValidationConfig": - """Read the validation flags from the live ``CFG``.""" - # Flags keep their names for experiment-yaml compatibility. - return cls( - rollouts=CFG.agent_plan_validation_rollouts, - rollouts_after_flaky=( - CFG.agent_plan_validation_rollouts_after_flaky), - fresh_env=CFG.agent_plan_validation_fresh_env, - physics_margin=CFG.agent_plan_validation_physics_margin, - rule_param_margin=CFG.agent_plan_validation_rule_param_margin, - necessity=CFG.agent_plan_validation_necessity, - use_journal=CFG.agent_solve_use_journal, ) diff --git a/predicators/agent_sdk/docker_agent_runner.py b/predicators/agent_sdk/docker_agent_runner.py deleted file mode 100644 index 560795a273..0000000000 --- a/predicators/agent_sdk/docker_agent_runner.py +++ /dev/null @@ -1,331 +0,0 @@ -"""Agent runner for Docker sandbox. - -Executed inside the Docker container by DockerSessionManager. Loads a -pickled ``QueryInput`` dict, creates a ``ClaudeSDKClient`` session with -both Claude built-in tools (Bash, Read, Write, Edit, Glob, Grep, Task*) -and custom predicator MCP tools, queries the agent, and pickles results -back to a shared directory. - -The predicators source tree is mounted read-only at ``/opt/predicators`` -(via ``PYTHONPATH``) for imports. Curated reference files are available -at ``/sandbox/reference/``. A writable sandbox is at ``/sandbox``. -PreToolUse hooks restrict the agent's built-in tools to ``/sandbox/``. - -Usage (inside Docker):: - - PYTHONPATH=/opt/predicators python3 \ - /opt/predicators/predicators/agent_sdk/docker_agent_runner.py \ - /data/query_input.pkl /data/query_output.pkl -""" -import asyncio -import logging -import sys -import traceback -from typing import Any, Dict, List, Optional - -import dill as pkl - -# Bootstrap: import predicators.utils before anything else so that Python -# resolves the circular import chain (structs → utils → image_patch_wrapper -# → structs) in the correct order. Without this, importing predicators.structs -# first causes image_patch_wrapper to try "from predicators.structs import Mask" -# while structs is still being initialized, raising an ImportError. -import predicators.utils # noqa: F401, E402 # pylint: disable=unused-import - -logging.basicConfig( - level=logging.INFO, - format="%(asctime)s [%(levelname)s] %(name)s: %(message)s", -) -logger = logging.getLogger(__name__) - -# pylint: disable=wrong-import-position -from predicators.agent_sdk.log_formatter import \ - format_conversation_markdown # noqa: E402 -from predicators.agent_sdk.session_base import build_agent_options, \ - build_sandbox_mcp, stream_agent_response # noqa: E402 - - -async def _run_query(query_input: Dict[str, Any]) -> Dict[str, Any]: - """Create a ClaudeSDKClient, query the agent, and collect responses.""" - from claude_agent_sdk import \ - ClaudeSDKClient # pylint: disable=import-outside-toplevel - - ctx = query_input["tool_context"] - tool_names: Optional[List[str]] = query_input.get("tool_names") - - # MCP server and options come from the same helpers the host-side - # managers use; every value is an explicit query_input entry (no CFG - # reads in-container). An invalid reasoning_effort raises here and - # surfaces as an error response, matching host-side validation. - mcp_server, allowed_tools = build_sandbox_mcp(ctx, tool_names) - options = build_agent_options( - system_prompt=query_input["system_prompt"], - model_name=query_input["model_name"], - allowed_tools=allowed_tools, - mcp_server=mcp_server, - max_turns=query_input.get("max_turns", 20), - # Sent by DockerSessionManager from its SessionConfig; the 20MB - # fallback only covers pickles from older hosts. - max_buffer_size=query_input.get("max_buffer_size", 20 * 1024 * 1024), - reasoning_effort=str(query_input.get("reasoning_effort", "")), - ) - - client = ClaudeSDKClient(options=options) - await client.connect() - - # Incremental log file path (on shared /data or /log volume) - log_path = query_input.get("log_path") - log_meta = {"query": query_input.get("message", "")} - - def _flush_log(collected: List[Dict[str, Any]]) -> None: - """Write current conversation state as markdown to the log file.""" - if not log_path: - return - try: - content = format_conversation_markdown(collected, - title="Docker Query", - meta=log_meta) - with open(log_path, "w", encoding="utf-8") as lf: - lf.write(content) - except Exception: # pylint: disable=broad-except - pass # Don't let logging errors break the agent - - # Docker-specific stderr reporting for real-time host visibility - # (the host streams container stderr into its own log). - def _report_block(dt: float, preview: str) -> None: - print(f"[+{dt:.2f}s] {preview}", file=sys.stderr, flush=True) - - def _report_result(entry: Dict[str, Any]) -> None: - print( - f"Agent iteration complete. " - f"Turns: {entry.get('num_turns', '?')}, " - f"Cost: ${entry.get('total_cost_usd', '?')}", - file=sys.stderr, - flush=True) - - try: - collected = await stream_agent_response( - client, - query_input["message"], - log_label="Docker runner", - report_block=_report_block, - on_result=_report_result, - flush=_flush_log, - ) - finally: - try: - await client.disconnect() - except Exception: # pylint: disable=broad-except - pass - - return { - "responses": collected, - } - - -def _rehash_objects_after_unpickle(ctx: Any) -> None: - """Fix stale Object hash caches after cross-process unpickling. - - ``Object.__hash__`` returns a ``cached_property`` (``_hash``) that - stores ``hash(str(self))``. Python randomises string hashes across - processes (PYTHONHASHSEED), so cached values from the *pickling* - process are stale here. When the option-model simulator later - creates fresh Objects (e.g. ``self._robot`` in ``_get_state``), - their hashes differ from the unpickled Objects, causing KeyError on - ``State.data`` dict lookups. - - Fix: clear every Object's cached ``_hash`` (and ``_str``) so it is - re-computed with the current process's hash seed, then rebuild every - ``State.data`` dict so its internal hash-table is consistent. - """ - from predicators.structs import \ - State # pylint: disable=import-outside-toplevel - - seen: set = set() - - def _clear(obj: Any) -> None: - oid = id(obj) - if oid in seen: - return - seen.add(oid) - obj.__dict__.pop("_hash", None) - obj.__dict__.pop("_str", None) - - def _process_state(state: Any) -> None: - if state is None or not isinstance(state, State): - return - for obj in list(state.data.keys()): - _clear(obj) - # Rebuild dict so Python re-hashes keys with current seed. A - # comprehension (not ``dict(...)``) is load-bearing: ``dict(d)`` - # copies each entry's stored hash without calling ``__hash__``, - # so it would preserve exactly the stale table this repairs. - # pylint: disable-next=unnecessary-comprehension - state.data = {obj: vals for obj, vals in state.data.items()} - - def _process_atoms(atoms: Any) -> None: - for atom in atoms: - for obj in atom.objects: - _clear(obj) - - def _process_task(task: Any) -> None: - # Task has .init (State) and .goal (Set[GroundAtom]) - # EnvironmentTask has .init_obs and .goal_description - if hasattr(task, "init"): - _process_state(task.init) - if hasattr(task, "init_obs"): - _process_state(task.init_obs) - for attr in ("goal", "alt_goal", "goal_description", "alt_goal_desc"): - atoms = getattr(task, attr, None) - # goal_description may be a plain NL string on - # EnvironmentTask; only atom collections carry Objects. - if atoms and not isinstance(atoms, str): - _process_atoms(atoms) - - # Train tasks - for task in getattr(ctx, "train_tasks", []): - _process_task(task) - - # Current task - if ctx.current_task is not None: - _process_task(ctx.current_task) - - # Example state - _process_state(getattr(ctx, "example_state", None)) - - # Trajectories - for traj in (getattr(ctx, "offline_trajectories", []) + - getattr(ctx, "online_trajectories", [])): - for state in traj.states: - _process_state(state) - - -def main() -> None: - """Entry point for Docker agent runner.""" - if len(sys.argv) != 3: - print(f"Usage: {sys.argv[0]} ", - file=sys.stderr) - sys.exit(1) - - input_path = sys.argv[1] - output_path = sys.argv[2] - - logger.info("Docker agent runner starting: input=%s output=%s", input_path, - output_path) - - # Load query input - with open(input_path, "rb") as f: - query_input = pkl.load(f) - - # Restore host CFG settings (arg-specific settings like - # max_num_steps_option_rollout are not set by default import) - if "cfg_snapshot" in query_input: - from predicators.settings import \ - CFG # pylint: disable=import-outside-toplevel - for k, v in query_input["cfg_snapshot"].items(): - setattr(CFG, k, v) - - # Fix stale Object hash caches from cross-process pickling. - ctx = query_input.get("tool_context") - if ctx is not None: - _rehash_objects_after_unpickle(ctx) - - # Recreate option model — the simulator (e.g. PyBullet physics - # server) is process-local and cannot survive pickling. - if ctx is not None and ctx.option_model is not None: - from predicators.option_model import \ - create_option_model # pylint: disable=import-outside-toplevel - from predicators.settings import \ - CFG as _cfg # pylint: disable=import-outside-toplevel - logger.info("Recreating option model (%s) inside Docker...", - _cfg.option_model_name) - ctx.option_model = create_option_model( - _cfg.option_model_name, - skip_residual_dynamics=_cfg.agent_planner_use_base_simulator) - # Sync with all options in context (GT + any previously proposed) - # after the model has its physics server set up. - ctx.option_model._name_to_parameterized_option = { # pylint: disable=protected-access - o.name: o - for o in ctx.options - } - - # Recreate SkillConfig in skill_factory_context — the robot's - # physics_client_id is process-local and stale after pickling. - if (ctx is not None - and ctx.skill_factory_context.get("skill_config") is not None): - from predicators.settings import \ - CFG as _cfg # pylint: disable=import-outside-toplevel - if _cfg.env.startswith("pybullet"): - try: - # pylint: disable=import-outside-toplevel,reimported - from predicators import utils as _utils - from predicators.envs.base_env import BaseEnv - from predicators.envs.pybullet_env import PyBulletEnv - from predicators.ground_truth_models.skill_factories import \ - SkillConfig - - # Find the PyBulletEnv subclass (envs already imported above - # by create_option_model → create_new_env). - env_cls = None - for cls in _utils.get_all_subclasses(BaseEnv): - if (not cls.__abstractmethods__ - and issubclass(cls, PyBulletEnv) - and cls.get_name() == _cfg.env): - env_cls = cls - break - - if env_cls is None: - logger.warning( - "Could not find PyBulletEnv for %s; " - "skill_config NOT recreated", _cfg.env) - else: - _, robot, _ = env_cls.initialize_pybullet(using_gui=False) - ctx.skill_factory_context["skill_config"] = SkillConfig( - robot=robot, - open_fingers_joint=robot.open_fingers, - closed_fingers_joint=robot.closed_fingers, - fingers_state_to_joint=( - env_cls._fingers_state_to_joint), # pylint: disable=protected-access - max_vel_norm=_cfg.pybullet_max_vel_norm, - ik_validate=_cfg.pybullet_ik_validate, - robot_init_tilt=getattr(env_cls, 'robot_init_tilt', - 0.0), - robot_init_wrist=getattr(env_cls, 'robot_init_wrist', - 0.0), - ) - logger.info( - "Recreated SkillConfig inside Docker for %s " - "(physics_client_id=%d)", _cfg.env, - robot.physics_client_id) - except Exception as e: # pylint: disable=broad-except - logger.error("Failed to recreate SkillConfig in Docker: %s", - e, - exc_info=True) - - logger.info("Loaded query input: message length=%d, model=%s", - len(query_input.get("message", "")), - query_input.get("model_name", "?")) - - # Run the query - try: - query_output = asyncio.run(_run_query(query_input)) - except Exception as e: # pylint: disable=broad-except - logger.error("Fatal error in agent runner: %s\n%s", e, - traceback.format_exc()) - query_output = { - "responses": [{ - "type": "error", - "error": str(e) - }], - } - - # Save output - with open(output_path, "wb") as f: - pkl.dump(query_output, f) - - logger.info("Docker agent runner finished: %d responses", - len(query_output.get("responses", []))) - - -if __name__ == "__main__": - main() diff --git a/predicators/agent_sdk/docker_sandbox.py b/predicators/agent_sdk/docker_sandbox.py deleted file mode 100644 index 69ec6dd8c6..0000000000 --- a/predicators/agent_sdk/docker_sandbox.py +++ /dev/null @@ -1,499 +0,0 @@ -"""Docker-sandboxed agent session manager. - -Runs ``ClaudeSDKClient`` inside a Docker container so that the agent's -built-in tools (Bash, Read, Write, Edit, Glob, Grep, Task*) all execute -in an isolated environment. Custom predicator MCP tools are created in-process -inside the container via the same ``create_mcp_tools()`` code used on -the host. - -The host predicators source tree is mounted read-only at -``/opt/predicators`` for Python imports (``PYTHONPATH``). PreToolUse -hooks block the agent's built-in tools (Read, Write, Edit, Glob, Grep) -from accessing anything outside ``/sandbox/``, so the agent cannot -browse environment source code or ground truth models directly. Curated -reference files are copied into ``/sandbox/reference/`` for the agent to -read. The agent can write and run Python scripts in ``/sandbox/``, and -``from predicators.structs import State`` works via the mount. - -Shared data (pickled context and results) passes through ``/data``. - -Behavioral notes relative to the shared base -(:mod:`predicators.agent_sdk.session_base`): - -- ``query()`` is a subprocess orchestrator: each call runs one fresh - container (no persistent client), so ``start_session``, ``close``, - and ``_recover_session`` are no-ops. -- The incremental markdown log is written in-container; the host only - prepends a metadata header afterwards. -- Cost accounting reuses the base delta scheme with the baseline reset - to zero per query, since every container session starts from zero. - -Usage ------ -When the ``agent_sdk_use_docker_sandbox`` flag is ``True``, the -``AgentSessionMixin`` creates a ``DockerSessionManager`` in place of the -normal ``AgentSessionManager``. The interface is identical:: - - manager = DockerSessionManager(...) - responses = await manager.query("Solve this task...") - await manager.close() - -Build the image first:: - - bash docker/build.sh -""" -import datetime -import json -import logging -import os -import shutil -import subprocess -import sys -import tempfile -import uuid -from pathlib import Path -from typing import Any, Dict, List, Optional - -import dill as pkl - -from predicators.agent_sdk.config import SessionConfig -from predicators.agent_sdk.sandbox_prompts import build_sandbox_system_prompt -from predicators.agent_sdk.session_base import SandboxSessionManagerBase -from predicators.agent_sdk.tools import ToolContext, session_log_filename -from predicators.settings import CFG - -logger = logging.getLogger(__name__) - -# Grace period past the per-query agent timeout before the container is -# force-killed (covers container startup + result pickling). -_CONTAINER_TIMEOUT_SLACK_S = 120 - -# Tail sizes for error reporting when a container run fails. -_STDIO_TAIL_CHARS = 2000 -_STDERR_TAIL_LINES = 20 - -# Build Docker-specific prompts from shared templates. -# CLAUDE.md (sandbox mechanics only; see build_claude_md) is written -# into the sandbox when it is populated. -_SANDBOX_SYSTEM_PROMPT = build_sandbox_system_prompt( - env_description="an isolated Docker sandbox", - workspace_description="/sandbox/", - ref_path="/sandbox/reference/", -) - -# --------------------------------------------------------------------------- -# Helper functions -# --------------------------------------------------------------------------- - - -def _get_claude_oauth_token() -> Optional[str]: - """Extract the Claude Code OAuth access token from the macOS Keychain. - - Returns ``None`` on non-macOS platforms or when the token cannot be - found. On macOS, ``claude login`` stores credentials under the - service name ``"Claude Code-credentials"``. - """ - if sys.platform != "darwin": - return None - try: # type: ignore[unreachable] - result = subprocess.run( - [ - "security", "find-generic-password", "-s", - "Claude Code-credentials", "-w" - ], - capture_output=True, - text=True, - timeout=5, - check=False, - ) - if result.returncode != 0: - return None - creds = json.loads(result.stdout.strip()) - return creds.get("claudeAiOauth", {}).get("accessToken") - except (subprocess.SubprocessError, json.JSONDecodeError, KeyError): - return None - - -# _flush_log stays unimplemented on purpose: logs flush inside the -# container, and a host-side call should fail loudly. -# pylint: disable-next=abstract-method -class DockerSessionManager(SandboxSessionManagerBase): - """Runs ClaudeSDKClient inside Docker with built-in + custom MCP tools. - - Matches the ``AgentSessionManager`` interface so that all agent-based - approaches work unchanged. Each ``query()`` call: - - 1. Serializes ``ToolContext`` + message to pickle in a temp directory. - 2. Runs ``docker run ...`` with the predicators source mounted at - ``/opt/predicators:ro`` (for Python imports) and a curated sandbox - at ``/sandbox`` (for agent file operations). - 3. Inside Docker, the runner script creates ``ClaudeSDKClient`` with - both built-in tools AND custom MCP tools, queries the agent, and - pickles back responses + mutated proposals. - 4. Host reads back the pickled results. - - PreToolUse hooks restrict the agent's built-in tools (Read, Write, - Edit, Glob, Grep) to ``/sandbox/`` only. Python imports via - ``PYTHONPATH`` are unaffected. - """ - - _log_label = "Docker" - - def __init__( - self, - system_prompt: str, - log_dir: str, - model_name: str, - tool_context: ToolContext, - tool_names: Optional[List[str]] = None, - image: str = "predicators-sandbox", - extra_reference_files: Optional[Dict[str, str]] = None, - phase: Optional[str] = None, - config: Optional[SessionConfig] = None, - ) -> None: - # Append sandbox instructions to the system prompt - super().__init__(system_prompt=system_prompt + _SANDBOX_SYSTEM_PROMPT, - log_dir=log_dir, - model_name=model_name, - tool_context=tool_context, - tool_names=tool_names, - extra_reference_files=extra_reference_files, - phase=phase, - config=config) - self._image = image - self._last_kind: str = "query" - - # -- Session lifecycle -- - - async def start_session(self) -> None: - """No-op: each query() is a fresh docker run.""" - - async def close(self) -> None: - """No-op: the sandbox directory is kept on disk for inspection.""" - - async def _recover_session(self) -> None: - """No-op: each query is independent.""" - - async def query(self, - message: str, - kind: str = "query") -> List[Dict[str, Any]]: - """Run the agent in Docker and return collected response messages. - - Returns the same ``List[Dict[str, Any]]`` format as - ``AgentSessionManager.query()``. - """ - self._query_count += 1 - self._tool_context.turn_id = self._query_count - self._last_kind = kind - - # Ensure sandbox is set up (lazy init, persists across queries) - self._ensure_sandbox_dir() - - # 1. Create temp directory for data exchange - tmp_dir = tempfile.mkdtemp(prefix="pred-docker-") - input_path = os.path.join(tmp_dir, "query_input.pkl") - output_path = os.path.join(tmp_dir, "query_output.pkl") - - # Compute final log filename upfront so the container can write - # directly to the log directory (incremental updates visible on host). - # Counter-first layout: alphabetical sort matches chronological - # order across mixed ``learn``/``test``/``explore`` phases. - timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S") - log_filename = session_log_filename( - self._query_count, kind, timestamp, - getattr(self._tool_context, "test_task_idx", None)) - if self._log_dir: - os.makedirs(self._log_dir, exist_ok=True) - incremental_log_path = os.path.join(self._log_dir, log_filename) - else: - incremental_log_path = os.path.join(tmp_dir, "query_log.md") - - try: - # 2. Pickle QueryInput - # Tell the container where to write the incremental log. - # If _log_dir is set, it's mounted at /log inside the container. - container_log_path = (f"/log/{log_filename}" - if self._log_dir else "/data/query_log.md") - query_input = { - "tool_context": self._tool_context, - "message": message, - "system_prompt": self._system_prompt, - "model_name": self._model_name, - "max_turns": self._config.max_turns, - "max_buffer_size": self._config.max_buffer_size, - "reasoning_effort": self._config.reasoning_effort, - "tool_names": self._tool_names, - "cfg_snapshot": dict(CFG.__dict__), - "log_path": container_log_path, - } - with open(input_path, "wb") as f: - pkl.dump(query_input, f) - - logger.info( - "Docker query %d: message length=%d, model=%s", - self._query_count, - len(message), - self._model_name, - ) - - # 3. Build docker run command. Resolve authentication once - # per query (the Keychain OAuth lookup is a subprocess call - # shared by the command and env builders). - api_key = os.environ.get("ANTHROPIC_API_KEY") - oauth_token = None if api_key else _get_claude_oauth_token() - container_name = f"pred-sandbox-{uuid.uuid4().hex[:8]}" - docker_cmd = self._build_docker_command(container_name, tmp_dir, - api_key, oauth_token) - - # 4. Run Docker container - logger.info( - "Starting Docker sandbox: container=%s image=%s", - container_name, - self._image, - ) - env = self._build_env(api_key, oauth_token) - - proc = subprocess.Popen( - docker_cmd, - env=env, - stdout=subprocess.PIPE, - stderr=subprocess.PIPE, - text=True, - ) - - # Stream stderr in real-time so tool calls / agent messages - # appear on the host terminal as they happen. - stderr_lines: List[str] = [] - try: - timeout_sec = (self._config.agent_timeout + - _CONTAINER_TIMEOUT_SLACK_S) - import threading # pylint: disable=import-outside-toplevel - - def _stream_stderr() -> None: - assert proc.stderr is not None - for line in proc.stderr: - line = line.rstrip("\n") - stderr_lines.append(line) - logger.info("%s", line) - - stderr_thread = threading.Thread(target=_stream_stderr, - daemon=True) - stderr_thread.start() - - # Wait for stdout (captured for error reporting) - stdout_data = proc.stdout.read() if proc.stdout else "" - proc.wait(timeout=timeout_sec) - stderr_thread.join(timeout=5) - except subprocess.TimeoutExpired: - proc.kill() - proc.wait() - logger.error("Docker container timed out after %ds", - timeout_sec) - stdout_data = "" - - if proc.returncode != 0: - logger.error( - "Docker container exited with code %d.\nstdout: %s\n" - "stderr (last 2000 chars): %s", - proc.returncode, - stdout_data[-_STDIO_TAIL_CHARS:] - if stdout_data else "(empty)", - "\n".join(stderr_lines)[-_STDIO_TAIL_CHARS:] - if stderr_lines else "(empty)", - ) - else: - logger.info("Docker container exited successfully.") - - # 5. Load query output - if os.path.exists(output_path): - with open(output_path, "rb") as f_in: - query_output = pkl.load(f_in) - - responses = query_output.get("responses", []) - # Track costs/turns via the base delta accounting. Each - # docker query is a fresh in-container session whose - # cumulative cost restarts from zero, so reset the delta - # baseline first: every result then charges its full - # cumulative cost. - self._last_cost_usd = 0.0 - for resp in responses: - if resp.get("type") == "result": - self._account_result(resp) - else: - logger.error( - "No output pickle found at %s. Container may have " - "crashed.", output_path) - responses = [{ - "type": - "error", - "error": - (f"Docker container failed (exit code " - f"{proc.returncode}). " - f"stderr: {''.join(stderr_lines[-_STDERR_TAIL_LINES:])}"), - }] - - # 7. Finalize query log - the incremental log was written - # directly to _log_dir as markdown (updated per-message). - # Prepend host metadata header now that the container is done. - if os.path.exists(incremental_log_path) and self._log_dir: - try: - with open(incremental_log_path, encoding="utf-8") as lf: - existing = lf.read() - header_lines = [ - f"- **Query:** {self._query_count}", - f"- **Timestamp:** {timestamp}", - f"- **Session:** {self._session_id}", - f"- **Image:** {self._image}", - "", - "", - ] - with open(incremental_log_path, "w", - encoding="utf-8") as lf: - lf.write("\n".join(header_lines) + existing) - logger.info("Finalized docker query/response at %s", - incremental_log_path) - except Exception: # pylint: disable=broad-except - logger.warning("Failed to enrich log at %s", - incremental_log_path, - exc_info=True) - else: - self._save_query_response_log(message, responses) - - # Track in-memory for conversation replay - self._conversation_log.append({ - "query": message, - "response": responses, - }) - - self._track_fatal_response(responses) - return responses - - finally: - # Cleanup temp data directory (sandbox persists across queries) - shutil.rmtree(tmp_dir, ignore_errors=True) - - def _session_info_extras(self) -> Dict[str, Any]: - """Extra session-info keys: manager type + container image.""" - return { - "session_type": "docker", - "docker_image": self._image, - } - - # -- Internal helpers -- - - def _build_docker_command(self, container_name: str, tmp_dir: str, - api_key: Optional[str], - oauth_token: Optional[str]) -> List[str]: - """Build the ``docker run`` command.""" - cmd = [ - "docker", - "run", - "--rm", - "--name", - container_name, - "--cap-add=NET_ADMIN", - "--cap-add=NET_RAW", - ] - - # Authentication: prefer ANTHROPIC_API_KEY, fall back to OAuth - if api_key: - cmd += ["-e", "ANTHROPIC_API_KEY"] - elif oauth_token: - # The token value itself is added to env in _build_env() - cmd += ["-e", "CLAUDE_CODE_OAUTH_TOKEN"] - else: - # Fall back to bind-mounting ~/.claude - claude_cfg = Path( - os.environ.get("CLAUDE_CONFIG_DIR", - str(Path.home() / ".claude"))) - cmd += ["-v", f"{claude_cfg}:/home/node/.claude"] - - # Mount predicators source for Python imports (hidden from agent - # tools by the PreToolUse hook - only Python's import system can - # read these files). - cmd += ["-v", f"{self._repo_root}:/opt/predicators:ro"] - cmd += ["-e", "PYTHONPATH=/opt/predicators"] - - # Mount curated sandbox directory - cmd += ["-v", f"{self._sandbox_dir}:/sandbox"] - - # Mount data exchange directory - cmd += ["-v", f"{tmp_dir}:/data"] - - # Mount log directory for incremental log updates visible on host - if self._log_dir: - log_dir_abs = os.path.abspath(self._log_dir) - cmd += ["-v", f"{log_dir_abs}:/log"] - - # Working directory - cmd += ["-w", "/sandbox"] - - # Image - cmd.append(self._image) - - # Command: run the agent runner script from the mounted source - cmd += [ - "python3", - "-u", - "/opt/predicators/predicators/agent_sdk/docker_agent_runner.py", - "/data/query_input.pkl", - "/data/query_output.pkl", - ] - - return cmd - - def _build_env(self, api_key: Optional[str], - oauth_token: Optional[str]) -> Dict[str, str]: - """Build environment dict for the docker subprocess.""" - # Pass through host env, stripping CLAUDECODE* vars - env = { - k: v - for k, v in os.environ.items() if not k.startswith("CLAUDECODE") - } - - # Ensure ANTHROPIC_API_KEY is passed through if set - if api_key: - env["ANTHROPIC_API_KEY"] = api_key - elif oauth_token: - env["CLAUDE_CODE_OAUTH_TOKEN"] = oauth_token - - return env - - def _save_query_response_log(self, query: str, - response: List[Dict[str, Any]]) -> None: - """Save query and response to a timestamped markdown file.""" - if not self._log_dir: - return - - timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S") - kind = self._last_kind - filename = session_log_filename( - self._query_count, kind, timestamp, - getattr(self._tool_context, "test_task_idx", None)) - filepath = os.path.join(self._log_dir, filename) - - lines = [ - f"- **Query:** {self._query_count}", - f"- **Timestamp:** {timestamp}", - f"- **Session:** {self._session_id}", - f"- **Image:** {self._image}", - "", - "# Docker Query", - "", - "## Prompt", - "", - query, - "", - "## Response", - "", - ] - for entry in response: - lines.append( - f"```json\n{json.dumps(entry, indent=2, default=str)}\n```") - lines.append("") - - os.makedirs(self._log_dir, exist_ok=True) - with open(filepath, "w", encoding="utf-8") as f: - f.write("\n".join(lines)) - - logger.info("Saved docker query/response to %s", filepath) diff --git a/predicators/agent_sdk/journal.py b/predicators/agent_sdk/journal.py index 6d45998f3b..16b101a9ad 100644 --- a/predicators/agent_sdk/journal.py +++ b/predicators/agent_sdk/journal.py @@ -1,35 +1,24 @@ -"""Persistent per-run solve journal and attempt log. - -Two markdown files in the sandbox that carry knowledge across solve -attempts, test tasks, and learning cycles: - -- ``journal.md`` is the AGENT's notebook. Solve and learn sessions - append to it with the ordinary file tools (no dedicated tool): short - factual entries - what was tried with exact parameters, what was - measured, what to try differently. The prompts ask for facts and - measurements rather than verdicts: a recorded "X is impossible" - from a failed attempt would re-import exactly the anchoring a fresh - context is meant to shed, while "tried yaws 0-15 deg at x in - [0.50, 0.54], all stopped >=5 cm short" steers the next attempt - without foreclosing it. -- ``attempts.md`` is the HARNESS's log, never edited by the agent: - each task's goal + initial state (once per task) and each - attempt's outcome and captured or best refused plan, so the - essentials of every attempt are on record even when the agent - writes nothing. - -Fresh-context solve sessions read both from their prompt (tail-capped -so recent attempts stay intact), so knowledge travels through these -curated channels instead of raw transcript history. - -Phase lifecycle: learning-phase content persists for the whole run -and accumulates across online-learning cycles, so every evaluation -starts from all learning knowledge so far. Test-phase additions live -only for their own evaluation: at ``end_test_phase`` the approach -archives both files to the run's log dir (outside the sandbox, so the -agent cannot read them) and rolls them back to their pre-test content -via :func:`read_raw` / :func:`restore` - entries written while -solving one evaluation's test tasks must not leak into the next. +"""Persistent per-run journal and round log. + +Two markdown files in the sandbox that carry knowledge across the rounds +and levels of a continual run: + +- ``journal.md`` is the AGENT's notebook. The agent appends to it with + the ordinary file tools (no dedicated tool): short factual entries - + what was tried with exact parameters, what was measured, what to try + differently. The prompts ask for facts and measurements rather than + verdicts: a recorded "X is impossible" from a failed attempt would + re-import exactly the anchoring a fresh context is meant to shed, + while "tried yaws 0-15 deg at x in [0.50, 0.54], all stopped >=5 cm + short" steers the next attempt without foreclosing it. +- ``attempts.md`` is the HARNESS's log, never edited by the agent: one + entry per round (what the agent did in the environment) and per model + change, so the essentials of every round are on record even when the + agent writes nothing. + +Each round's query injects both (tail-capped so recent entries stay +intact), so knowledge travels through these curated channels instead of +raw transcript history. """ from __future__ import annotations @@ -38,133 +27,33 @@ from typing import Optional JOURNAL_FILENAME = "journal.md" -# The harness-owned attempt log (task contexts, attempt outcomes). +# The harness-owned round log. ATTEMPTS_FILENAME = "attempts.md" -# Per-entry cap for harness attempt-log entries: the first entry per -# task embeds the init-state feature dict (the prompt's own -# representation) and a captured plan. The writer orders the layout -# block last, so tail truncation at this cap can only ever cut layout, -# never the outcome or the captured plan. +# Per-entry cap for harness log entries. MAX_ENTRY_CHARS = 4000 -MAX_AUTO_ENTRY_CHARS = MAX_ENTRY_CHARS -# Cap on how much of each file is injected into a solve prompt. -# Tail-biased: recent attempts (usually the same task) matter most. +# Cap on how much of each file is injected into a query. Tail-biased: +# recent rounds matter most. MAX_PROMPT_CHARS = 6000 -# The learn-phase-maintained domain strategy document. Unlike the -# append-only journal (facts and measurements), strategy.md is a LIVING -# document the learn agent rewrites freely each cycle: its best current -# natural-language account of how to solve tasks in this domain. Solve -# prompts inject it as explicitly-advisory reference. -STRATEGY_FILENAME = "strategy.md" - -# Cap on how much strategy is injected into a solve prompt. Head-biased -# (unlike the journal): the document is curated, so its lead carries the -# headline strategy and a tail truncation only cuts detail. -MAX_STRATEGY_PROMPT_CHARS = 4000 - - -def journal_path(sandbox_dir: str) -> str: - """Host path of the run's journal file.""" - return os.path.join(sandbox_dir, JOURNAL_FILENAME) - - -def attempts_path(sandbox_dir: str) -> str: - """Host path of the run's harness-owned attempt log.""" - return os.path.join(sandbox_dir, ATTEMPTS_FILENAME) - - -def strategy_path(sandbox_dir: str) -> str: - """Host path of the run's domain strategy document.""" - return os.path.join(sandbox_dir, STRATEGY_FILENAME) - - -def read_strategy(sandbox_dir: Optional[str], - max_chars: int = MAX_STRATEGY_PROMPT_CHARS) -> str: - """Strategy document content for prompt injection ("" when absent). - - Head-biased truncation: the document is curated by the learn agent, - so the front holds the headline strategy; a truncation notice marks - the cut so readers know detail was dropped. - """ - if not sandbox_dir: - return "" - path = strategy_path(sandbox_dir) - if not os.path.isfile(path): - return "" - with open(path, "r", encoding="utf-8") as f: - content = f.read().strip() - if len(content) > max_chars: - # Cut at a line boundary, never mid-word. - head = content[:max_chars] - cut = head.rfind("\n") - if cut > 0: - head = head[:cut] - content = (head.rstrip() + - "\n[strategy truncated at the prompt cap - read " - f"./{STRATEGY_FILENAME} for the rest]") - return content - def append_entry(sandbox_dir: str, header: str, body: str, max_chars: int = MAX_ENTRY_CHARS, - filename: str = ATTEMPTS_FILENAME) -> Optional[str]: - """Append one harness entry; returns a truncation notice or None. + filename: str = ATTEMPTS_FILENAME) -> None: + """Append one harness entry. ``header`` becomes a ``###
`` line; ``body`` is written - verbatim below it, truncated at ``max_chars`` (default - :data:`MAX_ENTRY_CHARS`; harness auto-entries pass - :data:`MAX_AUTO_ENTRY_CHARS`). + verbatim below it, truncated at ``max_chars``. """ os.makedirs(sandbox_dir, exist_ok=True) - note: Optional[str] = None body = body.strip() if len(body) > max_chars: body = body[:max_chars].rstrip() body += "\n[entry truncated at the per-entry size cap]" - note = (f"entry truncated to {max_chars} chars - keep journal " - "entries short and factual") with open(os.path.join(sandbox_dir, filename), "a", encoding="utf-8") as f: f.write(f"### {header.strip()}\n{body}\n\n") - return note - - -def read_raw(sandbox_dir: Optional[str], - filename: str = JOURNAL_FILENAME) -> Optional[str]: - """Exact file content, or None if the file does not exist. - - Unlike :func:`read_journal` there is no prompt trimming and the - absent-file case is distinguishable from an empty file, so the - result is a faithful snapshot for :func:`restore`. - """ - if not sandbox_dir: - return None - path = os.path.join(sandbox_dir, filename) - if not os.path.isfile(path): - return None - with open(path, "r", encoding="utf-8") as f: - return f.read() - - -def restore(sandbox_dir: str, - snapshot: Optional[str], - filename: str = JOURNAL_FILENAME) -> None: - """Reset the file to a :func:`read_raw` snapshot. - - A ``None`` snapshot means the file did not exist, so it is removed - if present. - """ - path = os.path.join(sandbox_dir, filename) - if snapshot is None: - if os.path.isfile(path): - os.remove(path) - return - os.makedirs(sandbox_dir, exist_ok=True) - with open(path, "w", encoding="utf-8") as f: - f.write(snapshot) def read_journal(sandbox_dir: Optional[str], diff --git a/predicators/agent_sdk/learn_prompts.py b/predicators/agent_sdk/learn_prompts.py deleted file mode 100644 index 6383757024..0000000000 --- a/predicators/agent_sdk/learn_prompts.py +++ /dev/null @@ -1,389 +0,0 @@ -"""Prompt construction for the learning (simulator synthesis) phase. - -Rendered from ``learn_system.md``, ``learn_message.md``, -``learn_predicate_invention.md``, and ``learn_partial_observability.md`` -in ``predicators/agent_sdk/prompts`` (see :mod:`prompt_templates`). The -approach classes gather the per-instance values (digests, data roster, -reports, paths) and call these pure builders, so every prompt can be -rendered and reviewed without a live session. -""" -import re -from typing import Any, Mapping, Sequence - -from predicators.agent_sdk.prompt_templates import render - -_BLANK_RUN_RE = re.compile(r"\n{3,}") - - -def _join(parts: Sequence[str]) -> str: - text = "\n\n".join(p.strip("\n") for p in parts if p and p.strip()) - return _BLANK_RUN_RE.sub("\n\n", text).strip("\n") + "\n" - - -# --------------------------------------------------------------------------- -# System prompt -# --------------------------------------------------------------------------- - - -def build_learn_system_prompt( - *, - partially_observable: bool, - residual_rule_signature: str, - scene_viz_hint: str, - physical_params_section: str = "", - extra_sections: Sequence[str] = (), - latent_extra_sections: Sequence[str] = (), - workflow_extra: str = "", - declared_params_only: bool = False, -) -> str: - """Compose the synthesis system prompt. - - The shared subclass contract supplies dynamics, fitting and optional - model-state guidance. Extra sections add predicate invention and - workflow requirements for the particular learning arm. - """ - # Retain the keyword arguments for callers outside this package. - del residual_rule_signature, scene_viz_hint - parts = [ - render("learn_system", "intro"), - render("subclass_model", "simulator"), - render("subclass_model", "dynamics"), - physical_params_section, - render("learn_system", "declared_params") - if declared_params_only else "", - render("subclass_model", "tools"), - *extra_sections, - ] - if partially_observable: - parts.append( - render("subclass_model", - "memory", - base_class="BaseSimulator", - estimate_errors="errors in the model and noisy input")) - parts.extend(latent_extra_sections) - parts += [ - render("learn_system", "plan_format"), - render("learn_system", "deliverables"), - render("learn_system", - "workflow", - workflow_extra=(" " + - workflow_extra) if workflow_extra else ""), - ] - return _join(parts) - - -def render_physical_params_section( - info: Mapping[str, Mapping[str, Any]]) -> str: - """The base-physics parameter menu for a revealed parameter menu. - - ``info`` maps a parameter name to its ``default``, ``lo``, ``hi``, - ``description``, and optional ``scale``; empty input renders - nothing, so envs without a menu never see the feature mentioned. - """ - if not info: - return "" - lines = [] - for name, meta in info.items(): - scale_note = (", fitted in log-space" - if meta.get("scale") == "log" else "") - lines.append(f"- `{name}` (built-in {meta['default']:.4g}, fit " - f"box [{meta['lo']:.4g}, {meta['hi']:.4g}]" - f"{scale_note}): {meta['description']}") - return render("subclass_model", - "physical_params", - param_list="\n".join(lines)) - - -def render_predicate_invention_section(scene_workbench: str) -> str: - """The predicate-invention system-prompt section.""" - return render("learn_predicate_invention", - "system", - scene_workbench=scene_workbench) - - -def render_predicate_latent_section() -> str: - """The predicate-side latent guidance (invention arms, PO only).""" - return render("learn_partial_observability", "predicates") - - -def render_predicate_workflow_extra() -> str: - """The invention arm's addition to the workflow's validation step.""" - return render("learn_predicate_invention", "workflow_extra") - - -# --------------------------------------------------------------------------- -# First message -# --------------------------------------------------------------------------- - - -def build_learn_message( - *, - n_trajs: int, - n_transitions: int, - n_demos: int, - n_interaction: int, - trajectory_listing: str, - structs_ref: str, - inferred_hint: str, - predicate_listing: str, - types_digest: str, - options_digest: str, - simulator_file: str, - objective_block: str = "", - prior_state_block: str = "", - divergence_block: str = "", - base_sim_block: str = "", - tools_block: str = "", - extra_messages: Sequence[str] = (), -) -> str: - """Compose the synthesis session's first message. - - Every block argument is already rendered (see the ``render_*`` - helpers below) or empty. ``extra_messages`` (predicate invention, - partial observability, sampler synthesis) are appended in order. - """ - body = render( - "learn_message", - "skeleton", - n_trajs=str(n_trajs), - n_transitions=str(n_transitions), - n_demos=str(n_demos), - n_interaction=str(n_interaction), - trajectory_listing=trajectory_listing.strip("\n"), - objective_block=objective_block, - prior_state_block=prior_state_block, - divergence_block=divergence_block, - structs_ref=structs_ref, - base_sim_block=base_sim_block, - inferred_hint=inferred_hint, - predicate_listing=predicate_listing, - types_digest=types_digest.strip("\n"), - options_digest=options_digest.strip("\n"), - tools_block=tools_block, - simulator_file=simulator_file, - ) - return _join([body, *extra_messages]) - - -def render_divergence_block(report: str, has_prior_model: bool) -> str: - """The start-of-session residual report section.""" - return render("learn_message", - "divergence_prior" if has_prior_model else "divergence_base", - report=report.strip("\n")) - - -def render_base_sim_block(refs: Sequence[str]) -> str: - """The base-simulator source listing, or empty.""" - if not refs: - return "" - return render("learn_message", - "base_sim", - ref_listing="\n".join(f" - {r}" for r in refs)) - - -def render_tools_block(tool_names: Sequence[str]) -> str: - """The session's tool roster, or empty.""" - if not tool_names: - return "" - return render("learn_message", - "tools", - tool_listing="\n".join(f" - {t}" for t in tool_names)) - - -def render_objective_block(description: str) -> str: - """The env's public task objective section, or empty.""" - if not description: - return "" - return render("learn_message", "objective", description=description) - - -def render_prior_state_block(prior_files: Sequence[str]) -> str: - """The prior-cycle-state paragraph for the artifacts found, or empty.""" - if not prior_files: - return "" - return render("learn_message", - "prior_state", - prior_files=" and ".join(prior_files)) - - -def render_predicate_invention_message(predicates_file: str, - goal_block: str) -> str: - """The invention arm's addition to the first message.""" - return render("learn_predicate_invention", - "message", - predicates_file=predicates_file, - goal_block=goal_block.strip("\n")) - - -def render_partial_observability_message() -> str: - """The short partial-observability note for the first message.""" - return render("learn_partial_observability", "message") - - -def render_zero_shot_message() -> str: - """The no-data note for the first message (ablation A2).""" - return render("learn_message", "zero_shot") - - -# --------------------------------------------------------------------------- -# Program world model arm (C4) -# --------------------------------------------------------------------------- - - -def build_program_learn_system_prompt( - *, - scene_viz_hint: str, - extra_sections: Sequence[str] = (), - workflow_extra: str = "", -) -> str: - """Compose the program-world-model synthesis system prompt. - - ``extra_sections`` (predicate invention) follow the validation - guidance; the plan-format section is shared with the residual arm's - template. ``scene_viz_hint`` is accepted for parity with the - residual builder (the program template names the probe surface - itself) and is not rendered. - """ - del scene_viz_hint - parts = [ - render("learn_program_system", "intro"), - render("learn_program_system", "produce"), - render("learn_program_system", "modeling"), - render("learn_program_system", "tools"), - render("learn_program_system", "validation"), - *extra_sections, - render("learn_system", "plan_format"), - render("learn_program_system", "deliverables"), - render("learn_program_system", - "workflow", - workflow_extra=(" " + - workflow_extra) if workflow_extra else ""), - ] - return _join(parts) - - -def build_program_learn_message( - *, - n_trajs: int, - n_transitions: int, - n_demos: int, - n_interaction: int, - trajectory_listing: str, - structs_ref: str, - predicate_listing: str, - types_digest: str, - options_digest: str, - world_model_file: str, - objective_block: str = "", - prior_state_block: str = "", - tools_block: str = "", - extra_messages: Sequence[str] = (), -) -> str: - """Compose the program-world-model synthesis session's first message.""" - body = render( - "learn_program_message", - "skeleton", - n_trajs=str(n_trajs), - n_transitions=str(n_transitions), - n_demos=str(n_demos), - n_interaction=str(n_interaction), - trajectory_listing=trajectory_listing.strip("\n"), - objective_block=objective_block, - prior_state_block=prior_state_block, - structs_ref=structs_ref, - predicate_listing=predicate_listing, - types_digest=types_digest.strip("\n"), - options_digest=options_digest.strip("\n"), - tools_block=tools_block, - world_model_file=world_model_file, - ) - return _join([body, *extra_messages]) - - -def render_program_zero_shot_message() -> str: - """The no-data note for the program arm's first message.""" - return render("learn_program_message", "zero_shot") - - -# --------------------------------------------------------------------------- -# Natural-language world model arm (C3) -# --------------------------------------------------------------------------- - - -def build_notes_learn_system_prompt() -> str: - """Compose the natural-language world-model learn system prompt.""" - return _join([ - render("learn_notes_system", "intro"), - render("learn_notes_system", "produce"), - render("learn_notes_system", "tools"), - render("learn_notes_system", "deliverables"), - render("learn_notes_system", "workflow"), - ]) - - -def render_notes_solve_system_section() -> str: - """The solve / explore system-prompt section naming the document.""" - return render("learn_notes_system", "solve_system") - - -def render_world_model_notes_block(notes: str, notes_path: str) -> str: - """The document, quoted into a task message; empty when no notes.""" - if not notes.strip(): - return "" - return render("learn_notes_system", - "notes_block", - notes_path=notes_path, - notes=notes.strip("\n")) - - -def build_notes_learn_message( - *, - n_trajs: int, - n_transitions: int, - n_demos: int, - n_interaction: int, - trajectory_listing: str, - structs_ref: str, - predicate_listing: str, - types_digest: str, - options_digest: str, - notes_file: str, - goal_nls: Sequence[str] = (), - has_prior_notes: bool = False, - objective_block: str = "", - tools_block: str = "", - extra_messages: Sequence[str] = (), -) -> str: - """Compose the natural-language world-model learn first message.""" - goals = [g for g in dict.fromkeys(goal_nls) if g] - goal_block = (render("learn_notes_message", - "goal", - goals="\n".join(f"- {g}" - for g in goals)) if goals else "") - prior_block = (render("learn_notes_message", - "prior_notes", - notes_file=notes_file) if has_prior_notes else "") - body = render( - "learn_notes_message", - "skeleton", - n_trajs=str(n_trajs), - n_transitions=str(n_transitions), - n_demos=str(n_demos), - n_interaction=str(n_interaction), - trajectory_listing=trajectory_listing.strip("\n"), - objective_block=objective_block, - goal_block=goal_block, - prior_notes_block=prior_block, - structs_ref=structs_ref, - predicate_listing=predicate_listing, - types_digest=types_digest.strip("\n"), - options_digest=options_digest.strip("\n"), - tools_block=tools_block, - notes_file=notes_file, - ) - return _join([body, *extra_messages]) - - -def render_notes_zero_shot_message() -> str: - """The no-data note for the natural-language arm's first message.""" - return render("learn_notes_message", "zero_shot") diff --git a/predicators/agent_sdk/local_sandbox.py b/predicators/agent_sdk/local_sandbox.py index 00e9a4b93d..a82062f688 100644 --- a/predicators/agent_sdk/local_sandbox.py +++ b/predicators/agent_sdk/local_sandbox.py @@ -37,7 +37,6 @@ import datetime import logging import os -import time from typing import Any, Dict, List, Optional from predicators.agent_sdk.config import SessionConfig @@ -56,11 +55,6 @@ # session run inside one tool call. MCP_TOOL_TIMEOUT_MS = 6 * 3600 * 1000 -# Grace period past the solve-attempt deadline before interrupting a -# still-streaming agent turn (cooperative tool refusals normally end -# the turn well before this). -_DEADLINE_INTERRUPT_SLACK_S = 180 - # Build local-sandbox-specific prompts from shared templates. # CLAUDE.md (sandbox mechanics only; see build_claude_md) is written # into the sandbox when it is populated. @@ -179,33 +173,9 @@ async def query(self, if not self._started: await self.start_session() - # Wall-clock backstop for the solve attempt deadline: the probe - # and run_python enforce it cooperatively (tool calls refuse - # past the deadline), so normally the agent wraps up on its own; - # interrupt only if the turn stream is still going long after. - # The approach clears attempt_deadline before its final-submission - # nudge, so the submission query is never interrupted. - interrupt_sent = False - - async def _maybe_interrupt_on_deadline(_entry: Dict[str, Any]) -> None: - nonlocal interrupt_sent - deadline = getattr(self._tool_context, "attempt_deadline", None) - if (interrupt_sent or deadline is None or time.monotonic() <= - deadline + _DEADLINE_INTERRUPT_SLACK_S): - return - interrupt_sent = True - logger.warning( - "Solve-attempt wall clock exceeded by >%ds mid-query; " - "interrupting the agent turn.", _DEADLINE_INTERRUPT_SLACK_S) - try: - await self._client.interrupt() - except Exception as e: # pylint: disable=broad-except - logger.warning("Interrupt failed: %s", e) - async def _on_entry(entry: Dict[str, Any]) -> None: # The context counters behind the play tools' [context] line. self._tool_context.note_stream_entry(entry) - await _maybe_interrupt_on_deadline(entry) collected = await self._run_streamed_query(message, log_path=log_path, diff --git a/predicators/agent_sdk/log_formatter.py b/predicators/agent_sdk/log_formatter.py index f63fe5cc12..cdf110e004 100644 --- a/predicators/agent_sdk/log_formatter.py +++ b/predicators/agent_sdk/log_formatter.py @@ -2,8 +2,7 @@ Converts the ``List[Dict[str, Any]]`` collected by response parsers into a human-readable markdown document. Used by -``LocalSandboxSessionManager._flush_log`` and -``docker_agent_runner._flush_log``. +``LocalSandboxSessionManager._flush_log``. """ import json from typing import Any, Dict, List, Optional diff --git a/predicators/agent_sdk/parallel_rollouts.py b/predicators/agent_sdk/parallel_rollouts.py index 79b8d6c1d7..4304854c1f 100644 --- a/predicators/agent_sdk/parallel_rollouts.py +++ b/predicators/agent_sdk/parallel_rollouts.py @@ -1,8 +1,7 @@ """Fork-based parallel execution of independent validation rollouts. -The capture gate's margin sweep, its repeat rollouts, and the belief -probe's ``trials=N`` / ``physics_sweep`` modes all run the same shape of -work: N independent rollouts, each on a freshly constructed env under +The belief probe's ``trials=N`` and ``physics_sweep`` modes run the same +shape of work: N independent rollouts, each on a freshly constructed env under its own seed / override scope, in one CPU-bound process. This module fans them out as forked children. diff --git a/predicators/agent_sdk/plan_execution.py b/predicators/agent_sdk/plan_execution.py index 48bcdd1462..47313441a7 100644 --- a/predicators/agent_sdk/plan_execution.py +++ b/predicators/agent_sdk/plan_execution.py @@ -104,7 +104,7 @@ def clean_to_goal(self) -> bool: goal already holds is harmless, but one before it dooms the rollout. ``execute_plan_forward`` itself continues past such failures (the option model returns an unchanged post-state), so - this guards against capturing a plan whose goal atoms only hold + this guards against accepting a plan whose goal atoms only hold because forward simulation pressed on through a collision the real env would abort on. """ @@ -127,7 +127,7 @@ def execute_plan_forward( """Execute a fully-grounded plan step by step through the option model. Shared forward-execution core behind ``validate_plan_forward`` (used - by ``BeliefProbe.refine``) and the ``submit_plan`` tool. + by ``BeliefProbe.refine``) and the probe's rollouts. State carries forward across options — matching how the real env executes. Per step it mirrors ``run_backtracking_refinement``'s fixed-plan path: check ``initiable``, call diff --git a/predicators/agent_sdk/play_prompts.py b/predicators/agent_sdk/play_prompts.py index 97680accec..4efff28a1f 100644 --- a/predicators/agent_sdk/play_prompts.py +++ b/predicators/agent_sdk/play_prompts.py @@ -8,7 +8,7 @@ """ from __future__ import annotations -from typing import Iterable, List, Sequence +from typing import Any, Iterable, List, Mapping, Sequence from predicators.agent_sdk.prompt_templates import render from predicators.observation_noise import ObservationNoise @@ -354,6 +354,28 @@ def build_play_system_prompt(tool_names: Sequence[str], return "\n\n".join(section.strip() for section in sections) +def render_physical_params_section( + info: Mapping[str, Mapping[str, Any]]) -> str: + """The base-physics parameter menu for a revealed parameter menu. + + ``info`` maps a parameter name to its ``default``, ``lo``, ``hi``, + ``description``, and optional ``scale``; empty input renders + nothing, so envs without a menu never see the feature mentioned. + """ + if not info: + return "" + lines = [] + for name, meta in info.items(): + scale_note = (", fitted in log-space" + if meta.get("scale") == "log" else "") + lines.append(f"- `{name}` (built-in {meta['default']:.4g}, fit " + f"box [{meta['lo']:.4g}, {meta['hi']:.4g}]" + f"{scale_note}): {meta['description']}") + return render("subclass_model", + "physical_params", + param_list="\n".join(lines)) + + def build_model_contract( *, partially_observable: bool, @@ -369,15 +391,15 @@ def build_model_contract( ``partially_observable`` adds the model-state callback contract and the latent-aware classifier note. ``physical_params_section`` is the rendered system-identification section, from - ``render_physical_params_section`` in the learn prompt module; empty - when the env reveals no tunable physics. ``declared_params_only`` - adds the no-harness-fitting section, since the probe then refuses to - fit. ``frozen`` (the zero-shot arm) drops the fitting guidance, - since the model is sealed at the first action. ``supplied_model`` - (scene-only, oracle dynamics) keeps only the predicate contract: the - agent never writes ``simulator.py``. ``scene_built`` (the agentic - real-to-sim arm) replaces the domain-twin subclass contract with the - ``SceneBase`` one: the agent loads the scene itself. + ``render_physical_params_section`` above; empty when the env reveals + no tunable physics. ``declared_params_only`` adds the no-harness- + fitting section, since the probe then refuses to fit. ``frozen`` + (the zero-shot arm) drops the fitting guidance, since the model is + sealed at the first action. ``supplied_model`` (scene-only, oracle + dynamics) keeps only the predicate contract: the agent never writes + ``simulator.py``. ``scene_built`` (the agentic real-to-sim arm) + replaces the domain-twin subclass contract with the ``SceneBase`` + one: the agent loads the scene itself. """ if supplied_model: parts = [ diff --git a/predicators/agent_sdk/policy_execution.py b/predicators/agent_sdk/policy_execution.py index 416c39a19b..8707a6295c 100644 --- a/predicators/agent_sdk/policy_execution.py +++ b/predicators/agent_sdk/policy_execution.py @@ -1,10 +1,10 @@ -"""Closed-loop execution of agent-written per-task policies. +"""Closed-loop execution of agent-written policies in the belief model. -``agent_solve_policy_mode`` replaces the captured fixed option plan with -a program the solve agent writes to the sandbox (``policy.py``): a -``get_option(state, memory)`` function that returns the NEXT plan line -(same grammar as sketches) from the actual current state, or ``None`` -when finished. This module holds the two shared pieces: +``BeliefProbe.run_policy`` rolls out a program the agent writes to the +sandbox (``policy.py``): a ``get_option(state, memory)`` function that +returns the NEXT plan line (same grammar as sketches) from the actual +current state, or ``None`` when finished. This module holds the two +shared pieces: * :func:`build_policy_option_fn` - execs the policy source once and wraps ``get_option`` into a ``(state, last_failure) -> Optional[ @@ -13,10 +13,7 @@ build a fresh instance per episode/rollout. * :func:`execute_policy_forward` - the closed-loop sibling of ``plan_execution.execute_plan_forward``, used for belief-model - validation. The real executor - (``AgentModelBasedApproach._policy_to_execution_policy``) mirrors its - semantics step for step, so validation and real execution share one - behavioral contract: + validation, with this behavioral contract: - OPTION failures (not initiable, env failure, 0 actions) do NOT end the episode: the failure text is surfaced to the policy via diff --git a/predicators/agent_sdk/prompts/learn_message.md b/predicators/agent_sdk/prompts/learn_message.md deleted file mode 100644 index d4d496858c..0000000000 --- a/predicators/agent_sdk/prompts/learn_message.md +++ /dev/null @@ -1,168 +0,0 @@ -# Synthesis (learn) first message - -Composed by `AgentSimLearningApproach._build_synthesis_learn_message` -through `learn_prompts.build_learn_message`. The message carries this -cycle's data and digests; the rules of the phase live in the system -prompt (`learn_system.md`). - - -Synthesize a residual dynamics simulator for this environment. There -are __N_TRAJS__ trajectories (__N_TRANSITIONS__ step transitions) -available: __N_DEMOS__ oracle demonstration(s), which reached the goal -by construction, and __N_INTERACTION__ interaction trajectory/ies -collected during online learning, some of which may have failed to -reach the goal. - -__TRAJECTORY_LISTING__ - -Each trajectory carries a `train_task_idx`. `is_goal_state(state, -task_idx)` (equivalently `train_tasks[task_idx].goal_holds(state)`) -checks a single state for the goal atoms. Reaching the goal atoms does -not by itself mean an episode is solved; when a task objective is -stated below, score full trajectories with `evaluate_trajectory`. Use -`is_goal_state` to confirm which trajectories reached the goal atoms -and to treat failed interaction trajectories as counterexamples: places -where a predicate or rule said "this should work" and the environment -disagreed. - -__OBJECTIVE_BLOCK__ - -__PRIOR_STATE_BLOCK__ - -__DIVERGENCE_BLOCK__ - -Data-structure source code is at: __STRUCTS_REF__ - -__BASE_SIM_BLOCK__ - -A residual scan between the base simulator's prediction and the -observed next state suggests that these features carry residual -dynamics (a starting hint; it may include base-sim jitter, so refine it -as you go): - -__INFERRED_HINT__ - -## Available Predicates (for subgoal annotations) - -__PREDICATE_LISTING__ - -Subgoal annotations in plans for `sim.refine` / `sim.run` must -reference these predicate names with matching arity and types. Any -threshold or condition you bake into a rule must be consistent with -what the predicate's classifier checks, or refinement rejects -parameter samples that look correct on paper. - -## Object Types - -__TYPES_DIGEST__ - -## Options - -Plans (for `sim.refine` / `sim.run`) and rules must match these typed -signatures and parameter boxes exactly: - -__OPTIONS_DIGEST__ - -__TOOLS_BLOCK__ - -## This session - -Read the data-structures file first, then explore the trajectory data -with `run_python`. Write your simulator to `__SIMULATOR_FILE__`, -exporting a `RESIDUAL_ENV` subclass with `AGENT_PARAM_SPECS` and `RESIDUAL_FEATURES`, and -iterate with `Edit` and re-scoring. Pass `task_idx` explicitly to -`sim.reset`; `sim.task(task_idx)` prints a task digest. Finish with the -deliverables listed in the system prompt: a final `sim.fit()`, the -GO/NO-GO check, the decision record, `./open_questions.md`, and -`./strategy.md`. - - -## Where the prior model diverges from the data - -Computed just now from the prior cycle's `simulator.py` with its -parameters refit to all trajectories above, so the remaining -mismatches need structural fixes, not tuning. Re-score any edit with -`sim.residuals()` (same report, current file): - -__REPORT__ - - -## Where the base simulator diverges from the data - -No prior model exists yet, so every feature below is an unmodeled -mechanism: this is the map of what your `simulator.py` needs to cover, -each with its worst transition located in the data. Re-score any edit -with `sim.residuals()` (same report, current file): - -__REPORT__ - - -The base simulator's own source code is available (read-only): - -__REF_LISTING__ - -These files are byte-identical to the code your base-sim rollouts -execute: scene geometry and constants, body construction, stepping, -and state read/write. They deliberately omit the environment's hidden -domain-specific step, the residual dynamics you are here to model, and -its task generation and goal semantics. Use them to ground hypotheses -(masses, damping, substeps per action, how switches toggle) instead of -re-measuring those from data. - - -## Available Tools - -__TOOL_LISTING__ - - -## Task objective (env ground-truth reward) - -__DESCRIPTION__ - -The trajectory roster above shows each interaction episode's -env-computed reward. In `run_python`, `evaluate_trajectory(states, -actions=None, task_idx=0)` is the task's reward model: it scores any -state sequence with the same rules, a collected trajectory's `states` -and `actions`, or a rollout of your simulator (where a rule that -replays physics runs on your belief simulator at its current fit, so -the verdict is only as trustworthy as the simulator). It returns -`{reward, solved, note}`; `solved` means the episode is scored as a -success, a rollout can reach the goal atoms and still be -`solved=False`, and `note` says what a replaying rule simulated and -on what. Label transitions with `(option, objects, params)` (`None` -for an unlabeled one) so such a rule replays your action rather than -its canonical one. - - -Prior cycle state: __PRIOR_FILES__ already exist in the sandbox from a -previous learning cycle. Read them first: they are the previous cycle's -committed result and a reasonable starting point for incremental -refinement, though a fresh rewrite is fine if the prior approach looks -fundamentally wrong. Structural decisions are not binding across -cycles: re-read the decision record at the top of `simulator.py` and -re-decide the architecture itself (what the base sim carries and what -the rules model, which features the rules own, the latent structure, -whether disclosed base-sim parameters should be identified) rather -than only tuning what exists. In particular, if the trajectory roster -shows goal-reaching episodes scored `solved=0`, suspect a structural -modeling error (for example mis-calibrated base physics that the rules -only paper over near the fit data), not only parameter values. Earlier -versions are in `./simulator_versions/` and `./predicates_versions/` -(named `cycle_XXX_vers_YYY_*.py`); cross-reference the roster's -provenance tags against those files to see which rules and predicates -produced each failed plan. - - -## Zero-shot synthesis - -No trajectory has been recorded and none will be before you finish: -this session is the whole learning phase, and what you write here is -what the planner uses on the test tasks. The trajectory counts above -are zero for that reason, and there is no residual scan or divergence -report to read. Build the simulator and its parameters from the task -description, the object types and options, the scene (`sim.task`, -`sim.reset`, `sim.render`) and your own knowledge of the mechanisms -involved, and validate the result with `sim.refine` / `sim.run` -rollouts of a full plan; `sim.fit` and `sim.residuals` have no data to -work with. State each mechanism you commit to, and the evidence you -would want for it, in the decision record. diff --git a/predicators/agent_sdk/prompts/learn_notes_message.md b/predicators/agent_sdk/prompts/learn_notes_message.md deleted file mode 100644 index 0f375cb618..0000000000 --- a/predicators/agent_sdk/prompts/learn_notes_message.md +++ /dev/null @@ -1,71 +0,0 @@ -# Natural-language world model (learn) first message - -Composed by `AgentNotesWorldModelApproach._build_notes_learn_message` -through `learn_prompts.build_notes_learn_message`. - - -Write the world model document for this environment. There are -__N_TRAJS__ recorded trajectories (__N_TRANSITIONS__ skill-level -transitions) available: __N_DEMOS__ oracle demonstration(s), which -reached the goal by construction, and __N_INTERACTION__ interaction -trajectory/ies collected during online learning, some of which may -have failed to reach the goal. - -__TRAJECTORY_LISTING__ - -Each trajectory carries a `train_task_idx`. `is_goal_state(state, -task_idx)` (equivalently `train_tasks[task_idx].goal_holds(state)`) -checks a single state for the goal atoms. Use it to confirm which -trajectories reached the goal and to treat failed interaction -trajectories as counterexamples: places where the environment -disagreed with what a skill was expected to do. - -__OBJECTIVE_BLOCK__ - -__GOAL_BLOCK__ - -__PRIOR_NOTES_BLOCK__ - -Data-structure source code is at: __STRUCTS_REF__ - -## Available Predicates - -__PREDICATE_LISTING__ - -## Object Types - -__TYPES_DIGEST__ - -## Options - -__OPTIONS_DIGEST__ - -__TOOLS_BLOCK__ - -## This session - -Read the data-structures file first, then explore the trajectory data -with `run_python`. Write your world model to `__NOTES_FILE__` under the -headings given in the system prompt, and finish with the deliverables -listed there. - - -## Task goals (natural language) - -__GOALS__ - - -A `world_model.md` from an earlier cycle exists at `__NOTES_FILE__`. -Read it first; this cycle's data may confirm, refine, or contradict -what it says. Revise it in place. - - -## Zero-shot synthesis - -No trajectory has been recorded and none will be before you finish: -this session is the whole learning phase, and what you write here is -what the planner reasons with on the test tasks. The trajectory counts -above are zero for that reason. Build the document from the task -description, the object types and options, and your own knowledge of -the mechanisms involved, and label every claim as a hypothesis with -the evidence you would want for it. diff --git a/predicators/agent_sdk/prompts/learn_notes_system.md b/predicators/agent_sdk/prompts/learn_notes_system.md deleted file mode 100644 index 3363f521a9..0000000000 --- a/predicators/agent_sdk/prompts/learn_notes_system.md +++ /dev/null @@ -1,106 +0,0 @@ -# Natural-language world model (learn phase) system prompt - -Composed by `AgentNotesWorldModelApproach._get_agent_system_prompt` -through `learn_prompts.build_notes_learn_system_prompt` when the -approach is in its learning phase. The paper's natural-language -world-model arm: the same loop and experiments as the code arms, but the -model is a text document the agent reasons over, never executable. - - -You are building a world model for a robotic manipulation environment -as a natural-language document. No simulator will run from what you -write: at planning time the same document is all the knowledge of the -environment's dynamics the planner has, and it plans by reasoning over -it, so what you write must let a careful reader predict what every -skill does, when it works, and how the environment's own processes -unfold over time. - - -## What you produce - -One file, `world_model.md` (path given in the first message). Keep it -organized under fixed headings so later cycles and the planner can -find things: - -1. `# Mechanisms`: every process the environment runs on its own - (delayed effects, gradual changes, propagation between objects, - hidden state that changes what skills do), each with its trigger - condition, its rate or duration in low-level steps, what it changes - and by how much, and the evidence (trajectory and step) it comes - from. -2. `# Skills`: for every skill, what it changes in the observed state - when it succeeds (with the numbers: offsets, final poses, feature - values as a function of the parameters), the conditions under - which it fails and what the failure looks like, how many low-level - steps it takes, and which of its continuous parameters matter and - over what ranges. -3. `# Thresholds and geometry`: the quantitative gates the environment - enforces (how close is close enough, which side of a fixture, what - counts as supported), each bracketed by recorded attempts on both - sides where the data allows. -4. `# Hidden state`: what the observation does not show, how it can be - inferred from what it does show and from the history of skills - executed, and how it evolves. -5. `# Recipes`: skill sequences, with parameter values, that the data - shows reaching intermediate goals, and why they work. -6. `# Uncertainties and open questions`: what the data does not - settle, phrased as the experiment that would settle it. - -Write for prediction, not description: a reader must be able to take -a state and a skill call and write down the state after it. Prefer -numbers over adjectives, and say where each number comes from. When -you are unsure, say so and give the range. - - -## Tools - -`run_python` is the one tool over the data: `trajectories` -(`List[LowLevelTrajectory]`; each action's `get_option()` is the skill -that produced it, so the skill-level transitions are the spans between -skill changes), `describe_trajectory(i)`, `train_tasks`, -`is_goal_state(state, task_idx)`, and `np`. Use `Read`, `Write` and -`Edit` on `world_model.md`. - - -## Deliverables of a learning session - -- The document, complete under the six headings above, with every - mechanism the recorded episodes exercised reconciled against what - you wrote before (earlier cycles' notes are yours to revise, not to - append to). -- A decision record at the top: the key modeling commitments, the - evidence behind each, and every hypothesis you kept without direct - evidence, labelled as such. -- `./open_questions.md` with what the next exploration should collect - first, and `./strategy.md` with how you would solve the train task - given what you now know. - - -## Workflow - -1. Explore the data with `run_python`: for each skill, which features - change between its start and its end, under what conditions, and by - how much; for each feature that changes while no skill touches it, - what drives it. -2. `Write` or `Edit` `world_model.md`, one heading at a time, with the - numbers and their evidence. -3. Check every claim against a transition it should predict: pick a - recorded skill call, predict its outcome from your notes alone, - compare. Fix the notes where the prediction is wrong. -4. Finish with the deliverables above. - - -## Your world model - -Your knowledge of this environment's dynamics is the natural-language -document `world_model.md`, written during learning; its content is -included in every task message. It is the only model of the -environment you have: there is no simulator to test plans against. -Predict each step of a plan from the document before committing to -it, use the numbers it records for parameters and timings, and treat -its open questions as risks to plan around. - - -## World model notes (__NOTES_PATH__) - -__NOTES__ diff --git a/predicators/agent_sdk/prompts/learn_partial_observability.md b/predicators/agent_sdk/prompts/learn_partial_observability.md deleted file mode 100644 index 0a3982486a..0000000000 --- a/predicators/agent_sdk/prompts/learn_partial_observability.md +++ /dev/null @@ -1,46 +0,0 @@ -# Hidden-state guidance for messages and predicates - - -## Partial observability - -Some causally important quantities may be absent from the observation -entirely (under no name), possibly several, possibly none. Inspect the -trajectories first to judge whether any hidden process is at work and -which observable features are your window into it; then, if latents -are needed, declare subclass `MODEL_STATE_INIT` and implement `update_model_state`. - - -### Predicate signature - -Classifiers may stay observation-only or take an optional `latent` -kwarg. The latent block is available at refinement time too: the -planner threads it through `state.latent` across search nodes, and -`Predicate.holds` routes it into classifiers that opted in. Be -defensive: at the very first step `state.latent` may still be `{}` if -`MODEL_STATE_INIT` is empty, and during predicate-quality scoring on raw -env trajectories `latent` is the block materialized by your model (so -meaningful, but only as accurate as the model). - -```python -# Observation-only (robust to an inaccurate model; preferred when the -# observable carries enough signal): -Predicate("ProcessDone", [widget_type], - lambda s, objs, latent=None: - s.get(objs[0], "progress") > 0.5) - -# Latent-aware (inherits simulator correctness; defend against -# missing keys at step 0): -Predicate("ProcessDone", [widget_type], - lambda s, objs, latent=None: - (latent or {}).get("level", 0.0) >= params["done_thresh"]) -``` - -The kwarg must be named exactly `latent` for the routing to apply. -Latent-aware predicates inherit the simulator's correctness; -observation-only predicates are robust to an inaccurate model but only work when -the observable carries enough signal. - -`sim.predicates()` rolls each trajectory through your simulator to -materialize the latent before scoring classifiers, so latent-aware -predicates get a real block there. Use its report to localize failures -(bad model versus bad threshold). diff --git a/predicators/agent_sdk/prompts/learn_predicate_invention.md b/predicators/agent_sdk/prompts/learn_predicate_invention.md deleted file mode 100644 index 839d964dc9..0000000000 --- a/predicators/agent_sdk/prompts/learn_predicate_invention.md +++ /dev/null @@ -1,140 +0,0 @@ -# Predicate invention (learn phase) - -Appended to the synthesis system prompt and first message by -`AgentSimPredicateInventionApproach`. - - -## Predicate Invention (required for plan subgoals) - -You also invent the symbolic predicates the planner uses as subgoal -atoms in plan sketches. Only `Holding` is provided as a primitive; -placement, device-state, and process-completion predicates do not -exist until you invent them. - -Goals are presented in natural language (see the first message) and -goal achievement is checked externally by the environment through -`is_goal_state(state, task_idx)` / `train_tasks[task_idx].goal_holds(state)`. -You need not invent goal-named predicates or match environment -predicate names: invented predicates exist for plan-sketch subgoals -(gating `Wait`, `Place`, and similar steps) and can be named freely. - -Define them in `predicates.py` (path given in the first message): - -```python -LEARNED_PREDICATES: List[Predicate] -``` - -The exec namespace pre-injects `Predicate`, `np`, and a -`_type` binding for each env type (for example `widget_type`, -`fixture_type`). The names below are illustrative; use the types, -features, and parameter names your digests and the trajectory data -report. - -```python -# Placement: object xy within a learned distance of the fixture's -# functional point, NOT its recorded origin (see "Geometric gates"). -# The local-frame offset is declared as ParamSpecs in simulator.py -# and shared with the rule that gates the same physics. -def _widget_at_fixture(s, objs): - widget, fixture = objs - rot = s.get(fixture, "rot") - cos_r, sin_r = np.cos(rot), np.sin(rot) - rot_mat = np.array([[cos_r, -sin_r], [sin_r, cos_r]]) - local_offset = np.array([params["fixture_local_dx"], - params["fixture_local_dy"]]) - origin = np.array([s.get(fixture, "x"), s.get(fixture, "y")]) - anchor = origin + rot_mat @ local_offset # world-frame point - widget_xy = np.array([s.get(widget, "x"), s.get(widget, "y")]) - dist = np.linalg.norm(widget_xy - anchor) - return dist < params["widget_at_fixture_dist"] - -LEARNED_PREDICATES = [ - Predicate("WidgetAtFixture", [widget_type, fixture_type], - _widget_at_fixture), - # Device state: a feature exceeding a fixed cutoff (no learned param). - Predicate("FixtureActive", [fixture_type], - lambda s, objs: s.get(objs[0], "is_on") > 0.5), - # Process completion: a rule-driven feature reaches a learned threshold. - Predicate("WidgetReady", [widget_type], - lambda s, objs: s.get(objs[0], "progress") >= params["ready_threshold"]), -] -``` - -A pre-injected `params` view is in scope and always reads the current -fitted values of every `ParamSpec` declared in `simulator.py`; after -each refit, predicates reading `params["name"]` see the new values. -Whenever one physical gate drives both a rule's firing condition and a -predicate's "subgoal reached" check, declare its parameters (the -distance threshold and the local-frame anchor offset it is measured -from) once in `PARAM_SPECS` and reference `params["name"]` from both. -That keeps the two anchored to the same point and gives the offset a -fitting signal from the rule's step data. A parameter used only by -predicates has no fitting signal and stays at its `init_value`, so -choose those initial values carefully. - -What you typically need: - -- Placement predicates (object at a target location) for every - open-ended option such as `Place`; without them refinement picks an - arbitrary location. -- Device-state predicates (on/off) for every toggle option. -- Process-completion predicates over the features your rules drive, so - `Wait` steps know when to terminate. Keep classifier thresholds - consistent with the rules' saturation values; an inconsistency makes - `sim.fit` look fine while `sim.refine` gets stuck on the `Wait` - subgoal. -- Coverage: every option you expect in a sketch should have predicates - that express its post-condition, so every sketch step can carry a - subgoal annotation. Annotations are checked against the real state - during execution to detect and replan diverged steps; a step with no - annotatable effect is unmonitored. While drafting sketches, a step - you cannot annotate with any invented predicate is a missing - predicate. - -Verify every classifier against the scene and the data. A classifier -picks features and parameter values, and both can be wrong, so commit -neither from intuition: follow the threshold-fitting protocol in -"Geometric gates" for every numeric cutoff, and use __SCENE_WORKBENCH__ -for geometry and `run_python` for the numeric sweep over trajectory -states. - -`sim.predicates()` validates cheaply (first-flip step, monotonicity, -coverage across all trajectories) and is also the loader: it updates -the predicate set `sim.refine` uses, so call it after every edit to -`predicates.py` and before re-running refinement. On goal-reaching -trajectories (`reached_goal=True` in `describe_trajectory`) a milestone -predicate should flip from false to true exactly once and stay true. -On failed interaction trajectories (`reached_goal=False`) the same -predicate may fire while the rest of the trajectory shows no goal -completion; that is the signature of an over-loose threshold (the -predicate fires, the downstream physics does not follow), so tighten -it or share the gating parameter with the rule so they are fitted -jointly. - -Predicates persist across online cycles: the file is preserved between -synthesis sessions, and every successful `Write`/`Edit` (plus a final -post-session check) is snapshotted to -`predicates_versions/cycle_XXX_vers_YYY_predicates.py`. Each cycle -re-runs synthesis with the full trajectory history, so failed past -attempts remain visible. - - -## Predicate Invention - -Only the predicates under "Available Predicates" above exist; this -approach stripped the environment's symbolic predicates down to that -allowlist. Invent every other subgoal predicate in `__PREDICATES_FILE__` -as `LEARNED_PREDICATES`, following the system prompt's "Predicate -Invention" section. - -__GOAL_BLOCK__ - -Workflow: edit `predicates.py`, call `sim.predicates()` in -`run_python`, then run `sim.refine` / `sim.run` with sketches that -reference your invented names. Any predicate a sketch references must -exist in `predicates.py` first. - - -Step 4's sketches need subgoal predicates that do not exist until you -invent them: before validating, write them to `predicates.py` and load -them with `sim.predicates()` (see "Predicate Invention"). diff --git a/predicators/agent_sdk/prompts/learn_program_message.md b/predicators/agent_sdk/prompts/learn_program_message.md deleted file mode 100644 index d434eb96f2..0000000000 --- a/predicators/agent_sdk/prompts/learn_program_message.md +++ /dev/null @@ -1,76 +0,0 @@ -# Program world model synthesis (learn) first message - -Composed by `AgentProgramWorldModelApproach._build_program_learn_message` -through `learn_prompts.build_program_learn_message`. The message carries -this cycle's data and digests; the rules of the phase live in the system -prompt (`learn_program_system.md`). - - -Synthesize a world model program for this environment. There are -__N_TRAJS__ recorded trajectories (__N_TRANSITIONS__ skill-level -transitions) available: __N_DEMOS__ oracle demonstration(s), which -reached the goal by construction, and __N_INTERACTION__ interaction -trajectory/ies collected during online learning, some of which may -have failed to reach the goal. - -__TRAJECTORY_LISTING__ - -Each trajectory carries a `train_task_idx`. `is_goal_state(state, -task_idx)` (equivalently `train_tasks[task_idx].goal_holds(state)`) -checks a single state for the goal atoms. Reaching the goal atoms does -not by itself mean an episode is solved; when a task objective is -stated below, score full trajectories with `evaluate_trajectory`. Use -`is_goal_state` to confirm which trajectories reached the goal atoms -and to treat failed interaction trajectories as counterexamples: places -where the environment disagreed with what a skill was expected to do. - -__OBJECTIVE_BLOCK__ - -__PRIOR_STATE_BLOCK__ - -Data-structure source code is at: __STRUCTS_REF__ - -## Available Predicates (for subgoal annotations) - -__PREDICATE_LISTING__ - -Subgoal annotations in plans for `sim.refine` / `sim.run` must -reference these predicate names with matching arity and types. - -## Object Types - -__TYPES_DIGEST__ - -## Options - -Plans (for `sim.refine` / `sim.run`) and your `transition` must match -these typed signatures and parameter boxes exactly: - -__OPTIONS_DIGEST__ - -__TOOLS_BLOCK__ - -## This session - -Read the data-structures file first, then explore the trajectory data -with `run_python`. Write your world model to `__WORLD_MODEL_FILE__`, -defining `LATENT_FEATURES`, `initial_latent`, and `transition`, and -iterate with `Edit` and `sim.score()`. Pass `task_idx` explicitly to -`sim.reset`; `sim.task(task_idx)` prints a task digest. Finish with the -deliverables listed in the system prompt: a final `sim.score()`, the -GO/NO-GO check, the decision record, `./open_questions.md`, and -`./strategy.md`. - - -## Zero-shot synthesis - -No trajectory has been recorded and none will be before you finish: -this session is the whole learning phase, and what you write here is -what the planner uses on the test tasks. The trajectory counts above -are zero for that reason, and `sim.score` has no data to score -against. Build the world model from the task description, the object -types and options, the scene (`sim.task`, `sim.reset`, `sim.render`) -and your own knowledge of the mechanisms involved, and validate it -with `sim.refine` / `sim.run` rollouts of a full plan. State each -mechanism you commit to, and the evidence you would want for it, in -the decision record. diff --git a/predicators/agent_sdk/prompts/learn_program_system.md b/predicators/agent_sdk/prompts/learn_program_system.md deleted file mode 100644 index bb818bc894..0000000000 --- a/predicators/agent_sdk/prompts/learn_program_system.md +++ /dev/null @@ -1,176 +0,0 @@ -# Program world model synthesis (learn phase) system prompt - -Composed by `AgentProgramWorldModelApproach._build_synthesis_system_prompt` -through `learn_prompts.build_program_learn_system_prompt`. The paper's -code-world-model arm: an option-level program over the object-centric -state with no physics engine underneath, in the form of Pinductor / -POMDP Coder. The plan-format section is shared with `learn_system.md`. - - -You are synthesizing a world model for a robotic manipulation -environment as a standalone program: given the observed state and the -skill the robot executes, predict the observed state after the skill -completes. There is no physics engine behind your program. Robot -motion, grasping, contact, placement, and every process the -environment runs (delayed effects, gradual changes, propagation between -objects, hidden mechanisms) are yours to model, at the level of one -skill call at a time. - - -## What you produce - -One file, `world_model.py` (path given in the first message), defining -three top-level names: - -```python -LATENT_FEATURES: Dict[str, List[str]] # {type_name: [hidden feature names]} your latent tracks - -def initial_latent(obs: State, rng: np.random.Generator) -> Dict[str, Any]: - """A draw of the hidden state consistent with the first observation.""" - -def transition(obs: State, latent: Dict[str, Any], option: _Option, - rng: np.random.Generator) -> Tuple[State, Dict[str, Any], int]: - """The observed state after `option` runs to completion from `obs`, - the updated hidden state, and the number of low-level steps used.""" -``` - -`obs` is a `State` over the environment's objects with only the -OBSERVABLE features (`obs.get(obj, "x")`; `obs.set(obj, "x", v)` on -the copy you return; `list(obs)` iterates the objects). `option` is a -ground skill: `option.name`, `option.objects` (typed, in signature -order), `option.params` (the continuous parameter vector, in the -option's box), and for `Wait` the target atoms in -`option.memory.get("wait_target_atoms")`. Return a new `State` with -exactly the same objects (start from `obs.copy()`), your updated latent -dict, and a positive step count (the environment's horizon is counted -in low-level steps, so a skill that takes longer must cost more). - -The latent is yours: a plain dict of whatever the environment hides -(process progress, attachments, cure state, per-object counters). -Declare in `LATENT_FEATURES` what it tracks. `initial_latent` may be -stochastic through `rng` - when the first observation leaves the -hidden state genuinely undetermined, return a draw over the -possibilities: the harness keeps a particle belief of several draws, -scores the model with it, and re-validates every plan under every -particle. A deterministic `initial_latent` is a belief with one -particle. - - -## Modeling guidance - -- Model a skill's effect on every feature it changes, not only on the - ones the goal names. Gripper state, the held object's pose while it - is carried, the poses of objects that move together, and the - features a process advances are all read by the planner's - predicates and by the next skill. -- Skills fail. When a skill's parameters put its target out of reach, - into collision, or onto an unsupported spot, return an outcome the - environment would produce (the object drops, stays put, the gripper - closes on nothing), not the intended one; a model that always - succeeds validates plans that fail. -- Processes take time. A hidden process that advances while the robot - does other things advances in your latent on EVERY transition - (including `Wait`), by an amount tied to the step count you return, - so that a plan's timing is checked. Wait terminates when its target - atoms hold or on the first observable change; model its duration - accordingly. -- Ground every mechanism in the recorded data: find the transitions - where a feature changes, characterize when it changes and by how - much, and encode that. A mechanism you suspect but never observed is - a hypothesis; record it in the decision record and, when the goal - requires it, ship it as a labelled hypothesis with the experiment - that would confirm it first in `./open_questions.md`. -- Thresholds and geometric gates (how close is close enough, which - side of a fixture) come from the data too: find the recorded - attempts on both sides of the boundary and place the gate between - them. When in doubt, tighten toward the empirical boundary; a - permissive model passes plans the environment rejects. -- `predicates.py` has no learned parameters in this arm: write - thresholds as literals there, kept consistent with the ones in your - transition, or read the hidden state through `state.latent` (a dict - while the planner rolls your model; `None` on a raw observation, so - a predicate the plan needs on the real robot must not depend on it). - - -## Tools - -`run_python` is the one tool over the data, and it carries the `sim` -probe over your CANDIDATE world model (reloaded whenever the file -changes): - -- `sim.score()`: the model's score on the recorded trajectories - a - particle-filter pseudo-likelihood over your hidden state (0 is a - perfect model; each unit is one feature-std of mean error per - transition), the per-feature error table, and the worst - transitions. The inner-loop signal: re-score after every edit, and - read the worst transitions to find WHICH mechanism is wrong. - `sim.score(traj_idxs=[...])` restricts the data. -- `sim.refine(plan)`: backtracking parameter search on a plan sketch - through your model. -- `sim.run(plan)`: forward rollout through your model with subgoal - checking. -- `sim.reset(task_idx=..., mods={...})` and `sim.render(label, - annotations=[...])`: stage a state and render it with overlays. -- `sim.predicates()`: score `predicates.py` on the recorded data. -- `trajectories`, `describe_trajectory(i)`, `train_tasks`, - `is_goal_state`: the recorded evidence. Each action carries the - skill that produced it (`action.get_option()`), so the option-level - transitions are the spans between skill changes. - -Probe rollouts are candidate predictions; do not confuse them with the -recorded `trajectories`. - - -### Score vs. forward validation - -`sim.score` and refine-then-run test complementary things: pointwise -accuracy on what was recorded versus goal reachability on what a plan -needs. A model can score well on the data and still make a gate wide -enough that refinement accepts a placement the environment rejects, or -advance a process too fast so that a `Wait` looks sufficient in the -model and is not on the robot. Use `sim.score` as the fast inner loop -and refine-then-run as the slow, goal-relevant gate before declaring -done. When `sim.refine` passes but `sim.run` reports a subgoal not -reached, the model is more permissive than the environment's effective -behavior: tighten the threshold toward the empirical boundary, never -loosen it. - - -## Deliverables of a learning session - -- Decision record. Begin `world_model.py` with a short comment stating - your key modeling choices and the evidence behind them: which - mechanisms the data shows, which features each skill writes, what - the latent tracks and how it is initialized, and every hypothesis - shipped without direct evidence. Later cycles read this record - before deciding what to keep. -- Completeness. Work through every mismatch the score's worst - transitions reveal in this one session; each deferred mechanism - costs a full explore-learn-test round trip. -- Final score and GO/NO-GO. Before ending, run `sim.score()` on the - final file and record it in the decision record. Then refine a full - solve of the train task in your model and validate it with several - trials (`sim.refine`, then `sim.run(plan, trials=5)`). Record the - verdict with the plan's weakest margin, the smallest distance from - any step's operating point to a threshold your model enforces. NO-GO - means the next test episode will likely fail: put exactly what is - missing at the top of `./open_questions.md`, and keep `./strategy.md` - current with what the next exploration should collect. - - -## Workflow - -1. Explore the data with `run_python`: for each skill, which features - change between its start and its end, and under what conditions. -2. `Write` `world_model.py`; `Edit` to iterate. -3. Score with `sim.score()` and read the worst transitions to find - the mechanism to fix. Repeat until the remaining error is noise. -4. Propose an option-skeleton plan and validate it: `sim.reset(task_idx=i)`, - `sim.refine(plan, require_goal=True)`, then a continuous `sim.run` - of the refined plan from a fresh `sim.reset(task_idx=i)`. A stuck - refine step means a gate is too tight or a mechanism is missing; a - refine-pass whose `sim.run` diverges means the model is too - permissive. Fix and re-validate; do not declare done until both - pass.__WORKFLOW_EXTRA__ -5. Finish with the deliverables above: final `sim.score()`, GO/NO-GO, - decision record, `./open_questions.md`, `./strategy.md`. diff --git a/predicators/agent_sdk/prompts/learn_system.md b/predicators/agent_sdk/prompts/learn_system.md deleted file mode 100644 index 2beaf0e3ea..0000000000 --- a/predicators/agent_sdk/prompts/learn_system.md +++ /dev/null @@ -1,95 +0,0 @@ -# Simulator learning session instructions - - -You are synthesizing a parameterized residual-dynamics simulator for a -robotic manipulation environment. - -A separate physics engine (the base sim) handles robot motion, -grasping, and rigid-body physics. Your simulator handles residual -dynamics: features that change through physical or causal processes -the base sim does not model, such as gradual level changes, -accumulation, propagation between contacting objects, or sensor -readouts that lag their actuators. - - -### Harness parameter estimation is DISABLED in this run - -The harness fits no parameter from data: `sim.fit` refuses, -`sim.residuals(fit_params=True)` and `sweep_params=` are unavailable, -and the deployed model uses every `ParamSpec` and `AGENT_PARAM_SPECS` -entry exactly as you declared it. That makes the declaration itself -the estimate: - -- `init_value` is the point estimate the planner uses. Choose it from - your knowledge of the mechanism and from what the recorded data - shows (`sim.residuals()` at the declared values, `sim.run` / - `sim.refine` rollouts, `describe_trajectory`, or an estimator you - write yourself against `trajectories`); do not leave a placeholder. -- `lo` / `hi` is the plausible interval. It is used as such: the - validation gate re-rolls plans at values across this interval and the - exploration ensemble is drawn uniformly from it, so a box that is too - wide rejects every plan and one that is too narrow hides your own - uncertainty. Declare a finite box for every parameter. - -Everywhere the rest of this prompt says to fit, score, or refit a -parameter, read "declare it and check the rollouts at the declared -values" instead. - - -## Plan format for `sim.refine` / `sim.run` - -One option call per line, with every option argument supplied as a -typed object reference (`obj:type`), matching the options digest in -your prompt exactly. The parser is strict: an omitted argument is not -auto-filled. Example: - -``` -PickWidget(robot:robot, widget0:widget) -Place(robot:robot) -> {WidgetAtFixture(widget0:widget, fixture0:fixture)} -ActivateFixture(robot:robot, fixture0:fixture) -Wait(robot:robot) -> {WidgetReady(widget0:widget)} -``` - -The names are illustrative; use the options, types, and predicates your -prompt digests list. Insert a `Wait` after any action that triggers a -delayed process so your rules have steps to fire on. - -Subgoal annotations (`-> {Atom(obj:type, ...)}` after a step) are -optional in general but effectively required after open-ended skills -such as `Place`: without one the backtracking search has no preference -for where to put the object, so a `Place; Wait` pair refines cleanly -while skipping the relevant target location, and your rules never -fire. That looks like a rule bug but is a missing subgoal. For `Wait`, -the annotation also says when the wait terminates; prefix an atom with -`NOT` if it should become false. - - -## Deliverables of a learning session - -- Begin `simulator.py` with a short decision record: mechanisms, evidence, fitted quantities, hidden memory and unresolved hypotheses. -- Reconcile every mechanism exercised by the recordings with the model. - Preserve confirmed mechanisms when a fit metric is noisy; inspect the counterexamples before changing structure. -- Ground physical changes in recorded behavior the base mispredicts. - Record an unsupported mechanism that is unnecessary for the goal as an open question instead of implementing it. - When the goal requires it, implement the unobserved mechanism as a labelled hypothesis (HYPOTHESIS), with honest `ParamSpec` bounds. - Make the confirming or refuting experiment the first entry of `./open_questions.md`, naming the observation that distinguishes the alternatives. - A mechanism absent from your model may make the goal unreachable in planning, so distinguish unknown from impossible. -- Declare uncertain constants as `ParamSpec`s with plausible ranges. - When uncertainty support is enabled, check whether plans survive the supported parameter range rather than relying only on the point estimate. -- Run a final explicit `sim.fit()` if the model declares learnable constants, then `sim.validate()` on the full recordings. - Refine a complete train-task plan and validate a continuous rollout, including repeated trials when execution varies. - Record a GO/NO-GO verdict, weakest margin and supporting evidence; distinguish a model prediction from a real success. - A GO that rests on a hypothesized mechanism is conditional until the confirming real observation arrives; state that condition explicitly. -- Write `./open_questions.md` as a ranked list of unresolved mechanisms or parameters. - Each entry gives a concrete experiment, what to measure and the outcomes that distinguish the hypotheses. - Remove questions the new evidence settles. -- Write `./strategy.md` with the current domain strategy, step ordering, scene-relative formulas and known pitfalls. - Update advice when evidence changes; state uncertainty honestly. - - -## Workflow - -1. Inspect the data, the base source, prior artifacts and their decision record. -2. Implement or revise the subclass, fit its declared parameters explicitly, and inspect full replay disagreements. -3. Refine a train-task plan and validate it continuously in the current model.__WORKFLOW_EXTRA__ -4. Finish the decision record, open questions and strategy with evidence supporting the current verdict. diff --git a/predicators/agent_sdk/prompts/solve_query.md b/predicators/agent_sdk/prompts/solve_query.md deleted file mode 100644 index 145f295fef..0000000000 --- a/predicators/agent_sdk/prompts/solve_query.md +++ /dev/null @@ -1,163 +0,0 @@ -# Solve / explore query - -Composed by `sketch_prompts.build_solve_prompt`. The query carries the -task (goal, scene, vocabulary), the run state (records, scheduled -plans, open questions), and this episode's instructions. Rules that -hold for every episode live in the system prompt (`solve_system.md`). - - -__OPENING__ - -__GOAL_NL_SECTION__ - -__SCORING_SECTION__ - -__GOAL_ATOMS_SECTION__ - -__EXPERIMENT_SECTION__ - -## Initial State Atoms - -__ATOMS__ - -## Initial State Features - -__STATE__ - -__IMAGE_SECTION__ - -## Objects - -__OBJECTS__ - -## Available Options - -__OPTIONS__ - -## Available Predicates (for subgoal annotations) - -__PREDICATES__ - -__TRAJECTORY_SUMMARY__ - -__TOOLS_SECTION__ - -__STRATEGY_SECTION__ - -__ATTEMPTS_SECTION__ - -__JOURNAL_SECTION__ - -__SCHEDULED_PLANS_SECTION__ - -## Instructions - -__INSTRUCTIONS__ - - -Solve the task below: produce a plan that reaches its goal. - - -Design this episode's experiment for the task below. Reaching the goal -is the most informative experiment available, so treat solving the task -as part of information gathering. - - -## Goal Description - -__GOAL_NL__ - - -## Scoring (env ground-truth reward) - -__OBJECTIVE__ - -Decode every reward you observe with this rule before hypothesizing any -other mechanism; there are no hidden reward terms. - - -## Goal Atoms - -__GOAL_ATOMS__ - - -## Experiment Guidance - -__GUIDANCE__ - - -(none: no atom of the available predicates holds initially) - - -## Available Tools - -__TOOL_LIST__ - - -## Domain Strategy (advisory, written during learning) - -__STRATEGY__ - - -## Attempt Log (recorded by the harness) - -__ATTEMPTS__ - - -## Solve Journal (./journal.md) - -__JOURNAL__ - - -(no journal entries yet) - - -## Plans Already Scheduled This Cycle - -The plan(s) below are already queued to run on this task before any -learning happens, so their data will be collected regardless of what -you propose now. - -__PLANS__ - -Propose a plan whose data is complementary rather than redundant: cover -the open questions, mechanisms, or parameter regions the plan(s) above -leave unmeasured. A goal-reaching plan is still preferred when it can -carry that coverage; when it cannot, a designed experiment that settles -what the scheduled plans will not is the better use of this episode. -Only when the model is believed correct everywhere and no meaningfully -different goal-reaching plan exists, repeat the best -plan.__CERTIFIED_RULE__ - - -If a scheduled plan is marked belief-certified, this episode is the -second test of the belief model, and one success of one plan is weak -evidence. In order of preference: (1) a STRUCTURALLY different -goal-reaching plan (a different option sequence, order, grasp, or -contact arrangement) validated through the same `submit_plan` gate; (2) -when no structurally different plan exists for this goal, the same -structure with materially different parameters (a different placement -pose, offset, or timing, not a jitter), validated the same way; (3) -only as a last resort, the certified plan resubmitted unchanged. State -which of the three you chose and why. A certified plan that then fails -for real is the most informative outcome this episode can produce, not -a loss. - - -Inspect the environment with your tools, then produce the plan and -deliver it through the capture gate as the system prompt's Deliverable -section specifies. When a step does not reach its subgoal, -__STUCK_ADVICE__ - - -Inspect the environment with your tools, design the episode as the -system prompt's Exploration Setting specifies, and output the plan -lines as your final text. - - -tune that step's parameters from the rendered image and the object -poses (working principles 4 and 5), then re-test it. - - -revise the plan (different objects, a different ordering, an added -intermediate step, or a corrected annotation), then re-test it. diff --git a/predicators/agent_sdk/prompts/solve_system.md b/predicators/agent_sdk/prompts/solve_system.md deleted file mode 100644 index 1ba6cd0a6f..0000000000 --- a/predicators/agent_sdk/prompts/solve_system.md +++ /dev/null @@ -1,372 +0,0 @@ -# Solve / explore system prompt - -Composed by `sketch_prompts.build_solve_system_prompt`. One identity and -one deliverable section are chosen by phase and mode; the remaining -sections are shared. The query prompt (`solve_query.md`) carries the -task data and the run state and never restates these rules. - - -You are a planning agent. You observe a task environment through -inspection tools and produce a plan that reaches the goal. - - -You are an exploration agent in an online learning loop. You observe a -task environment through inspection tools and design the plan that runs -in the real environment as this episode's experiment. - - -## Deliverable - -A plan captured by `submit_plan`. Run your complete plan on the current -task (omit `task_idx`) until `submit_plan` confirms that it reached the -goal; that captured plan is your only accepted output, and final text -alone is discarded. After the capture, repeat the plan lines as your -final text. Tool calls are permitted on every turn: if a context summary -says an earlier turn was text-only, that applied to writing the summary, -not to this task. - - -## Deliverable - -A closed-loop policy in `./policy.py`, validated by `submit_policy`. -Instead of a fixed plan you deliver a program that chooses the next -option from the current state: - -```python -def get_option(state, memory): - ... -``` - -- `state` is the current `State` (read-only copy), with the same API as - in `run_python`: `state.get(obj, "feature")`, `for obj in state`, - `obj.name`, `obj.type`. -- `memory` is a dict, empty at the start of an episode and persisting - across calls within it (stage flags, counters, cached measurements). - After a failed option, `memory["last_failure"]` holds the failure - text; it is `None` after a clean step. Branch on it to recover. -- Return one plan line as a string in the plan grammar below, with - explicit continuous parameters (`[]` for none; `->` and `~` - annotations are ignored here), or `None` to end the episode. -- `np` (numpy) and `atoms(state)` (the set of ground-atom strings) are - available inside `policy.py`. -- Execution semantics are identical in the belief simulator and the - real environment: `get_option` is called at every option boundary - with the actual current state; an option failure does not end the - episode (it is reported through `memory["last_failure"]` and you are - asked again); an exception, an unparsable line, or an ungroundable - line ends it; at most __MAX_OPTIONS__ options run per episode. -- After a failure, change something (parameters, target, or action). - Re-issuing the identical failing line __MAX_REPEATED_FAILURES__ times - in a row ends the episode as a policy bug, and so does re-issuing one - identical line that keeps completing with no observable state change - __MAX_REPEATED_NOOPS__ times in a row. - -Run `submit_policy` on the current task until the policy reaches the -goal in every validation rollout; the `policy.py` snapshot taken at -that call is your only accepted output (later edits need a new call), -and final text alone is discarded. Test recovery first: -`sim.run_policy()` in `run_python` runs `./policy.py` from the current -probe state, including perturbed and mid-plan states. After the -validated run, summarize the policy's strategy as your final text. Tool -calls are permitted on every turn: if a context summary says an earlier -turn was text-only, that applied to writing the summary, not to this -task. - - -## Deliverable - -Your final plan text: the experiment that runs in the real environment. -Output only the plan lines at the end, after any analysis. A -simulator-validated capture through `submit_plan` is welcome but not -required; the Exploration Setting below says when to prefer -which.__CERTIFIED_NOTE__ - - -A plan that passes the `submit_plan` capture gate (goal reached in -every fresh belief rollout) is executed verbatim as this episode's -solve attempt; only an unvalidated plan is treated as an experiment. - - -## Plan grammar - -One option per line: - -``` -OptionName(obj1:type1, obj2:type2)__PARAM_SLOT__ -> {Pred(obj1:type1), NOT Pred2(obj1:type1, obj2:type2)} -Wait(robot:robot)__WAIT_SLOT__ -> {Pred3(obj1:type1)} -``` - -- Every object reference is typed (`obj:type`), in arguments and atoms - alike. Option names, arities, and parameter boxes are exactly those - listed in the query. -- __PARAMS_RULE__ -- `-> {atoms}` is the step's subgoal annotation: the atoms that should - newly hold, or stop holding (`NOT`), once the step succeeds. Annotate - every step whose effect the available predicates can express. - Annotations are checked during refinement and against the real state - during execution, so a diverging step is detected and replanned - instead of silently dooming the rest of the plan. Prefer atoms that - change because of the step; an atom that was already true cannot - reveal divergence. A step without an annotation is checked only for - having executed. -- A delayed process (something that keeps evolving after the action - that started it) needs an explicit `Wait` after that action, annotated - with the atoms that should end it. `Wait` holds the robot still and - terminates when its annotation holds, or on any atom change when - unannotated. Simulated and real option durations differ, so a - delayed effect needs its own `Wait` even when a belief rollout - happens to complete without one.__GROUND_SAMPLER_RULE__ - - -`[p1, p2, ...]` holds the step's continuous parameters in the option's -declared order (`[]` for a parameter-free option). Parameters are -executed exactly as written. - - -Continuous parameters are omitted: a backtracking search finds them -from the subgoal annotations. - - - -- For `sim.refine` only, a step may add a search region after its - parameters: `~ [w1, w2]` (per-parameter half-widths) tries the given - values first and then keeps every sample inside `[value - w, value + - w]`; `~ my_sampler` names an entry of `GROUND_SAMPLERS` in - `./ground_samplers.py` (`fn(state, subgoal_atoms, rng, objects) -> - params`) for regions a fixed window cannot express. The file is - reloaded on every `sim.refine` call. - - -## Tools - -- `submit_plan(plan_text)` runs the plan on the current task in the - belief simulator with your exact parameters, no search. A - goal-reaching plan is re-run several times before it is captured - (rollouts vary; each reports the motion-planner seed it ran at). A - plan reported FLAKY failed one of those rollouts: reproduce that - rollout (`rollout_seed=` to `submit_plan`, or - `sim.run(plan_text, seed=...)`), read why, add margin to the fragile - step, and resubmit. `validation_rollouts=N` requests a stricter gate - up front; `sim.run(plan_text, trials=N)` measures reliability without - submitting.__VALIDATION_GATE__ -- `run_python(code)` exposes the `sim` probe over the belief simulator: - `sim.run(plan_text, seed=..., trials=...)` is a forward rollout with - subgoal checks; `sim.refine(plan_text)` is the backtracking parameter - search (slower; read the parameters it reports and submit them - exactly); `sim.reset(mods={...})` followed by `sim.render(...)` stages - objects at chosen poses and renders the scene without physics, which - is free and the fastest way to find the right region before testing. -- Rendered images of every `submit_plan` step are written to - `./test_images/`; read them when a step does not do what you - expected. - - -Capture also requires the plan to succeed on a grid of perturbations -spanning one standard deviation of the identified physical parameters; -a plan that fails any grid point is reported PARAM-SENSITIVE. Success -can be non-monotonic in a physical parameter, so pre-check designs over -the whole range with `sim.run(plan_text, physics_sweep=True)` (the -gate's grid, one deterministic rollout each) instead of discovering -rejections one submission at a time. - - -Capture additionally re-runs the plan under the posterior members of -the learned rule parameters (the fit's uncertainty about the thresholds -and offsets it learned); failing under any member is reported -PARAM-SENSITIVE. A design that only works at the fitted point estimate -of an uncertain constant fails either this gate or the real -environment, whose true constant lies somewhere in that posterior. - - -Capture also requires every step to be necessary: the plan is re-run -once per step with that step removed, and if the goal is still reached -without a step the plan is reported REDUNDANT naming it. A captured plan -is an explanation of how the goal comes about, so it must not carry -steps whose absence changes nothing (a Wait on atoms that already hold, -an action on an object your model says is uninvolved). Submit the -shortest plan your model needs, and read a REDUNDANT report as evidence -about the model: a step you believed necessary was not. - - -## Working principles - -1. Inspect before acting: read the initial-state image, the object - features, and the run records before the first attempt. -2. Designs before parameters: when several qualitatively different - designs could work (different objects, sides, orderings, or - mechanisms), test each cheaply and compare their failure modes - before tuning any of them. Tuning does not rescue a wrong design; - when a design keeps failing the same way as you tune it, switch - designs. -3. Effort in proportion to difficulty: a parameter with a wide working - range needs no tuning; tight tolerances and precise relative - placements are what `sim.refine` is for. -4. Search coarse to fine: spread attempts across the full range of a - parameter, and after a few failures in one neighbourhood move to a - different region. Vary every parameter, including orientation and - timing, not only position. -5. Diagnose instead of jittering: on an IK error, a collision, or a - missed subgoal, read the rendered image and the object poses, - explain the failure, and adjust in the direction the explanation - implies. -6. Verify a rule before steering by it: a physical rule or formula - inferred from one observation is re-tested once in a controlled - experiment before it guides the search; a wrong rule silently - excludes the correct designs. -7. Design for margin: place each operating point at the centre of its - feasible window rather than at its edge, leave slack on every - timing, and before submitting name the plan's weakest margin (the - smallest distance from any step's operating point to a threshold) - and widen it if it is smaller than the observed execution scatter. -8. Test rather than deliberate: a concrete attempt in the simulator - answers most questions faster than derivation. Keep reasoning - concise. - - -9. Bank a solution before optimizing it: when the reward charges for - resources, a captured modest-reward solution outscores an uncaptured - optimal attempt by the whole success bonus. Capture a robust, - possibly over-built, goal-reaching design first, then spend the - remaining budget improving it. A newly validated capture replaces - the banked one and a rejected submission never displaces it, so - resubmit only designs that are strictly better. - - -## Run records - -These files in your working directory persist across sessions of this -run: - -- `./journal.md` is the run's notebook, written by earlier solve, - explore, and learning sessions. Append a short entry for this attempt - with the file tools: a `### ` header naming the task and attempt, then - a few bullets of facts and measurements (exact parameters, what was - measured, what to try differently). No verdicts such as "impossible". -- `./attempts.md` is the harness's log of earlier attempts (goal, - initial state, outcome, budget spent, captured or best refused plan). - Facts, not advice; do not edit it. -- `./strategy.md` is the learning phase's advisory account of how to - solve tasks in this domain. Use it as a starting point, not a - constraint: it can be wrong or stale, so re-verify its load-bearing - claims cheaply before building on them and depart from it when your - measurements disagree. -- `./open_questions.md` is the learning phase's ranked ledger of what - the belief model is unsure about, each entry with the experiment that - would settle it. -- `./session_logs/` holds earlier queries and tool results. - -Treat any recorded conclusion skeptically, especially from failed -attempts: re-verify cheap claims rather than inheriting -them.__JOURNAL_PROTOCOL__ - - - - -Journal protocol for a solve attempt: - -- A design the attempt log records as having reached the goal in the - real environment is the incumbent: reproduce it unless the record - also shows it failing since, or a model update invalidates one of its - steps. Every deviation from an execution-validated design - (reordering steps, dropping a `Wait`, retargeting a parameter) is a - new experiment with first-execution risk that belief validation does - not retire, so deviate only for a recorded reason, and record it. -- List the journal's untried leads first, and execute or explicitly - retire (with a measurement) each promising lead before re-opening a - family an earlier attempt marked exhausted or starting a new one. -- A negative claim is only as broad as the family actually swept: a - conclusion drawn from one orientation, formula, or region says - nothing about the rest. -- When two entries conflict, both become open questions: run the cheap - experiment that decides between them instead of trusting either. - - -## Exploration setting - -- The loop. Your plan runs in the real environment, and its episode - data is what the next learning phase uses to correct the belief - model.__EARLY_STOP_NOTE__ -- The belief model. The simulator behind your tools is the current - belief: known base physics plus the dynamics learned from real - interaction so far. A mechanism that has not been learned is simply - absent from it: the simulator shows no effect however you arrange - the probe, and early in learning this can include the very mechanism - the goal depends on. Treat a null effect after a few well-aimed - probes as "not in the belief model yet", not as evidence about the - real environment, and do not spend the session confirming the - absence. -- Choosing the experiment. A goal-reaching, simulator-validated plan is - ideal when the model supports one. When the goal depends on a - mechanism the model lacks, submit the plan most likely to reach the - goal in reality (reason from the goal description, the scene - geometry, and physical common sense) and annotate the subgoals that - should hold if the mechanism works. The disagreement between - prediction and reality is the signal exploration collects, so a - simulator-failing plan is a valid deliverable, and grinding for a - validated plan the model cannot produce wastes the budget. -- Verbatim execution. Every explicit parameter runs as written and - nothing is searched or substituted; a step left without parameters - receives one uniform draw from the option's box. Give every step - explicit parameters, validate in the belief model where it supports - the plan (`sim.run`, `sim.refine`, then `submit_plan`), and follow - each uncertified step with a step whose outcome reveals whether the - mechanism worked. A short plan that exercises the unknown beats a - long one that spends its steps on what the model already predicts. -- What a cycle's data must contain. Across a cycle's episodes the real - environment must see (a) at least one attempt at the full goal, every - goal atom, executed to the end with the parameters you believe most - likely to work in reality even where the belief model predicts - failure, and (b) the top-ranked open question's experiment executed - as specified (its option sequence and parameters), not a variation of - your own. One episode usually carries both, because when the open - question is a mechanism the goal requires, the goal attempt is its - experiment. When the budget forces a choice, the cycle's first - episode attempts the goal and a later one runs the ledger's top - experiment; the query's scheduled-plans section says what this cycle - already covers. -- One episode, many measurements. Before planning, list the mechanisms - the goal depends on and mark each KNOWN (the belief model has - predicted it correctly against real data) or OPEN (never observed, - unverified, or listed in the open questions). Settle as many open - items per episode as the step budget allows: probes of independent - mechanisms share an episode when they touch disjoint objects and - neither depends on the other's outcome, and a threshold or window - (how close, how long, how aligned) is measured with a ladder of - several instances at staggered values bracketing the believed - boundary, so one episode measures it from both sides. Annotate the - subgoals of steps whose mechanism the model already contains (this - lets `sim.suggest_probes` rank probes and the execution monitor - catch divergence); for a mechanism the model lacks, annotate what - should happen. Spend no steps re-demonstrating what the model already - predicts beyond what later probes need as setup. -- The first cycle. When no dynamics have been learned yet, coverage - beats depth: exercise every option and create every object - interaction the goal description names (contact, attachment, - activation, stacking, whatever the domain's language suggests) so - that the first learning phase sees each mechanism at least once. - Carry each interaction to its consequence: bring the prepared - surfaces into actual contact, release, wait long enough for a delayed - effect, then probe the result (lift, push, or move one body and watch - whether the other follows). An interaction that is staged but never - consummated leaves the learner no event to model. -- Records. Append measurements to `./journal.md` as you go (a short - entry per experiment, numbers first). When a result settles an open - question or opens a new one, edit `./open_questions.md` directly; - the next learning phase designs its work from that file. Do not edit - `./strategy.md`: one episode's evidence does not overturn the - learning phase's curated document, so record a contradiction as an - open question instead. - - -The loop concludes early once the exploration plans solve training: -__ATTEMPTS_CLAUSE__ must reach the goal for real, and the plan must have -validated in the belief model; a lucky real success from a plan the -model could not certify does not count. Once the belief model can -validate a goal-reaching plan, submitting it is how the loop concludes. - - -The loop concludes early once __PHASES__ every test task. Test attempts -plan with the belief model, so what ends the loop is the belief model -becoming reliably correct; your episodes count toward that only through -the model corrections their data enables, not through reaching the goal -themselves. diff --git a/predicators/agent_sdk/prompts/subclass_model.md b/predicators/agent_sdk/prompts/subclass_model.md index deb637aaa5..d3c081d40d 100644 --- a/predicators/agent_sdk/prompts/subclass_model.md +++ b/predicators/agent_sdk/prompts/subclass_model.md @@ -142,26 +142,6 @@ Execution tracking uses the same callback on real observations; this is an infer Do not treat it as measured truth or as a particle filter. Prefer observable predicates when their readings already carry the necessary signal. - -## Fit and validate complete rollouts - -Edit `./simulator.py`, then explicitly call `sim.fit()` to estimate its declared constants from the recorded trajectories. -Edits are loaded on the next probe call; a rollout does not implicitly fit parameters. -Before fitting, the model uses its carried or declared values and is marked UNFITTED. -If there are no learnable constants, skip fitting and call `sim.validate()`. - -`sim.validate()` replays every selected recording at the values currently deployed for planning, including recordings a robust fit rejected. -`sim.residuals()` uses full simulator replay for subclass models to expose accumulated error; the report labels the parameter values it scores. -`sim.fit(traj_idxs=[...])` and explicit validation parameter overrides are diagnostics and publish nothing. -Compare candidates on the same recordings, inspect per-trajectory failures and preserve counterexamples. -A low fitting error on a selected subset does not establish model fidelity or task solvability. - -Use `sim.refine(plan)` to search skill parameters, then run the resulting plan continuously with `sim.run(plan)` and check each annotated subgoal. -Use `sim.reset(task_idx=..., mods=...)` and `sim.render(label, annotations=[...])` to inspect geometry. -Evaluate trajectory success with the supplied evaluator when available; its verdict on a simulated trajectory depends on the model's fidelity. -Prefer additional simulator checks over spending real steps on a prediction that disagrees with recorded evidence. -Keep speculative mechanisms labeled as hypotheses and state what observation would distinguish competing explanations. - ## Supplied physical parameter menu diff --git a/predicators/agent_sdk/proposal_exec.py b/predicators/agent_sdk/proposal_exec.py index a671599252..c8cc75333b 100644 --- a/predicators/agent_sdk/proposal_exec.py +++ b/predicators/agent_sdk/proposal_exec.py @@ -31,7 +31,7 @@ def _load_sampler_dict( var_name: str, key_error: Callable[[Any], Optional[str]], ) -> Tuple[Dict[str, Any], List[str], Optional[str]]: - """Shared core of the two sampler loaders. + """Core of the ground-sampler loader. Execs ``code`` and validates its ``var_name`` dict. ``key_error`` returns a skip reason for an invalid key (None = valid). Returns @@ -62,27 +62,6 @@ def _load_sampler_dict( return valid, warnings, None -def load_learned_samplers( - code: str, - context: Dict[str, Any], - option_names: Set[str], -) -> Tuple[Dict[str, Any], List[str], Optional[str]]: - """Exec sampler code and validate its ``LEARNED_SAMPLERS`` dict. - - The single loader behind both ``sim.samplers()`` and - ``SamplerLearningMixin._load_samplers_from_module_file``, so the - two cannot drift. Keys must be known option names. - """ - - def key_error(name: Any) -> Optional[str]: - if name not in option_names: - return (f"Skipped '{name}' (not a known option name; known: " - f"{', '.join(sorted(option_names))}).") - return None - - return _load_sampler_dict(code, context, "LEARNED_SAMPLERS", key_error) - - def load_ground_samplers( code: str, context: Dict[str, Any], diff --git a/predicators/agent_sdk/response_parser.py b/predicators/agent_sdk/response_parser.py index c12d52cf84..4dac4be672 100644 --- a/predicators/agent_sdk/response_parser.py +++ b/predicators/agent_sdk/response_parser.py @@ -3,8 +3,8 @@ Converts ``claude_agent_sdk`` message types (``AssistantMessage``, ``UserMessage``, ``ResultMessage``, and the ``SystemMessage`` that marks a context compaction) into plain dicts suitable for logging and -serialization. Used by ``AgentSessionManager``, -``LocalSandboxSessionManager``, and ``docker_agent_runner``. +serialization. Used by ``AgentSessionManager`` and +``LocalSandboxSessionManager``. """ from typing import Any, Dict, List, Optional, Sequence diff --git a/predicators/agent_sdk/sandbox_setup.py b/predicators/agent_sdk/sandbox_setup.py index e41abbfe3a..cebec25855 100644 --- a/predicators/agent_sdk/sandbox_setup.py +++ b/predicators/agent_sdk/sandbox_setup.py @@ -208,7 +208,7 @@ def deny(reason): # this vetoes the interpreter's own import and open events, so a script # cannot reach the hidden predicators modules, their source, or the # harness's run artifacts through computed paths or ``importlib``. Still -# best effort: OS-level isolation (the docker sandbox) is the hard line. +# best effort: it is not OS-level isolation. # --------------------------------------------------------------------------- PYGUARD_DIRNAME = os.path.join(".claude", "pyguard") diff --git a/predicators/agent_sdk/session_base.py b/predicators/agent_sdk/session_base.py index 5365995a2e..f2cd5b5934 100644 --- a/predicators/agent_sdk/session_base.py +++ b/predicators/agent_sdk/session_base.py @@ -897,13 +897,14 @@ def _track_fatal_response(self, response: List[Dict[str, Any]]) -> None: """Terminate the run after consecutive fatally-broken queries. Each fatal query (see :func:`query_fatal_error`) returns in ~1 s - at $0.00 and is otherwise indistinguishable from a no-capture - attempt, so without this check the solve restart / replan / - online-cycle budgets grind through hundreds of instant failures. - The counter is process-wide (class attribute) and any healthy - query resets it; at ``agent_sdk_max_consecutive_fatal_queries`` - the raised :class:`AgentSessionFatalError` propagates past the - per-task handlers and ends the run. + at $0.00 and is otherwise indistinguishable from an attempt that + produced nothing, so without this check the solve restart / + replan / online-cycle budgets grind through hundreds of instant + failures. The counter is process-wide (class attribute) and any + healthy query resets it; at + ``agent_sdk_max_consecutive_fatal_queries`` the raised + :class:`AgentSessionFatalError` propagates past the per-task + handlers and ends the run. """ limit = CFG.agent_sdk_max_consecutive_fatal_queries if limit <= 0: diff --git a/predicators/agent_sdk/sketch_prompts.py b/predicators/agent_sdk/sketch_prompts.py deleted file mode 100644 index 6959e85448..0000000000 --- a/predicators/agent_sdk/sketch_prompts.py +++ /dev/null @@ -1,426 +0,0 @@ -"""Prompt construction for the solve and explore phases. - -Two builders, both rendered from the Markdown templates in -``predicators/agent_sdk/prompts`` (see :mod:`prompt_templates`): - -- :func:`build_solve_system_prompt` composes ``solve_system.md``: the - agent's identity, its deliverable contract, the plan grammar, the - tool semantics, the working principles, the run-record protocol, - and (explore phase) the exploration setting. Everything that holds - for every episode of a phase lives here. -- :func:`build_solve_prompt` composes ``solve_query.md``: the task - (goal, scene, vocabulary), the run state (records, scheduled plans, - open questions), and this episode's instructions. - -The query never restates a rule from the system prompt; the split is -what keeps each rule stated exactly once. -""" -import re -from typing import List, Optional, Sequence, Set - -from predicators import utils -from predicators.agent_sdk.prompt_templates import render -from predicators.settings import CFG -from predicators.structs import LowLevelTrajectory, ParameterizedOption, \ - Predicate, Task - -_BLANK_RUN_RE = re.compile(r"\n{3,}") - - -def _join(parts: Sequence[str]) -> str: - """Join non-empty prompt parts with blank lines, squeezing runs of blank - lines that empty optional sections leave behind.""" - text = "\n\n".join(p for p in parts if p) - return _BLANK_RUN_RE.sub("\n\n", text).strip("\n") + "\n" - - -def build_early_stop_note() -> str: - """The exploration setting's early-stop sentence, from ``CFG``. - - Empty when early stopping is off. Describes the rule the run uses - (train-driven: every attempt must be certified and solve for real; - test-driven: consecutive perfect test phases). - """ - if not CFG.online_learning_early_stopping: - return "" - if CFG.online_learning_early_stopping_by_test_solve_rate: - n_perfect = CFG.online_learning_early_stopping_consecutive_perfect_tests - phases = ("the next test phase solves" if n_perfect <= 1 else - f"{n_perfect} consecutive test phases each solve") - return render("solve_system", "early_stop_test", phases=phases) - attempts_clause = ("every episode of a cycle" if - CFG.online_learning_early_stopping_require_all_attempts - else "each train task's first episode of a cycle") - return render("solve_system", - "early_stop_train", - attempts_clause=attempts_clause) - - -def build_solve_system_prompt( - *, - explore: bool, - policy_mode: bool = False, - propose_params: bool = True, - ground_samplers: bool = False, - physics_margin: bool = False, - necessity: bool = False, - rule_param_margin: bool = False, - use_journal: bool = True, - execute_certified_plan: bool = True, - early_stop_note: str = "", - policy_max_options: int = 0, - policy_max_repeated_failures: int = 0, - policy_max_repeated_noops: int = 0, -) -> str: - """Compose the solve-phase or explore-phase system prompt. - - ``explore`` selects the exploration identity, deliverable, and - setting (``early_stop_note`` and ``execute_certified_plan`` only - matter there); otherwise ``policy_mode`` selects the closed-loop - policy deliverable over the captured plan. ``propose_params`` and - ``ground_samplers`` shape the grammar; ``physics_margin`` and - ``rule_param_margin`` describe the capture gate's extra checks; - ``use_journal`` includes the run-record protocol. - """ - assert not (explore and policy_mode), ( - "exploration delivers a plan sketch even in policy-mode runs") - identity = render("solve_system", - "identity_explore" if explore else "identity_solve") - if explore: - certified_note = "" - if execute_certified_plan: - certified_note = " " + render("solve_system", "certified_note") - deliverable = render("solve_system", - "deliverable_explore", - certified_note=certified_note) - elif policy_mode: - deliverable = render( - "solve_system", - "deliverable_policy", - max_options=str(policy_max_options), - max_repeated_failures=str(policy_max_repeated_failures), - max_repeated_noops=str(policy_max_repeated_noops)) - else: - deliverable = render("solve_system", "deliverable_plan") - - params_rule = render( - "solve_system", - "params_rule_propose" if propose_params else "params_rule_search") - ground_sampler_rule = "" - if propose_params and ground_samplers: - ground_sampler_rule = "\n" + render("solve_system", - "ground_sampler_rule") - grammar = render("solve_system", - "grammar", - param_slot="[p1, p2]" if propose_params else "", - wait_slot="[]" if propose_params else "", - params_rule=params_rule, - ground_sampler_rule=ground_sampler_rule) - - validation_gate = "" - if physics_margin: - validation_gate += " " + render("solve_system", - "validation_gate_physics") - if rule_param_margin: - validation_gate += " " + render("solve_system", - "validation_gate_rule_params") - if necessity: - validation_gate += " " + render("solve_system", - "validation_gate_necessity") - tools = render("solve_system", "tools", validation_gate=validation_gate) - - principles = render("solve_system", "principles") - if not explore: - principles += "\n" + render("solve_system", "banking") - - run_records = "" - if use_journal: - journal_protocol = "" - if not explore: - journal_protocol = "\n\n" + render("solve_system", - "journal_protocol_solve") - run_records = render("solve_system", - "run_records", - journal_protocol=journal_protocol) - - setting = "" - if explore: - note = (" " + early_stop_note.strip()) if early_stop_note else "" - setting = render("solve_system", - "exploration_setting", - early_stop_note=note) - - return _join([ - identity, deliverable, grammar, tools, principles, run_records, setting - ]) - - -def build_solve_prompt( - task: Task, - *, - all_predicates: Set[Predicate], - all_options: Set[ParameterizedOption], - trajectory_summary: str = "", - tool_names: Optional[Sequence[str]] = None, - experiment_guidance: str = "", - scheduled_plans: Optional[Sequence[str]] = None, - initial_image_section: str = "", - propose_params: bool = False, - require_tool_validation: bool = False, - explore_mode: bool = False, - journal: str = "", - strategy: str = "", - attempts: str = "", -) -> str: - """Compose the solve or explore query for ``task``. - - ``propose_params`` labels each option's parameter box as proposed by - the agent (otherwise as found by the search) and picks the stuck- - step advice. ``require_tool_validation`` selects the capture-gate - instructions; ``explore_mode`` the experiment instructions (the two - are exclusive). ``scheduled_plans`` lists the plans this cycle - already queued, so the agent is asked for a complementary one. - ``journal``, ``attempts``, and ``strategy`` are the run records' - contents (``journal.py``); the protocol for using them is in the - system prompt. - """ - assert not (explore_mode and require_tool_validation), ( - "explore_mode accepts an uncaptured experiment sketch, which " - "contradicts the hard capture gate of require_tool_validation") - - init_state = task.init - objects = sorted(init_state, key=lambda o: o.name) - obj_lines = [f" {obj.name}: {obj.type.name}" for obj in objects] - - # Only expose goal atoms whose predicate is in the agent's current - # predicate set: approaches that strip env predicates rely on - # goal_nl and must not leak atoms of predicates the agent invents. - goal_atoms = [ - str(a) for a in sorted(task.goal, key=str) - if a.predicate in all_predicates - ] - - option_lines = [] - for opt in sorted(all_options, key=lambda o: o.name): - type_sig = ", ".join(t.name for t in opt.types) - params_dim = opt.params_space.shape[0] - param_info = "" - if params_dim > 0: - low = opt.params_space.low.tolist() - high = opt.params_space.high.tolist() - label = "params" if propose_params else "auto-searched params" - desc = (", ".join(opt.params_description) - if opt.params_description else f"{params_dim}d") - param_info = f" [{label}: {desc}, range {low} to {high}]" - option_lines.append(f" {opt.name}({type_sig}){param_info}") - - atoms = utils.abstract(init_state, all_predicates) - atom_lines = [str(a) for a in sorted(atoms, key=str)] - - pred_lines = [] - any_latent = False - for pred in sorted(all_predicates, key=lambda p: p.name): - type_sig = ", ".join(t.name for t in pred.types) - line = f" {pred.name}({type_sig})" - if pred.accepts_latent: - line += " [reads belief latent]" - any_latent = True - if pred.natural_language_assertion is not None: - names = [t.name for t in pred.types] - line += f": {pred.natural_language_assertion(names)}" - pred_lines.append(line) - if any_latent: - pred_lines.append( - " ([reads belief latent]: its truth can depend on belief-only " - "state that real observations never carry, so an annotation " - "using it may be dropped from closed-loop monitoring at " - "capture as execution-unverifiable - prefer predicates over " - "observable features for monitored subgoals.)") - - goal_nl_section = "" - if task.goal_nl: - goal_nl_section = render("solve_query", - "goal_nl", - goal_nl=task.goal_nl) - - # The env's public reward form (success condition plus costs) when - # the task ships an evaluator that states one; never oracle values. - scoring_section = "" - evaluator = getattr(task, "evaluator", None) - if evaluator is not None and evaluator.objective_description(): - scoring_section = render("solve_query", - "scoring", - objective=evaluator.objective_description()) - - goal_atoms_section = "" - if goal_atoms: - goal_atoms_section = render("solve_query", - "goal_atoms", - goal_atoms="\n".join(goal_atoms)) - - experiment_section = "" - if experiment_guidance: - experiment_section = render("solve_query", - "experiment_guidance", - guidance=experiment_guidance) - - tools_section = "" - if tool_names: - tools_section = render("solve_query", - "tools", - tool_list="\n".join(f" - {t}" - for t in tool_names)) - - strategy_section = "" - if strategy: - strategy_section = render("solve_query", "strategy", strategy=strategy) - - attempts_section = "" - if attempts: - attempts_section = render("solve_query", "attempts", attempts=attempts) - - journal_section = "" - if journal or attempts: - journal_section = render("solve_query", - "journal", - journal=journal - or render("solve_query", "no_journal")) - - scheduled_section = "" - if scheduled_plans: - plans = "\n".join(f"Plan {i + 1}:\n{p}" - for i, p in enumerate(scheduled_plans)) - certified_rule = "" - if any("belief-certified" in p for p in scheduled_plans): - certified_rule = "\n\n" + render("solve_query", "certified_rule") - scheduled_section = render("solve_query", - "scheduled_plans", - plans=plans, - certified_rule=certified_rule) - - if explore_mode: - instructions = render("solve_query", "instructions_explore") - elif require_tool_validation: - stuck = render("solve_query", - "stuck_tune" if propose_params else "stuck_revise") - instructions = render("solve_query", - "instructions_capture", - stuck_advice=stuck) - else: - instructions = render("solve_query", - "instructions_capture", - stuck_advice=render("solve_query", - "stuck_revise")) - - body = render( - "solve_query", - "skeleton", - opening=render("solve_query", - "opening_explore" if explore_mode else "opening_solve"), - goal_nl_section=goal_nl_section, - scoring_section=scoring_section, - goal_atoms_section=goal_atoms_section, - experiment_section=experiment_section, - atoms="\n".join(atom_lines) if atom_lines else render( - "solve_query", "no_atoms"), - # 4 decimals (0.1 mm at metric scale), matching the per-step - # state dumps: the default 2 dp misstated mm-scale geometry - # (a 0.025 half-extent printed as 0.03) and sent agents to - # sim.task()/sim.state() to re-derive every number. - state=init_state.dict_str(indent=2, num_decimal_points=4), - image_section=initial_image_section.strip("\n"), - objects="\n".join(obj_lines), - options="\n".join(option_lines), - predicates="\n".join(pred_lines), - trajectory_summary=trajectory_summary.strip("\n"), - tools_section=tools_section, - strategy_section=strategy_section, - attempts_section=attempts_section, - journal_section=journal_section, - scheduled_plans_section=scheduled_section, - instructions=instructions, - ) - return _join([body]) - - -def _executed_options(traj: LowLevelTrajectory) -> List[str]: - """The option sequence a trajectory executed, consecutive repeats - collapsed, from the options its actions carry; empty when the actions carry - none (a demo replayed from raw actions, or a pickle that dropped them).""" - out: List[str] = [] - for act in traj.actions: - if not act.has_option(): - continue - opt = act.get_option() - text = f"{opt.name}({', '.join(o.name for o in opt.objects)})" - if not out or out[-1] != text: - out.append(text) - return out - - -def summarize_trajectories( - trajectories: Sequence[LowLevelTrajectory], - predicates: Set[Predicate], - train_tasks: Optional[Sequence[Task]] = None) -> str: - """The ``## Trajectory Summary`` query section, or "". - - Per recent trajectory (the last - ``CFG.agent_sdk_max_trajectories_in_context``): the option plan it - executed, the env's verdict on it (reward and whether the goal - atoms held at the end, when the episode was evaluated), and the - atoms gained and lost between its first and last state. Shared by - the solve approach and the explorers. - - Trajectories are numbered by their index in the full list, so a - number means the same episode in every session of a run. The - verdict lines are what an agent with no learn phase otherwise never - sees: the model-free arm's cycle-1 session used to read only - "Trajectory 0: 42 steps" for an episode whose plan had failed. - """ - if not trajectories: - return "" - max_trajs = CFG.agent_sdk_max_trajectories_in_context - recent = trajectories[-max_trajs:] - first = len(trajectories) - len(recent) - lines = [ - f"\n## Trajectory Summary ({len(trajectories)} total, " - f"showing last {len(recent)})" - ] - for i, traj in enumerate(recent): - n_steps = len(traj.actions) - init_atoms = utils.abstract(traj.states[0], predicates) - final_atoms = utils.abstract(traj.states[-1], predicates) - new_atoms = final_atoms - init_atoms - lost_atoms = init_atoms - final_atoms - lines.append(f"\nTrajectory {first + i}: {n_steps} steps") - executed = _executed_options(traj) - if executed: - lines.append(" Executed: " + " -> ".join(executed)) - verdict: List[str] = [] - reward = traj.env_reward - if reward is not None: - verdict.append(f"env reward {reward:.2f}") - terminated = traj.env_terminated - if terminated is not None: - verdict.append("goal atoms held at the end" - if terminated else "goal NOT reached") - elif train_tasks is not None: - try: - task = train_tasks[traj.train_task_idx] - except (IndexError, ValueError, AttributeError): - task = None - if task is not None: - held = task.goal_holds(traj.states[-1]) - verdict.append("goal atoms held at the end" - if held else "goal NOT reached") - if verdict: - lines.append(" Outcome: " + ", ".join(verdict)) - if new_atoms: - lines.append( - " Gained: " + - f"{', '.join(str(a) for a in sorted(new_atoms, key=str))}") - if lost_atoms: - lines.append( - " Lost: " + - f"{', '.join(str(a) for a in sorted(lost_atoms, key=str))}") - return "\n".join(lines) diff --git a/predicators/agent_sdk/sketch_refinement.py b/predicators/agent_sdk/sketch_refinement.py index ea72664c75..5513445915 100644 --- a/predicators/agent_sdk/sketch_refinement.py +++ b/predicators/agent_sdk/sketch_refinement.py @@ -25,8 +25,8 @@ from predicators.agent_sdk.sketch_types import SketchStep from predicators.option_model import _OptionModelBase from predicators.planning import run_backtracking_refinement -from predicators.structs import GroundAtom, ParameterizedOption, \ - ParameterizedSampler, Predicate, State, Task, _Option +from predicators.structs import GroundAtom, ParameterizedOption, Predicate, \ + State, Task, _Option # Signature of an info-gain scorer: given a candidate post-state and the # atoms whose truth the step is meant to establish, return a scalar where @@ -152,7 +152,6 @@ class _RefineContext: deepest_failure_holder: Optional[List[DeepestFailure]] info_scorer: Optional[InfoScorer] info_n_feasible_target: int - parameterized_samplers: Optional[Dict[str, ParameterizedSampler]] solved_check: Optional[Callable[[List[State], List[Any], bool], Tuple[bool, str]]] # Proposed continuous params are decisions, not seeds: a step that @@ -202,7 +201,6 @@ class _RefinementState: # Options whose synthesized sampler already misbehaved once - so the # per-draw fallback warning fires at most once per option, not on every # one of the (potentially thousands of) draws during backtracking. - sampler_warned: Set[str] = dataclasses.field(default_factory=set) # Step indices whose LLM-proposed initial_params have already been used -- # tried directly on the plain path, or seeded into the info-seeking pool. @@ -248,47 +246,18 @@ class _RefinementState: elapsed: List[float] = dataclasses.field(default_factory=list) -def _draw_params(search: _RefinementState, ctx: _RefineContext, - step: SketchStep, state: State, +def _draw_params(step: SketchStep, state: State, rng_: np.random.Generator) -> np.ndarray: - """Draw continuous params for a step's option. - - Precedence, most specific first: the step's ground sampler (a ``~`` - annotation compiled into a ``GroundSampler``: uniform window or - named code fn), then the option's learned parameterized sampler - (keyed by option name), then uniform ``sample_params`` - the - fallback also on a sampler error or wrong-shaped return (a - misbehaving ground fn falls all the way to uniform, not to the - parameterized sampler, mirroring the parameterized fallback). - """ + """Draw continuous params for a step's option: from the step's ground + sampler (a ``~`` annotation compiled into a ``GroundSampler``: uniform + window or named code fn) when it has one and it draws, uniformly with + ``sample_params`` otherwise.""" if step.ground_sampler is not None: drawn = step.ground_sampler.draw(state, rng_, step.option.params_space, step.objects, step.subgoal_atoms or set()) if drawn is not None: return drawn - return sample_params(step.option, rng_) - sampler = (ctx.parameterized_samplers.get(step.option.name) - if ctx.parameterized_samplers else None) - if sampler is not None: - box = step.option.params_space - expected = box.shape[0] - try: - raw = sampler(state, step.subgoal_atoms or set(), rng_, - list(step.objects)) - params = np.asarray(raw, dtype=np.float32).reshape(-1) - if params.shape == (expected, ): - return np.clip(params, box.low, box.high) - reason = (f"returned shape {params.shape}, " - f"expected ({expected},)") - except Exception as e: # pylint: disable=broad-except - reason = f"raised {type(e).__name__}: {e}" - if step.option.name not in search.sampler_warned: - search.sampler_warned.add(step.option.name) - logging.warning( - "[%s] synthesized sampler for %s %s; falling back to " - "uniform sampling for this option.", ctx.run_id, - step.option.name, reason) return sample_params(step.option, rng_) @@ -347,21 +316,16 @@ def _is_pinned(ctx: _RefineContext, step: SketchStep) -> bool: and step.option.params_space.shape[0] > 0) -def _is_deterministic(ctx: _RefineContext, step: SketchStep) -> bool: - """Whether the step's sampler flags itself as returning constant params.""" - # A sampler may flag itself as returning constant params (ignoring - # state/rng); re-drawing it yields the identical option, so its step - # gets a single attempt -- backtracking then skips straight past it - # instead of wasting the full budget re-descending through it. - if step.ground_sampler is not None: - # A ground-sampler step bypasses the parameterized sampler, - # so a deterministic sampler flag must not collapse it to - # one attempt. An all-zero window pins every draw to the - # center, which IS deterministic - one attempt suffices. - return step.ground_sampler.deterministic - sampler = (ctx.parameterized_samplers.get(step.option.name) - if ctx.parameterized_samplers else None) - return bool(getattr(sampler, "deterministic", False)) +def _is_deterministic(step: SketchStep) -> bool: + """Whether the step's ground sampler returns constant params. + + Re-drawing it yields the identical option, so its step gets a single + attempt: backtracking then skips straight past it instead of wasting + the full budget re-descending through it. An all-zero window pins + every draw to the center, which is deterministic. + """ + return (step.ground_sampler is not None + and step.ground_sampler.deterministic) def _sample_info_seeking(search: _RefinementState, ctx: _RefineContext, @@ -493,8 +457,7 @@ def _consider(grounded: _Option) -> None: if len(scored) > n_pooled_before else "infeasible — not pooled") while len(scored) < ctx.info_n_feasible_target and n_draws < draw_cap: - grounded = ground_step(step, - _draw_params(search, ctx, step, state, rng_)) + grounded = ground_step(step, _draw_params(step, state, rng_)) n_draws += 1 _consider(grounded) pool.spent += n_draws @@ -573,7 +536,7 @@ def _sample_step(search: _RefinementState, ctx: _RefineContext, idx: int, logging.debug("[%s] step %d %s: trying LLM-proposed params %s", ctx.run_id, idx, step.option.name, params.tolist()) return ground_step(step, params) - return ground_step(step, _draw_params(search, ctx, step, state, rng_)) + return ground_step(step, _draw_params(step, state, rng_)) def _validate_step(search: _RefinementState, ctx: _RefineContext, idx: int, @@ -703,7 +666,6 @@ def refine_sketch( deepest_failure_holder: Optional[List[DeepestFailure]] = None, info_scorer: Optional[InfoScorer] = None, info_n_feasible_target: int = 1, - parameterized_samplers: Optional[Dict[str, ParameterizedSampler]] = None, strip_latent_wait_targets: bool = True, solved_check: Optional[Callable[[List[State], List[Any], bool], Tuple[bool, str]]] = None, @@ -798,17 +760,9 @@ def refine_sketch( that ``WaitOption`` terminates on the intended atom change rather than the first incidental one. - ``parameterized_samplers`` maps an option name to a parameterized - (per-skill) sampler ``(state, subgoal_atoms, rng, objects) -> - params`` (the NSRTSampler signature, with the step subgoal in the - atoms slot), used on both plain and info-seeking draws to aim that - option's parameters at the subgoal instead of drawing uniformly. - The return is clipped to the option's box; a missing or misbehaving - sampler falls back to uniform sampling. A step whose sketch line - carries a ``~ [widths]`` region annotation bypasses the sampler - entirely: after the one-shot center try, its draws come from the - step's ``GroundSampler``, the most specific prior winning - ground - sampler, then parameterized sampler, then uniform. + A step whose sketch line carries a ``~ [widths]`` region annotation + draws, after the one-shot center try, from the step's + ``GroundSampler``; other steps draw uniformly. """ if not sketch: return RefineOutcome(plan=[], @@ -833,7 +787,6 @@ def refine_sketch( deepest_failure_holder=deepest_failure_holder, info_scorer=info_scorer, info_n_feasible_target=info_n_feasible_target, - parameterized_samplers=parameterized_samplers, solved_check=solved_check, pin_proposed_params=pin_proposed_params, pinned_step_retries=max(1, pinned_step_retries)) @@ -853,7 +806,7 @@ def refine_sketch( for _step in sketch: if _step.option.params_space.shape[0] == 0: max_tries.append(1) - elif _is_deterministic(ctx, _step): + elif _is_deterministic(_step): max_tries.append(1) elif _is_pinned(ctx, _step): max_tries.append(ctx.pinned_step_retries) @@ -1045,7 +998,6 @@ def suggest_probes( rng: np.random.Generator, max_draws: int = 20, top_k: int = 3, - parameterized_samplers: Optional[Dict[str, ParameterizedSampler]] = None, on_rollout: Optional[Callable[[], None]] = None, plan_scorer: Optional[Callable[[List[_Option], Set[GroundAtom]], Tuple[float, Dict[str, float]]]] = None, @@ -1084,26 +1036,6 @@ def _score(option: _Option, nxt: State, for a in sorted(atoms, key=str) } - ctx = _RefineContext(task=task, - sketch=sketch, - option_model=option_model, - predicates=predicates, - max_samples_per_step=max_draws, - check_subgoals=True, - check_final_goal=False, - log_state=False, - run_id="suggest_probes", - on_step_fail=None, - deepest_failure_holder=None, - info_scorer=info_scorer, - info_n_feasible_target=1, - parameterized_samplers=parameterized_samplers, - solved_check=None, - pin_proposed_params=True, - pinned_step_retries=1) - search = _RefinementState(step_pools=[None] * len(sketch), - step_trajs=[None] * len(sketch), - step_samples_cumulative=[0] * len(sketch)) suggestions: List[StepProbeSuggestion] = [] notes: List[str] = [] state = task.init @@ -1147,8 +1079,7 @@ def _roll(grounded: _Option) -> Optional[State]: best_option: Optional[_Option] = None if has_params and atoms: for _ in range(max_draws): - grounded = ground_step( - step, _draw_params(search, ctx, step, state, rng)) + grounded = ground_step(step, _draw_params(step, state, rng)) n_draws += 1 nxt = _roll(grounded) if nxt is None or not atoms.issubset( @@ -1232,7 +1163,6 @@ def refine_and_validate_report( max_samples_per_step: int, check_subgoals: bool, log_state: bool = False, - parameterized_samplers: Optional[Dict[str, ParameterizedSampler]] = None, run_id: str = "refine", timeout_source: str = "explicit", extra_summary_lines: Optional[List[str]] = None, @@ -1278,7 +1208,6 @@ def refine_and_validate_report( check_subgoals=check_subgoals, log_state=log_state, run_id=run_id, - parameterized_samplers=parameterized_samplers, solved_check=solved_check, strip_latent_wait_targets=strip_latent_wait_targets, ) diff --git a/predicators/agent_sdk/sketch_types.py b/predicators/agent_sdk/sketch_types.py index 113c624572..4cba67acb3 100644 --- a/predicators/agent_sdk/sketch_types.py +++ b/predicators/agent_sdk/sketch_types.py @@ -21,15 +21,11 @@ class GroundSampler: """Per-step (ground) sampler compiled from a sketch annotation. - The ground level of the two-level sampler hierarchy that - ``_draw_params`` consults: ground sampler (this, most specific) > - learned parameterized sampler (``parameterized_samplers``, keyed by - option name) > uniform. A parameterized sampler is authored once - and sees every ground call of its option; a ground sampler is - declared inline for ONE step of ONE sketch and dies with the call - - it lives on the ``SketchStep`` rather than in the option-name-keyed - registry, which could not hold different distributions for two - same-option steps in one sketch. + ``_draw_params`` draws a step's params from its ground sampler when + it has one, uniformly otherwise. A ground sampler is declared inline + for ONE step of ONE sketch and dies with the call: it lives on the + ``SketchStep``, so two same-option steps in one sketch can draw from + different distributions. Two kinds, one per instance: - window (``center`` + ``width`` set): the uniform box a @@ -37,9 +33,9 @@ class GroundSampler: proposed params; - code (``fn`` + ``name`` set): an agent-written function that a ``~ my_sampler`` annotation references by name (loaded fresh per - refine call from the sandbox's ``GROUND_SAMPLERS``); it shares - the parameterized-sampler call signature, so it can shape any - state-conditioned distribution. + refine call from the sandbox's ``GROUND_SAMPLERS``) with the + signature ``(state, subgoal_atoms, rng, objects) -> params``, so + it can shape any state-conditioned distribution. """ center: Optional[np.ndarray] = None width: Optional[np.ndarray] = None @@ -58,7 +54,7 @@ def draw(self, state: State, rng: np.random.Generator, box: Box, Returns ``None`` when a code fn misbehaves (raises or returns a wrong-shaped array); the caller falls back to uniform sampling - for that draw, mirroring the parameterized-sampler fallback. + for that draw. """ if self.fn is not None: try: diff --git a/predicators/agent_sdk/synthesis_backend.py b/predicators/agent_sdk/synthesis_backend.py index 7acc85fbd9..5e84e594bb 100644 --- a/predicators/agent_sdk/synthesis_backend.py +++ b/predicators/agent_sdk/synthesis_backend.py @@ -2,8 +2,8 @@ :class:`SynthesisBackend` declares exactly the approach surface that the synthesis tool factories in :mod:`predicators.agent_sdk.tools` -(``create_synthesis_tools``, ``make_predicate_quality_loader``, -``make_sampler_loader``) and the approach-layer validation +(``create_synthesis_tools``, ``make_predicate_quality_loader``) and +the approach-layer validation glue in :mod:`predicators.approaches.synthesis_validation` dereference. It exists so those modules can be typed against the contract instead of importing the concrete ``AgentSimLearningApproach`` - the import that @@ -22,8 +22,7 @@ from predicators.code_sim_learning.utils import LearnedSimulator from predicators.option_model import _OracleOptionModel from predicators.structs import Action, LowLevelTrajectory, \ - ParameterizedOption, ParameterizedSampler, Predicate, State, Task, \ - Type + ParameterizedOption, Predicate, State, Task, Type class SynthesisBackend(Protocol): @@ -60,8 +59,6 @@ class SynthesisBackend(Protocol): float]]]] # ── State written by the tools ─────────────────────────────── - # Per-skill samplers keyed by option name. - _synthesized_samplers: Dict[str, ParameterizedSampler] # Candidate simulator state published during validation so the # recurrent combined simulator sees the rules under evaluation. _residual_rules: Optional[List] @@ -94,9 +91,6 @@ def _get_all_predicates(self) -> Set[Predicate]: def _get_all_options(self) -> Set[ParameterizedOption]: ... - def _get_all_samplers(self) -> Dict[str, ParameterizedSampler]: - ... - def _group_triples_by_trajectory( self, triples: List[Tuple[State, Action, State]], @@ -137,29 +131,6 @@ def previous_fit_evidence( self, version_tag: str) -> Optional[Dict[str, LaplaceEvidence]]: """The previous version's recorded evidence, as a one-entry dict.""" - def _record_sysid_diagnostics(self, report: Dict[str, Dict[str, Any]], - physical_names: Sequence[str], - num_survivors: int, num_segments: int, - rms: List[float]) -> None: - ... - - def _fit_parameters_recurrent( - self, - rules: List, - specs: List[ParamSpec], - base_pred_triples: List[Tuple[State, Action, State]], - residual_features: Dict[str, List[str]], - ) -> Tuple[FitResult, float]: - ... - - def _fit_parameters_joint_rollout( - self, - rules: List, - rule_specs: List[ParamSpec], - residual_features: Dict[str, List[str]], - ) -> Tuple[FitResult, float]: - ... - def _build_combined_simulator( self, learned_simulator: LearnedSimulator, @@ -195,23 +166,3 @@ class PredicateSynthesisBackend(SynthesisBackend, Protocol): # Initial predicates that survived retraction, used to build the # exec namespace the agent's predicate code runs in. _kept_initial_predicates: Set[Predicate] - - -class SamplerSynthesisBackend(Protocol): - """The narrow surface ``make_sampler_loader`` needs. - - ``SamplerLearningMixin`` satisfies this directly (its declared host- - class contract covers every member), so the mixin can pass ``self`` - without seeing the full backend. - """ - - _fitted_params: Dict[str, float] - _train_tasks: List[Task] - _types: Set[Type] - _synthesized_samplers: Dict[str, ParameterizedSampler] - - def _get_all_predicates(self) -> Set[Predicate]: - ... - - def _get_all_options(self) -> Set[ParameterizedOption]: - ... diff --git a/predicators/agent_sdk/tools/__init__.py b/predicators/agent_sdk/tools/__init__.py index 6bdd8e555e..c31998d24f 100644 --- a/predicators/agent_sdk/tools/__init__.py +++ b/predicators/agent_sdk/tools/__init__.py @@ -3,22 +3,21 @@ This package replaces the former single-module ``tools.py``. Layout: - ``registry``: tool-name rosters and the session tool-list surface. -- ``context``: ``ToolContext`` / ``PlanCapture`` shared session state. +- ``context``: ``ToolContext``, the shared session state. - ``results``: tool-result formatting and sandbox-file helpers. - ``sandbox_guard``: sandbox-escape screening for agent-supplied text. -- ``budget``: solve-attempt budget footer and watchdog. +- ``budget``: the round's budget footer and the run_python watchdog. - ``scene``: scene rendering and state-manipulation helpers. - ``verdicts``: task-evaluator verdicts and ground-sampler loading. -- ``testing`` / ``exploration``: the static MCP - tool builders, assembled by ``assembly.create_mcp_tools``. +- ``exploration``: the static ``run_python`` builder, assembled by + ``assembly.create_mcp_tools``. - ``digests``: the type / option / task / trajectory digest renderers shared by the prompts and the probe. - ``snapshots``: versioned write-time snapshots of agent-edited files. - ``python_exec``: shared python-exec core (run_python / run_python). - ``synthesis`` / ``params_view``: the synthesis-session tool factory. -- ``predicate_synthesis`` / ``sampler_synthesis``: the loaders behind - ``sim.predicates()`` / ``sim.samplers()``. +- ``predicate_synthesis``: the loader behind ``sim.predicates()``. This facade re-exports the package's public surface (plus a few underscore names kept for pre-split imports); new code should import @@ -26,17 +25,15 @@ """ # pylint: disable=unused-import from predicators.agent_sdk.tools.assembly import create_mcp_tools -from predicators.agent_sdk.tools.context import PlanCapture, ToolContext +from predicators.agent_sdk.tools.context import ToolContext from predicators.agent_sdk.tools.params_view import _ParamsView from predicators.agent_sdk.tools.predicate_synthesis import \ make_predicate_quality_loader from predicators.agent_sdk.tools.registry import ALL_TOOL_NAMES, \ BUILTIN_TOOLS, EXPLORATION_TOOL_NAMES, MCP_SERVER_NAME, \ - SYNTHESIS_TOOL_NAMES, TESTING_TOOL_NAMES, get_allowed_tool_list, \ - list_session_tool_names + SYNTHESIS_TOOL_NAMES, get_allowed_tool_list, list_session_tool_names from predicators.agent_sdk.tools.results import _make_coercing_tool, \ _make_spilling_text_result, session_log_filename -from predicators.agent_sdk.tools.sampler_synthesis import make_sampler_loader from predicators.agent_sdk.tools.sandbox_guard import \ SANDBOX_HIDDEN_MODULES_PATTERN, SANDBOX_INTROSPECTION, \ SANDBOX_SYSTEM_ROOTS, _screen_text_for_sandbox_escape @@ -58,8 +55,6 @@ "SANDBOX_INTROSPECTION", "SANDBOX_SYSTEM_ROOTS", "SYNTHESIS_TOOL_NAMES", - "TESTING_TOOL_NAMES", - "PlanCapture", "ToolContext", "agent_render_resolution", "apply_state_modifications", @@ -73,7 +68,6 @@ "list_session_tool_names", "load_ground_sampler_fns", "make_predicate_quality_loader", - "make_sampler_loader", "make_solved_check", "make_write_snapshot_hook", "render_pybullet_image", diff --git a/predicators/agent_sdk/tools/assembly.py b/predicators/agent_sdk/tools/assembly.py index 29d8ce3116..e8a856d997 100644 --- a/predicators/agent_sdk/tools/assembly.py +++ b/predicators/agent_sdk/tools/assembly.py @@ -5,7 +5,6 @@ from predicators.agent_sdk.tools.exploration import _build_exploration_tools from predicators.agent_sdk.tools.results import _make_coercing_tool, \ _make_spilling_text_result -from predicators.agent_sdk.tools.testing import _build_testing_tools def create_mcp_tools(ctx: ToolContext, @@ -35,11 +34,8 @@ def create_mcp_tools(ctx: ToolContext, # probe in one namespace), so the solve-phase instance is neither # built nor offered there. extra_names = {getattr(t, "name", "") for t in ctx.extra_mcp_tools} - _all = { - **_build_testing_tools(ctx, _text_result, tool), - **({} if "run_python" in extra_names else _build_exploration_tools( - ctx, _text_result, tool)), - } + _all = ({} if "run_python" in extra_names else _build_exploration_tools( + ctx, _text_result, tool)) if tool_names is None: tools = list(_all.values()) else: diff --git a/predicators/agent_sdk/tools/budget.py b/predicators/agent_sdk/tools/budget.py index 7540a32a0e..329f63dce3 100644 --- a/predicators/agent_sdk/tools/budget.py +++ b/predicators/agent_sdk/tools/budget.py @@ -8,25 +8,20 @@ def _budget_footer(ctx: ToolContext, rollouts_before: int = 0) -> str: - """``[budget]`` line appended to tool results during a solve attempt. + """``[budget]`` line appended to tool results during a play round. - Shows attempt wall-clock (elapsed, and the budget when one is set) - and the attempt's cumulative sim-rollout count (plus this call's - delta). Agents pace well when they can see a clock and terribly when - they can't: the 47k-rollout single-call sweep of run_20260717_230436 - ran 7 h with zero cost feedback. Empty when no attempt is in flight. + Shows the round's elapsed wall clock and its cumulative sim-rollout + count (plus this call's delta). Agents pace well when they can see a + clock and terribly when they can't: the 47k-rollout single-call + sweep of run_20260717_230436 ran 7 h with zero cost feedback. Empty + when no attempt is in flight. """ start = ctx.attempt_start if start is None: return "" parts = [] elapsed_min = (time.monotonic() - start) / 60.0 - deadline = ctx.attempt_deadline - if deadline is not None: - total_min = (deadline - start) / 60.0 - parts.append(f"attempt time {elapsed_min:.1f}/{total_min:.0f} min") - else: - parts.append(f"attempt time {elapsed_min:.1f} min") + parts.append(f"attempt time {elapsed_min:.1f} min") total_rollouts = ctx.attempt_rollout_count delta = total_rollouts - rollouts_before rollout_part = f"sim rollouts this attempt: {total_rollouts}" diff --git a/predicators/agent_sdk/tools/capture.py b/predicators/agent_sdk/tools/capture.py deleted file mode 100644 index 6b70de8462..0000000000 --- a/predicators/agent_sdk/tools/capture.py +++ /dev/null @@ -1,174 +0,0 @@ -"""The ``submit_plan`` capture decision, as a pure function. - -:func:`_decide_capture` encodes the run-verified capture gates in one -side-effect-free place: the handler in ``testing.py`` computes the -inputs, then performs the ctx mutations and message formatting its -decision calls for. Keeping the policy pure makes every guard -combination directly unit-testable -(``tests/agent_sdk/test_capture_decision.py``). -""" -import enum -from dataclasses import dataclass -from typing import Optional - - -class CaptureDecision(enum.Enum): - """What ``submit_plan`` does with the submitted plan.""" - # Goal reached, evaluator-certified, every validation rollout passed: - # captured and marked as a validated solve. - VALIDATED_CAPTURE = "validated_capture" - # Final-submission nudge: captured to execute for its honest reward, - # but not marked as a solve (see BestEffortReason). - BEST_EFFORT_CAPTURE = "best_effort_capture" - # Goal reached on rollout 1 but a validation repeat failed: refused, - # and later submissions on the task face the escalated gate. - FLAKY_NO_CAPTURE = "flaky_no_capture" - # Execution validation passed, but a rollout at +-1-posterior-sigma - # perturbed physical params failed: the plan has no margin to the - # physics fit's parameter error, so it is refused (the real env may - # sit anywhere in that range). - PARAM_SENSITIVE_NO_CAPTURE = "param_sensitive_no_capture" - # Every gate above passed, but the plan still reaches the goal with - # one of its steps removed: that step is padding, the plan is not an - # explanation of the goal, and it is refused naming the step. - REDUNDANT_NO_CAPTURE = "redundant_no_capture" - # Goal atoms reached via a route the task evaluator scores as a - # non-solve: refused. - REWARD_HACK_NO_CAPTURE = "reward_hack_no_capture" - # Goal reached on a train task instead of the current task: not - # captured, flagged loudly. - WRONG_TASK_NOTE = "wrong_task_note" - # Nothing to capture and nothing to flag. - NO_CAPTURE = "no_capture" - - -class BestEffortReason(enum.Enum): - """Why a best-effort capture cannot count as a validated solve.""" - GOAL_NOT_REACHED = "goal_not_reached" # honest shortfall - REWARD_HACK = "reward_hack" # evaluator scores the rollout a non-solve - FLAKY = "flaky" # a validation repeat failed - PARAM_SENSITIVE = "param_sensitive" # failed at perturbed physics - REDUNDANT = "redundant" # reaches the goal without one of its steps - - -@dataclass(frozen=True) -class CaptureOutcome: - """A capture decision; ``best_effort_reason`` is set iff the decision is - ``BEST_EFFORT_CAPTURE``.""" - decision: CaptureDecision - best_effort_reason: Optional[BestEffortReason] = None - - @property - def captured(self) -> bool: - """Whether the plan is captured as the current answer.""" - return self.decision in (CaptureDecision.VALIDATED_CAPTURE, - CaptureDecision.BEST_EFFORT_CAPTURE) - - -def _decide_capture(*, - capture_enabled: bool, - is_current_task: bool, - have_plan: bool, - goal_achieved: bool, - evaluator_rejected: bool, - reward_hack: bool, - flaky: bool, - best_effort_mode: bool, - have_validated_capture: bool, - param_sensitive: bool = False, - redundant: bool = False) -> CaptureOutcome: - """Decide what ``submit_plan`` does with an evaluated plan. - - Pure: no ctx access, no I/O - the caller supplies exactly what the - gates read and applies the side effects the decision calls for. - - - ``capture_enabled``: ``ctx.capture_goal_reaching_plans`` - only - approaches that consume captured plans set it. - - ``is_current_task``: the plan ran on the current solve/explore - task (no ``task_idx`` given), the only task whose captures count. - - ``have_plan``: the parsed plan grounded to at least one step. - - ``goal_achieved``: goal reached, clean to goal, and within the - episode horizon on the first rollout. - - ``evaluator_rejected``: a non-coarse evaluator verdict scored the - first rollout illegitimate. - - ``reward_hack``: ``evaluator_rejected`` and the rollout actually - reached the goal atoms (terminated) - an illegitimate route, as - opposed to an honest shortfall. - - ``flaky``: a validation repeat failed. - - ``best_effort_mode``: ``ctx.capture_best_effort_plan`` - set only - for the final-submission nudge after an attempt exhausted its - turn budget. - - ``have_validated_capture``: a validated-solve capture already - exists (``ctx.solved_plan_reached_goal``). - - ``param_sensitive``: execution validation passed but a rollout at - +-1-posterior-sigma perturbed physical params failed - the plan - has no margin to the physics fit's parameter error. - - ``redundant``: every other gate passed but the plan still reaches - the goal with one of its steps removed - that step is padding. - - With ``best_effort_mode`` (final-submission nudge after turn-cap - exhaustion) capture the submission unconditionally: honest - shortfall, evaluator-rejected rollout, and flaky repeat alike. The - budget is spent, so executing the agent's best plan for its honest - reward beats forfeiting the task (run_20260714_145053 task 4: a - goal-reaching but certificate-rejected final submission was refused - and the task forfeited, scoring n/a instead of its honest reward). - Only a validated solve is marked as one; everything else is a - best-effort capture that executes but cannot count as a solve - the - certificate still protects the score. - """ - # A best-effort capture never displaces a validated-solve capture. - best_effort_capture = best_effort_mode and not have_validated_capture - validated_solve = (goal_achieved and not reward_hack and not flaky - and not param_sensitive and not redundant) - if (capture_enabled and is_current_task - and (validated_solve or best_effort_capture) and have_plan): - if validated_solve: - return CaptureOutcome(CaptureDecision.VALIDATED_CAPTURE) - if not goal_achieved: - reason = BestEffortReason.GOAL_NOT_REACHED - elif reward_hack: - reason = BestEffortReason.REWARD_HACK - elif flaky: - reason = BestEffortReason.FLAKY - elif param_sensitive: - reason = BestEffortReason.PARAM_SENSITIVE - else: - reason = BestEffortReason.REDUNDANT - return CaptureOutcome(CaptureDecision.BEST_EFFORT_CAPTURE, reason) - if (capture_enabled and is_current_task and goal_achieved - and not evaluator_rejected and flaky): - # Loudly refuse a flaky capture: the agent still has this - # session to add margin and resubmit, which beats discovering - # the flakiness as a failed real episode. - return CaptureOutcome(CaptureDecision.FLAKY_NO_CAPTURE) - if (capture_enabled and is_current_task and goal_achieved - and not evaluator_rejected and param_sensitive): - # Loudly refuse a physics-margin failure: the fitted params are - # uncertain at the reported sigma, and the real env may sit - # anywhere in that range (run_20260723_091108: a capture - # validated 8/8 at the fitted friction failed deterministically - # at the true value just outside the design's success band). - return CaptureOutcome(CaptureDecision.PARAM_SENSITIVE_NO_CAPTURE) - if (capture_enabled and is_current_task and goal_achieved - and not evaluator_rejected and redundant): - # Loudly refuse padding: the plan reaches the goal without one of - # its steps, so that step explains nothing and only spends real - # episode steps (run_20260902_152811: three presses and a release - # for a goal the model reached with two presses and a Wait). - return CaptureOutcome(CaptureDecision.REDUNDANT_NO_CAPTURE) - if capture_enabled and is_current_task and reward_hack: - # Loudly refuse a reward hack: the rollout reaches the goal atoms - # but the evaluator's certificate rejects the route (e.g. the - # target was knocked over directly), and the real evaluator - # applies the same certificate, so it can never count as a solve. - # (Under a best-effort final submission the same plan is instead - # captured, flagged as a non-solve, to execute for its honest - # reward.) - return CaptureOutcome(CaptureDecision.REWARD_HACK_NO_CAPTURE) - if capture_enabled and not is_current_task and goal_achieved: - # Loudly flag a success that cannot count: agents have burned - # whole sessions validating on a train task, believing they - # were done (run_20260707_112310 test task 0, session 3). - return CaptureOutcome(CaptureDecision.WRONG_TASK_NOTE) - return CaptureOutcome(CaptureDecision.NO_CAPTURE) diff --git a/predicators/agent_sdk/tools/clearance.py b/predicators/agent_sdk/tools/clearance.py deleted file mode 100644 index 33cf81bb41..0000000000 --- a/predicators/agent_sdk/tools/clearance.py +++ /dev/null @@ -1,193 +0,0 @@ -"""Robot-clearance measurement for the capture gate. - -Certification's validation rollouts sample the belief's own execution -variability, but not the gap between the belief's realized poses and -the real executor's: every move phase terminates anywhere within its -pose tolerance, grasp heights vary within the descend tolerance, and -placed objects land with scatter. A plan whose robot passes within -that slop of a bystander certifies 8/8 by luck and loses the ninth -draw for real (2026-09-02 bridge seed3: two belief-certified explore -plans died on 6.5 mm and 9.9 mm real contacts against a block every -rollout had cleared). - -:class:`RobotClearanceProbe` measures, along each validation rollout's -low-level trajectory, the minimum distance between the robot's links -and every body the executor would treat as a bystander - excluding the -held object and its welded attachments (the skill factory's own -planning-scene reconstruction decides that), the current option's -argument objects (a pick's fingers straddle their target by design), -the bodies the option's skill declares it contacts (a push strikes its -appliance's switch, a separate body: boil/fan 2026-09-02 refused every -plan at -3 to -14 mm against the switch being pushed) and the object -held when the option starts (a Place releases it and retreats past it -at a distance set by the grasp, not by the executor's slop: domino -2026-09-02 refused every plan at ~5 mm against the domino just placed). -The verdict compares that minimum against the executor's pose slop, -``sqrt(move_to_pose_tol)``, read from the plan's own skills, so the -bar is the executor's and not a domain constant. -""" -from __future__ import annotations - -import logging -import math -from typing import Any, List, Optional, Sequence, Tuple - -import pybullet as p - -from predicators.structs import _Option - -# Only bodies within this distance of a robot link are queried; anything -# farther cannot be the minimum that matters. -_QUERY_DISTANCE = 0.05 - - -def phase_skill_of(options: Sequence[_Option]) -> Optional[Any]: - """The first PhaseSkill (duck-typed) behind a grounded plan's options. - - A PhaseSkill's built option binds its policy as a bound method, so - the skill - with its planning simulator and executor tolerances - - is reachable through it. ``None`` when no option is skill-factory - built or none carries a planning simulator. - """ - for opt in options: - skill = getattr(getattr(opt.parent, "policy", None), "__self__", None) - config = getattr(skill, "_config", None) - if (skill is not None and hasattr(skill, "_sim_collision_context") - and getattr(config, "simulator", None) is not None): - return skill - return None - - -class RobotClearanceProbe: - """Minimum robot-to-bystander clearance over validation rollouts. - - Feed each executed option's low-level trajectory to - :meth:`observe`; :attr:`min_dist` and :attr:`where` then describe - the closest approach seen so far. ``stride`` subsamples trajectory - states (a contact event spans many control steps, so every third - state loses nothing the verdict needs); the final state of every - option is always probed - phase goals are where the executor's - refusals were measured. - """ - - def __init__(self, skill: Any, stride: int = 3) -> None: - self._skill = skill - self._stride = max(1, stride) - self.min_dist = math.inf - self.where = "" - self.num_probes = 0 - - @property - def threshold(self) -> float: - """The executor's pose slop: its move-phase terminal tolerance.""" - return float(math.sqrt(self._skill._config.move_to_pose_tol)) # pylint: disable=protected-access - - def observe(self, label: str, option: _Option, - states: Sequence[Any]) -> None: - """Probe one option's trajectory states (best-effort diagnostic).""" - exempt = {o.name for o in option.objects} - if states: - exempt |= self._designed_contacts(option, states[0]) - last = len(states) - 1 - for k, state in enumerate(states): - if k % self._stride and k != last: - continue - try: - dist, body = self._min_robot_clearance(state, exempt) - except Exception as e: # pylint: disable=broad-except - logging.debug("Clearance probe skipped a state: %s", e) - continue - self.num_probes += 1 - if dist < self.min_dist: - self.min_dist = dist - self.where = (f"{label}, step {option.name}" - f"({', '.join(o.name for o in option.objects)})" - f" vs {body}") - - def _designed_contacts(self, option: _Option, state: Any) -> set: - """Names of the bodies this option touches by design: the ones its - skill declares (a push's switch, see ``PhaseSkill.contact_objects``) - and the object it holds when it starts (a Place's fingers open around - the block they release, and their retreat clears it by the grasp - geometry, not by the executor's pose slop). - - Best-effort: a failing lookup exempts - nothing. - """ - names: set = set() - skill = getattr(getattr(option.parent, "policy", None), "__self__", - None) - contacts = getattr(skill, "contact_objects", None) - if contacts is not None: - try: - names |= {o.name for o in contacts(state, option.objects)} - except Exception as e: # pylint: disable=broad-except - logging.debug("Clearance probe: contact lookup failed: %s", e) - try: - held = self._held_object_name(state) - except Exception as e: # pylint: disable=broad-except - logging.debug("Clearance probe: held lookup failed: %s", e) - held = None - if held is not None: - names.add(held) - return names - - def _held_object_name(self, state: Any) -> Optional[str]: - """The name of the object held in ``state``, if any.""" - _, _, names, held, _ = self._skill._sim_collision_context(state) # pylint: disable=protected-access - return None if held is None else names.get(held) - - def _min_robot_clearance(self, state: Any, - exempt: set) -> Tuple[float, str]: - skill = self._skill - sim = skill._config.simulator # pylint: disable=protected-access - _, bodies, names, _, _ = skill._sim_collision_context(state) # pylint: disable=protected-access - robot = sim._pybullet_robot # pylint: disable=protected-access - robot.set_joints(state.joint_positions) - client = sim._physics_client_id # pylint: disable=protected-access - best = _QUERY_DISTANCE - best_body = "" - for body in bodies: - name = names.get(body, str(body)) - if name in exempt: - continue - points = p.getClosestPoints(robot.robot_id, - body, - _QUERY_DISTANCE, - physicsClientId=client) - if not points: - continue - dist = min(pt[8] for pt in points) - if dist < best: - best, best_body = dist, name - return best, best_body - - def verdict(self) -> Tuple[bool, str, str]: - """``(ok, summary line, detail)``. - - ``ok`` is False when the closest approach is inside the - executor's pose slop; ``detail`` then names it for the refusal - message (empty when ok or when nothing was probed). - """ - if self.num_probes == 0: - return True, "", "" - thr = self.threshold - if not math.isfinite(self.min_dist): - return True, (f"min robot clearance: >{_QUERY_DISTANCE * 1e3:.0f}" - " mm"), "" - summary = (f"min robot clearance: {self.min_dist * 1e3:.1f} mm " - f"({self.where}; executor pose slop {thr * 1e3:.0f} mm)") - if self.min_dist < thr: - return False, summary, ( - f"the robot passes within {self.min_dist * 1e3:.1f} mm of a " - f"bystander ({self.where}), inside the executor's " - f"{thr * 1e3:.0f} mm pose slop") - return True, summary, "" - - -def clearance_lines(probe: Optional[RobotClearanceProbe]) -> List[str]: - """Report lines for a probe, empty when no probe ran.""" - if probe is None: - return [] - _, summary, _ = probe.verdict() - return [summary] if summary else [] diff --git a/predicators/agent_sdk/tools/context.py b/predicators/agent_sdk/tools/context.py index e91941e226..fe957abbf8 100644 --- a/predicators/agent_sdk/tools/context.py +++ b/predicators/agent_sdk/tools/context.py @@ -9,26 +9,7 @@ from predicators.option_model import _OptionModelBase from predicators.settings import CFG from predicators.structs import CausalProcess, LowLevelTrajectory, \ - ParameterizedOption, ParameterizedSampler, Predicate, State, Task, Type - - -@dataclass(frozen=True) -class PlanCapture: - """A captured plan popped off a :class:`ToolContext` in one piece. - - Returned by :meth:`ToolContext.take_plan_capture` so consumers see - the four ``solved_plan*`` fields as the single value they are: - ``plan`` is falsy when nothing was captured. - """ - plan: Optional[Any] - sketch: Optional[Any] - reached_goal: Optional[bool] - eval_reward: Optional[float] - validation_summary: Optional[str] = None - # Closed-loop policy mode: the validated policy.py SOURCE snapshot - # (mutually exclusive with ``plan``). The capture is falsy when both - # are None. - policy_source: Optional[str] = None + ParameterizedOption, Predicate, State, Task, Type @dataclass @@ -69,10 +50,9 @@ class ToolContext: # need the real engine). A comparison arm lists what its surface # withholds; the probe raises on a listed call instead of serving it. probe_disabled: FrozenSet[str] = frozenset() - # Synthesis-session loaders behind ``sim.predicates()`` and - # ``sim.samplers()``: each reloads the agent-authored file fresh - # (predicates.py / samplers.py), installs the result into the - # approach so refinement sees the draft, and returns the report + # Synthesis-session loader behind ``sim.predicates()``: it reloads + # the agent-authored predicates.py fresh, installs the result into + # the approach so refinement sees the draft, and returns the report # text. Empty in sessions that do not offer the artifact. probe_artifact_loaders: Dict[str, Callable[..., @@ -111,18 +91,11 @@ class ToolContext: probe_score_provider: Optional[Callable[..., str]] = None # Active-experiment info-gain scorer, synced from the learning # approach when info-seeking exploration is on: - # ``(state, atoms) -> disagreement``. The agent_model_based explorer - # passes it into refinement so continuous-parameter search prefers - # candidates that straddle the learned model's decision boundaries. - # None ⇒ plain feasibility search (default). + # ``(state, atoms) -> disagreement``. The probe passes it into + # refinement so continuous-parameter search prefers candidates that + # straddle the learned model's decision boundaries. None ⇒ plain + # feasibility search (default). atom_disagreement_fn: Optional[Callable[[State, Any], float]] = None - # Synthesized per-skill samplers (option name -> sampler), synced from - # the learning approach when agent_sim_learn_parameterized_samplers is on. - # The agent_model_based explorer and synthesis tools pass these into - # refinement so continuous-parameter search aims at each step's subgoal - # instead of drawing uniformly. Empty ⇒ uniform sampling (default). - parameterized_samplers: Dict[str, ParameterizedSampler] = field( - default_factory=dict) current_task: Optional[Task] = None # The last real observation of the level in progress (continual # play: every env tool result and the session query refresh it), so @@ -153,8 +126,6 @@ class ToolContext: joint_draw_scope: Optional[Callable[[Dict[str, float]], Any]] = None episode_prefix_provider: Optional[Callable[[], Tuple[List[State], List[Any]]]] = None - skill_factory_context: Dict[str, Any] = field(default_factory=dict) - proposals_disabled: bool = False # set True during test-time solving log_dir: Optional[str] = None env: Optional[Any] = None # simulator env reference (for rendering) image_save_dir: Optional[str] = None # sandbox path for rendered images @@ -167,7 +138,7 @@ class ToolContext: # main.py's ``test_task_idx``. None outside the test phase. Threaded into # the saved session-log filename so test queries are attributable to a task. test_task_idx: Optional[int] = None - test_call_id: int = 0 # incremented per submit_plan call + test_call_id: int = 0 # incremented per probe rollout call # 0-based learning cycle (matching main.py's "ONLINE LEARNING CYCLE i"; # -1 = the offline pass) while a synthesis (learn) session is active, # None otherwise. Set/cleared around the synthesis query so tools that @@ -184,165 +155,49 @@ class ToolContext: # frozen for the session's lifetime. Subclasses set this before # opening a fresh session and clear it on close. extra_session_hooks: Dict[str, list] = field(default_factory=dict) - # Populated by AgentModelBasedExplorer so learning approaches can diff - # mental-model subgoals against real trajectories. - # TODO(sim-learning): consume these in learn_from_interaction_results. - last_sketch_subgoals: Optional[Any] = None - # Agent-session phase the tools are serving: "explore", "solve" - # (test-time) or "synthesis" (learn). Set by the session mixin when - # it builds the session; None before any session exists. + # Agent-session phase the tools are serving: "solve" or + # "synthesis" (the model-writing rounds). Set by the session mixin + # when it builds the session; None before any session exists. phase: Optional[str] = None # True when the approach will track the simulator's latent block at # execution (code_sim_learning.latent_tracker), so latent-reading # atoms are evaluable on real observations and refinement must keep # them as Wait targets; False (bare observations) strips them. latent_tracking_available: bool = False - last_sketch_options: Optional[Any] = None - # Set by AgentModelBasedExplorer per request: did the mental model reach - # the task goal during refinement? Read by get_interaction_requests to - # stamp InteractionRequest.mental_model_solved (None ⇒ no verdict). - last_mental_model_solved: Optional[bool] = None - # Sketch-line descriptions of the exploration plans already generated - # this online-learning cycle (a cycle's requests are all generated - # before any executes). Cleared by get_interaction_requests per cycle, - # appended by AgentModelBasedExplorer per request, and shown in the next - # explore prompt so the agent proposes a complementary plan instead of - # repeating the identical one for every request. - cycle_scheduled_plans: List[str] = field(default_factory=list) - # Digest of the latest rollout system-ID fit's weak spots - # (unexplainable segments, unidentified/insensitive params, - # cross-cycle conflicts), synced from the sim-learning approach. - # The agent_model_based explorer appends it to its experiment guidance - # so the next exploration targets the gaps. None ⇒ no fit ran yet - # (or it had no weak spots). - sysid_diagnostics: Optional[str] = None - # The natural-language world-model arm's document (world_model.md - # content) and its agent-visible path: the solve prompt and the - # model-free explorer quote it into every task message. Empty - # everywhere else. - world_model_notes: str = "" - world_model_notes_path: str = "" - # Set by submit_plan / submit_policy when a plan is verified - # to reach the goal on the CURRENT solve task: the simulator-verified plan - # (grounded options with found params) and the parallel subgoal sketch. - # The bilevel approach returns this directly instead of re-refining, so - # the agent's tool-validated answer is exactly what gets executed. None ⇒ - # nothing captured this query. - solved_plan: Optional[Any] = None - solved_sketch: Optional[Any] = None - # Whether the captured solved_plan counts as a validated solve in its - # belief-sim rollout(s): goal reached, evaluator-certified, and every - # validation rollout passed. False ⇒ it was a best-effort capture (see - # below). Cleared together with solved_plan. - solved_plan_reached_goal: Optional[bool] = None - # Gate for the above: only approaches that consume captured plans - # (AgentModelBasedApproach) set this True. Keeps the open-loop - # planner, which also uses submit_plan, from recording - # spurious captures. - capture_goal_reaching_plans: bool = False - # Set (with capture_goal_reaching_plans) only for the final-submission - # nudge after an attempt exhausted its turn budget: submit_plan - # then captures the agent's submitted plan on the current task even if it - # does not reach the goal, is scored a non-solve by the task evaluator, - # or is flaky, so the approach executes the best-effort plan (for its - # honest reward) instead of paying for another full-budget attempt. A - # best-effort capture never displaces a validated-solve capture. - capture_best_effort_plan: bool = False - # Fresh-physics scope for capture-validation rollouts: a callable - # returning a context manager. While entered, ``ctx.option_model`` - # simulates on a freshly constructed env instance instead of the shared - # session env, whose reset cannot reconstruct state exactly (solver - # warm-start state, velocity residuals), making repeated rollouts - # correlated with each other and systematically offset from the fresh - # real env. Accepts an optional ``physical_overrides`` keyword (a - # param-name -> value dict applied to the fresh env on top of the - # identified params) for the physics-margin rollouts. Installed by + # Fresh-physics scope for validation rollouts: a callable returning a + # context manager. While entered, ``ctx.option_model`` simulates on a + # freshly constructed env instance instead of the shared session env, + # whose reset cannot reconstruct state exactly (solver warm-start + # state, velocity residuals), making repeated rollouts correlated + # with each other and systematically offset from the fresh real env. + # Accepts an optional ``physical_overrides`` keyword (a param-name -> + # value dict applied to the fresh env on top of the identified + # params) for the physics-sweep rollouts. Installed by # AgentSimLearningApproach (see ``_fresh_validation_env_scope``); - # None ⇒ validation rollouts share the session env. Gated by + # None ⇒ probe rollouts share the session env. Gated by # agent_plan_validation_fresh_env. validation_env_scope: Optional[Callable[..., Any]] = None # Candidate-aware counterpart: loads the deployed candidate before # cloning its physics and rebinds its option model for the whole rollout. - # Never substitute the solve-time model for a synthesis candidate. + # Never substitute the deployed model for a synthesis candidate. probe_validation_env_scope: Optional[Callable[..., Any]] = None - # Physics-margin points for the capture gate: a zero-arg callable - # returning the current grid of perturbations spanning +-1 posterior - # sigma of the identified physical params (full override dicts, - # ascending; empty when no fit with nonzero posterior width is - # deployed). A callable rather than a stored list so the points - # always track the LATEST applied fit. Installed by - # AgentSimLearningApproach; consumed by submit_plan under - # agent_plan_validation_physics_margin and by the sim.run physics - # sweep. + # Physics-margin points: a zero-arg callable returning the current grid + # of perturbations spanning +-1 posterior sigma of the identified + # physical params (full override dicts, ascending; empty when no fit + # with nonzero posterior width is deployed). A callable rather than a + # stored list so the points always track the LATEST applied fit. + # Installed by AgentSimLearningApproach; consumed by the sim.run + # physics sweep. physics_margin_provider: Optional[Callable[[], List[Dict[str, float]]]] = None - # Rule-parameter margin points for the capture gate: a zero-arg - # callable returning the calibrated rule-parameter ensemble (full - # fitted-param dicts drawn from the fit posterior; the same members - # info-seeking exploration scores with). The gate re-rolls a - # capture-eligible submission under each member so a plan that - # survives only at the point estimate of an uncertain learned - # constant is rejected as PARAM-SENSITIVE. Installed by - # AgentSimLearningApproach; consumed under - # agent_plan_validation_rule_param_margin. - # How the capture gate names one rule-param margin point and the - # set it came from in its reports. The program-world-model arm - # sweeps belief particles over the model's hidden state through - # the same gate and relabels them here. - rule_param_margin_label: str = "rule-param ensemble member" - rule_param_margin_note: str = ( - "calibrated posterior members of the learned rule parameters") - rule_param_margin_provider: Optional[Callable[[], - List[Dict[str, - float]]]] = None - # Context manager applying one rule-parameter override dict for the - # duration of a validation rollout: swap-and-restore of the live - # fitted-params mapping that the learned rules and frozen predicate - # classifiers read through (the score_atom_disagreement pattern). - # The gate enters it BEFORE the fresh-env scope so anything bound at - # env construction sees the override too. - rule_param_override_scope: Optional[Callable[[Dict[str, float]], - Any]] = None - # Capture-task keys (see ``_capture_task_key``) that have produced a - # FLAKY rejection in submit_plan. A flaky submission is direct - # evidence the agent is tuning in a marginal region where a lucky - # streak can pass the base rollout gate (run_20260717_182321: a - # 20/20-swept placement validated 3/3, then failed the real episode), - # so subsequent captures on these tasks must clear the escalated - # agent_plan_validation_rollouts_after_flaky gate instead. - flaky_capture_task_keys: Set[Any] = field(default_factory=set) - # Task-evaluator reward of the rollout that produced the current - # solved_plan capture (None when no evaluator verdict was computed). - # The restart loop ranks best-effort captures across attempts by it. - # Cleared together with solved_plan. - solved_plan_eval_reward: Optional[float] = None - # One-line record of the capture-time validation outcome (rollout - # tally, first failing step, physics-margin tally), for the journal - # auto-entry. Cleared together with solved_plan. - solved_plan_validation_summary: Optional[str] = None - # Closed-loop policy mode (CFG.agent_solve_policy_mode): the captured - # policy.py source, SNAPSHOTTED at submit_policy call time so a - # later edit of the file cannot swap unvalidated code into the - # executed artifact. Mutually exclusive with solved_plan; cleared - # together with it. - solved_policy_source: Optional[str] = None - # True while the current solve attempt's deliverable is a policy: - # submit_plan keeps its probing role but its CAPTURE gate is - # disabled, and submit_policy requires it. Set by _solve_attempt. - policy_capture_mode: bool = False - # Restart-loop attempt bookkeeping, set by AgentModelBasedApproach._solve - # around each attempt. ``attempt_start``/``attempt_deadline`` are - # time.monotonic() values; the deadline is enforced cooperatively by - # the probe (every sim call) and run_python, and surfaced in tool - # results as a budget footer. None ⇒ no attempt in flight / no wall - # clock. The deadline is cleared before the final-submission nudge so - # nothing blocks the submission itself. - attempt_index: int = 0 + # Round bookkeeping (a continual play round is one attempt): + # ``attempt_start`` is a time.monotonic() value, surfaced in tool + # results as the budget footer's elapsed time. None ⇒ no round in + # flight. attempt_start: Optional[float] = None - attempt_deadline: Optional[float] = None - # Count of full-plan belief-sim rollouts this attempt (probe runs, - # trials, capture-validation repeats). Reset per attempt; shown in - # the budget footer so sweeps carry a visible price. + # Count of full-plan belief-sim rollouts this round (probe runs, + # trials). Reset per round; shown in the budget footer so sweeps + # carry a visible price. attempt_rollout_count: int = 0 # The run's conversation as the play tools show it in the [context] # line (continual protocol): the prompt size of the latest assistant @@ -354,22 +209,14 @@ class ToolContext: context_turns: int = 0 context_compactions: int = 0 context_window_tokens: Optional[int] = None - # Best submission on the current task this attempt that - # submit_plan evaluated but refused to capture (evaluator - # scored it a non-solve, or it was flaky), ranked by evaluator - # reward. Reset per attempt; the journal auto-entry records it so a - # later attempt (or the final best-effort nudge) can resubmit it - # instead of the attempt's work vanishing with its context. - best_uncaptured_plan_lines: Optional[List[str]] = None - best_uncaptured_reward: Optional[float] = None # Per-call deadline for the run_python call currently executing - # (agent_sdk_python_call_timeout); enforced at the same - # probe checkpoints as attempt_deadline. None ⇒ no call in flight. + # (agent_sdk_python_call_timeout); enforced at the probe's + # checkpoints. None ⇒ no call in flight. python_call_deadline: Optional[float] = None # Adaptive info-seeking trigger (agent_explorer_info_seeking_adaptive): - # set True the first time submit_plan's rule-param margin gate refuses - # a plan as PARAM-SENSITIVE, cleared when a plan is captured. While - # True the proactive info-seeking apparatus (suggest_probes ranking, + # set True the first time the probe's physics sweep finds a plan's + # success straddling the belief interval. While True the proactive + # info-seeking apparatus (suggest_probes ranking, # disagreement guidance) is active; while False, and # under the adaptive flag, it stays dormant so easy levels pay no # info-seeking step tax. Ignored unless the adaptive flag is on. @@ -387,10 +234,11 @@ def info_seeking_active(self) -> bool: Off when info-seeking exploration is disabled outright. On whenever it is enabled and the adaptive flag is off (the original always-on behaviour). Under the adaptive flag it turns - on only once the capture gate has refused a plan as PARAM- - SENSITIVE this run (``param_sensitive_refusal_pending``), so the - agent spends real steps reducing uncertainty only after a - fragile plan has actually been caught. + on only once the probe's physics sweep has found a plan whose + success straddles the belief interval this run + (``param_sensitive_refusal_pending``), so the agent spends real + steps reducing uncertainty only after a fragile plan has + actually been found. """ if not CFG.agent_explorer_info_seeking: return False @@ -413,84 +261,26 @@ def note_stream_entry(self, entry: Dict[str, Any]) -> None: elif kind == "system" and entry.get("subtype") == "compact_boundary": self.context_compactions += 1 - def begin_attempt(self, index: int, wall_clock: float) -> None: - """Start restart-loop bookkeeping for solve attempt ``index``. - - Resets everything scoped to a single attempt (rollout count, - best refused submission) and arms the wall-clock deadline - (``wall_clock <= 0`` ⇒ no deadline). The matching teardown stays - in ``AgentModelBasedApproach._solve``'s finally block, - interleaved with its journal write. - """ - self.attempt_index = index + def begin_attempt(self) -> None: + """Start a round's bookkeeping: its clock and rollout count.""" self.attempt_rollout_count = 0 - self.best_uncaptured_plan_lines = None - self.best_uncaptured_reward = None self.attempt_start = time.monotonic() - self.attempt_deadline = (self.attempt_start + - wall_clock if wall_clock > 0 else None) def pause_attempt_clock(self, seconds: float) -> None: """Push every armed wall-clock mark ``seconds`` into the future. Called by the session manager after it slept out a usage limit, - so the wait is charged to neither the attempt's budget nor the - run_python call in flight, and the budget footer's elapsed time - stays honest. + so the wait is charged to neither the round nor the run_python + call in flight, and the budget footer's elapsed time stays + honest. """ if seconds <= 0: return if self.attempt_start is not None: self.attempt_start += seconds - if self.attempt_deadline is not None: - self.attempt_deadline += seconds if self.python_call_deadline is not None: self.python_call_deadline += seconds - def clear_plan_capture(self) -> None: - """Clear the four ``solved_plan*`` fields together. - - They form one value (see :class:`PlanCapture`); clearing any of - them individually would leave a stale mix. - """ - self.solved_plan = None - self.solved_sketch = None - self.solved_plan_reached_goal = None - self.solved_plan_eval_reward = None - self.solved_plan_validation_summary = None - self.solved_policy_source = None - - def take_plan_capture(self) -> PlanCapture: - """Pop the captured plan, clearing it so it cannot be reused. - - The returned capture's ``plan`` is falsy when nothing was - captured since the last clear. - """ - capture = PlanCapture( - plan=self.solved_plan, - sketch=self.solved_sketch, - reached_goal=self.solved_plan_reached_goal, - eval_reward=self.solved_plan_eval_reward, - validation_summary=self.solved_plan_validation_summary, - policy_source=self.solved_policy_source) - self.clear_plan_capture() - return capture - - -def _capture_task_key(ctx: ToolContext) -> Any: - """Stable identity of the task behind a ``task_idx="current"`` capture. - - Keys ``ctx.flaky_capture_task_keys`` so a FLAKY rejection escalates - the validation gate for later submissions on the SAME task only. - Test-time solves are keyed by the test task index (stable across the - sessions and replans of one task); exploration/synthesis captures - fall back to the learning iteration, which at worst escalates - conservatively across that cycle's tasks. - """ - if ctx.test_task_idx is not None: - return ("test", ctx.test_task_idx) - return ("iter", ctx.iteration_id) - @contextmanager def decorrelated_rollout_seed(rollout_idx: int) -> Iterator[None]: diff --git a/predicators/agent_sdk/tools/exploration.py b/predicators/agent_sdk/tools/exploration.py index af0a6dccd6..099520a969 100644 --- a/predicators/agent_sdk/tools/exploration.py +++ b/predicators/agent_sdk/tools/exploration.py @@ -184,9 +184,9 @@ def belief_probe_blurb(synthesis_probe: bool, run_desc = ( "`sim.run(plan_text, render=True, draws=None, contacts=False, " "physics_sweep=False, seed=None)` rehearses an option plan " - "FROM THE CURRENT STATE (same grammar as submit_plan) on the " - f"{int(CFG.belief_joint_draws)} joint draws of the belief - a " - "parameter draw, a state draw and the memory those parameters " + "FROM THE CURRENT STATE (same grammar as skills_execute_plan) on " + f"the {int(CFG.belief_joint_draws)} joint draws of the belief - " + "a parameter draw, a state draw and the memory those parameters " "imply, each on a fresh env with its own planner seed - and " "reports the success estimate P-hat with its standard error, " "scored by the TASK EVALUATOR on the episode so far followed " @@ -207,7 +207,7 @@ def belief_probe_blurb(synthesis_probe: bool, "`sim.refine(sketch_text, timeout=60, require_goal=False, " "require_solved=False)` runs " "backtracking parameter search from the belief mean (same " - "grammar as submit_plan" + "grammar as skills_execute_plan" f"{_region_syntax_blurb()}; " "success = each step establishes its `-> {subgoals}` " "annotation), scores up to " @@ -228,13 +228,13 @@ def belief_probe_blurb(synthesis_probe: bool, "your sketch forward on your own parameters and, per `-> " "{subgoals}`-annotated step with continuous params, ranks " "feasible alternatives by the learned model's ensemble " - "disagreement on those atoms (advice only: what you submit " + "disagreement on those atoms (advice only: what you execute " "runs as written); " if surface.uncertainty else "") run_desc = ( "`sim.run(plan_text, render=True, trials=1, solved=False, " "contacts=False)` executes an option " "plan FROM THE CURRENT " - "STATE (same grammar as submit_plan; print the result " + "STATE (same grammar as skills_execute_plan; print the result " "for per-step outcomes incl. saved per-step scene-image paths - " "view them with the Read tool; pass render=False inside tight " "sweep loops) and advances the state; `-> {subgoals}` " @@ -251,13 +251,13 @@ def belief_probe_blurb(synthesis_probe: bool, "unmodified reset() state) also scores each trial with the " "TASK EVALUATOR (per-trial solved/reward) - reaching the goal " "atoms is NOT the same as being scored a solve, so check this " - f"BEFORE submitting; {belief_desc}" + f"BEFORE executing the plan; {belief_desc}" "contacts=True (single run) reports, per ") refine_desc = ( "`sim.refine(sketch_text, timeout=60, require_goal=False, " "require_solved=False)` runs " "backtracking parameter search FROM THE CURRENT STATE (same " - "grammar as submit_plan" + "grammar as skills_execute_plan" f"{_region_syntax_blurb()}; " "success = each step establishes its `-> {subgoals}` " "annotation, and the result's Verdict line states what it " @@ -315,10 +315,8 @@ def _build_exploration_tools(ctx: ToolContext, _text_result: Callable, The namespace is the probe facade, numpy, and the collected real trajectories as read-only evidence (see ``build_probe_namespace`` - nothing evaluator-shaped beyond the probe's gated paths): the probe - reuses the exact machinery behind ``submit_plan`` (same - plan grammar, same option-model executor, same renderer) but - carries no scoring surface - nothing run here can be captured as - the answer, so it is safe to hand the agent as a freely composable + carries no scoring surface and nothing run here acts in the + environment, so it is safe to hand the agent as a freely composable physics probe. Synthesis sessions attach their own ``run_python`` (fit data + the candidate-simulator probe in one namespace; see ``_get_synthesis_tool_names``), and ``create_mcp_tools`` skips this @@ -328,12 +326,9 @@ def _build_exploration_tools(ctx: ToolContext, _text_result: Callable, # pylint: disable-next=import-outside-toplevel from predicators.agent_sdk.belief_probe import build_probe_namespace - submit_desc = ( - "EXPLORATORY " - "ONLY: nothing run here is captured as your answer - preview " - "the evaluator's verdict with sim.run(solved=True), then " - "validate and submit the final plan via submit_plan " - "from the true initial state.") + exploratory_desc = ( + "EXPLORATORY ONLY: nothing run here acts in the environment; " + "preview the evaluator's verdict with sim.run(solved=True).") run_python = _make_python_exec_tool( tool, name="run_python", @@ -359,8 +354,9 @@ def _build_exploration_tools(ctx: ToolContext, _text_result: Callable, "for sim-free code; printed output up to the stop is " "returned): budget sweeps accordingly - " "prefer coarse-to-fine over exhaustive grids, and print " - "intermediate bests so partial results survive a stop. " if - surface_cfg.python_call_timeout > 0 else "") + f"{submit_desc}"), + "intermediate bests so partial results survive a stop. " + if surface_cfg.python_call_timeout > 0 else "") + + f"{exploratory_desc}"), exec_ns=build_probe_namespace(ctx), sandbox_dir=ctx.sandbox_dir, text_result=_text_result, diff --git a/predicators/agent_sdk/tools/python_exec.py b/predicators/agent_sdk/tools/python_exec.py index d315a49b7c..aa79ccc7fd 100644 --- a/predicators/agent_sdk/tools/python_exec.py +++ b/predicators/agent_sdk/tools/python_exec.py @@ -151,15 +151,6 @@ async def python_exec(args: Dict[str, Any]) -> Dict[str, Any]: rollouts_before = 0 if budget_ctx is not None: rollouts_before = budget_ctx.attempt_rollout_count - attempt_dl = budget_ctx.attempt_deadline - if (attempt_dl is not None and time.monotonic() > attempt_dl - and not budget_ctx.capture_best_effort_plan): - return text_result( - "The attempt's wall-clock exploration budget is " - "exhausted - this call was not run. Submit your single " - "best plan NOW via submit_plan on the current " - "task (omit task_idx)." + - _budget_footer(budget_ctx, rollouts_before)) call_timeout = ToolSurfaceConfig.from_cfg().python_call_timeout if budget_ctx.probe_option_model_provider is not None: # Synthesis sessions probe the CANDIDATE simulator, whose @@ -184,9 +175,6 @@ def _footer() -> str: if budget_ctx is not None: if budget_ctx.python_call_deadline is not None: wd_deadlines.append(budget_ctx.python_call_deadline) - if (budget_ctx.attempt_deadline is not None - and not budget_ctx.capture_best_effort_plan): - wd_deadlines.append(budget_ctx.attempt_deadline) if call_timeout_s is not None and call_timeout_s > 0: wd_deadlines.append(time.monotonic() + call_timeout_s) if wd_deadlines: diff --git a/predicators/agent_sdk/tools/registry.py b/predicators/agent_sdk/tools/registry.py index b4ef0ff788..6932f3b75b 100644 --- a/predicators/agent_sdk/tools/registry.py +++ b/predicators/agent_sdk/tools/registry.py @@ -20,31 +20,22 @@ "TaskList", ] -TESTING_TOOL_NAMES = [ - "submit_plan", - # Closed-loop policy mode (agent_solve_policy_mode): validates and - # captures the agent-written policy.py. Only offered on solve - # rosters when the mode is on. - "submit_policy", -] -# The one code-execution tool. Solve sessions get the static instance +# The one code-execution tool. Play sessions get the static instance # built by ``create_mcp_tools`` (namespace = the BeliefProbe facade over # the deployed belief model, predicators/agent_sdk/belief_probe.py); # synthesis sessions attach their own instance under the same name # (fit data + the probe over the candidate simulator), which replaces -# the static one at assembly. Offered to every session that has a -# simulator to probe (see ``AgentModelFreeApproach._get_solve_tool_names``). +# the static one at assembly. EXPLORATION_TOOL_NAMES = [ "run_python", ] -ALL_TOOL_NAMES = TESTING_TOOL_NAMES + EXPLORATION_TOOL_NAMES +ALL_TOOL_NAMES = list(EXPLORATION_TOOL_NAMES) # Name of the tool ``create_synthesis_tools`` builds for a synthesis # session (the same ``run_python`` name as the solve-phase instance - # see EXPLORATION_TOOL_NAMES). ``tests/agent_sdk/test_tool_registry.py`` -# asserts that the factory output matches this tuple. Predicate and -# sampler drafts are loaded through the probe (``sim.predicates()`` / -# ``sim.samplers()``), not through tools. +# asserts that the factory output matches this tuple. Predicate drafts +# are loaded through the probe (``sim.predicates()``), not through tools. SYNTHESIS_TOOL_NAMES = ("run_python", ) diff --git a/predicators/agent_sdk/tools/sampler_synthesis.py b/predicators/agent_sdk/tools/sampler_synthesis.py deleted file mode 100644 index 0cbbc231bc..0000000000 --- a/predicators/agent_sdk/tools/sampler_synthesis.py +++ /dev/null @@ -1,172 +0,0 @@ -"""The ``sim.samplers()`` loader for sampler-synthesis sessions.""" -from typing import Any, Callable, Dict, List, Optional, Tuple - -import numpy as np - -from predicators.agent_sdk.proposal_exec import build_exec_context, \ - load_learned_samplers -from predicators.agent_sdk.synthesis_backend import SamplerSynthesisBackend -from predicators.agent_sdk.tools.params_view import _ParamsView -from predicators.agent_sdk.tools.sandbox_guard import _scrub_host_paths -from predicators.agent_sdk.tools.snapshots import _ArtifactSnapshotter -from predicators.settings import CFG -from predicators.structs import Object - - -def make_sampler_loader( - samplers_file: str, - samplers_versions_dir: str, - approach: SamplerSynthesisBackend, - cycle_index_provider: Optional[Callable[[], int]] = None, -) -> Callable[[], str]: - """Build the ``sim.samplers()`` loader for one synthesis session. - - On each call the loader loads - ``samplers.py`` fresh (snapshotting into ``samplers_versions_dir``), - validates the ``LEARNED_SAMPLERS`` dict (option name -> callable), - installs it into ``approach._synthesized_samplers`` so refinement - uses it, and reports a per-option shape/in-box sanity check. - - Args: - samplers_file: Host path to the agent-edited ``samplers.py``. - samplers_versions_dir: Directory for per-call snapshots. - approach: The ``AgentSimLearningApproach`` instance. - cycle_index_provider: Returns the current 0-based cycle - (negative = the offline pass, rendered as ``offline``). - """ - # pylint: disable=import-outside-toplevel - import traceback # pylint: disable=redefined-outer-name,reimported - - from predicators.code_sim_learning.fit_space import ParamSpec - - # pylint: enable=import-outside-toplevel - _snapshotter = _ArtifactSnapshotter( - live_file=samplers_file, - versions_dir=samplers_versions_dir, - artifact_name="samplers", - cycle_index_provider=cycle_index_provider, - missing_file_hint=("Use Write to create it with " - "LEARNED_SAMPLERS = {\"OptionName\": fn, ...}."), - ) - params_view = _ParamsView(approach._fitted_params) # pylint: disable=protected-access - - def _snapshot_and_load_samplers( - path: str, - ) -> Tuple[Dict[str, Any], Optional[str], Optional[str], List[str]]: - """Snapshot ``path`` then exec it into a fresh namespace. - - Returns ``(samplers, version_tag, error_msg, warnings)``. - Entries keyed by an unknown option name, or whose value is not - callable, are skipped and described in ``warnings``. On success, - mutates ``approach._synthesized_samplers`` to the validated - dict. - """ - raw, version_tag, err = _snapshotter.snapshot(path) - if err is not None: - return {}, None, err, [] - assert raw is not None and version_tag is not None - - ctx = build_exec_context( - types=approach._types, # pylint: disable=protected-access - predicates=approach._get_all_predicates(), # pylint: disable=protected-access - options=approach._get_all_options(), # pylint: disable=protected-access - extra_context={ - "params": params_view, - "ParamSpec": ParamSpec, - }) - option_names = {o.name for o in approach._get_all_options()} # pylint: disable=protected-access - valid, warnings, err = load_learned_samplers(raw.decode("utf-8"), ctx, - option_names) - if err is not None: - return {}, version_tag, (f"[{version_tag}] Error executing " - f"{path}:\n{err}"), [] - - # Mutate approach state so sim.refine / test-time - # refinement draw from the agent's draft samplers. - approach._synthesized_samplers = valid # pylint: disable=protected-access - return valid, version_tag, None, warnings - - def _sanity_check(name: str, fn: Any) -> str: - """Draw a few params from a representative state; report shape/box.""" - # pylint: disable=protected-access - options_by_name = {o.name: o for o in approach._get_all_options()} - opt = options_by_name[name] - train_tasks = approach._train_tasks - if not train_tasks: - return f" {name}: no train task to sanity-check against." - state = train_tasks[0].init - # Pick the first object of each option-arg type present in the state. - objs: List[Object] = [] - for t in opt.types: - match = next((o for o in state if o.type.name == t.name), None) - if match is None: - return (f" {name}: no object of type '{t.name}' in the " - "train-task state to sanity-check against.") - objs.append(match) - box = opt.params_space - expected = box.shape[0] - rng = np.random.default_rng(CFG.seed) - in_box = 0 - n_draws = 3 - for _ in range(n_draws): - try: - raw = fn(state, set(), rng, objs) - arr = np.asarray(raw, dtype=np.float32).reshape(-1) - except Exception: # pylint: disable=broad-except - last = traceback.format_exc().strip().splitlines()[-1] - return (f" {name}: ERROR — sampler raised: {last} " - "(note: this check, and refinement at steps with " - "no subgoal annotation, call the sampler with " - "subgoal_atoms=set(); it must not crash on an " - "empty set — fall back to a default or uniform " - "draw).") - if arr.shape != (expected, ): - return (f" {name}: ERROR — returned shape {arr.shape}, " - f"expected ({expected},).") - if bool(np.all(arr >= box.low - 1e-6)) and \ - bool(np.all(arr <= box.high + 1e-6)): - in_box += 1 - return (f" {name}: OK — {n_draws} draws, {in_box}/{n_draws} " - f"within the params box.") - - def sampler_report() -> str: - """Reload samplers.py, install LEARNED_SAMPLERS, and report the per- - option sanity check.""" - try: - samplers, version_tag, err, warnings = ( - _snapshot_and_load_samplers(samplers_file)) - except Exception: # pylint: disable=broad-except - return (f"Error loading samplers.py:\n" - f"{_scrub_host_paths(traceback.format_exc())}") - - if err is not None: - return err - - prefix = f"[{version_tag}]" - lines = [ - f"{prefix} Sampler report — {len(samplers)} per-skill " - f"sampler(s) installed.", - ] - if warnings: - lines.append("") - lines.append("Warnings (entries skipped during load):") - for w in warnings: - lines.append(f" - {w}") - - if not samplers: - lines.append("") - lines.append("LEARNED_SAMPLERS is empty — add " - "{\"OptionName\": fn} entries to samplers.py.") - return "\n".join(lines) - - lines.append("") - lines.append("Sanity check (representative train-task state):") - for name in sorted(samplers): - lines.append(_sanity_check(name, samplers[name])) - lines.append("") - lines.append("Now call sim.refine with a sketch that " - "uses these options to measure the samples-to-refine " - "improvement.") - return "\n".join(lines) - - return sampler_report diff --git a/predicators/agent_sdk/tools/synthesis.py b/predicators/agent_sdk/tools/synthesis.py index ee53e936a2..f5610ac365 100644 --- a/predicators/agent_sdk/tools/synthesis.py +++ b/predicators/agent_sdk/tools/synthesis.py @@ -546,9 +546,6 @@ def _evaluate_rollout_fit(rules: list, sse=float(outcome.pre_sse), pinned=True, coverage=(0, len(rollouts))) - if hasattr(approach, "_record_sysid_diagnostics"): - approach._record_sysid_diagnostics( # pylint: disable=protected-access - {}, physical_names, 0, len(rollouts), outcome.traj_rms) rms_str = ", ".join(f"{r:.4g}" for r in outcome.traj_rms) trim_threshold = (CFG.code_sim_learning_rollout_trim_rms_factor * DEFAULT_NOISE_SIGMA) @@ -591,9 +588,9 @@ def _evaluate_rollout_fit(rules: list, sse=post_sse, applied_physical=dict(applied), coverage=(outcome.num_survivors, len(rollouts)), - # Physics-margin points for the capture gate, restored - # when this fit is deployed as the cycle's model; the - # joint belief replaces them with its own draws. + # The physics sweep's +-1-sigma grid, restored when this + # fit is deployed as the cycle's model; under the joint + # belief the sweep reads the belief's interval ends. sigma_points=(physics_sigma_points( applied, ident_report, @@ -601,10 +598,6 @@ def _evaluate_rollout_fit(rules: list, num_points=CFG.agent_plan_validation_physics_margin_points) if outcome.belief is None else []), belief=outcome.belief) - if hasattr(approach, "_record_sysid_diagnostics"): - approach._record_sysid_diagnostics( # pylint: disable=protected-access - ident_report, physical_names, outcome.num_survivors, - len(rollouts), outcome.traj_rms) kept_at_init = sorted(n for n in physical_names if applied[n] != fitted[n]) # Like-for-like SSE headline: the % reduction is measured on the @@ -751,8 +744,8 @@ def _evaluate_rollout_fit(rules: list, "Applied to the planning base env: the most likely value of " "every parameter the data moved (identified, weakly " "identified or wide posterior). " + kept_note + - "submit_plan and sim.run(plan, physics_sweep=True) certify " - "a plan across each deployed parameter's belief interval; " + "sim.run(plan, physics_sweep=True) certifies a plan " + "across each deployed parameter's belief interval; " "a plan that passes only part of an interval is reported " "with the passing and failing ranges, which is the cue " "that one real experiment narrowing that parameter is " @@ -787,7 +780,8 @@ def _evaluate_rollout_fit(rules: list, surface = probe_surface or ProbeSurface() if surface.fit: protocol = ( - " Nothing the probe runs is captured; the validation protocol " + " Nothing the probe runs acts in the environment; the " + "validation protocol " "before declaring the simulator done is: `sim.fit()` (canonical " "fit report), `sim.refine(plan, require_goal=True)` (params " "exist that reach each subgoal), then a continuous `sim.run` of " @@ -795,7 +789,8 @@ def _evaluate_rollout_fit(rules: list, "means a rule is more permissive than the data).") elif surface.edit_model: protocol = ( - " Nothing the probe runs is captured; the validation protocol " + " Nothing the probe runs acts in the environment; the " + "validation protocol " "before relying on the simulator is: `sim.validate()` (does the " "model explain the recordings at its declared values), " "`sim.refine(plan, require_goal=True)` (params exist that reach " @@ -804,7 +799,8 @@ def _evaluate_rollout_fit(rules: list, "permissive than the data).") else: protocol = ( - " Nothing the probe runs is captured; the rehearsal protocol " + " Nothing the probe runs acts in the environment; the " + "rehearsal protocol " "before acting is: `sim.refine(plan, require_goal=True)` " "(params exist that reach each subgoal), then a continuous " "`sim.run` of the refined plan.") diff --git a/predicators/agent_sdk/tools/tasks.py b/predicators/agent_sdk/tools/tasks.py deleted file mode 100644 index 133b38f751..0000000000 --- a/predicators/agent_sdk/tools/tasks.py +++ /dev/null @@ -1,59 +0,0 @@ -"""Shared resolution of a tool's ``task_idx`` argument to a task. - -Every task-scoped tool follows the same convention: an int ``task_idx`` -indexes the train tasks (bounds-checked), and omitting it falls back to -the current solve/explore task. :func:`_resolve_task` is the single -implementation of that convention. -""" -from dataclasses import dataclass -from typing import Any, Dict, Optional, Tuple, Union - -from predicators.agent_sdk.tools.context import ToolContext -from predicators.agent_sdk.tools.results import _error_result -from predicators.structs import Task - - -@dataclass(frozen=True) -class ResolvedTask: - """A tool call's resolved task plus how to refer to it. - - ``label`` follows the tools' display convention: the int train-task - index, or the string ``"current"`` for the current solve/explore - task. Handlers interpolate it directly into report text and pass it - to ``_resolve_task_evaluator``; ``is_current`` is the boolean the - capture guards read (it replaces the old ``task_idx == "current"`` - comparisons). - """ - task: Task - label: Union[int, str] - is_current: bool - - @property - def description(self) -> str: - """Human-readable phrase: ``train task 3`` or ``current task``.""" - if self.is_current: - return "current task" - return f"train task {self.label}" - - -def _resolve_task( - ctx: ToolContext, task_idx: Optional[int] -) -> Tuple[Optional[ResolvedTask], Optional[Dict[str, Any]]]: - """Resolve a tool's ``task_idx`` argument (None ⇒ current task). - - Returns ``(resolved, error)`` with exactly one of the two set; - ``error`` is a ready-to-return tool error result. - """ - if task_idx is not None: - if task_idx < 0 or task_idx >= len(ctx.train_tasks): - return None, _error_result( - f"Invalid task_idx {task_idx}. " - f"Available: 0-{len(ctx.train_tasks)-1}") - return ResolvedTask(task=ctx.train_tasks[task_idx], - label=task_idx, - is_current=False), None - if ctx.current_task is not None: - return ResolvedTask(task=ctx.current_task, - label="current", - is_current=True), None - return None, _error_result("No task_idx provided and no current_task set.") diff --git a/predicators/agent_sdk/tools/testing.py b/predicators/agent_sdk/tools/testing.py deleted file mode 100644 index fec44e9e72..0000000000 --- a/predicators/agent_sdk/tools/testing.py +++ /dev/null @@ -1,1597 +0,0 @@ -"""Testing tools, including the submit_plan capture surface.""" -import contextlib -import functools -import logging -import os -from typing import Any, Callable, Dict, List, Optional, Sequence, Set, Tuple - -import numpy as np - -from predicators import utils -from predicators.agent_sdk import bilevel_sketch -from predicators.agent_sdk.config import RefinementConfig, ValidationConfig -from predicators.agent_sdk.parallel_rollouts import \ - prefetch_parallel as _prefetch_parallel -from predicators.agent_sdk.tools.budget import _budget_footer -from predicators.agent_sdk.tools.capture import BestEffortReason, \ - CaptureDecision, _decide_capture -from predicators.agent_sdk.tools.clearance import RobotClearanceProbe, \ - phase_skill_of -from predicators.agent_sdk.tools.context import ToolContext, \ - _capture_task_key, decorrelated_rollout_seed -from predicators.agent_sdk.tools.results import _error_result -from predicators.agent_sdk.tools.scene import format_object_poses, \ - render_pybullet_image -from predicators.agent_sdk.tools.tasks import _resolve_task -from predicators.agent_sdk.tools.verdicts import _EvalStateCollector, \ - _format_evaluator_verdict, _resolve_task_evaluator, _sandbox_base, \ - evaluate_states_with, load_ground_sampler_fns -from predicators.code_sim_learning.identifiability import straddle_summary -from predicators.settings import CFG -from predicators.structs import GroundAtom, State, Task - -# Ceiling on agent-requested validation rollouts per submission -# (validation_rollouts): the agent pays for rollouts from its budget, but -# a typo'd request should not silently torch it. -_MAX_REQUESTED_ROLLOUTS = 25 - - -def _missing_goal_atoms(task: Task, state: State) -> Set[GroundAtom]: - """Goal atoms that do NOT hold in ``state`` by their own classifiers. - - Evaluated per atom with the goal predicates' own classifiers (the - same ones ``goal_holds`` runs), never by abstracting the state with - the agent's predicate set: the env's goal predicates are not in - that set under predicate invention, so every goal atom then read - as missing whenever the goal was not reached - including the ones - that held - and one agent concluded the goal atoms could never be - made True in the belief and abandoned a working route - (2026-08-27 bridge policy seed 0, cycle 3). - """ - return {a for a in task.goal if not a.holds(state)} - - -def _policy_source_path(ctx: ToolContext) -> Optional[str]: - """Host path of the agent-editable ``policy.py`` (policy mode).""" - base = _sandbox_base(ctx) - if not base: - return None - return os.path.join(base, "policy.py") - - -def _parameter_margin_sweep( - ctx: ToolContext, validation_cfg: ValidationConfig, - fresh_scope: Callable[..., Any], rollout: Callable[[], Tuple[bool, - str]], - subject: str) -> Tuple[List[str], Optional[str], str]: - """Margin sweep over BOTH parameter-uncertainty sources of one. - - capture-eligible submission - the single code path behind the - physics-margin and rule-parameter gates of ``submit_plan`` - and ``submit_policy``. - - The execution repeats before this all run AT the fitted parameters, - so they cannot see a submission whose success band excludes the - fit's own error (run_20260723_091108: a capture validated 8/8 at - fitted lateral_friction 0.5319 failed deterministically at true - 0.5). Two sources express that error: - - * identified PHYSICAL params: the fit posterior's sigma grid, - applied as construction overrides on a fresh env (perturbing the - shared env would leak into later tool calls), at the BASE planner - seed so a failure is attributable to the perturbation alone; - * learned RULE params: the calibrated posterior ensemble (the same - members info-seeking exploration scores with), applied by - swapping the live fitted-params view - entered BEFORE the fresh - env so values bound at construction also see the member. - - ``rollout`` runs one validation rollout and returns ``(ok, why)``. - Returns ``(outcome lines, param-sensitive detail or None, suffix - for the validation note)``; any failing point sets the detail, - which rejects the submission as PARAM-SENSITIVE. - """ - outcomes: List[str] = [] - detail: Optional[str] = None - note = "" - if (validation_cfg.physics_margin - and ctx.physics_margin_provider is not None): - points = ctx.physics_margin_provider() or [] - - def _physics_rollout(point: Dict[str, float]) -> Tuple[bool, str]: - with fresh_scope(physical_overrides=point): - return rollout() - - prefetched = _prefetch_parallel( - [functools.partial(_physics_rollout, point) for point in points], - f"{subject} physics margin") - passed: List[bool] = [] - for point_idx, point in enumerate(points): - ctx.attempt_rollout_count += 1 - pre = prefetched[point_idx] - ok, why = pre if pre is not None else _physics_rollout(point) - passed.append(bool(ok)) - desc = ", ".join(f"{k}={v:.4g}" for k, v in sorted(point.items())) - if ok: - outcomes.append(f"physics point ({desc}): goal reached") - else: - outcomes.append(f"physics point ({desc}): FAILED - {why}") - if detail is None: - detail = f"at {desc}: {why}" - if (detail is not None and any(passed) - and CFG.code_sim_learning_interval_belief): - # Certification over the belief interval (interval belief): - # a mixed sweep is the interval straddling the plan's success - # boundary, and the passing/failing ranges say which way. - straddle = straddle_summary(points, passed) - detail += (f" ({sum(passed)}/{len(passed)} belief-interval " - f"points passed" + - (f"; {straddle}" if straddle else "") + ")") - if points and detail is None: - note += ( - f" Physics-margin check passed: the {subject} also reached " - f"the goal at all {len(points)} grid points spanning +-1 " - "sigma of the identified physical parameters.") - if (validation_cfg.rule_param_margin and detail is None - and ctx.rule_param_margin_provider is not None - and ctx.rule_param_override_scope is not None): - rule_points = ctx.rule_param_margin_provider() or [] - override_scope = ctx.rule_param_override_scope - - def _member_rollout(point: Dict[str, float]) -> Tuple[bool, str]: - assert override_scope is not None - with override_scope(point), fresh_scope(): - return rollout() - - # Prefetching runs every member even though the sequential loop - # below still breaks at the first failure - the extra results - # are discarded, keeping the report identical with the flag on - # or off (failures are rare enough that the prepaid tail is - # cheaper than serializing the common all-pass case). - member_prefetched = _prefetch_parallel( - [functools.partial(_member_rollout, pt) for pt in rule_points], - f"{subject} rule-param margin") - for member_idx, point in enumerate(rule_points): - ctx.attempt_rollout_count += 1 - pre = member_prefetched[member_idx] - ok, why = pre if pre is not None else _member_rollout(point) - desc = (f"{ctx.rule_param_margin_label} " - f"{member_idx + 1}/{len(rule_points)}") - if ok: - outcomes.append(f"{desc}: goal reached") - else: - shown = ", ".join(f"{k}={_fmt_point_value(v)}" - for k, v in sorted(point.items())[:8]) - if len(point) > 8: - shown += ", ..." - outcomes.append(f"{desc}: FAILED - {why}") - detail = f"under {desc} ({shown}): {why}" - break - if rule_points and detail is None: - note += ( - f" Rule-parameter margin check passed: the {subject} also " - f"reached the goal under all {len(rule_points)} " - f"{ctx.rule_param_margin_note}.") - return outcomes, detail, note - - -def _necessity_sweep( - ctx: ToolContext, fresh_scope: Callable[..., Any], - rollout_without: Callable[[int], Tuple[bool, str]], - step_names: Sequence[str]) -> Tuple[List[str], Optional[str], str]: - """Necessity gate: refuse a plan that still reaches the goal with one of - its steps removed. - - ``rollout_without(k)`` runs the plan with step ``k`` deleted and - returns ``(goal reached, why not)``. A step whose removal leaves the - goal reached is padding: it explains nothing about how the goal - comes about, spends real episode steps, and if the model is wrong - about it can break the plan for real (run_20260902_152811: a - validated capture pressed three of four buttons and released one - that was never on, for a goal its own model reached with two - presses and a Wait). - - Returns ``(outcome lines, redundant detail or None, suffix for the - validation note)``. A one-step plan has nothing to ablate. - """ - outcomes: List[str] = [] - detail: Optional[str] = None - if len(step_names) < 2: - return outcomes, detail, "" - - def _ablated(k: int) -> Tuple[bool, str]: - with fresh_scope(): - return rollout_without(k) - - prefetched = _prefetch_parallel( - [functools.partial(_ablated, k) for k in range(len(step_names))], - "plan necessity") - for k, name in enumerate(step_names): - ctx.attempt_rollout_count += 1 - pre = prefetched[k] - ok, why = pre if pre is not None else _ablated(k) - if ok: - outcomes.append(f"without step {k} ({name}): goal STILL reached") - if detail is None: - detail = f"step {k} ({name})" - else: - outcomes.append(f"without step {k} ({name}): {why}") - note = "" - if detail is None: - note = (f" Necessity check passed: removing any one of the " - f"{len(step_names)} steps loses the goal.") - return outcomes, detail, note - - -def _fmt_point_value(value: Any) -> str: - """A margin point's value for a report: numbers compactly, anything else (a - belief particle's nested latent) by its repr, truncated.""" - if isinstance(value, (int, float, np.floating, np.integer)): - return f"{float(value):.4g}" - text = repr(value) - return text if len(text) <= 40 else text[:37] + "..." - - -def _build_testing_tools(ctx: ToolContext, _text_result: Callable, - tool: Callable) -> Dict[str, Any]: - """Evaluation tools (option plans / policies against tasks).""" - - # Tool descriptions bake config values at BUILD time (session open); - # the handlers below re-read config at CALL time. - _gs_eval_doc = ( - "Runs your exact params with NO sampling (a `~` ground-sampler " - "annotation - `~ [w1, w2]` region or `~ my_sampler` - is accepted " - "but IGNORED here; only `sim.refine` uses it). " - if RefinementConfig.from_cfg().ground_samplers else - "Runs your exact params with NO sampling. ") - - @tool( - "submit_plan", - "SUBMIT a fully-specified plan as your answer for the CURRENT task. " - "`plan` is text - one option per line, same grammar as `sim.run` / " - "`sim.refine`: `Option(obj1:type1, obj2:type2)[param1, param2] -> " - "{Atom(obj:type), ...}` (typed object refs; EXACT continuous params " - "in `[]`, `[]` for none; optional `-> {atoms}` subgoals, prefix NOT " - "to require false). " + _gs_eval_doc + - "The plan is rolled out from the task's TRUE initial state through " - "the belief model and reported step by step (include_states/" - "include_atoms control the report). If it reaches the goal it is " - "captured as your answer, and the per-step subgoals make it execute " - "closed-loop (monitored, with replan-on-divergence). Capture is " - "gated: a goal-reaching plan is re-run several times (simulation " - "varies across runs; each rollout reports the motion-planner seed " - "it ran at) and a FLAKY plan is reported instead of captured. The " - "gate's rollout set is exactly what `sim.run(plan, trials=N)` runs " - "(fresh env per rollout, same planner seeds), so measure " - "reliability there BEFORE submitting, and reproduce one failed " - "rollout with `sim.run(plan, seed=S, fresh=True)`; then add " - "margin and resubmit. `validation_rollouts` " - "requests a STRICTER gate for this submission (more rollouts; never " - "fewer than configured). This is the ONLY path that captures an " - "answer: explore (other tasks, modified states, partial plans, " - "parameter sweeps, seeded reproductions) with `sim` in run_python, " - "then submit the final plan here. " - "When identified physical parameters are active, it is also re-run " - "at a grid of perturbations spanning +-1 sigma of those parameters " - "(the physics fit's own uncertainty); a PARAM-SENSITIVE plan is " - "reported instead of captured - add design margin so it succeeds " - "across the whole range. " - "When the necessity gate is on, it is also re-run once per step " - "with that step removed; a plan that still reaches the goal without " - "one of its steps is reported REDUNDANT naming the step, not " - "captured - submit the shortest plan your model needs. " - "When the task has an evaluator, a goal-reaching plan the evaluator " - "still scores as a non-solve (no success credit in its reward) is " - "NOT captured (the real env applies the same scoring, so it could " - "never count as a solve).", - { - "type": "object", - "properties": { - "plan": { - "type": - "string", - "description": - "Plan text, one option per line: " - "`Option(obj1:type1, obj2:type2)[p1, p2] -> " - "{Atom(obj:type), ...}` (exact params in `[]`; `[]` for " - "none; optional `-> {atoms}` subgoals, NOT-prefix to " - "require false).", - }, - "include_states": { - "type": - "boolean", - "description": - "Include the full low-level state feature dict after each " - "step", - "default": - True - }, - "include_atoms": { - "type": "boolean", - "description": - "Include atoms added/deleted after each step", - "default": True - }, - "validation_rollouts": { - "type": - "integer", - "description": - "Request a stricter capture gate: total validation " - "rollouts a goal-reaching submission must pass. The " - "effective count is max(configured gate, this) - it can " - "raise the gate but never lower it. Use before " - "committing a plan you suspect is marginal.", - }, - }, - "required": ["plan"], - }, - ) - async def submit_plan(args: Dict[str, Any]) -> Dict[str, Any]: - refine_cfg = RefinementConfig.from_cfg() - validation_cfg = ValidationConfig.from_cfg() - ctx.test_call_id += 1 - # Snapshot for the [budget] footer's per-call delta; this handler - # increments the counter itself (initial rollout + validation - # repeats), and without the snapshot the footer reports the - # attempt's cumulative total as "+N this call". - rollouts_before = ctx.attempt_rollout_count - - if ctx.option_model is None: - return _error_result("No option model available in ToolContext.") - - all_options = ctx.options - opt_map = {o.name: o for o in all_options} - model = ctx.option_model - model._name_to_parameterized_option = ( # type: ignore[attr-defined] # pylint: disable=protected-access - opt_map) - - plan_text = (args.get("plan") or "").strip() - include_states = args.get("include_states", False) - include_atoms = args.get("include_atoms", True) - requested_rollouts = args.get("validation_rollouts") - if requested_rollouts is not None and (not isinstance( - requested_rollouts, int) or requested_rollouts < 1): - return _error_result( - "validation_rollouts must be a positive integer.") - - # Always the CURRENT task from its true initial state: this is - # the submission path, and exploration on other tasks or from - # modified states lives on the probe (sim.run). - resolved, task_err = _resolve_task(ctx, None) - if task_err is not None: - return task_err - assert resolved is not None - task = resolved.task - task_label = resolved.label - - lines = [f"Testing option plan on task {task_label}:"] - saved_image_paths: List[str] = [] - - all_predicates = ctx.predicates - - if not plan_text: - return _error_result("`plan` is required (option plan text).") - # Parse the text plan into a sketch (options + objects + exact params + - # subgoals) using the SAME grammar/parser as sim.refine. - types = set(ctx.types) - for opt in all_options: - types.update(opt.types) - for pred in all_predicates: - types.update(pred.types) - types.update(o.type for o in task.init) - try: - # strict: the `plan` argument is pure plan text, so a line that - # fails to parse is an error the agent must see - silently - # dropping it (the freeform default) executes a different plan - # than the agent asked for. - gs_fns, gs_err = load_ground_sampler_fns(ctx) - if gs_err is not None: - return _error_result(gs_err) - parse_notices: List[str] = [] - sketch_steps = bilevel_sketch.parse_sketch_from_text( - plan_text, - task, - predicates=all_predicates, - options=all_options, - types=types, - parse_continuous_params=True, - strict=True, - parse_ground_samplers=refine_cfg.ground_samplers, - ground_sampler_fns=gs_fns or None, - notices=parse_notices) - except Exception as e: # pylint: disable=broad-except - return _error_result(f"Could not parse plan: {e}") - lines.extend(f"NOTE: {n}" for n in parse_notices) - if not sketch_steps: - return _error_result( - "Parsed empty plan. Each line must be " - "`Option(obj:type, ...)[params] -> {subgoals}` with a known " - "option, typed object refs, and exact params in `[]`.") - # Ground each step with its parsed exact params, via the same - # helper the refine path uses: an annotated Wait gets its - # wait_target_atoms installed, so it waits for the annotated - # atoms here exactly as in refine and in real execution - - # grounding directly made the same Wait terminate on the first - # incidental atom change in this rollout but wait for its - # targets in refine, two different durations for one plan. - grounded_plan: List[Any] = [] - for step_idx, st in enumerate(sketch_steps): - params = (st.initial_params if st.initial_params is not None else - np.array([], dtype=np.float32)) - try: - grounded_plan.append( - bilevel_sketch.ground_step( - st, np.asarray(params, dtype=np.float32))) - except Exception as e: # pylint: disable=broad-except - return _error_result(f"Failed to ground step {step_idx} " - f"({st.option.name}): {e}") - - # Per-low-level-step states + option labels for the task-evaluator - # verdict below (see _EvalStateCollector for why per-step states). - eval_collector = _EvalStateCollector(model, task.init) - - # Robot-clearance probe (see tools/clearance.py): rollout 1 and - # every validation repeat feed it their low-level trajectories, - # and its verdict joins the margin gates below. None when the - # plan's skills carry no planning simulator to measure on. - clearance_probe: Optional[RobotClearanceProbe] = None - probe_skill = phase_skill_of(grounded_plan) - if probe_skill is not None: - clearance_probe = RobotClearanceProbe(probe_skill) - rollout_counter = [1] - - def _probe_clearance(label: str, outcome: Any) -> None: - if clearance_probe is None: - return - traj = getattr(model, "last_trajectory", None) - states = getattr(traj, "states", None) - if states: - clearance_probe.observe(label, outcome.option, states) - - # Per-step report callback, driven by the shared forward executor. - def _report_step(i: int, outcome: Any) -> None: - eval_collector.collect(outcome) - _probe_clearance("rollout 1", outcome) - opt = outcome.option - sig = f"{opt.name}({[o.name for o in opt.objects]})" - if not outcome.initiable: - atoms = utils.abstract(outcome.pre_state, ctx.predicates) - atoms_str = ", ".join(str(a) for a in sorted(atoms)) - lines.append(f"Step {i}: {sig} - NOT INITIABLE\n" - f" Current atoms: {{{atoms_str}}}\n" - f" Object poses at failure:\n" - f"{format_object_poses(outcome.pre_state)}") - return - step_line = f"Step {i}: {sig} ({outcome.num_actions} actions)" - if (opt.name == "Wait" and outcome.failure_reason is None - and outcome.num_actions >= utils.wait_rollout_step_cap()): - step_line += ( - "\n NOTE: this Wait ran to its step cap - " - "its wait-target atoms never became true in the " - "belief (and no other atom changed). Check whether " - "the awaited change is modeled, or drop the Wait.") - if outcome.failure_reason is not None: - step_line += (f"\n FAILURE REASON: {outcome.failure_reason}" - "\n Object poses at failure:\n" - f"{format_object_poses(outcome.pre_state)}") - post = outcome.post_state - if post is not None and include_atoms: - before = utils.abstract(outcome.pre_state, ctx.predicates) - after = utils.abstract(post, ctx.predicates) - added_s = ", ".join(str(a) for a in sorted(after - before)) - del_s = ", ".join(str(a) for a in sorted(before - after)) - step_line += (f"\n Added: {{{added_s}}}" - f"\n Deleted: {{{del_s}}}") - if post is not None and include_states: - step_line += ("\n State:\n" + - post.dict_str(indent=4, num_decimal_points=4)) - lines.append(step_line) - # Render from the outcome's state explicitly: the rollout - # runs on the gate's fresh env (see fresh_scope below) while - # the renderer draws the shared session env, so without the - # state the images would show a stale scene. - img_block = render_pybullet_image( - ctx, - f"step_{i}_{opt.name}", - state=post if post is not None else outcome.pre_state) - if img_block and img_block.get("saved_path"): - saved_image_paths.append(img_block["saved_path"]) - - # One substrate for the WHOLE gate: rollout 1 (the capture - # rollout) runs on the same freshly constructed env as the - # validation repeats, at the base planner seed. This makes the - # gate reproducible from inside the session - sim.run(plan, - # trials=N) runs the identical rollout set (fresh env per trial, - # planner seeds base..base+N-1) - and stops a submission from - # advancing the shared session env. Rollout 1 on the warm shared - # env was a different physics substrate from every repeat: the - # 2026-08-30 bridge runs tuned plans to 27/27 on one substrate - # that then scored 1/10 on the other, with no way to reproduce - # the gate's rollouts. - fresh_scope = (ctx.validation_env_scope - if validation_cfg.fresh_env else None) - # Execute exactly like the real closed-loop executor: abort at the - # first failing option (0-action collision / not-initiable / env - # failure) instead of pressing on. Otherwise forward simulation can - # continue past a collision and report a goal that the real rollout — - # which ends the episode at that failed option — never reaches. - ctx.attempt_rollout_count += 1 - with (fresh_scope() - if fresh_scope is not None else contextlib.nullcontext()): - result = bilevel_sketch.execute_plan_forward( - task, - grounded_plan, - ctx.option_model, - predicates=all_predicates, - sketch=sketch_steps, - on_step=_report_step, - stop_on_failure=True) - final_atoms = utils.abstract(result.final_state, ctx.predicates) - # Task-evaluator verdict on this belief-sim rollout, computed - # BEFORE capture and INSIDE the scope (certificate probes must - # judge on the env the rollout ran on): the real evaluator - # applies the same certificate, so a goal-reaching but - # illegitimate plan can never count as a solve and must not be - # captured as the answer (run_20260712_173955 tasks 1-2: - # flagged-illegitimate captures stood all session and were - # executed only to be rejected). Failure-tolerant: verdict - # stays None when the task has no evaluator or nothing - # executed. A coarse verdict (option-boundary states only) can - # falsely reject a legitimate cascade, so it never blocks - # capture. - evaluator = _resolve_task_evaluator(ctx, task_label) - verdict: Optional[Dict[str, Any]] = None - if evaluator is not None and len(eval_collector.states) > 1: - try: - verdict = evaluate_states_with(evaluator, - eval_collector.states, - eval_collector.labels, - sim_env=getattr( - ctx.option_model, - "sim_env", None)) - except Exception as e: # pylint: disable=broad-except - logging.debug("Task-evaluator verdict failed: %s", e) - # Use the env's goal-check (its own classifiers); robust to invented - # predicates that don't reuse env names. - goal_reached = result.goal_reached - # One more real-executor constraint the option model doesn't enforce: - # the episode is capped at the phase's step budget (the horizon, or - # the interaction-request cap for explore episodes). A plan whose - # goal is reached only after more steps than that will time out in - # real rollout, so don't count it as achieved/captured. - horizon = ctx.execution_step_budget() - within_horizon = (result.actions_to_goal is not None - and result.actions_to_goal <= horizon) - goal_achieved = (goal_reached and result.clean_to_goal - and within_horizon) - evaluator_rejected = (verdict is not None and not verdict["legitimate"] - and not eval_collector.coarse) - # An evaluator rejection only disqualifies a capture when the goal - # atoms actually hold via an illegitimate route - a reward hack (e.g. - # the agent knocked the target over directly). An honest shortfall, - # where the rollout simply fails to reach the goal, is ALSO - # legitimate=False (there is no genuine cascade to certify), but that - # is exactly what a best-effort submission is meant to capture, so it - # must not be conflated with a reward hack. - reward_hack = (evaluator_rejected and verdict is not None - and verdict["terminated"]) - - # Multi-rollout validation of a capture candidate. The shared sim - # env is nondeterministic across repeats (motion-planner sampling, - # physics-solver state), which is the same variability the real - # rollout will sample - a plan that only sometimes succeeds here is - # a margin-free plan that will likely fail on the real env - # (run_20260712_192457 task 1: a sim-validated 2-hop relay died on a - # ~9mm placement drift). So a goal-reaching plan is captured only - # after every one of validation_cfg.rollouts total - # rollouts succeeds; a flaky repeat is reported to the agent, who - # still has the session to add margin and resubmit. - def _validation_rollout() -> Tuple[bool, str, List[Optional[State]]]: - """One extra rollout of the exact plan. - - Returns ``(ok, failure detail, per-step post-states)``; the - post-state list is padded with ``None`` to the plan length - so a truncated (failed) rollout still indexes safely. - Passing rollouts' post-states feed the captured-annotation - intersection filter. - """ - v_collector = _EvalStateCollector(model, task.init) - rollout_counter[0] += 1 - rollout_label = f"rollout {rollout_counter[0]}" - - def _on_validation_step(i: int, outcome: Any) -> None: - v_collector.on_step(i, outcome) - _probe_clearance(rollout_label, outcome) - - r = bilevel_sketch.execute_plan_forward( - task, - grounded_plan, - model, - predicates=all_predicates, - sketch=sketch_steps, - on_step=_on_validation_step, - stop_on_failure=True) - posts: List[Optional[State]] = [s.post_state for s in r.steps] - posts += [None] * (len(grounded_plan) - len(posts)) - if r.first_failure_idx is not None: - fr = r.steps[r.first_failure_idx].failure_reason - opt = r.steps[r.first_failure_idx].option - return False, (f"step {r.first_failure_idx} " - f"({opt.name}) failed: {fr}"), posts - if not r.goal_reached: - missing = _missing_goal_atoms(task, r.final_state) - missing_str = ", ".join(str(a) for a in sorted(missing)) - detail = f" (missing: {{{missing_str}}})" if missing else "" - return False, f"goal not reached{detail}", posts - if not (r.actions_to_goal is not None - and r.actions_to_goal <= horizon): - return False, (f"goal reached only after " - f"{r.actions_to_goal} low-level steps, past " - f"the episode horizon ({horizon})"), posts - # Same legitimacy rule as the first rollout: a non-coarse - # illegitimate verdict fails the validation. - if (evaluator is not None and len(v_collector.states) > 1 - and not v_collector.coarse): - try: - v = evaluate_states_with(evaluator, - v_collector.states, - v_collector.labels, - sim_env=getattr( - ctx.option_model, "sim_env", - None)) - if not v["legitimate"]: - return False, ( - "this rollout reached the goal atoms but the " - "task evaluator scored it as a non-solve " - f"(solved=False, reward={v['reward']:.2f})"), posts - except Exception as e: # pylint: disable=broad-except - logging.debug("Validation-rollout verdict failed: %s", e) - return True, "", posts - - flaky_detail: Optional[str] = None - validation_note = "" - n_rollouts = max(1, validation_cfg.rollouts) - # Escalated gate once this task has produced a FLAKY rejection: the - # agent is provably tuning in a marginal region, where a lucky - # streak passes the base gate and dies on the single real episode - # (run_20260717_182321: a 20/20-swept relay placement validated 3/3, - # then missed the target for real). - capture_task_key = _capture_task_key(ctx) - if capture_task_key in ctx.flaky_capture_task_keys: - n_rollouts = max(n_rollouts, validation_cfg.rollouts_after_flaky) - # The agent may request a STRICTER gate for this submission (a - # plan it suspects is marginal); it can never lower the - # configured gate - that would let a lucky draw bypass it. - capped_request: Optional[int] = None - if requested_rollouts is not None: - capped_request = min(requested_rollouts, _MAX_REQUESTED_ROLLOUTS) - if capped_request < requested_rollouts: - lines.append( - f"NOTE: validation_rollouts={requested_rollouts} capped " - f"at {_MAX_REQUESTED_ROLLOUTS}.") - n_rollouts = max(n_rollouts, capped_request) - # fresh_scope (computed above, shared with rollout 1): repeats on - # the shared env are correlated (its reset cannot reconstruct - # state exactly), so only fresh envs sample the same distribution - # the real episode will. - rollout_outcomes: List[str] = [] - # Per-step post-states of PASSING validation rollouts, for the - # captured-annotation intersection filter. Failing rollouts are - # excluded on purpose: they are off-track by definition, so their - # post-states are not evidence about what holds on a successful - # execution (using them would prune annotations that hold in - # every on-track run). Physics-margin rollouts are likewise - # excluded: they run under deliberately perturbed physics. - passing_validation_posts: List[List[Optional[State]]] = [] - base_planner_seed = CFG.seed - if (ctx.capture_goal_reaching_plans and goal_achieved - and not evaluator_rejected and grounded_plan - and n_rollouts > 1): - # Run ALL validation rollouts even after a failure: the - # per-rollout outcome list distinguishes failure modes (a - # physics-tail fizzle vs. an IK stall vs. a certificate - # rejection) and yields a reliability estimate - a bare - # "rollout k FAILED" left agents guessing which - # (run_20260717_182040 seed0 turn 214). - # decorrelated_rollout_seed: a fresh env alone gives - # bit-identical repeats (motion planning reads the - # constant CFG.seed at call time), so without it the - # validation repeats re-run the capture rollout verbatim - # and detect nothing. The capture rollout itself keeps - # the base seed; repeats sample execution variability. - def _repeat_rollout( - repeat_idx: int - ) -> Tuple[bool, str, List[Optional[State]]]: - with (fresh_scope() if fresh_scope is not None else - contextlib.nullcontext()), \ - decorrelated_rollout_seed(repeat_idx - 1): - return _validation_rollout() - - repeat_indices = list(range(2, n_rollouts + 1)) - repeat_prefetched = _prefetch_parallel([ - functools.partial(_repeat_rollout, k) for k in repeat_indices - ], "capture repeat rollouts") - for pos, repeat_idx in enumerate(repeat_indices): - ctx.attempt_rollout_count += 1 - pre = repeat_prefetched[pos] - ok, why, repeat_posts = (pre if pre is not None else - _repeat_rollout(repeat_idx)) - repeat_seed = base_planner_seed + repeat_idx - 1 - if ok: - passing_validation_posts.append(repeat_posts) - rollout_outcomes.append( - f"rollout {repeat_idx} (planner seed " - f"{repeat_seed}): goal reached") - else: - rollout_outcomes.append( - f"rollout {repeat_idx} (planner seed " - f"{repeat_seed}): FAILED - {why}") - if flaky_detail is None: - flaky_detail = (f"rollout {repeat_idx}/{n_rollouts} " - f"(planner seed {repeat_seed}) " - f"FAILED: {why}") - if flaky_detail is None: - fresh_note = (", each on a freshly constructed simulator " - "instance" if fresh_scope is not None else "") - validation_note = ( - f" Validated {n_rollouts}/{n_rollouts} rollouts " - f"(planner seeds {base_planner_seed}-" - f"{base_planner_seed + n_rollouts - 1}; the " - "simulator's motion planning and physics stepping vary " - "across runs; repeats sample that execution " - f"variability{fresh_note}; sim.run(plan, " - f"trials={n_rollouts}) reruns this exact rollout set).") - - # Parameter-margin gates (see _parameter_margin_sweep): the - # execution repeats above all run AT the fitted parameters, so - # they cannot see a plan whose success band excludes the fit's - # own error - in the identified physical params or the learned - # rule constants. - param_sensitive_detail: Optional[str] = None - margin_outcomes: List[str] = [] - if (fresh_scope is not None and ctx.capture_goal_reaching_plans - and goal_achieved and not evaluator_rejected and grounded_plan - and flaky_detail is None): - margin_outcomes, param_sensitive_detail, margin_note = \ - _parameter_margin_sweep( - ctx, validation_cfg, fresh_scope, - lambda: _validation_rollout()[:2], "plan") - validation_note += margin_note - - # Clearance gate (see tools/clearance.py): the rollouts above - # certify the plan against the belief's own execution - # variability, not against the real executor's realization slop; - # a robot link passing inside that slop of a bystander is a - # margin-free plan whether or not every rollout cleared it. - clearance_detail: Optional[str] = None - clearance_summary = "" - if (clearance_probe is not None and goal_achieved - and not evaluator_rejected and grounded_plan - and flaky_detail is None): - clearance_ok, clearance_summary, clearance_why = \ - clearance_probe.verdict() - if not clearance_ok: - clearance_detail = clearance_why - any_margin_detail = param_sensitive_detail or clearance_detail - - # Necessity gate (see _necessity_sweep): only a plan that cleared - # every gate above is worth ablating, and only a plan with more - # than one step can be. - redundant_detail: Optional[str] = None - necessity_outcomes: List[str] = [] - if (validation_cfg.necessity and fresh_scope is not None - and ctx.capture_goal_reaching_plans and goal_achieved - and not evaluator_rejected and grounded_plan - and flaky_detail is None and any_margin_detail is None): - - def _rollout_without(k: int) -> Tuple[bool, str]: - ablated_plan = grounded_plan[:k] + grounded_plan[k + 1:] - ablated_sketch = (sketch_steps[:k] + sketch_steps[k + 1:] - if sketch_steps is not None else None) - r = bilevel_sketch.execute_plan_forward( - task, - ablated_plan, - model, - predicates=all_predicates, - sketch=ablated_sketch, - stop_on_failure=True) - if r.first_failure_idx is not None: - fr = r.steps[r.first_failure_idx].failure_reason - return False, (f"step {r.first_failure_idx} of the " - f"shortened plan failed: {fr}") - if not r.goal_reached: - missing = _missing_goal_atoms(task, r.final_state) - missing_str = ", ".join(str(a) for a in sorted(missing)) - return False, ("goal not reached " - f"(missing: {{{missing_str}}})") - return True, "" - - necessity_outcomes, redundant_detail, necessity_note = \ - _necessity_sweep( - ctx, fresh_scope, _rollout_without, - [f"{g.name}({', '.join(o.name for o in g.objects)})" - for g in grounded_plan]) - validation_note += necessity_note - - def _stash_uncaptured_submission() -> None: - """Remember the best refused submission of this attempt. - - The journal auto-entry records it at attempt end, so the - plan (and its honest evaluator reward) survives the fresh- - context restart and the final best-effort nudge can resubmit - it instead of the attempt's work vanishing with its context. - """ - reward = float(verdict["reward"]) if verdict is not None else None - prev = ctx.best_uncaptured_reward - if ctx.best_uncaptured_plan_lines is not None and ( - reward is None or (prev is not None and reward <= prev)): - return - ctx.best_uncaptured_reward = reward - ctx.best_uncaptured_plan_lines = list( - bilevel_sketch.format_plan_lines(grounded_plan)) - - # The capture decision itself is pure (see _decide_capture, which - # also documents the best-effort-mode semantics); the branches - # below apply its ctx mutations and format its messages. - capture_outcome = _decide_capture( - # In policy mode the deliverable is policy.py (via - # submit_policy); this tool remains a probe but can no - # longer capture the answer. - capture_enabled=(ctx.capture_goal_reaching_plans - and not ctx.policy_capture_mode), - is_current_task=True, - have_plan=bool(grounded_plan), - goal_achieved=goal_achieved, - evaluator_rejected=evaluator_rejected, - reward_hack=reward_hack, - flaky=flaky_detail is not None, - best_effort_mode=ctx.capture_best_effort_plan, - have_validated_capture=bool(ctx.solved_plan_reached_goal), - param_sensitive=any_margin_detail is not None, - redundant=redundant_detail is not None) - decision = capture_outcome.decision - captured = capture_outcome.captured - # Adaptive info-seeking trigger (ctx.info_seeking_active): a plan - # refused because a learned/physical parameter's uncertainty - # threatens it is the signal that the agent now needs to spend - # real steps reducing that uncertainty; a captured plan clears it. - # Clearance (bystander slop) is a separate concern and does not - # arm info-seeking, so key on the parameter-margin detail only. - if captured: - ctx.param_sensitive_refusal_pending = False - elif param_sensitive_detail is not None: - ctx.param_sensitive_refusal_pending = True - if captured: - # Capture the plan with a sketch that keeps only the subgoals - # that actually held (so the closed-loop monitor won't flag a - # spurious divergence on a wrong annotation). An annotation - # must hold in rollout 1 AND in every PASSING validation - # rollout: an atom that held once by luck under the sim's own - # nondeterminism would otherwise survive into the executed - # sketch and kill the real episode on a spurious divergence. - # (With zero passing repeats this reduces to the rollout-1 - # filter.) - validated_solve = decision is CaptureDecision.VALIDATED_CAPTURE - captured_sketch = [] - - def _held_in_passing_repeats(atom: Any, i: int, - want_held: bool) -> bool: - for posts in passing_validation_posts: - post_i = posts[i] if i < len(posts) else None - # bool(): classifiers may return numpy bools, which - # fail identity checks against Python bools. - if post_i is None or bool(atom.holds(post_i)) != want_held: - return False - return True - - # Execution-verifiability probe. The closed-loop monitor - # evaluates the captured annotations on REAL observations, - # which carry no latent (``State.latent`` is None outside - # belief rollouts). An atom whose truth in the certifying - # post-state depends on the belief latent therefore reads - # false at execution no matter what physically happens, and - # a single such annotation aborts a healthy episode (a - # latent-only SeamBonded killed two runs whose bonds had in - # fact formed). Certify each surviving positive atom on the - # same post-state with the latent stripped and drop the - # ones that fail; a classifier that RAISES without a latent - # would crash the monitor, so it is dropped from either - # polarity the same way (a stripped negative atom that - # merely evaluates is kept - it cannot fire spuriously). - unverifiable_dropped: List[str] = [] - - def _probe_without_latent(atom: Any, - post: State) -> Optional[bool]: - stripped = State(post.data) - try: - return bool(atom.holds(stripped)) - except Exception: # pylint: disable=broad-except - return None - - for i, st in enumerate(sketch_steps): - post = (result.steps[i].post_state - if i < len(result.steps) else None) - if post is not None: - after = utils.abstract(post, all_predicates) - pos_held = { - a - for a in (st.subgoal_atoms or set()) - if a in after and _held_in_passing_repeats(a, i, True) - } - neg_held = { - a - for a in (st.subgoal_neg_atoms or set()) - if a not in after - and _held_in_passing_repeats(a, i, False) - } - pos_drop = { - a - for a in pos_held - if _probe_without_latent(a, post) is not True - } - neg_drop = { - a - for a in neg_held - if _probe_without_latent(a, post) is None - } - unverifiable_dropped.extend( - f"step {i} ({st.option.name}): {a}" - for a in sorted(pos_drop | neg_drop, key=str)) - pos_held -= pos_drop - neg_held -= neg_drop - else: - pos_held, neg_held = set(), set() - captured_sketch.append( - bilevel_sketch.SketchStep(option=st.option, - objects=st.objects, - subgoal_atoms=pos_held or None, - subgoal_neg_atoms=neg_held - or None)) - # Re-align each Wait's target atoms with the FILTERED - # sketch: the real executor waits on exactly the monitored - # (execution-verifiable) atoms. A latent-only target (e.g. a - # belief Bonded) reads false on every real observation, so - # leaving it in the grounded option's memory would stall the - # real Wait to its step-cap backstop no matter what happens. - # With every target filtered away the Wait falls back to - # any-atom-change, the same rule the belief rollout then - # shares. - for g_opt, cap_step in zip(grounded_plan, captured_sketch): - if g_opt.name != "Wait": - continue - g_opt.memory.pop("wait_target_atoms", None) - g_opt.memory.pop("wait_target_neg_atoms", None) - if cap_step.subgoal_atoms: - g_opt.memory["wait_target_atoms"] = cap_step.subgoal_atoms - if cap_step.subgoal_neg_atoms: - g_opt.memory["wait_target_neg_atoms"] = \ - cap_step.subgoal_neg_atoms - ctx.solved_plan = grounded_plan - ctx.solved_sketch = captured_sketch - ctx.solved_plan_reached_goal = validated_solve - ctx.solved_plan_eval_reward = (float(verdict["reward"]) - if verdict is not None else None) - summary_bits = [ - f"validation: " - f"{1 + sum(1 for o in rollout_outcomes if 'FAILED' not in o)}" - f"/{1 + len(rollout_outcomes)} rollouts ok" - ] - if flaky_detail is not None: - summary_bits.append(f"first failure: {flaky_detail}") - if margin_outcomes: - n_margin_ok = sum(1 for o in margin_outcomes - if "FAILED" not in o) - summary_bits.append(f"physics margin: {n_margin_ok}/" - f"{len(margin_outcomes)} points ok") - if clearance_summary: - summary_bits.append(clearance_summary) - ctx.solved_plan_validation_summary = "; ".join(summary_bits) - n_annot = sum(1 for s in captured_sketch - if s.subgoal_atoms or s.subgoal_neg_atoms) - reason = capture_outcome.best_effort_reason - if reason is None: - best_effort_note = "" - elif reason is BestEffortReason.GOAL_NOT_REACHED: - best_effort_note = (" (best-effort: goal NOT reached, " - "accepted because the attempt budget is " - "exhausted; it executes for its honest " - "reward but will not count as a solve)") - elif reason is BestEffortReason.REWARD_HACK: - best_effort_note = (" (best-effort: the rollout reaches the " - "goal atoms but the task evaluator " - "scores it as a non-solve, and the real " - "env applies the same scoring; accepted " - "because the attempt budget is exhausted " - "- it executes for its honest reward but " - "will not count as a solve)") - elif reason is BestEffortReason.FLAKY: - best_effort_note = (f" (best-effort: {flaky_detail}; " - "accepted because the attempt budget is " - "exhausted - it executes for its honest " - "reward but may not reproduce its " - "solve)") - elif reason is BestEffortReason.PARAM_SENSITIVE: - best_effort_note = (" (best-effort: failed " - f"{any_margin_detail}; accepted " - "because the attempt budget is exhausted " - "- it executes for its honest reward but " - "may fail under the true physics)") - else: - assert reason is BestEffortReason.REDUNDANT - best_effort_note = (" (best-effort: the plan also reaches " - f"the goal without {redundant_detail}; " - "accepted because the attempt budget is " - "exhausted - it executes for its honest " - "reward, padding included)") - unverifiable_note = "" - if unverifiable_dropped: - dropped_lines = "\n".join(f" {d}" - for d in unverifiable_dropped) - unverifiable_note = ( - f"\nNOTE: {len(unverifiable_dropped)} annotation(s) " - "cannot be verified from a real observation (their " - "truth here depends on the belief latent, which real " - "env states do not carry) and were excluded from " - "closed-loop monitoring - the plan still executes " - "them, they just cannot trigger a replan:\n" - f"{dropped_lines}\n" - "Prefer annotating with predicates whose classifiers " - "read observable features.") - logging.info( - "Capture: excluded %d execution-unverifiable " - "annotation(s) from the monitored sketch:\n%s", - len(unverifiable_dropped), dropped_lines) - lines.append(f"Captured as the current answer{best_effort_note}: " - f"{len(grounded_plan)} steps, " - f"{n_annot} with subgoal annotations for closed-loop " - f"monitoring.{validation_note}{unverifiable_note}") - elif decision is CaptureDecision.FLAKY_NO_CAPTURE: - # Record the task so later submissions face the escalated - # gate - flakiness here is evidence the whole parameter - # region is marginal, not just this point. - ctx.flaky_capture_task_keys.add(capture_task_key) - _stash_uncaptured_submission() - escalated_n = max(max(1, validation_cfg.rollouts), - validation_cfg.rollouts_after_flaky) - n_ok = 1 + sum(1 for o in rollout_outcomes if "FAILED" not in o) - per_rollout = "\n".join(f" {o}" for o in rollout_outcomes) - lines.append( - f"FLAKY (plan NOT captured): the plan reached the goal on " - f"rollout 1 but {flaky_detail}. Per-rollout outcomes " - f"(estimated reliability {n_ok}/{n_rollouts}):\n" - f" rollout 1 (planner seed {base_planner_seed}): " - f"goal reached\n{per_rollout}\n" - "The simulator's motion " - "planning and physics stepping vary across runs, and the " - "real environment samples the same variability - a plan " - "that only sometimes succeeds in simulation will likely " - "fail for real. This gate is reproducible in run_python: " - f"sim.run(plan, trials={n_rollouts}) runs the identical " - "rollout set (fresh env per trial, same planner seeds), " - "and sim.run(plan, seed=, fresh=True) re-runs one failed rollout exactly " - "with full per-step reporting (without fresh=True the " - "warm session env is a different, optimistic substrate). " - "Then add margin (e.g. tighter spacing, aim " - "impacts closer to the middle of the fall path) and " - "resubmit. Because this task has now produced a flaky " - f"submission, captures require {escalated_n}/{escalated_n} " - "successful rollouts: fix the margin rather than " - "resubmitting near-identical parameters.") - elif decision is CaptureDecision.REDUNDANT_NO_CAPTURE: - _stash_uncaptured_submission() - per_step = "\n".join(f" {o}" for o in necessity_outcomes) - lines.append( - "REDUNDANT (plan NOT captured): every validation rollout " - f"reached the goal, but so does the plan without " - f"{redundant_detail}. A captured plan is an explanation of " - "how the goal comes about, and a step whose absence changes " - "nothing explains nothing: it only spends real episode " - "steps, and if your model is wrong about it, it can break " - "the plan for real. Per-step ablation:\n" - f"{per_step}\n" - "Drop the unnecessary step(s) and resubmit the shortest " - "plan your model needs. If you believed that step was " - "necessary, your model disagrees: check the belief before " - "resubmitting.") - elif (decision is CaptureDecision.PARAM_SENSITIVE_NO_CAPTURE - and param_sensitive_detail is None): - _stash_uncaptured_submission() - lines.append( - "CLEARANCE-SENSITIVE (plan NOT captured): every validation " - f"rollout reached the goal, but {clearance_detail}. The " - "real executor realizes each move anywhere within that " - "slop (pose tolerance, grasp height, landing scatter), so " - "a plan this tight succeeds in the belief by the luck of " - "the draw and fails for real on the next one. Widen the " - "margin at that step: stage and place objects farther " - "apart than the gripper's footprint, hover higher above " - "faces, or grasp farther from the neighbor - then " - "resubmit.") - elif decision is CaptureDecision.PARAM_SENSITIVE_NO_CAPTURE: - _stash_uncaptured_submission() - per_point = "\n".join(f" {o}" for o in margin_outcomes) - lines.append( - "PARAM-SENSITIVE (plan NOT captured): the plan passed " - "execution validation at the fitted physical parameters " - f"but FAILED {param_sensitive_detail}.\n" - "Physics-margin rollouts (a grid spanning +-1 sigma of " - "the identified physical parameters, the sysID fit's " - f"own uncertainty):\n{per_point}\n" - "The fitted values are uncertain at this scale and the " - "real environment may sit anywhere in that range - " - "including BETWEEN passing points: success can be " - "non-monotonic in a physical parameter, so a design must " - "hold across the whole range, not just at the values you " - "tuned at. Add margin to the DESIGN (not the execution) - " - "e.g. tighter spacing or impacts nearer the middle of the " - "fall path - then resubmit.") - if CFG.code_sim_learning_interval_belief and any( - "goal reached" in o for o in margin_outcomes): - lines.append( - "The belief interval straddles this plan's success " - "boundary (the passing and failing ranges are named " - "above). One real experiment that narrows that " - "parameter is worth more than more planning here: call " - "sim.suggest_probes(plan_text) to rank probes on your " - "sketch, run the best one for real, refit, and " - "resubmit - or find a design that holds across the " - "whole interval.") - elif decision is CaptureDecision.REWARD_HACK_NO_CAPTURE: - assert verdict is not None - _stash_uncaptured_submission() - lines.append( - "NOT CAPTURED: the rollout reaches the goal atoms but the " - "task evaluator scores it as a non-solve (solved=False, " - f"reward={verdict['reward']:.2f}). The real env applies the " - "same scoring, so executing this plan cannot count as a " - "solve. Find a plan whose rollout the evaluator scores " - "solved=True.") - if result.first_failure_idx is not None: - fr = result.steps[result.first_failure_idx].failure_reason - lines.append( - f"\nPlan FAILED at step {result.first_failure_idx}: {fr}") - final_atoms_str = ", ".join(str(a) for a in sorted(final_atoms)) - lines.append(f"\nFinal atoms: {{{final_atoms_str}}}") - if task.goal_nl: - lines.append(f"Goal (natural language): {task.goal_nl}") - else: - goal_str = ", ".join(str(g) for g in sorted(task.goal)) - lines.append(f"Goal: {{{goal_str}}}") - lines.append(f"Goal achieved: {goal_achieved}") - # Task-evaluator verdict line (verdict computed above, before the - # capture decision it gates). On a FLAKY rejection this verdict is - # rollout 1's only - printing it unlabeled next to a failing - # rollout's non-solve read as two contradictory verdicts in one - # message (run_20260717_182040 seed1 turn 96). - if verdict is not None: - vline = _format_evaluator_verdict(verdict, - coarse=eval_collector.coarse) - if flaky_detail is not None and not captured: - vline += (" [rollout 1 only - NOT the operative outcome; " - "this submission was rejected as FLAKY above]") - lines.append(vline) - # Goal atoms hold but the plan needs more low-level steps than the - # episode horizon allows: say so and that it was NOT captured, so the - # agent shortens the plan instead of stopping on a false positive. - # (A best-effort capture still happens above; then only warn.) - if goal_reached and not within_horizon and not captured: - lines.append( - f"NOT EXECUTABLE (plan was NOT captured): reaching the goal " - f"takes {result.actions_to_goal} low-level steps but the " - f"episode horizon is {horizon}. The real executor will run " - f"out of steps — shorten the plan (fewer or quicker steps) " - f"before resubmitting.") - elif goal_reached and not within_horizon: - lines.append( - f"WARNING: reaching the goal takes {result.actions_to_goal} " - f"low-level steps but the episode horizon is {horizon}, so " - f"the real executor will run out of steps before the goal.") - # Print the missing goal atoms even when the goal is stated in - # natural language: "Goal achieved: False" with no per-atom - # diagnosis left agents unable to tell a near-miss from a - # non-starter, and validation-rollout failures already name the - # missing atoms - this just makes rollout 1 report the same way. - if not goal_reached: - missing = _missing_goal_atoms(task, result.final_state) - missing_str = ", ".join(str(a) for a in sorted(missing)) - lines.append(f"Missing goal atoms: {{{missing_str}}}") - - # Append image save paths to text output - if saved_image_paths: - lines.append("\nSaved images:") - for p in saved_image_paths: - lines.append(f" {p}") - - # Build result with text only (images are saved to disk) - return _text_result("\n".join(lines) + - _budget_footer(ctx, rollouts_before)) - - @tool( - "submit_policy", - "Validate ./policy.py - your closed-loop `get_option(state, memory)` " - "program - on the CURRENT task and capture it as your answer. The " - "policy source is SNAPSHOTTED at call time (later edits need a new " - "call). Each rollout runs the policy closed-loop through the belief " - "model: get_option is called at every option boundary with the " - "actual current state; option failures (not initiable, motion-" - "planning refusal, 0 actions) do NOT end the episode - the failure " - "text arrives in memory['last_failure'] and get_option is asked " - "again, so RECOVERY is your policy's job; exceptions in get_option, " - "unparsable/ungroundable lines, re-issuing one identical " - "failing line repeatedly (the stuck-loop guard), and re-issuing " - "one identical line that keeps completing with no state change " - "(its no-op livelock twin) DO end it. Capture " - "is gated like " - "submit_plan: the goal-reaching rollout is repeated " - "several times (fresh simulator env + varied planner seed per " - "repeat, fresh memory per episode) and a FLAKY policy is reported " - "instead of captured; physics-margin perturbations apply too. " - "`validation_rollouts` requests a stricter gate. Test recovery " - "behavior first with sim.run_policy() in " - "run_python, which runs ./policy.py from the CURRENT probe " - "state (including perturbed or mid-plan states).", - { - "type": "object", - "properties": { - "validation_rollouts": { - "type": - "integer", - "description": - "Request a stricter capture gate: total validation " - "rollouts a goal-reaching policy must pass (effective " - "count is max(configured, this); never fewer).", - }, - "include_atoms": { - "type": "boolean", - "description": - "Include atoms added/deleted after each step", - "default": True - }, - }, - }, - ) - async def submit_policy(args: Dict[str, Any]) -> Dict[str, Any]: - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.policy_execution import \ - build_policy_option_fn, execute_policy_forward - validation_cfg = ValidationConfig.from_cfg() - ctx.test_call_id += 1 - rollouts_before = ctx.attempt_rollout_count - if not ctx.policy_capture_mode: - return _error_result( - "submit_policy is only available in policy mode " - "(agent_solve_policy_mode); submit plans via " - "submit_plan instead.") - if ctx.option_model is None: - return _error_result("No option model available in ToolContext.") - all_options = ctx.options - model = ctx.option_model - model._name_to_parameterized_option = ( # type: ignore[attr-defined] # pylint: disable=protected-access - {o.name: o - for o in all_options}) - requested_rollouts = args.get("validation_rollouts") - include_atoms = args.get("include_atoms", True) - if requested_rollouts is not None and (not isinstance( - requested_rollouts, int) or requested_rollouts < 1): - return _error_result( - "validation_rollouts must be a positive integer.") - - resolved, task_err = _resolve_task(ctx, None) - if task_err is not None: - return task_err - assert resolved is not None - task = resolved.task - task_label = resolved.label - - policy_path = _policy_source_path(ctx) - if policy_path is None or not os.path.isfile(policy_path): - return _error_result( - "No ./policy.py found. Write your closed-loop policy there " - "first: `def get_option(state, memory): ...` returning one " - "plan line (sketch grammar) or None for DONE.") - with open(policy_path, "r", encoding="utf-8") as f: - policy_source = f.read() - - all_predicates = ctx.predicates - types = set(ctx.types) - for opt in all_options: - types.update(opt.types) - for pred in all_predicates: - types.update(pred.types) - types.update(o.type for o in task.init) - - def _fresh_option_fn() -> Tuple[Optional[Any], Optional[str]]: - # Fresh instance per episode: memory must reset per rollout. - return build_policy_option_fn(policy_source, - task, - predicates=all_predicates, - options=all_options, - types=types) - - option_fn, load_err = _fresh_option_fn() - if load_err is not None or option_fn is None: - return _error_result(load_err or "policy.py failed to load.") - max_opts = CFG.agent_policy_max_options - horizon = ctx.execution_step_budget() - - lines = [f"Testing policy.py on task {task_label}:"] - saved_image_paths: List[str] = [] - eval_collector = _EvalStateCollector(model, task.init) - - def _report_step(i: int, outcome: Any) -> None: - eval_collector.collect(outcome) - opt = outcome.option - sig = f"{opt.name}({[o.name for o in opt.objects]})" - step_line = f"Step {i}: {sig} ({outcome.num_actions} actions)" - if outcome.failure_reason is not None: - step_line += ( - f"\n OPTION FAILURE (surfaced to the policy as " - f"memory['last_failure']): {outcome.failure_reason}") - post = outcome.post_state - if post is not None and include_atoms: - before = utils.abstract(outcome.pre_state, ctx.predicates) - after = utils.abstract(post, ctx.predicates) - added_s = ", ".join(str(a) for a in sorted(after - before)) - del_s = ", ".join(str(a) for a in sorted(before - after)) - step_line += (f"\n Added: {{{added_s}}}" - f"\n Deleted: {{{del_s}}}") - lines.append(step_line) - # Explicit state: the rollout runs on the gate's fresh env - # while the renderer draws the shared session env (see - # submit_plan's _report_step). - img_block = render_pybullet_image( - ctx, - f"policy_step_{i}_{opt.name}", - state=post if post is not None else outcome.pre_state) - if img_block and img_block.get("saved_path"): - saved_image_paths.append(img_block["saved_path"]) - - # One substrate for the whole gate, as in submit_plan: rollout 1 - # runs on the same fresh env as the validation repeats, at the - # base planner seed, so the gate is reproducible in-session. - fresh_scope = (ctx.validation_env_scope - if validation_cfg.fresh_env else None) - ctx.attempt_rollout_count += 1 - with (fresh_scope() - if fresh_scope is not None else contextlib.nullcontext()): - result = execute_policy_forward(task, - option_fn, - model, - predicates=all_predicates, - max_policy_options=max_opts, - on_step=_report_step) - # Verdict INSIDE the scope: certificate probes must judge on - # the env the rollout ran on. - evaluator = _resolve_task_evaluator(ctx, task_label) - verdict: Optional[Dict[str, Any]] = None - if evaluator is not None and len(eval_collector.states) > 1: - try: - verdict = evaluate_states_with(evaluator, - eval_collector.states, - eval_collector.labels, - sim_env=getattr( - ctx.option_model, - "sim_env", None)) - except Exception as e: # pylint: disable=broad-except - logging.debug("Task-evaluator verdict failed: %s", e) - - goal_reached = result.goal_reached - within_horizon = (result.actions_to_goal is not None - and result.actions_to_goal <= horizon) - if result.policy_error is not None: - lines.append(f"POLICY ERROR (ended the episode): " - f"{result.policy_error}") - n_surfaced = sum(1 for s in result.steps - if s.failure_reason is not None) - if n_surfaced: - lines.append( - f"{n_surfaced} option failure(s) were surfaced to the " - "policy during this rollout (recovery attempts included " - "above).") - # Closed-loop: recovered option failures do NOT disqualify - the - # policy handling them is the point. Only the goal, the horizon, - # and policy-code errors gate. - goal_achieved = (goal_reached and within_horizon - and result.policy_error is None) - evaluator_rejected = (verdict is not None and not verdict["legitimate"] - and not eval_collector.coarse) - reward_hack = (evaluator_rejected and verdict is not None - and verdict["terminated"]) - - def _policy_validation_rollout() -> Tuple[bool, str]: - fn, err = _fresh_option_fn() - if err is not None or fn is None: - return False, f"policy failed to load: {err}" - v_collector = _EvalStateCollector(model, task.init) - r = execute_policy_forward(task, - fn, - model, - predicates=all_predicates, - max_policy_options=max_opts, - on_step=v_collector.on_step) - if r.policy_error is not None: - return False, f"policy error: {r.policy_error}" - if not r.goal_reached: - missing = _missing_goal_atoms(task, r.final_state) - missing_str = ", ".join(str(a) for a in sorted(missing)) - detail = f" (missing: {{{missing_str}}})" if missing else "" - return False, f"goal not reached{detail}" - if not (r.actions_to_goal is not None - and r.actions_to_goal <= horizon): - return False, (f"goal reached only after {r.actions_to_goal} " - f"low-level steps, past the episode horizon " - f"({horizon})") - if (evaluator is not None and len(v_collector.states) > 1 - and not v_collector.coarse): - try: - v = evaluate_states_with(evaluator, - v_collector.states, - v_collector.labels, - sim_env=getattr( - ctx.option_model, "sim_env", - None)) - if not v["legitimate"]: - return False, ( - "this rollout reached the goal atoms but the " - "task evaluator scored it as a non-solve " - f"(solved=False, reward={v['reward']:.2f})") - except Exception as e: # pylint: disable=broad-except - logging.debug("Validation-rollout verdict failed: %s", e) - return True, "" - - flaky_detail: Optional[str] = None - validation_note = "" - n_rollouts = max(1, validation_cfg.rollouts) - capture_task_key = _capture_task_key(ctx) - if capture_task_key in ctx.flaky_capture_task_keys: - n_rollouts = max(n_rollouts, validation_cfg.rollouts_after_flaky) - if requested_rollouts is not None: - n_rollouts = max(n_rollouts, - min(requested_rollouts, _MAX_REQUESTED_ROLLOUTS)) - # fresh_scope computed above, shared with rollout 1. - rollout_outcomes: List[str] = [] - base_planner_seed = CFG.seed - if (ctx.capture_goal_reaching_plans and goal_achieved - and not evaluator_rejected and n_rollouts > 1): - - def _policy_repeat_rollout(repeat_idx: int) -> Tuple[bool, str]: - with (fresh_scope() if fresh_scope is not None else - contextlib.nullcontext()), \ - decorrelated_rollout_seed(repeat_idx - 1): - return _policy_validation_rollout() - - repeat_indices = list(range(2, n_rollouts + 1)) - repeat_prefetched = _prefetch_parallel([ - functools.partial(_policy_repeat_rollout, k) - for k in repeat_indices - ], "policy repeat rollouts") - for pos, repeat_idx in enumerate(repeat_indices): - ctx.attempt_rollout_count += 1 - pre = repeat_prefetched[pos] - ok, why = (pre if pre is not None else - _policy_repeat_rollout(repeat_idx)) - repeat_seed = base_planner_seed + repeat_idx - 1 - if ok: - rollout_outcomes.append( - f"rollout {repeat_idx} (planner seed " - f"{repeat_seed}): goal reached") - else: - rollout_outcomes.append( - f"rollout {repeat_idx} (planner seed " - f"{repeat_seed}): FAILED - {why}") - if flaky_detail is None: - flaky_detail = (f"rollout {repeat_idx}/{n_rollouts} " - f"(planner seed {repeat_seed}) " - f"FAILED: {why}") - if flaky_detail is None: - validation_note = ( - f" Validated {n_rollouts}/{n_rollouts} rollouts " - f"(planner seeds {base_planner_seed}-" - f"{base_planner_seed + n_rollouts - 1}; fresh env and " - "fresh policy memory per rollout).") - - # Parameter-margin gates, mirroring submit_plan (one - # shared code path: see _parameter_margin_sweep). - param_sensitive_detail: Optional[str] = None - margin_outcomes: List[str] = [] - if (fresh_scope is not None and ctx.capture_goal_reaching_plans - and goal_achieved and not evaluator_rejected - and flaky_detail is None): - margin_outcomes, param_sensitive_detail, margin_note = \ - _parameter_margin_sweep(ctx, validation_cfg, fresh_scope, - _policy_validation_rollout, "policy") - validation_note += margin_note - - capture_outcome = _decide_capture( - capture_enabled=(ctx.capture_goal_reaching_plans - and ctx.policy_capture_mode), - is_current_task=True, - have_plan=True, - goal_achieved=goal_achieved, - evaluator_rejected=evaluator_rejected, - reward_hack=reward_hack, - flaky=flaky_detail is not None, - best_effort_mode=ctx.capture_best_effort_plan, - have_validated_capture=bool(ctx.solved_plan_reached_goal), - param_sensitive=param_sensitive_detail is not None) - decision = capture_outcome.decision - captured = capture_outcome.captured - if captured: - validated_solve = decision is CaptureDecision.VALIDATED_CAPTURE - ctx.solved_plan = None - ctx.solved_sketch = None - ctx.solved_policy_source = policy_source - ctx.solved_plan_reached_goal = validated_solve - ctx.solved_plan_eval_reward = (float(verdict["reward"]) - if verdict is not None else None) - summary_bits = [ - f"validation: " - f"{1 + sum(1 for o in rollout_outcomes if 'FAILED' not in o)}" - f"/{1 + len(rollout_outcomes)} rollouts ok" - ] - if flaky_detail is not None: - summary_bits.append(f"first failure: {flaky_detail}") - if margin_outcomes: - n_margin_ok = sum(1 for o in margin_outcomes - if "FAILED" not in o) - summary_bits.append(f"physics margin: {n_margin_ok}/" - f"{len(margin_outcomes)} points ok") - ctx.solved_plan_validation_summary = "; ".join(summary_bits) - reason = capture_outcome.best_effort_reason - best_effort_note = "" - if reason is not None: - best_effort_note = ( - " (best-effort: accepted because the attempt budget is " - "exhausted; it executes for its honest reward but may " - "not count as a solve)") - lines.append( - f"Captured policy.py as the current answer{best_effort_note}" - f": {len(result.steps)} option(s) in the capture rollout." - f"{validation_note}") - elif decision is CaptureDecision.FLAKY_NO_CAPTURE: - ctx.flaky_capture_task_keys.add(capture_task_key) - per_rollout = "\n".join(f" {o}" for o in rollout_outcomes) - n_ok = 1 + sum(1 for o in rollout_outcomes if "FAILED" not in o) - lines.append( - f"FLAKY (policy NOT captured): rollout 1 reached the goal " - f"but {flaky_detail}. Per-rollout outcomes (estimated " - f"reliability {n_ok}/{n_rollouts}):\n{per_rollout}\n" - "A closed-loop policy that cannot recover in some rollouts " - "needs better feedback handling - inspect the failing " - "seeds with rollout_seed and strengthen the recovery " - "branches.") - elif decision is CaptureDecision.PARAM_SENSITIVE_NO_CAPTURE: - per_point = "\n".join(f" {o}" for o in margin_outcomes) - lines.append(f"PARAM-SENSITIVE (policy NOT captured): failed " - f"{param_sensitive_detail}. Per-point outcomes:\n" - f"{per_point}") - elif decision is CaptureDecision.REWARD_HACK_NO_CAPTURE: - lines.append( - "NOT captured: the rollout reaches the goal atoms but the " - "task evaluator scores it as a non-solve, and the real env " - "applies the same scoring.") - - lines.append(f"Goal achieved: {goal_reached}") - if verdict is not None: - lines.append( - _format_evaluator_verdict(verdict, - coarse=eval_collector.coarse)) - if goal_reached and not within_horizon: - lines.append( - f"NOT EXECUTABLE: reaching the goal takes " - f"{result.actions_to_goal} low-level steps but the episode " - f"horizon is {horizon}.") - if not goal_reached: - missing = _missing_goal_atoms(task, result.final_state) - missing_str = ", ".join(str(a) for a in sorted(missing)) - lines.append(f"Missing goal atoms: {{{missing_str}}}") - if saved_image_paths: - lines.append("\nSaved images:") - lines.extend(f" {p}" for p in saved_image_paths) - return _text_result("\n".join(lines) + - _budget_footer(ctx, rollouts_before)) - - return { - "submit_plan": submit_plan, - "submit_policy": submit_policy, - } diff --git a/predicators/agent_sdk/tools/verdicts.py b/predicators/agent_sdk/tools/verdicts.py index 771283d477..63cf08d087 100644 --- a/predicators/agent_sdk/tools/verdicts.py +++ b/predicators/agent_sdk/tools/verdicts.py @@ -16,11 +16,11 @@ class _EvalStateCollector: """Per-step states + option labels of one rollout, for evaluator verdicts. The single collector behind every surface that scores a belief-sim - rollout (``submit_plan``'s first and validation rollouts, - ``_belief_rollout_verdict``). The cascade certificate needs per-step - states (topple-onset analysis); option-boundary states give garbage - verdicts, so prefer the option model's ``last_trajectory`` and flag - the verdict as coarse (``self.coarse``) when it is unavailable. + rollout (``_belief_rollout_verdict``). The cascade certificate needs + per-step states (topple-onset analysis); option-boundary states give + garbage verdicts, so prefer the option model's ``last_trajectory`` + and flag the verdict as coarse (``self.coarse``) when it is + unavailable. """ def __init__(self, option_model: Any, init_state: State) -> None: @@ -65,12 +65,11 @@ def evaluate_states_with(evaluator: Any, ``note`` is the agent-facing sentence naming what a physics- replaying certificate simulated and on what substrate (see ``TaskEvaluator.verdict_note``); ``legitimate``/``reason`` are - HARNESS-INTERNAL (capture gating, - logs), and ``terminated`` is agent-computable from the public goal - atoms: agent-facing surfaces expose only the public (solved, - reward) pair - the standard RL end-of-episode observables - so the - agent must infer the scoring rules from the stated objective and - the outcomes its rollouts earn. + HARNESS-INTERNAL (logs), and ``terminated`` is agent-computable from + the public goal atoms: agent-facing surfaces expose only the public + (solved, reward) pair - the standard RL end-of-episode observables - + so the agent must infer the scoring rules from the stated objective + and the outcomes its rollouts earn. """ ok, reason = evaluator._certify(states, step_options, sim_env=sim_env) # pylint: disable=protected-access note_fn = getattr(evaluator, "verdict_note", None) @@ -174,6 +173,14 @@ def _sandbox_base(ctx: ToolContext) -> Optional[str]: return ctx.log_dir or None +def _policy_source_path(ctx: ToolContext) -> Optional[str]: + """Host path of the agent-editable ``policy.py``.""" + base = _sandbox_base(ctx) + if not base: + return None + return os.path.join(base, "policy.py") + + def _ground_samplers_path(ctx: ToolContext) -> Optional[str]: """Host path of the agent-editable ``ground_samplers.py``.""" base = _sandbox_base(ctx) diff --git a/predicators/approaches/__init__.py b/predicators/approaches/__init__.py index d1e5300eb2..502e46736b 100644 --- a/predicators/approaches/__init__.py +++ b/predicators/approaches/__init__.py @@ -1,6 +1,5 @@ """Handle creation of approaches.""" -import logging from typing import List, Set from typing import Type as TypingType @@ -16,20 +15,8 @@ # Find the subclasses. utils.import_submodules(__path__, __name__) -# Deprecated CLI aliases from before the agent approaches' model-free / -# model-based rename (2026-08-30); old commands and configs still pass -# them. -_DEPRECATED_APPROACH_ALIASES = { - "agent_planner": "agent_model_free", - "agent_bilevel": "agent_model_based", -} - def _get_approach_cls_from_name(name: str) -> TypingType[BaseApproach]: - if name in _DEPRECATED_APPROACH_ALIASES: - logging.warning("Approach name %r is deprecated; use %r.", name, - _DEPRECATED_APPROACH_ALIASES[name]) - name = _DEPRECATED_APPROACH_ALIASES[name] for cls in utils.get_all_subclasses(BaseApproach): if not cls.__abstractmethods__ and cls.get_name() == name: return cls diff --git a/predicators/approaches/agent_continual_ablation_approach.py b/predicators/approaches/agent_continual_ablation_approach.py index 88c3f924e2..e55ed6f2e6 100644 --- a/predicators/approaches/agent_continual_ablation_approach.py +++ b/predicators/approaches/agent_continual_ablation_approach.py @@ -50,8 +50,6 @@ def __init__(self, *args: Any, **kwargs: Any) -> None: disabled = ( "continual_uncertainty_decisions", "agent_sim_learn_param_uncertainty", - "agent_plan_validation_rule_param_margin", - "agent_plan_validation_physics_margin", "agent_explorer_info_seeking", "agent_explorer_info_seeking_adaptive", "agent_explorer_info_seeking_noise_aware", diff --git a/predicators/approaches/agent_continual_approach.py b/predicators/approaches/agent_continual_approach.py index d4f628425f..a07a43f84e 100644 --- a/predicators/approaches/agent_continual_approach.py +++ b/predicators/approaches/agent_continual_approach.py @@ -293,8 +293,7 @@ def _round_extra_tools(self, session: ProtocolSession) -> List[Any]: self._install_extra_synthesis_surfaces(exec_ns, base_pred_triples, inferred_hint, extra_paths) candidate_provider = self._make_candidate_probe_model_provider( - paths.simulator_file, trajectories, base_pred_triples, - inferred_hint) + paths.simulator_file, trajectories) def probe_model() -> _OptionModelBase: # Before a candidate exists the probe runs the real skill @@ -437,7 +436,7 @@ def _deploy_session_model(self, session: ProtocolSession, inferred_hint: Dict[str, List[str]], paths: Any, extra_paths: Dict[str, str]) -> None: loaded = self._load_synthesis_artifacts(trajectories, inferred_hint, - paths, extra_paths, {}) + paths, extra_paths) if loaded is None: # No loadable simulator.py this session: the prior model, if # any, still stands. diff --git a/predicators/approaches/agent_continual_real_to_sim_approach.py b/predicators/approaches/agent_continual_real_to_sim_approach.py index 4e8dd3e97e..202f6863cb 100644 --- a/predicators/approaches/agent_continual_real_to_sim_approach.py +++ b/predicators/approaches/agent_continual_real_to_sim_approach.py @@ -27,8 +27,6 @@ _UNCERTAINTY_FLAGS = ( "continual_uncertainty_decisions", "agent_sim_learn_param_uncertainty", - "agent_plan_validation_rule_param_margin", - "agent_plan_validation_physics_margin", "agent_explorer_info_seeking", "agent_explorer_info_seeking_adaptive", "agent_explorer_info_seeking_noise_aware", diff --git a/predicators/approaches/agent_model_based_approach.py b/predicators/approaches/agent_model_based_approach.py deleted file mode 100644 index 1b91dad592..0000000000 --- a/predicators/approaches/agent_model_based_approach.py +++ /dev/null @@ -1,1299 +0,0 @@ -"""Agent model-based approach: the agent delivers a simulator-validated plan. - -The agent plans a sequence of parameterized skills with object bindings, -subgoal atoms after each step, and continuous parameters, and must -DELIVER it as an ``submit_plan`` capture on the current task - -nothing it did not validate in the simulator (the model) is ever -executed. A backtracking parameter search remains available to the agent -as a probe method (``sim.refine``) and to mid-episode -suffix replans, but there is no approach-side refinement of unvalidated -sketches. - -Registered under the CLI approach name ``agent_model_based`` -(``agent_bilevel`` is kept as a deprecated alias). - -Example command:: - - python predicators/main.py --env pybullet_domino \ - --approach agent_model_based --seed 0 \ - --num_train_tasks 1 --num_test_tasks 1 \ - --num_online_learning_cycles 1 --explorer agent_model_free -""" -import dataclasses -import hashlib -import logging -import os -import time -from typing import Any, Callable, Dict, List, Optional, Sequence, Set, Tuple - -import numpy as np - -from predicators import utils -from predicators.agent_sdk import bilevel_sketch -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error -from predicators.agent_sdk.sketch_types import SketchStep as _SketchStep -from predicators.agent_sdk.tools import BUILTIN_TOOLS, load_ground_sampler_fns -from predicators.approaches import ApproachFailure -from predicators.approaches.agent_model_free_approach import \ - AgentModelFreeApproach -from predicators.execution_monitoring.subgoal_annotations_monitor import \ - SubgoalExecutionStatus -from predicators.settings import CFG -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, _Option - -# Fraction of agent_solve_attempt_wall_clock below which the attempt's -# budget counts as spent, not merely close to it: agents that watch the -# [budget] footer wrap up shortly BEFORE the deadline, and a query whose -# remaining tools would only refuse has nothing left to give. Used to -# label why an attempt ended (see _attempt_end_reason). -_SPENT_WALL_FRACTION = 0.2 - -# Cap on the natural-language goal text in a task's journal entry: long -# enough for any real goal_nl, short enough that a runaway goal string -# cannot crowd the journal's entry budget. -_JOURNAL_GOAL_MAX_CHARS = 400 - -# Final-submission nudge (see _nudge_final_submission). It runs once per -# task, on the LAST attempt, so it accepts a plan that falls short of a -# validated solve: the restarts are gone and a partial plan beats -# forfeiting the task. -_FINAL_SUBMIT_NUDGE = ( - "You are out of exploration budget for this attempt. Do NOT explore " - "further. In as few tool calls as possible, submit your single best " - "plan NOW via submit_plan on the current task (omit " - "task_idx), using the best parameters you have already validated. " - "It is captured as your answer even if it does not fully reach the " - "goal or does not score as a solve; then finish.") - -# Policy-mode variant: a best-effort POLICY is genuinely better than a -# best-effort plan - it is closed-loop, so whatever recovery logic it -# carries still applies at execution. -_FINAL_SUBMIT_POLICY_NUDGE = ( - "You are out of exploration budget for this attempt. Do NOT explore " - "further. In as few tool calls as possible, submit your current best " - "./policy.py NOW via submit_policy on the current task. It is " - "captured as your answer even if it does not fully reach the goal or " - "does not score as a solve; then finish.") - - -@dataclasses.dataclass -class _CaptureInfo: - """Metadata of the most recently consumed captured plan. - - Recorded by :meth:`AgentModelBasedApproach._consume_validated_plan` so - the restart loop can distinguish a validated solve (return - immediately) from a best-effort capture (bank it, rank across - attempts by evaluator reward) and journal the plan. - """ - validated: bool - reward: Optional[float] - plan_lines: List[str] - # One-line capture-time validation record (rollout tally, first - # failing step, physics-margin tally), journaled so a later - # fresh-context attempt sees HOW reliable the capture was (e.g. - # "8/10 rollouts ok; first failure: ... step 25 (Place) ...") - # instead of only that it exists. - validation_summary: Optional[str] = None - - -class AgentModelBasedApproach(AgentModelFreeApproach): - """Model-based planning: the agent plans a skeleton with subgoals and - parameters and submits it as a simulator-validated capture. - - Extends AgentModelFreeApproach - reuses agent session, tools, - trajectory management, exploration, save/load. Overrides solving - with the capture-only query loop plus the restart/journal machinery. - """ - - def __init__(self, *args: Any, **kwargs: Any) -> None: - super().__init__(*args, **kwargs) - if CFG.agent_bilevel_max_execution_replans > 0 and \ - CFG.execution_monitor != "subgoal_annotations": - raise ValueError( - "agent_bilevel_max_execution_replans > 0 requires " - "--execution_monitor subgoal_annotations (got " - f"{CFG.execution_monitor!r}): divergence detection lives " - "in the execution monitor, so without it test execution " - "is silently open-loop.") - if CFG.agent_solve_policy_mode: - if CFG.agent_bilevel_max_execution_replans > 0: - raise ValueError( - "agent_solve_policy_mode is mutually exclusive with " - "agent_bilevel_max_execution_replans > 0: the policy " - "OWNS closed-loop recovery (option failures are " - "surfaced to it), so the sketch-divergence replan " - "machinery must be off.") - if not CFG.agent_planner_use_simulator: - raise ValueError( - "agent_solve_policy_mode requires " - "agent_planner_use_simulator: the policy is validated " - "in the belief model before execution.") - # Live status of the currently executing annotated plan, exported - # to the subgoal_annotations execution monitor. None whenever no - # monitored plan is active (exploration, replanning disabled). - self._exec_status: Optional[SubgoalExecutionStatus] = None - # The grounded option plan behind _exec_status, kept so a - # divergence with no refinable suffix can resume the remaining - # not-yet-executed options open-loop (the dispensed policy holds - # them only in its closure). Set/cleared alongside _exec_status. - self._exec_plan: Optional[List[_Option]] = None - # Per-episode replan budget, refreshed by reset_for_new_episode. - self._exec_replans_left = 0 - # Whether the most recent sketch query ended because the agent hit - # agent_sdk_max_agent_turns_per_iteration. Set by - # _query_agent_for_plan_sketch; every query ending is terminal for - # its attempt, so this only labels WHY the attempt ended (see - # _attempt_end_reason) in the logs and the journal. - self._last_sketch_query_hit_turn_cap = False - # Why the last capture-less attempt ended (_attempt_end_reason), - # sampled inside the attempt while its budget signals are still - # armed - _solve clears them before writing the journal entry. - self._last_attempt_end_reason = "" - # Metadata of the last capture consumed by - # _consume_validated_plan; read by the restart loop in _solve. - self._last_capture_info: Optional[_CaptureInfo] = None - # Tasks whose goal + init-state journal entry is already written - # (one context entry per task, at the top of its section). - self._journal_task_context_recorded: Set[Any] = set() - # Snapshot of that set at begin_test_phase: test-task keys are - # rolled back with the journal itself, so a later evaluation - # (whose entries were removed) re-writes its context entries. - self._pre_test_journal_context_keys: Optional[Set[Any]] = None - - @classmethod - def get_name(cls) -> str: - return "agent_model_based" - - # ------------------------------------------------------------------ # - # Execution monitoring (closed-loop test execution) - # ------------------------------------------------------------------ # - - def reset_for_new_episode(self) -> None: - super().reset_for_new_episode() - self._exec_status = None - self._exec_plan = None - self._exec_replans_left = CFG.agent_bilevel_max_execution_replans - # Optionally give each test solve a fresh agent conversation. reset() - # fires once per test task (not on mid-episode replans, which go - # through step()); the next query lazily rebuilds the session with the - # same sandbox + artifacts but empty chat context. Test-phase only, so - # exploration episodes keep their shared session. - if CFG.agent_fresh_session_per_test_task and self._in_test_phase: - self._close_agent_session() - - def get_execution_monitoring_info(self) -> List[Any]: - if self._exec_status is None: - return [] - return [self._exec_status] - - def begin_test_phase(self) -> None: - super().begin_test_phase() - self._pre_test_journal_context_keys = set( - self._journal_task_context_recorded) - - def end_test_phase(self) -> None: - super().end_test_phase() - # The journal rollback removed this evaluation's entries, so its - # task-context dedup keys must go too - the same test tasks are - # re-solved next evaluation and need fresh goal + init entries. - if self._pre_test_journal_context_keys is not None: - self._journal_task_context_recorded = \ - self._pre_test_journal_context_keys - self._pre_test_journal_context_keys = None - - # ------------------------------------------------------------------ # - # Agent session hooks - # ------------------------------------------------------------------ # - - def _get_synthesis_tool_names(self) -> Optional[List[str]]: - """No synthesis phase in this approach - declare an empty set.""" - return [] - - # ------------------------------------------------------------------ # - # System prompt - # ------------------------------------------------------------------ # - - def _get_agent_system_prompt(self) -> str: - # Sessions are per-phase (see _ensure_agent_session: a phase - # change closes and rebuilds the session), so a solve session - # only ever receives solve queries and an explore session only - # explore queries: each phase's system prompt states just its - # own deliverable contract and rules. The query carries the task - # data and the run state (sketch_prompts.build_solve_prompt). - return bilevel_sketch.build_solve_system_prompt( - explore=self._explore_phase, - policy_mode=CFG.agent_solve_policy_mode, - propose_params=CFG.agent_bilevel_use_llm_initial_params, - ground_samplers=CFG.agent_bilevel_ground_samplers, - physics_margin=CFG.agent_plan_validation_physics_margin, - rule_param_margin=CFG.agent_plan_validation_rule_param_margin, - necessity=CFG.agent_plan_validation_necessity, - use_journal=CFG.agent_solve_use_journal, - execute_certified_plan=CFG.agent_explorer_execute_certified_plan, - early_stop_note=bilevel_sketch.build_early_stop_note(), - policy_max_options=CFG.agent_policy_max_options, - policy_max_repeated_failures=CFG. - agent_policy_max_repeated_failures, - policy_max_repeated_noops=CFG.agent_policy_max_repeated_noops, - ) - - # ------------------------------------------------------------------ # - # Solve prompt (no continuous params, subgoal format) - # ------------------------------------------------------------------ # - - def _build_solve_prompt(self, task: Task) -> str: - """Build prompt asking for a plan sketch without continuous params.""" - journal_text = "" - strategy_text = "" - attempts_text = "" - if CFG.agent_solve_use_journal: - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk import journal as journal_mod - journal_text = journal_mod.read_journal( - self._tool_context.sandbox_dir) - attempts_text = journal_mod.read_journal( - self._tool_context.sandbox_dir, - filename=journal_mod.ATTEMPTS_FILENAME) - strategy_text = journal_mod.read_strategy( - self._tool_context.sandbox_dir) - return bilevel_sketch.build_solve_prompt( - task, - all_predicates=self._get_all_predicates(), - all_options=self._get_all_options(), - trajectory_summary=self._build_trajectory_summary(), - tool_names=self._solve_prompt_tool_names(), - initial_image_section=self._initial_image_section(), - propose_params=CFG.agent_bilevel_use_llm_initial_params, - require_tool_validation=True, - journal=journal_text, - strategy=strategy_text, - attempts=attempts_text, - ) - - def _solve_prompt_tool_names(self) -> Optional[List[str]]: - """Tool list advertised in the solve prompt's "Available Tools". - - Mirrors what the explore prompt lists (the explorer renders - ``agent_session.tool_names``): the same MCP subset *plus* the - sandbox's built-in tools (Bash/Read/Write/...). The built-ins are - only actually granted under the local or docker sandbox -- which - is exactly when ``LocalSandboxSessionManager.tool_names`` prepends - them -- so they are advertised only then. Without a sandbox the - list is the bare MCP subset, unchanged. - """ - names = self._get_solve_tool_names() - if names is None: - return None - if CFG.agent_sdk_use_local_sandbox or CFG.agent_sdk_use_docker_sandbox: - return list(BUILTIN_TOOLS) + names - return names - - # ------------------------------------------------------------------ # - # Solving - # ------------------------------------------------------------------ # - - def _solve(self, task: Task, timeout: int) -> Callable[[State], Action]: - replan_policy = self._maybe_replan_from_divergence(task, timeout) - if replan_policy is not None: - return replan_policy - ctx = self._tool_context - self._record_task_context_in_journal(task) - max_attempts = max(1, CFG.agent_solve_max_attempts) - wall_clock = CFG.agent_solve_attempt_wall_clock - # Best best-effort capture across attempts, ranked by evaluator - # reward. A validated (evaluator-solved) capture returns - # immediately; only when no attempt produces one does the best - # banked policy execute for its honest reward. - best_policy: Optional[Callable[[State], Action]] = None - best_reward = -float("inf") - last_failure: Optional[ApproachFailure] = None - for attempt in range(1, max_attempts + 1): - if CFG.agent_solve_fresh_context: - # Fresh conversation per attempt (and per test task): a - # failed attempt's context carries its confidently wrong - # world model (run_20260717_230436 seed1's "hard collision - # boundary" that its identical sibling placed through), so - # a restart is the cheapest de-anchoring mechanism. Curated - # knowledge travels through the solve journal instead. - self._close_agent_session() - ctx.begin_attempt(attempt, wall_clock) - self._last_capture_info = None - policy: Optional[Callable[[State], Action]] = None - unexpected: Optional[Exception] = None - try: - policy = self._solve_attempt(task) - except ApproachFailure as e: - last_failure = e - except AgentSessionFatalError: - # The session backend is unusable (auth/billing/config); - # neither a banked capture nor further restarts can help. - # Re-raise so the run terminates (the finally still runs - # for bookkeeping). - raise - except Exception as e: # pylint: disable=broad-except - # ApproachTimeout is a SIBLING of ApproachFailure (both - # subclass ExceptionWithInfo), and env/SDK errors can - # also escape - none of them may skip the cleanup below, - # and a banked capture from an earlier attempt should - # still execute rather than be forfeited (handled after - # the finally). - unexpected = e - finally: - # Attempt bookkeeping must not outlive the attempt on ANY - # exit path (including KeyboardInterrupt): stale fields - # would append bogus [budget] footers and mislabel journal - # entries in later sessions sharing this ToolContext. - ctx.attempt_deadline = None - # Policy mode is scoped to solve attempts: left armed, it - # would silently disable submit_plan's capture - # gate for the EXPLORER's queries, which deliver sketches - # even in policy-mode configs. - ctx.policy_capture_mode = False - info = self._take_capture_info() - self._record_attempt_in_journal(attempt, max_attempts, policy, - info) - ctx.attempt_start = None - ctx.attempt_index = 0 - if unexpected is not None: - if best_policy is not None: - logging.warning( - "[%s] Solve attempt %d/%d raised %s; executing the " - "banked best-effort capture instead of forfeiting.", - self._run_id, attempt, max_attempts, unexpected) - return best_policy - raise unexpected - if policy is not None and (info is None or info.validated): - # Defensive: every capture path records metadata, so a - # missing record is treated as a validated solve rather - # than banked at unknown reward. - return policy - if policy is not None: - assert info is not None - reward = (info.reward - if info.reward is not None else -float("inf")) - if best_policy is None or reward > best_reward: - best_policy = policy - best_reward = reward - if attempt < max_attempts: - logging.info( - "[%s] Solve attempt %d/%d ended without a validated " - "solve%s; restarting with %s context.", self._run_id, - attempt, max_attempts, " (best-effort capture banked)" - if policy is not None else "", - "fresh" if CFG.agent_solve_fresh_context else "the same") - if best_policy is not None: - logging.info( - "[%s] No validated solve in %d attempt(s); executing the " - "best best-effort capture (evaluator reward %s).", - self._run_id, max_attempts, - f"{best_reward:.2f}" if best_reward > -float("inf") else "n/a") - return best_policy - if last_failure is not None: - raise last_failure - raise ApproachFailure( - f"Bilevel solve produced no captured plan in {max_attempts} " - "attempt(s).") - - def _attempt_wall_spent(self) -> bool: - """Whether the attempt's wall clock is spent (or nearly so). - - True once less than :data:`_SPENT_WALL_FRACTION` of the wall - clock remains, not merely once the deadline passes: agents that - watch the [budget] footer wrap up shortly BEFORE the deadline - (run_20260718_125643 ended an attempt with 2 minutes left), and - an attempt whose remaining tools would only refuse is spent in - every sense that matters. Tool-side refusals keep using the - exact deadline, so the closing minutes still allow submissions. - """ - deadline = self._tool_context.attempt_deadline - if deadline is None: - return False - floor = _SPENT_WALL_FRACTION * CFG.agent_solve_attempt_wall_clock - return time.monotonic() > deadline - floor - - def _attempt_end_reason(self) -> str: - """Why the attempt's single query ended, for logs and journal.""" - if self._last_sketch_query_hit_turn_cap: - return "turn cap" - if self._attempt_wall_spent(): - return "wall clock spent" - return "no submission" - - def _take_capture_info(self) -> Optional[_CaptureInfo]: - """Pop the metadata _consume_validated_plan recorded (or None). - - An accessor rather than a bare attribute read: the attribute is - set as a side effect of _solve_attempt, which mypy cannot see, - so reading it directly right after ``= None`` is flagged - unreachable. - """ - info = self._last_capture_info - self._last_capture_info = None - return info - - def _append_journal_auto_entry(self, header: str, - body_lines: List[str]) -> bool: - """Best-effort append of a harness-written solve-journal entry. - - Callers guard with :meth:`_journal_active`. Returns True on a - successful write; a failed write is logged, never raised - the - journal must not be able to fail a solve. - """ - sandbox_dir = self._tool_context.sandbox_dir - assert self._journal_active() and sandbox_dir - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk import journal as journal_mod - try: - journal_mod.append_entry( - sandbox_dir, - header, - "\n".join(body_lines), - max_chars=journal_mod.MAX_AUTO_ENTRY_CHARS, - filename=journal_mod.ATTEMPTS_FILENAME) - except OSError as e: - logging.warning("Journal entry %r failed: %s", header, e) - return False - return True - - def _journal_task_label(self) -> str: - """The journal's label for the task currently being solved. - - Carries the HARNESS cycle number so the attempt log anchors a - canonical numbering: without it agents invented their own cycle - counts (a journal's "cycle 4" was the harness's cycle 1), and - cross-referencing run logs against the journal needed a mental - offset. - """ - idx = self._tool_context.test_task_idx - task_part = f"task {idx}" if idx is not None else "task ?" - return f"cycle {self._tool_context.iteration_id} {task_part}" - - def _record_task_context_in_journal(self, task: Task) -> None: - """Append the task's goal + init-state entry, once per task. - - Written at the START of the task's first attempt so it tops the - task's journal section - above even the agent's own in-attempt - notes - keeping every later entry interpretable: a recorded - plan's geometry is only meaningful relative to its layout. The - init dict uses the exact representation the solve prompt shows - (including its excluded-objects filtering). - """ - if not self._journal_active(): - return - key = (self._tool_context.test_task_idx, id(task)) - if key in self._journal_task_context_recorded: - return - if task.goal_nl: - goal_txt = " ".join(task.goal_nl.split()) - if len(goal_txt) > _JOURNAL_GOAL_MAX_CHARS: - goal_txt = goal_txt[:_JOURNAL_GOAL_MAX_CHARS].rstrip() + "..." - else: - goal_txt = ", ".join(str(a) for a in sorted(task.goal, key=str)) - body = [f"- goal: {goal_txt}", "- initial state features:"] - body.extend(f" {line}" - for line in task.init.dict_str(indent=2).splitlines()) - header = f"{self._journal_task_label()} goal + initial state (auto)" - if self._append_journal_auto_entry(header, body): - self._journal_task_context_recorded.add(key) - - def _record_attempt_in_journal(self, attempt: int, max_attempts: int, - policy: Optional[Any], - info: Optional[_CaptureInfo]) -> None: - """Auto-append this attempt's factual record to the attempt log. - - The harness-written record (outcome, budget spent, captured or - best refused plan) guarantees the essentials of every attempt - are on record even when the agent writes nothing; the agent's - own lessons live in journal.md, which it edits directly. - """ - if not self._journal_active(): - return - ctx = self._tool_context - body = [f"- outcome: {self._attempt_outcome_text(policy, info)}"] - if ctx.attempt_start is not None: - elapsed_min = (time.monotonic() - ctx.attempt_start) / 60.0 - body.append(f"- budget spent: {elapsed_min:.1f} min, " - f"{ctx.attempt_rollout_count} sim rollouts") - if info is not None and info.validation_summary: - body.append(f"- {info.validation_summary}") - if info is not None and info.plan_lines: - body.append("- captured plan:") - body.extend(f" {line}" for line in info.plan_lines) - elif ctx.best_uncaptured_plan_lines: - # Nothing captured, but the attempt's best refused submission - # (evaluator non-solve or flaky) is worth carrying: a later - # attempt - or the final best-effort nudge - can resubmit it - # instead of the work vanishing with the attempt's context. - reward_txt = (f"evaluator reward {ctx.best_uncaptured_reward:.2f}" - if ctx.best_uncaptured_reward is not None else - "no evaluator verdict") - body.append(f"- best refused submission ({reward_txt}, " - "not captured):") - body.extend(f" {line}" for line in ctx.best_uncaptured_plan_lines) - header = (f"{self._journal_task_label()} attempt " - f"{attempt}/{max_attempts} (auto)") - self._append_journal_auto_entry(header, body) - - def _attempt_outcome_text(self, policy: Optional[Any], - info: Optional[_CaptureInfo]) -> str: - """One-line outcome for an attempt's journal record.""" - if policy is None: - # The reason is the one fact a fresh-context restart cannot - # rediscover: it tells the next attempt whether the last one - # ran out of budget or talked itself out of submitting. - if self._last_attempt_end_reason: - return f"no capture ({self._last_attempt_end_reason})" - return "no capture" - if info is None or info.validated: - return "SOLVED (validated capture)" - if info.reward is not None: - return f"best-effort capture (evaluator reward {info.reward:.2f})" - return "best-effort capture (no evaluator verdict)" - - def _solve_attempt(self, task: Task) -> Callable[[State], Action]: - """One full solve attempt: a single agent query on one session. - - The attempt's budgets are the wall clock - (``agent_solve_attempt_wall_clock``) and the query's turn cap; - the only deliverable is an ``submit_plan`` capture - (consumed via :meth:`_consume_validated_plan`). - - However that query ends - a spent budget, an unparseable sketch, - or a session that simply never submitted - the attempt is over. - A second query on the same conversation would re-explore from a - context that already contains whatever went wrong - (run_20260808_113951 queries 004-009: three full-price queries - restating the same "no plan can reach the goal" argument), so the - fresh-context restart is the only retry. See :meth:`_end_attempt` - for the one exception, on the final attempt. - """ - self._sync_tool_context() - self._tool_context.current_task = task - # Let submit_plan record a goal-reaching - # plan on this task into solved_plan/solved_sketch (consumed below). - self._tool_context.capture_goal_reaching_plans = True - # Policy mode: the deliverable is policy.py via submit_policy; - # submit_plan stays available for probing but cannot - # capture. - self._tool_context.policy_capture_mode = CFG.agent_solve_policy_mode - # LLM-free bypass: a prewritten policy.py as the captured - # artifact (smoke tests / debugging the execution path). - if CFG.agent_solve_policy_mode and CFG.agent_policy_file: - with open(CFG.agent_policy_file, "r", encoding="utf-8") as f: - self._tool_context.solved_policy_source = f.read() - self._tool_context.solved_plan_reached_goal = True - policy = self._consume_validated_plan() - assert policy is not None - return policy - # Render the initial state so the agent can see the scene layout. - self._render_initial_state_image(task) - # Whether later fresh-context restarts exist after this attempt; - # decides whether this attempt pays for the final-submission nudge - # (see _end_attempt). attempt_index == 0 (no restart loop in - # flight) behaves like a final attempt. - restarts_remain = (0 < self._tool_context.attempt_index < max( - 1, CFG.agent_solve_max_attempts)) - # Clear any prior capture so we only act on this query's result. - self._tool_context.clear_plan_capture() - self._last_sketch_query_hit_turn_cap = False - self._last_attempt_end_reason = "" - try: - self._query_agent_for_plan_sketch(task) - except AgentSessionFatalError: - raise - except Exception as e: # pylint: disable=broad-except - # The agent may have validated a working plan via - # submit_plan even if its final text didn't parse. - policy = self._consume_validated_plan() - if policy is not None: - return policy - logging.warning("[%s] Solve query failed: %s", self._run_id, e) - else: - # Fast path: the agent already refined + forward-validated - # a plan on this task via submit_plan - return it - # directly instead of re-refining the (possibly different) - # final-text sketch. - policy = self._consume_validated_plan() - if policy is not None: - return policy - # The agent must itself reach a confirmed - # submit_plan capture (consumed above) so we - # never execute a plan it didn't verify. - logging.info("[%s] Query ended without a validated plan.", - self._run_id) - # Sample the end reason before _end_attempt: the nudge suspends - # the attempt deadline, and _solve clears it outright before the - # journal entry is written. - self._last_attempt_end_reason = self._attempt_end_reason() - policy = self._end_attempt(restarts_remain) - if policy is not None: - return policy - raise ApproachFailure("Bilevel solve failed: the attempt's agent " - "query produced no captured plan " - f"({self._last_attempt_end_reason}).") - - def _end_attempt( - self, - restarts_remain: bool) -> Optional[Callable[[State], Action]]: - """End an attempt that produced no capture. - - With later fresh-context restarts remaining, end with NO nudge - (return None): the restart is the retry, and the journal - auto-entry already records the attempt's best refused - submission. On the final attempt the best-effort submission - nudge is the ultimate fallback - return whatever policy it - captures (None when even that yields nothing, giving up on the - task). - """ - if restarts_remain: - return None - return self._nudge_final_submission() - - # ------------------------------------------------------------------ # - # Plan sketch extraction - # ------------------------------------------------------------------ # - - def _query_agent_for_plan_sketch(self, task: Task) -> List[_SketchStep]: - """Query agent for a plan sketch and parse it.""" - sketch_file = CFG.agent_bilevel_plan_sketch_file - if sketch_file: - # An absolute path is used as-is; a bare filename is resolved - # against the configured plan-sketch directory under scripts/. - if os.path.isabs(sketch_file): - filepath = sketch_file - else: - filepath = ( - f"{utils.get_path_to_predicators_root()}/scripts/" - f"{CFG.agent_bilevel_plan_sketch_dir}/{sketch_file}") - with open(filepath, "r", encoding="utf-8") as f: - plan_text = f.read().strip() - logging.info("Loaded plan sketch from file: %s", sketch_file) - else: - prompt = self._build_solve_prompt(task) - responses = self._query_agent_sync(prompt, kind="test") - dead = query_fatal_error(responses) - if dead is not None: - # An outage is not a failed attempt: recording 0/1 here - # would write a bogus eval datapoint. Stop the run; the - # relaunch re-runs this cycle's test. - raise AgentSessionFatalError( - "test query died without the agent doing any work " - f"({dead}); not recording this attempt as a failure.") - # Record cap-exhaustion before parsing: a capped session usually - # has no final text, so the "empty plan text" failure below is - # still attributable to the turn cap by _attempt_end_reason. - self._last_sketch_query_hit_turn_cap = \ - self._responses_hit_turn_cap(responses) - plan_text = self._extract_option_plan_text(responses) - - if not plan_text: - raise ApproachFailure("Agent returned empty plan text.") - - # Tolerant parse of the agent's final text; named `~ my_sampler` - # references resolve against the sandbox's ground_samplers.py (a - # broken file just drops the annotations here - this is the - # best-effort fallback path, not the strict tool path). - gs_fns, gs_err = load_ground_sampler_fns(self._tool_context) - if gs_err is not None: - logging.warning("[%s] %s", self._run_id, gs_err) - sketch = bilevel_sketch.parse_sketch_from_text( - plan_text, - task, - predicates=self._get_all_predicates(), - options=self._get_all_options(), - types=self._types, - parse_continuous_params=CFG.agent_bilevel_use_llm_initial_params, - parse_ground_samplers=CFG.agent_bilevel_ground_samplers, - ground_sampler_fns=gs_fns or None, - ) - - if not sketch: - option_names = sorted(o.name for o in self._get_all_options()) - raise ApproachFailure(f"Parsed empty plan sketch from agent.\n" - f" Plan text:\n{plan_text}\n" - f" Available option names: {option_names}") - - logging.info( - "[%s] Agent produced sketch with %d steps, %d with " - "subgoals.", self._run_id, len(sketch), - sum(1 for s in sketch if s.subgoal_atoms)) - return sketch - - @staticmethod - def _responses_hit_turn_cap(responses: List[Dict[str, Any]]) -> bool: - """Whether a query's response stream ended on the SDK turn cap. - - The SDK reports the cap as result subtype ``error_max_turns``; - the num_turns comparison is a fallback for backends whose result - entries lack the subtype field. - """ - max_turns = CFG.agent_sdk_max_agent_turns_per_iteration - for entry in responses: - if entry.get("type") != "result": - continue - if entry.get("subtype") == "error_max_turns": - return True - num_turns = entry.get("num_turns") - if num_turns is not None and num_turns >= max_turns: - return True - return False - - # ------------------------------------------------------------------ # - # Backtracking refinement (used by mid-episode suffix replans) - # ------------------------------------------------------------------ # - - def _refine_sketch( - self, - task: Task, - sketch: List[_SketchStep], - timeout: float, - attempt: int = 0, - on_step_fail: Optional[Callable[[int, List[Optional[_Option]], str], - None]] = None, - ) -> Tuple[List[_Option], bool]: - """Backtracking search over continuous parameters for a plan sketch. - - Returns ``(plan, success)``. On success, ``plan`` is a list of - grounded options that achieves the task goal. On failure, - ``plan`` is the longest partial refinement found. - - This is the approach-flavored entry to - ``bilevel_sketch.refine_sketch`` (which stays settings-free): - it gathers the approach-owned inputs (option model, predicates, - samplers, run id), reads the search knobs from ``CFG``, and - first passes the task through :meth:`_attach_initial_latent` so - partially-observable approaches can seed ``task.init.latent`` - with the initial latent block. Used by mid-episode suffix - replans and by the offline replay scripts under - ``scripts/domino_debug/``. - - ``attempt`` perturbs the RNG so retries explore different - samples - without it, refinement is deterministic in - ``CFG.seed`` and a forward-validation failure would loop on - the identical plan. ``on_step_fail`` is forwarded to the search - (called with the step index, the partial plan, and the failure - reason whenever a step fails to refine). - """ - task = self._attach_initial_latent(task) - assert self._option_model is not None, \ - "agent_bilevel requires a simulator " \ - "(agent_planner_use_simulator=True)." - outcome = bilevel_sketch.refine_sketch( - task, - sketch, - self._option_model, - predicates=self._get_all_predicates(), - timeout=timeout, - rng=np.random.default_rng(CFG.seed + attempt), - max_samples_per_step=CFG.agent_bilevel_max_samples_per_step, - check_subgoals=CFG.agent_bilevel_check_subgoals, - log_state=CFG.agent_bilevel_log_state, - run_id=self._run_id, - parameterized_samplers=self._get_all_samplers(), - on_step_fail=on_step_fail, - strip_latent_wait_targets=( - not self._tool_context.latent_tracking_available), - ) - return outcome.plan, outcome.success - - def _attach_initial_latent(self, task: Task) -> Task: - """Hook for partial-observability approaches to seed the latent. - - Subclasses that thread a ``latent`` state block through the - simulator (``AgentSimLearningApproach`` when the loaded rules - are recurrent) override this to attach an initial latent to - ``task.init.latent`` before refinement begins. The default - returns ``task`` unchanged - fully-observable approaches need do - nothing. - """ - return task - - def _sample_params(self, option: ParameterizedOption, _state: State, - rng: np.random.Generator) -> np.ndarray: - """Sample continuous parameters for an option.""" - return bilevel_sketch.sample_params(option, rng) - - def _parse_subgoal_annotations( - self, - text: str, - predicates: Set[Predicate], - objects: Sequence[Object], - ) -> List[Optional[Tuple[Set[GroundAtom], Set[GroundAtom]]]]: - """Shim over ``bilevel_sketch.parse_subgoal_annotations``.""" - option_names = {o.name for o in self._get_all_options()} - return bilevel_sketch.parse_subgoal_annotations( - text, predicates, objects, option_names) - - # ------------------------------------------------------------------ # - # Helpers - # ------------------------------------------------------------------ # - - def _maybe_replan_from_divergence( - self, task: Task, - timeout: int) -> Optional[Callable[[State], Action]]: - """Handle a mid-episode re-solve triggered by the subgoal_annotations - execution monitor. - - CogMan calls solve() identically at episode start and on a - monitor-triggered replan; ``_exec_status`` distinguishes them - (non-None only while a monitored plan executes; - reset_for_new_episode clears it at episode start). On a replan - ``task.init`` is the real state where the just-finished step's - annotation failed. Divergence is usually a continuous-execution - problem (a sampled parameter whose real outcome differed from - the option-model rollout), not a wrong skeleton, so we first try - to resume a suffix of the executed sketch (cheap, no agent - query; see :meth:`_replan_suffix`). When no suffix refines - or - the episode's replan budget is spent - the remaining - not-yet-executed options resume OPEN-LOOP instead of failing the - episode: an annotation is the agent's prediction, not proof the - goal is out of reach, and aborting a plan whose remaining - settle/cure steps might still deliver turns a maybe-fail into a - certain fail. The divergence stays in the log and the goal check - decides the episode. Set - ``CFG.agent_bilevel_replan_agent_fallback`` to instead fall - through to a fresh agent sketch query when no suffix refines. - """ - status = self._exec_status - if status is None or status.steps_initiated == 0: - return None - self._exec_status = None - exec_plan = self._exec_plan or [] - self._exec_plan = None - failed_idx = status.steps_initiated - 1 - steps = list(status.sketch) - failed_name = steps[failed_idx].option.name - if self._exec_replans_left > 0: - self._exec_replans_left -= 1 - logging.info( - "Subgoal divergence after step %d (%s). Replanning from the " - "current state (%d execution replans left).", failed_idx, - failed_name, self._exec_replans_left) - policy = self._replan_suffix(task.init, task, steps, failed_idx, - timeout) - if policy is not None: - return policy - if CFG.agent_bilevel_replan_agent_fallback: - # No suffix of the executed skeleton refines from here; - # fall through to pay for a fresh agent sketch. - logging.info("Suffix replan failed; querying the agent for " - "a fresh sketch.") - return None - reason = "no suffix of the executed sketch refines from here" - else: - reason = "no execution replans left" - remaining = list(exec_plan[failed_idx + 1:]) - logging.warning( - "Subgoal divergence after step %d (%s): %s. Resuming the " - "remaining %d step(s) open-loop; the divergence stands " - "recorded and the goal check decides the episode.", failed_idx, - failed_name, reason, len(remaining)) - return self._plan_to_policy(remaining, sketch=steps[failed_idx + 1:]) - - def _nudge_final_submission(self) -> Optional[Callable[[State], Action]]: - """One short follow-up query on the LAST attempt, after its query ended - with no captured plan: tell the agent to submit its best plan now. - - A session that hits the turn cap mid-iteration contributes - nothing, even when it has a near-working plan in context; this - converts that dead end into a submission attempt at the cost of a - few turns. - - The submitted plan is captured and executed even if its belief - rollout does not reach the goal, is scored a non-solve by the - task evaluator, or is flaky: there are no restarts left, and a - partial plan beats forfeiting the task. - """ - nudge = (_FINAL_SUBMIT_POLICY_NUDGE - if CFG.agent_solve_policy_mode else _FINAL_SUBMIT_NUDGE) - if CFG.agent_solve_use_journal: - nudge += (" If an earlier attempt's entry in the Attempt Log " - "records a better plan (captured or refused) than " - "anything from this attempt, resubmit that plan instead." - " After the submission, append ONE short factual entry " - "to ./journal.md for later fresh-context attempts and " - "tasks: what you tried (exact parameters), the key " - "measurements, and what to try differently - facts and " - "measurements only, no verdicts like 'impossible'.") - # SUSPEND (not clear) the attempt deadline for the nudge query: - # its cooperative refusals and the sandbox interrupt backstop - # must not block the submission (or the journal entry) itself. - # Restore it afterwards rather than leaking the None: _solve's - # per-attempt bookkeeping owns clearing the deadline, and a - # helper that silently disarms the wall clock is a trap for any - # future caller that runs mid-attempt. - saved_deadline = self._tool_context.attempt_deadline - self._tool_context.attempt_deadline = None - self._tool_context.capture_best_effort_plan = True - try: - nudge_responses = self._query_agent_sync(nudge, kind="test") - dead = query_fatal_error(nudge_responses) - if dead is not None: - raise AgentSessionFatalError( - "final-submission nudge died without the agent doing " - f"any work ({dead}); not recording this attempt as a " - "failure.") - except AgentSessionFatalError: - raise - except Exception as e: # pylint: disable=broad-except - logging.warning("Final-submission nudge failed: %s", e) - finally: - self._tool_context.capture_best_effort_plan = False - self._tool_context.attempt_deadline = saved_deadline - policy = self._consume_validated_plan() - if policy is not None: - logging.info( - "[%s] Final-submission nudge produced a validated plan.", - self._run_id) - return policy - - def _consume_validated_plan(self) -> Optional[Callable[[State], Action]]: - """Return a policy from an agent-validated plan, or None. - - ``submit_plan`` records a captured (goal-reaching, validated) - plan on the current solve task into the tool context. Returning - that exact simulator-verified plan guarantees the agent's tool- - validated answer is what executes, and avoids a fresh refinement - that with a different seed might not reproduce it. - """ - capture = self._tool_context.take_plan_capture() - if capture.policy_source: - return self._policy_capture_to_policy(capture) - if not capture.plan: - return None - # A capture with reached_goal False was accepted under the - # best-effort nudge; anything else is a validated solve. - validated = capture.reached_goal is not False - lines = list( - bilevel_sketch.format_plan_lines(capture.plan, - sketch=capture.sketch)) - self._last_capture_info = _CaptureInfo( - validated=validated, - reward=capture.eval_reward, - plan_lines=lines, - validation_summary=capture.validation_summary) - verdict = ("simulator-verified" if validated else - "best-effort: not a validated solve in the belief rollout") - # Log the full plan (options + continuous params + subgoal - # annotations) so the run log shows exactly what will execute. - logging.info( - "[%s] Using agent-validated plan from capture " - "(%d steps, %s):\n%s", self._run_id, len(capture.plan), verdict, - "\n".join(lines)) - return self._plan_to_policy(capture.plan, sketch=capture.sketch) - - def _policy_capture_to_policy(self, - capture: Any) -> Callable[[State], Action]: - """Turn a captured policy.py source into the execution policy. - - Policy-mode counterpart of the plan branch below: records the - capture metadata (the journal gets the source hash + validation - summary rather than plan lines), composes the SNAPSHOTTED source - against the real task and vocabulary, and wraps it in the - closed-loop executor. No ``SubgoalExecutionStatus`` is ever - published: the monitor and the divergence-replan path stay inert - (the constructor enforces replans == 0 in policy mode). - """ - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.policy_execution import \ - build_policy_option_fn - source = capture.policy_source - validated = capture.reached_goal is not False - sha = hashlib.sha256(source.encode("utf-8")).hexdigest()[:12] - n_lines = len(source.splitlines()) - self._last_capture_info = _CaptureInfo( - validated=validated, - reward=capture.eval_reward, - plan_lines=[ - f"" - ], - validation_summary=capture.validation_summary) - verdict = ("simulator-verified" if validated else - "best-effort: not a validated solve in the belief rollout") - logging.info( - "[%s] Using agent-validated POLICY from capture (sha=%s, " - "%d lines, %s).", self._run_id, sha, n_lines, verdict) - task = self._tool_context.current_task - assert task is not None - option_fn, err = build_policy_option_fn( - source, - task, - predicates=self._get_all_predicates(), - options=self._get_all_options(), - types=self._types) - if err is not None or option_fn is None: - raise ApproachFailure( - f"Captured policy.py failed to load for execution: {err}") - return self._policy_to_execution_policy(option_fn) - - def _policy_to_execution_policy( - self, option_fn: Any) -> Callable[[State], Action]: - """Closed-loop real executor for a composed policy option fn. - - Mirrors ``execute_policy_forward``'s semantics on the real env: - option execution failures (non-initiable, a skill raising - mid-execution - e.g. a motion-planning refusal - or an option - step-cap timeout) do NOT end the episode; the failure text is - surfaced to the policy via ``memory['last_failure']`` and the - next option is requested from the current state, bounded by - ``CFG.agent_policy_max_options`` total options, - ``CFG.agent_policy_max_repeated_failures`` consecutive failures - of one identical command (the stuck-loop guard), and - ``CFG.agent_policy_max_repeated_noops`` consecutive clean - completions of one identical command that changed nothing - observable (the guard's livelock twin). ``get_option`` - bugs and DONE end the episode via ``ApproachFailure`` (harmless - when the goal already holds). - - Implementation note: each issued option still runs through - ``utils.option_policy_to_policy`` (per-option step caps and the - Wait atom-change machinery live there), but the wrapper is - REBUILT after every surfaced failure - the failed option is - stuck inside the old wrapper's closure (its terminal never - holds), so a fresh wrapper is what makes the next call request a - new option. - """ - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.policy_execution import PolicyError, \ - option_repeat_key, repeated_failure_message, \ - repeated_noop_message, states_features_allclose - predicates = self._get_all_predicates() - - def _abstract(s: State) -> Set[GroundAtom]: - return utils.abstract(s, predicates) - - issued = 0 - last_failure: Optional[str] = None - repeat_key: Optional[Any] = None - repeat_count = 0 - issued_option: Optional[_Option] = None - issued_state: Optional[State] = None - noop_key: Optional[Any] = None - noop_count = 0 - - class _PolicyFatal(utils.OptionExecutionFailure): - """DONE / policy bug / budget: never surfaced, ends episode.""" - - def _option_policy(state: State) -> _Option: - nonlocal issued, last_failure, repeat_key, repeat_count, \ - issued_option, issued_state, noop_key, noop_count - if last_failure is None and issued > 0: - # The previous option completed cleanly: the policy is - # making progress, so the stuck-loop counter resets. - repeat_key = None - repeat_count = 0 - # ...unless the clean completion changed nothing - # observable: an identical command re-completing as a - # no-op K times is a livelock the failure guard cannot - # see (mirrors execute_policy_forward). - if issued_option is not None and issued_state is not None \ - and states_features_allclose(issued_state, state): - key = option_repeat_key(issued_option) - noop_count = noop_count + 1 if key == noop_key else 1 - noop_key = key - if noop_count >= CFG.agent_policy_max_repeated_noops: - raise _PolicyFatal( - repeated_noop_message(issued_option, noop_count)) - else: - noop_key = None - noop_count = 0 - if issued >= CFG.agent_policy_max_options: - raise _PolicyFatal( - "Policy exhausted its option budget " - f"({CFG.agent_policy_max_options} options) without " - "signalling DONE.") - try: - nxt = option_fn(state, last_failure) - except PolicyError as e: - raise _PolicyFatal(f"policy error: {e}") from e - if nxt is None: - logging.info("Policy signaled DONE after %d options.", issued) - raise _PolicyFatal("Policy signaled DONE.") - issued += 1 - last_failure = None - logging.info("Executing policy option %d/%d: %s", issued, - CFG.agent_policy_max_options, nxt.simple_str()) - if not nxt.initiable(state): - # Same text and attribution as execute_policy_forward - # (option_policy_to_policy's own "Unsound option policy" - # raise would name the PREVIOUS option in its info). - raise utils.OptionExecutionFailure( - "not initiable", info={"last_failed_option": nxt}) - issued_option = nxt - issued_state = state - return nxt - - def _fresh_inner() -> Callable[[State], Action]: - return utils.option_policy_to_policy( - _option_policy, - max_option_steps=CFG.max_num_steps_option_rollout, - abstract_function=_abstract) - - inner_box = {"inner": _fresh_inner()} - - def _execution_policy(state: State) -> Action: - nonlocal last_failure, repeat_key, repeat_count - while True: - try: - return inner_box["inner"](state) - except _PolicyFatal as e: - raise ApproachFailure(str(e)) from e - except utils.OptionExecutionFailure as e: - # Surface to the policy and continue: the failed - # option is stuck in the old wrapper, so rebuild. - failed = getattr(e, "info", {}).get("last_failed_option") - prefix = (f"{failed.name}: " if failed is not None else "") - last_failure = f"{prefix}{e}" - logging.info("Option failure surfaced to the policy: %s", - last_failure) - # Mirrors execute_policy_forward: K consecutive - # failures of one identical command are a policy - # bug, not recovery - end the episode attributably - # instead of burning the remaining option budget. - if failed is not None: - key = option_repeat_key(failed) - repeat_count = (repeat_count + - 1 if key == repeat_key else 1) - repeat_key = key - if (repeat_count >= - CFG.agent_policy_max_repeated_failures): - raise ApproachFailure( - repeated_failure_message(failed, - repeat_count)) from e - else: - repeat_key = None - repeat_count = 0 - inner_box["inner"] = _fresh_inner() - - return _execution_policy - - def _plan_to_policy( - self, - plan: List[_Option], - sketch: Optional[List[_SketchStep]] = None, - ) -> Callable[[State], Action]: - """Wrap a grounded option plan into a step-by-step policy. - - With ``CFG.agent_bilevel_max_execution_replans > 0`` and a full - per-step sketch, the policy also publishes a live - ``SubgoalExecutionStatus`` (via - ``get_execution_monitoring_info``) that the subgoal_annotations - execution monitor reads to check, at each option boundary, that - the just-finished step's annotation holds in the REAL state. On - divergence the monitor makes CogMan re-invoke solve(), which - lands in :meth:`_maybe_replan_from_divergence`. - """ - predicates = self._get_all_predicates() - - def _abstract(s: State) -> Set[GroundAtom]: - return utils.abstract(s, predicates) - - monitored = (CFG.agent_bilevel_max_execution_replans > 0 - and sketch is not None and len(sketch) == len(plan)) - - queue = list(plan) - total = len(queue) - status: Optional[SubgoalExecutionStatus] = None - if monitored: - assert sketch is not None - status = SubgoalExecutionStatus(sketch=list(sketch)) - self._exec_status = status - self._exec_plan = list(plan) - - def _option_policy(state: State) -> _Option: - del state # unused - if not queue: - logging.info("Option plan exhausted after %d options.", total) - # See the twin of this raise in utils.option_plan_to_policy: - # a finished plan and a failed option arrive as the same - # exception type, so the normal terminus is flagged. - raise utils.OptionExecutionFailure( - "Option plan exhausted!", info={"plan_exhausted": True}) - option = queue.pop(0) - num_done = total - len(queue) - if status is not None: - status.steps_initiated = num_done - status.current_option = option - next_option = None if not queue else queue[0].simple_str() - logging.info("Executing option %d/%d: %s (remaining=%d, next=%s)", - num_done, total, option.simple_str(), len(queue), - next_option) - return option - - inner = utils.option_policy_to_policy( - _option_policy, - max_option_steps=CFG.max_num_steps_option_rollout, - abstract_function=_abstract) - return self._wrap_option_failures(inner) - - def _replan_suffix( - self, - state: State, - task: Task, - sketch: List[_SketchStep], - failed_idx: int, - timeout: int, - ) -> Optional[Callable[[State], Action]]: - """Cheap-first recovery: re-refine a suffix of the current sketch. - - Divergence is usually a continuous-execution problem (a sampled - parameter whose real outcome differed from the option-model - rollout), not a wrong skeleton, so before paying for a fresh - agent sketch we retry the one we have. Candidate resume points - run from the failed step backward to just after the latest - earlier annotated step whose subgoals still hold in the current - state. The holds-check only bounds the walk-back - annotations - are optional and can hold coincidentally (e.g. a final - SwitchOff's {Off} atom holds before the switch was ever touched) - - so every candidate suffix must still refine AND forward- - validate from the current state before we trust it. Returns None - when no suffix candidate validates. - """ - assert self._option_model is not None - sub_task = Task(state, task.goal) - resume_floor = 0 - for j in range(failed_idx - 1, -1, -1): - step = sketch[j] - if step.subgoal_atoms is None and step.subgoal_neg_atoms is None: - continue - pos_ok = all(a.holds(state) for a in (step.subgoal_atoms or set())) - neg_ok = not any( - a.holds(state) for a in (step.subgoal_neg_atoms or set())) - if pos_ok and neg_ok: - resume_floor = j + 1 - break - start = time.perf_counter() - for j in range(failed_idx, resume_floor - 1, -1): - remaining = timeout - (time.perf_counter() - start) - if remaining <= 0: - break - suffix = list(sketch[j:]) - plan, success = self._refine_sketch(sub_task, - suffix, - remaining, - attempt=j) - if not success: - logging.info( - "Suffix replan: refinement failed resuming at " - "step %d.", j) - continue - ok, reason = bilevel_sketch.validate_plan_forward( - sub_task, - plan, - self._option_model, - predicates=self._get_all_predicates(), - sketch=suffix, - run_id=self._run_id, - ) - if ok: - logging.info( - "Suffix replan: resuming executed sketch at step %d " - "(%d steps).", j, len(plan)) - return self._plan_to_policy(plan, sketch=suffix) - logging.info( - "Suffix replan: forward validation failed resuming at " - "step %d: %s", j, reason) - return None diff --git a/predicators/approaches/agent_model_free_approach.py b/predicators/approaches/agent_model_free_approach.py index 90e9c43e63..9c9b08913e 100644 --- a/predicators/approaches/agent_model_free_approach.py +++ b/predicators/approaches/agent_model_free_approach.py @@ -1,63 +1,32 @@ -"""Agent model-free approach: fixed-vocabulary open-loop planning. +"""The base of the agent arms: the agent session, its tool context, the +recorded trajectories and the checkpoints. -The agent plans directly from its own world knowledge - no simulator -(model) is required to validate a plan before execution. Combines online -trajectory collection (via AgentModelFreeExplorer) with open-loop option plan -generation (via Claude Agent SDK). No predicate/process/type invention - -just stores trajectories and generates plans. - -Registered under the CLI approach name ``agent_model_free`` -(``agent_planner`` is kept as a deprecated alias). - -Example command: - python predicators/main.py --env pybullet_domino \ - --approach agent_model_free --seed 0 \ - --num_train_tasks 1 --num_test_tasks 1 \ - --num_online_learning_cycles 1 --explorer agent_model_free +The arms themselves (``agent_continual_approach`` and its siblings) add +the play loop of the continual protocol through ``ContinualPlayMixin``; +this class and the simulator-learning classes between it and the arms +carry the machinery they share. None of them solves a task on its own. """ -import copy import datetime import logging import os -import subprocess -from typing import Any, Callable, Dict, List, Optional, Sequence, Set, Tuple, \ - cast +from typing import Any, Callable, Dict, List, Optional, Set, cast import dill as pkl -import numpy as np from gym.spaces import Box from predicators import utils -from predicators.agent_sdk.rendering import save_task_state_image -from predicators.agent_sdk.response_parser import extract_final_text -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error -from predicators.agent_sdk.sketch_prompts import summarize_trajectories -from predicators.agent_sdk.tools import agent_render_resolution -from predicators.agent_sdk.tools.digests import render_options_digest, \ - render_types_digest from predicators.approaches import ApproachFailure from predicators.approaches.agent_session_mixin import AgentSessionMixin from predicators.approaches.base_approach import BaseApproach -from predicators.explorers import create_explorer -from predicators.explorers.base_explorer import BaseExplorer -from predicators.ground_truth_models import \ - augment_state_with_helper_objects, augment_task_with_helper_objects, \ - merge_gt_helper_predicates, merge_gt_helper_types from predicators.option_model import _OptionModelBase, create_option_model from predicators.settings import CFG -from predicators.structs import Action, Dataset, GroundAtom, \ - InteractionRequest, InteractionResult, LowLevelTrajectory, Object, \ - ParameterizedOption, ParameterizedSampler, Predicate, State, Task, Type +from predicators.structs import Action, Dataset, LowLevelTrajectory, \ + ParameterizedOption, Predicate, State, Task, Type class AgentModelFreeApproach(AgentSessionMixin, BaseApproach): - """Fixed-vocabulary open-loop planning via Claude Agent SDK. - - - Collects trajectories online using AgentModelFreeExplorer - - At solve time, queries the agent for an option plan - - No predicate/process/type invention - """ + """The agent session, tool context, trajectories and checkpoints the agent + arms share.""" def __init__(self, initial_predicates: Set[Predicate], @@ -70,16 +39,6 @@ def __init__(self, **kwargs: Any) -> None: super().__init__(initial_predicates, initial_options, types, action_space, train_tasks, *args, **kwargs) - # Optionally hand the agent the ground-truth helper scaffolding (e.g. - # the domino/fan grid loc/side types and grid predicates) so an - # "agent-with-grid" ablation plans over the oracle's vocabulary. Opt-in - # via CFG.use_gt_helpers; no-op for envs without a helper factory or - # when off. Merge here (before the agent session inits below) so the - # session, solve-time abstraction, and _get_all_predicates see them. - if self._use_gt_helpers(): - self._types = merge_gt_helper_types(self._types, CFG.env) - self._initial_predicates = merge_gt_helper_predicates( - self._initial_predicates, CFG.env) self._offline_dataset = Dataset([]) self._online_trajectories: List[LowLevelTrajectory] = [] self._option_model: Optional[_OptionModelBase] = ( @@ -87,65 +46,24 @@ def __init__(self, self._create_planner_option_model()) # Terminate Wait on atom change using the approach's predicates (which # may include invented ones), looked up lazily so the lambda picks up - # predicates invented after __init__. When the grid ablation is on, - # re-derive helper objects first so grid predicates (e.g. BallAtLoc) - # stay evaluable on the otherwise helper-free execution states. + # predicates invented after __init__. if self._option_model is not None and \ CFG.wait_option_terminate_on_atom_change: cast( # pylint: disable=protected-access Any, self._option_model)._abstract_function = ( - lambda s: utils.abstract(self._maybe_augment_state(s), - self._get_all_predicates())) - self._online_learning_cycle = 0 - # Synthesized per-skill samplers (option name -> sampler). Empty for - # the base planner; learning subclasses populate it. Threaded into - # bilevel refinement via _get_all_samplers() so continuous-parameter - # search aims at each step's subgoal instead of drawing uniformly. - self._synthesized_samplers: Dict[str, ParameterizedSampler] = {} - self._requests_train_task_idxs: Optional[List[int]] = None + lambda s: utils.abstract(s, self._get_all_predicates())) self._run_id = datetime.datetime.now().strftime("%Y%m%d_%H%M%S") - self._pre_test_conversation_log: Optional[List[Dict[str, Any]]] = None - # True only between begin_test_phase / end_test_phase, so per-episode - # hooks can act on test solves without touching exploration episodes. - self._in_test_phase = False - # 0-based index of the test task being solved, mirroring main.py's - # ``test_task_idx``. Incremented per test solve; threaded into the - # session-log filename via the ToolContext. - self._test_task_idx = -1 - # Solve-journal snapshot taken at begin_test_phase (None = no - # journal file existed) and whether it was captured successfully. - # end_test_phase archives the full test-phase journal outside the - # sandbox and rolls the file back to this snapshot, so learning - # entries persist across cycles while one evaluation's test-task - # entries never leak into the next evaluation. - self._pre_test_journal: Optional[str] = None - # Sandbox commit taken at begin_test_phase; end_test_phase - # archives what the phase added and resets the tree to it. - self._pre_test_sandbox_rev: Optional[str] = None - self._pre_test_attempts: Optional[str] = None - self._pre_test_journal_valid = False - # Scene renders attempted this episode. The first is the true initial - # state; later ones come from mid-episode replans and get distinct - # filenames so they don't overwrite the init snapshot. Reset in - # reset_for_new_episode. - self._episode_scene_renders = 0 - # Filename of the most recently saved scene render, consumed by - # _initial_image_section so the prompt references the image matching - # the state the agent is actually planning from. - self._last_scene_image_name: Optional[str] = None - # Initializes _tool_context and _agent_session_id (see mixin). Use the - # (possibly helper-augmented) vocabulary so the agent session exposes - # the grid types/predicates when CFG.use_gt_helpers is on. + # Initializes _tool_context and _agent_session_id (see mixin). self._init_agent_session_state(self._types, self._initial_predicates, initial_options, train_tasks) # Capture the underlying env once at construction. The initial option # model wraps ``env.simulate`` (a bound method), so ``__self__`` is the - # env. Later cycles may rebuild ``_option_model`` with a plain learned - # simulator that has no ``__self__``; pinning the env reference here - # keeps scene rendering (the probe's sim.render) working in every - # synthesis/solve cycle. + # env. A later model may rebuild ``_option_model`` with a plain + # learned simulator that has no ``__self__``; pinning the env + # reference here keeps scene rendering (the probe's sim.render) + # working in every round. env_self = getattr(getattr(self._option_model, '_simulator', None), '__self__', None) if env_self is not None: @@ -170,29 +88,6 @@ def _get_log_dir(self) -> str: # Overridable helpers (for subclass customisation) # ------------------------------------------------------------------ # - def _use_gt_helpers(self) -> bool: - """Whether to hand the agent the ground-truth helper scaffolding. - - Opt-in via ``CFG.use_gt_helpers`` (the process-planning - approaches read it too). When on, the grid helper - types/predicates are merged into the agent's vocabulary and the - solved task is augmented with the grid objects + oracle goal - (see ``__init__`` / ``_solve``). - """ - return CFG.use_gt_helpers - - def _maybe_augment_state(self, state: State) -> State: - """Re-derive GT helper objects on a state when the ablation is on. - - Executed states are helper-free (the grid is injected only into - the planning task), so this keeps helper predicates evaluable - during execution and Wait-on-atom-change termination. No-op when - helpers are disabled or the env has no helper factory. - """ - if self._use_gt_helpers(): - return augment_state_with_helper_objects(state, CFG.env) - return state - def _get_all_options(self) -> Set[ParameterizedOption]: """Return the full set of options available for planning.""" return self._initial_options @@ -201,28 +96,17 @@ def _get_all_predicates(self) -> Set[Predicate]: """Return the full set of predicates for abstraction.""" return self._initial_predicates - def _get_all_samplers(self) -> Dict[str, ParameterizedSampler]: - """Return synthesized per-skill samplers (option name -> sampler). - - Empty by default; learning subclasses populate the backing - field. Threaded into bilevel refinement so parameter search aims - at each step's subgoal. - """ - return self._synthesized_samplers - def _get_all_trajectories(self) -> List[LowLevelTrajectory]: """Return all trajectories (offline + online).""" return self._offline_dataset.trajectories + self._online_trajectories def _create_planner_option_model(self) -> Optional[_OptionModelBase]: - """Build the option model the planner tests plans against. + """Build the option model the tools roll plans out through. Honors two CFG knobs: * ``agent_planner_use_simulator`` -- when False, returns ``None`` - so the agent gets no ``submit_plan`` rollouts and must - plan open-loop from data + LLM reasoning (the model-free - baseline). + so the agent has no simulator to roll plans out in. * ``agent_planner_use_base_simulator`` -- when True (and a simulator is used), wraps the *base* env (``skip_residual_dynamics=True``), denying the planner the delayed @@ -234,875 +118,24 @@ def _create_planner_option_model(self) -> Optional[_OptionModelBase]: CFG.option_model_name, skip_residual_dynamics=CFG.agent_planner_use_base_simulator) - # ------------------------------------------------------------------ # - # AgentSessionMixin hooks - # ------------------------------------------------------------------ # - - # -- Prompt building blocks ----------------------------------------- # - - _SYSTEM_PROMPT_BASE = ( - "You are a planning agent. You observe task environments through " - "inspection tools and generate option plans to achieve goals. " - "You have access to read-only tools to inspect predicates, " - "options, trajectories, and training tasks. Use these to " - "understand the environment and generate effective plans.\n\n" - "Some effects may not be immediate - if an action triggers a " - "delayed process (e.g. water filling, dominoes cascading, " - "heating), insert a Wait after it so the effect has time " - "to occur before the next action. The Wait action holds the " - "robot's current pose. You can annotate Wait with target atoms " - "using `-> {atoms}` to specify exactly when it should terminate " - "(e.g. `Wait(robot:Robot)[] -> {Boiled(water:water_type)}`). " - "Use `NOT Pred(...)` for atoms that should become false. " - "If no annotation is provided, the Wait terminates on any atom " - "change. Without a Wait, the robot will proceed to the next " - "action before the delayed effect has occurred, which might " - "cause the plan to fail.") - - _SCRATCHPAD_SECTION = """ -## Scratchpad - CRITICAL -You MUST maintain `./notes.md` as your working memory. \ -**Read it at the very start of the session** and **read it \ -again before every submit_plan call** to remind yourself \ -what you already tried. **Update it immediately after every \ -submit_plan call** - no exceptions. - -Use this exact format for each option you are tuning: - -``` -## - Parameter Search -| # | params | outcome | notes | -|---|--------|---------|-------| -| 1 | [x, y, ...] | IK fail | ... | -| 2 | [x, y, ...] | success, JugNotAt... | ... | -``` - -After every test, append a row and update these summary fields: -- **Confirmed working params**: (list any that achieve the desired atoms) -- **Explored ranges**: e.g. "x: 0.9–1.05, y: 1.4–1.55" - look for GAPS -- **Unreachable region**: e.g. "y > 1.47 always IK-fails" -- **Next hypothesis**: what to try and why - -The cycle is: Read notes → plan next experiment → run test → \ -update notes → repeat. Without this loop you WILL forget what \ -you tried and repeat the same failed parameters. Treat notes.md \ -as your lab notebook - write after every single experiment. - -**If you notice you have NOT updated notes after a test, STOP \ -and update before doing anything else.**""" - - # -- System prompt --------------------------------------------------- # - - def _get_agent_system_prompt(self) -> str: - use_scratchpad = CFG.agent_planner_use_scratchpad - - sections = [self._SYSTEM_PROMPT_BASE] - - # Scratchpad - if use_scratchpad: - sections.append(self._SCRATCHPAD_SECTION) - - # Tuning workflow (numbered steps, dynamic) - steps = [] - if use_scratchpad: - steps.append( - "**Read `./notes.md` before every test**, then **update it " - "immediately after every submit_plan call**. Record " - "what you tried, what happened, and what you learned. " - "This is your memory - without it you will repeat failures.") - steps += [ - "**Review past session logs** in `./session_logs/` if available. " - "Previous queries and tool results from earlier sessions are " - "saved there. Read them to build on prior knowledge.", - "**Inspect rendered images** from `./test_images/` when " - "something goes wrong to understand the actual outcome.", - "**Expect geometric offsets.** The target position for " - "options is often offset from the reference object's reported " - "position due to object geometry. Explore a wide range around " - "the object's coordinates, not just values close to the " - "reported position.", - "**Search coarse-to-fine.** For each continuous parameter, " - "start with a WIDE grid spanning most of the valid range " - "(e.g. test 4–5 spread-out values across [low, high]). " - "Identify which coarse region works, THEN refine within it. " - "Never spend more than 3 attempts tweaking values in a small " - "neighborhood - if none work, jump to a different region. " - "Check your notes for gaps in the explored range.", - "**Vary ALL params, not just position.** Orientation and " - "other parameters change offsets and feasibility. If an " - "action fails at a position, try different values for the " - "other parameters before giving up on that region. Test at " - "least 2-3 values for each non-position parameter.", - ] - numbered = "\n".join(f"{i}. {s}" for i, s in enumerate(steps, 1)) - sections.append( - f"\n## Continuous Parameter Tuning\nFollow this workflow:\n" - f"{numbered}") - - return "\n".join(sections) - def _get_sandbox_reference_files(self) -> Dict[str, str]: """Document public control semantics without exporting implementation.""" return {"skills.md": "predicators/agent_sdk/prompts/public_skills.md"} - def _get_solve_tool_names(self) -> Optional[List[str]]: - # Type / option digests are static per session, so the solve - # prompt injects them directly (see _build_solve_prompt); the - # trajectory and task digests live in run_python's namespace - # (`trajectories` / `describe_trajectory` / `sim.task()`). - # Every remaining tool needs a simulator: submit_plan - # rolls fully-specified plans out through the option model and - # run_python probes it, so a planner without a simulator - # gets neither. - tools = [] - if CFG.agent_planner_use_simulator: - tools.append("submit_plan") - # Closed-loop policy mode: the delivery gate for the - # agent-written policy.py (submit_plan stays as a - # probe but no longer captures). - if CFG.agent_solve_policy_mode: - tools.append("submit_policy") - tools.append("run_python") - return tools - - # ------------------------------------------------------------------ # - # Learning - # ------------------------------------------------------------------ # - - def learn_from_offline_dataset(self, dataset: Dataset) -> None: - self._offline_dataset = dataset - self._tool_context.offline_trajectories = dataset.trajectories - if dataset.trajectories: - self._tool_context.example_state = \ - dataset.trajectories[0].states[0] - # Post-offline checkpoint: main.py's --load_approach path loads - # cycle None before the online loop, which previously had no - # file to read for this approach family (save only ran at the - # end of each online cycle). Hook so subclasses that learn more - # afterwards checkpoint once, after their own learning. - self._checkpoint_after_offline_learning() - - def get_interaction_requests(self) -> List[InteractionRequest]: - # Explore sessions carry their own phase tag (see the mixin's - # ``_explore_phase``) so their system prompt logs separately - # from the solve and synthesis ones. - self._explore_phase = True - try: - explorer = self._create_explorer() - requests: List[InteractionRequest] = [] - self._requests_train_task_idxs = [] - # A cycle's requests are all generated before any executes, so - # the explorer shows each query the plans already scheduled this - # cycle and asks for a complementary one. Fresh list per cycle. - self._tool_context.cycle_scheduled_plans = [] - for _ in range(CFG.online_nsrt_learning_requests_per_cycle): - task_idx = self._rng.choice(len(self._train_tasks)) - # Clear so a planning explorer's verdict is read fresh per - # request; non-planning explorers leave it None (no verdict). - self._tool_context.last_mental_model_solved = None - policy, termination_function = \ - explorer.get_exploration_strategy(task_idx, CFG.timeout) - req = InteractionRequest( - train_task_idx=task_idx, - act_policy=policy, - query_policy=lambda s: None, - termination_function=termination_function, - mental_model_solved=self._tool_context. - last_mental_model_solved) - requests.append(req) - self._requests_train_task_idxs.append(task_idx) - return requests - finally: - self._explore_phase = False - - def restore_interaction_requests(self, train_task_idxs: List[int]) -> None: - # A resume that reuses the cycle's persisted episodes never calls - # get_interaction_requests, which is what pairs each result with - # its train task below (run_20260828_173451 asserted here after a - # preemption mid-learn). - self._requests_train_task_idxs = list(train_task_idxs) - - def learn_from_interaction_results( - self, results: Sequence[InteractionResult]) -> None: - assert self._requests_train_task_idxs is not None - # Subclasses (e.g. AgentSimLearningApproach) may track the snapshot - # tags of the simulator/predicates files in effect when the explorer - # generated these plans. Tag each new trajectory so the next - # learn-phase prompt can surface provenance. ``None`` for approaches - # that don't track versions. - sim_version: Optional[str] = getattr(self, - "_current_simulator_version", - None) - preds_version: Optional[str] = getattr(self, - "_current_predicates_version", - None) - samplers_version: Optional[str] = getattr(self, - "_current_samplers_version", - None) - for i, result in enumerate(results): - task_idx = self._requests_train_task_idxs[i] - traj = LowLevelTrajectory( - result.states, - result.actions, - _train_task_idx=task_idx, - _source_simulator_version=sim_version, - _source_predicates_version=preds_version, - _source_samplers_version=samplers_version, - _env_reward=result.episode_reward, - _env_terminated=result.episode_terminated, - ) - self._online_trajectories.append(traj) - - # Update tool context - self._sync_tool_context() - - logging.info( - "[Run %s] Cycle %s: collected %d trajectories, %d total online.", - self._run_id, self._online_learning_cycle, len(results), - len(self._online_trajectories)) - - # Hook (default: save now) so subclasses that learn more after - # this method can checkpoint ONCE, after their learning, instead - # of writing a pre-learn file under the same cycle name that a - # preemption between the two writes would leave looking complete. - self._checkpoint_after_interaction_results(self._online_learning_cycle) - self._online_learning_cycle += 1 - - # ------------------------------------------------------------------ # - # Solving - # ------------------------------------------------------------------ # - - @staticmethod - def _wrap_option_failures( - policy: Callable[[State], Action]) -> Callable[[State], Action]: - """Wrap a policy so OptionExecutionFailure surfaces as ApproachFailure. - - Bilevel planning and the base open-loop planner both build a - low-level policy from a grounded option plan; this adapter is - their single place to translate the harness's option-execution - exception into the ApproachFailure CogMan expects. - """ - - def _policy(s: State) -> Action: - try: - return policy(s) - except utils.OptionExecutionFailure as e: - raise ApproachFailure(e.args[0], e.info) - - return _policy - def _solve(self, task: Task, timeout: int) -> Callable[[State], Action]: - self._sync_tool_context() - # When enabled, plan over the oracle's grid-augmented task: inject the - # grid loc/side objects and rewrite the goal to the grid BallAtLoc so - # the agent sees the oracle's scaffolding. Augmentation preserves - # goal_nl. No-op otherwise. - if self._use_gt_helpers(): - task = augment_task_with_helper_objects(task, CFG.env) - self._tool_context.current_task = task - # Render the initial state so the agent can see the scene layout. - self._render_initial_state_image(task) - try: - option_plan = self._query_agent_for_option_plan(task) - except AgentSessionFatalError: - # An ApproachFailure would be absorbed per-task; the broken - # session backend must terminate the run instead. - raise - except Exception as e: - raise ApproachFailure(f"Agent failed to produce option plan: {e}") - - preds = self._get_all_predicates() - policy = utils.option_plan_to_policy( - option_plan, - max_option_steps=CFG.max_num_steps_option_rollout, - abstract_function=lambda s: utils.abstract( - self._maybe_augment_state(s), preds)) - - return self._wrap_option_failures(policy) - - def _render_initial_state_image(self, task: Task) -> Optional[str]: - """Render the state this solve starts from and save to the sandbox. - - The first render of an episode is the true initial state - (``task{N:03d}_initial_state.png``); later renders come from - mid-episode replans and are saved as - ``task{N:03d}_replan{K}_state.png`` so they don't overwrite the - init snapshot (the replan "task" is rooted at the current, - partially-executed state). - - Returns the saved image path, or None if rendering is unavailable. - """ - self._last_scene_image_name = None - env = self._tool_context.env - if env is None: - return None - try: - # The session/sandbox (and thus ``image_save_dir`` on the - # ToolContext) is created lazily on the first agent query. This - # render runs *before* that query in ``_solve``, so on the very - # first test task the dir would still be None and task0's image - # would be silently skipped; ensure the session (and dir) exist - # first. Inside the try so a session-creation hiccup leaves - # rendering best-effort rather than crashing the solve. - self._ensure_agent_session() - except Exception as e: # pylint: disable=broad-except - logging.warning("Failed to render initial state image: %s", e) - return None - save_dir = self._tool_context.image_save_dir - if save_dir is None: - return None - task_id = self._tool_context.test_task_idx - replan_idx = self._episode_scene_renders - # Count attempts, not successes: if the init render fails, a later - # replan render still must not masquerade as the init image. - self._episode_scene_renders += 1 - if task_id is not None: - stem = f"task{task_id:03d}" - else: - stem = "" - if replan_idx == 0: - filename = f"{stem}_initial_state.png" if stem \ - else "initial_state.png" - else: - filename = f"{stem}_replan{replan_idx}_state.png" if stem \ - else f"replan{replan_idx}_state.png" - with agent_render_resolution(): - saved_path = save_task_state_image(env, task, save_dir, filename) - if saved_path is not None: - self._last_scene_image_name = filename - return saved_path - - def _initial_image_section(self) -> str: - """Return a prompt section pointing at the current solve's rendered - scene image, or an empty string if none was rendered. - - ``_render_initial_state_image`` must have been called first; - this references whichever file that call saved (init or replan - snapshot), so replan queries point at the current scene rather - than the stale episode-init image. - """ - save_dir = self._tool_context.image_save_dir - img_name = self._last_scene_image_name - if not save_dir or img_name is None: - return "" - if not os.path.exists(os.path.join(save_dir, img_name)): - return "" - # cwd of the agent is the sandbox root, so reference test_images/. - return ("\n## Initial State Image\n" - "A rendering of the scene this plan starts from has been " - f"saved to `./test_images/{img_name}`. **Read this image " - "first** to understand the spatial layout before " - "planning.\n") - - # ------------------------------------------------------------------ # - # Test phase lifecycle - # ------------------------------------------------------------------ # - - def begin_test_phase(self) -> None: - """Snapshot the learning conversation log and solve journal.""" - self._in_test_phase = True - self._test_task_idx = -1 - if self._agent_session is not None: - self._pre_test_conversation_log = copy.deepcopy( - self._agent_session.conversation_log) - else: - self._pre_test_conversation_log = None - self._snapshot_journal_for_test_phase() - self._snapshot_sandbox_for_test_phase() - - def end_test_phase(self) -> None: - """Restore the conversation log and journal to pre-test state.""" - self._in_test_phase = False - self._tool_context.test_task_idx = None - if self._agent_session is not None \ - and self._pre_test_conversation_log is not None: - # In-place restore through the public property (it returns - # the live list), so any other holder of the reference sees - # the rollback too. - log = self._agent_session.conversation_log - log[:] = self._pre_test_conversation_log - self._pre_test_conversation_log = None - self._archive_and_rollback_test_journal() - self._archive_and_rollback_test_sandbox() - - def _eval_phase_label(self) -> str: - """Name of the evaluation phase now running, for archives. - - The 0-based cycle whose learning it evaluates (matching - main.py's "ONLINE LEARNING CYCLE i"): the counter has already - advanced past that cycle's learn, so subtract 1; the pre- - learning initial test is "initial". - """ - eval_cycle = self._online_learning_cycle - 1 - return "initial" if eval_cycle < 0 else f"cycle{eval_cycle}" - - def _snapshot_sandbox_for_test_phase(self) -> None: - """Commit the sandbox tree so end_test_phase can restore it.""" - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk import sandbox_setup - self._pre_test_sandbox_rev = None - try: - self._pre_test_sandbox_rev = sandbox_setup.snapshot_sandbox( - self._tool_context.sandbox_dir) - except (OSError, subprocess.SubprocessError) as e: - logging.warning( - "[%s] Failed to snapshot the sandbox at test-phase start; " - "its test-phase files will NOT be rolled back: %s", - self._run_id, e) - - def _archive_and_rollback_test_sandbox(self) -> None: - """Archive everything the test phase wrote into the sandbox, then - restore the pre-test tree. - - The journal rollback above covers two files; this covers the - rest (test-phase session logs, scene images, notes, scripts and - plan files the agent wrote), so no evaluation leaves anything a - later learn/explore session or evaluation can read. The archive - lives in the run's log dir, outside the sandbox. - """ - rev = self._pre_test_sandbox_rev - self._pre_test_sandbox_rev = None - sandbox_dir = self._tool_context.sandbox_dir - if rev is None or not sandbox_dir: - return - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk import sandbox_setup - archive_dir = os.path.join(self._get_log_dir(), - f"sandbox_eval_{self._eval_phase_label()}") - try: - archived = sandbox_setup.rollback_sandbox(sandbox_dir, rev, - archive_dir) - except (OSError, subprocess.SubprocessError) as e: - logging.warning( - "[%s] Failed to roll back the test-phase sandbox: %s", - self._run_id, e) - return - logging.info( - "[%s] Rolled the sandbox back to its pre-test snapshot; %d " - "test-phase file(s) archived to %s", self._run_id, len(archived), - archive_dir) - - def _journal_active(self) -> bool: - """Whether solve-journal entries can be written at all.""" - return bool(CFG.agent_solve_use_journal - and self._tool_context.sandbox_dir) - - def _snapshot_journal_for_test_phase(self) -> None: - """Capture the learning-only journal and attempt log at test start. - - The snapshots are what ``end_test_phase`` rolls both files back - to. A failed capture leaves ``_pre_test_journal_valid`` False so - the rollback is skipped rather than destroying learning entries. - """ - self._pre_test_journal = None - self._pre_test_attempts = None - self._pre_test_journal_valid = False - if not self._journal_active(): - return - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk import journal as journal_mod - sandbox_dir = self._tool_context.sandbox_dir - try: - self._pre_test_journal = journal_mod.read_raw(sandbox_dir) - self._pre_test_attempts = journal_mod.read_raw( - sandbox_dir, filename=journal_mod.ATTEMPTS_FILENAME) - self._pre_test_journal_valid = True - except OSError as e: - logging.warning( - "[%s] Failed to snapshot the solve journal at test-phase " - "start; test-phase entries will NOT be rolled back: %s", - self._run_id, e) - - def _archive_and_rollback_test_journal(self) -> None: - """Archive the test-phase journal and attempt log, then roll back. - - Each evaluation must be independent of previous evaluations: - content written while solving test tasks (harness attempt-log - entries and the agent's own journal notes) would otherwise leak - this evaluation's test tasks into the next one. Learning content - - the pre-test snapshots - persists across cycles. Before the - rollback, both files (learning + this evaluation's additions) - are copied to the run's log dir, which lives outside the sandbox - so the agent cannot read them, for later inspection. - """ - if not self._pre_test_journal_valid: - return - snapshots = { - "journal": self._pre_test_journal, - "attempts": self._pre_test_attempts, - } - self._pre_test_journal = None - self._pre_test_attempts = None - self._pre_test_journal_valid = False - sandbox_dir = self._tool_context.sandbox_dir - if not self._journal_active(): - return - assert sandbox_dir is not None - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk import journal as journal_mod - filenames = { - "journal": journal_mod.JOURNAL_FILENAME, - "attempts": journal_mod.ATTEMPTS_FILENAME, - } - # One archive per evaluation phase, named by the 0-based cycle - # whose learning it evaluates (matching main.py's "ONLINE - # LEARNING CYCLE i"). The counter has already advanced past that - # cycle's learn, so subtract 1; the pre-learning initial test - # archives as "initial". A same-cycle re-eval overwrites its own - # file. - label = self._eval_phase_label() - try: - for kind, filename in filenames.items(): - content = journal_mod.read_raw(sandbox_dir, filename=filename) - if content is not None: - archive_path = os.path.join(self._get_log_dir(), - f"{kind}_eval_{label}.md") - with open(archive_path, "w", encoding="utf-8") as f: - f.write(content) - logging.info("[%s] Archived the test-phase %s to %s", - self._run_id, filename, archive_path) - journal_mod.restore(sandbox_dir, - snapshots[kind], - filename=filename) - except OSError as e: - logging.warning( - "[%s] Failed to archive/roll back the test-phase solve " - "journal: %s", self._run_id, e) - - def reset_for_new_episode(self) -> None: - """Advance the test-task counter at each test episode start. - - CogMan calls this exactly once per test task (via - ``cogman.reset`` in main.py's ``_solve_task``) and never on a - replan inside an episode, so the counter stays in lockstep with - main.py's ``test_task_idx``. The index reaches the sandbox via - the ToolContext and lands in the session-log filename. No-op - outside the test phase. - """ - super().reset_for_new_episode() - # New episode -> the next scene render is a true init snapshot. - self._episode_scene_renders = 0 - if self._in_test_phase: - self._test_task_idx += 1 - self._tool_context.test_task_idx = self._test_task_idx - - def _query_agent_for_option_plan(self, task: Task) -> list: - """Query the agent for an option plan and parse it.""" - prompt = self._build_solve_prompt(task) - responses = self._query_agent_sync(prompt, kind="test") - dead = query_fatal_error(responses) - if dead is not None: - # An outage is not a failed attempt: recording 0/1 here - # would write a bogus eval datapoint. Stop the run; the - # relaunch re-runs this cycle's test. - raise AgentSessionFatalError( - "test query died without the agent doing any work " - f"({dead}); not recording this attempt as a failure.") - plan_text = self._extract_option_plan_text(responses) - - if not plan_text: - # Log the raw responses for debugging. - n_responses = len(responses) - types = [r.get("type") for r in responses] - raise ApproachFailure( - f"Agent returned empty plan text. " - f"Got {n_responses} responses with types: {types}") - - return self._parse_and_ground_plan(plan_text, task) - - def _solve_prompt_visualize_line(self) -> str: - """The stuck-step visualization bullet: the probe's staging + render is - the only visualization surface, so the bullet appears only when - run_python is offered.""" - if CFG.agent_planner_use_simulator: - return ( - "- **Use run_python when stuck** - after 3+ failures on " - "the same step, STOP testing and use run_python " - "(`sim.reset(mods={...})`, then `sim.render(...)`) to move " - "the object to several candidate positions and " - "orientations. It's free (no physics). Find the right " - "region visually, then test.\n") - return "" - - def _solve_prompt_scratchpad_line(self) -> str: - """Return the notes.md bullet for the solve prompt, or empty.""" - if CFG.agent_planner_use_scratchpad: - return ( - "- **Read `./notes.md` before every " - "submit_plan call** " - "and **update it immediately after each call** - append a " - "row to the parameter table and update the explored-ranges " - "summary. If you realize you forgot to update, STOP and " - "update before doing anything else.\n") - return "" - - def _build_solve_prompt(self, task: Task) -> str: - """Build the prompt for generating an option plan.""" - init_state = task.init - objects = list(init_state) - - # Objects - obj_strs = [] - for obj in sorted(objects, key=lambda o: o.name): - obj_strs.append(f" {obj.name}: {obj.type.name}") - - # Goal. Only expose goal atoms whose predicate is in the agent's - # current predicate set (same filter as the bilevel sketch - # prompt): approaches that strip env predicates rely on goal_nl - # to communicate the goal. - visible_preds = self._get_all_predicates() - goal_strs = [ - str(a) for a in sorted(task.goal, key=str) - if a.predicate in visible_preds - ] - - # Types and options: static per-session digests, injected here - # instead of costing a tool turn (see _get_solve_tool_names). - types_digest = render_types_digest(self._tool_context.types) - options_digest = render_options_digest( - self._get_all_options(), - gt_options_ref_path=self._tool_context.gt_options_ref_path) - - # Current atoms - atoms = utils.abstract(init_state, self._get_all_predicates()) - atom_strs = [str(a) for a in sorted(atoms, key=str)] - - # Trajectory summary - traj_summary = self._build_trajectory_summary() - - # State features (compact) - state_str = init_state.dict_str(indent=2) - - # Available tools - tool_names = self._get_solve_tool_names() - tools_str = "" - if tool_names: - tool_list = "\n".join(f" - {t}" for t in tool_names) - tools_str = f"\n## Available Tools\n{tool_list}\n" - - # Natural language goal description (if available) - goal_nl_section = "" - if task.goal_nl: - goal_nl_section = f""" -## Goal Description -{task.goal_nl} -""" - - # Initial state image reference - initial_image_section = self._initial_image_section() - - if CFG.agent_planner_use_simulator: - instructions_intro = ( - "Use your available tools to inspect the environment and " - "test your plan before committing to it.") - else: - instructions_intro = ( - "You do NOT have a simulator to test plans against. Inspect " - "the trajectory data and reason carefully about the dynamics, " - "then commit to your best open-loop plan.") - - prompt = f"""You are solving a task. \ -Generate an option plan to achieve the goal. -{goal_nl_section} -## Goal Atoms -{chr(10).join(goal_strs)} - -## Initial State Atoms -{chr(10).join(atom_strs)} - -## Initial State Features -{state_str} -{initial_image_section}{self._solve_prompt_extra_sections()} -## Objects -{chr(10).join(obj_strs)} - -## Object Types -{types_digest} - -## Available Options -{options_digest} -{traj_summary}{tools_str} -## Instructions -{instructions_intro} - -Based on the task information and any past trajectory data, output an option plan to achieve the goal. - -After any action whose desired subgoal depends on a delayed process (e.g. water \ -filling, dominoes cascading, heating), insert a Wait action to let the process \ -complete before proceeding. You can annotate Wait with target atoms using \ -`-> {{atoms}}` to specify exactly when it should terminate. Use `NOT Pred(...)` for \ -atoms that should become false. If no annotation is provided, the Wait terminates on \ -any atom change. Only use Wait when there is a genuine delayed effect; do not insert \ -it between actions with immediate effects (e.g. Pick, Place). - -For Wait with target atoms: `Wait(robot:Robot)[] -> {{Boiled(water:water_type)}}` -For negated targets: `Wait(robot:Robot)[] -> {{NOT Touching(a:block, b:block)}}` - -**Important - parameter tuning workflow:** -- When a step fails or produces unexpected results, inspect the rendered images \ -in `./test_images/` to see what actually happened in the scene. -{self._solve_prompt_scratchpad_line()}\ -- Review past session logs in `./session_logs/` if available - they contain prior queries and results. -- When a step fails (e.g. IK error), use the image + object poses to reason about \ -WHY and adjust params directionally. Don't just try random nearby values. -{self._solve_prompt_visualize_line()}\ -- **Vary all parameters, not just position** - orientation and other params affect \ -both the outcome and whether the action succeeds. Try 2-3 values for each \ -non-position parameter per target region. -- **Search coarse-to-fine**: spread initial attempts across the full parameter range. \ -If 3 nearby values all fail the same way, jump to a very different region instead of \ -continuing to tweak. Check your notes for gaps in explored ranges. - -Output the plan with one option per line in this exact format: - OptionName(obj1:type1, obj2:type2)[param1, param2] - -If an option has no continuous parameters, use empty brackets: OptionName(obj1:type1)[] - -Output ONLY the option plan lines at the end, after any analysis.""" - - return prompt - - def _build_trajectory_summary(self) -> str: - """Summarize trajectory data for context.""" - return summarize_trajectories(self._get_all_trajectories(), - self._get_all_predicates(), - train_tasks=self._train_tasks) - - def _solve_prompt_extra_sections(self) -> str: - """Subclass hook: sections inserted into the solve prompt after the - initial-state image reference (empty by default).""" - return "" - - @staticmethod - def _extract_option_plan_text(responses: List[Dict[str, Any]]) -> str: - """Extract plan text from the last assistant text response.""" - return extract_final_text(responses) - - @staticmethod - def _strip_code_fences(text: str) -> str: - """Strip markdown code fences wrapping the plan text.""" - lines = text.split('\n') - # Remove leading/trailing ``` lines (with optional language tag). - while lines and lines[0].strip().startswith('```'): - lines.pop(0) - while lines and lines[-1].strip().startswith('```'): - lines.pop() - return '\n'.join(lines) - - def _parse_wait_annotations( - self, - text: str, - predicates: Set[Predicate], - objects: Sequence[Object], - ) -> List[Tuple[Set[GroundAtom], Set[GroundAtom]]]: - """Parse ``-> {atoms}`` annotations from plan lines. - - Returns a list parallel to the option lines in the text. Each - entry is ``(positive_atoms, negative_atoms)`` for Wait lines - with annotations, or ``(set(), set())`` otherwise. - """ - results: List[Tuple[Set[GroundAtom], Set[GroundAtom]]] = [] - option_names = {o.name for o in self._get_all_options()} - for line in text.split('\n'): - stripped = line.strip() - if not stripped: - continue - first_token = stripped.split('(')[0] - if first_token not in option_names: - if results: - break - continue - if first_token == "Wait" and '->' in stripped: - pos, neg = utils.parse_wait_target_annotations( - stripped, predicates, objects) - results.append((pos, neg)) - else: - results.append((set(), set())) - return results - - def _parse_and_ground_plan(self, plan_text: str, task: Task) -> list: - """Parse option plan text and ground into executable options.""" - objects = list(task.init) - all_options = self._get_all_options() - option_names = sorted(o.name for o in all_options) - - # Strip markdown code fences that agents often wrap plans in. - cleaned_text = self._strip_code_fences(plan_text) - - # Extract Wait target annotations before stripping them. - wait_annotations = self._parse_wait_annotations( - cleaned_text, self._get_all_predicates(), objects) - - # Strip annotations so the option plan parser doesn't choke. - parseable_text = utils.strip_wait_annotations(cleaned_text) - - parsed = utils.parse_model_output_into_option_plan( - parseable_text, - objects, - self._types, - all_options, - parse_continuous_params=True) - if not parsed: - raise ApproachFailure(f"Parsed empty option plan from agent.\n" - f" Plan text:\n{plan_text}\n" - f" Available option names: {option_names}") - - grounded = [] - for i, (option, objs, params) in enumerate(parsed): - try: - params_arr = np.array(params, dtype=np.float32) - ground_opt = option.ground(objs, params_arr) - # Inject Wait target atoms from annotations. - if (ground_opt.name == "Wait" and i < len(wait_annotations)): - pos, neg = wait_annotations[i] - if pos: - ground_opt.memory["wait_target_atoms"] = pos - if neg: - ground_opt.memory["wait_target_neg_atoms"] = neg - grounded.append(ground_opt) - except Exception as e: # pylint: disable=broad-except - logging.warning("[Run %s] Failed to ground option " - "%s: %s", self._run_id, option.name, e) - break - - if not grounded: - raise ApproachFailure("No options successfully grounded.") - logging.info("[Run %s] Agent produced plan with %d options.", - self._run_id, len(grounded)) - return grounded - - # ------------------------------------------------------------------ # - # Explorer - # ------------------------------------------------------------------ # - - def _create_explorer(self) -> BaseExplorer: - """Create explorer for interaction requests.""" - if CFG.explorer in ("agent_model_free", "agent_model_based", - "agent_plan", "agent_bilevel"): - self._sync_tool_context() - return self._create_agent_explorer( - self._get_all_predicates(), - self._get_all_options(), - name=CFG.explorer, - ) - return create_explorer( - CFG.explorer, - self._get_all_predicates(), - self._get_all_options(), - self._types, - self._action_space, - self._train_tasks, - ) + raise ApproachFailure( + f"{self.get_name()} has no task solver: the agent arms play " + "levels under the continual protocol " + "(ContinualPlayMixin.play_level).") def _sync_tool_context(self) -> None: """Push current approach state into the shared ToolContext. - The MCP tools (submit_plan, run_python, etc.) read from the + The MCP tools (run_python and the play tools) read from the ToolContext dataclass, not the approach directly. This keeps them in sync after mutations (e.g. new trajectories collected, - options added). Called before each solve and learning - interaction. Subclasses should call super() and then set + options added). Subclasses should call super() and then set additional fields (e.g. skill_factory_context). """ self._tool_context.types = self._types @@ -1120,9 +153,6 @@ def _sync_tool_context(self) -> None: self._tool_context.online_trajectories = self._online_trajectories self._tool_context.log_dir = self._get_log_dir() self._tool_context.option_model = self._option_model - # Synthesized samplers, so the explorer and synthesis tools thread the - # same per-skill samplers into refinement that the approach uses. - self._tool_context.parameterized_samplers = self._get_all_samplers() # Wire the active-experiment info-gain scorer when a learning subclass # exposes one and info-seeking exploration is on. Syncing the bound # method (not a snapshot) keeps it pointed at the latest fit/ensemble. @@ -1174,24 +204,9 @@ def _load_extra_save_state(self, save_dict: Dict[str, Any]) -> None: been refreshed, but before the tool context is re-synced. """ - def _checkpoint_after_offline_learning(self) -> None: - """Checkpoint hook at the end of offline learning (see above).""" - self.save(None) - - def _checkpoint_after_interaction_results(self, cycle: int) -> None: - """Checkpoint hook at the end of an online cycle's data collection, - BEFORE the cycle counter increments (see above).""" - self.save(cycle) - def save(self, online_learning_cycle: Optional[int] = None) -> None: - """Save approach state to disk. - - The pickled ``online_learning_cycle`` is the cycle the FILE - denotes (``_c`` means "cycle c completed"), so a subclass that - saves after the counter already advanced still writes a - consistent checkpoint; ``_None`` (post-offline) records the live - counter (0). - """ + """Save approach state to disk; the continual runner names the + checkpoint after the level it closes (``online_learning_cycle``).""" save_path = utils.get_approach_save_path_str() path = f"{save_path}_{online_learning_cycle}.{self._save_suffix}" save_dict = { @@ -1199,9 +214,6 @@ def save(self, online_learning_cycle: Optional[int] = None) -> None: self._offline_dataset, "online_trajectories": self._online_trajectories, - "online_learning_cycle": - (online_learning_cycle if online_learning_cycle is not None else - self._online_learning_cycle), "run_id": self._run_id, "agent_session_id": @@ -1220,13 +232,6 @@ def load(self, online_learning_cycle: Optional[int] = None) -> None: self._offline_dataset = save_dict["offline_dataset"] self._online_trajectories = save_dict["online_trajectories"] - # ``_c`` means cycle c completed -> resume at c+1; the - # post-offline ``_None`` file means no online cycle completed -> - # resume at its recorded counter (0), NOT +1 (which would write - # cycle 0's checkpoint as ``_1`` and break the next load). - saved_cycle = save_dict["online_learning_cycle"] - self._online_learning_cycle = (saved_cycle + 1 if online_learning_cycle - is not None else saved_cycle) # pylint: disable=attribute-defined-outside-init # (_agent_session_id is initialized via the agent-session mixin.) self._agent_session_id = save_dict.get("agent_session_id") diff --git a/predicators/approaches/agent_nl_world_model_approach.py b/predicators/approaches/agent_nl_world_model_approach.py deleted file mode 100644 index 081b32ff14..0000000000 --- a/predicators/approaches/agent_nl_world_model_approach.py +++ /dev/null @@ -1,363 +0,0 @@ -"""The natural-language world model baseline (paper arm C3). - -The same loop and experiments as the code arms - the agent explores the -train tasks, learns after every cycle, and solves the test tasks - but -the learned model is a natural-language document, ``world_model.md``, -never executable code. The learn session writes it from the recorded -data with the same data tools the code arms get; the solve and explore -sessions receive it quoted into every task message and plan by -reasoning over it, with no simulator to test plans against (the -model-free planner's open-loop solve). Predicates are the env's kept -initial ones (the same allowlist as the code arms) and goals arrive as -natural language. -""" -from __future__ import annotations - -import logging -import os -from typing import Any, Dict, List, Optional, Sequence, Set - -import numpy as np - -from predicators.agent_sdk import learn_prompts -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error -from predicators.agent_sdk.tools import _SnapshotTarget -from predicators.agent_sdk.tools.digests import render_options_digest, \ - render_trajectory_digest, render_types_digest -from predicators.agent_sdk.tools.python_exec import _make_python_exec_tool -from predicators.agent_sdk.tools.results import _make_coercing_tool, \ - _make_spilling_text_result -from predicators.agent_sdk.tools.snapshots import finalize_versioned_snapshot -from predicators.approaches.agent_model_free_approach import \ - AgentModelFreeApproach -from predicators.code_sim_learning.program_world_model import \ - option_transitions -from predicators.settings import CFG -from predicators.structs import Dataset, InteractionResult, \ - LowLevelTrajectory, Predicate - -logger = logging.getLogger(__name__) - -_NOTES_FILE = "world_model.md" -_NOTES_VERSIONS_DIR = "world_model_versions" - -_RUN_PYTHON_DESCRIPTION = ( - "Execute Python code (`code`, or `path` to a .py file you wrote in " - "the sandbox) for data exploration. Available variables: " - "trajectories (List[LowLevelTrajectory]; each has `is_demo`, " - "`train_task_idx`, `states`, `actions`; each action's `get_option()` " - "is the skill that produced it), train_tasks (List[Task]; each has " - "`init`, `goal`, `goal_holds(state)`), is_goal_state (callable: " - "state, task_idx -> bool), describe_trajectory(traj_idx, " - "include_states=True, include_atoms=False, max_timesteps=10), and " - "np. print() output is returned; the namespace persists across " - "calls; oversize output is saved under `tool_outputs/run_python/` " - "and previewed. There is no simulator in this session: the world " - "model you write is a document, checked by predicting recorded " - "transitions from it by hand.") - - -class AgentNotesWorldModelApproach(AgentModelFreeApproach): - """Model-free agentic planning over a learned natural-language world model - document.""" - - _save_suffix: str = "AgentNotesWM" - - def __init__(self, *args: Any, **kwargs: Any) -> None: - super().__init__(*args, **kwargs) - self._notes: str = "" - self._notes_version: Optional[str] = None - missing = [i for i, t in enumerate(self._train_tasks) if not t.goal_nl] - assert not missing, ( - f"{type(self).__name__} presents goals in natural language, so " - f"every train task must set `goal_nl`. Missing on task " - f"indices: {missing}") - - @classmethod - def get_name(cls) -> str: - return "agent_nl_world_model" - - # ── Vocabulary ─────────────────────────────────────────────── - - def _get_all_predicates(self) -> Set[Predicate]: - """The env predicates, restricted to the configured allowlist - (``agent_sim_learn_kept_predicates_names``) like the code arms.""" - preds = super()._get_all_predicates() - kept = CFG.agent_sim_learn_kept_predicates_names - if kept: - preds = {p for p in preds if p.name in set(kept)} - return preds - - # ── Session surface ────────────────────────────────────────── - - def _get_synthesis_tool_names(self) -> Optional[List[str]]: - return ["run_python"] - - def _get_agent_system_prompt(self) -> str: - if self._learning_mode: - return learn_prompts.build_notes_learn_system_prompt() - return "\n\n".join([ - super()._get_agent_system_prompt(), - learn_prompts.render_notes_solve_system_section(), - ]) - - def _solve_prompt_extra_sections(self) -> str: - return learn_prompts.render_world_model_notes_block( - self._notes, - self._notes_paths()["notes_file_for_agent"]) - - def _sync_tool_context(self) -> None: - super()._sync_tool_context() - self._tool_context.world_model_notes = self._notes - self._tool_context.world_model_notes_path = \ - self._notes_paths()["notes_file_for_agent"] - - # ── Learning ───────────────────────────────────────────────── - - def learn_from_offline_dataset(self, dataset: Dataset) -> None: - super().learn_from_offline_dataset(dataset) - self._learn_notes() - self.save(None) - - def learn_from_interaction_results( - self, results: Sequence[InteractionResult]) -> None: - cycle = self._online_learning_cycle - super().learn_from_interaction_results(results) - self._learn_notes() - self.save(cycle) - - def _checkpoint_after_offline_learning(self) -> None: - """No-op: this class checkpoints after its own learning.""" - - def _checkpoint_after_interaction_results(self, cycle: int) -> None: - """No-op: this class checkpoints after its own learning.""" - del cycle - - def _learn_notes(self) -> None: - """Run one document-writing session over all recorded data.""" - trajectories = self._get_all_trajectories() - if not trajectories and not CFG.agent_sim_learn_zero_shot: - logger.warning("No recorded trajectories; skipping the world " - "model document session.") - return - if not trajectories: - logger.info("Zero-shot synthesis: the agent writes the world " - "model document without data.") - self._run_notes_session(trajectories) - - def _notes_paths(self) -> Dict[str, str]: - """Host and agent-visible paths of the document (the residual arm's - sandbox mapping).""" - if CFG.agent_sdk_use_local_sandbox: - sandbox_dir: Optional[str] = os.path.abspath( - os.path.join(self._get_log_dir(), "sandbox")) - else: - sandbox_dir = self._tool_context.sandbox_dir - base = sandbox_dir or self._get_log_dir() - notes_file = os.path.join(base, _NOTES_FILE) - if CFG.agent_sdk_use_local_sandbox: - notes_file_for_agent = f"./{_NOTES_FILE}" - sandbox_dir_for_agent: Optional[str] = "." - elif sandbox_dir: - notes_file_for_agent = f"/sandbox/{_NOTES_FILE}" - sandbox_dir_for_agent = "/sandbox" - else: - notes_file_for_agent = notes_file - sandbox_dir_for_agent = None - return { - "base": base, - "notes_file": notes_file, - "versions_dir": os.path.join(base, _NOTES_VERSIONS_DIR), - "notes_file_for_agent": notes_file_for_agent, - "sandbox_dir_for_agent": sandbox_dir_for_agent or "", - } - - def _build_notes_exec_ns( - self, trajectories: List[LowLevelTrajectory]) -> Dict[str, Any]: - predicates = self._get_all_predicates() - train_tasks = self._train_tasks - - def describe_trajectory(traj_idx: int, - include_states: bool = True, - include_atoms: bool = False, - max_timesteps: int = 10) -> str: - return render_trajectory_digest(trajectories, - train_tasks, - predicates, - traj_idx, - include_states=include_states, - include_atoms=include_atoms, - max_timesteps=max_timesteps) - - return { - "trajectories": - trajectories, - "train_tasks": - train_tasks, - "is_goal_state": - lambda state, task_idx: train_tasks[task_idx].goal_holds(state), - "describe_trajectory": - describe_trajectory, - "np": - np, - } - - def _run_notes_session(self, - trajectories: List[LowLevelTrajectory]) -> None: - # pylint: disable=import-outside-toplevel - from claude_agent_sdk import tool as _sdk_tool - - from predicators.approaches.agent_sim_learning_approach import \ - AgentSimLearningApproach - - # pylint: enable=import-outside-toplevel - paths = self._notes_paths() - os.makedirs(paths["base"], exist_ok=True) - # A restored document is written back so the agent can Read it. - if self._notes and not os.path.isfile(paths["notes_file"]): - with open(paths["notes_file"], "w", encoding="utf-8") as f: - f.write(self._notes) - exec_ns = self._build_notes_exec_ns(trajectories) - run_python = _make_python_exec_tool( - _make_coercing_tool(_sdk_tool), - name="run_python", - description=_RUN_PYTHON_DESCRIPTION, - exec_ns=exec_ns, - sandbox_dir=paths["base"], - sandbox_dir_for_agent=paths["sandbox_dir_for_agent"] or None, - text_result=_make_spilling_text_result( - paths["base"], - agent_prefix=paths["sandbox_dir_for_agent"] or None), - call_timeout_s=CFG.agent_sdk_synthesis_python_call_timeout, - ) - ctx = self._tool_context - ctx.extra_mcp_tools = [run_python] - ctx.learn_cycle_index = self._online_learning_cycle - targets = [ - _SnapshotTarget( - live_file=paths["notes_file"], - versions_dir=paths["versions_dir"], - artifact_name="world_model_notes", - cycle_index_provider=lambda: self._online_learning_cycle, - ) - ] - build_hooks = ( - AgentSimLearningApproach._build_synthesis_session_hooks # pylint: disable=protected-access - ) - ctx.extra_session_hooks = build_hooks(targets, paths["base"]) - self._learning_mode = True - self._close_agent_session() - try: - self._ensure_agent_session() - message = self._build_notes_learn_message(trajectories, paths) - responses = self._query_agent_sync(message, kind="learn") - dead = query_fatal_error(responses) - if dead is not None: - raise AgentSessionFatalError( - "The learn session died without the agent doing any " - f"work ({dead}); refusing to checkpoint this cycle as " - "learned.") - finally: - ctx.extra_session_hooks = {} - ctx.extra_mcp_tools = [] - ctx.learn_cycle_index = None - self._learning_mode = False - self._close_agent_session() - self._load_notes(paths) - - def _build_notes_learn_message(self, - trajectories: List[LowLevelTrajectory], - paths: Dict[str, str]) -> str: - predicates = self._get_all_predicates() - n_trajs = len(trajectories) - n_demos = sum(1 for t in trajectories if t.is_demo) - n_transitions = sum( - len(option_transitions(t, predicates)) for t in trajectories) - session_tool_names = (self._agent_session.tool_names - if self._agent_session is not None else []) - extra_messages: List[str] = [] - if not trajectories and CFG.agent_sim_learn_zero_shot: - extra_messages.append( - learn_prompts.render_notes_zero_shot_message()) - listing = "\n".join( - f" [{i}] {'demo' if t.is_demo else 'interaction'}, task " - f"{t.train_task_idx}" for i, t in enumerate(trajectories)) - signatures = "\n".join( - f"- {p.name}({', '.join(t.name for t in p.types)})" - for p in sorted(predicates, key=lambda p: p.name)) - objective = next((t.evaluator.objective_description() - for t in self._train_tasks if t.evaluator is not None - and t.evaluator.objective_description()), "") - return learn_prompts.build_notes_learn_message( - n_trajs=n_trajs, - n_transitions=n_transitions, - n_demos=n_demos, - n_interaction=n_trajs - n_demos, - trajectory_listing=listing, - structs_ref=self._structs_reference_path(), - predicate_listing=signatures or "(none)", - types_digest=render_types_digest(self._tool_context.types), - options_digest=render_options_digest( - self._get_all_options(), - gt_options_ref_path=self._tool_context.gt_options_ref_path), - notes_file=paths["notes_file_for_agent"], - goal_nls=[t.goal_nl or "" for t in self._train_tasks], - has_prior_notes=os.path.isfile(paths["notes_file"]), - objective_block=learn_prompts.render_objective_block(objective), - tools_block=learn_prompts.render_tools_block(session_tool_names), - extra_messages=extra_messages, - ) - - def _structs_reference_path(self) -> str: - """Write the data-structures source into the sandbox reference dir (the - residual arm's convention) and return its agent-visible path.""" - # pylint: disable-next=import-outside-toplevel - import inspect - - # pylint: disable-next=import-outside-toplevel - from predicators import structs - paths = self._notes_paths() - ref_dir = os.path.join(paths["base"], "reference") - os.makedirs(ref_dir, exist_ok=True) - with open(os.path.join(ref_dir, "structs.py"), "w", - encoding="utf-8") as f: - f.write(inspect.getsource(structs)) - if paths["sandbox_dir_for_agent"]: - return f"{paths['sandbox_dir_for_agent']}/reference/structs.py" - return os.path.join(ref_dir, "structs.py") - - def _load_notes(self, paths: Dict[str, str]) -> None: - tag = finalize_versioned_snapshot( - paths["notes_file"], - paths["versions_dir"], - cycle_idx=self._online_learning_cycle, - artifact_name="world_model_notes") - if not os.path.isfile(paths["notes_file"]): - logger.warning( - "The session left no %s; the previous document " - "stands.", _NOTES_FILE) - return - with open(paths["notes_file"], "r", encoding="utf-8") as f: - self._notes = f.read() - self._notes_version = tag - self._sync_tool_context() - logger.info("Loaded the world model document (%d chars, %s).", - len(self._notes), tag or "unversioned") - - # ── Checkpointing ──────────────────────────────────────────── - - def _extra_save_state(self) -> Dict[str, Any]: - return { - "world_model_notes": self._notes, - "world_model_notes_version": self._notes_version, - } - - def _load_extra_save_state(self, save_dict: Dict[str, Any]) -> None: - self._notes = str(save_dict.get("world_model_notes") or "") - self._notes_version = save_dict.get("world_model_notes_version") - if self._notes: - paths = self._notes_paths() - os.makedirs(paths["base"], exist_ok=True) - with open(paths["notes_file"], "w", encoding="utf-8") as f: - f.write(self._notes) diff --git a/predicators/approaches/agent_program_world_model_approach.py b/predicators/approaches/agent_program_world_model_approach.py index bde3e62ea8..4c22ae77fd 100644 --- a/predicators/approaches/agent_program_world_model_approach.py +++ b/predicators/approaches/agent_program_world_model_approach.py @@ -2,37 +2,26 @@ invented predicates (paper arm C4: a code world model with no engine underneath, in the form of Pinductor / POMDP Coder). -Everything about the loop is the residual arm's - the explorer, the -sketch / refine / run tools, the capture gate, the solve pipeline, -predicate invention - except the model artifact: instead of residual -rules over a physics engine the agent writes ``world_model.py``, an -option-level transition program with its own hidden state (see +Everything about the loop is the residual arm's - the sketch / refine / +run tools and predicate invention - except the model artifact: instead of +residual rules over a physics engine the agent writes ``world_model.py``, +an option-level transition program with its own hidden state (see :mod:`code_sim_learning.program_world_model`). There is no parameter -fit; the learn session scores the program with the Pinductor +fit; the synthesis tools score the program with the Pinductor particle-filter kernel pseudo-likelihood (``sim.score``) and the agent -edits it. The belief over the hidden state is a particle set drawn from -the program's ``initial_latent``; the capture gate re-rolls every -submission under every particle through the residual arm's -rule-parameter margin channel. +edits it. """ from __future__ import annotations -import copy import logging import os from contextlib import contextmanager from typing import Any, Callable, Dict, FrozenSet, Iterator, List, Optional, \ - Tuple, cast + Tuple import numpy as np -from predicators import utils -from predicators.agent_sdk import learn_prompts -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error from predicators.agent_sdk.tools import _SnapshotTarget -from predicators.agent_sdk.tools.digests import render_options_digest, \ - render_types_digest from predicators.agent_sdk.tools.program_synthesis import CandidateLoader, \ create_program_synthesis_tools from predicators.agent_sdk.tools.snapshots import finalize_versioned_snapshot @@ -41,9 +30,9 @@ AgentSimPredicateInventionApproach from predicators.code_sim_learning.program_world_model import \ ProgramOptionModel, ProgramWorldModel, load_program_world_model, \ - option_transitions, roll_program_latents + roll_program_latents from predicators.settings import CFG -from predicators.structs import LowLevelTrajectory, State, Task +from predicators.structs import LowLevelTrajectory, State logger = logging.getLogger(__name__) @@ -63,15 +52,6 @@ def __init__(self, *args: Any, **kwargs: Any) -> None: super().__init__(*args, **kwargs) self._program: Optional[ProgramWorldModel] = None self._program_model: Optional[ProgramOptionModel] = None - # The belief particles ride the capture gate's rule-parameter - # margin channel: every submission is re-rolled from each - # particle's hidden state. - ctx = self._tool_context - ctx.rule_param_margin_provider = self._belief_particles - ctx.rule_param_override_scope = self._particle_override_scope - ctx.rule_param_margin_label = "belief particle" - ctx.rule_param_margin_note = ( - "particles of the belief over the world model's hidden state") @classmethod def get_name(cls) -> str: @@ -95,27 +75,6 @@ def _world_model_paths(paths: _SynthesisPaths) -> Dict[str, str]: # ── Learning ───────────────────────────────────────────────── - def _learn_simulator(self, trajectories: List[LowLevelTrajectory]) -> None: - """Run one program-synthesis session and deploy what it wrote.""" - self._fit_trajectories = list(trajectories) - self._persist_fit_trajectories("recorded") - usable = [ - t for t in trajectories if t.actions and t.actions[0].has_option() - ] - if not usable and not CFG.agent_sim_learn_zero_shot: - logger.warning("No skill-level transitions; skipping world " - "model synthesis.") - return - if not usable: - logger.info("Zero-shot synthesis: no skill-level transitions; " - "the agent writes the world model without data.") - program = self._run_program_synthesis_session(trajectories) - if program is None: - logger.warning("Synthesis produced no loadable world model; " - "the previous model stands.") - return - self._install_program(program) - def _install_program(self, program: ProgramWorldModel) -> None: self._program = program self._program_model = ProgramOptionModel(program, seed=CFG.seed) @@ -123,46 +82,6 @@ def _install_program(self, program: ProgramWorldModel) -> None: logger.info("Deployed the program world model (latent over %s).", dict(program.latent_features) or "nothing") - def _run_program_synthesis_session( - self, trajectories: List[LowLevelTrajectory] - ) -> Optional[ProgramWorldModel]: - paths = self._resolve_synthesis_paths() - wm_paths = self._world_model_paths(paths) - extra_paths = self._compute_extra_synthesis_paths(paths.base) - exec_ns = self._build_synthesis_exec_ns(trajectories) - self._attach_program_session_state(exec_ns, trajectories, paths, - wm_paths, extra_paths) - # Fresh session so the synthesis prompt + tools take effect. - self._close_agent_session() - self._ensure_agent_session() - structs_ref = self._write_structs_reference() - message = self._build_program_learn_message(trajectories, paths, - wm_paths, structs_ref, - extra_paths) - try: - responses = self._query_agent_sync(message, kind="learn") - dead = query_fatal_error(responses) - if dead is not None: - raise AgentSessionFatalError( - "The learn session died without the agent doing any " - f"work ({dead}); refusing to checkpoint this cycle as " - "learned.") - finally: - ctx = self._tool_context - ctx.extra_session_hooks = {} - ctx.extra_mcp_tools = [] - ctx.probe_artifact_loaders.clear() - ctx.probe_option_model_provider = None - ctx.probe_fit_provider = None - ctx.probe_validation_provider = None - ctx.probe_residuals_provider = None - ctx.probe_score_provider = None - ctx.probe_param_status = None - ctx.learn_cycle_index = None - self._learning_mode = False - self._close_agent_session() - return self._load_program_artifacts(wm_paths, extra_paths) - def _attach_program_session_state( self, exec_ns: Dict[str, Any], @@ -283,59 +202,6 @@ def _provider() -> ProgramOptionModel: return _provider - def _build_program_learn_message( - self, - trajectories: List[LowLevelTrajectory], - paths: _SynthesisPaths, - wm_paths: Dict[str, str], - structs_ref: str, - extra_paths: Dict[str, str], - ) -> str: - predicates = self._get_all_predicates() - n_trajs = len(trajectories) - n_demos = sum(1 for t in trajectories if t.is_demo) - n_transitions = sum( - len(option_transitions(t, predicates)) for t in trajectories) - prior: List[str] = [] - if os.path.isfile(wm_paths["world_model_file"]): - prior.append("`./world_model.py`") - if os.path.isfile(os.path.join(paths.base, "predicates.py")): - prior.append("`./predicates.py`") - session_tool_names = (self._agent_session.tool_names - if self._agent_session is not None else []) - extra_messages: List[str] = [] - if not trajectories and CFG.agent_sim_learn_zero_shot: - extra_messages.append( - learn_prompts.render_program_zero_shot_message()) - extra_message = self._extra_synthesis_message(extra_paths) - if extra_message: - extra_messages.append(extra_message) - return learn_prompts.build_program_learn_message( - n_trajs=n_trajs, - n_transitions=n_transitions, - n_demos=n_demos, - n_interaction=n_trajs - n_demos, - trajectory_listing=self._format_trajectory_listing(trajectories), - structs_ref=structs_ref, - predicate_listing=self._format_predicate_signatures(predicates), - types_digest=render_types_digest(self._tool_context.types), - options_digest=render_options_digest( - self._tool_context.options, - gt_options_ref_path=self._tool_context.gt_options_ref_path), - world_model_file=wm_paths["world_model_file_for_agent"], - objective_block=self._format_objective_block(), - prior_state_block=learn_prompts.render_prior_state_block(prior), - tools_block=learn_prompts.render_tools_block(session_tool_names), - extra_messages=extra_messages, - ) - - def _build_synthesis_system_prompt(self) -> str: - return learn_prompts.build_program_learn_system_prompt( - scene_viz_hint=self._scene_viz_hint(), - extra_sections=self._extra_synthesis_system_prompt_sections(), - workflow_extra=self._synthesis_workflow_extra(), - ) - def _load_program_artifacts( self, wm_paths: Dict[str, str], extra_paths: Dict[str, str]) -> Optional[ProgramWorldModel]: @@ -371,59 +237,6 @@ def _load_program_file( # ── Belief over the hidden state ───────────────────────────── - def _attach_initial_latent(self, task: Task) -> Task: - """Seed ``task.init.latent`` with the nominal particle (a seeded draw - from the program's ``initial_latent``).""" - if self._program_model is None: - return task - init_state = task.init.copy() - init_state.latent = self._program_model.initial_latent( - task.init, rng=np.random.default_rng(CFG.seed)) - return Task(init=init_state, - goal=task.goal, - alt_goal=task.alt_goal, - goal_nl=task.goal_nl) - - def _belief_particles(self) -> List[Dict[str, float]]: - """Distinct draws from ``initial_latent`` for the current task: the - capture gate's margin points (empty until a model exists).""" - model = self._program_model - if model is None: - return [] - task = self._tool_context.current_task - if task is None: - if not self._train_tasks: - return [] - task = self._train_tasks[0] - rng = np.random.default_rng(CFG.seed + 7919) - particles: List[Dict[str, Any]] = [] - seen = set() - for _ in range(CFG.agent_program_belief_particles): - try: - latent = model.initial_latent(task.init, rng=rng) - except utils.OptionExecutionFailure as e: - logger.warning("Belief particles unavailable: %s", e) - break - key = repr(sorted(latent.items(), key=lambda kv: str(kv[0]))) - if key in seen: - continue - seen.add(key) - particles.append(latent) - return cast(List[Dict[str, float]], particles) - - @contextmanager - def _particle_override_scope(self, - particle: Dict[str, float]) -> Iterator[None]: - """Roll every latent-less start from ``particle`` while entered.""" - model = self._program_model - assert model is not None - prev = model.initial_latent_override - model.initial_latent_override = copy.deepcopy(particle) - try: - yield - finally: - model.initial_latent_override = prev - def materialise_latent( self, traj: LowLevelTrajectory) -> List[Optional[Dict[str, Any]]]: if self._program is None: diff --git a/predicators/approaches/agent_session_mixin.py b/predicators/approaches/agent_session_mixin.py index 14e9e4ad1c..c83dc77fee 100644 --- a/predicators/approaches/agent_session_mixin.py +++ b/predicators/approaches/agent_session_mixin.py @@ -1,8 +1,8 @@ """Mixin providing shared agent session infrastructure. Extracts common code for ToolContext initialization, lazy -AgentSessionManager creation, async-to-sync bridging, and agent explorer -creation shared by AgentModelFreeApproach and its subclasses. +AgentSessionManager creation and async-to-sync bridging shared by +AgentModelFreeApproach and its subclasses. """ import logging import os @@ -15,8 +15,6 @@ SessionManagerProtocol, run_async_sync, run_query_sync from predicators.agent_sdk.tools import ALL_TOOL_NAMES, ToolContext, \ create_mcp_tools, get_allowed_tool_list -from predicators.explorers import create_explorer -from predicators.explorers.base_explorer import BaseExplorer from predicators.settings import CFG from predicators.structs import ParameterizedOption, Predicate, Task, Type @@ -31,7 +29,7 @@ class AgentSessionMixin: And may optionally override: - _get_solve_tool_names() -- complete tool surface for - solve / explore sessions. May mix static MCP tool names with + play sessions. May mix static MCP tool names with names of dynamic ``SdkMcpTool`` instances. ``None`` = all static MCP tools, ``[]`` = none. - _get_synthesis_tool_names() -- complete tool surface for @@ -54,11 +52,6 @@ class AgentSessionMixin: # by the sim-learning approach around synthesis sessions; the class # default keeps plain solve-only hosts working without declaring it. _learning_mode: bool = False - # Flipped around explorer creation (``get_interaction_requests``) so - # explore sessions carry their own phase tag: their system prompt is - # saved as ``system_prompt_explore.md`` next to the solve and - # synthesis ones instead of overwriting the solve copy. - _explore_phase: bool = False # Phase tag the live ``_agent_session`` was created with; a query # under a different phase closes and rebuilds the session so the # saved prompt, tools, and CLAUDE.md always match the active phase. @@ -101,7 +94,7 @@ def _get_agent_system_prompt(self) -> str: raise NotImplementedError def _get_solve_tool_names(self) -> Optional[List[str]]: - """Return the complete tool surface for solve / explore sessions. + """Return the complete tool surface for play sessions. May mix static MCP tool names with names of dynamic ``SdkMcpTool`` instances. ``None`` means "all static MCP tools"; @@ -123,7 +116,7 @@ def _get_synthesis_tool_names(self) -> Optional[List[str]]: return [] def _get_sandbox_reference_files(self) -> Dict[str, str]: - """Return extra reference files for the docker sandbox. + """Return extra reference files for the sandbox. Maps destination paths (relative to ``/sandbox/reference/``) to source paths (relative to the repo root). Override in @@ -138,17 +131,12 @@ def _get_sandbox_reference_files(self) -> Dict[str, str]: def _ensure_agent_session(self) -> None: """Create the agent session manager if needed. - When ``SessionConfig.use_docker_sandbox`` is ``True``, creates a - ``DockerSessionManager`` that runs ``ClaudeSDKClient`` inside a - Docker container with full built-in tools (Bash, Read, Write, - …). Otherwise creates the normal in-process - ``AgentSessionManager``. + When ``SessionConfig.use_local_sandbox`` is ``True``, creates a + ``LocalSandboxSessionManager`` that runs the agent in a sandbox + directory with its built-in tools (Bash, Read, Write, ...). + Otherwise creates the in-process ``AgentSessionManager``. """ - phase = ("synthesis" if self._learning_mode else - "explore" if self._explore_phase else "solve") - # The tools read the phase for phase-dependent facts such as - # the real episode's step budget (explore episodes are capped by - # max_num_steps_interaction_request, tests by the horizon). + phase = "synthesis" if self._learning_mode else "solve" self._tool_context.phase = phase if self._agent_session is not None: if self._agent_session_phase == phase: @@ -220,21 +208,7 @@ def _ensure_agent_session(self) -> None: logger.info("\n".join(lines)) session: SessionManagerProtocol - if config.use_docker_sandbox: - from predicators.agent_sdk.docker_sandbox import \ - DockerSessionManager # pylint: disable=import-outside-toplevel - session = DockerSessionManager( - system_prompt=self._get_agent_system_prompt(), - log_dir=self._get_log_dir(), - model_name=config.model_name, - tool_context=self._tool_context, - tool_names=tool_names, - image=config.docker_image, - extra_reference_files=self._get_sandbox_reference_files(), - phase=phase, - config=config, - ) - elif config.use_local_sandbox: + if config.use_local_sandbox: from predicators.agent_sdk.local_sandbox import \ LocalSandboxSessionManager # pylint: disable=import-outside-toplevel session = LocalSandboxSessionManager( @@ -279,7 +253,7 @@ def _ensure_agent_session(self) -> None: # suffix) as full_system_prompt_{phase}.md during sandbox setup; for # in-process sessions the approach prompt IS the full prompt, so # save it under the same name here and nothing else. - if not (config.use_docker_sandbox or config.use_local_sandbox): + if not config.use_local_sandbox: log_dir = self._get_log_dir() os.makedirs(log_dir, exist_ok=True) prompt_path = os.path.join(log_dir, @@ -323,22 +297,3 @@ def _query_agent_sync(self, message: str, self._ensure_agent_session() assert self._agent_session is not None return run_query_sync(self._agent_session, message, **query_kwargs) - - def _create_agent_explorer( - self, - predicates: Set[Predicate], - options: Set[ParameterizedOption], - name: str = "agent_model_free", - ) -> BaseExplorer: - """Create an agent explorer with tool_context and agent_session.""" - self._ensure_agent_session() - return create_explorer( - name, - predicates, - options, - self._types, - self._action_space, - self._train_tasks, - tool_context=self._tool_context, - agent_session=self._agent_session, - ) diff --git a/predicators/approaches/agent_sim_learning_approach.py b/predicators/approaches/agent_sim_learning_approach.py index b1f3925825..e5227a1915 100644 --- a/predicators/approaches/agent_sim_learning_approach.py +++ b/predicators/approaches/agent_sim_learning_approach.py @@ -1,26 +1,17 @@ -"""Agent sim-learning approach: learns a simulator program online. - -Extends AgentModelBasedApproach to learn residual dynamics via an -agent-synthesized step-level simulator with parameterized process -rules. Parameters are fitted by Levenberg-Marquardt (fitting.py). - -The approach creates a base oracle (PyBullet with process -dynamics disabled) and composes it with the learned step-level -dynamics into a single simulator function, plugged into a standard -_OracleOptionModel for true per-step interleaving. - -Example command:: - - python predicators/main.py --env pybullet_boil \ - --approach agent_sim_learning --seed 0 \ - --num_train_tasks 10 --num_test_tasks 5 \ - --num_online_learning_cycles 5 --explorer agent_model_free +"""The model half of the EMPIRIC arms: the agent's simulator program, its +fitted parameters and the belief over them. + +The agent writes a simulator (a subclass of the base env, or residual +rules over it); the harness loads it, fits or deploys its parameters, +keeps their uncertainty, and composes it with a base oracle (PyBullet +with process dynamics disabled) into the option model plans are +rehearsed in. The continual arms (``agent_continual_approach``) drive +all of it from their play rounds. """ import copy import dataclasses import hashlib -import inspect import logging import math import os @@ -29,24 +20,20 @@ from typing import Any, Callable, Collection, ContextManager, Dict, \ FrozenSet, Iterator, List, Optional, Sequence, Set, Tuple -import dill as pkl import numpy as np import pybullet from gym.spaces import Box from predicators import utils -from predicators.agent_sdk import learn_prompts from predicators.agent_sdk.fit_status import format_fit_status -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - max_session_log_number, query_fatal_error +from predicators.agent_sdk.play_prompts import render_physical_params_section +from predicators.agent_sdk.session_base import max_session_log_number from predicators.agent_sdk.tools import SYNTHESIS_TOOL_NAMES, \ - _SnapshotTarget, create_synthesis_tools, evaluate_states_with, \ - finalize_versioned_snapshot, make_write_snapshot_hook -from predicators.agent_sdk.tools.digests import render_options_digest, \ - render_trajectory_digest, render_types_digest -from predicators.approaches.agent_model_based_approach import \ - AgentModelBasedApproach -from predicators.approaches.sampler_learning_mixin import SamplerLearningMixin + _SnapshotTarget, evaluate_states_with, finalize_versioned_snapshot, \ + make_write_snapshot_hook +from predicators.agent_sdk.tools.digests import render_trajectory_digest +from predicators.approaches.agent_model_free_approach import \ + AgentModelFreeApproach from predicators.approaches.synthesis_validation import \ build_candidate_option_model, carry_over_params from predicators.code_sim_learning.active_experiment import laplace_ensemble, \ @@ -57,18 +44,14 @@ from predicators.code_sim_learning.fit_space import FitResult, ParamSpec, \ declared_interval_fit_result, declared_interval_report from predicators.code_sim_learning.fitting import FIT_NOISE_SIGMA, \ - compute_sse, compute_sse_recurrent, fit_rule_parameters, \ - fit_rule_parameters_latent, log_param_changes, log_sse_breakdown -from predicators.code_sim_learning.identifiability import Verdict, \ - format_identifiability, physics_sigma_points + compute_sse, compute_sse_recurrent, log_sse_breakdown +from predicators.code_sim_learning.identifiability import physics_sigma_points from predicators.code_sim_learning.latent_tracker import LatentTracker, \ make_latent_tracker, make_subclass_latent_tracker from predicators.code_sim_learning.model_state import has_model_state -from predicators.code_sim_learning.orchestrator import \ - prior_parameter_belief, run_rollout_sysid +from predicators.code_sim_learning.orchestrator import prior_parameter_belief from predicators.code_sim_learning.parameter_belief import BeliefConfig, \ ParameterBelief, stable_seed -from predicators.code_sim_learning.physical_sysid import fit_params_rollout from predicators.code_sim_learning.rollout_env import RolloutTrajectory, \ dispose_env, physical_param_anchors from predicators.code_sim_learning.rollout_objective import compute_rollout_sse @@ -80,13 +63,12 @@ observation_view, read_latent_init, read_physical_param_specs, \ read_residual_env, read_simulator_components, stamp_physical_spec_scales from predicators.envs import create_new_env -from predicators.ground_truth_models import get_gt_simulator from predicators.observation_noise import ObservationNoise from predicators.option_model import _OptionModelBase, _OracleOptionModel from predicators.settings import CFG -from predicators.structs import Action, Dataset, DerivedPredicate, \ - GroundAtom, InteractionResult, LowLevelTrajectory, ParameterizedOption, \ - Predicate, State, Task, Type, step_option_labels +from predicators.structs import Action, DerivedPredicate, GroundAtom, \ + LowLevelTrajectory, ParameterizedOption, Predicate, State, Task, Type, \ + step_option_labels logger = logging.getLogger(__name__) @@ -193,23 +175,13 @@ def residual_hint_from_hits(hits: Dict[Tuple[str, str], int], return {t: sorted(fs) for t, fs in out.items()} -class AgentSimLearningApproach(SamplerLearningMixin, AgentModelBasedApproach): - """Bilevel planning with a learned step-level simulator. - - During online learning: - 1. Collect trajectories (inherited from AgentModelBasedApproach) - 2. Segment into option-level transitions - 3. Synthesize parameterized residual rules via Claude agent - 4. Fit rule parameters via Levenberg-Marquardt - 5. Compose with base oracle into a combined simulator - 6. Build _OracleOptionModel with the combined simulator - - During solving: - - Uses the learned model for plan validation in backtracking - refinement. +class AgentSimLearningApproach(AgentModelFreeApproach): + """The agent-written simulator, its parameters and their belief. - Per-skill sampler learning (mode resolution, synthesis session - plumbing, loading) lives in :class:`SamplerLearningMixin`. + Loads the simulator the agent writes, deploys its parameters (the + agent's published ``sim.fit`` or its declared values), keeps the + belief over them, and composes it with the base oracle into the + ``_OracleOptionModel`` the probe rehearses plans in. """ # Allowlist of env predicate names surfaced to the agent; None keeps @@ -247,29 +219,18 @@ def __init__(self, *args, option_model=option_model, **kwargs) - # Capture-validation rollouts each run on a freshly constructed env - # (see ToolContext.validation_env_scope): repeats on the shared + # Probe trials and sweep rollouts each run on a freshly constructed + # env (see ToolContext.validation_env_scope): repeats on the shared # ``_base_env`` are correlated across resets, so only fresh envs # sample the distribution the real episode will. self._tool_context.validation_env_scope = \ self._fresh_validation_env_scope - # Physics-margin points for the capture gate (+-1 posterior sigma - # of the latest applied fit): a callable so the tool always sees - # the current fit, not the one deployed when the session opened. - # Under agent_sim_learn_param_uncertainty False the points are - # never built (see _physics_margin_points), so this returns []. + # The sim.run physics sweep's stress points (see _stress_points): + # a callable so the probe always sees the current fit, not the + # one deployed when the session opened. Under + # agent_sim_learn_param_uncertainty False the legacy grid is never + # built (see _physics_margin_points). self._tool_context.physics_margin_provider = self._stress_points - # Rule-parameter margin points for the capture gate: the - # calibrated ensemble the info-seeking explorer scores with - # (posterior subsample / Laplace / jitter, see - # _select_param_ensemble) doubles as the uncertainty sweep over - # LEARNED rule constants - a submission must survive every - # member, not just the fitted point estimate. Callables so the - # gate always sees the latest fit's ensemble. - self._tool_context.rule_param_margin_provider = \ - lambda: [dict(m) for m in self._param_ensemble] - self._tool_context.rule_param_override_scope = \ - self._rule_param_override_scope # The joint belief's rehearsals split a draw's parameters into the # ones the env applies and the rule parameters read through the # draw scope. @@ -342,7 +303,6 @@ def __init__(self, # (consumed in the next learn-phase prompt). self._current_simulator_version: Optional[str] = None self._current_predicates_version: Optional[str] = None - self._init_sampler_learning_state() # Partial-observability latent block: loaded from a simulator's # LATENT_INIT export (None ⇒ no latent state). When the loaded # rules use the recurrent 5-arg signature, fitting, the combined @@ -371,7 +331,7 @@ def __init__(self, # delta against the previous version reads from here. self._fit_evidence_history: Dict[str, Dict[str, float]] = {} # +-1-posterior-sigma perturbations of the applied params (the - # capture gate's physics-margin points). Set only by the joint + # legacy physics sweep's grid). Set only by the joint # rollout fit, which has the identifiability report; cleared by # every _apply_identified_physical_params call so points can # never outlive the fit they were derived from. @@ -391,33 +351,6 @@ def __init__(self, # declaration/data signature); values are # orchestrator._FitComputation bundles. self._sysid_fit_cache: Dict[Tuple, Any] = {} - # Final per-cycle fit history for the cross-cycle consistency - # check: name -> (map_value, posterior_std_fit_space, scale). - # Mutually-incompatible confident fits across cycles are the - # signature of an overconfident probe; flagged, and the verdict - # downgraded, rather than silently trusted. - self._sysid_fit_history: Dict[str, Tuple[float, float, str]] = {} - # A rejected (INCONSISTENT) fit awaiting confirmation: - # name -> (map_value, posterior_std_fit_space). If the NEXT - # cycle's independent fit lands within the consistency band of - # the pending value, the jump is accepted as real (two - # independent fits agree); until then the trusted history value - # holds. Without this, a genuinely-updated fit would read - # INCONSISTENT against stale history forever. - self._sysid_pending_fit: Dict[str, Tuple[float, float]] = {} - # The applied physical params as of the last CYCLE-LEVEL fit - - # the reference the INCONSISTENT hold policy reverts to. - # Deliberately not _identified_physical_params: the agent's - # in-session sim.fit calls mutate that dict, so "hold the - # currently-applied value" was a no-op that held the very fit - # it refused to trust (run_20260724_232411 seed2 cycle 2: - # "holding the currently-applied 0.6267" - 0.6267 WAS the - # distrusted new fit, applied minutes earlier in-session). - self._cycle_applied_physical: Dict[str, float] = {} - # Agent-facing digest of the latest rollout fit (unexplainable - # segments, unidentified/insensitive params, cross-cycle - # conflicts); surfaced to the explorer as experiment objectives. - self._last_sysid_diagnostics: str = "" @classmethod def get_name(cls) -> str: @@ -459,27 +392,6 @@ def _compute_kept_initial_predicates(self) -> Set[Predicate]: # ── Agent session hooks ────────────────────────────────────── - def _get_agent_system_prompt(self) -> str: - if self._learning_mode: - return self._build_synthesis_system_prompt() - prompt = super()._get_agent_system_prompt() - base_sim_refs = self._base_sim_reference_paths() - if base_sim_refs: - ref_listing = "\n".join(f" - {r}" for r in base_sim_refs) - prompt += ( - "\n\n## Base Simulator Source\n" - "The environment simulator's own source code is " - "available (read-only):\n" - f"{ref_listing}\n" - "It covers the observable sim core: scene geometry and " - "constants, body construction, physics stepping, and " - "state read/write. It deliberately omits the hidden " - "domain-specific dynamics, task generation, and goal " - "semantics. Read it to ground your spatial and physical " - "reasoning (dimensions, contact geometry, actuation) " - "instead of guessing from images or trial and error.\n") - return prompt - def _get_sandbox_reference_files(self) -> Dict[str, str]: files = super()._get_sandbox_reference_files() # Base-sim source rides the standard reference channel so every @@ -493,46 +405,30 @@ def _get_synthesis_tool_names(self) -> Optional[List[str]]: """Complete tool surface for the synthesis agent. The names of the dynamic synthesis callables (just - ``run_python``) attached to ``ctx.extra_mcp_tools`` inside - :meth:`_synthesize_with_agent`. The mixin asserts the attached - instances and this list agree. Fitting, residual reports, and - plan validation are NOT tools: they live on the ``sim`` probe - (``sim.fit`` / ``sim.residuals`` / ``sim.refine`` / - ``sim.run``) inside ``run_python``. - - No inspect tools: the type/option digests are injected into the - learn message (see :meth:`_build_synthesis_learn_message`) and - trajectory access lives in ``run_python`` (``trajectories`` + - ``describe_trajectory``). The probe rides inside that same - ``run_python`` namespace as ``sim`` (one exec namespace per - session - a helper defined next to the data is visible to probe - sweeps; the solve-phase instance of the tool is not built when - this one is attached). In the - agent-synthesis session the probe runs against the CANDIDATE - simulator.py via ctx.probe_option_model_provider (installed in - _synthesize_with_agent); in the oracle-sim-program sampler - session no provider is installed and the probe falls back to - ctx.option_model, which there IS the deployed belief model. + ``run_python``); a continual round keeps only the toolkit tools + named here. Fitting, residual reports, and plan validation are + NOT tools: they live on the ``sim`` probe (``sim.fit`` / + ``sim.residuals`` / ``sim.refine`` / ``sim.run``) inside + ``run_python``. + + No inspect tools: trajectory access lives in ``run_python`` + (``trajectories`` + ``describe_trajectory``). The probe rides + inside that same ``run_python`` namespace as ``sim`` (one exec + namespace per session - a helper defined next to the data is + visible to probe sweeps) and runs against the CANDIDATE + simulator.py via ctx.probe_option_model_provider. """ names: List[str] = list(SYNTHESIS_TOOL_NAMES) return names # ── Subclass hooks ────────────────────────────────────────── # Default implementations are no-ops so subclasses can add - # predicate-invention (or other) extensions without copying - # _synthesize_with_agent. + # predicate-invention (or other) extensions. def _learning_cycle_index(self) -> int: - """0-based cycle index used in versioned snapshot filenames. - - Matches main.py's "ONLINE LEARNING CYCLE i" numbering exactly: - ``_online_learning_cycle`` is incremented before this class's - online simulator learn runs, so subtracting 1 recovers the - cycle the session belongs to. The offline (pre-cycle-0) learn - yields -1, which the snapshot/journal formatters render as - "offline" - keeping it distinct from cycle 0's online pass. - """ - return self._online_learning_cycle - 1 + """Index used in versioned snapshot filenames; the continual arms + number snapshots by level.""" + return 0 def _compute_extra_synthesis_paths(self, base: str) -> Dict[str, str]: """Return extra path bindings for the synthesis sandbox.""" @@ -555,34 +451,6 @@ def _install_extra_synthesis_surfaces( """ del exec_ns, base_pred_triples, inferred_hint, extra_paths - def _extra_synthesis_message(self, extra_paths: Dict[str, str]) -> str: - """Return text to append to the agent's first synthesis message. - - Under ``CFG.partially_observable`` this is the short partial- - observability note; subclasses that override MUST chain via - ``super()`` so the note survives. - """ - del extra_paths - if CFG.partially_observable: - return learn_prompts.render_partial_observability_message() - return "" - - def _extra_synthesis_system_prompt_sections(self) -> List[str]: - """Sections a subclass adds to the synthesis system prompt. - - Inserted after the validation guidance and before the recurrent - rules tutorial (partial observability) and the plan format. - Subclasses that override MUST chain via ``super()``. - """ - return [] - - def _extra_synthesis_latent_sections(self) -> List[str]: - """Sections a subclass adds after the recurrent-rules tutorial. - - Only rendered under ``CFG.partially_observable``. - """ - return [] - def _post_synthesis_loading( self, extra_paths: Dict[str, str], @@ -649,33 +517,6 @@ def _build_synthesis_session_hooks( # ── Learning ──────────────────────────────────────────────── - def learn_from_offline_dataset(self, dataset: Dataset) -> None: - super().learn_from_offline_dataset(dataset) - self._learn_simulator(self._get_all_trajectories()) - # The single post-offline checkpoint, AFTER the simulator learn - # (the base hook is a no-op for this class, see below). - self.save(None) - - def learn_from_interaction_results( - self, results: Sequence[InteractionResult]) -> None: - # Capture the index BEFORE super() increments it: the checkpoint - # below must be the one this cycle's filename denotes. - cycle = self._online_learning_cycle - super().learn_from_interaction_results(results) - self._learn_simulator(self._get_all_trajectories()) - # The single per-cycle checkpoint, AFTER this cycle's simulator - # learning, so a resume never re-pays a completed learn and never - # mistakes a pre-learn file for a completed cycle. (The base hook - # that would have saved pre-learn is a no-op for this class.) - self.save(cycle) - - def _checkpoint_after_offline_learning(self) -> None: - """No-op: this class checkpoints after its own simulator learn.""" - - def _checkpoint_after_interaction_results(self, cycle: int) -> None: - """No-op: this class checkpoints after its own simulator learn.""" - del cycle - # ── Checkpointing ──────────────────────────────────────────── # The base checkpoint (AgentModelFreeApproach.save/load) persists # the datasets + cycle counter. This approach's real state is split @@ -684,21 +525,19 @@ def _checkpoint_after_interaction_results(self, cycle: int) -> None: # which are embedded as file CONTENTS - run dirs are minted per run # and pruned, so a path reference to the old run's sandbox would be # fragile. Closures (_residual_rules, _learned_simulator, the option - # model, learned predicates/samplers) are never pickled: they are + # model, learned predicates) are never pickled: they are # rebuilt from the restored files in _rehydrate_from_artifacts. _save_suffix: str = "AgentSimLearner" _CHECKPOINT_SANDBOX_FILES: Tuple[str, ...] = ("simulator.py", "predicates.py", - "samplers.py", "ground_samplers.py", "notes.md", "journal.md", "attempts.md", "strategy.md", "open_questions.md") _CHECKPOINT_SANDBOX_DIRS: Tuple[str, ...] = ("simulator_versions", - "predicates_versions", - "samplers_versions") + "predicates_versions") _CHECKPOINT_MAX_FILE_BYTES = 2 * 1024 * 1024 def _checkpoint_sandbox_dir(self) -> str: @@ -783,16 +622,12 @@ def _extra_save_state(self) -> Dict[str, Any]: dict(self._fit_evidence_history), "identified_physical_sigma_points": list(self._identified_physical_sigma_points), - "sysid_fit_history": - dict(self._sysid_fit_history), "residual_features": dict(self._residual_features), "current_simulator_version": self._current_simulator_version, "current_predicates_version": self._current_predicates_version, - "current_samplers_version": - self._current_samplers_version, "sandbox_files": self._collect_sandbox_artifacts(), "git_describe": @@ -812,8 +647,8 @@ def _load_extra_save_state(self, save_dict: Dict[str, Any]) -> None: "is at %s - resuming across code versions is untested.", saved_rev, current_rev) self._resume_query_count = int(save_dict.get("agent_query_count", 0)) - # In-place update: _ParamsView holders (invented predicate and - # sampler closures) alias this exact dict object. + # In-place update: _ParamsView holders (invented predicate + # closures) alias this exact dict object. self._fitted_params.clear() self._fitted_params.update(save_dict.get("fitted_params") or {}) self._fit_sse = save_dict.get("fit_sse", float("inf")) @@ -831,18 +666,12 @@ def _load_extra_save_state(self, save_dict: Dict[str, Any]) -> None: save_dict.get("carried_physical_prior") or {}) self._fit_evidence_history = dict( save_dict.get("fit_evidence_history") or {}) - self._sysid_fit_history = dict( - save_dict.get("sysid_fit_history") or {}) self._residual_features = dict( save_dict.get("residual_features") or {}) self._current_simulator_version = save_dict.get( "current_simulator_version") self._current_predicates_version = save_dict.get( "current_predicates_version") - # pylint: disable-next=attribute-defined-outside-init - # (initialized by SamplerLearningMixin's init hook) - self._current_samplers_version = save_dict.get( - "current_samplers_version") self._restore_sandbox_artifacts(save_dict.get("sandbox_files") or {}) self._rehydrate_from_artifacts() # AFTER rehydration: _apply_identified_physical_params clears @@ -860,8 +689,7 @@ def _rehydrate_from_artifacts(self) -> None: Order matters: simulator.py first (rules + latent init + physical specs), then the option model, then identified physics onto the base env, then subclass artifacts (predicates read the - already- restored ``_fitted_params``), then samplers and the - ensemble. + already- restored ``_fitted_params``), then the ensemble. """ paths = self._resolve_synthesis_paths() if not os.path.isfile(paths.simulator_file): @@ -934,126 +762,12 @@ def _step_fn(s: State, c: Any) -> Any: self._apply_identified_physical_params( self._identified_physical_params) self._rehydrate_extra_artifacts(paths.base) - if self._samplers_enabled(): - sampler_paths = self._sampler_paths(paths.base) - self._synthesized_samplers = self._load_samplers_from_module_file( - sampler_paths["samplers_file"]) self._rebuild_param_ensemble() logger.info( "Rehydrated learned simulator from checkpoint artifacts " - "(%d rules, %d fitted params, %d learned predicates, " - "%d samplers).", len(rules), len(self._fitted_params), - len(getattr(self, "_learned_predicates", set()) or set()), - len(self._synthesized_samplers)) - - def _learn_simulator(self, trajectories: List[LowLevelTrajectory]) -> None: - """Synthesize rules, fit parameters, and build the option model.""" - # Cache for recurrent fitting: lets _group_triples_by_trajectory - # slice the flat base_pred_triples back into per-trajectory chunks - # (latent threads within a trajectory, not across). Harmless for - # fully-observable (legacy) simulators, which never regroup. - self._fit_trajectories = list(trajectories) - # Dumped HERE, where the data arrives, rather than only inside the - # sysID fit: a cycle where the agent declines to fit is exactly the - # one worth post-morteming, and that is the branch that never ran. - # run_20260817_171402 declined on a sweep that returned one identical - # SSE for every value of five parameters, and left nothing on disk to - # explain it -- the episode had to be written off. - self._persist_fit_trajectories("recorded") - # New data invalidates the memoized explainability verdicts and - # the memoized whole fits. - self._explainability_cache.clear() - self._sysid_fit_cache.clear() - # Decide how samplers are obtained this cycle: ground-truth (if - # requested and available for the env) else agent synthesis. GT - # samplers are static, so install them up front, independent of - # whether simulator learning runs below (it is skipped when there - # are no step transitions and no oracle sim program to fall - # back on, e.g. when every demo failed). - self._maybe_install_oracle_samplers() - # Two parallel triple lists drive the rest of this method: - # * obs_triples - raw (s_t, a, s_{t+1}) from the data. - # * base_pred_triples - same triples but s_t replaced by the - # base sim's one-step prediction. The rules run on top of that - # prediction; SSE compares against s_{t+1}. - obs_triples = self._extract_obs_triples(trajectories) - if (not obs_triples and not CFG.agent_sim_learn_oracle_sim_program - and not CFG.agent_sim_learn_zero_shot): - logger.warning("No step transitions; skipping simulator learning.") - return - if obs_triples: - # Headless env for the pre-compute: reusing the GUI base_env - # corrupts its visual-shape state after a few hundred steps. - fit_env = create_new_env(CFG.env, - do_cache=False, - use_gui=False, - skip_residual_dynamics=True) - logger.info("Pre-computing base states for %d transitions.", - len(obs_triples)) - try: - base_pred_triples = self._compute_base_pred_triples( - obs_triples, fit_env) - finally: - # This env is rebuilt every learning cycle; dispose it - # (main client AND any secondary probe world) or each - # cycle leaks a full physics world (~145MB for the - # domino env). - dispose_env(fit_env) - inferred_hint = self._infer_residual_features_from_scan( - obs_triples, base_pred_triples) - logger.info("Residual features (data-driven hint): %s", - inferred_hint) - elif CFG.agent_sim_learn_oracle_sim_program: - # The oracle sim program is data-free (rules and parameter - # inits come from get_gt_simulator), so a run whose every - # demo failed still gets a working option model; the fit - # below degrades to the declared inits. - logger.warning("No step transitions; loading oracle sim " - "program without data.") - base_pred_triples = [] - inferred_hint = {} - else: - # Zero-shot synthesis (ablation A2): the session runs with - # nothing recorded, so the artifacts come from the task - # description, the scene and the agent's own knowledge; the - # params deploy at their declared inits. - logger.info("Zero-shot synthesis: no step transitions; the " - "agent writes its artifacts without data.") - base_pred_triples = [] - inferred_hint = {} - - self._synthesize_with_agent(trajectories, obs_triples, - base_pred_triples, inferred_hint) - - if self._residual_rules is not None and self._fitted_params: - rules, params = self._residual_rules, self._fitted_params - self._learned_simulator = LearnedSimulator( - step_fn=lambda s, c, _r=rules, _p=params: # type: ignore[misc] - apply_rules(s, _r, _p, cmds=c), - name="agent_synthesized") - elif self._learned_simulator is None: - logger.warning("Synthesis produced no simulator, skipping.") - return - - combined_sim = self._build_combined_simulator(self._learned_simulator) - self._option_model = self._build_option_model(combined_sim) - logger.info("Built learned option model (SSE: %.6f).", self._fit_sse) - - # When the simulator came from the oracle short-circuit no agent - # session ran above, so per-skill samplers (if enabled) get their - # own session here, after the option model is built so the - # session's probe (sim.refine) has a working simulator. When - # the agent *did* synthesize the simulator, samplers already rode - # along in that session and this is skipped. - if self._do_synthesize_samplers and \ - CFG.agent_sim_learn_oracle_sim_program: - if base_pred_triples: - self._synthesize_samplers_standalone(trajectories, - base_pred_triples, - inferred_hint) - else: - logger.warning("No step transitions; skipping standalone " - "sampler synthesis.") + "(%d rules, %d fitted params, %d learned predicates).", len(rules), + len(self._fitted_params), + len(getattr(self, "_learned_predicates", set()) or set())) def _build_option_model( self, @@ -1070,7 +784,7 @@ def _build_option_model( # so that env is the one physics-needing task-evaluator # certificates (the domino counterfactual push probe) must run # against. Without this the probe is silently unavailable in the - # sandbox and captures are accepted on the pure rules only. + # sandbox and plans are scored on the pure rules only. model.sim_env = self._base_env # Belief-side verdicts predict the real evaluator, so the # certificate's verification replay must run the agent's FULL @@ -1305,8 +1019,6 @@ def _make_candidate_probe_model_provider( self, simulator_file: str, trajectories: List[LowLevelTrajectory], - base_pred_triples: List[Tuple[State, Action, State]], - inferred_hint: Dict[str, List[str]], ) -> Callable[[], _OracleOptionModel]: """Lazy option-model builder behind the synthesis run_python. @@ -1342,16 +1054,13 @@ def _provider() -> _OracleOptionModel: if cache.get("digest") == digest: return cache["model"] fit_state = self._probe_fit_state() - rules, specs, features, ns = \ - self._load_simulator_from_module_file( - simulator_file, trajectories) + rules, specs, _, ns = self._load_simulator_from_module_file( + simulator_file, trajectories) if rules is None or specs is None: raise RuntimeError( "run_python probe: ./simulator.py failed to load " "(exec error or missing simulator exports) - " "fix the file and probe again.") - residual_features = (features - if features is not None else inferred_hint) # The candidate's rule parameters, which the joint belief # covers beside AGENT_PARAM_SPECS (see parameter_belief). setattr(self, "_probe_rule_specs", list(specs)) @@ -1378,14 +1087,8 @@ def _provider() -> _OracleOptionModel: # runs a fit it may not afford (sketch seed1 learn 011: an # implicit refit inside a probe hit the call cap and came # back as an empty "param fitting failed:"). - model, params, _ = build_candidate_option_model( - self, - rules, - specs, - residual_features, - base_pred_triples, - latent_init=latent_init, - fit=False) + model, params = build_candidate_option_model( + self, rules, specs, latent_init=latent_init) if CFG.agent_sim_learn_declared_params_only: status = ("at the DECLARED values of the current " "simulator.py (harness parameter estimation is " @@ -1411,29 +1114,15 @@ def _provider() -> _OracleOptionModel: # ── Active-experiment ensemble (info-seeking exploration) ──── - def _info_seeking_active(self) -> bool: - """Whether the proactive info-seeking apparatus should run now. - - Delegates to the run context's adaptive gate. Partial unit-test - objects have no ``_tool_context``; there, fall back to the plain - flag (adaptive gating needs the run-scoped refusal signal the - context carries). - """ - ctx = getattr(self, "_tool_context", None) - if ctx is not None: - return ctx.info_seeking_active() - return CFG.agent_explorer_info_seeking - def _rebuild_param_ensemble(self) -> None: """Rebuild the learned model's rule-parameter ensemble. - Two consumers share it: info-seeking exploration (off under - ablation A6) and the capture gate's rule-param margin (off under - ablation A7). Built when either is on, a fit has populated - ``_fitted_params`` and parameter uncertainty is in use; cleared - otherwise. The ensemble can use an exploration-only posterior - even when solver params remain at the global-budget point - estimate. + Its consumer is info-seeking exploration + (:meth:`score_atom_disagreement`). Built when that is on, a fit + has populated ``_fitted_params`` and parameter uncertainty is in + use; cleared otherwise. The ensemble can use an exploration-only + posterior even when solver params remain at the global-budget + point estimate. Picks the most *calibrated* ensemble the fit affords, preferring spreads that reflect real posterior uncertainty over uniform @@ -1453,22 +1142,10 @@ def _rebuild_param_ensemble(self) -> None: "Rule-parameter ensemble: %d draws of the joint belief.", len(self._param_ensemble)) return - wanted = (CFG.agent_explorer_info_seeking - or CFG.agent_plan_validation_rule_param_margin) - if (not wanted or not self._fitted_params + if (not CFG.agent_explorer_info_seeking or not self._fitted_params or not CFG.agent_sim_learn_param_uncertainty): self._param_ensemble = [] return - if CFG.agent_sim_learn_oracle_sim_params: - # Oracle params carry no uncertainty: no fit ran, so the only - # ensemble on offer would be box jitter around the truth, - # which manufactures wrong models (a zero rate, a rewired - # lamp) that the capture gate would then demand every plan - # survive. Nothing to hedge against, so no ensemble. - self._param_ensemble = [] - logger.info("Oracle sim params: no rule-parameter ensemble " - "(nothing uncertain to sweep).") - return num_members = CFG.agent_explorer_info_ensemble_size self._param_ensemble, method = self._select_param_ensemble(num_members) logger.info( @@ -1536,9 +1213,8 @@ def score_atom_disagreement(self, state: State, i.e. an informative experiment. Returns 0.0 when the ensemble is trivial (<=1 member) or no atoms are given. - Wired into refinement as the info-scorer for the agent_model_based - explorer; a read-only query that leaves ``_fitted_params`` - unchanged on return. + Wired into the probe's refinement as the info-scorer; a read-only + query that leaves ``_fitted_params`` unchanged on return. Under ``agent_explorer_info_seeking_noise_aware`` with a declared observation-noise channel, each member reads the atoms from @@ -1606,9 +1282,7 @@ def _rule_param_override_scope( swap changes their gates for the wrapped validation rollout and the restore returns the deployed fit untouched - the same pattern :meth:`score_atom_disagreement` uses for ensemble - scoring. Installed on the tool context as - ``rule_param_override_scope`` for the capture gate's - rule-parameter margin sweep. + scoring. """ saved = dict(self._fitted_params) self._fitted_params.clear() @@ -1621,153 +1295,6 @@ def _rule_param_override_scope( # ── Agent-based synthesis ──────────────────────────────────── - def _synthesize_with_agent( - self, - trajectories: List[LowLevelTrajectory], - obs_triples: List[Tuple[State, Action, State]], - base_pred_triples: List[Tuple[State, Action, State]], - inferred_hint: Dict[str, List[str]], - ) -> None: - """Obtain RESIDUAL_RULES / PARAM_SPECS / RESIDUAL_FEATURES, then fit. - - ``inferred_hint`` is passed to the agent as a starting point and - used as the eval/test scope until it declares its own - ``RESIDUAL_FEATURES``. CFG flag - ``agent_sim_learn_oracle_sim_program`` short-circuits the agent - session by loading the GT simulator instead (and - ``agent_sim_learn_oracle_sim_params`` additionally skips the - parameter fit; see :meth:`_fit_params_after_synthesis`). - """ - if CFG.agent_sim_learn_oracle_sim_program: - rules, specs, residual_features = \ - self._load_oracle_sim_program(inferred_hint) - else: - loaded = self._run_agent_synthesis_session(trajectories, - obs_triples, - base_pred_triples, - inferred_hint) - if loaded is None: - return - rules, specs, residual_features = loaded - self._residual_rules = rules - self._residual_features = residual_features - self._fit_params_after_synthesis(rules, specs, base_pred_triples, - residual_features) - - def _load_oracle_sim_program( - self, inferred_hint: Dict[str, List[str]] - ) -> Tuple[List, List[ParamSpec], Dict[str, List[str]]]: - """Load the ground-truth simulator instead of running an agent. - - ``get_gt_simulator`` dispatches by observability: in - partially-observable mode it returns the PO GT simulator - (gt_simulator_po.py - latent heat threaded across steps, - surfaced as the observable bubbling_level), which predicts only - observable features; otherwise it returns the fully-observable - gt_simulator.py (which reads/writes heat_level as a State - feature). The two factories gate on CFG.partially_observable so - the env-name dispatch resolves to exactly one module per run. - - Unless ``agent_sim_learn_oracle_sim_params`` also holds, the - declared parameter inits are perturbed so the subsequent fit - starts from a miscalibrated - not oracle - belief. - """ - rules, specs, residual_features = get_gt_simulator(CFG.env) - self._log_feature_set_diff(inferred_hint, residual_features, - "inferred", "oracle") - if not CFG.agent_sim_learn_oracle_sim_params: - specs = self._perturb_spec_inits(specs) - logger.info("Loaded oracle sim program (%d rules, %d params).", - len(rules), len(specs)) - return rules, specs, residual_features - - @staticmethod - def _perturb_spec_inits(specs: List[ParamSpec]) -> List[ParamSpec]: - """Perturb each spec's init with multiplicative Gaussian noise. - - Used when the oracle sim PROGRAM is loaded but its param VALUES - must still be learned: the fit then starts from a plausible but - wrong belief instead of the answer. Each perturbed init is - clipped to its spec's box. - """ - rng = np.random.default_rng(CFG.seed) - noise_scale = CFG.agent_sim_learn_oracle_sim_param_noise_scale - if noise_scale < 0.0: - raise ValueError("agent_sim_learn_oracle_sim_param_noise_scale " - "must be non-negative.") - perturbed = [] - for s in specs: - val = float( - np.clip(s.init_value * (1.0 + rng.normal(0, noise_scale)), - s.lo, s.hi)) - perturbed.append( - ParamSpec(s.name, - val, - lo=s.lo, - hi=s.hi, - scale=getattr(s, "scale", "linear"))) - return perturbed - - def _run_agent_synthesis_session( - self, - trajectories: List[LowLevelTrajectory], - obs_triples: List[Tuple[State, Action, State]], - base_pred_triples: List[Tuple[State, Action, State]], - inferred_hint: Dict[str, List[str]], - ) -> Optional[Tuple[List, List[ParamSpec], Dict[str, List[str]]]]: - """Run one agent synthesis session and load what it committed. - - Returns ``(rules, specs, residual_features)``, or None when the - session left no loadable simulator artifact. Per-skill samplers - (when enabled) ride along in the same session. - """ - paths = self._resolve_synthesis_paths() - extra_paths = self._compute_extra_synthesis_paths(paths.base) - sampler_paths = (self._sampler_paths(paths.base) - if self._do_synthesize_samplers else {}) - exec_ns = self._build_synthesis_exec_ns(trajectories) - self._attach_synthesis_session_state(exec_ns, trajectories, - base_pred_triples, inferred_hint, - paths, extra_paths, sampler_paths) - # Fresh session so the synthesis prompt + tools take effect. - self._close_agent_session() - self._ensure_agent_session() - structs_ref = self._write_structs_reference() - base_sim_refs = self._base_sim_reference_paths() - message = self._build_synthesis_learn_message( - trajectories, obs_triples, inferred_hint, paths, structs_ref, - extra_paths, sampler_paths, base_sim_refs) - try: - responses = self._query_agent_sync(message, kind="learn") - dead = query_fatal_error(responses) - if dead is not None: - # The synthesis session never ran (usage limit, auth, - # transport): nothing was learned, so this cycle must - # not be checkpointed as learned. The cycle's explore - # episodes are stashed (main._save_inflight_interactions), - # so a relaunch resumes at exactly this learn. Silently - # continuing once wrote a byte-identical checkpoint and - # burned a whole cycle (2026-08-27 run_20260827_121111). - raise AgentSessionFatalError( - "The learn session died without the agent doing any " - f"work ({dead}); refusing to checkpoint this cycle as " - "learned.") - finally: - self._tool_context.extra_session_hooks = {} - self._tool_context.extra_mcp_tools = [] - self._tool_context.probe_artifact_loaders.clear() - self._tool_context.probe_option_model_provider = None - self._tool_context.probe_fit_provider = None - self._tool_context.probe_validation_provider = None - self._tool_context.probe_param_status = None - self._tool_context.probe_residuals_provider = None - self._tool_context.learn_cycle_index = None - self._learning_mode = False - self._close_agent_session() - return self._load_synthesis_artifacts(trajectories, inferred_hint, - paths, extra_paths, - sampler_paths) - def _resolve_synthesis_paths(self) -> _SynthesisPaths: """Host- and agent-visible paths for one synthesis session. @@ -1776,10 +1303,9 @@ def _resolve_synthesis_paths(self) -> _SynthesisPaths: in __init__, but it isn't constructed until ``_ensure_agent_session()`` runs later in the session setup. - The agent-visible paths differ by sandbox backend: cwd-relative - for local-sandbox (the validation hook resolves against cwd and - rejects literal ``/sandbox/...`` paths), the docker mount point - for docker, the absolute host path otherwise. + The agent-visible paths are cwd-relative in the local sandbox + (the validation hook resolves against cwd and rejects literal + ``/sandbox/...`` paths) and absolute host paths otherwise. """ if CFG.agent_sdk_use_local_sandbox: sandbox_dir: Optional[str] = os.path.abspath( @@ -1791,9 +1317,6 @@ def _resolve_synthesis_paths(self) -> _SynthesisPaths: if CFG.agent_sdk_use_local_sandbox: simulator_file_for_agent = "./simulator.py" sandbox_dir_for_agent: Optional[str] = "." - elif sandbox_dir: - simulator_file_for_agent = "/sandbox/simulator.py" - sandbox_dir_for_agent = "/sandbox" else: simulator_file_for_agent = simulator_file sandbox_dir_for_agent = None @@ -1847,174 +1370,12 @@ def describe_trajectory(traj_idx: int, self._make_evaluate_trajectory_fn() return exec_ns - def _attach_synthesis_session_state( - self, - exec_ns: Dict[str, Any], - trajectories: List[LowLevelTrajectory], - base_pred_triples: List[Tuple[State, Action, State]], - inferred_hint: Dict[str, List[str]], - paths: _SynthesisPaths, - extra_paths: Dict[str, str], - sampler_paths: Dict[str, str], - ) -> None: - """Install this synthesis session's state on the tool context. - - Everything installed here is cleared by the caller's ``finally`` - once the session query returns. - """ - # Label tool output (e.g. attempt-log headers) with the - # learning cycle for the duration of this session. - self._tool_context.learn_cycle_index = self._learning_cycle_index() - # Build dynamic synthesis tools and attach them to the tool - # context *before* opening the session. The attached set is - # filtered against ``_get_synthesis_tool_names`` so that method - # is the single source of truth for what the agent sees: - # anything a builder constructs but the names list omits is - # dropped here. - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.belief_probe import _check_time_budget - toolkit = create_synthesis_tools( - exec_ns, - base_pred_triples, - inferred_hint, - simulator_file=paths.simulator_file, - versions_dir=paths.versions_dir, - approach=self, - sandbox_dir=paths.base, - sandbox_dir_for_agent=paths.sandbox_dir_for_agent, - cycle_index_provider=self._learning_cycle_index, - budget_check=lambda: _check_time_budget(self._tool_context), - ) - tools = list(toolkit.tools) - self._install_extra_synthesis_surfaces(exec_ns, base_pred_triples, - inferred_hint, extra_paths) - if self._do_synthesize_samplers: - self._install_sampler_surface(sampler_paths) - declared = set(self._get_synthesis_tool_names() or ()) - self._tool_context.extra_mcp_tools = [ - t for t in tools if getattr(t, "name", "") in declared - ] - # Point the probe at the CANDIDATE simulator for this session - # (never the stale pre-synthesis option model; on cycle 1 that - # wraps the real env), then merge the probe facade into - # run_python's namespace: synthesis sessions offer ONE exec - # namespace, so helpers defined next to the data are visible to - # probe sweeps (create_mcp_tools skips the solve-phase instance - # when this one is attached). Unconditional: with fit / refine / - # forward-validation all living on ``sim``, the probe IS the - # validation surface, so a synthesis session without it would - # have no way to test what it writes. Only ``sim``/``BeliefProbe`` - # are taken from the probe namespace: ``trajectories`` already - # binds the fit list and solve-only extras do not apply. - self._tool_context.probe_option_model_provider = \ - self._make_candidate_probe_model_provider( - paths.simulator_file, trajectories, base_pred_triples, - inferred_hint) - self._tool_context.probe_fit_provider = toolkit.fit_runner - self._tool_context.probe_validation_provider = toolkit.validation_runner - self._tool_context.probe_residuals_provider = \ - toolkit.residuals_runner - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.belief_probe import build_probe_namespace - probe_ns = build_probe_namespace(self._tool_context) - exec_ns["sim"] = probe_ns["sim"] - exec_ns["BeliefProbe"] = probe_ns["BeliefProbe"] - self._learning_mode = True - # PostToolUse hook: snapshot simulator.py / predicates.py on - # every successful Write/Edit/MultiEdit, so the version history - # covers everything the agent committed to file (not just - # states that happened to coincide with an eval call). Only - # active for this synthesis session. - snapshot_targets = self._build_write_snapshot_targets( - paths.simulator_file, paths.versions_dir, extra_paths) - if self._do_synthesize_samplers: - snapshot_targets.append( - self._sampler_snapshot_target(sampler_paths)) - self._tool_context.extra_session_hooks = ( - self._build_synthesis_session_hooks(snapshot_targets, paths.base)) - - def _build_synthesis_learn_message( - self, - trajectories: List[LowLevelTrajectory], - obs_triples: List[Tuple[State, Action, State]], - inferred_hint: Dict[str, List[str]], - paths: _SynthesisPaths, - structs_ref: str, - extra_paths: Dict[str, str], - sampler_paths: Dict[str, str], - base_sim_refs: Optional[List[str]] = None, - ) -> str: - """Compose the synthesis session's first user message. - - Gathers this cycle's data roster, digests, and reports and - renders ``learn_message.md``. Reads the just-opened session's - tool names, so the session must be open before this is called. - """ - n_trajs = len(trajectories) - n_demos = sum(1 for t in trajectories if t.is_demo) - # Start-of-session divergence report: when a prior model exists, - # score it (params refit to ALL data, so what remains is the - # structural gap) before the agent's first turn - the session - # then starts from "here is where the model breaks" instead of - # spending turns rediscovering it. The same report stays callable - # as `sim.residuals()` against every subsequent edit. With no - # prior model the "prior" is the bare base simulator and every - # mismatch is an unmodeled mechanism. - prior_state_block = self._format_prior_state_block(paths.base) - divergence_block = "" - if (self._tool_context.probe_residuals_provider is not None - and obs_triples): - try: - report = self._tool_context.probe_residuals_provider( - max_transitions=100000, - fit_params=not CFG.agent_sim_learn_declared_params_only) - divergence_block = learn_prompts.render_divergence_block( - report, has_prior_model=bool(prior_state_block)) - except Exception as e: # pylint: disable=broad-except - logger.warning("Skipping start-of-session residual report: %s", - e) - session_tool_names = (self._agent_session.tool_names - if self._agent_session is not None else []) - extra_messages = [] - if not trajectories and CFG.agent_sim_learn_zero_shot: - extra_messages.append(learn_prompts.render_zero_shot_message()) - extra_message = self._extra_synthesis_message(extra_paths) - if extra_message: - extra_messages.append(extra_message) - if self._do_synthesize_samplers: - extra_messages.append( - self._sampler_synthesis_message(sampler_paths)) - return learn_prompts.build_learn_message( - n_trajs=n_trajs, - n_transitions=len(obs_triples), - n_demos=n_demos, - n_interaction=n_trajs - n_demos, - trajectory_listing=self._format_trajectory_listing(trajectories), - structs_ref=structs_ref, - inferred_hint=str(inferred_hint), - predicate_listing=self._format_predicate_signatures( - self._get_all_predicates()), - types_digest=render_types_digest(self._tool_context.types), - options_digest=render_options_digest( - self._tool_context.options, - gt_options_ref_path=self._tool_context.gt_options_ref_path), - simulator_file=paths.simulator_file_for_agent, - objective_block=self._format_objective_block(), - prior_state_block=prior_state_block, - divergence_block=divergence_block, - base_sim_block=learn_prompts.render_base_sim_block(base_sim_refs - or []), - tools_block=learn_prompts.render_tools_block(session_tool_names), - extra_messages=extra_messages, - ) - def _load_synthesis_artifacts( self, trajectories: List[LowLevelTrajectory], inferred_hint: Dict[str, List[str]], paths: _SynthesisPaths, extra_paths: Dict[str, str], - sampler_paths: Dict[str, str], ) -> Optional[Tuple[List, List[ParamSpec], Dict[str, List[str]]]]: """Load the artifacts the finished session committed to disk. @@ -2078,8 +1439,6 @@ def _load_synthesis_artifacts( logger.info("Agent synthesized %d rules, %d params.", len(rules), len(specs)) self._post_synthesis_loading(extra_paths, specs) - if self._do_synthesize_samplers: - self._finalize_and_load_samplers(sampler_paths) return rules, specs, residual_features def _fit_params_after_synthesis( @@ -2089,30 +1448,12 @@ def _fit_params_after_synthesis( base_pred_triples: List[Tuple[State, Action, State]], residual_features: Dict[str, List[str]], ) -> None: - """Deploy the agent's parameters; only oracle programs fit here.""" + """Deploy the agent's parameters; the harness never fits here.""" if getattr(self, "_residual_env_cls", None) is not None and \ not specs and not self._physical_param_specs: self._fitted_params.clear() self._last_fit_result = None self._fit_sse = float("inf") - elif CFG.agent_sim_learn_oracle_sim_params: - self._fitted_params.clear() - self._fitted_params.update({s.name: s.init_value for s in specs}) - if self._physical_param_specs: - # Oracle mode: trust the agent-declared physical inits. - self._apply_identified_physical_params( - {s.name: s.init_value - for s in self._physical_param_specs}) - # No fit ran; the ensemble falls back to uniform perturbation. - self._last_fit_result = None - if base_pred_triples: - self._fit_sse = self._oracle_param_sse(rules, - base_pred_triples, - residual_features, - FIT_NOISE_SIGMA) - else: - logger.info("No transitions; skipping oracle-param SSE.") - self._fit_sse = float("inf") elif CFG.agent_sim_learn_declared_params_only: self._deploy_declared_params(rules, specs, base_pred_triples, residual_features) @@ -2158,33 +1499,14 @@ def _fit_params_after_synthesis( version, len(expected), self._fit_sse) applied = self._probe_fit_state().get("applied_physical") if self._physical_param_specs and applied: - # Mirror _fit_parameters_joint_rollout's deploy: the - # cycle-level applied snapshot and the physics-margin - # sigma points come from the published fit (applying - # resets the points, so set them after). + # The physics-margin sigma points come from the + # published fit (applying resets the points, so set + # them after). self._apply_identified_physical_params(dict(applied)) - self._cycle_applied_physical = dict(applied) if CFG.agent_sim_learn_param_uncertainty and int( CFG.belief_joint_draws) <= 0: self._identified_physical_sigma_points = list( self._probe_fit_state().get("sigma_points") or []) - elif CFG.agent_sim_learn_oracle_sim_program and base_pred_triples: - # This baseline supplies a program without an agent - # session. Fitting is its explicitly configured protocol. - logger.info("Oracle sim program: fitting its " - "parameters on the harness side.") - if self._physical_param_specs or has_physics_rules(rules): - fit_result, self._fit_sse = ( - self._fit_parameters_joint_rollout( - rules, specs, residual_features)) - elif has_latent_rules(rules): - fit_result, self._fit_sse = \ - self._fit_parameters_recurrent( - rules, specs, base_pred_triples, - residual_features) - else: - fit_result, self._fit_sse = fit_rule_parameters( - rules, specs, base_pred_triples, residual_features) else: self._deploy_unfitted_params(specs) if fit_result is not None: @@ -2212,7 +1534,6 @@ def _deploy_unfitted_params(self, specs: List[ParamSpec]) -> None: if physical or self._identified_physical_params: self._apply_identified_physical_params(physical) self._identified_physical_sigma_points = [] - self._cycle_applied_physical = dict(physical) # Retain the historical published fit for provenance and file # reversions, but do not use its SSE or posterior for this model. self._last_fit_result = None @@ -2227,7 +1548,7 @@ def _physics_margin_points( report: Dict[str, Dict[str, Any]], physical_specs: List[ParamSpec], ) -> List[Dict[str, float]]: - """The capture gate's physics-margin grid for ``applied``. + """The physics sweep's +-1-sigma grid for ``applied``. Empty under ``agent_sim_learn_param_uncertainty`` False (ablations A6+A7 combined: point estimates only, so there is no width to @@ -2268,7 +1589,6 @@ def _deploy_declared_params( if physical_specs: applied = {s.name: s.init_value for s in physical_specs} self._apply_identified_physical_params(applied) - self._cycle_applied_physical = dict(applied) self._identified_physical_sigma_points = \ self._physics_margin_points( applied, declared_interval_report(physical_specs), @@ -2456,47 +1776,6 @@ def _rollout_fit_trajectories( "keeping the whole trajectories.") return rollouts - def _persist_fit_trajectories(self, label: str = "fitted") -> None: - """Dump the raw fit trajectories for offline post-mortems. - - ``label`` distinguishes the two moments this is called from: - ``recorded`` when a cycle's data arrives (always), ``fitted`` - when a sysID fit has just run and the payload's identified - params mean something. The first is what makes a cycle that - declined to fit replayable at all. - - The rollout sysID fit data otherwise exists only in memory: - when run_20260724_232411 shipped friction fits 2-7 sigma from - the truth, the failing fits could not be replayed offline - the - episodes had to be approximately re-executed from logged plans, - which cannot reproduce mid-episode replans or the warm-env - recording context (exactly the suspected corruption channel). - One pickle per cycle-level fit under ``/fit_data/``; - never raises - persistence must not take down a run. - """ - if not CFG.code_sim_learning_persist_fit_data: - return - try: - out_dir = os.path.join(self._get_log_dir(), "fit_data") - os.makedirs(out_dir, exist_ok=True) - idx = len([f for f in os.listdir(out_dir) if f.endswith(".pkl")]) - path = os.path.join(out_dir, - f"fit_trajectories_{idx:03d}_{label}.pkl") - payload = { - "trajectories": - list(self._fit_trajectories), - "physical_param_specs": - list(self._physical_param_specs), - "identified_physical_params": - dict(self._identified_physical_params), - } - with open(path, "wb") as f: - pkl.dump(payload, f) - logger.info("Persisted %d fit trajectories to %s", - len(self._fit_trajectories), path) - except Exception as e: # pylint: disable=broad-except - logger.warning("Could not persist fit trajectories: %s", e) - def _apply_identified_physical_params( self, identified: Dict[str, float]) -> None: """Publish identified physical params into the planning base env. @@ -2591,392 +1870,8 @@ def previous_fit_evidence( } return None - def _fit_parameters_joint_rollout( - self, - rules: List, - rule_specs: List[ParamSpec], - residual_features: Dict[str, List[str]], - ) -> Tuple[FitResult, float]: - """Joint physical+rule fit against free-running base-sim rollouts. - - Reached when the artifact declares ``PHYSICAL_PARAM_SPECS``. Consumes - the RAW observed trajectories rather than ``base_pred_triples``: - physical parameters only manifest when momentum free-runs, which - the teacher-forced triples destroy (``State`` has no - velocities). One theta = physical + rule params, one joint fit, - so rules cannot silently absorb physics error; with no rules - this degenerates to pure identification. The identified physical - values are applied in place to the planning base env, and the - per-parameter identifiability report (posterior contraction) is - logged so null parameters are visible rather than silently - trusted. - """ - physical_specs = self._physical_param_specs - physical_names = [s.name for s in physical_specs] - # Factory, not an instance: every rollout runs in a fresh env. - fit_env = self._get_rollout_fit_env() - self._persist_fit_trajectories() - rollouts = self._rollout_fit_trajectories(residual_features) - init_params = { - s.name: s.init_value - for s in physical_specs + rule_specs - } - anchors = self.fit_prior_anchors(physical_specs) - if not rollouts: - logger.warning( - "No complete trajectories for rollout sysID; keeping the " - "declared physical-param inits unfitted.") - result = fit_params_rollout(fit_env, [], - physical_specs, - residual_features, - rules=rules, - rule_specs=rule_specs, - latent_init=self._latent_init, - anchors=anchors) - self._apply_identified_physical_params( - {n: init_params[n] - for n in physical_names}) - return result, float("nan") - - # The adjuster's third argument is the fit's own SSE-at-theta - # probe (survivor set + shared scaling - the objective the fit - # minimized), which the cross-cycle consistency check uses to - # arbitrate flagged jumps on evidence. Deliberately NOT a - # full-set SSE: trimmed (unexplainable) segments would add the - # same large error to both candidates and dilute the ratio. - outcome = run_rollout_sysid( - fit_env, - rollouts, - physical_specs, - residual_features, - rules=rules, - rule_specs=rule_specs, - latent_init=self._latent_init, - anchors=anchors, - rms_cache=self._explainability_cache, - report_adjuster=lambda result, report, sse_fn: - (self._check_cross_cycle_consistency( - result, report, physical_names, pooled_sse=sse_fn)), - held={ - **self._identified_physical_params, - **self._cycle_applied_physical - }) - if outcome.num_survivors == 0: - # NO fit ran (the result is pinned at the declared inits). - # Apply nothing: the planner keeps its standing belief - - # the previous cycle's applied values - rather than being - # reverted to baselines (or moved to this call's declared - # inits) by data that supports neither. - logger.warning( - "Rollout sysID: no explainable segments this cycle; " - "leaving the planner's physical params untouched.") - self._record_sysid_diagnostics({}, physical_names, 0, - len(rollouts), outcome.traj_rms) - return outcome.fit_result, float("nan") - inference = outcome.inference - logger.info("Identifiability (posterior/prior contraction):\n%s", - format_identifiability(inference.parameter_diagnostics)) - log_param_changes(init_params, inference.point_estimate) - self._apply_identified_physical_params(inference.selected_parameters) - self.note_carried_posterior(inference.selected_parameters, - inference.parameter_diagnostics) - if outcome.evidence is not None: - self.note_fit_evidence( - self._current_simulator_version or "harness", outcome.evidence) - # Snapshot the cycle-level decision: this (not whatever the - # agent's in-session sim.fit last applied) is what a future - # INCONSISTENT verdict holds on to. - self._cycle_applied_physical = dict(inference.selected_parameters) - # Physics-margin points for the capture gate: the fit's posterior - # widths (floored, see identifiability_report) turned into a grid - # of perturbations spanning +-1 sigma of the applied values. - self._identified_physical_sigma_points = self._physics_margin_points( - inference.selected_parameters, inference.parameter_diagnostics, - physical_specs) - if self._identified_physical_sigma_points: - logger.info("Physics-margin points for capture validation: %s", - [{k: f"{v:.4f}" - for k, v in pt.items()} - for pt in self._identified_physical_sigma_points]) - self._record_sysid_diagnostics(inference.parameter_diagnostics, - physical_names, outcome.num_survivors, - len(rollouts), outcome.traj_rms) - return outcome.fit_result, outcome.post_sse - - def _check_cross_cycle_consistency( - self, - result: FitResult, - report: Dict[str, Dict[str, Any]], - physical_names: Sequence[str], - pooled_sse: Optional[Callable[[Dict[str, float]], - float]] = None) -> None: - """Flag params whose confident MAP jumped since the previous cycle. - - The curvature probe measures local *precision*: a biased - objective yields precisely-wrong values that the probe still - stamps "identified" (observed: per-cycle friction fits 0.0585, - 0.0614, 0.0919, 0.0794, each with posterior_std ~0.003 - - mutually incompatible by many sigmas). Comparing successive - final fits in FIT space (log for log-scale params) catches - exactly this: a jump above - ``CFG.code_sim_learning_rollout_cross_cycle_sigma`` combined - sigmas sets ``Verdict.INCONSISTENT``, and the trust selection - then HOLDS the currently-applied value instead of hopping to - the new fit - neither of two mutually-incompatible confident - fits can be preferred on this evidence, and hopping churned the - belief env for whole runs (run_20260721_205821 seed1: - restitution 0.71 -> 0.52 -> 0.02 -> 0.32 -> 0.02). History - records only these final per-cycle fits, not the agent's - in-session tool fits, whose param sets churn. - - Sigma distance alone cannot tell a real correction from probe - churn, and successive cycle fits are NOT independent equals: - the new fit minimized the objective over a superset of the old - fit's data. So before holding, a flagged jump is arbitrated on - evidence via ``pooled_sse`` (the fit's own SSE-at-theta probe - over its surviving segments - the objective it minimized): - when the held value explains that data decisively worse than - the new fit - (``CFG.code_sim_learning_rollout_consistency_sse_ratio``), the - jump is accepted. Without a decisive gap the hold stands - (run_20260724_232411-style subset disagreement stays held + - hull-swept). Motivated by run_20260727_210827 seed1: a sharp - but biased 2-trajectory cycle-0 fit (0.9313, true 0.5) was held - over the 4-trajectory refit (0.4748, pooled SSE 0.14 vs ~4.4) - for the rest of the run. - """ - k = CFG.code_sim_learning_rollout_cross_cycle_sigma - fitted = result.point_estimate - scales = result.scales or ["linear"] * len(result.names) - for i, name in enumerate(result.names): - if name not in physical_names: - continue - if report.get(name, {}).get("verdict") is Verdict.ANCHORED: - # An ablation-reverted param's point estimate IS the - # baseline, not a fit. Recording it would make the next - # cycle's genuine fit read as a many-sigma jump (and - # spuriously downgrade it); keep the previous history - # entry, which holds the last real fit. - continue - post = float( - report.get(name, {}).get("posterior_std", float("nan"))) - scale = scales[i] - value = fitted[name] - prev = self._sysid_fit_history.get(name) - flagged = False - if (k > 0 and prev is not None and np.isfinite(post) - and np.isfinite(prev[1])): - prev_val, prev_std, prev_scale = prev - if prev_scale == scale: - dist = _fit_space_dist(value, prev_val, scale) - combined = float(np.sqrt(post**2 + prev_std**2)) - if combined > 0 and dist / combined > k: - n_sigma = dist / combined - pending = self._sysid_pending_fit.get(name) - if pending is not None and _fit_space_dist( - value, pending[0], scale) / max( - float(np.sqrt(post**2 + pending[1]**2)), - 1e-12) <= k: - # Two INDEPENDENT cycles agree on the new - # value: the jump was real, not probe - # overconfidence - accept it. - logger.info( - "Rollout sysID cross-cycle consistency: " - "%s jump to ~%.4f confirmed by an " - "independent refit (pending %.4f); " - "accepting the new value.", name, value, - pending[0]) - self._sysid_pending_fit.pop(name, None) - elif self._arbitrate_cross_cycle_jump( - name, fitted, prev_val, pooled_sse): - # Pooled evidence decisively prefers the - # new fit over the held value (the helper - # logs the SSE gap); accept it now instead - # of waiting a cycle for confirmation. - self._sysid_pending_fit.pop(name, None) - else: - flagged = True - logger.warning( - "Rollout sysID cross-cycle consistency: %s " - "moved %.4f -> %.4f (%.1f combined sigmas > " - "%g) since the previous cycle; the posterior " - "is overconfident.", name, prev_val, value, - n_sigma, k) - entry = report.get(name) - if (entry is not None and - entry["verdict"] is Verdict.IDENTIFIED): - entry["verdict"] = Verdict.INCONSISTENT - entry["note"] = ( - f"{prev_val:.4f} -> {value:.4f} is " - f"{n_sigma:.1f} combined sigmas; holding " - "the last trusted value, margin sweep " - "spans both") - # Both incompatible fits become hull - # candidates so the margin sweep covers - # the whole disagreement - the interval - # [0.3236, 0.6267] contained the true - # 0.5 in run_20260724_232411 seed2. - cands = set(entry.get("candidate_values", ())) - cands.update((float(prev_val), float(value))) - if pending is not None: - cands.add(float(pending[0])) - entry["candidate_values"] = sorted(cands) - self._sysid_pending_fit[name] = (value, post) - if flagged: - # Keep the trusted value as the comparison reference; - # the rejected fit waits in _sysid_pending_fit for an - # independent confirmation. Recording the rejected fit - # here would make it the NEXT cycle's reference, i.e. - # accept the hop one cycle late without any new - # evidence. - continue - self._sysid_pending_fit.pop(name, None) - self._sysid_fit_history[name] = (value, post, scale) - - @staticmethod - def _arbitrate_cross_cycle_jump( - name: str, fitted: Dict[str, float], prev_val: float, - pooled_sse: Optional[Callable[[Dict[str, float]], float]]) -> bool: - """Settle a flagged cross-cycle jump by pooled-data evidence. - - Evaluates ``pooled_sse`` (the fit's own objective over its - surviving segments) under the new joint fit and under the same - fit with ``name`` swapped back to the held value. Returns True - (accept the jump) only when the held - value's explanation is decisively worse - at least - ``CFG.code_sim_learning_rollout_consistency_sse_ratio`` times - the new fit's SSE. Anything short of decisive (including an SSE - evaluation failure) returns False and leaves the hold-and- - hull-sweep behavior in charge. - """ - ratio = CFG.code_sim_learning_rollout_consistency_sse_ratio - if pooled_sse is None or ratio <= 0: - return False - try: - sse_new = pooled_sse(dict(fitted)) - held_theta = dict(fitted) - held_theta[name] = prev_val - sse_held = pooled_sse(held_theta) - except Exception: # pylint: disable=broad-except - logger.warning( - "Rollout sysID cross-cycle arbitration: pooled SSE " - "evaluation failed for %s; holding the trusted value.", - name, - exc_info=True) - return False - if not (np.isfinite(sse_new) and np.isfinite(sse_held)): - return False - decisive = sse_held > ratio * sse_new - logger.info( - "Rollout sysID cross-cycle arbitration: %s pooled SSE %.4g at " - "the new fit %.4g vs %.4g at the held value %.4g - %s.", name, - sse_new, fitted[name], sse_held, prev_val, - ("decisively better, accepting the jump" - if decisive else "not decisive, holding")) - return decisive - - def _record_sysid_diagnostics(self, report: Dict[str, Dict[str, Any]], - physical_names: Sequence[str], - num_survivors: int, num_segments: int, - rms: List[float]) -> None: - """Digest the fit's weak spots for the next explore phase. - - Generic (domain-free) statements of what the data could not - support - unexplainable segments, parameters the rollouts do not - constrain, cross-cycle conflicts - phrased as experiment - objectives. The explorer appends this to its guidance so the - agent designs interactions that fill the gaps, instead of - relying on whatever manipulation data the tasks happen to - produce. - """ - lines: List[str] = [] - dropped = num_segments - num_survivors - if dropped > 0: - lines.append( - f"- {dropped} of {num_segments} recorded motion segments " - "were unexplainable at ANY physical parameters (best RMS " - f"{[f'{r:.3g}' for r in rms]}): their dynamics are not " - "repeatable under replay. Prefer experiments whose outcome " - "is dominated by object dynamics rather than prolonged " - "robot-object contact: actuate cleanly, then let the scene " - "evolve and settle on its own.") - for name in physical_names: - entry = report.get(name, {}) - verdict = entry.get("verdict", Verdict.UNKNOWN) - note = entry.get("note", "") - label = verdict.value + (f" ({note})" if note else "") - if verdict is Verdict.IDENTIFIED: - cands = entry.get("candidate_values", ()) - if len(cands) > 1: - lines.append( - f"- physical param '{name}': identified, but the " - "recorded segments preferred mutually-incompatible " - f"explanations spanning [{min(cands):.4g}, " - f"{max(cands):.4g}] (the physics-margin sweep " - "covers that whole hull). A clean, repeatable " - "interaction that excites this parameter and " - "little else would collapse the hull.") - continue - belief = entry.get("belief_interval") - if verdict is Verdict.WIDE and belief is not None: - anchor = entry.get("anchor") - if anchor is None: - where = "" - elif anchor < belief[0] or anchor > belief[1]: - where = (f"; the baseline {anchor:.4g} lies outside it, " - "so the data already exclude the baseline") - else: - where = f"; the baseline {anchor:.4g} lies inside it" - lines.append( - f"- physical param '{name}': the data moved it to " - f"{entry.get('map', float('nan')):.4g} but only weakly; " - f"the planner's belief is the interval [{belief[0]:.4g}, " - f"{belief[1]:.4g}]{where}. Plans are certified across " - "the whole interval, so an experiment whose observable " - "outcome DIFFERS across it would narrow the belief and " - "widen the set of certifiable plans.") - continue - interval = entry.get("flat_interval") - if (verdict in (Verdict.WEAKLY_IDENTIFIED, Verdict.NOT_IDENTIFIED) - and interval is not None and interval[0] != interval[1]): - lines.append( - f"- physical param '{name}': the data cannot " - f"distinguish values in [{interval[0]:.4g}, " - f"{interval[1]:.4g}]. An experiment whose observable " - "outcome DIFFERS across this interval would pin it " - "down.") - continue - if verdict is Verdict.ANCHORED: - # Anchor ablation handled this param correctly (the move - # was compensatory; the baseline is applied) - it is NOT - # a failed identification, so don't advise dropping it. - lines.append( - f"- physical param '{name}': the fitted move was " - "compensatory (a refit with it at its baseline explains " - "the data equally well), so the baseline was kept. An " - "experiment that excites this parameter SPECIFICALLY " - "(not jointly with the others) would distinguish the " - "two explanations.") - continue - if verdict is Verdict.INCONSISTENT: - lines.append( - f"- physical param '{name}': successive cycles produced " - f"confident but mutually-incompatible fits ({note}). " - "The objective is biased somewhere: collect a clean, " - "repeatable interaction that excites this parameter and " - "little else, so one of the two values can be refuted.") - continue - lines.append( - f"- physical param '{name}': {label}. An experiment whose " - "observable outcome CHANGES when this parameter changes " - "would identify it; if none exists, drop it from " - "PHYSICAL_PARAM_SPECS.") - self._last_sysid_diagnostics = ("\n".join(lines) if lines else "") - def _sync_tool_context(self) -> None: super()._sync_tool_context() - self._tool_context.sysid_diagnostics = (self._last_sysid_diagnostics - or None) self._tool_context.latent_tracking_available = \ self._latent_tracking_available() @@ -3073,38 +1968,6 @@ def _group_triples_by_trajectory( idx += n return groups - def _fit_parameters_recurrent( - self, - rules: List, - specs: List[ParamSpec], - base_pred_triples: List[Tuple[State, Action, State]], - residual_features: Dict[str, List[str]], - lm_seed: Optional[Tuple[np.ndarray, Optional[np.ndarray]]] = None, - ) -> Tuple[FitResult, float]: - """LM fit over the recurrent (per-trajectory) SSE. - - Counterpart to :func:`fitting.fit_rule_parameters` for rules - that carry a latent block. Re-groups the flat - ``base_pred_triples`` into per-trajectory chunks (latent threads - within a trajectory, not across) via the lengths cached in - ``self._fit_trajectories``; falls back to a single trajectory if - no grouping info exists. Delegates the actual fit/log to - :func:`fitting.fit_rule_parameters_latent` so the agent's - ``sim.fit`` surface scores latent rules through the exact - same path. - """ - groups = self._group_triples_by_trajectory(base_pred_triples) - if not groups: - logger.warning("No trajectory groups for recurrent fitting; " - "falling back to single-trajectory rollout.") - groups = [base_pred_triples] - return fit_rule_parameters_latent(rules, - specs, - groups, - self._latent_init, - residual_features, - lm_seed=lm_seed) - def _oracle_param_sse_recurrent( self, rules: List, @@ -3168,37 +2031,6 @@ def _oracle_param_sse_rollout( logger.info(" %-30s %.4f", name, val) return sse - def _attach_initial_latent(self, task: Task) -> Task: - """Seed ``task.init.latent`` with the initial latent block. - - Refinement starts at ``task.init`` (the planner's ``traj[0]``), so - the combined simulator must find a well-formed latent there. If no - ``LATENT_INIT`` was loaded (or the resulting block is empty), leave - the task alone so downstream code keeps the legacy - ``state.latent is None`` behaviour. Overrides the no-op default in - :class:`AgentModelBasedApproach`. - """ - tracker = self.make_latent_tracker() if getattr( - self, "_residual_env_cls", None) is not None else None - if tracker is not None: - return Task(init=tracker.attach(task.init, None), - goal=task.goal, - alt_goal=task.alt_goal, - goal_nl=task.goal_nl, - evaluator=task.evaluator) - if self._latent_init is None: - return task - initial_latent = init_latent(self._latent_init, self._fitted_params - or {}) - if not initial_latent: - return task - init_state = task.init.copy() - init_state.latent = initial_latent - return Task(init=init_state, - goal=task.goal, - alt_goal=task.alt_goal, - goal_nl=task.goal_nl) - def materialise_latent( self, traj: LowLevelTrajectory, @@ -3325,25 +2157,6 @@ def _compute_base_pred_triples( return [(base_env.simulate(s, a), a, s_next) for s, a, s_next in obs_triples] - @staticmethod - def _infer_residual_features_from_scan( - obs_triples: List[Tuple[State, Action, State]], - base_pred_triples: List[Tuple[State, Action, State]], - abs_tol: float = 1e-4, - rel_tol: float = 1e-3, - min_hits: int = 3, - ) -> Dict[str, List[str]]: - """Features whose base-sim prediction diverges from observation. - - Flags ``(type, feat)`` if ``|pred - obs| > rel_tol*|obs| + abs_tol`` - on at least ``min_hits`` triples. The ``min_hits`` floor keeps - one-off PyBullet jitter from leaking base-handled features into the set. - """ - del obs_triples # objects are identical across both triple lists - hits: Dict[Tuple[str, str], int] = {} - count_residual_hits(base_pred_triples, hits, abs_tol, rel_tol) - return residual_hint_from_hits(hits, min_hits) - @staticmethod def _log_feature_set_diff( a: Dict[str, List[str]], @@ -3366,23 +2179,6 @@ def _log_feature_set_diff( if only_b: logger.info(" only in %s: %s", b_label, only_b) - @staticmethod - def _format_predicate_signatures(predicates: Set[Predicate]) -> str: - """Pretty-print predicates as ``Name(type1, type2)`` lines. - - Mirrors the ``## Available Predicates`` block in - ``bilevel_sketch.build_solve_prompt``. - """ - lines = [] - for pred in sorted(predicates, key=lambda p: p.name): - type_sig = ", ".join(t.name for t in pred.types) - line = f" {pred.name}({type_sig})" - if pred.natural_language_assertion is not None: - names = [t.name for t in pred.types] - line += f" - {pred.natural_language_assertion(names)}" - lines.append(line) - return "\n".join(lines) - def _make_evaluate_trajectory_fn(self) -> Any: """Build the ``evaluate_trajectory`` helper exposed in the synthesis exec namespace (next to ``is_goal_state``). @@ -3428,8 +2224,8 @@ def evaluate_trajectory(states: Sequence[State], ``physics_sweep=True`` also scores the sequence at every point of the identified physical parameters' belief - interval (the same grid ``sim.run(physics_sweep=True)`` and - the capture gate use), each on a fresh env at that physics, + interval (the same grid ``sim.run(physics_sweep=True)`` + uses), each on a fresh env at that physics, and adds ``sweep``: the per-point verdicts and the fraction scored solved. A verdict that replays physics can flip across the interval; a sequence is certified only when it @@ -3511,72 +2307,6 @@ def _sweep_evaluation(self, evaluator: Any, states: List[State], "certified": solved == len(per_point), } - @staticmethod - def _format_trajectory_listing( - trajectories: List[LowLevelTrajectory]) -> str: - """Render a per-trajectory listing with provenance tags. - - Each interaction trajectory shows the simulator / predicates - snapshot used to generate the plan that collected it (if - tracked). Demo trajectories list as ``demo``. Listed in the same - order the agent sees them via the ``trajectories`` var. - """ - if not trajectories: - return "" - lines = ["Trajectory roster (matches the `trajectories` list):"] - for idx, traj in enumerate(trajectories): - kind = "demo" if traj.is_demo else "interaction" - try: - task_str = f"task {traj.train_task_idx}" - except AssertionError: - task_str = "task ?" - provenance: List[str] = [] - sim_v = traj.source_simulator_version - preds_v = traj.source_predicates_version - if sim_v: - provenance.append(f"sim {sim_v}") - if preds_v: - provenance.append(f"predicates {preds_v}") - tail = (f" - generated using {', '.join(provenance)}" - if provenance else "") - if traj.env_reward is not None: - solved = int( - bool(traj.env_terminated) and not traj.env_rejected) - tail += (f" - env reward={traj.env_reward:.2f} " - f"(solved={solved})") - lines.append(f" [{idx}] {kind}, {task_str}{tail}") - return "\n".join(lines) + "\n" - - def _format_objective_block(self) -> str: - """The env's public task objective (reward form), or empty. - - Emitted when a train task's evaluator states an objective. The - statement is public by design: it contains the reward FORM - (success condition + costs), never oracle quantities like the - true minimum block count. - """ - description = next( - (t.evaluator.objective_description() - for t in self._train_tasks if t.evaluator is not None - and t.evaluator.objective_description()), "") - return learn_prompts.render_objective_block(description) - - def _format_prior_state_block(self, base: str) -> str: - """Tell the agent about any simulator/predicates left over from a - previous learning cycle. - - Returns a paragraph the agent can act on (read the files first - and treat this cycle as incremental refinement) or an empty - string if no prior state exists. The base sandbox dir is scanned - for ``simulator.py`` / ``predicates.py``. - """ - prior: List[str] = [] - if os.path.isfile(os.path.join(base, "simulator.py")): - prior.append("`./simulator.py`") - if os.path.isfile(os.path.join(base, "predicates.py")): - prior.append("`./predicates.py`") - return learn_prompts.render_prior_state_block(prior) - def _simulator_load_namespace(self) -> Dict[str, Any]: """The names pre-injected when ``simulator.py`` is exec'd. @@ -3662,35 +2392,6 @@ def _load_simulator_from_module_file( # ── Static helpers ─────────────────────────────────────────── - def _write_structs_reference(self) -> str: - """Write key struct sources to the sandbox; return the agent-visible - path.""" - # pylint: disable=import-outside-toplevel,reimported - from predicators.structs import Action as _Action - from predicators.structs import LowLevelTrajectory as _LLT - from predicators.structs import Object as _Object - from predicators.structs import State as _State - from predicators.structs import Type as _Type - - source = "\n\n".join( - inspect.getsource(cls) - for cls in [_Type, _Object, _State, _Action, _LLT]) - - base = self._tool_context.sandbox_dir or self._get_log_dir() - ref_dir = os.path.join(base, "reference") - os.makedirs(ref_dir, exist_ok=True) - ref_path = os.path.join(ref_dir, "structs.py") - with open(ref_path, "w", encoding="utf-8") as f: - f.write(source) - - # Same backend-dependent agent-visible path mapping as - # _resolve_synthesis_paths. - if CFG.agent_sdk_use_local_sandbox: - return "./reference/structs.py" - if self._tool_context.sandbox_dir: - return "/sandbox/reference/structs.py" - return ref_path - def _base_sim_reference_paths(self) -> List[str]: """Agent-visible paths of the provisioned base-sim sources. @@ -3713,25 +2414,11 @@ def _base_sim_reference_paths(self) -> List[str]: type(self._base_env).__name__) return [] names = [os.path.basename(rel) for rel in src_files] - # Same backend-dependent path mapping as _write_structs_reference. + # Reference copies exist only in the local sandbox. if CFG.agent_sdk_use_local_sandbox: return [f"./reference/base_sim/{n}" for n in names] - if CFG.agent_sdk_use_docker_sandbox: - return [f"/sandbox/reference/base_sim/{n}" for n in names] return [] - @staticmethod - def _extract_obs_triples( - trajectories: List[LowLevelTrajectory], - ) -> List[Tuple[State, Action, State]]: - """Extract observed (s_t, action_t, s_{t+1}) triples.""" - triples: List[Tuple[State, Action, State]] = [] - for traj in trajectories: - for i in range(len(traj.actions)): - triples.append( - (traj.states[i], traj.actions[i], traj.states[i + 1])) - return triples - def _recreate_base_env(self) -> None: """Reconnect after a PyBullet physics-server crash.""" self._rebuild_base_env("PyBullet physics client crashed; recreating " @@ -3870,17 +2557,16 @@ def _fresh_model_env_scope( float]] = None) -> Iterator[None]: """Run the option model on a freshly constructed base env. - ``physical_overrides`` (the capture gate's physics-margin - rollouts) is applied to the fresh env ON TOP of the identified - params, so the rollout runs at a perturbed physics; the shared - session env is never touched. - - Installed as ``ToolContext.validation_env_scope`` so - ``submit_plan``'s capture-validation rollouts each sample - a fresh physics world. The shared ``_base_env``'s reset cannot - reconstruct state exactly (solver warm-start state, velocity - residuals, near-matching bodies skipped by the reconstruction diff - - the same mechanism measured in :func:`rollout_states`), so + ``physical_overrides`` (a physics-sweep point) is applied to the + fresh env ON TOP of the identified params, so the rollout runs at + a perturbed physics; the shared session env is never touched. + + Installed as ``ToolContext.validation_env_scope`` so the probe's + trials and sweep rollouts each sample a fresh physics world. The + shared ``_base_env``'s reset cannot reconstruct state exactly + (solver warm-start state, velocity residuals, near-matching bodies + skipped by the reconstruction diff - the same mechanism measured + in :func:`rollout_states`), so repeats on it are correlated with each other and systematically offset from the fresh env the real episode runs in (run_20260717_182321: a placement swept 20/20 on the shared env @@ -4088,50 +2774,6 @@ def step(state: State, action: Action) -> State: return make_stepper - def _build_synthesis_system_prompt(self) -> str: - """Compose the synthesis system prompt from the templates. - - Per-instance choices: the rule signature (flag-gated on - ``CFG.partially_observable``, which also swaps the env's - observation and the GT simulator module, so prompt and world - never disagree; under the flag only the recurrent 5-arg form is - shown), the optional PHYSICAL_PARAM_SPECS section (env parameter - menu), the scene-visualization hint, and the subclass extras. - """ - return learn_prompts.build_learn_system_prompt( - partially_observable=CFG.partially_observable, - residual_rule_signature=self._residual_rule_signature(), - scene_viz_hint=self._scene_viz_hint(), - physical_params_section=self._physical_params_prompt_section(), - extra_sections=self._extra_synthesis_system_prompt_sections(), - latent_extra_sections=self._extra_synthesis_latent_sections(), - workflow_extra=self._synthesis_workflow_extra(), - declared_params_only=CFG.agent_sim_learn_declared_params_only, - ) - - def _synthesis_workflow_extra(self) -> str: - """Extra text appended to the Workflow list. - - Subclasses with additional deliverables (e.g. invented - predicates) extend the workflow here so the numbered list stays - the single authoritative loop description. - """ - return "" - - @staticmethod - def _scene_viz_hint() -> str: - """The find-the-anchor-offset sentence. - - The probe is unconditional in synthesis sessions, so the hint - always names its staging + overlay surface. - """ - return ("use the `sim` probe in `run_python`: " - "`sim.reset(task_idx=..., " - "mods={...})` to stage a representative state from each " - "bucket and `sim.render(label, annotations=[...])` to " - "overlay, on one render, the recorded origin and the " - "positions where the effect did vs. did not fire") - def _physical_params_prompt_section(self) -> str: """Markdown for the optional PHYSICAL_PARAM_SPECS (system-ID) block. @@ -4147,18 +2789,4 @@ def _physical_params_prompt_section(self) -> str: getter = getattr(base_env, "get_physical_param_info", None) if callable(getter): info = getter() or {} - return learn_prompts.render_physical_params_section(info) - - def _residual_rule_signature(self) -> str: - """The ``def`` line used in the geometric-gate example. - - Matches the canonical rule signature the prompt advertises - (``CFG.partially_observable`` selects it) so the worked example - doesn't contradict it. - """ - if CFG.partially_observable: - # The example's body reads `state`; bind it so the recurrent - # form is a runnable rule, not a signature over a foreign name. - return ("def residual_rule(observation, latent, history, " - "updates, params):\n state = observation") - return "def residual_rule(state, updates, params):" + return render_physical_params_section(info) diff --git a/predicators/approaches/agent_sim_predicate_invention_approach.py b/predicators/approaches/agent_sim_predicate_invention_approach.py index 4e07a04cb3..4a242e1676 100644 --- a/predicators/approaches/agent_sim_predicate_invention_approach.py +++ b/predicators/approaches/agent_sim_predicate_invention_approach.py @@ -9,39 +9,29 @@ option model's abstraction function, and every other caller asking the approach for its current predicates. -Predicates persist across online learning cycles: ``predicates.py`` is -preserved at the sandbox root, and every version evaluated during -synthesis (plus a final snapshot of post-eval edits) is saved to -``predicates_versions/`` as ``cycle_XXX_vers_YYY_predicates.py``. - -Partial observability is not a separate approach: like every -sim-learning arm, the synthesis prompt follows -``CFG.partially_observable`` (see ``AgentSimLearningApproach``) - under -the flag the agent is taught the recurrent 5-arg rule signature -``rule(observation, latent, history, updates, params)`` and -``LATENT_INIT``, and this module appends the predicate-side latent -guidance (classifiers may take an optional ``latent`` kwarg, -auto-routed by ``Predicate.holds``). The latent *mechanics* (recurrent -LM fitting, the latent-threaded combined simulator riding +Predicates persist across rounds: ``predicates.py`` is preserved at the +sandbox root, and every version evaluated during synthesis (plus a final +snapshot of post-eval edits) is saved to ``predicates_versions/`` as +``cycle_XXX_vers_YYY_predicates.py``. + +Partial observability is not a separate approach. The latent mechanics +(recurrent LM fitting, the latent-threaded combined simulator riding ``State.latent`` so backtracking restores it per search node, ``LATENT_INIT`` loading and initial-latent seeding) live in ``AgentSimLearningApproach`` and activate automatically whenever the -loaded rules use the 5-arg signature, independent of the flag. - -Example command (partially observable):: +loaded rules use the 5-arg signature ``rule(observation, latent, +history, updates, params)``; classifiers may take an optional +``latent`` kwarg, auto-routed by ``Predicate.holds``. - python predicators/main.py --env pybullet_boil \ - --approach agent_sim_predicate_invention --seed 0 \ - --num_train_tasks 10 --num_test_tasks 5 \ - --partially_observable True \ - --num_online_learning_cycles 2 --explorer agent_model_free +The continual arms (``AgentContinualApproach`` and its variants) build +on this class; launch them through +``scripts/configs/empiric/benchmark.yaml``. """ import logging import os from typing import Any, Dict, FrozenSet, List, Set, Tuple -from predicators.agent_sdk import learn_prompts from predicators.agent_sdk.tools import _SnapshotTarget, \ finalize_versioned_snapshot, make_predicate_quality_loader from predicators.approaches.agent_sim_learning_approach import \ @@ -162,61 +152,6 @@ def _build_write_snapshot_targets( )) return targets - def _extra_synthesis_message(self, extra_paths: Dict[str, str]) -> str: - message = learn_prompts.render_predicate_invention_message( - extra_paths["predicates_file_for_agent"], - self._format_goal_nl_block()) - return message + self._chained_extra_message(extra_paths) - - def _chained_extra_message(self, extra_paths: Dict[str, str]) -> str: - """The base class's extra message (the partial-observability note under - ``CFG.partially_observable``), separated for appending.""" - base = super()._extra_synthesis_message(extra_paths) - return "\n\n" + base if base else "" - - def _format_goal_nl_block(self) -> str: - """Render the deduped natural-language goals for the train tasks. - - Returns an empty string only if every task is missing a - ``goal_nl``, but ``__init__`` asserts they're present, so in - practice this always returns a non-empty block. - """ - seen: List[str] = [] - for task in self._train_tasks: - nl = task.goal_nl - if nl and nl not in seen: - seen.append(nl) - if not seen: - return "" - if len(seen) == 1: - return f"Goal (natural language): {seen[0]}\n\n" - bullets = "\n".join(f" - {g}" for g in seen) - return f"Goals across train tasks (natural language):\n{bullets}\n\n" - - def _synthesis_workflow_extra(self) -> str: - # The base workflow's validation step depends on invented - # predicates: sketches can only reference predicates that exist. - return learn_prompts.render_predicate_workflow_extra() - - def _extra_synthesis_system_prompt_sections(self) -> List[str]: - # The scene workbench is the sim probe inside run_python (the - # probe is unconditional in synthesis sessions). - workbench = ("the `sim` probe in `run_python` as scene workbench " - "(`sim.reset(task_idx=..., mods={...})` to stage " - "states, `sim.render(label, annotations=[...])` " - "to render with overlays)") - sections = super()._extra_synthesis_system_prompt_sections() - sections.append( - learn_prompts.render_predicate_invention_section(workbench)) - return sections - - def _extra_synthesis_latent_sections(self) -> List[str]: - # The predicate-side latent guidance belongs to invention arms - # only and follows the simulator-side tutorial it refers to. - sections = super()._extra_synthesis_latent_sections() - sections.append(learn_prompts.render_predicate_latent_section()) - return sections - def _post_synthesis_loading( self, extra_paths: Dict[str, str], diff --git a/predicators/approaches/continual_play_mixin.py b/predicators/approaches/continual_play_mixin.py index 8fd77bca64..0706ca1610 100644 --- a/predicators/approaches/continual_play_mixin.py +++ b/predicators/approaches/continual_play_mixin.py @@ -28,17 +28,15 @@ A reset starts a new episode on the same level without ending the round. A level can span several rounds if the agent returns before settling it. -Why a mixin. The arms' learning and session machinery live in the -phased approach classes (``AgentModelFreeApproach`` and its +Why a mixin. The arms' model and session machinery live in the +approach classes below them (``AgentModelFreeApproach`` and its ``AgentSimPredicateInventionApproach`` descendant), where the simulator -synthesis, the parameter fit, predicate invention, the sandbox and the +loading, the parameter belief, predicate invention, the sandbox and the session managers are implemented. An arm keeps that class as its base -and mixes this loop in front of it, the way ``AgentSessionMixin`` and -``SamplerLearningMixin`` add their concerns; the phased loop's own entry -points (``_solve``, the explorers, the learning sessions) are simply -unused under the protocol. The mixin has no base class of its own, so -there is no diamond, and what it needs from its host is declared below -as the host contract. +and mixes this loop in front of it, the way ``AgentSessionMixin`` adds +its concerns. The mixin has no base class of its own, so there is no +diamond, and what it needs from its host is declared below as the host +contract. The harness never chooses for the agent: whether to act, reset, model or give up is decided inside the conversation; the loop only services @@ -323,8 +321,8 @@ def _run_conversation_round(self, session: ProtocolSession) -> PlayState: kind = "continue" query = self._build_query(session, kind) round_number = self._rounds_played + 1 - # No per-round clock: the run's wall-clock cap is the only clock. - ctx.begin_attempt(round_number, 0.0) + # No per-round budget: the run's wall-clock cap is the only one. + ctx.begin_attempt() self._round_in_flight = True self.save(session.level_index) entries_before = len(session.index_entries()) @@ -333,7 +331,6 @@ def _run_conversation_round(self, session: ProtocolSession) -> PlayState: responses = self._query_agent_sync(query, kind=SESSION_KIND) finally: ctx.attempt_start = None - ctx.attempt_deadline = None self._round_in_flight = False self._agent_session.resume_session_id = None self._after_round(session, state) @@ -542,8 +539,6 @@ def _episodes_to_trajectories(self, episodes: Sequence[Dict[str, Any]], self, "_current_simulator_version", None), _source_predicates_version=getattr( self, "_current_predicates_version", None), - _source_samplers_version=getattr( - self, "_current_samplers_version", None), _env_reward=ep.get("reward"), _env_terminated=ep.get("terminated"), )) diff --git a/predicators/approaches/gnn_dynamics_approach.py b/predicators/approaches/gnn_dynamics_approach.py deleted file mode 100644 index 84cc24bd4e..0000000000 --- a/predicators/approaches/gnn_dynamics_approach.py +++ /dev/null @@ -1,459 +0,0 @@ -"""A neural dynamics baseline (paper arm C5): a GNN transition model over the -object-centric state features, trained on option-level transitions from the -same interaction data the agent arms collect, and planned against by random -shooting through the learned model. - -The model predicts, for a state and a ground option, the change of every -object feature and the number of low-level steps the option takes. The -last ``gnn_dynamics_history_len`` pre-option states ride along as node -features so a mechanism that is not visible in one observation (a cure -that started a few options ago) can be inferred from the recent past. -Planning samples random applicable options, rolls them through the -model, and accepts the first sequence whose predicted final state -satisfies the goal under the env's (oracle) predicates; by default the -plan is re-shot from the observed state every time an option terminates. -""" - -import logging -import time -from typing import Any, Callable, Dict, List, Optional, Sequence, Set, Tuple - -import dill as pkl -import numpy as np -import torch -import torch.nn -import torch.optim -from gym.spaces import Box -from torch.utils.data import DataLoader - -from predicators import utils -from predicators.approaches import ApproachFailure, ApproachTimeout, \ - BaseApproach -from predicators.explorers import create_explorer -from predicators.gnn.gnn import EncodeProcessDecode, setup_graph_net -from predicators.gnn.gnn_utils import GraphDictDataset, compute_normalizers, \ - get_single_model_prediction, graph_batch_collate, normalize_graph, \ - train_model -from predicators.nsrt_learning.segmentation import segment_trajectory -from predicators.settings import CFG -from predicators.structs import Action, Dataset, DummyOption, \ - InteractionRequest, InteractionResult, LowLevelTrajectory, \ - ParameterizedOption, Predicate, State, Task, Type, _Option - -# One training example: the recent pre-option states (oldest first), the -# pre-option state, the ground option, the post-option state, and the -# option's low-level step count. -_Example = Tuple[List[State], State, _Option, State, int] - - -class GNNDynamicsShootingApproach(BaseApproach): - """GNN transition model over object features + shooting planner.""" - - def __init__(self, initial_predicates: Set[Predicate], - initial_options: Set[ParameterizedOption], types: Set[Type], - action_space: Box, train_tasks: List[Task]) -> None: - super().__init__(initial_predicates, initial_options, types, - action_space, train_tasks) - self._sorted_options = sorted(self._initial_options, - key=lambda o: o.name) - self._trajectories: List[LowLevelTrajectory] = [] - self._online_learning_cycle = 0 - self._requests_train_task_idxs: List[int] = [] - self._gnn: Optional[EncodeProcessDecode] = None - self._data_exemplar: Optional[Tuple[Dict, Dict]] = None - self._input_normalizers: Optional[Dict] = None - self._target_normalizers: Optional[Dict] = None - # Node feature layout: type one-hot, current features, history - # features (per lag), option-argument slot one-hot. - self._type_to_index: Dict[str, int] = {} - self._feat_to_index: Dict[str, int] = {} - self._max_option_objects = 0 - self._max_option_params = 0 - self._mse_loss = torch.nn.MSELoss() - - @classmethod - def get_name(cls) -> str: - return "gnn_dynamics_shooting" - - @property - def is_learning_based(self) -> bool: - return True - - # ── Data ───────────────────────────────────────────────────── - - def _generate_examples(self) -> List[_Example]: - """Option-level transitions from every stored trajectory.""" - examples: List[_Example] = [] - history_len = CFG.gnn_dynamics_history_len - for traj in self._trajectories: - if not traj.actions or not traj.actions[0].has_option(): - continue - segments = segment_trajectory(traj, self._initial_predicates) - pre_states: List[State] = [] - for segment in segments: - state = segment.states[0] - history = pre_states[-history_len:] if history_len else [] - examples.append((list(history), state, segment.get_option(), - segment.states[-1], len(segment.actions))) - pre_states.append(state) - return examples - - def _setup_fields(self, examples: Sequence[_Example]) -> None: - types: Set[str] = set() - feats: Set[str] = set() - max_objects = 0 - max_params = 0 - for _, state, option, _, _ in examples: - for obj in state: - types.add(obj.type.name) - feats.update(obj.type.feature_names) - max_objects = max(max_objects, len(option.objects)) - max_params = max(max_params, option.params.shape[0]) - # Every option's argument count and parameter box, not just the - # ones seen: a test-time sample of an unseen option must graphify. - for param_opt in self._sorted_options: - max_objects = max(max_objects, len(param_opt.types)) - max_params = max(max_params, param_opt.params_space.shape[0]) - self._type_to_index = {t: i for i, t in enumerate(sorted(types))} - self._feat_to_index = {f: i for i, f in enumerate(sorted(feats))} - self._max_option_objects = max_objects - self._max_option_params = max_params - - def _feature_row(self, state: State, obj: Any) -> np.ndarray: - row = np.zeros(len(self._feat_to_index)) - for feat, val in zip(obj.type.feature_names, state[obj]): - row[self._feat_to_index[feat]] = val - return row - - def _graphify_input(self, history: Sequence[State], state: State, - option: _Option) -> Tuple[Dict, Dict[Any, int]]: - objects = list(state) - object_to_node = {o: i for i, o in enumerate(objects)} - n_obj = len(objects) - n_types = len(self._type_to_index) - n_feats = len(self._feat_to_index) - history_len = CFG.gnn_dynamics_history_len - n_node_feats = (n_types + n_feats + history_len * n_feats + - self._max_option_objects) - nodes = np.zeros((n_obj, n_node_feats)) - arg_slots = {o: i for i, o in enumerate(option.objects)} - for obj in objects: - i = object_to_node[obj] - nodes[i, self._type_to_index[obj.type.name]] = 1 - current = self._feature_row(state, obj) - nodes[i, n_types:n_types + n_feats] = current - # History as differences to the current features, most - # recent lag first; missing lags stay zero (no change). - for lag in range(history_len): - if lag < len(history): - past = history[-1 - lag] - if obj in past.data: - past_row = self._feature_row(past, obj) - start = n_types + n_feats * (1 + lag) - nodes[i, start:start + n_feats] = past_row - current - if obj in arg_slots: - nodes[i, n_types + n_feats * (1 + history_len) + - arg_slots[obj]] = 1 - # Fully connected (no self loops) so effects can flow between - # objects; edge features: constant, both endpoints are arguments. - senders, receivers, edges = [], [], [] - for s in range(n_obj): - for r in range(n_obj): - if s == r: - continue - senders.append(s) - receivers.append(r) - both = float(objects[s] in arg_slots - and objects[r] in arg_slots) - edges.append([1.0, both]) - n_edge = len(edges) - onehot = np.zeros(len(self._sorted_options)) - onehot[self._sorted_options.index(option.parent)] = 1 - params = np.zeros(self._max_option_params) - params[:option.params.shape[0]] = option.params - graph = { - "n_node": np.array(n_obj), - "nodes": nodes, - "n_edge": np.reshape(n_edge, [1]).astype(np.int64), - "edges": np.reshape(edges, [n_edge, 2]), - "senders": np.reshape(senders, [n_edge]).astype(np.int64), - "receivers": np.reshape(receivers, [n_edge]).astype(np.int64), - "globals": np.r_[onehot, params], - } - return graph, object_to_node - - def _graphify_target(self, state: State, next_state: State, - num_actions: int, graph_input: Dict, - object_to_node: Dict[Any, int]) -> Dict: - n_obj = len(object_to_node) - nodes = np.zeros((n_obj, len(self._feat_to_index))) - for obj, i in object_to_node.items(): - nodes[i] = (self._feature_row(next_state, obj) - - self._feature_row(state, obj)) - return { - "n_node": graph_input["n_node"], - "nodes": nodes, - "n_edge": graph_input["n_edge"], - "edges": np.zeros((int(graph_input["n_edge"][0]), 1)), - "senders": graph_input["senders"], - "receivers": graph_input["receivers"], - "globals": np.array([num_actions / float(max(CFG.horizon, 1))]), - } - - # ── Learning ───────────────────────────────────────────────── - - def learn_from_offline_dataset(self, dataset: Dataset) -> None: - self._trajectories = list(dataset.trajectories) - self._learn_model() - self._save(None) - - def get_interaction_requests(self) -> List[InteractionRequest]: - explorer = create_explorer(CFG.explorer, self._initial_predicates, - self._initial_options, self._types, - self._action_space, self._train_tasks) - requests: List[InteractionRequest] = [] - self._requests_train_task_idxs = [] - for _ in range(CFG.online_nsrt_learning_requests_per_cycle): - task_idx = int(self._rng.choice(len(self._train_tasks))) - policy, termination_fn = explorer.get_exploration_strategy( - task_idx, CFG.timeout) - requests.append( - InteractionRequest(train_task_idx=task_idx, - act_policy=policy, - query_policy=lambda s: None, - termination_function=termination_fn)) - self._requests_train_task_idxs.append(task_idx) - return requests - - def restore_interaction_requests(self, train_task_idxs: List[int]) -> None: - self._requests_train_task_idxs = list(train_task_idxs) - - def learn_from_interaction_results( - self, results: Sequence[InteractionResult]) -> None: - assert len(results) == len(self._requests_train_task_idxs) - for task_idx, result in zip(self._requests_train_task_idxs, results): - self._trajectories.append( - LowLevelTrajectory(result.states, - result.actions, - _is_demo=False, - _train_task_idx=task_idx)) - self._learn_model() - self._save(self._online_learning_cycle) - self._online_learning_cycle += 1 - - def _learn_model(self) -> None: - examples = self._generate_examples() - if not examples: - logging.warning("GNN dynamics: no option-level transitions yet; " - "keeping the previous model.") - return - self._setup_fields(examples) - graph_inputs, graph_targets = [], [] - for history, state, option, next_state, num_actions in examples: - graph_input, object_to_node = self._graphify_input( - history, state, option) - graph_inputs.append(graph_input) - graph_targets.append( - self._graphify_target(state, next_state, num_actions, - graph_input, object_to_node)) - self._data_exemplar = (graph_inputs[0], graph_targets[0]) - self._gnn = setup_graph_net(GraphDictDataset([graph_inputs[0]], - [graph_targets[0]]), - num_steps=CFG.gnn_num_message_passing, - layer_size=CFG.gnn_layer_size) - if CFG.gnn_do_normalization: - self._input_normalizers = compute_normalizers(graph_inputs) - self._target_normalizers = compute_normalizers(graph_targets) - graph_inputs = [ - normalize_graph(g, self._input_normalizers) - for g in graph_inputs - ] - graph_targets = [ - normalize_graph(g, self._target_normalizers) - for g in graph_targets - ] - num_validation = (max(1, int(len(examples) * 0.1)) - if CFG.gnn_use_validation_set else 0) - train_set = GraphDictDataset(graph_inputs[num_validation:], - graph_targets[num_validation:]) - val_set = GraphDictDataset(graph_inputs[:num_validation], - graph_targets[:num_validation]) - dataloaders = { - "train": - DataLoader(train_set, - batch_size=CFG.gnn_batch_size, - shuffle=False, - num_workers=0, - collate_fn=graph_batch_collate), - "val": - DataLoader(val_set, - batch_size=CFG.gnn_batch_size, - shuffle=False, - num_workers=0, - collate_fn=graph_batch_collate), - } - optimizer = torch.optim.Adam(self._gnn.parameters(), - lr=CFG.gnn_learning_rate, - weight_decay=CFG.gnn_weight_decay) - logging.info( - "Training GNN dynamics on %d option transitions from " - "%d trajectories.", len(examples), len(self._trajectories)) - best = train_model(self._gnn, - dataloaders, - optimizer=optimizer, - criterion=self._mse_loss, - global_criterion=self._mse_loss, - num_epochs=CFG.gnn_num_epochs, - do_validation=CFG.gnn_use_validation_set) - self._gnn.load_state_dict(best) - - # ── Checkpointing ──────────────────────────────────────────── - - def _save(self, online_learning_cycle: Optional[int]) -> None: - info = { - "trajectories": - self._trajectories, - "online_learning_cycle": - self._online_learning_cycle, - "exemplar": - self._data_exemplar, - "state_dict": - (self._gnn.state_dict() if self._gnn is not None else None), - "type_to_index": - self._type_to_index, - "feat_to_index": - self._feat_to_index, - "max_option_objects": - self._max_option_objects, - "max_option_params": - self._max_option_params, - "input_normalizers": - self._input_normalizers, - "target_normalizers": - self._target_normalizers, - } - path = (f"{utils.get_approach_save_path_str()}_" - f"{online_learning_cycle}.gnn") - with open(path, "wb") as f: - pkl.dump(info, f) - - def load(self, online_learning_cycle: Optional[int]) -> None: - path = (f"{utils.get_approach_load_path_str()}_" - f"{online_learning_cycle}.gnn") - with open(path, "rb") as f: - info = pkl.load(f) - self._trajectories = info["trajectories"] - self._online_learning_cycle = info["online_learning_cycle"] - self._data_exemplar = info["exemplar"] - self._type_to_index = info["type_to_index"] - self._feat_to_index = info["feat_to_index"] - self._max_option_objects = info["max_option_objects"] - self._max_option_params = info["max_option_params"] - self._input_normalizers = info["input_normalizers"] - self._target_normalizers = info["target_normalizers"] - if info["state_dict"] is not None and self._data_exemplar is not None: - example_input, example_target = self._data_exemplar - self._gnn = setup_graph_net(GraphDictDataset([example_input], - [example_target]), - num_steps=CFG.gnn_num_message_passing, - layer_size=CFG.gnn_layer_size) - self._gnn.load_state_dict(info["state_dict"]) - - # ── Prediction ─────────────────────────────────────────────── - - def predict_next_state(self, history: Sequence[State], state: State, - option: _Option) -> Tuple[State, int]: - """The model's post-option state and low-level step count.""" - assert self._gnn is not None, "Learn a model before predicting." - graph_input, object_to_node = self._graphify_input( - history, state, option) - if CFG.gnn_do_normalization: - assert self._input_normalizers is not None - graph_input = normalize_graph(graph_input, self._input_normalizers) - out = get_single_model_prediction(self._gnn, graph_input) - if CFG.gnn_do_normalization: - assert self._target_normalizers is not None - out = normalize_graph(out, self._target_normalizers, invert=True) - next_state = state.copy() - for obj, i in object_to_node.items(): - delta = out["nodes"][i] - for feat in obj.type.feature_names: - next_state.set( - obj, feat, - state.get(obj, feat) + - float(delta[self._feat_to_index[feat]])) - num_actions = max( - 1, int(round(float(out["globals"][0]) * max(CFG.horizon, 1)))) - return next_state, num_actions - - # ── Planning ───────────────────────────────────────────────── - - def _shoot(self, task: Task, init_state: State, history: Sequence[State], - deadline: float) -> Optional[List[_Option]]: - """Random shooting through the learned model. - - Returns the first sampled option sequence whose predicted final - state satisfies the goal, or None when the tries or the deadline - run out. - """ - for _ in range(CFG.gnn_dynamics_shooting_max_tries): - if time.perf_counter() > deadline: - return None - state = init_state - past = list(history) - plan: List[_Option] = [] - num_actions = 0 - for _ in range(CFG.gnn_dynamics_max_plan_length): - if task.goal_holds(state): - return plan - option = utils.sample_applicable_option( - self._sorted_options, state, self._rng) - if option is None: - break - next_state, k = self.predict_next_state(past, state, option) - plan.append(option) - past.append(state) - state = next_state - num_actions += k - if num_actions > CFG.horizon: - break - if task.goal_holds(state) and plan: - return plan - return None - - def _solve(self, task: Task, timeout: int) -> Callable[[State], Action]: - if self._gnn is None: - raise ApproachFailure("GNN dynamics model has not been learned.") - deadline = time.perf_counter() + timeout - observed: List[State] = [] - plan: List[_Option] = [] - cur_option: _Option = DummyOption - replan = CFG.gnn_dynamics_replan_every_option - - def _next_option(state: State) -> _Option: - nonlocal plan - if replan or not plan: - shot = self._shoot(task, state, observed, deadline) - if shot is None: - if time.perf_counter() > deadline: - raise ApproachTimeout( - "GNN dynamics shooting timed out.") - raise ApproachFailure( - "GNN dynamics shooting found no plan that reaches " - "the goal under the learned model.") - plan = list(shot) - option = plan.pop(0) - if not option.initiable(state): - raise ApproachFailure( - f"Planned option {option.name} is not initiable in the " - "observed state.") - return option - - def _policy(state: State) -> Action: - nonlocal cur_option - if cur_option is DummyOption or cur_option.terminal(state): - cur_option = _next_option(state) - observed.append(state) - return cur_option.policy(state) - - return _policy diff --git a/predicators/approaches/sampler_learning_mixin.py b/predicators/approaches/sampler_learning_mixin.py deleted file mode 100644 index 023ba2c429..0000000000 --- a/predicators/approaches/sampler_learning_mixin.py +++ /dev/null @@ -1,431 +0,0 @@ -"""Parameterized (per-skill) sampler learning for the sim-learning approach. - -A parameterized sampler is keyed by option name and authored once; the -ground level of the sampler hierarchy (per-step ``GroundSampler`` from a -sketch ``~ [widths]`` region annotation) is not learned and lives in -``bilevel_sketch``, overriding the parameterized sampler per step. - -Samplers are a first-class artifact of the base sim-learning approach -(gated by ``CFG.agent_sim_learn_parameterized_samplers``), not a subclass -extension like predicates — so they are woven into -``AgentSimLearningApproach._synthesize_with_agent`` and -``_learn_simulator`` directly rather than via the ``_extra_synthesis_*`` -hooks, which keeps them independent of the predicate subclass's -(non-super-calling) hook overrides. When a sim-synthesis session runs -(``oracle_sim_program=False``) the sampler tool/snapshot/message ride -along in it; when none runs (``oracle_sim_program=True``) they get a -dedicated session via :meth:`_synthesize_samplers_standalone`. - -This mixin owns everything sampler-specific: mode resolution (learn vs. -ground truth), sandbox path bindings, the synthesis tool/snapshot/message -builders, loading ``LEARNED_SAMPLERS`` from file, and the standalone -synthesis session. The host approach keeps only the call sites. -""" -import logging -import os -from typing import TYPE_CHECKING, Any, Dict, List, Optional, Set, Tuple, cast - -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error -from predicators.agent_sdk.tools import _SnapshotTarget, \ - create_synthesis_tools, finalize_versioned_snapshot, make_sampler_loader -from predicators.agent_sdk.tools.digests import render_options_digest -from predicators.code_sim_learning.fit_space import ParamSpec -from predicators.ground_truth_models import get_gt_samplers -from predicators.settings import CFG -from predicators.structs import Action, LowLevelTrajectory, \ - ParameterizedOption, ParameterizedSampler, Predicate, State, Task, Type - -if TYPE_CHECKING: - from predicators.agent_sdk.synthesis_backend import SynthesisBackend - from predicators.agent_sdk.tools import ToolContext - -logger = logging.getLogger(__name__) - - -class SamplerLearningMixin: - """Per-skill sampler synthesis, loading, and oracle installation. - - Mixed into :class:`AgentSimLearningApproach`. Holds the - sampler-learning state (``_do_synthesize_samplers``, - ``_current_samplers_version``) — the host ``__init__`` must call - :meth:`_init_sampler_learning_state`. - """ - - # ── Host-class contract ───────────────────────────────────── - # Everything below is provided by the host approach (its - # AgentSessionMixin / BaseApproach ancestry or the host class - # itself). Declared under TYPE_CHECKING only, so these never - # shadow the real implementations in the MRO at runtime. - if TYPE_CHECKING: - _tool_context: "ToolContext" - _train_tasks: List[Task] - _types: Set[Type] - _fitted_params: Dict[str, float] - _learning_mode: bool - _synthesized_samplers: Dict[str, ParameterizedSampler] - - def _learning_cycle_index(self) -> int: - raise NotImplementedError - - def _get_log_dir(self) -> str: - raise NotImplementedError - - def _get_all_predicates(self) -> Set[Predicate]: - raise NotImplementedError - - def _get_all_options(self) -> Set[ParameterizedOption]: - raise NotImplementedError - - def _get_synthesis_tool_names(self) -> Optional[List[str]]: - raise NotImplementedError - - def _build_synthesis_exec_ns( - self, - trajectories: List[LowLevelTrajectory]) -> Dict[str, Any]: - raise NotImplementedError - - def _query_agent_sync(self, message: str, - **query_kwargs: Any) -> List[Dict[str, Any]]: - raise NotImplementedError - - def _ensure_agent_session(self) -> None: - raise NotImplementedError - - def _close_agent_session(self) -> None: - raise NotImplementedError - - @staticmethod - def _build_synthesis_session_hooks( - targets: List[_SnapshotTarget], - sandbox_dir: str) -> Dict[str, list]: - raise NotImplementedError - - @staticmethod - def _format_predicate_signatures(predicates: Set[Predicate]) -> str: - raise NotImplementedError - - def _init_sampler_learning_state(self) -> None: - """Initialize sampler state; called from the host ``__init__``.""" - # Snapshot tag of the most recent samplers file committed by the - # synthesis agent — used to stamp newly collected online - # trajectories with their source-version provenance. - self._current_samplers_version: Optional[str] = None - # Whether this run learns samplers (vs. using ground-truth ones). - # Refined per cycle in _learn_simulator once GT availability is - # known; this default is what the synthesis-session tool surface - # reads. - self._do_synthesize_samplers: bool = ( - CFG.agent_sim_learn_parameterized_samplers - and not CFG.agent_sim_learn_oracle_samplers) - - @staticmethod - def _samplers_enabled() -> bool: - """Whether per-skill samplers are used at all this run.""" - return CFG.agent_sim_learn_parameterized_samplers - - def _maybe_install_oracle_samplers(self) -> None: - """Resolve sampler mode for this cycle and install GT ones if used. - - Sets ``self._do_synthesize_samplers`` (learn vs. use ground - truth). When ``agent_sim_learn_oracle_samplers`` is on and the - env provides ground-truth samplers, installs them and skips - synthesis; if none exist, warns and falls back to synthesis. - """ - gt_samplers = None - if self._samplers_enabled() and CFG.agent_sim_learn_oracle_samplers: - gt_samplers = get_gt_samplers(CFG.env) - if gt_samplers: - self._synthesized_samplers = dict(gt_samplers) - self._current_samplers_version = "oracle" - logger.info("Using %d ground-truth sampler(s): %s", - len(gt_samplers), ", ".join(sorted(gt_samplers))) - else: - logger.warning( - "agent_sim_learn_oracle_samplers=True but no ground-truth " - "samplers for env %s; falling back to synthesis.", CFG.env) - self._do_synthesize_samplers = (self._samplers_enabled() - and not gt_samplers) - - def _sampler_paths(self, base: str) -> Dict[str, str]: - """Sandbox path bindings for samplers.py (host + agent-visible).""" - samplers_file = os.path.join(base, "samplers.py") - samplers_versions_dir = os.path.join(base, "samplers_versions") - if CFG.agent_sdk_use_local_sandbox: - samplers_file_for_agent = "./samplers.py" - elif self._tool_context.sandbox_dir: - samplers_file_for_agent = "/sandbox/samplers.py" - else: - samplers_file_for_agent = samplers_file - return { - "samplers_file": samplers_file, - "samplers_versions_dir": samplers_versions_dir, - "samplers_file_for_agent": samplers_file_for_agent, - } - - def _install_sampler_surface(self, paths: Dict[str, str]) -> None: - """Register the ``sim.samplers()`` loader for a synthesis session.""" - self._tool_context.probe_artifact_loaders["samplers"] = \ - make_sampler_loader( - samplers_file=paths["samplers_file"], - samplers_versions_dir=paths["samplers_versions_dir"], - approach=self, - cycle_index_provider=self._learning_cycle_index, - ) - - def _sampler_snapshot_target(self, paths: Dict[str, - str]) -> _SnapshotTarget: - """Snapshot target that versions samplers.py on every Write/Edit.""" - return _SnapshotTarget( - live_file=paths["samplers_file"], - versions_dir=paths["samplers_versions_dir"], - artifact_name="samplers", - cycle_index_provider=self._learning_cycle_index, - ) - - def _sampler_synthesis_message(self, paths: Dict[str, str]) -> str: - """Instructions appended to the agent's first synthesis message.""" - path = paths["samplers_file_for_agent"] - # The ground channel exists only when its flag is on; do not - # describe it to sessions that cannot use it. - ground_note = "" - if CFG.agent_bilevel_ground_samplers: - ground_note = ( - "\nSamplers here are the reusable cross-task prior: " - "refinement uses yours on every draw of that option, in " - "every sketch and every task. A sketch step that carries " - "its own `~` ground-sampler annotation (a `~ [widths]` " - "window or `~ name` from ground_samplers.py) bypasses " - "yours for that step (precedence: ground sampler > " - "parameterized sampler > uniform).") - return f"""\ -## Per-Skill Sampler Synthesis - -Backtracking refinement draws each option's continuous parameters \ -*uniformly* from its params box by default. When a sketch step's subgoal \ -pins the parameters into a tiny region (e.g. a placement that must land \ -within a few cm of an exact point and at a specific orientation), uniform \ -sampling almost never hits it and refinement exhausts its budget. Fix this \ -by writing per-skill samplers to `{path}` as a dict \ -`LEARNED_SAMPLERS = {{"OptionName": sampler_fn, ...}}` keyed by option name. - -Each sampler has signature \ -`fn(state, subgoal_atoms, rng, objects) -> params` (the same signature as \ -the env's NSRT samplers) where: -- `state` is the current `State` (read object features with `state.get(obj, "feat")`), -- `subgoal_atoms` is the set of `GroundAtom`s the step must establish — \ -read the target relation here (e.g. an `InFront`/at-target atom names the \ -two objects whose geometry the placement must satisfy) and compute the \ -parameters that achieve it. At steps with NO subgoal annotation this set \ -is EMPTY — the sampler must not crash on `set()`; fall back to a default \ -or uniform draw, -- `rng` is a `numpy` `Generator` (use it for small jitter so retries differ), -- `objects` is the list of typed objects bound to this option call. -Return a `float32` array whose length matches the option's params box \ -(see the Options digest in your prompt for the dimension and ranges); \ -refinement clips it to that box, so stay within the ranges. -{ground_note} - -Aim the parameters at the subgoal geometrically (then add a little `rng` \ -jitter); do NOT just return uniform draws. Read the option signatures \ -from the Options digest in your prompt and the predicate classifiers \ -(for the subgoal geometry) with the predicate listing above. - -Workflow: write `{path}`, call `sim.samplers()` (snapshots + installs \ -them and sanity-checks shape/box), then call `sim.refine` \ -with a sketch using those options — the samples-to-refine count should \ -drop sharply versus uniform. Iterate with `Edit` and re-run. Every \ -successful Write/Edit of `{path}` is snapshotted to `samplers_versions/` \ -as `cycle_XXX_vers_YYY_samplers.py`.""" - - def _finalize_and_load_samplers(self, paths: Dict[str, str]) -> None: - """Snapshot the final samplers.py and load it into approach state.""" - tag = finalize_versioned_snapshot( - paths["samplers_file"], - paths["samplers_versions_dir"], - cycle_idx=self._learning_cycle_index(), - artifact_name="samplers", - ) - if tag is not None: - self._current_samplers_version = tag - logger.info("Final samplers snapshot: %s", tag) - loaded = self._load_samplers_from_module_file(paths["samplers_file"]) - self._synthesized_samplers = loaded - logger.info("Loaded %d per-skill sampler(s) from %s.", len(loaded), - paths["samplers_file"]) - for name in sorted(loaded): - logger.info(" sampler: %s", name) - - def _load_samplers_from_module_file( - self, path: str) -> Dict[str, ParameterizedSampler]: - """Load LEARNED_SAMPLERS from ``path``; validate each entry. - - Mirrors ``_load_predicates_from_module_file``. Returns an empty - dict on missing file or exec failure (samplers are optional). - Validation (unknown option names, non-callables) is shared with - ``sim.samplers()`` via ``load_learned_samplers``. - """ - # pylint: disable=import-outside-toplevel - from predicators.agent_sdk.proposal_exec import build_exec_context, \ - load_learned_samplers - from predicators.agent_sdk.tools import _ParamsView - - # pylint: enable=import-outside-toplevel - # ParamSpec is imported at module scope (used by exec'd samplers - # that close over learned params, mirroring the predicate loader). - - if not os.path.isfile(path): - logger.info("No samplers file at %s; sampler set is empty.", path) - return {} - - with open(path, "r", encoding="utf-8") as f: - code = f.read() - - ctx = build_exec_context(types=self._types, - predicates=self._get_all_predicates(), - options=self._get_all_options(), - extra_context={ - "params": - _ParamsView(self._fitted_params), - "ParamSpec": ParamSpec, - }) - - option_names = {o.name for o in self._get_all_options()} - valid, warnings, err = load_learned_samplers(code, ctx, option_names) - if err is not None: - logger.warning("Failed to load %s:\n%s", path, err) - return {} - for warning in warnings: - logger.warning("%s: %s", path, warning) - return valid - - def _synthesize_samplers_standalone( - self, trajectories: List[LowLevelTrajectory], - base_pred_triples: List[Tuple[State, Action, State]], - inferred_hint: Dict[str, List[str]]) -> None: - """Run a dedicated sampler-synthesis session. - - Used when oracle_sim_program short-circuits the sim-synthesis - session, so samplers still get learned. Reuses that session's - sandbox/snapshot/tool machinery. Called from _learn_simulator - after the option model is built, so the session's probe has a - working simulator. - """ - if CFG.agent_sdk_use_local_sandbox: - sandbox_dir: Optional[str] = os.path.abspath( - os.path.join(self._get_log_dir(), "sandbox")) - else: - sandbox_dir = self._tool_context.sandbox_dir - base = sandbox_dir or self._get_log_dir() - - if CFG.agent_sdk_use_local_sandbox: - sandbox_dir_for_agent: Optional[str] = "." - elif sandbox_dir: - sandbox_dir_for_agent = "/sandbox" - else: - sandbox_dir_for_agent = None - - paths = self._sampler_paths(base) - simulator_file = os.path.join(base, "simulator.py") - versions_dir = os.path.join(base, "simulator_versions") - - # Same namespace the main synthesis session gets (trajectories, - # train_tasks, is_goal_state, describe_trajectory, np, ParamSpec, - # evaluate_trajectory when the env defines evaluators). - exec_ns: Dict[str, Any] = self._build_synthesis_exec_ns(trajectories) - # The probe's `sim.refine` gives the agent the samples-to-refine - # feedback signal; the sampler tool installs + sanity-checks the - # samplers. - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.belief_probe import _check_time_budget - toolkit = create_synthesis_tools( - exec_ns, - base_pred_triples, - inferred_hint, - simulator_file=simulator_file, - versions_dir=versions_dir, - # The host class (AgentSimLearningApproach) provides the - # full backend surface; the mixin's own type covers only - # the sampler slice. - approach=cast("SynthesisBackend", self), - sandbox_dir=base, - sandbox_dir_for_agent=sandbox_dir_for_agent, - cycle_index_provider=self._learning_cycle_index, - budget_check=lambda: _check_time_budget(self._tool_context), - ) - tools = list(toolkit.tools) - self._install_sampler_surface(paths) - # Use the same declared surface as the mixin will assert against - # (_get_synthesis_tool_names already includes the sampler tool since - # _do_synthesize_samplers is True here). The rule-fitting surface is - # exposed but irrelevant — the message steers the agent to samplers. - declared = set(self._get_synthesis_tool_names() or ()) - self._tool_context.extra_mcp_tools = [ - t for t in tools if getattr(t, "name", "") in declared - ] - # The probe here runs the DEPLOYED belief model (no candidate - # provider: ctx.option_model already wraps the oracle sim - # program), which is exactly what samplers must speed up. The - # fit runner still targets simulator.py for consistency. - self._tool_context.probe_fit_provider = toolkit.fit_runner - self._tool_context.probe_validation_provider = toolkit.validation_runner - self._tool_context.probe_residuals_provider = \ - toolkit.residuals_runner - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.belief_probe import build_probe_namespace - probe_ns = build_probe_namespace(self._tool_context) - exec_ns["sim"] = probe_ns["sim"] - exec_ns["BeliefProbe"] = probe_ns["BeliefProbe"] - self._learning_mode = True - self._tool_context.extra_session_hooks = ( - self._build_synthesis_session_hooks( - [self._sampler_snapshot_target(paths)], base)) - - self._close_agent_session() - self._ensure_agent_session() - - predicate_listing = self._format_predicate_signatures( - self._get_all_predicates()) - options_digest = render_options_digest( - self._tool_context.options, - gt_options_ref_path=self._tool_context.gt_options_ref_path) - message = f"""\ -Synthesize per-skill samplers for this environment's options. The \ -simulator dynamics are already fixed (oracle/learned); your only job is \ -to make backtracking refinement land each option's continuous parameters \ -on its sketch-step subgoal instead of drawing them uniformly. - -## Available Predicates (subgoal geometry) -{predicate_listing} - -## Options -{options_digest} - -Explore the trajectory data with `run_python` (variables: \ -`trajectories`, `train_tasks`, `is_goal_state`, \ -`describe_trajectory(traj_idx)`, `np`, `ParamSpec`, plus the `sim` \ -probe over the deployed simulator - `sim.refine` is your \ -samples-to-refine feedback signal).""" - message = message + "\n\n" + self._sampler_synthesis_message(paths) - - try: - responses = self._query_agent_sync(message, kind="learn") - dead = query_fatal_error(responses) - if dead is not None: - # Nothing was synthesized: stop before this cycle is - # checkpointed as learned (see the simulator learn). - raise AgentSessionFatalError( - "The sampler-synthesis session died without the agent " - f"doing any work ({dead}); refusing to checkpoint this " - "cycle as learned.") - finally: - self._tool_context.extra_session_hooks = {} - self._tool_context.extra_mcp_tools = [] - self._tool_context.probe_artifact_loaders.clear() - self._tool_context.probe_fit_provider = None - self._tool_context.probe_validation_provider = None - self._tool_context.probe_residuals_provider = None - self._learning_mode = False - self._close_agent_session() - - self._finalize_and_load_samplers(paths) diff --git a/predicators/approaches/synthesis_validation.py b/predicators/approaches/synthesis_validation.py index 4a759fc6ec..e518fed271 100644 --- a/predicators/approaches/synthesis_validation.py +++ b/predicators/approaches/synthesis_validation.py @@ -16,10 +16,8 @@ from typing import TYPE_CHECKING, Any, Dict, List, Tuple from predicators.code_sim_learning.fit_space import ParamSpec -from predicators.code_sim_learning.fitting import fit_rule_parameters from predicators.code_sim_learning.utils import LearnedSimulator, \ - apply_rules, has_latent_rules, has_physics_rules -from predicators.structs import Action, State + apply_rules, has_latent_rules if TYPE_CHECKING: from predicators.agent_sdk.synthesis_backend import SynthesisBackend @@ -31,26 +29,20 @@ def build_candidate_option_model( approach: "SynthesisBackend", rules: List, specs: List[ParamSpec], - residual_features: Dict[str, List[str]], - base_pred_triples: List[Tuple[State, Action, State]], latent_init: Any = None, - fit: bool = True, -) -> Tuple[Any, Dict[str, float], float]: - """Fit ``specs`` (unless ``fit=False``) and build the candidate's option - model. - - ``fit=False`` builds the candidate at :func:`carry_over_params` - (the last published fit where a spec still exists and the value - lies in its box, the declared init value otherwise) and returns - ``nan`` for the SSE: fitting is the agent's explicit ``sim.fit`` - call, never a side effect of probing (see +) -> Tuple[Any, Dict[str, float]]: + """Build the candidate's option model at :func:`carry_over_params`. + + The parameters are the last published fit's where a spec still + exists and the value lies in its box, the declared init value + otherwise: fitting is the agent's explicit ``sim.fit`` call, never a + side effect of probing (see ``AgentSimLearningApproach._make_candidate_probe_model_provider``). The front half of the synthesis-session probe: every rollout must exercise the candidate simulator at its *deployed* (fitted) parameters, never at init_value. Returns ``(option_model, - fitted_params, fit_sse)``; raises ``RuntimeError`` when fitting - fails. + params)``. Publishes side effects onto ``approach`` exactly once, here, so the two surfaces can never disagree: the candidate ``rules`` / @@ -59,12 +51,6 @@ def build_candidate_option_model( place* (invented predicates hold a ``_ParamsView`` over it - the gating rule and the gating predicate must anchor to the same values). - - Recurrent (latent-declaring, 5-arg) rules are fit with the latent - threaded per trajectory; fully-observable rules take the legacy - per-transition path. Dispatch keys off the candidate rule - signatures (:func:`has_latent_rules`), as everywhere else in the - fitting stack. """ # pylint: disable=protected-access latent = has_latent_rules(rules) @@ -80,37 +66,10 @@ def build_candidate_option_model( if latent: approach._latent_init = latent_init - if not fit: - params = carry_over_params(approach._fitted_params, specs) - approach._fitted_params.clear() - approach._fitted_params.update(params) - return _finish_candidate_model(approach, rules, params), params, \ - float("nan") - try: - if has_physics_rules(rules): - # Physics-command rules act through engine stepping, so the - # teacher-forced objectives below cannot see them; fit - # against free-running rollouts instead (the same routing - # sim.fit uses). The joint fit also covers any declared - # PHYSICAL_PARAM_SPECS, which _load_simulator_from_module_file - # published onto the approach before this runs. - fit_result, fit_sse = approach._fit_parameters_joint_rollout( - rules, specs, residual_features) - elif latent: - fit_result, fit_sse = approach._fit_parameters_recurrent( - rules, specs, base_pred_triples, residual_features) - else: - fit_result, fit_sse = fit_rule_parameters(rules, specs, - base_pred_triples, - residual_features) - params = fit_result.point_estimate - except Exception as e: - raise RuntimeError(f"param fitting failed:\n{e}") from e - - # In place (clear + update, never replace): see docstring. + params = carry_over_params(approach._fitted_params, specs) approach._fitted_params.clear() approach._fitted_params.update(params) - return _finish_candidate_model(approach, rules, params), params, fit_sse + return _finish_candidate_model(approach, rules, params), params def carry_over_params(fitted: Dict[str, float], diff --git a/predicators/code_sim_learning/identifiability.py b/predicators/code_sim_learning/identifiability.py index 24908be060..fa40238b43 100644 --- a/predicators/code_sim_learning/identifiability.py +++ b/predicators/code_sim_learning/identifiability.py @@ -378,12 +378,10 @@ def physics_sigma_points(applied: Dict[str, float], num_points: int = 2) -> List[Dict[str, float]]: """A grid of perturbations spanning +-1 posterior sigma of the fit. - Consumed by the capture gate's physics-margin check - (``agent_plan_validation_physics_margin``) and by the ``sim.run`` - physics sweep: validation rollouts AT the fitted values sample - execution variability only, so a plan can pass them all while - having zero margin to the fit's parameter error - (run_20260723_091108: a capture validated 8/8 at fitted + Consumed by the ``sim.run`` physics sweep: rollouts AT the fitted + values sample execution variability only, so a plan can pass them + all while having zero margin to the fit's parameter error + (run_20260723_091108: a plan validated 8/8 at fitted lateral_friction 0.5319 failed deterministically at true 0.5). Each param whose FITTED value was deployed (``Verdict.applies_fitted``, which under the interval belief includes ``Verdict.WIDE``: a moved diff --git a/predicators/code_sim_learning/program_world_model.py b/predicators/code_sim_learning/program_world_model.py index 80f064df34..38ef65d139 100644 --- a/predicators/code_sim_learning/program_world_model.py +++ b/predicators/code_sim_learning/program_world_model.py @@ -136,9 +136,8 @@ class ProgramOptionModel(_OptionModelBase): The latent rides on ``State.latent``; a state without one (a task's initial state, a probe reset) is seeded from the program's - ``initial_latent`` - or from ``initial_latent_override`` while the - capture gate sweeps the belief particles. Program exceptions surface - as :class:`utils.OptionExecutionFailure` with the traceback tail in + ``initial_latent``. Program exceptions surface as + :class:`utils.OptionExecutionFailure` with the traceback tail in ``last_execution_failure``, the same channel the engine-backed model uses, so every consumer reports them the same way. """ @@ -147,7 +146,6 @@ def __init__(self, program: ProgramWorldModel, seed: int = 0) -> None: super().__init__() self._program = program self._rng = np.random.default_rng(seed) - self.initial_latent_override: Optional[Dict[str, Any]] = None self.last_execution_failure: Optional[str] = None self.last_trajectory: Optional[LowLevelTrajectory] = None self.sim_env: Optional[Any] = None @@ -161,10 +159,8 @@ def initial_latent( self, state: State, rng: Optional[np.random.Generator] = None) -> Dict[str, Any]: - """A latent for ``state``: the override if one is set, else a draw from - the program's ``initial_latent`` (with ``rng`` or the model's own).""" - if self.initial_latent_override is not None: - return copy.deepcopy(self.initial_latent_override) + """A latent for ``state``: a draw from the program's ``initial_latent`` + (with ``rng`` or the model's own).""" latent = self._program.initial_latent( _observation(state), rng if rng is not None else self._rng) if not isinstance(latent, dict): diff --git a/predicators/envs/pybullet_domino/env.py b/predicators/envs/pybullet_domino/env.py index e95ab76d35..ce62bcd36b 100644 --- a/predicators/envs/pybullet_domino/env.py +++ b/predicators/envs/pybullet_domino/env.py @@ -270,23 +270,18 @@ def _configure_instance_physics(self) -> None: # automatically after every reset_state. friction = CFG.domino_true_friction if self._skip_domain_specific_dynamics and \ - CFG.domino_planning_friction is not None and \ - not CFG.agent_sim_learn_oracle_sim_params: - # agent_sim_learn_oracle_sim_params grants the planner the - # TRUE friction (oracle upper bound) while task generation keeps - # using domino_planning_friction for the differentiation filter. + CFG.domino_planning_friction is not None: friction = CFG.domino_planning_friction if self._domino_component is not None and abs( friction - self._domino_component.domino_friction) > 1e-9: self.set_domino_physical_params(lateral_friction=friction) # Heavy-block tasks: planning sims BELIEVE the heavy gray blocks # are ordinary dominoes (normal mass), so their rollouts propagate - # a chain straight through one. The eval env (and the oracle- - # params planner) keeps the true heavy mass, asserted at reset. + # a chain straight through one. The eval env keeps the true heavy + # mass, asserted at reset. if CFG.domino_heavy_block_tasks \ and self._domino_component is not None \ - and self._skip_domain_specific_dynamics \ - and not CFG.agent_sim_learn_oracle_sim_params: + and self._skip_domain_specific_dynamics: self.set_domino_physical_params(block_mass=self.domino_mass) def _create_robot_predicates(self) -> None: diff --git a/predicators/envs/pybullet_domino/task_generators/min_block_generation.py b/predicators/envs/pybullet_domino/task_generators/min_block_generation.py index 2c4607dded..1ec0cc17ee 100644 --- a/predicators/envs/pybullet_domino/task_generators/min_block_generation.py +++ b/predicators/envs/pybullet_domino/task_generators/min_block_generation.py @@ -69,8 +69,9 @@ def _domino_code_digest() -> str: # Default L-shape leg-sampling bands (entry_leg, exit_leg) for turn tasks # when no explicit differentiating band is configured - the natural-corner # region shared by the plain and heavy turn variants. Explicit differentiating -# bands come from CFG.domino_min_block_turn_* (probe with -# scripts/domino_debug/probe_min_block_bands.py when the friction pair moves). +# bands come from CFG.domino_min_block_turn_* (re-probe them when the friction +# pair moves, with scripts/domino_debug/probe_min_block_bands.py from tag +# iclr-empiric-submission). _DEFAULT_TURN_ENTRY_BAND = (0.26, 0.34) _DEFAULT_TURN_EXIT_BAND = (0.18, 0.26) # Heavy turn variant: legs from the 2026-07-25 canonical-anchor design @@ -557,11 +558,12 @@ def _make_turn_task(env: "PyBulletDominoComposedEnv", # finite (entry, exit) cells - the data for any future retuning. direction = _planning_mismatch_direction() if CFG.domino_min_block_turn_entry_lo is not None: - # Explicitly configured bands - probe with - # scripts/domino_debug/probe_min_block_bands.py when the - # friction pair changes (the differentiating cells move with the - # frictions). The shipping under_reach arm's bands live in - # scripts/configs/predicatorv3/envs/all.yaml. + # Explicitly configured bands - re-probe them when the friction + # pair changes (the differentiating cells move with the + # frictions), with scripts/domino_debug/probe_min_block_bands.py + # from tag iclr-empiric-submission. The benchmark's bands live in + # the domino_high_friction_turn entry of + # scripts/configs/empiric/envs.yaml. assert CFG.domino_min_block_turn_entry_hi is not None assert CFG.domino_min_block_turn_exit_lo is not None assert CFG.domino_min_block_turn_exit_hi is not None @@ -580,9 +582,11 @@ def _make_turn_task(env: "PyBulletDominoComposedEnv", # which LLM planner arms cannot discover (retune 2026-07-12). raise ValueError( "under_reach turn tasks require explicit " - "domino_min_block_turn_{entry,exit}_{lo,hi} flags (probe with " - "scripts/domino_debug/probe_min_block_bands.py; see the " - "domino_high_friction block in envs/all.yaml).") + "domino_min_block_turn_{entry,exit}_{lo,hi} flags (see the " + "domino_high_friction_turn entry of " + "scripts/configs/empiric/envs.yaml; probe new bands with " + "scripts/domino_debug/probe_min_block_bands.py from tag " + "iclr-empiric-submission).") else: entry_leg = round(float(rng.uniform(*_DEFAULT_TURN_ENTRY_BAND)), 2) exit_leg = round(float(rng.uniform(*_DEFAULT_TURN_EXIT_BAND)), 2) diff --git a/predicators/execution_monitoring/subgoal_annotations_monitor.py b/predicators/execution_monitoring/subgoal_annotations_monitor.py deleted file mode 100644 index 93659e817e..0000000000 --- a/predicators/execution_monitoring/subgoal_annotations_monitor.py +++ /dev/null @@ -1,99 +0,0 @@ -"""An execution monitor that checks plan-sketch subgoal annotations at option -boundaries and suggests replanning on divergence.""" - -import logging -from dataclasses import dataclass -from typing import Any, Optional, Sequence - -from predicators.execution_monitoring.base_execution_monitor import \ - BaseExecutionMonitor -from predicators.structs import State, _Option - - -@dataclass -class SubgoalExecutionStatus: - """Live execution status of an annotated plan, exported by an approach via - ``get_execution_monitoring_info``. - - ``sketch`` items are duck-typed sketch steps exposing - ``subgoal_atoms``, ``subgoal_neg_atoms`` and ``option`` (see - ``agent_sdk.bilevel_sketch.SketchStep``); the type is kept loose so - the monitoring layer does not import agent_sdk. The approach's - dispensed policy mutates ``steps_initiated``/``current_option`` as - it executes, so the monitor always sees the live values. - """ - sketch: Sequence[Any] - steps_initiated: int = 0 - current_option: Optional[_Option] = None - - -class SubgoalAnnotationsExecutionMonitor(BaseExecutionMonitor): - """Suggest replanning when the step that just finished has a subgoal - annotation that does not hold in the real state. - - The check happens at the exact option boundary: when the currently - executing option's terminal condition is true in the given state, - the step it completes is checked before the policy advances to the - next option. Forward validation only proves a plan works in the - option model; real execution can still diverge (e.g. a place whose - drop-settle is chaotic lands off-target), after which the remaining - open-loop plan is doomed — it burns the episode horizon waiting for - effects that can no longer occur. Two boundaries are not caught: - divergence that only manifests inside a non-terminating option, and - Wait steps ended by the atom-change path in - ``utils.option_policy_to_policy`` (those terminate exactly when - their target atoms — derived from the same annotation — hold, so the - check would pass anyway). - """ - - @classmethod - def get_name(cls) -> str: - return "subgoal_annotations" - - def step(self, state: State) -> bool: - # No active annotated plan (e.g. exploration with an override - # policy, or replanning disabled): never suggest replanning. - if not self._approach_info: - return False - status = self._approach_info[0] - if not isinstance(status, SubgoalExecutionStatus): - return False - option = status.current_option - if option is None or status.steps_initiated <= 0: - return False - # Note: terminal() is also called by the policy machinery on - # this same state; skill terminal functions read memory but do - # not mutate it, so the double call is safe. - if not option.terminal(state): - return False - step_idx = status.steps_initiated - 1 - step = status.sketch[step_idx] - - def _holds(atom: Any) -> Optional[bool]: - # An atom the classifier cannot evaluate on this observation - # (e.g. an invented classifier that indexes a latent real - # states do not carry) is unverifiable, not diverged: skip - # it with a warning instead of crashing the episode or - # aborting on a claim no observation can settle. - try: - return bool(atom.holds(state)) - except Exception as e: # pylint: disable=broad-except - logging.warning( - "Subgoal atom %s is unverifiable on the real " - "observation (%s: %s); skipping it.", atom, - type(e).__name__, e) - return None - - unsat = [ - str(a) for a in (step.subgoal_atoms or set()) if _holds(a) is False - ] - unsat += [ - f"NOT {a}" for a in (step.subgoal_neg_atoms or set()) - if _holds(a) is True - ] - if not unsat: - return False - logging.info( - "Subgoal divergence after step %d (%s): unsatisfied %s. " - "Suggesting replan.", step_idx, step.option.name, sorted(unsat)) - return True diff --git a/predicators/explorers/__init__.py b/predicators/explorers/__init__.py index 31995dffbc..3a930dbbe7 100644 --- a/predicators/explorers/__init__.py +++ b/predicators/explorers/__init__.py @@ -1,13 +1,11 @@ """Handle creation of explorers.""" -import logging -from typing import TYPE_CHECKING, Callable, Dict, List, Optional, Set +from typing import Callable, Dict, List, Optional, Set from gym.spaces import Box from predicators import utils from predicators.competence_models import SkillCompetenceModel -from predicators.explorers.agent_explorer_base import AgentExplorerBase from predicators.explorers.base_explorer import BaseExplorer from predicators.explorers.bilevel_planning_explorer import \ BilevelPlanningExplorer @@ -18,10 +16,6 @@ NSRTSamplerWithEpsilonIndicator, ParameterizedOption, Predicate, State, \ Task, Type, _GroundSTRIPSOperator -if TYPE_CHECKING: - from predicators.agent_sdk.session_manager import SessionManagerProtocol - from predicators.agent_sdk.tools import ToolContext - __all__ = ["BaseExplorer"] # Find the subclasses. @@ -49,23 +43,10 @@ def create_explorer( seen_train_task_idxs: Optional[Set[int]] = None, pursue_task_goal_first: Optional[bool] = None, maple_q_function: Optional[MapleQFunction] = None, - tool_context: Optional["ToolContext"] = None, - agent_session: Optional["SessionManagerProtocol"] = None, ) -> BaseExplorer: """Create an explorer given its name.""" if max_steps_before_termination is None: max_steps_before_termination = CFG.max_num_steps_interaction_request - # Deprecated aliases from before the explorers' model-free / - # model-based rename (2026-08-30); old launch commands and - # requeued jobs still pass them. - aliases = { - "agent_plan": "agent_model_free", - "agent_bilevel": "agent_model_based", - } - if name in aliases: - logging.warning("Explorer name %r is deprecated; use %r.", name, - aliases[name]) - name = aliases[name] for cls in utils.get_all_subclasses(BaseExplorer): if not cls.__abstractmethods__ and cls.get_name() == name: # Special case GLIB because it uses babble predicates and an atom @@ -122,13 +103,6 @@ def create_explorer( action_space, train_tasks, max_steps_before_termination, nsrts, maple_q_function) - elif issubclass(cls, AgentExplorerBase): - assert tool_context is not None - assert agent_session is not None - explorer = cls(initial_predicates, initial_options, types, - action_space, train_tasks, - max_steps_before_termination, tool_context, - agent_session) else: explorer = cls(initial_predicates, initial_options, types, action_space, train_tasks, diff --git a/predicators/explorers/agent_explorer_base.py b/predicators/explorers/agent_explorer_base.py deleted file mode 100644 index 90b7eee828..0000000000 --- a/predicators/explorers/agent_explorer_base.py +++ /dev/null @@ -1,73 +0,0 @@ -"""Shared plumbing for the agent explorers. - -Both agent explorers query a Claude agent session for a plan and roll it -out in the real environment; they differ in what the agent is given (the -model-based explorer exposes the learned belief simulator through tools, -the model-free one only the task description). This base class holds the -common session wiring: construction, the trajectory summary shown to the -agent, final-text extraction, and the random-options fallback. -""" - -from typing import Any, Dict, List, Optional, Set - -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk.learn_prompts import render_world_model_notes_block -from predicators.agent_sdk.response_parser import extract_final_text -from predicators.agent_sdk.session_manager import SessionManagerProtocol -from predicators.agent_sdk.sketch_prompts import summarize_trajectories -from predicators.agent_sdk.tools import ToolContext -from predicators.explorers.base_explorer import BaseExplorer -from predicators.structs import Action, ExplorationStrategy, \ - ParameterizedOption, Predicate, State, Task, Type - - -class AgentExplorerBase(BaseExplorer): - """Base class for explorers that query a Claude agent session.""" - - def __init__(self, predicates: Set[Predicate], - options: Set[ParameterizedOption], types: Set[Type], - action_space: Box, train_tasks: List[Task], - max_steps_before_termination: int, tool_context: ToolContext, - agent_session: SessionManagerProtocol) -> None: - super().__init__(predicates, options, types, action_space, train_tasks, - max_steps_before_termination) - self._tool_context = tool_context - self._agent_session = agent_session - - def _agent_tool_names(self) -> Optional[List[str]]: - """Return tool names exposed by the current session, if any.""" - return getattr(self._agent_session, "tool_names", None) - - def _random_options_fallback(self) -> ExplorationStrategy: - """Fall back to random option sampling.""" - - def fallback_policy(state: State) -> Action: - del state - raise utils.RequestActPolicyFailure( - "Random option sampling failed!") - - policy = utils.create_random_option_policy(self._options, self._rng, - fallback_policy) - return policy, lambda _: False - - def _build_trajectory_summary(self) -> str: - """Summarize trajectory data for the agent.""" - all_trajs = (self._tool_context.offline_trajectories + - self._tool_context.online_trajectories) - return summarize_trajectories(all_trajs, - self._predicates, - train_tasks=self._train_tasks) - - def _world_model_notes_block(self) -> str: - """The natural-language world model quoted into the explore prompt - (empty unless the approach keeps one).""" - return render_world_model_notes_block( - self._tool_context.world_model_notes, - self._tool_context.world_model_notes_path) - - @staticmethod - def _extract_option_plan_text(responses: List[Dict[str, Any]]) -> str: - """Extract plan text from the last assistant text response.""" - return extract_final_text(responses) diff --git a/predicators/explorers/agent_model_based_explorer.py b/predicators/explorers/agent_model_based_explorer.py deleted file mode 100644 index 2bb31d7285..0000000000 --- a/predicators/explorers/agent_model_based_explorer.py +++ /dev/null @@ -1,512 +0,0 @@ -"""Agent model-based explorer: the agent sketches an experiment against the -learned belief model, and it runs as written. - -Queries a Claude agent for a fully parameterized plan sketch and rolls -it out for real exactly as written. The agent refines and validates -in-session against the currently-learned belief model (``sim.refine``, -``sim.run``, ``submit_plan``); a plan that passed the capture -gate executes as a belief-certified solve attempt, anything else -executes as an experiment with no harness-side parameter search or -substitution. When the belief model disagrees with reality (e.g. a -subgoal atom it expected after a Wait doesn't actually hold), the -trajectory is a targeted learning signal for online simulator -synthesis. - -Parallels ``AgentModelBasedApproach`` for the sketch workflow; the -session plumbing lives in ``AgentExplorerBase``. - -Registered under the CLI explorer name ``agent_model_based`` -(``agent_bilevel`` is kept as a deprecated alias). -""" - -import logging -import os -from typing import Any, Callable, Dict, List, Optional, Sequence - -import numpy as np - -from predicators import utils -from predicators.agent_sdk import bilevel_sketch -from predicators.agent_sdk.rendering import save_task_state_image -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error -from predicators.agent_sdk.session_manager import run_query_sync -from predicators.agent_sdk.tools import PlanCapture, agent_render_resolution, \ - load_ground_sampler_fns -from predicators.explorers.agent_explorer_base import AgentExplorerBase -from predicators.settings import CFG -from predicators.structs import Action, ExplorationStrategy, State, Task - - -class AgentModelBasedExplorer(AgentExplorerBase): - """Queries a Claude agent for a plan sketch and executes it as written.""" - - @classmethod - def get_name(cls) -> str: - return "agent_model_based" - - # ------------------------------------------------------------------ # - # Exploration strategy - # ------------------------------------------------------------------ # - - def _get_exploration_strategy(self, train_task_idx: int, - timeout: int) -> ExplorationStrategy: - task = self._train_tasks[train_task_idx] - # The approach syncs tool_context.option_model right before building - # this explorer, so reading here picks up the latest learned model. - option_model = self._tool_context.option_model - assert option_model is not None, \ - "agent_model_based explorer needs a synced option_model" - - # Reset the per-request mental-model verdict so a stale value can't - # leak if the query below throws or falls back to random before - # producing one. - self._tool_context.last_mental_model_solved = None - - # Point the agent's interactive tools (submit_plan, the - # sim probe) at the EXPLORE task. They - # default to ctx.current_task when the agent omits task_idx, and - # test-time _solve leaves current_task on the last TEST task. - # Without this the agent tunes/validates its exploration plan against - # the wrong task (e.g. a test goal referencing objects this task - # lacks), so parameter search is meaningless and only tasks solvable - # without tuning get solved. - # - # Enable the capture path too (keyed to current_task == this explore - # task): the agent often submits + simulator-validates a goal-reaching - # plan via submit_plan but ends with a - # prose summary whose final text doesn't parse into a sketch. Without - # capture that productive solve is lost to the random-options fallback; - # with it we recover the captured plan below (see _sketch_from_capture) - # and execute it at its captured params. Clear any - # stale capture first; the next test _solve re-points current_task and - # clears capture again, so an exploration plan can't leak into a test - # solve. - self._tool_context.current_task = task - self._tool_context.capture_goal_reaching_plans = True - # Exploration delivers plan sketches even in policy-mode configs: - # make sure a solve attempt's policy_capture_mode never bleeds - # into this query (it would disable the plan capture gate). - self._tool_context.policy_capture_mode = False - self._tool_context.clear_plan_capture() - - try: - prompt = bilevel_sketch.build_solve_prompt( - task, - all_predicates=self._predicates, - all_options=self._options, - trajectory_summary=self._build_trajectory_summary(), - tool_names=self._agent_tool_names(), - experiment_guidance=self._build_experiment_guidance(), - # Plans generated by this cycle's earlier requests: ask for a - # complementary plan instead of the identical one repeated. - scheduled_plans=list(self._tool_context.cycle_scheduled_plans), - initial_image_section=self._initial_image_section( - task, train_task_idx), - propose_params=CFG.agent_bilevel_use_llm_initial_params, - # The sketch is the experiment; a capture-gate-validated - # plan is welcome (it counts as a solve) but not required. - require_tool_validation=False, - # Explore contract: the sketch is a real-env experiment, and - # the belief model may lack goal-critical dynamics, so a - # simulator-failing sketch is a valid deliverable. - explore_mode=True, - ) - responses = run_query_sync(self._agent_session, - prompt, - kind="explore") - dead = query_fatal_error(responses) - if dead is not None: - # The session backend refused the query (usage limit, - # auth, transport): a random-options episode here would - # be junk data that the cycle then learns from (2026-08-28 - # run_20260827_171610 cycle 3). Terminate the run instead; - # the relaunch re-explores this cycle. - raise AgentSessionFatalError( - "explore query died without the agent doing any work " - f"({dead}); not falling back to random exploration.") - plan_text = self._extract_option_plan_text(responses) - # The session's tool capture: a goal-reaching plan the agent - # validated in the belief through the capture gate - # (submit_plan, N fresh - # rollouts). ``reached_goal`` is the gate's verdict. - capture = self._tool_context.take_plan_capture() - if CFG.agent_explorer_execute_certified_plan and capture.plan \ - and capture.reached_goal is True: - # Certified: the mental model solves the task with THIS - # plan, so run it verbatim as a solve attempt instead of - # re-searching (or boundary-probing) its parameters. A - # real success now counts for early stopping. The cycle's - # later requests see it under "plans already scheduled" - # and are asked for a DIFFERENT certified plan (a second - # test of the model), resubmitting this one only as a - # last resort. - plan = list(capture.plan) - logging.info( - "agent_model_based explorer: the agent's tool-validated " - "plan passed the belief's capture gate (%s); executing " - "it verbatim as this episode's solve attempt (mental " - "model solved the goal).", capture.validation_summary - or "goal reached") - if capture.sketch: - self._tool_context.last_sketch_subgoals = [ - (s.subgoal_atoms, s.subgoal_neg_atoms) - for s in capture.sketch - ] - self._tool_context.last_sketch_options = [ - (s.option.name, [o.name for o in s.objects]) - for s in capture.sketch - ] - self._tool_context.last_mental_model_solved = True - self._tool_context.cycle_scheduled_plans.append( - self._format_plan(plan) + - "\n NOTE: belief-certified; executes verbatim as a " - "solve attempt.") - return self._certified_plan_strategy(plan) - if not plan_text and not capture.plan: - raise ValueError("agent returned empty plan text") - - gs_fns, gs_err = load_ground_sampler_fns(self._tool_context) - if gs_err is not None: - logging.warning("[explore] %s", gs_err) - sketch = bilevel_sketch.parse_sketch_from_text( - plan_text, - task, - predicates=self._predicates, - options=self._options, - types=self._types, - parse_continuous_params=CFG. - agent_bilevel_use_llm_initial_params, - parse_ground_samplers=CFG.agent_bilevel_ground_samplers, - ground_sampler_fns=gs_fns or None, - ) if plan_text else [] - if not sketch: - sketch = self._sketch_from_capture(capture) or [] - if not sketch: - raise ValueError("parsed empty plan sketch") - self._tool_context.last_sketch_subgoals = [ - (s.subgoal_atoms, s.subgoal_neg_atoms) for s in sketch - ] - self._tool_context.last_sketch_options = [ - (s.option.name, [o.name for o in s.objects]) for s in sketch - ] - # The agent's sketch IS the experiment: it runs in the real - # environment exactly as written. The harness does no belief - # refinement of it - the agent refines and validates - # in-session (sim.refine / sim.run / submit_plan), - # and only a capture-gate-certified plan (handled above) - # counts as a belief-validated solve for early stopping. - plan = self._ground_sketch_verbatim(sketch) - self._tool_context.last_mental_model_solved = False - record = self._format_sketch(sketch, plan) - logging.info( - "agent_model_based explorer: executing the agent's sketch " - "verbatim for train task %d (%d steps; not " - "belief-certified):\n%s", train_task_idx, len(plan), record) - self._tool_context.cycle_scheduled_plans.append( - record + "\n NOTE: executes as written, without " - "belief-model certification; a real success will NOT " - "count toward early stopping.") - policy = utils.option_plan_to_policy( - plan, - abstract_function=lambda s: utils.abstract( - s, self._predicates)) - return self._wrap_policy(policy), lambda _: False - except AgentSessionFatalError: - # A random fallback would hide the broken session backend; - # re-raise so the run terminates. - raise - except Exception as e: # pylint: disable=broad-except - logging.warning(f"agent_model_based explorer failed: {e}. " - "Falling back to random options.") - - if not CFG.agent_explorer_fallback_to_random: - raise utils.RequestActPolicyFailure( - "agent_model_based explorer failed and fallback disabled.") - return self._random_options_fallback() - - # ------------------------------------------------------------------ # - # Helpers - # ------------------------------------------------------------------ # - - def _sketch_from_capture( - self, - capture: PlanCapture) -> Optional[List[bilevel_sketch.SketchStep]]: - """Rebuild a sketch from a captured, tool-validated plan, or None. - - ``submit_plan`` stashes a forward-validated plan on the explore - task into ``solved_plan`` (grounded options with continuous - params) and ``solved_sketch`` (the option skeleton plus the - subgoals that actually held). We reconstruct a sketch from that - skeleton and graft each captured option's continuous params onto - the step's ``initial_params``, so the plan executes at exactly - the values the agent validated. The capture was already taken - (consumed) by the caller. - """ - plan = capture.plan - captured_sketch = capture.sketch - if not plan or not captured_sketch: - return None - seeded: List[bilevel_sketch.SketchStep] = [] - for i, step in enumerate(captured_sketch): - params = None - if i < len(plan): - params = np.asarray(plan[i].params, dtype=np.float32) - seeded.append( - bilevel_sketch.SketchStep( - option=step.option, - objects=step.objects, - subgoal_atoms=step.subgoal_atoms, - subgoal_neg_atoms=step.subgoal_neg_atoms, - initial_params=params)) - logging.info( - "agent_model_based explorer: final text didn't parse, recovered " - "the " - "agent's tool-validated plan from capture (%d steps); executing " - "it at the captured params.", len(seeded)) - return seeded - - def _ground_sketch_verbatim( - self, sketch: Sequence[bilevel_sketch.SketchStep]) -> List[Any]: - """Ground each sketch step at the agent's proposed parameters. - - Nothing is searched or substituted: the executed plan is the - agent's. A step left without parameters (or with the wrong - arity) gets ONE uniform draw from the option's box and a - warning. Wait steps carry their annotated subgoals as - ``wait_target_atoms`` so the option terminates on the intended - atom change. - """ - plan: List[Any] = [] - for i, step in enumerate(sketch): - dim = step.option.params_space.shape[0] - params = step.initial_params - if params is None or len(params) != dim: - if dim > 0: - logging.warning( - "agent_model_based explorer: step %d (%s) has no " - "usable proposed params (%s); drawing one sample " - "from the option's box - propose every " - "parameter explicitly.", i, step.option.name, - None if params is None else list(params)) - params = bilevel_sketch.sample_params(step.option, self._rng) - plan.append( - bilevel_sketch.ground_step( - step, np.asarray(params, dtype=np.float32))) - return plan - - @staticmethod - def _format_sketch(sketch: Sequence[bilevel_sketch.SketchStep], - plan: Sequence[Any]) -> str: - """One indented ``i: Option(objs)[params] -> {atoms}`` line per - grounded step (params as executed, atoms as annotated).""" - lines = [] - for i, (step, opt) in enumerate(zip(sketch, plan)): - objs = ", ".join(o.name for o in opt.objects) - par = ", ".join(f"{p:.4f}" for p in opt.params) - line = f" {i}: {opt.name}({objs})[{par}]" - atoms = sorted(str(a) for a in (step.subgoal_atoms or set())) - atoms += sorted(f"NOT {a}" - for a in (step.subgoal_neg_atoms or set())) - if atoms: - line += f" -> {{{', '.join(atoms)}}}" - lines.append(line) - return "\n".join(lines) - - @staticmethod - def _format_plan(plan: Sequence[Any]) -> str: - """One indented ``i: Option(objs)[params]`` line per grounded - option.""" - lines = [] - for i, opt in enumerate(plan): - obj_s = ", ".join(o.name for o in opt.objects) - par_s = ", ".join(f"{p:.4f}" for p in opt.params) - lines.append(f" {i}: {opt.name}({obj_s})[{par_s}]") - return "\n".join(lines) - - def _certified_plan_strategy(self, - plan: Sequence[Any]) -> ExplorationStrategy: - """Execute a belief-certified grounded plan verbatim.""" - logging.info("agent_model_based explorer: certified plan:\n%s", - self._format_plan(plan)) - policy = utils.option_plan_to_policy( - list(plan), - abstract_function=lambda s: utils.abstract(s, self._predicates)) - return self._wrap_policy(policy), lambda _: False - - def _wrap_policy( - self, policy: Callable[[State], - Action]) -> Callable[[State], Action]: - """Convert OptionExecutionFailure into RequestActPolicyFailure. - - Lets the main loop cleanly terminate the episode when the - refined plan finishes or fails mid-execution (which is exactly - the disagreement signal we want to collect). - """ - - def _wrapped(state: State) -> Action: - try: - return policy(state) - except utils.OptionExecutionFailure as e: - raise utils.RequestActPolicyFailure(e.args[0], e.info) from e - - return _wrapped - - def _initial_image_section(self, task: Task, train_task_idx: int) -> str: - """Render the explore task's initial state and return a prompt section - pointing at it, mirroring what test-time solves get. - - Saved as ``train_task{N:03d}_initial_state.png`` so train-task - scenes are inspectable alongside the test-task init images. - Empty string when rendering is unavailable (e.g. the sandbox - isn't created yet, so ``image_save_dir`` is unset). - """ - env = self._tool_context.env - save_dir = self._tool_context.image_save_dir - if env is None or save_dir is None: - return "" - img_name = f"train_task{train_task_idx:03d}_initial_state.png" - with agent_render_resolution(): - saved = save_task_state_image(env, task, save_dir, img_name) - if saved is None: - return "" - # cwd of the agent is the sandbox root, so reference test_images/. - return ("\n## Initial State Image\n" - "A rendering of the initial scene has been saved to " - f"`./test_images/{img_name}`. **Read this image first** to " - "understand the spatial layout before planning.\n") - - def _build_experiment_guidance(self) -> str: - """LLM-proposal half of active-experiment design. - - Always injects the learn phase's open-questions ledger (the - ranked experiment specs it wrote for exploration to run) when - one exists in the sandbox. When info-seeking is on, additionally - point the agent at ``sim.suggest_probes`` and - when an ensemble - scorer is wired - at the predicates the learned model is - currently most internally uncertain about. - """ - parts = [] - ledger = self._read_open_questions() - if ledger: - parts.append( - "The learning phase left this ranked ledger of OPEN " - "QUESTIONS - uncertainties it could not settle from the " - "data collected so far, each with the experiment that " - "would settle it. The TOP entry is mandatory for this " - "cycle: run its experiment as specified (its option " - "sequence and parameters) unless a plan already " - "scheduled this cycle covers it, and fold in as many " - "lower entries as the step budget allows:\n" + ledger) - if self._tool_context.info_seeking_active(): - parts.append( - "Your explicit continuous parameters execute exactly as " - "written. To find the parameters a step could be run at " - "to teach the model most, call " - "`sim.suggest_probes(plan_text)`: it rolls your sketch " - "forward on your own parameters and, per annotated step, " - "ranks feasible alternatives by the learned model's " - "ensemble disagreement on the step's subgoal atoms. Adopt " - "one by writing it into your sketch, only on a step whose " - "failure the episode can afford; annotate steps with the " - "geometry/timing you are least sure the model has right.") - disagreement = self._build_disagreement_summary() - if disagreement: - parts.append(disagreement) - # System-ID gaps from the previous learn phase (synced by the - # sim-learning approach): what the collected data could NOT - # support, phrased as experiment objectives. Exploration is the - # only place those gaps can be filled. - sysid = getattr(self._tool_context, "sysid_diagnostics", None) - if sysid: - parts.append( - "The previous system-identification fit left gaps that " - "only new interaction data can close:\n" + sysid) - return "\n\n".join(parts) - - _MAX_OPEN_QUESTIONS_CHARS = 4000 - - def _read_open_questions(self) -> str: - """The learn phase's ./open_questions.md ledger, or "". - - The ledger is ranked, so on overflow the head is kept. Read - directly from the sandbox instead of asking the session to go - find it: the explore query should START from the ledger, not - spend its budget rediscovering it. - """ - sandbox_dir = self._tool_context.sandbox_dir - if not sandbox_dir: - return "" - path = os.path.join(sandbox_dir, "open_questions.md") - try: - with open(path, "r", encoding="utf-8") as f: - text = f.read().strip() - except OSError: - return "" - if len(text) > self._MAX_OPEN_QUESTIONS_CHARS: - # Cut at an entry boundary, never mid-sentence, and point at - # the file: a mid-sentence cut silently dropped the entries - # the header told the agent to fold in (run_20260830). - head = text[:self._MAX_OPEN_QUESTIONS_CHARS] - cut = head.rfind("\n#") - if cut <= 0: - cut = head.rfind("\n\n") - if cut > 0: - head = head[:cut] - text = (head.rstrip() + - "\n[... ledger truncated at the prompt cap - read " - "./open_questions.md for the remaining entries]") - return text - - def _build_disagreement_summary(self) -> str: - """Name the predicates the ensemble disagrees most about. - - Scans a bounded sample of recent-trajectory states, scores each - abstract atom's ensemble disagreement via the wired scorer, and - reports the highest-disagreement predicates. Grounded in the - actual ensemble, so it points the agent at genuinely-uncertain - dynamics rather than guesses. Empty when no scorer/trajectories. - """ - fn = self._tool_context.atom_disagreement_fn - if fn is None: - return "" - all_trajs = (self._tool_context.offline_trajectories + - self._tool_context.online_trajectories) - if not all_trajs: - return "" - recent = all_trajs[-CFG.agent_sdk_max_trajectories_in_context:] - states: List[State] = [] - for traj in recent: - n = len(traj.states) - if n == 0: - continue - stride = max(1, n // 6) # <= ~6 states/trajectory to bound cost - states.extend(traj.states[::stride]) - best: Dict[str, float] = {} - for s in states: - for atom in utils.abstract(s, self._predicates): - try: - d = float(fn(s, {atom})) - except Exception: # pylint: disable=broad-except - continue - name = atom.predicate.name - if d > best.get(name, 0.0): - best[name] = d - # One log line with the full ranking (scope note: abstract() yields - # true atoms only, so a predicate absent here was never measured, not - # necessarily agreed-upon). All values <= 0.05 -> no guidance: the - # ensemble is internally confident (or too tight) everywhere. - all_ranked = sorted(((v, k) for k, v in best.items()), reverse=True) - logging.info( - "agent_model_based explorer: per-predicate max ensemble " - "disagreement " - "over %d states - %s.", len(states), - ", ".join(f"{k}={v:.4f}" for v, k in all_ranked) or "(none)") - ranked = [(v, k) for v, k in all_ranked if v > 0.05][:4] - if not ranked: - return "" - named = ", ".join(f"{k} (disagreement {v:.2f})" for v, k in ranked) - return ("Across recent trajectories, the learned model is most " - f"internally uncertain about: {named}. A sketch that puts " - "these predicates on the critical path will be most " - "informative.") diff --git a/predicators/explorers/agent_model_free_explorer.py b/predicators/explorers/agent_model_free_explorer.py deleted file mode 100644 index bfcfd04cbb..0000000000 --- a/predicators/explorers/agent_model_free_explorer.py +++ /dev/null @@ -1,175 +0,0 @@ -"""Agent model-free explorer: Claude agent generates grounded option plans -without a learned world model. - -Produces fully-grounded option plans (including continuous parameters) -from a one-shot task description and rolls them out in the real -environment. Unlike ``AgentModelBasedExplorer``, the agent has no belief -simulator to validate against and no backtracking refinement; it must -supply complete parameters itself. - -Registered under the CLI explorer name ``agent_model_free`` -(``agent_plan`` is kept as a deprecated alias). -""" - -import logging - -import numpy as np - -from predicators import utils -from predicators.agent_sdk.session_base import AgentSessionFatalError, \ - query_fatal_error -from predicators.agent_sdk.session_manager import run_query_sync -from predicators.explorers.agent_explorer_base import AgentExplorerBase -from predicators.settings import CFG -from predicators.structs import ExplorationStrategy, Task - - -class AgentModelFreeExplorer(AgentExplorerBase): - """Queries a Claude agent to produce grounded option plans.""" - - @classmethod - def get_name(cls) -> str: - return "agent_model_free" - - def _get_exploration_strategy(self, train_task_idx: int, - timeout: int) -> ExplorationStrategy: - task = self._train_tasks[train_task_idx] - try: - prompt = self._build_exploration_prompt(train_task_idx) - responses = run_query_sync(self._agent_session, - prompt, - kind="explore") - dead = query_fatal_error(responses) - if dead is not None: - # See agent_model_based_explorer: never explore at random - # because the session backend is down. - raise AgentSessionFatalError( - "explore query died without the agent doing any work " - f"({dead}); not falling back to random exploration.") - plan_text = self._extract_option_plan_text(responses) - if plan_text: - option_plan = self._parse_and_ground_plan(plan_text, task) - if option_plan: - # The Wait wrapper needs the abstraction to end a - # Wait on its target atoms. - policy = utils.option_plan_to_policy( - option_plan, - abstract_function=lambda s: utils.abstract( - s, self._predicates)) - return policy, lambda _: False - logging.info("Agent explorer: no valid plan, falling back to " - "random options.") - except AgentSessionFatalError: - # A random fallback would hide the broken session backend; - # re-raise so the run terminates. - raise - except Exception as e: # pylint: disable=broad-except - logging.warning(f"Agent explorer failed: {e}. " - "Falling back to random options.") - - if not CFG.agent_explorer_fallback_to_random: - raise utils.RequestActPolicyFailure( - "Agent explorer failed and fallback disabled.") - return self._random_options_fallback() - - def _build_exploration_prompt(self, train_task_idx: int) -> str: - """Build a prompt for the agent to produce an option plan.""" - task = self._train_tasks[train_task_idx] - init_state = task.init - - objects = list(init_state) - obj_strs = [] - for obj in sorted(objects, key=lambda o: o.name): - obj_strs.append(f" {obj.name}: {obj.type.name}") - - # Goal atoms - goal_strs = [str(a) for a in sorted(task.goal, key=str)] - - # Available options with signatures. - all_options = self._options - option_strs = [] - for opt in sorted(all_options, key=lambda o: o.name): - type_sig = ", ".join(t.name for t in opt.types) - params_dim = opt.params_space.shape[0] - if params_dim > 0: - low = opt.params_space.low.tolist() - high = opt.params_space.high.tolist() - param_info = (f", params_dim={params_dim}, " - f"low={low}, high={high}") - else: - param_info = "" - option_strs.append(f" {opt.name}({type_sig}{param_info})") - - # Current atoms - atoms = utils.abstract(init_state, self._predicates) - atom_strs = [str(a) for a in sorted(atoms, key=str)] - - # Trajectory summary - traj_summary = self._build_trajectory_summary() - - # Available tools - tools_str = "" - tool_names = self._agent_tool_names() - if tool_names: - tool_list = "\n".join(f" - {t}" for t in tool_names) - tools_str = f"\n## Available Tools\n{tool_list}\n" - - task_intro = ("You are exploring a task environment. " - f"Generate an option plan to explore task " - f"{train_task_idx}.") - prompt = f"""{task_intro} - -## Goal -{chr(10).join(goal_strs)} - -## Initial State Atoms -{chr(10).join(atom_strs)} - -## Objects -{chr(10).join(obj_strs)} - -## Available Options -{chr(10).join(option_strs)} -{traj_summary}{self._world_model_notes_block()}{tools_str} -## Instructions -Use your available tools to inspect the environment and test your plan before committing to it. - -Output an option plan, one option per line, in this exact format: -OptionName(obj1:type1, obj2:type2)[param1, param2] - -If an option has no continuous parameters, use empty brackets: OptionName(obj1:type1)[] - -Output ONLY the option plan lines at the end, after any analysis.""" - - return prompt - - def _parse_and_ground_plan(self, plan_text: str, task: Task) -> list: - """Parse option plan text and ground into executable options.""" - objects = list(task.init) - all_options = self._options - parsed = utils.parse_model_output_into_option_plan( - plan_text, - objects, - self._types, - all_options, - parse_continuous_params=True) - if not parsed: - logging.info("Agent explorer: parsed empty option plan.") - return [] - - grounded = [] - for option, objs, params in parsed: - try: - ground_opt = option.ground(objs, - np.array(params, dtype=np.float32)) - grounded.append(ground_opt) - except Exception as e: # pylint: disable=broad-except - logging.info(f"Agent explorer: failed to ground " - f"option {option.name}: {e}") - break - - if not grounded: - logging.info("Agent explorer: no options successfully grounded.") - else: - logging.info(f"Agent explorer: grounded {len(grounded)} options.") - return grounded diff --git a/predicators/explorers/fixed_plan_explorer.py b/predicators/explorers/fixed_plan_explorer.py index 85c8ae9b44..2ad08d0164 100644 --- a/predicators/explorers/fixed_plan_explorer.py +++ b/predicators/explorers/fixed_plan_explorer.py @@ -28,7 +28,7 @@ from predicators.settings import CFG from predicators.structs import ExplorationStrategy, Object, State, _Option -# Same grammar as scripts/domino_debug/replay_plan.py. +# One option per line: ``Name(obj, ...) [param, ...]``. _LINE = re.compile(r"^\s*(\w+)\s*\(([^)]*)\)\s*\[([^\]]*)\]") diff --git a/predicators/ground_truth_models/skill_factories/base.py b/predicators/ground_truth_models/skill_factories/base.py index 601a98ca3e..828caca57d 100644 --- a/predicators/ground_truth_models/skill_factories/base.py +++ b/predicators/ground_truth_models/skill_factories/base.py @@ -218,9 +218,6 @@ def _fmt_option_params(params: Array) -> str: # (state, objects, params, config) -> (x, y, z, yaw) TargetPoseFn = Callable[[State, Sequence[Object], Array, SkillConfig], Tuple[float, float, float, float]] -# Objects a skill contacts by design beyond its arguments, for a -# grounding; see ``PhaseSkill.contact_objects``. -ContactObjectsFn = Callable[[State, Sequence[Object]], Set[Object]] # --------------------------------------------------------------------------- # Internal type aliases for Phase target functions @@ -492,16 +489,14 @@ class PhaseSkill: option = PhaseSkill("Pick", types, params_space, config, phases).build() """ - def __init__( - self, - name: str, - types: Sequence[Type], - params_space: Box, - config: SkillConfig, - phases: List[Phase], - params_description: Optional[Tuple[str, ...]] = None, - base_mode: Optional[str] = None, - contact_objects_fn: Optional[ContactObjectsFn] = None) -> None: + def __init__(self, + name: str, + types: Sequence[Type], + params_space: Box, + config: SkillConfig, + phases: List[Phase], + params_description: Optional[Tuple[str, ...]] = None, + base_mode: Optional[str] = None) -> None: assert len(phases) > 0 self._name = name self._types = types @@ -509,9 +504,6 @@ def __init__( self._config = config self._phases = phases self._params_description = params_description - # Objects the skill touches by design beyond its arguments (a - # push's switch); see ``contact_objects``. - self._contact_objects_fn = contact_objects_fn # Mobile-base positioning mode for this skill (None disables it): # "home" park at the robot's home base (good offset to press a # switch; diagonal fixed-base reach for far targets). @@ -538,22 +530,6 @@ def build(self) -> ParameterizedOption: params_description=self._params_description, ) - def contact_objects(self, state: State, - objects: Sequence[Object]) -> Set[Object]: - """Objects this skill contacts by design that are not among its - arguments, for the given grounding. - - A push skill's argument is the appliance (faucet, burner, fan) - while the body its finger strikes is that appliance's switch, a - separate object. Robot-clearance checks (the capture gate's - bystander probe) exempt these along with the arguments: the - contact is the skill's purpose, not a margin-free near miss. - Empty when the skill declares none. - """ - if self._contact_objects_fn is None: - return set() - return set(self._contact_objects_fn(state, objects)) - def _initiable(self, state: State, memory: Dict, objects: Sequence[Object], params: Array) -> bool: del state, objects, params # unused diff --git a/predicators/ground_truth_models/skill_factories/push.py b/predicators/ground_truth_models/skill_factories/push.py index 7521f3d3c2..fcb856c36b 100644 --- a/predicators/ground_truth_models/skill_factories/push.py +++ b/predicators/ground_truth_models/skill_factories/push.py @@ -47,8 +47,7 @@ def _get_domino_pose(state, objects, params, config): ) """ -import math -from typing import Callable, List, Optional, Sequence, Set, Tuple +from typing import Callable, List, Optional, Sequence, Tuple import numpy as np @@ -75,35 +74,6 @@ def _get_domino_pose(state, objects, params, config): "pass over a short target)", 0.0, 0.11), ] -# A push target pose is computed from the pose of the body it strikes, -# so that body sits within a few millimetres of it; anything farther is -# not the push's target. -CONTACT_TARGET_RADIUS = 0.02 - - -def object_at_pose(state: State, - pose: Tuple[float, float, float], - exclude: Set[Object], - radius: float = CONTACT_TARGET_RADIUS) -> Optional[Object]: - """The posed object nearest ``pose`` within ``radius``, or None. - - Objects without x/y/z features and those in ``exclude`` (a - grounding's own arguments, the robot) are skipped. - """ - best: Optional[Object] = None - best_dist = radius - for obj in state: - if obj in exclude: - continue - feats = obj.type.feature_names - if not {"x", "y", "z"}.issubset(feats): - continue - dist = math.sqrt( - sum((state.get(obj, f) - v)**2 for f, v in zip("xyz", pose))) - if dist < best_dist: - best, best_dist = obj, dist - return best - def resolve_ee_yaw_offset(config: SkillConfig) -> float: """The EE yaw offset Push should use, in radians. @@ -212,15 +182,6 @@ def create_push_skill( list(_PUSH_PARAMS) + list(extra_params)) _empty = np.array([], dtype=np.float32) - def _contact_objects(state: State, - objects: Sequence[Object]) -> Set[Object]: - # The body at the target pose is what the push is for (a - # faucet's switch, a fan's switch): it is exempt from clearance - # checks like an argument, see PhaseSkill.contact_objects. - x, y, z, _ = get_target_pose_fn(state, objects, _empty, config) - target = object_at_pose(state, (x, y, z), exclude=set(objects)) - return set() if target is None else {target} - # -- Standard 4-waypoint trajectory ---------------------------------- def _waypoints( @@ -375,5 +336,4 @@ def _get_target( config, phases, params_description=params_description, - base_mode="home", - contact_objects_fn=_contact_objects).build() + base_mode="home").build() diff --git a/predicators/settings.py b/predicators/settings.py index a069d68386..591c4b5ca6 100644 --- a/predicators/settings.py +++ b/predicators/settings.py @@ -1097,9 +1097,10 @@ class GlobalSettings: # from [lo, hi] per attempt. None (default) = the legacy over_reach band # hardcoded in _make_turn_task (the low-friction arm); under_reach arms # MUST set these explicitly - the legacy under_reach band shipped - # agent-intractable pair-corner tasks. Probe candidate bands with - # scripts/domino_debug/probe_min_block_bands.py whenever the friction - # pair changes (the differentiating cells move with the frictions). + # agent-intractable pair-corner tasks. Probe candidate bands whenever + # the friction pair changes (the differentiating cells move with the + # frictions), with scripts/domino_debug/probe_min_block_bands.py from + # tag iclr-empiric-submission. domino_min_block_turn_entry_lo: Optional[float] = None domino_min_block_turn_entry_hi: Optional[float] = None domino_min_block_turn_exit_lo: Optional[float] = None @@ -1166,8 +1167,8 @@ class GlobalSettings: fan_train_num_pos_y = 3 # The historical 6 x 6 uniform test split. The loc bounds in # pybullet_fan.py admit at most 10 x 9 cells at the 8 cm pitch, which - # fills the arena up to the fan rows; the maze split in - # scripts/configs/predicatorv3/envs/all.yaml uses that full grid. + # fills the arena up to the fan rows; the historical maze split used + # that full grid. fan_test_num_pos_x = 6 fan_test_num_pos_y = 6 fan_train_num_walls_per_task = [1] @@ -1584,17 +1585,6 @@ class GlobalSettings: gnn_use_validation_set = True # parameters for GNN option policy approach - # GNN dynamics + shooting baseline (gnn_dynamics_shooting, paper arm - # C5): how many previous pre-option states ride along as node - # features (so a hidden mechanism is inferable from the recent - # past), the longest option sequence one shooting try samples, how - # many tries a plan query gets before failing, and whether the plan - # is re-shot from the observed state after every option (MPC) or - # executed open-loop. - gnn_dynamics_history_len = 2 - gnn_dynamics_max_plan_length = 30 - gnn_dynamics_shooting_max_tries = 200 - gnn_dynamics_replan_every_option = True gnn_option_policy_solve_with_shooting = True gnn_option_policy_shooting_variance = 0.1 gnn_option_policy_shooting_max_samples = 100 @@ -2087,8 +2077,9 @@ class GlobalSettings: # or a stream error before the first tool call) before the run # terminates with AgentSessionFatalError. Such failures make every # future query hopeless, but each one returns in ~1 s at $0.00 and is - # otherwise indistinguishable from a no-capture attempt, so without - # this check the solve restart / replan / online-cycle budgets grind + # otherwise indistinguishable from an attempt that produced nothing, + # so without this check the solve restart / replan / online-cycle + # budgets grind # through hundreds of instant failures (run_20260721_161159: 300 # "organization has disabled Claude subscription access" queries # across 10 cycles, agent never ran). 0 disables the check. @@ -2115,22 +2106,15 @@ class GlobalSettings: # maximum buffer size". 20 MB comfortably fits full-res scene images. agent_sdk_max_buffer_size = 20 * 1024 * 1024 agent_sdk_resume_session = True # resume previous session if available - agent_sdk_max_trajectories_in_context = 3 - agent_sdk_log_agent_responses = True # Sandbox settings for agent SDK - agent_sdk_use_docker_sandbox = False # run agent inside Docker container - agent_sdk_docker_image = "predicators-sandbox" # Docker image name - # sandbox dir with built-in tools, no Docker + # sandbox dir with built-in tools agent_sdk_use_local_sandbox = False - # Agent explorer settings - agent_explorer_fallback_to_random = True # fall back to random on failure - # Agent planner approach settings agent_planner_use_scratchpad = False # include notes.md scratchpad # Whether the planner is given a simulator to test candidate plans with - # (the submit_plan tool / option-model rollouts). When False, the + # (option-model rollouts). When False, the # agent must plan open-loop from trajectory data and LLM reasoning alone # -- the genuinely model-free baseline. agent_planner_use_simulator = True @@ -2144,85 +2128,24 @@ class GlobalSettings: # Agent bilevel approach settings agent_bilevel_max_samples_per_step = 50 # param samples per step - agent_bilevel_check_subgoals = True # check subgoal atoms after each step - # When True, the agent proposes per-step continuous parameters inside the - # plan sketch (`Option(obj:type)[p1, p2] -> {subgoals}`). Refinement tries - # the proposed params first, then falls back to the registered sampler / - # uniform backtracking on failure. Default False keeps the param-free - # sketch (search finds all continuous params). - agent_bilevel_use_llm_initial_params = False # When True, sketch steps may carry GROUND samplers - per-step, per-call - # sampling priors that override any learned parameterized sampler for - # that step (precedence: ground > parameterized > uniform). Two forms - # after a step's `[params]`: a uniform window `~ [w1, w2]` + # sampling priors that replace uniform sampling for that step. Two + # forms after a step's `[params]`: a uniform window `~ [w1, w2]` # (per-dimension half-widths around the proposed params) or a named # code sampler `~ my_sampler` referencing GROUND_SAMPLERS in the # sandbox's ground_samplers.py, loaded fresh on each refine call - # (signature (state, subgoal_atoms, rng, objects) -> params, same as a - # parameterized sampler, so any state-conditioned region is - # expressible). Default False hides the grammar from the agent and - # rejects the annotations, keeping baseline arms free of the channel. + # (signature (state, subgoal_atoms, rng, objects) -> params, so any + # state-conditioned region is expressible). Default False hides the + # grammar from the agent and rejects the annotations, keeping baseline + # arms free of the channel. agent_bilevel_ground_samplers = False - # When True, close the agent SDK session at the start of each test task - # so every test solve begins with a FRESH conversation (no context from - # earlier test tasks). The sandbox filesystem and learned artifacts are - # untouched. Default False keeps the current behavior: all test tasks - # share one continuous agent conversation. - agent_fresh_session_per_test_task = False - # Restart loop for test-task solving. Solve-time outcomes are close to - # heavy-tailed in agent-search quality (run_20260717 family split: the - # same tasks solved in 9-32 min in one launch and burned 2-11 h without - # solving in its identical sibling, anchored on wrong conclusions), so - # several short, independent attempts beat one long one. Each attempt - # above the first starts from a fresh conversation; the solve journal - # (below) carries curated knowledge across attempts. An attempt ends - # early with a validated (evaluator-solved) capture; otherwise its - # best-effort capture is banked and the best across attempts executes. - # Each attempt is exactly ONE agent query: however that query ends - - # a spent budget, an unparseable sketch, or a session that simply - # never submitted - the fresh-context restart is the only retry, so - # this is the sole knob controlling how many shots a task gets. Only - # the final attempt (no restart left) pays for the best-effort - # submission nudge. - agent_solve_max_attempts = 1 - # Wall-clock budget per solve attempt, in seconds (0 disables). The - # turn cap bounds turns, not compute - one run_python sweep hid - # 47k rollouts (~7 h) inside a single turn. On expiry, exploration - # tools refuse with a submit-now message and the approach runs the - # same best-effort submission flow as turn-cap exhaustion. - agent_solve_attempt_wall_clock = 0.0 - # When True, every solve attempt (including the first, i.e. every test - # task) begins with a fresh agent conversation; cross-attempt and - # cross-task knowledge travels through the solve journal instead of - # raw transcript history, which also carries the *wrong* conclusions - # of failed attempts. - agent_solve_fresh_context = False - # Persistent per-run solve journal: the harness logs each attempt's - # outcome + captured plan to /attempts.md, the agent keeps - # its own lessons in /journal.md with the file tools, and - # both are injected into every solve prompt. The prompts ask for - # facts/measurements rather than verdicts, so failed attempts steer - # later ones away from repeated sweeps without re-importing their - # anchoring mistakes. - agent_solve_use_journal = False - # Closed-loop policy mode: the solve agent's deliverable is a per-task - # PROGRAM (/policy.py with get_option(state, memory) -> next - # plan line or None) validated in the belief model via the - # submit_policy tool and executed at test time WITHOUT an LLM in - # the loop. Option failures are surfaced to the policy (via - # memory["last_failure"]) instead of ending the episode, so recovery - # (re-place a drifted block, re-aim after a BiRRT refusal) is the - # policy's job - which is why this mode is mutually exclusive with - # the sketch-divergence replan machinery - # (agent_bilevel_max_execution_replans must be 0). - agent_solve_policy_mode = False - # Total options a policy episode may issue (belief validation AND real - # execution): the anti-oscillation bound that converts a retry loop - # that never progresses into a bounded, attributable failure. + # Total options one policy rollout of sim.run_policy may issue: the + # anti-oscillation bound that converts a retry loop that never + # progresses into a bounded, attributable failure. agent_policy_max_options = 50 # Consecutive identical failures (same option, objects, and params) - # after which a policy episode ends with a fatal stuck-loop error, in - # belief validation AND real execution. An unchanged command that + # after which a sim.run_policy rollout ends with a fatal stuck-loop + # error. An unchanged command that # just failed fails the same way; re-issuing it is a policy bug, not # recovery (the 2026-08-22 policy-arm tests burned 20+ of their 50 # options on one identical colliding PickBlock). 3 still allows a @@ -2236,18 +2159,14 @@ class GlobalSettings: # cannot see (nothing fails). 3 tolerates a benign settle-in-place # step without letting a livelock burn the budget. agent_policy_max_repeated_noops = 3 - # LLM-free bypass: path to a prewritten policy.py used as the captured - # artifact for every test task (mirrors the sketch-file bypass). For - # smoke tests and debugging the execution path. - agent_policy_file = "" # --auto_resume only resumes from checkpoints modified within this # many hours. The checkpoint path ignores the run timestamp, so a # RELAUNCH of a finished experiment under the same experiment_id # would otherwise silently continue the old run; a Slurm requeue or # a prompt resubmission of a live run is always recent. auto_resume_max_age_hours = 36.0 - # Per-call wall-clock limit for solve-session run_python calls, in - # seconds (0 disables). Enforced cooperatively at every probe sim + # Per-call wall-clock limit for run_python calls, in seconds (0 + # disables). Enforced cooperatively at every probe sim # call, plus a hard async-exception watchdog for sim-free code (a # pure-Python loop blocks the event loop, so nothing else can stop # it), so a combinatorial sweep stops with its printed output @@ -2273,64 +2192,9 @@ class GlobalSettings: # new params) and report UNFITTED until the agent fits the current # file. 0 disables the cap. agent_sdk_fit_call_timeout = 3600.0 - # Test-time closed-loop recovery. After each option in the refined plan - # finishes, the subgoal_annotations execution monitor checks the - # sketch's subgoal annotation for that step against the REAL state; on - # divergence (execution left the option-model rollout — e.g. a place - # that settled off-target), CogMan re-invokes solve(), which resumes a - # re-refined suffix of the executed sketch from the current state, - # instead of running the rest of the stale plan open-loop. Value = - # recoveries per test episode, shared across chained replans; when no - # suffix refines (or the budget is spent) the remaining plan resumes - # open-loop rather than failing the episode. 0 disables (legacy - # open-loop execution). Requires --execution_monitor - # subgoal_annotations (enforced at approach construction). - agent_bilevel_max_execution_replans = 0 - # When an execution replan's suffix refinement fails, whether to fall - # back to querying the agent for a fresh sketch - a brand-new - # full-turn-budget session. Default False: the cheap suffix replan is - # the only recovery, and when no suffix of the executed sketch - # refines from the diverged state the remaining plan resumes - # open-loop (the divergence is logged; the goal check decides the - # episode). Re-opening the agent budget is especially wasteful after - # a best-effort (non-solve) capture, whose execution diverges by - # construction. - agent_bilevel_replan_agent_fallback = False - # log state pretty_str before/after each step - agent_bilevel_log_state = False - # Load a plan sketch from a file instead of querying the LLM. The dir is - # under scripts/; the file may be a bare name or an absolute path. - agent_bilevel_plan_sketch_dir = "plan_sketches" - agent_bilevel_plan_sketch_file = "" - # When a sketch refinement runs without an explicit timeout, the - # caller computes - # max(_min, _per_step * len(sketch)) - # so plans with more steps automatically get more wall-clock budget. - agent_bilevel_refinement_timeout_per_step = 30.0 # seconds per step - agent_bilevel_refinement_timeout_min = 30.0 # floor on auto-scaled timeout - # Total number of belief-sim rollouts a goal-reaching plan must pass in - # submit_plan before it is captured as the agent's answer. The - # shared sim env is nondeterministic across repeats (motion-planner - # sampling, physics-solver state), so repeats sample the same execution - # variability the real rollout will - a flaky plan is reported to the - # agent in-session (where it can add margin and resubmit) instead of - # captured and discovered as a failed real episode. 1 disables repeats. - # An n-rollout gate passes a plan with per-rollout success rate p with - # probability p^n, so small n lets marginal plans through: at p=0.85, - # 3 rollouts pass 61% and 5 pass 44% (bridge run_20260819_053515: a - # knife-edge grasp offset validated 3/3, then cammed out on the real - # episode). The agent can request a stricter gate per submission via - # the tool's validation_rollouts argument; it can never lower this. - agent_plan_validation_rollouts = 5 - # Escalated rollout count once a task has produced a FLAKY rejection - # (see the p^n math above; run_20260717_182321: a 20/20-swept relay - # placement validated 3/3, then missed the target for real). A FLAKY - # rejection is direct evidence the agent is tuning in a marginal - # region, so subsequent captures on that task must clear this stricter - # gate instead. Never lowers the base count. - agent_plan_validation_rollouts_after_flaky = 10 - # Run each validation rollout inside ``ctx.validation_env_scope`` (a - # freshly constructed sim env) when the approach installs one. A shared + # Run each probe trial and sweep rollout inside + # ``ctx.validation_env_scope`` (a freshly constructed sim env) when + # the approach installs one. A shared # env's reset provably cannot reconstruct state exactly (solver # warm-start state, velocity residuals, near-matching bodies skipped by # the reconstruction diff - see rollout_states in physical_sysid.py), so @@ -2338,70 +2202,25 @@ class GlobalSettings: # the fresh real env; fresh envs make them honest i.i.d. samples of what # the real episode will draw. Costs one env construction per rollout. agent_plan_validation_fresh_env = True - # Physics-margin gate on captures: after a goal-reaching plan passes - # the execution-validation rollouts, re-run it at a grid of - # perturbations spanning +-1 sigma of the identified physical - # parameters (sigma = the posterior width the sysID fit reported, - # floored by code_sim_learning_rollout_min_posterior_width). The - # execution repeats above only sample motion-planner/physics-stepping - # variability AT the fitted values; a plan can pass them all and - # still have zero margin to the fit's parameter error - # (run_20260723_091108: a capture validated 8/8 at fitted - # lateral_friction 0.5319 failed deterministically at true 0.5 - - # the design's success band started at the fitted value). A failing - # perturbed rollout refuses the capture as PARAM-SENSITIVE so the - # agent adds design margin in-session. Runs only when the approach - # installs a fresh-env scope (perturbing the shared env would leak) - # and a fit with nonzero posterior width has been applied. Default - # False so existing arms keep their behavior; the main arm - # (approaches/all.yaml sim_predicator) - # turns it on. - agent_plan_validation_physics_margin = False - # Number of grid points the margin gate (and the sim.run physics - # sweep) spreads evenly across the +-1-sigma range, endpoints - # included. Endpoints alone (2) are provably insufficient: near a - # feasibility boundary success is a SPECKLED function of the - # params, and run_20260724_140531's capture passed both +-1-sigma - # endpoints (lateral_friction 0.4295/0.5246) while failing + # Number of grid points the sim.run physics sweep spreads evenly + # across the +-1-sigma range of the last applied fit, endpoints + # included (without the joint belief; with it the sweep reads the + # belief's interval ends). Endpoints alone (2) are provably + # insufficient: near a feasibility boundary success is a SPECKLED + # function of the params, and run_20260724_140531's plan passed both + # +-1-sigma endpoints (lateral_friction 0.4295/0.5246) while failing # deterministically at the true 0.5 between them. Replaying that - # capture mapped the speckle: a hazard band [~0.494, 0.511] holding + # plan mapped the speckle: a hazard band [~0.494, 0.511] holding # ~30% failures at ~0.001 grain, so ANY even grid is a # probabilistic detector - a 16-point grid's two in-band points - # both passed (would still have captured), while the 32-point - # grid's 0.5046 fails (rejects it). Per-point rollouts are - # deterministic measurements costing one rollout (~seconds), and - # captures are infrequent, so density is cheap sensitivity; designs - # with real margin pass every density identically. + # both passed, while the 32-point grid's 0.5046 fails. Per-point + # rollouts are deterministic measurements costing one rollout + # (~seconds), so density is cheap sensitivity; designs with real + # margin pass every density identically. agent_plan_validation_physics_margin_points = 32 - # Rule-parameter margin gate: after the physics points, re-run each - # capture-eligible submission under the calibrated rule-parameter - # ensemble members (the same posterior draws info-seeking - # exploration scores with), rejecting as PARAM-SENSITIVE a plan - # that survives only at the point estimate of an uncertain LEARNED - # constant (a gate threshold, a geometric offset). The physics - # sweep cannot catch these: it perturbs identified base-physics - # params, while a learned rule constant baked near a data boundary - # carries its own posterior uncertainty. No-op unless the approach - # installs the ensemble providers (see rule_param_margin_provider); - # this flag alone is enough for the ensemble to be built. - agent_plan_validation_rule_param_margin = False - # Necessity gate on captures: after a goal-reaching plan passes every - # other gate, re-run it once per step with that step removed. If the - # goal is still reached without a step, the plan is refused as - # REDUNDANT naming that step. A captured plan is an explanation of the - # goal, and a step whose absence changes nothing explains nothing: it - # is padding (a Wait on atoms that already hold, a press of a button - # the model says does nothing) that costs real episode steps and, when - # the model is wrong about the step, can break the plan for real. - # run_20260902_152811: a validated capture pressed three of four - # buttons and released one that was never on, for a goal its own - # model reached with two presses and a Wait. Costs one rollout per - # plan step, run in parallel with the other sweeps' workers. - agent_plan_validation_necessity = False - # Fork-parallel rollouts: the capture gate's repeat rollouts, its - # physics/rule-param margin sweeps, the belief probe's - # trials/physics_sweep modes, and the rollout-sysID objective (each - # candidate theta scores N trajectory segments) all run N + # Fork-parallel rollouts: the belief probe's trials/physics_sweep + # modes and the rollout-sysID objective (each candidate theta scores + # N trajectory segments) all run N # INDEPENDENT fresh-env rollouts; with a value W > 1, up to W run # concurrently as forked children (see # agent_sdk/parallel_rollouts.py). Verdict/fit semantics are @@ -2413,33 +2232,22 @@ class GlobalSettings: # enable in experiment configs sized to the job's CPU allocation # (e.g. 6 with --cpus-per-task=8). agent_validation_parallel_workers = 0 - # Agent bilevel explorer settings. Separate from the solve-path budget - # above because the explorer runs full backtracking while looking for - # the deepest subgoal-failure to truncate at. Denominated in - # option-model rollouts per search node: plain steps spend one per - # backtracking attempt (classic semantics); info-seeking steps spend - # the same budget pooling candidates (see refine_sketch). - # Active-experiment-design exploration: build the learned model's # parameter ensemble so the agent can rank candidate probes by the # ensemble's disagreement on a step's subgoal atoms - # (sim.suggest_probes) and the capture gate can sweep the rule-param - # margin. The agent decides what to run; the harness never moves - # its parameters. Off => the ensemble is built only when the - # rule-param margin gate asks for it. + # (sim.suggest_probes). The agent decides what to run; the harness + # never moves its parameters. Off => no ensemble is built. agent_explorer_info_seeking = False # Adaptive info-seeking: with this on (and agent_explorer_info_seeking # on), the proactive half of info-seeking - the probe-ranking # sim.suggest_probes result and the disagreement guidance - stays - # dormant until the - # capture gate has refused a plan as PARAM-SENSITIVE - # (ctx.param_sensitive_refusal_pending). The rule-param margin gate - # and its (Laplace) ensemble stay always on, so the FIRST refusal can - # still fire; only then does the agent start spending real steps to - # reduce the uncertainty the gate named. This removes the info-seeking - # step tax on easy levels no plan is ever refused on, while keeping - # the robustness on levels where a fragile plan is caught. Off => - # info-seeking is always active (the original behaviour). + # dormant until the sim.run physics sweep finds a plan whose success + # straddles the belief interval (ctx.param_sensitive_refusal_pending); + # only then does the agent start spending real steps to reduce that + # uncertainty. This removes the info-seeking step tax on easy levels + # while keeping the robustness on levels where a fragile plan is + # found. Off => info-seeking is always active (the original + # behaviour). agent_explorer_info_seeking_adaptive = False # Noise-aware probe value (docs/uncertainty/design.md, section # 3.5): under a declared observation-noise channel the ensemble @@ -2457,18 +2265,6 @@ class GlobalSettings: # Ensemble size used to estimate disagreement. 1 disables scoring # (every candidate scores 0) and reduces to first-feasible. agent_explorer_info_ensemble_size = 6 - # A plan the explore session validated through the capture gate - # (submit_plan: goal reached in - # agent_plan_validation_rollouts fresh belief rollouts) is executed - # verbatim as the episode's solve attempt with mental_model_solved= - # True. The cycle's remaining requests on that task still query the - # agent, which sees the certified plan among the plans already - # scheduled and is asked for a different certified plan (a second, - # independent test of the belief), resubmitting the same one only as - # a last resort; every certified attempt solving for real satisfies - # the train-driven early-stop rule. Off feeds the capture into the - # experiment search as seeds instead. - agent_explorer_execute_certified_plan = True # Per-parameter jitter as a fraction of the ParamSpec box width, for # the uniform-fallback ensemble only (see calibrated flag below). agent_explorer_info_perturb_frac = 0.15 @@ -2478,16 +2274,6 @@ class GlobalSettings: # fit runs). agent_explorer_info_calibrated_ensemble = True - # Code sim-learning parameter fitting settings. - # Persist the raw rollout-fit trajectories (states + actions per - # recorded episode) to /fit_data/ at every cycle-level - # fit. The fit data otherwise lives only in memory, which made the - # wrong fits of run_20260724_232411 (lateral_friction 1.0358 / - # 0.3236 vs true 0.5) impossible to replay offline: approximate - # re-execution from logged plans cannot reproduce mid-episode - # replans or the warm-env recording context, the very channel - # suspected of corrupting the fits. Cost: one small pickle per fit. - code_sim_learning_persist_fit_data = True # Truncate each rollout-fit trajectory once the scored features have # settled (physical_sysid.truncate_settled_tail): keep everything up # to the last observed motion plus a margin, drop the static tail. @@ -2542,9 +2328,9 @@ class GlobalSettings: # only for the ones whose posterior contracted below a fixed # fraction of the prior. A parameter that moved but whose posterior # stayed wide gets the dedicated verdict "wide posterior" - # (Verdict.WIDE): its most likely value is deployed, and the capture - # gate's physics-margin sweep and sim.run(physics_sweep=True) - # certify plans across its whole interval. A mixed sweep (some + # (Verdict.WIDE): its most likely value is deployed, and + # sim.run(physics_sweep=True) certifies plans across its whole + # interval. A mixed sweep (some # points pass, some fail) is reported as the interval straddling the # plan's success boundary, with the passing and failing ranges, and # arms adaptive info-seeking (the probe trigger). The fit report @@ -2589,7 +2375,7 @@ class GlobalSettings: # run_20260723_091108, 0.1414 vs 0.1 on run_20260708_213258). The # default 0.1 (~+-10% for log params) brackets the typical bias; # the consumers of the reported width (verdict contraction, the - # capture gate's physics-margin sigma points) inherit the floor. + # physics sweep's sigma points) inherit the floor. # 0 disables. code_sim_learning_rollout_min_posterior_width = 0.1 # Post-fit anchor-ablation backward elimination (False disables): @@ -2789,12 +2575,6 @@ class GlobalSettings: # such and its anchor (env-registry baseline) is kept instead of # the fitted value. 0 disables the screen. code_sim_learning_rollout_sensitivity_factor = 2.0 - # Cross-cycle consistency check on the final per-cycle fit: a param - # whose MAP moved more than this many combined posterior sigmas - # since the previous cycle's fit is flagged (and its "identified" - # verdict downgraded) - mutually-incompatible confident fits are - # the signature of an overconfident probe. 0 disables. - code_sim_learning_rollout_cross_cycle_sigma = 3.0 # Pooled-evidence arbitration of a cross-cycle conflict: when the # new fit is flagged (see above) but explains the fit's surviving # segments with an SSE at least this factor smaller than the @@ -2822,18 +2602,11 @@ class GlobalSettings: # per fit. code_sim_learning_warm_start_with_lm = True - # Sim-learning oracle flags (for ablation / debugging). - # When True, load GT residual rules instead of running agent synthesis. - # Parameters init_values are perturbed so the fit still has work to do. - agent_sim_learn_oracle_sim_program = False - # Relative scale for perturbing oracle parameter init_values before the fit. - agent_sim_learn_oracle_sim_param_noise_scale = 0.2 # Ablations A6+A7 combined ("no uncertainty"): when False, nothing # consumes a posterior over the model parameters. The physics-margin sigma - # points are never built (so the capture gate's physics margin and - # the probe's physics_sweep have nothing to sweep) and the - # rule-parameter ensemble stays empty (so the rule-param margin and - # the info-seeking disagreement score have nothing to score). Fits + # points are never built (so the probe's physics_sweep has nothing + # to sweep) and the rule-parameter ensemble stays empty (so the + # info-seeking disagreement score has nothing to score). Fits # still run; only their point estimates are used. agent_sim_learn_param_uncertainty = True # Ablation A4 ("no harness parameter fitting"): when True, no @@ -2850,49 +2623,12 @@ class GlobalSettings: # Program world model arm (agent_program_world_model, paper arm C4 in # the form of Pinductor): the belief over the program's hidden state # is a particle set of this size (drawn from the program's - # initial_latent; the capture gate re-rolls every submission under - # each particle), the score's distance kernel is + # initial_latent), the score's distance kernel is # exp(-distance / bandwidth) with distances in feature-std units, # and the score report shows this many worst transitions. agent_program_belief_particles = 6 agent_program_kernel_bandwidth = 0.2 agent_program_score_max_examples = 3 - # Ablation A2 ("zero-shot synthesis"): when True, the synthesis - # session runs even when no transition has been recorded, so the - # agent writes its artifacts from the task description, the scene - # and its own knowledge. Pair with no demos and - # num_online_learning_cycles 0 for one learn, one solve, done. - agent_sim_learn_zero_shot = False - # When True, use GT parameter values directly, skipping the fit. - # Also grants planning base sims the TRUE physical params (e.g. the true - # domino friction even when domino_planning_friction is set) — as if all - # param learning, rule-level and physical, had already succeeded. Task - # generation still reads domino_planning_friction for the - # differentiation filter, so the oracle, the no-learning baseline, and - # the sysID learner all see IDENTICAL tasks (and share the task cache: - # this agent_ flag is outside the cache key's - # domino_/pybullet_/skill_phase_ prefixes on purpose). - agent_sim_learn_oracle_sim_params = False - # When True, the agent learns PARAMETERIZED samplers - per-option - # (lifted-skill) functions that aim continuous option parameters at each - # sketch step's subgoal, instead of bilevel refinement drawing them - # uniformly from the option's box. The agent authors a versioned - # ``samplers.py`` (LEARNED_SAMPLERS keyed by option name) and tunes it - # with ``sim.samplers()``. Sampler learning rides along in - # the sim/predicate synthesis session when one runs - # (oracle_sim_program=False); when no synthesis session runs - # (oracle_sim_program=True) it gets a dedicated session of its own. - # The GROUND level of the sampler hierarchy needs no flag: a sketch - # step's ``~ [widths]`` region annotation compiles to a per-step - # GroundSampler that overrides the parameterized sampler for that step - # (ground > parameterized > uniform). - agent_sim_learn_parameterized_samplers = False - # When True (and parameterized_samplers is on), use ground-truth - # per-skill samplers from the env's GroundTruthSamplerFactory instead of - # having the agent learn them — if such samplers exist for the env; - # otherwise warn and fall back to synthesis. Mirrors - # agent_sim_learn_oracle_sim_program. - agent_sim_learn_oracle_samplers = False # Allowlist of env predicate names surfaced to the agent for # agent_sim_learning and its subclasses (e.g. diff --git a/predicators/structs.py b/predicators/structs.py index da9656d298..6ac2c671c1 100644 --- a/predicators/structs.py +++ b/predicators/structs.py @@ -554,7 +554,7 @@ def accepts_latent(self) -> bool: Such a predicate's truth can depend on belief-only state that real observations do not carry, so it may be unverifiable at - execution time (see the capture gate's latent-stripped probe). + execution time. """ return _classifier_accepts_latent(self._classifier) @@ -2118,7 +2118,6 @@ class LowLevelTrajectory: _train_task_idx: Optional[int] = field(default=None) _source_simulator_version: Optional[str] = field(default=None) _source_predicates_version: Optional[str] = field(default=None) - _source_samplers_version: Optional[str] = field(default=None) _env_reward: Optional[float] = field(default=None) _env_terminated: Optional[bool] = field(default=None) @@ -2163,12 +2162,6 @@ def source_predicates_version(self) -> Optional[str]: collected this trajectory, or ``None`` if not tracked.""" return self._source_predicates_version - @property - def source_samplers_version(self) -> Optional[str]: - """Snapshot tag of the per-skill samplers used to generate the plan - that collected this trajectory, or ``None`` if not tracked.""" - return self._source_samplers_version - @property def env_rejected(self) -> bool: """Whether the supervisor (the environment's evaluator) rejected the @@ -2635,8 +2628,7 @@ class InteractionRequest: # explorers); online learning treats ``False`` as not-solved for # early stopping even if real-env execution happens to reach the # goal, so a model that executes-but-mispredicts isn't certified as - # trained. See AgentModelBasedExplorer / - # run.online_learning.generate_interaction_results. + # trained. See run.online_learning.generate_interaction_results. mental_model_solved: Optional[bool] = None diff --git a/predicators/utils.py b/predicators/utils.py index cf670e652b..c5a871c504 100644 --- a/predicators/utils.py +++ b/predicators/utils.py @@ -1612,18 +1612,13 @@ def __str__(self) -> str: def real_episode_step_budget(phase: Optional[str]) -> int: - """Low-level steps a real episode of this ``phase`` may use. - - Explore (interaction-request) episodes are capped by - ``max_num_steps_interaction_request`` on top of the horizon; test - and any other episodes by ``horizon`` alone. The belief tools quote - this number to the agent so its plans are sized for the budget the - real executor enforces (a 1000-step explore cap once went unstated - while the tools quoted the 3000-step horizon, and half the bridge - experiments were cut mid-plan). + """Low-level steps a real episode may use: the horizon, in every agent + session phase. + + The belief tools quote this number to the agent so its plans are + sized for the budget the real executor enforces. """ - if phase == "explore": - return int(min(CFG.horizon, CFG.max_num_steps_interaction_request)) + del phase return int(CFG.horizon) @@ -1679,8 +1674,7 @@ def strip_latent_wait_targets(options: Sequence[_Option], negative one is trivially satisfied (the Wait ends at once). Both are removed from every Wait's ``memory`` in place; with no targets left the Wait falls back to any-atom-change termination. Returns a - description per dropped target for logging. Mirrors the capture's - execution-verifiability filter on subgoal annotations. + description per dropped target for logging. """ dropped: List[str] = [] for i, option in enumerate(options): diff --git a/scripts/cluster_utils.py b/scripts/cluster_utils.py index 5b982f0144..52ac17eef1 100644 --- a/scripts/cluster_utils.py +++ b/scripts/cluster_utils.py @@ -2,10 +2,11 @@ import copy import os +import re import shlex import subprocess from dataclasses import dataclass -from typing import Any, Dict, Iterator, List, Optional, Tuple +from typing import Any, Collection, Dict, Iterator, List, Optional, Tuple import yaml @@ -131,30 +132,36 @@ def _resolve_config(config_filepath: str) -> Dict[str, Any]: return merged -def _resolve_extends(section: Dict[str, Any]) -> Dict[str, Any]: - """Expand ``EXTENDS`` entries of an ENVS or APPROACHES section. +def parse_seed_range(text: str) -> Tuple[int, int]: + """Parse a seed range as (start, count): "3" is seed 3 alone, "2-4" is + seeds 2, 3 and 4.""" + match = re.fullmatch(r"(\d+)(?:-(\d+))?", text.strip()) + if match is None: + raise ValueError(f"seed range {text!r} is not N or N-M") + start = int(match.group(1)) + stop = int(match.group(2)) if match.group(2) is not None else start + if stop < start: + raise ValueError(f"seed range {text!r} ends before it starts") + return start, stop - start + 1 - An entry ``{EXTENDS: base, FLAGS: {...}}`` is the ``base`` entry of - the same section with the entry's own keys deep-merged on top, and - it is un-parked (``SKIP: False``) unless it says otherwise. This is - how a launcher gives a menu arm a round-specific experiment id: the - id is the entry's key, the definition stays in the menu. + +def _select(section: Dict[str, Any], keys: Optional[Collection[str]], + kind: str) -> Dict[str, Any]: + """The entries of an ENVS or APPROACHES section that a launch runs: + + the ``keys`` named on the command line, whether or not the config + parks them, else every entry the config does not SKIP. """ - resolved: Dict[str, Any] = {} - for key, entry in section.items(): - base_key = entry.get("EXTENDS") - if base_key is None: - resolved[key] = entry - continue - if base_key not in section: - raise ValueError(f"{key} EXTENDS unknown entry {base_key}") - if "EXTENDS" in section[base_key]: - raise ValueError(f"{key} EXTENDS {base_key}, which itself " - "EXTENDS another entry; extend menu entries") - derived = {k: v for k, v in entry.items() if k != "EXTENDS"} - merged = _deep_merge(section[base_key], {"SKIP": False}) - resolved[key] = _deep_merge(merged, derived) - return resolved + if keys is None: + return { + key: entry + for key, entry in section.items() if not entry.get("SKIP", False) + } + unknown = sorted(set(keys) - set(section)) + if unknown: + raise ValueError(f"unknown {kind} {unknown}; the config defines " + f"{sorted(section)}") + return {key: entry for key, entry in section.items() if key in keys} def parse_configs(config_filename: str) -> Iterator[Dict[str, Any]]: @@ -177,11 +184,25 @@ def parse_configs(config_filename: str) -> Iterator[Dict[str, Any]]: def generate_run_configs(config_filename: str, - batch_seeds: bool = False) -> Iterator[RunConfig]: - """Generate run configs from a (local path) config file.""" + batch_seeds: bool = False, + round_name: Optional[str] = None, + envs: Optional[Collection[str]] = None, + approaches: Optional[Collection[str]] = None, + seeds: Optional[Tuple[int, int]] = None, + require_round: bool = False) -> Iterator[RunConfig]: + """Generate run configs from a (local path) config file. + + The experiment id is ``-``, suffixed with + ``_`` when the launch names a round: ``round_name``, else the + config's ROUND key. ``envs`` and ``approaches`` pick entries by key + and ``seeds`` is a (start, count) pair that overrides START_SEED and + NUM_SEEDS. With ``require_round``, a continual-protocol run without + a round is an error: its runs auto-resume from their run folders, so + an unnamed relaunch would resume the previous launch. + """ for config in parse_configs(config_filename): - start_seed = config["START_SEED"] - num_seeds = config["NUM_SEEDS"] + start_seed, num_seeds = seeds or (config["START_SEED"], + config["NUM_SEEDS"]) args = config["ARGS"] flags = config["FLAGS"] if "USE_GPU" in config.keys(): @@ -196,20 +217,23 @@ def generate_run_configs(config_filename: str, train_refinement_estimator = config["TRAIN_REFINEMENT_ESTIMATOR"] else: train_refinement_estimator = False - approaches = _resolve_extends(config["APPROACHES"]) - envs = _resolve_extends(config["ENVS"]) + launch_round = round_name or config.get("ROUND") + if launch_round is not None and not re.fullmatch( + r"[A-Za-z0-9][A-Za-z0-9_]*", str(launch_round)): + raise ValueError(f"round {launch_round!r} must be letters, " + "digits and underscores") + suffix = f"_{launch_round}" if launch_round is not None else "" + selected_approaches = _select(config["APPROACHES"], approaches, + "approaches") + selected_envs = _select(config["ENVS"], envs, "envs") # Loop over approaches. - for approach_exp_id, approach_config in approaches.items(): - if approach_config.get("SKIP", False): - continue + for approach_exp_id, approach_config in selected_approaches.items(): approach = approach_config["NAME"] # Loop over envs. - for env_exp_id, env_config in envs.items(): - if env_config.get("SKIP", False): - continue + for env_exp_id, env_config in selected_envs.items(): env = env_config["NAME"] # Create the experiment ID, args, and flags. - experiment_id = f"{env_exp_id}-{approach_exp_id}" + experiment_id = f"{env_exp_id}-{approach_exp_id}{suffix}" run_args = list(args) if "ARGS" in approach_config: run_args.extend(approach_config["ARGS"]) @@ -220,6 +244,13 @@ def generate_run_configs(config_filename: str, run_flags.update(approach_config["FLAGS"]) if "FLAGS" in env_config: run_flags.update(env_config["FLAGS"]) + if (require_round and not suffix and + run_flags.get("experiment_protocol") == "continual"): + raise ValueError( + f"{experiment_id} runs the continual protocol, " + "whose runs auto-resume from their run folders: " + "name a round (--round or the config's ROUND) so " + "this launch cannot resume an earlier one") # Loop or batch over seeds. if batch_seeds: yield BatchSeedRunConfig(experiment_id, approach, env, diff --git a/scripts/configs/predicatorv3/random_actions_pybullet.yaml b/scripts/configs/ExoPredicator/random_actions_pybullet.yaml similarity index 96% rename from scripts/configs/predicatorv3/random_actions_pybullet.yaml rename to scripts/configs/ExoPredicator/random_actions_pybullet.yaml index 150d7fec43..5da4d1b9cb 100644 --- a/scripts/configs/predicatorv3/random_actions_pybullet.yaml +++ b/scripts/configs/ExoPredicator/random_actions_pybullet.yaml @@ -1,6 +1,6 @@ # Random actions agent on all PyBullet environments - generates test videos. # Usage: -# PYTHONPATH=. python scripts/local/launch_simp.py -c mara2/random_actions_pybullet.yaml +# PYTHONPATH=. python scripts/local/launch_simp.py -c ExoPredicator/random_actions_pybullet.yaml --- APPROACHES: random_actions_pybullet: diff --git a/scripts/configs/empiric/approaches.yaml b/scripts/configs/empiric/approaches.yaml new file mode 100644 index 0000000000..264cf3a4b7 --- /dev/null +++ b/scripts/configs/empiric/approaches.yaml @@ -0,0 +1,81 @@ +# The seven arms of the EMPIRIC benchmark. Each key is the approach half +# of the experiment id (-_). The shared flags in +# common.yaml give every arm the principled joint belief; No uncertainty +# turns it off. Arms gated on a deployable model on the test level set +# continual_require_model_on_test. +APPROACHES: + # EMPIRIC: the agent writes a model of the world's physics, the harness + # fits its parameters and keeps their uncertainty. + mb_opus: + NAME: agent_continual + FLAGS: + agent_sdk_model_name: claude-opus-5 + continual_require_model_on_test: true + # Model-free: the same agent with the model-side machinery switched + # off, so the two differ only in the model. + mf_opus: + NAME: agent_continual_model_free + FLAGS: &mf_flags + agent_sdk_model_name: claude-opus-5 + code_sim_learning_interval_belief: false + agent_explorer_info_seeking_noise_aware: false + code_sim_learning_rollout_noise_filter: false + code_sim_learning_carry_posterior: false + code_sim_learning_fit_evidence: false + continual_belief_frame: false + agent_model_repair: false + agent_planner_use_simulator: false + continual_uncertainty_decisions: false + # Model-free with the scene files an agentic real-to-sim agent would + # build from: the engine wrapper, the scene manifest and the URDF and + # mesh files as read-only references. Its prompt only asks it to solve + # the levels; nothing asks for a simulator or model. + mf_scene_package_opus: + NAME: agent_continual_model_free + FLAGS: + <<: *mf_flags + continual_provide_scene_package: true + # Standalone: the agent writes and owns a program world model, with no + # harness fitting. + standalone_opus: + NAME: agent_continual_program_world_model + FLAGS: + agent_sdk_model_name: claude-opus-5 + # Oracle dynamics: the ground-truth simulator stands in for the learned + # model, with no automatic execution gate or shadow audit (see + # docs/comparisons/oracle-repair-notes.md). + oracle_dynamics_opus: + NAME: agent_continual_oracle_dynamics + FLAGS: + agent_sdk_model_name: claude-opus-5 + continual_skill_preflight: false + continual_validation_audit: false + # No fitting: EMPIRIC with only its declared parameters and no harness + # parameter fitting (the agent may fit in its own sandbox code). + no_fitting_opus: + NAME: agent_continual_no_fitting + FLAGS: + agent_sdk_model_name: claude-opus-5 + continual_require_model_on_test: true + agent_sim_learn_declared_params_only: true + # No uncertainty: EMPIRIC fitting raw noisy observations, with no state + # smoothing, uncertainty margins, information seeking or joint belief. + no_uncertainty_opus: + NAME: agent_continual_no_uncertainty + FLAGS: + agent_sdk_model_name: claude-opus-5 + continual_require_model_on_test: true + code_sim_learning_rollout_noise_filter: false + continual_belief_frame: false + continual_uncertainty_decisions: false + agent_sim_learn_param_uncertainty: false + agent_explorer_info_seeking: false + agent_explorer_info_seeking_adaptive: false + agent_explorer_info_seeking_noise_aware: false + code_sim_learning_interval_belief: false + code_sim_learning_carry_posterior: false + belief_joint_draws: 0 + # Nothing about the observation noise is declared to this arm: no + # noise section, no [noise] line, and the harness fit models none + # of it. + continual_obs_noise_declared: false diff --git a/scripts/configs/empiric/benchmark.yaml b/scripts/configs/empiric/benchmark.yaml new file mode 100644 index 0000000000..2e03f5f637 --- /dev/null +++ b/scripts/configs/empiric/benchmark.yaml @@ -0,0 +1,29 @@ +# The EMPIRIC benchmark: the seven arms (approaches.yaml) on the five +# settings (envs.yaml), five seeds each, with the principled joint belief +# (common.yaml). +# +# Every launch names a round, which suffixes the experiment ids +# (-_, the run folders under continual_runs_dir), so a +# new round never auto-resumes an old one: +# +# python scripts/engaging/launch.py -c empiric/benchmark.yaml \ +# --round r2 --partition mit_preemptable --accounts b,c,d +# +# --envs, --approaches and --seeds launch a subset: +# +# python scripts/engaging/launch.py -c empiric/benchmark.yaml \ +# --round fan_fix_r1 --envs fan --approaches mb_opus --seeds 2-4 +# +# A launch worth keeping on record gets its own file next to this one, +# which includes this file and sets ROUND (and SKIP entries or seeds): +# +# includes: [benchmark.yaml] +# ROUND: fan_fix_r1 +# START_SEED: 2 +# NUM_SEEDS: 3 +includes: + - common.yaml + - envs.yaml + - approaches.yaml +START_SEED: 0 +NUM_SEEDS: 5 diff --git a/scripts/configs/predicatorv3/continual_common.yaml b/scripts/configs/empiric/common.yaml similarity index 57% rename from scripts/configs/predicatorv3/continual_common.yaml rename to scripts/configs/empiric/common.yaml index 2e29ed87c7..c014b67f38 100644 --- a/scripts/configs/predicatorv3/continual_common.yaml +++ b/scripts/configs/empiric/common.yaml @@ -1,32 +1,14 @@ -# Shared base of every continual-protocol launch (the agent plays levels -# of one environment, train then test, through the sandboxed skills). -# A launcher includes this file plus the menus envs/continual.yaml and -# approaches/continual.yaml, un-parks one env and the arms it compares, -# and gives each arm a round-specific experiment id with EXTENDS: -# -# includes: [continual_common.yaml, envs/continual.yaml, -# approaches/continual.yaml] -# ENVS: -# balloons: {SKIP: False} -# APPROACHES: -# mb_opus_compose_r2: {EXTENDS: mb_opus} -# mf_opus_compose_r2: {EXTENDS: mf_opus} -# -# The experiment id is "-" (the run directory under -# continual_runs_dir). Env FLAGS override approach FLAGS, which override -# these. Launch with -# python scripts/engaging/launch.py -c predicatorv3/.yaml --partition mit_preemptable --accounts b,c -START_SEED: 0 -NUM_SEEDS: 1 +# Arguments and flags every EMPIRIC run shares: the continual protocol +# (the agent plays the levels of one environment, train then test, +# through the sandboxed skills), observation noise, and the principled +# joint belief that every arm but No uncertainty carries. +# benchmark.yaml includes this file with envs.yaml and approaches.yaml. +# Env FLAGS override approach FLAGS, which override these. ARGS: - debug -- make_failure_videos -- make_test_videos -- make_interaction_videos - auto_resume FLAGS: skill_phase_use_motion_planning: true - max_num_steps_interaction_request: 500 pretrained_model_service_provider: openrouter llm_model_name: google/gemini-2.5-pro llm_openai_max_response_tokens: 1e6 @@ -57,12 +39,9 @@ FLAGS: partially_observable: true agent_sdk_max_agent_turns_per_iteration: 10000 agent_sdk_image_max_px: 900 - agent_solve_use_journal: true continual_max_idle_rounds: 5 - agent_bilevel_use_llm_initial_params: true agent_validation_parallel_workers: 6 code_sim_learning_rollout_min_posterior_width: 0.1 - agent_plan_validation_rule_param_margin: true agent_explorer_info_seeking: true agent_explorer_info_seeking_adaptive: true code_sim_learning_interval_belief: true @@ -78,8 +57,3 @@ FLAGS: continual_belief_frame: true agent_model_repair: true continual_uncertainty_decisions: true - num_online_learning_cycles: 5 - online_learning_early_stopping: true - online_learning_early_stopping_require_all_attempts: true - online_learning_early_stopping_skip_redundant_test: true - online_nsrt_learning_requests_per_cycle: 2 diff --git a/scripts/configs/empiric/envs.yaml b/scripts/configs/empiric/envs.yaml new file mode 100644 index 0000000000..0018744631 --- /dev/null +++ b/scripts/configs/empiric/envs.yaml @@ -0,0 +1,134 @@ +# The five benchmark settings. Each key is the env half of the experiment +# id (-_). Env FLAGS override the approach's and the +# common ones, so per-domain noise levels and level budgets live here. +ENVS: + # Balloons: composition test levels (every test level composes lifts + # the agent measured on training racks; the weakest-first release + # bursts) with the 25-step goal dwell. + balloons: + NAME: pybullet_balloons + FLAGS: + max_initial_demos: 0 + horizon: 1500 + sesame_check_expected_atoms: false + pybullet_birrt_path_subsample_ratio: 2 + balloons_require_jam_decoy: false + balloons_goal_dwell_steps: 25 + continual_obs_noise_position: 0.01 + continual_obs_noise_orientation: 0.02 + continual_obs_noise_scalar: 0.0 + continual_steps_per_level: 5000 + num_train_tasks: 2 + # Bridge: a three-span training row and a four-span test row, with the + # Sept 16 four-span repairs (rigid grasp, lift-first transit, a + # certificate that waits for the robot to withdraw). + bridge: + NAME: pybullet_bridge + FLAGS: + max_initial_demos: 0 + horizon: 3000 + skill_place_settle_preload_force: 3.0 + process_planning_heuristic_weight: 10.0 + process_planning_max_execution_replans: 3 + wait_option_max_steps: 120 + pybullet_birrt_contact_margin: -0.005 + pybullet_pin_held_weld_assemblies: true + continual_obs_noise_position: 0.005 + continual_obs_noise_orientation: 0.02 + continual_obs_noise_scalar: 0.0 + continual_steps_per_level: 10000 + num_train_tasks: 1 + bridge_train_span_blocks: 3 + bridge_test_span_blocks: 4 + pybullet_grasp_max_force: 10000.0 + bridge_lift_before_transit: true + bridge_goal_robot_clearance: 0.01 + # Boil: one jug in training, two jugs on the test level, with the 5 cm + # faucet tolerance. + boil: + NAME: pybullet_boil + FLAGS: + max_initial_demos: 0 + excluded_objects_in_state_str: switch + max_num_steps_option_rollout: 100 + horizon: 500 + boil_goal: simple + boil_require_jug_full_to_heatup: true + script_option_file_name: boil.txt + boil_water_fill_speed: 0.0015 + pybullet_birrt_path_subsample_ratio: 2 + boil_num_jugs_train: [1] + boil_num_jugs_test: [2] + boil_num_burner_train: [1] + boil_num_burner_test: [1] + boil_faucet_align_threshold: 0.05 + continual_obs_noise_position: 0.0125 + continual_obs_noise_orientation: 0.05 + continual_obs_noise_scalar: 0.07 + continual_steps_per_level: 5000 + num_train_tasks: 1 + # Fan: uniform training positions, and a test level that turns on the + # exposed, inertial and ramp transfers (a 3 mm ramp with a longer + # landing). + fan: + NAME: pybullet_fan + FLAGS: + max_initial_demos: 0 + excluded_objects_in_state_str: switch + terminate_on_goal_reached: true + process_planning_heuristic_weight: 10.0 + horizon: 500 + pybullet_birrt_path_subsample_ratio: 2 + process_planning_max_execution_replans: 3 + fan_train_num_pos_x: 3 + fan_train_num_pos_y: 3 + fan_exposed_transfer: true + fan_inertial_transfer: true + fan_ramp_transfer: true + fan_ramp_rise: 0.003 + fan_ramp_landing_extension: 0.10 + fan_train_num_walls_per_task: [0] + fan_test_num_walls_per_task: [0] + fan_test_num_pos_x: 3 + fan_test_num_pos_y: 3 + fan_train_task_generation: uniform + fan_test_task_generation: uniform + continual_obs_noise_position: 0.005 + continual_obs_noise_orientation: 0.02 + continual_obs_noise_scalar: 0.0 + continual_steps_per_level: 5000 + num_train_tasks: 1 + # Domino: min-block friction tasks with a turn on the test level, true + # friction above the planner's belief. + domino_high_friction_turn: + NAME: pybullet_domino + FLAGS: + max_initial_demos: 0 + excluded_objects_in_state_str: loc,rot,angle,direction + horizon: 500 + domino_initialize_at_finished_state: false + domino_use_domino_blocks_as_target: true + domino_use_continuous_place: true + process_planning_heuristic_weight: 2.0 + domino_has_glued_dominos: false + keep_failed_demos: true + predicate_invent_invent_derived_predicates: true + pybullet_birrt_extend_num_interp: 20 + pybullet_birrt_path_subsample_ratio: 2 + domino_min_block_tasks: true + domino_true_friction: 0.5 + domino_planning_friction: 0.1 + domino_min_block_span_lo: 0.29 + domino_min_block_span_hi: 0.31 + domino_min_block_turn_entry_lo: 0.21 + domino_min_block_turn_entry_hi: 0.24 + domino_min_block_turn_exit_lo: 0.17 + domino_min_block_turn_exit_hi: 0.2 + domino_min_block_num_blues: 4 + domino_block_cost: 0.1 + domino_test_turn_ratio: 1.0 + continual_obs_noise_position: 0.01 + continual_obs_noise_orientation: 0.04 + continual_obs_noise_scalar: 0.0 + continual_steps_per_level: 5000 + num_train_tasks: 1 diff --git a/scripts/configs/predicatorv3/approaches/all.yaml b/scripts/configs/predicatorv3/approaches/all.yaml deleted file mode 100644 index 97b0f0748c..0000000000 --- a/scripts/configs/predicatorv3/approaches/all.yaml +++ /dev/null @@ -1,388 +0,0 @@ -# Canonical arm menu (paper ids on each entry, C7 dropped); all parked, the -# exp_*.yaml launchers un-skip. List-valued FLAGS live only here (merges concat). -APPROACHES: - - # ---- OURS ---- - - # C1: PO, no GT; learns predicates, hybrid sim and params, with the - # physics-margin gate. Its FLAGS block is the anchor the ABLATIONS merge. - sim_predicator: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: &ours_flags - demonstrator: "oracle_process_planning" - explorer: "agent_model_based" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - agent_sim_learn_kept_predicates_names: ["Holding"] - partially_observable: True - agent_explorer_info_seeking: True - execution_monitor: "subgoal_annotations" - # Open-loop execution: suffix replanning never recovered a bridge - # divergence (0/4, 2026-09-01). 0 also disarms the divergence monitor. - agent_bilevel_max_execution_replans: 0 - agent_bilevel_use_llm_initial_params: True # LLM proposes params - agent_sdk_max_agent_turns_per_iteration: 200 - agent_sdk_image_max_px: 900 - agent_solve_max_attempts: 1 - agent_solve_attempt_wall_clock: 2700 - agent_solve_fresh_context: True - agent_solve_use_journal: True - agent_plan_validation_physics_margin: True - agent_plan_validation_rule_param_margin: True - agent_explorer_info_mcmc_steps: 0 - agent_validation_parallel_workers: 6 - bilevel_plan_without_sim: True # for the demonstrator - code_sim_learning_rollout_min_posterior_width: 0.1 - - # C1 in policy mode: the solve deliverable is a per-task closed-loop - # policy.py that owns recovery, so the replan machinery is off. - sim_predicator_policy: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - explorer: "agent_model_based" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - agent_sim_learn_kept_predicates_names: ["Holding"] - partially_observable: True - agent_explorer_info_seeking: True - execution_monitor: "subgoal_annotations" - agent_bilevel_max_execution_replans: 0 - agent_solve_policy_mode: True - agent_bilevel_use_llm_initial_params: True # LLM proposes params - agent_sdk_max_agent_turns_per_iteration: 200 - agent_sdk_image_max_px: 900 - agent_solve_max_attempts: 1 - agent_solve_attempt_wall_clock: 2700 - agent_solve_fresh_context: True - agent_solve_use_journal: True - agent_plan_validation_physics_margin: True - agent_plan_validation_rule_param_margin: True - agent_explorer_info_mcmc_steps: 0 - agent_validation_parallel_workers: 6 - bilevel_plan_without_sim: True # for the demonstrator - code_sim_learning_rollout_min_posterior_width: 0.1 - - # ---- ORACLES (given a GT world model, ordered from most GT handed in to least) ---- - - # GT monolithic sim + GT predicates; the agent plans against it with - # submit_plan / sim.refine and the sketch scaffolding. - agent_oracle_mono_sim: - NAME: "agent_model_based" - SKIP: True - FLAGS: - explorer: "agent_model_free" - demonstrator: "oracle_process_planning" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - option_model_use_gui: True - agent_bilevel_log_state: False - agent_bilevel_plan_sketch_file: "tests/approaches/test_data/boil_plan_sketch.txt" - # U1 (upper bound): GT hybrid sim + true physical params, as if learning had - # succeeded; same solve budget as C1 (5 time-boxed attempts + journal). - agent_oracle_hybrid_sim: - NAME: "agent_sim_learning" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - bilevel_plan_without_sim: True # for the demonstrator - explorer: "agent_model_based" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - option_model_use_gui: False - agent_bilevel_log_state: False - agent_sim_learn_oracle_sim_program: True - agent_sim_learn_oracle_sim_params: True - agent_bilevel_use_llm_initial_params: True - num_online_learning_cycles: 0 - execution_monitor: "subgoal_annotations" - agent_bilevel_max_execution_replans: 2 - agent_sdk_max_agent_turns_per_iteration: 200 - agent_sdk_image_max_px: 900 - agent_solve_max_attempts: 1 - agent_solve_attempt_wall_clock: 2700 - agent_solve_fresh_context: True - agent_solve_use_journal: True - agent_plan_validation_rule_param_margin: True - agent_explorer_info_mcmc_steps: 0 - agent_validation_parallel_workers: 6 - # GT hybrid sim + GT predicates; learn only the params. - agent_param_learning: - NAME: "agent_sim_learning" - SKIP: True - FLAGS: - explorer: "agent_model_based" - demonstrator: "oracle_process_planning" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - option_model_use_gui: True - agent_bilevel_log_state: False - agent_bilevel_plan_sketch_file: "tests/approaches/test_data/boil_plan_sketch.txt" - agent_sim_learn_oracle_sim_program: True - agent_sim_learn_oracle_sim_params: False - agent_sim_learn_oracle_sim_param_noise_scale: 1.0 # 0.8 gives a satisficing plan - code_sim_learning_num_mcmc_steps: 0 - # GT predicates; learn the hybrid sim and its params. - agent_sim_learning: - NAME: "agent_sim_learning" - SKIP: True - FLAGS: - explorer: "agent_model_based" - demonstrator: "oracle_process_planning" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - option_model_use_gui: True - agent_bilevel_log_state: False - agent_bilevel_plan_sketch_file: "tests/approaches/test_data/boil_plan_sketch.txt" - agent_sim_learn_oracle_sim_program: False - agent_sim_learn_oracle_sim_params: False - code_sim_learning_num_mcmc_steps: 0 - # Fully-observable OURS: full state, no GT models. - agent_fo_predicate_invention: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - explorer: "agent_model_based" - demonstrator: "oracle_process_planning" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - option_model_use_gui: False - agent_bilevel_log_state: False - online_learning_early_stopping: True - agent_sim_learn_oracle_sim_program: False - agent_sim_learn_oracle_sim_params: False - code_sim_learning_num_mcmc_steps: 0 - agent_sim_learn_kept_predicates_names: ["Holding"] - - # ---- BASELINES (lower bounds / alternative learners) ---- - - # C2: the same agent with no simulator (agent_planner_use_simulator False); - # memory carries across trials via the session, notes.md and the journal. - # Same observation as OURS (partially_observable): the tier-1 C2 runs - # of 2026-09-03 predate this line and saw the hidden features. - agent_model_free_planning: - NAME: "agent_model_free" - SKIP: True - FLAGS: - explorer: "agent_model_free" - demonstrator: "oracle_process_planning" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - partially_observable: True - agent_planner_use_simulator: False - agent_planner_use_scratchpad: True - agent_solve_use_journal: True - bilevel_plan_without_sim: True # for the demonstrator - agent_solve_max_attempts: 1 - agent_solve_attempt_wall_clock: 2700 - agent_sdk_max_agent_turns_per_iteration: 200 - agent_sdk_image_max_px: 900 - # C6: NSRT learning with oracle predicates, samplers and full observability - # (flags from configs/ExoPredicator/causal_predicator_baselines.yaml). - operator_learning: - NAME: "online_nsrt_learning" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - bilevel_plan_without_sim: True - explorer: "exploit_planning" - terminate_on_goal_reached_and_option_terminated: True - bilevel_planning_explorer_enumerate_plans: True - exploit_bilevel_planning_explorer_fallback_explorer: "RandomNSRTs" - online_learning_assert_no_exclude_pred: False - disable_harmlessness_check: True - sampler_learner: "oracle" - option_learner: "no_learning" - clustering_learner_check_effect_equality: False - # C3: OURS' loop, but the model is a text document (world_model.md) the - # agent writes and reasons over; no code, no simulator, same inputs as OURS. - nl_world_model: - NAME: "agent_nl_world_model" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - explorer: "agent_model_free" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - agent_planner_use_simulator: False - agent_planner_use_scratchpad: True - agent_sim_learn_kept_predicates_names: ["Holding"] - partially_observable: True - agent_sdk_max_agent_turns_per_iteration: 200 - agent_sdk_image_max_px: 900 - agent_solve_max_attempts: 1 - agent_solve_attempt_wall_clock: 2700 - agent_solve_fresh_context: True - agent_solve_use_journal: True - bilevel_plan_without_sim: True # for the demonstrator - # C4 (Pinductor form): OURS' loop, but the model is an option-level program - # with a hidden state, scored by sim.score; belief particles drive the gate. - code_world_model: - NAME: "agent_program_world_model" - SKIP: True - FLAGS: - <<: *ours_flags - agent_explorer_info_seeking: False - agent_plan_validation_physics_margin: False - # C5: a GNN transition model over the object features (history-conditioned), - # same interaction budget as the agent arms, planned by shooting. - gnn_dynamics: - NAME: "gnn_dynamics_shooting" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - bilevel_plan_without_sim: True - explorer: "random_options" - terminate_on_goal_reached_and_option_terminated: True - partially_observable: True # same observation as OURS - gnn_num_epochs: 5000 - gnn_use_validation_set: True - gnn_do_normalization: True - gnn_dynamics_history_len: 2 - # C8: MAPLE-Q with oracle operators and samplers, same interaction budget as - # the agent arms (ExoPredicator flags minus its data sizing). - maple_q: - NAME: "maple_q_with_process" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - bilevel_plan_without_sim: True - explorer: "maple_q" - strips_learner: "oracle" - sampler_learner: "oracle" - only_learn_exogenous_processes: True - online_learning_assert_no_exclude_pred: False - maple_q_same_hla_option_param_space: False - mlp_regressor_max_itr: 640000 - active_sampler_learning_batch_size: 512 - # Oracle residual program with LLM-guessed params, no learning (not a paper arm). - agent_base_sim_no_learning: - NAME: "agent_sim_learning" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - bilevel_plan_without_sim: True # for the demonstrator - explorer: "agent_model_based" - terminate_on_goal_reached_and_option_terminated: True - agent_sdk_use_local_sandbox: True - option_model_terminate_on_repeat: False - option_model_use_gui: False - agent_bilevel_log_state: False - agent_sim_learn_oracle_sim_program: True - agent_sim_learn_oracle_sim_params: False - agent_bilevel_use_llm_initial_params: True - num_online_learning_cycles: 0 - execution_monitor: "subgoal_annotations" - agent_bilevel_max_execution_replans: 2 - - # ---- ABLATIONS (OURS minus one component, paper A1-A8; merge *ours_flags) ---- - - # A1: no learning; plans are made, checked and run on the base engine. - # The sweeps measure it as C1's pre-loop test; parked here for standalone reruns. - sim_predicator_no_learning: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - num_online_learning_cycles: 0 - # A2: one data-free synthesis session, then solve (no online learning). - sim_predicator_zero_shot: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - agent_sim_learn_zero_shot: True - num_online_learning_cycles: 0 - # A3: scripted random-options exploration instead of the agent explorer. - sim_predicator_undirected_explore: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - explorer: "random_options" - agent_explorer_info_seeking: False - # A4: no parameter fitting; declared inits are the estimate, [lo, hi] the - # interval the margins and the ensemble sample from. - sim_predicator_no_param_fit: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - agent_sim_learn_declared_params_only: True - # A5: execute the first goal-reaching capture (1 rollout, no margins). - sim_predicator_no_validation: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - agent_plan_validation_rollouts: 1 - agent_plan_validation_rollouts_after_flaky: 1 - agent_plan_validation_physics_margin: False - agent_plan_validation_rule_param_margin: False - # A6: experiments under the point estimate (no ensemble-disagreement - # probe scoring); the validation gate keeps its margins. - sim_predicator_explore_no_disagreement: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - agent_explorer_info_seeking: False - # A7: validation under the point estimate (no perturbed rollouts, no - # ensemble members); experiments keep the ensemble. - sim_predicator_validation_no_uncertainty: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - agent_plan_validation_physics_margin: False - agent_plan_validation_rule_param_margin: False - # A6+A7 combined: point estimates only; every consumer of a posterior is off. - sim_predicator_no_uncertainty: - NAME: "agent_sim_predicate_invention" - SKIP: True - FLAGS: - <<: *ours_flags - agent_sim_learn_param_uncertainty: False - agent_plan_validation_physics_margin: False - agent_plan_validation_rule_param_margin: False - agent_explorer_info_seeking: False - # A8: no invented predicates (the non-inventing class, Holding + goal_nl). - sim_predicator_no_predicates: - NAME: "agent_sim_learning" - SKIP: True - FLAGS: - <<: *ours_flags - - # ---- DEMONSTRATOR / SCRIPTED (the demo source every agent arm consumes) ---- - - # Oracle process planning + bilevel refinement with GT models; oracle.yaml - # runs it standalone. - oracle: - NAME: "oracle_process_planning" - SKIP: True - FLAGS: - demonstrator: "oracle_process_planning" - terminate_on_goal_reached_and_option_terminated: True - sesame_check_expected_atoms: False - bilevel_plan_without_sim: True - # Scripted human-in-the-loop option control (manual baseline). - human_interaction: - NAME: "human_interaction" - SKIP: True - FLAGS: - human_option_control_approach_use_scripted_option: True - human_option_control_approach_use_all_options: True - scripted_option_dir: "scripted_option_policies" - skill_phase_use_motion_planning: True - terminate_on_goal_reached_and_option_terminated: True diff --git a/scripts/configs/predicatorv3/approaches/continual.yaml b/scripts/configs/predicatorv3/approaches/continual.yaml deleted file mode 100644 index b7664ecaac..0000000000 --- a/scripts/configs/predicatorv3/approaches/continual.yaml +++ /dev/null @@ -1,203 +0,0 @@ -# Continual-protocol arm menu. Every entry is parked (SKIP); a launcher -# un-parks arms, usually under a round-specific key with EXTENDS (see -# continual_common.yaml). MB arms are gated on a deployable model on the -# test level (continual_require_model_on_test). MF arms switch off the -# model-side machinery so the two differ only in the model. The skill -# preflight (continual_skill_preflight) is off by default since Sept 18, -# 2026; a launcher that wants the Sept 17 Sonnet setting sets it to true. -APPROACHES: - from_assets_opus: - NAME: agent_continual_from_assets - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - continual_require_model_on_test: true - agent_sim_learn_declared_params_only: false - continual_uncertainty_decisions: true - mb_opus: - NAME: agent_continual - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - continual_require_model_on_test: true - mb_sonnet: - NAME: agent_continual - SKIP: true - FLAGS: - agent_sdk_model_name: claude-sonnet-5 - continual_require_model_on_test: true - # EMPIRIC with what the agentic real-to-sim arm receives (Sept 18, - # 2026): the engine wrapper, the scene manifest and the URDF and mesh - # files, plus the twin's own core module where the domain declares one - # (Fan and Balloons). The domain twin still backs the model. - mb_scene_package_opus: - NAME: agent_continual - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - continual_require_model_on_test: true - agent_sim_provide_base_sim_source: true - continual_provide_scene_package: true - mf_opus: - NAME: agent_continual_model_free - SKIP: true - FLAGS: &mf_flags - agent_sdk_model_name: claude-opus-5 - code_sim_learning_interval_belief: false - agent_explorer_info_seeking_noise_aware: false - code_sim_learning_rollout_noise_filter: false - code_sim_learning_carry_posterior: false - code_sim_learning_fit_evidence: false - continual_belief_frame: false - agent_model_repair: false - agent_planner_use_simulator: false - continual_uncertainty_decisions: false - mf_sonnet: - NAME: agent_continual_model_free - SKIP: true - FLAGS: - <<: *mf_flags - agent_sdk_model_name: claude-sonnet-5 - # The direct agent with the scene files the agentic real-to-sim arm - # builds from (Sept 19, 2026): the engine wrapper, the scene manifest - # and the URDF and mesh files as read-only references. Its prompt only - # asks it to solve the levels; nothing asks for a simulator or model. - mf_scene_package_opus: - NAME: agent_continual_model_free - SKIP: true - FLAGS: - <<: *mf_flags - continual_provide_scene_package: true - # ---- The other six arms of the eight-agent comparison - # (continual_eight_agent_noisy_sweep.yaml). Their approach classes - # came over from the continual-comparisons line in the Sept 17 merge - # (agent_continual_program_approach.py, agent_continual_frozen_approach.py, - # agent_continual_ablation_approach.py). Ids follow the mb_/mf_ - # pattern: _. - # Standalone: the agent writes and owns a program world model, no harness - # fitting. - standalone_opus: - NAME: agent_continual_program_world_model - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - standalone_sonnet: - NAME: agent_continual_program_world_model - SKIP: true - FLAGS: - agent_sdk_model_name: claude-sonnet-5 - # Oracle dynamics: the ground-truth simulator stands in for the learned - # model. Use the repaired r2 runtime with no automatic execution gate - # or shadow audit. See docs/comparisons/oracle-repair-notes.md. - oracle_dynamics_opus: - NAME: agent_continual_oracle_dynamics - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - continual_skill_preflight: false - continual_validation_audit: false - oracle_dynamics_sonnet: - NAME: agent_continual_oracle_dynamics - SKIP: true - FLAGS: - agent_sdk_model_name: claude-sonnet-5 - continual_skill_preflight: false - continual_validation_audit: false - # Scene-only: the exact scene twin with corrected base calibration and - # no mechanism code, frozen for the run (formerly "oracle scene"). - scene_only_opus: - NAME: agent_continual_scene_only - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - scene_only_sonnet: - NAME: agent_continual_scene_only - SKIP: true - FLAGS: - agent_sdk_model_name: claude-sonnet-5 - # Zero shot: no training level, the test level is played cold. - zero_shot_opus: - NAME: agent_continual_zero_shot - SKIP: true - FLAGS: - agent_sdk_model_name: claude-opus-5 - zero_shot_sonnet: - NAME: agent_continual_zero_shot - SKIP: true - FLAGS: - agent_sdk_model_name: claude-sonnet-5 - # No fitting: the MB agent with only its declared parameters, no - # harness parameter fitting (the agent may fit in its own sandbox - # code). Gated like the MB arms. - no_fitting_opus: - NAME: agent_continual_no_fitting - SKIP: true - FLAGS: &no_fitting_flags - agent_sdk_model_name: claude-opus-5 - continual_require_model_on_test: true - agent_sim_learn_declared_params_only: true - no_fitting_sonnet: - NAME: agent_continual_no_fitting - SKIP: true - FLAGS: - <<: *no_fitting_flags - agent_sdk_model_name: claude-sonnet-5 - # No uncertainty handling: fit raw noisy observations, with no state - # smoothing, uncertainty margins or information seeking. Gated like MB. - no_uncertainty_opus: - NAME: agent_continual_no_uncertainty - SKIP: true - FLAGS: &no_uncertainty_flags - agent_sdk_model_name: claude-opus-5 - continual_require_model_on_test: true - code_sim_learning_rollout_noise_filter: false - continual_belief_frame: false - continual_uncertainty_decisions: false - agent_sim_learn_param_uncertainty: false - agent_plan_validation_rule_param_margin: false - agent_plan_validation_physics_margin: false - agent_explorer_info_seeking: false - agent_explorer_info_seeking_adaptive: false - agent_explorer_info_seeking_noise_aware: false - code_sim_learning_interval_belief: false - code_sim_learning_carry_posterior: false - belief_joint_draws: 0 - # Nothing about the observation noise is declared to this arm: no - # noise section, no [noise] line, and the harness fit models none - # of it (Sept 18, 2026). - continual_obs_noise_declared: false - no_uncertainty_sonnet: - NAME: agent_continual_no_uncertainty - SKIP: true - FLAGS: - <<: *no_uncertainty_flags - agent_sdk_model_name: claude-sonnet-5 - # Agentic real-to-sim baseline (Sept 18, 2026): no domain twin. The - # agent gets the generic PyBulletEnv, a domain-agnostic SceneBase, the - # scene manifest and the asset files, and builds its own simulator; - # the harness fits nothing and runs no uncertainty machinery. Gated - # like the MB arms. Runs with the same skill library as the other - # arms (composite by default); set skill_library: primitive for the - # robot-stack variant. - real_to_sim_opus: - NAME: agent_continual_real_to_sim - SKIP: true - FLAGS: &real_to_sim_flags - agent_sdk_model_name: claude-opus-5 - continual_require_model_on_test: true - agent_sim_learn_declared_params_only: true - continual_uncertainty_decisions: false - agent_sim_learn_param_uncertainty: false - agent_plan_validation_rule_param_margin: false - agent_plan_validation_physics_margin: false - agent_explorer_info_seeking: false - agent_explorer_info_seeking_adaptive: false - agent_explorer_info_seeking_noise_aware: false - code_sim_learning_interval_belief: false - code_sim_learning_carry_posterior: false - real_to_sim_sonnet: - NAME: agent_continual_real_to_sim - SKIP: true - FLAGS: - <<: *real_to_sim_flags - agent_sdk_model_name: claude-sonnet-5 diff --git a/scripts/configs/predicatorv3/common.yaml b/scripts/configs/predicatorv3/common.yaml deleted file mode 100644 index 7e130607fd..0000000000 --- a/scripts/configs/predicatorv3/common.yaml +++ /dev/null @@ -1,35 +0,0 @@ -ARGS: - - "debug" - # - "use_gui" - - "make_failure_videos" - - "make_test_videos" - - "make_interaction_videos" - # - "make_demo_videos" - # - "make_demo_images" # support images - # - "make_failure_images" # query images - # - "make_test_images" # query images - # - "save_atoms" -FLAGS: - num_online_learning_cycles: 5 - online_learning_early_stopping: True - online_learning_early_stopping_require_all_attempts: True - online_learning_early_stopping_skip_redundant_test: True - online_nsrt_learning_requests_per_cycle: 2 - skill_phase_use_motion_planning: True - max_num_steps_interaction_request: 500 - pretrained_model_service_provider: "openrouter" - llm_model_name: "google/gemini-2.5-pro" - llm_openai_max_response_tokens: 1e6 - terminate_on_goal_reached: False - pybullet_ik_validate: False - num_train_tasks: 1 - num_test_tasks: 1 - video_fps: 20 - pybullet_camera_height: 900 - pybullet_camera_width: 900 - planning_filter_unreachable_nsrt: False - timeout: 600 - log: 'logs/' - no_repeated_arguments_in_grounding: True -START_SEED: 0 -NUM_SEEDS: 1 \ No newline at end of file diff --git a/scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml b/scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml deleted file mode 100644 index 10ce8bc357..0000000000 --- a/scripts/configs/predicatorv3/continual_benchmark_five_seeds.yaml +++ /dev/null @@ -1,38 +0,0 @@ -# Extend the paper's six comparison arms with seeds 3 and 4. -# Freeze from f2ed37aef: repaired runtime, original benchmark geometry. -# Do not use the later Balloons layout change for this cohort. -# Launch with --partition mit_preemptable --requeue --accounts a,b,c,d. -# Account dat (the user's "e") is backup only, outside normal rotation. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 3 -NUM_SEEDS: 2 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - oracle_dynamics_opus_benchmark_r2: - EXTENDS: oracle_dynamics_opus - mf_opus_benchmark_r2: - EXTENDS: mf_opus - mf_scene_package_opus_benchmark_r1: - EXTENDS: mf_scene_package_opus - standalone_opus_benchmark_r2: - EXTENDS: standalone_opus - no_fitting_opus_benchmark_r1: - EXTENDS: no_fitting_opus - no_uncertainty_opus_benchmark_r2: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml b/scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml deleted file mode 100644 index df14befdd2..0000000000 --- a/scripts/configs/predicatorv3/continual_direct_scene_files_benchmark_r1.yaml +++ /dev/null @@ -1,28 +0,0 @@ -# The direct agent with the scene files (Sept 19, 2026): the agentic -# real-to-sim arm's engine wrapper, scene manifest and URDF and mesh -# files as read-only references, and a prompt that only asks it to solve -# the levels (mf_scene_package_opus). Opus, three seeds each: 5 domains x -# 3 seeds = 15 runs, so the logs land under -# logs/agent_continual_model_free/-mf_scene_package_opus_benchmark_r1/seed. -# -# python scripts/engaging/launch.py -c predicatorv3/continual_direct_scene_files_benchmark_r1.yaml --partition mit_preemptable,mit_normal --requeue --accounts b,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - mf_scene_package_opus_benchmark_r1: - EXTENDS: mf_scene_package_opus diff --git a/scripts/configs/predicatorv3/continual_eight_agent_noisy_sweep.yaml b/scripts/configs/predicatorv3/continual_eight_agent_noisy_sweep.yaml deleted file mode 100644 index 0aad4ddd63..0000000000 --- a/scripts/configs/predicatorv3/continual_eight_agent_noisy_sweep.yaml +++ /dev/null @@ -1,55 +0,0 @@ -# The eight-agent benchmark sweep (Sept 18, 2026): the two main agents, -# model-based (EMPIRIC, agent_continual) and model-free (the direct -# agent, agent_continual_model_free), plus the six ablation arms -# (standalone program world model, oracle dynamics, scene-only, zero -# shot, no harness fitting, no explicit uncertainty), all on Opus, on -# the five benchmark test settings, three seeds each: 8 approaches x 5 -# domains x 3 seeds = 120 runs. Envs and arms come from the menus -# (envs/continual.yaml, approaches/continual.yaml): Balloons composition -# test levels with the 25-step dwell, Bridge four-span test row, Boil -# two-jug test level, Fan maze test level, Domino high-friction turn. -# The skill preflight is off (the runtime default). The round keys keep -# these logs apart from the Sept 17 rounds -# (logs//-/seed). -# -# The Sept 12-14 version of this sweep ran the same eight arms on the -# Sept 14 settings (three-span Bridge, uniform Fan, one-jug Boil, the -# original Balloons distribution) under runtime 091d8c5db. -# -# Forty Slurm arrays of three seeds; the Opus arms saturate one Claude -# account's session limit in about an hour, so spread the accounts: -# python scripts/engaging/launch.py -c predicatorv3/continual_eight_agent_noisy_sweep.yaml --partition mit_preemptable,mit_normal --requeue --accounts b,c -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - mb_opus_benchmark_r1: - EXTENDS: mb_opus - mf_opus_benchmark_r1: - EXTENDS: mf_opus - standalone_opus_benchmark_r1: - EXTENDS: standalone_opus - oracle_dynamics_opus_benchmark_r1: - EXTENDS: oracle_dynamics_opus - scene_only_opus_benchmark_r1: - EXTENDS: scene_only_opus - zero_shot_opus_benchmark_r1: - EXTENDS: zero_shot_opus - no_fitting_opus_benchmark_r1: - EXTENDS: no_fitting_opus - no_uncertainty_opus_benchmark_r1: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml b/scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml deleted file mode 100644 index e9985a865a..0000000000 --- a/scripts/configs/predicatorv3/continual_empiric_benchmark_r2.yaml +++ /dev/null @@ -1,26 +0,0 @@ -# Two prospective seeds per domain, with bounded nonblocking validation. -# Freeze this runtime before submitting; never resume historical EMPIRIC. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 3 -NUM_SEEDS: 2 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - mb_opus_benchmark_r2: - EXTENDS: mb_opus - FLAGS: - continual_skill_preflight: false - continual_validation_audit: true - continual_validation_audit_seconds: 600.0 diff --git a/scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml b/scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml deleted file mode 100644 index 04111f0f08..0000000000 --- a/scripts/configs/predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml +++ /dev/null @@ -1,29 +0,0 @@ -# EMPIRIC rerun on the benchmark runtime (Sept 18, 2026), with what the -# agentic real-to-sim arm receives: the engine wrapper, the scene -# manifest and the asset files, plus the twin's core module on Fan and -# Balloons (mb_scene_package_opus). Opus, three seeds each: 5 domains x -# 3 seeds = 15 runs, on the same menus as the other benchmark arms, so -# the logs land under -# logs/agent_continual/-mb_scene_package_opus_benchmark_r1/seed. -# -# python scripts/engaging/launch.py -c predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml --partition mit_preemptable,mit_normal --requeue --accounts b,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - mb_scene_package_opus_benchmark_r1: - EXTENDS: mb_scene_package_opus diff --git a/scripts/configs/predicatorv3/continual_fan_inertial_baselines_r1.yaml b/scripts/configs/predicatorv3/continual_fan_inertial_baselines_r1.yaml deleted file mode 100644 index e51f4ff1ef..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_inertial_baselines_r1.yaml +++ /dev/null @@ -1,20 +0,0 @@ -# Five additional arms on the original frozen, no-ramp inertial domain. -includes: - - /home/ycliang/predicators/logs/fan-inertial-baselines-runtime-20260921/scripts/configs/predicatorv3/continual_fan_inertial_pilot_r1.yaml -START_SEED: 0 -NUM_SEEDS: 5 -APPROACHES: - mb_opus_inertial_pilot_r1: - SKIP: true - mf_opus_inertial_pilot_r1: - SKIP: true - oracle_dynamics_opus_inertial_r1: - EXTENDS: oracle_dynamics_opus - mf_scene_package_opus_inertial_r1: - EXTENDS: mf_scene_package_opus - standalone_opus_inertial_r1: - EXTENDS: standalone_opus - no_fitting_opus_inertial_r1: - EXTENDS: no_fitting_opus - no_uncertainty_opus_inertial_r1: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_fan_inertial_confirmation_r1.yaml b/scripts/configs/predicatorv3/continual_fan_inertial_confirmation_r1.yaml deleted file mode 100644 index 108eb42fcd..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_inertial_confirmation_r1.yaml +++ /dev/null @@ -1,15 +0,0 @@ -# Fresh confirmation seeds; launch only after the pilot screening rule passes. -# Execute from the frozen inertial runtime, not a revised environment. -includes: - - continual_fan_inertial_pilot_r1.yaml -START_SEED: 2 -NUM_SEEDS: 3 -APPROACHES: - mb_opus_inertial_pilot_r1: - SKIP: true - mf_opus_inertial_pilot_r1: - SKIP: true - mb_opus_inertial_confirmation_r1: - EXTENDS: mb_opus - mf_opus_inertial_confirmation_r1: - EXTENDS: mf_opus diff --git a/scripts/configs/predicatorv3/continual_fan_inertial_pilot_r1.yaml b/scripts/configs/predicatorv3/continual_fan_inertial_pilot_r1.yaml deleted file mode 100644 index 1287c7c63d..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_inertial_pilot_r1.yaml +++ /dev/null @@ -1,27 +0,0 @@ -# Separate candidate, not a replacement for either existing Fan cohort. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 2 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - fan_inertial: - EXTENDS: fan_maze - FLAGS: - fan_exposed_transfer: true - fan_inertial_transfer: true - fan_train_num_walls_per_task: "[0]" - fan_test_num_walls_per_task: "[0]" - fan_test_num_pos_x: 3 - fan_test_num_pos_y: 3 - fan_train_task_generation: uniform - fan_test_task_generation: uniform -APPROACHES: - mb_opus_inertial_pilot_r1: - EXTENDS: mb_opus - mf_opus_inertial_pilot_r1: - EXTENDS: mf_opus diff --git a/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2.yaml b/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2.yaml deleted file mode 100644 index 34c7f8a337..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2.yaml +++ /dev/null @@ -1,17 +0,0 @@ -# Fresh matched seeds with aligned state/timing/discrepancy guidance. -# Fan is the reviewed 3 mm ramp with the extended landing. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 5 -ENVS: - fan: - SKIP: false -APPROACHES: - oracle_dynamics_opus_fan_prompt_r2: - EXTENDS: oracle_dynamics_opus - FLAGS: - continual_skill_preflight: false - continual_validation_audit: false diff --git a/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2_retry_seed0.yaml b/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2_retry_seed0.yaml deleted file mode 100644 index 3d4b8ce09e..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2_retry_seed0.yaml +++ /dev/null @@ -1,5 +0,0 @@ -# Retry after account a rejected Claude subscription access before doing work. -includes: - - continual_fan_oracle_prompt_r2.yaml -START_SEED: 0 -NUM_SEEDS: 1 diff --git a/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2_retry_seed3.yaml b/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2_retry_seed3.yaml deleted file mode 100644 index e7e3e11e8c..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_oracle_prompt_r2_retry_seed3.yaml +++ /dev/null @@ -1,5 +0,0 @@ -# Retry after account a rejected Claude subscription access before doing work. -includes: - - continual_fan_oracle_prompt_r2.yaml -START_SEED: 3 -NUM_SEEDS: 1 diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_confirmation_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_confirmation_r1.yaml deleted file mode 100644 index 4630f1a11e..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_confirmation_r1.yaml +++ /dev/null @@ -1,15 +0,0 @@ -# Fresh seeds after the ramp pilot screening gate passes. -# Execute from the unchanged frozen ramp runtime. -includes: - - continual_fan_ramp_pilot_r1.yaml -START_SEED: 2 -NUM_SEEDS: 3 -APPROACHES: - mb_opus_ramp_pilot_r1: - SKIP: true - mf_opus_ramp_pilot_r1: - SKIP: true - mb_opus_ramp_confirmation_r1: - EXTENDS: mb_opus - mf_opus_ramp_confirmation_r1: - EXTENDS: mf_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_long_landing_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_long_landing_r1.yaml deleted file mode 100644 index 7e5df47da0..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_long_landing_r1.yaml +++ /dev/null @@ -1,16 +0,0 @@ -# Five development seeds on the visually reviewed 10 cm landing extension. -includes: - - continual_fan_ramp_pilot_r1.yaml -START_SEED: 0 -NUM_SEEDS: 5 -ENVS: - fan_ramp: - FLAGS: - fan_ramp_landing_extension: 0.10 -APPROACHES: - mb_opus_ramp_pilot_r1: - SKIP: true - mf_opus_ramp_pilot_r1: - SKIP: true - mb_opus_ramp_long_landing_r1: - EXTENDS: mb_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_confirmation_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_confirmation_r1.yaml deleted file mode 100644 index e3f5984148..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_confirmation_r1.yaml +++ /dev/null @@ -1,5 +0,0 @@ -# Additional EMPIRIC seeds on the unchanged lower-drop runtime. -includes: - - /home/ycliang/predicators/logs/fan-ramp-low-drop-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_r1.yaml -START_SEED: 2 -NUM_SEEDS: 3 diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_r1.yaml deleted file mode 100644 index 83bf1ea7bb..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_r1.yaml +++ /dev/null @@ -1,17 +0,0 @@ -# Visually reviewed 3 mm ramp, 10 cm longer landing, EMPIRIC screening first. -includes: - - continual_fan_ramp_pilot_r1.yaml -START_SEED: 0 -NUM_SEEDS: 2 -ENVS: - fan_ramp: - FLAGS: - fan_ramp_landing_extension: 0.10 - fan_ramp_rise: 0.003 -APPROACHES: - mb_opus_ramp_pilot_r1: - SKIP: true - mf_opus_ramp_pilot_r1: - SKIP: true - mb_opus_ramp_low_drop_r1: - EXTENDS: mb_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r1.yaml deleted file mode 100644 index 6275901a45..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r1.yaml +++ /dev/null @@ -1,10 +0,0 @@ -# Two additional seeds, preserving the existing frozen Oracle ramp setup. -includes: - - /home/ycliang/predicators/logs/fan-ramp-skill-repair-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml -START_SEED: 5 -NUM_SEEDS: 2 -APPROACHES: - mb_opus_ramp_skill_repair_r1: - SKIP: true - oracle_dynamics_opus_ramp_skill_repair_r1: - EXTENDS: oracle_dynamics_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r2.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r2.yaml deleted file mode 100644 index f9b0b2e97a..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r2.yaml +++ /dev/null @@ -1,10 +0,0 @@ -# Two further seeds, preserving the frozen Oracle Dynamics Fan + Ramp setup. -includes: - - /home/ycliang/predicators/logs/fan-ramp-skill-repair-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml -START_SEED: 7 -NUM_SEEDS: 2 -APPROACHES: - mb_opus_ramp_skill_repair_r1: - SKIP: true - oracle_dynamics_opus_ramp_skill_repair_r1: - EXTENDS: oracle_dynamics_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml deleted file mode 100644 index ed4f65f528..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_oracle_extra_r3.yaml +++ /dev/null @@ -1,10 +0,0 @@ -# Two further seeds of the unchanged frozen Oracle Fan + Ramp cohort. -includes: - - /home/ycliang/predicators/logs/fan-ramp-skill-repair-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml -START_SEED: 9 -NUM_SEEDS: 2 -APPROACHES: - mb_opus_ramp_skill_repair_r1: - SKIP: true - oracle_dynamics_opus_ramp_skill_repair_r1: - EXTENDS: oracle_dynamics_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_pilot_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_pilot_r1.yaml deleted file mode 100644 index 92a8b26f04..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_pilot_r1.yaml +++ /dev/null @@ -1,28 +0,0 @@ -# User approved two seeds per agent. Separate from previous Fan cohorts. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 2 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - fan_ramp: - EXTENDS: fan_maze - FLAGS: - fan_exposed_transfer: true - fan_inertial_transfer: true - fan_ramp_transfer: true - fan_train_num_walls_per_task: "[0]" - fan_test_num_walls_per_task: "[0]" - fan_test_num_pos_x: 3 - fan_test_num_pos_y: 3 - fan_train_task_generation: uniform - fan_test_task_generation: uniform -APPROACHES: - mb_opus_ramp_pilot_r1: - EXTENDS: mb_opus - mf_opus_ramp_pilot_r1: - EXTENDS: mf_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml deleted file mode 100644 index f51a824177..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_baselines_r1.yaml +++ /dev/null @@ -1,20 +0,0 @@ -# Six comparison arms on exactly the repaired EMPIRIC ramp runtime. -includes: - - /home/ycliang/predicators/logs/fan-ramp-skill-repair-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml -START_SEED: 0 -NUM_SEEDS: 5 -APPROACHES: - mb_opus_ramp_skill_repair_r1: - SKIP: true - oracle_dynamics_opus_ramp_skill_repair_r1: - EXTENDS: oracle_dynamics_opus - mf_opus_ramp_skill_repair_r1: - EXTENDS: mf_opus - mf_scene_package_opus_ramp_skill_repair_r1: - EXTENDS: mf_scene_package_opus - standalone_opus_ramp_skill_repair_r1: - EXTENDS: standalone_opus - no_fitting_opus_ramp_skill_repair_r1: - EXTENDS: no_fitting_opus - no_uncertainty_opus_ramp_skill_repair_r1: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml deleted file mode 100644 index 020b931104..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml +++ /dev/null @@ -1,11 +0,0 @@ -# Five matched development seeds on the reviewed 3 mm ramp / longer landing. -# Freeze the original domain and EMPIRIC configuration; only skills change. -includes: - - /home/ycliang/predicators/logs/fan-ramp-skill-repair-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_low_drop_r1.yaml -START_SEED: 0 -NUM_SEEDS: 5 -APPROACHES: - mb_opus_ramp_low_drop_r1: - SKIP: true - mb_opus_ramp_skill_repair_r1: - EXTENDS: mb_opus diff --git a/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml b/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml deleted file mode 100644 index c015e1a5de..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_seed0_resume.yaml +++ /dev/null @@ -1,5 +0,0 @@ -# Resume only the interrupted seed, preserving its cohort and runtime. -includes: - - /home/ycliang/predicators/logs/fan-ramp-skill-repair-runtime-20260921/scripts/configs/predicatorv3/continual_fan_ramp_skill_repair_r1.yaml -START_SEED: 0 -NUM_SEEDS: 1 diff --git a/scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml b/scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml deleted file mode 100644 index 058e2b3f59..0000000000 --- a/scripts/configs/predicatorv3/continual_fan_transfer_pilot_r1.yaml +++ /dev/null @@ -1,27 +0,0 @@ -# Separate pilot, not a replacement for the paper's Fan maze results. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 2 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - fan_transfer: - EXTENDS: fan_maze - FLAGS: - fan_exposed_transfer: true - # Quoted CLI literals replace the inherited lists; YAML lists append. - fan_train_num_walls_per_task: "[0]" - fan_test_num_walls_per_task: "[0]" - fan_test_num_pos_x: 3 - fan_test_num_pos_y: 3 - fan_train_task_generation: uniform - fan_test_task_generation: uniform -APPROACHES: - mb_opus_transfer_pilot_r1: - EXTENDS: mb_opus - mf_opus_transfer_pilot_r1: - EXTENDS: mf_opus diff --git a/scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml b/scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml deleted file mode 100644 index 566a16693c..0000000000 --- a/scripts/configs/predicatorv3/continual_five_ablations_benchmark_r1.yaml +++ /dev/null @@ -1,42 +0,0 @@ -# The five remaining comparison arms of the eight-agent benchmark sweep -# (Sept 18, 2026), launched after the scene-only round: standalone -# program world model, oracle dynamics, zero shot, no harness fitting, -# no explicit uncertainty, all on Opus, on the five benchmark settings, -# three seeds each: 5 arms x 5 domains x 3 seeds = 75 runs. Envs and -# arms come from the menus (envs/continual.yaml, approaches/ -# continual.yaml); the round keys match the eight-agent sweep -# (continual_eight_agent_noisy_sweep.yaml), so the logs land under -# logs//-_opus_benchmark_r1/seed either way. -# The Opus MB, MF and scene-only rounds ran earlier under the same keys. -# -# Launched from a frozen worktree so later edits to the main tree do not -# reach running or requeued jobs: -# python scripts/engaging/launch.py -c predicatorv3/continual_five_ablations_benchmark_r1.yaml --partition mit_preemptable,mit_normal --requeue --accounts a,b,c,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - standalone_opus_benchmark_r1: - EXTENDS: standalone_opus - oracle_dynamics_opus_benchmark_r1: - EXTENDS: oracle_dynamics_opus - zero_shot_opus_benchmark_r1: - EXTENDS: zero_shot_opus - no_fitting_opus_benchmark_r1: - EXTENDS: no_fitting_opus - no_uncertainty_opus_benchmark_r1: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml b/scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml deleted file mode 100644 index 666575ff7f..0000000000 --- a/scripts/configs/predicatorv3/continual_from_assets_expansion_r1.yaml +++ /dev/null @@ -1,29 +0,0 @@ -# Only the eight new runs: Bridge is already running, Fan maze is cancelled. -# Pin menus to the tested runtime rather than the mutable working checkout. -includes: - - /home/ycliang/predicators/logs/empiric-from-assets-runtime-20260921/scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml -ENVS: - fan: - SKIP: true - bridge: - SKIP: true - fan_ramp: - EXTENDS: fan - FLAGS: - fan_exposed_transfer: true - fan_inertial_transfer: true - fan_ramp_transfer: true - fan_ramp_landing_extension: 0.10 - fan_ramp_rise: 0.003 - fan_train_num_walls_per_task: "[0]" - fan_test_num_walls_per_task: "[0]" - fan_test_num_pos_x: 3 - fan_test_num_pos_y: 3 - fan_train_task_generation: uniform - fan_test_task_generation: uniform - domino_high_friction_turn: - SKIP: false - balloons: - SKIP: false - boil: - SKIP: false diff --git a/scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml b/scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml deleted file mode 100644 index 56e420ec6b..0000000000 --- a/scripts/configs/predicatorv3/continual_from_assets_pilot_r1.yaml +++ /dev/null @@ -1,37 +0,0 @@ -# Two seeds each on Fan + ramp, Bridge, Domino, Balloons and Boil. -# No environment changes, no mandatory preflight, separate result cohort. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 2 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - fan_ramp: - EXTENDS: fan - FLAGS: - fan_exposed_transfer: true - fan_inertial_transfer: true - fan_ramp_transfer: true - fan_ramp_landing_extension: 0.10 - fan_ramp_rise: 0.003 - fan_train_num_walls_per_task: "[0]" - fan_test_num_walls_per_task: "[0]" - fan_test_num_pos_x: 3 - fan_test_num_pos_y: 3 - fan_train_task_generation: uniform - fan_test_task_generation: uniform - bridge: - SKIP: false - domino_high_friction_turn: - SKIP: false - balloons: - SKIP: false - boil: - SKIP: false -APPROACHES: - from_assets_opus_pilot_r1: - EXTENDS: from_assets_opus diff --git a/scripts/configs/predicatorv3/continual_no_uncertainty_raw_obs_seed0_r1.yaml b/scripts/configs/predicatorv3/continual_no_uncertainty_raw_obs_seed0_r1.yaml deleted file mode 100644 index df101ac6e6..0000000000 --- a/scripts/configs/predicatorv3/continual_no_uncertainty_raw_obs_seed0_r1.yaml +++ /dev/null @@ -1,23 +0,0 @@ -# Updated no-explicit-uncertainty ablation: one seed across the five benchmark -# domains. This round fits and plans from raw noisy observations, with state -# smoothing and uncertainty-aware decisions disabled. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 1 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - no_uncertainty_raw_obs_opus_r1: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_oracle_validation_r2.yaml b/scripts/configs/predicatorv3/continual_oracle_validation_r2.yaml deleted file mode 100644 index 0a9ec8b201..0000000000 --- a/scripts/configs/predicatorv3/continual_oracle_validation_r2.yaml +++ /dev/null @@ -1,19 +0,0 @@ -# Repair pilot only. Retain r1 recordings; never resume them with new code. -# Launch from a frozen worktree after compute-node validation: -# python scripts/engaging/launch.py -c predicatorv3/continual_oracle_validation_r2.yaml --partition mit_preemptable --requeue --accounts a,b,c,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - domino_high_friction_turn: - SKIP: false - bridge: - SKIP: false -APPROACHES: - oracle_dynamics_opus_benchmark_r2: - EXTENDS: oracle_dynamics_opus - FLAGS: - continual_skill_preflight: false diff --git a/scripts/configs/predicatorv3/continual_principled_belief_r1.yaml b/scripts/configs/predicatorv3/continual_principled_belief_r1.yaml deleted file mode 100644 index 417073b6c8..0000000000 --- a/scripts/configs/predicatorv3/continual_principled_belief_r1.yaml +++ /dev/null @@ -1,44 +0,0 @@ -# The paper's seven arms rerun on the principled joint belief -# (principled-belief branch, Sept 26, 2026): five benchmark settings, -# five seeds, current environments (the Balloons chute aligned with the -# rack since Sept 19, the ramp Fan, Boil's Place fix). Every arm but No -# uncertainty carries 16 joint draws (continual_common.yaml). Launch -# from a frozen worktree with --partition mit_preemptable --requeue -# --accounts b,c,d (account a's organization blocks Claude Code; dat is -# backup only). Seeds 0-1 go first; override START_SEED and NUM_SEEDS in -# a second launch for seeds 2-4. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 2 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - mb_opus_joint_r1: - EXTENDS: mb_opus - mf_opus_joint_r1: - EXTENDS: mf_opus - mf_scene_package_opus_joint_r1: - EXTENDS: mf_scene_package_opus - standalone_opus_joint_r1: - EXTENDS: standalone_opus - oracle_dynamics_opus_joint_r1: - EXTENDS: oracle_dynamics_opus - no_fitting_opus_joint_r1: - EXTENDS: no_fitting_opus - no_uncertainty_opus_joint_r1: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/continual_principled_belief_smoke_r1.yaml b/scripts/configs/predicatorv3/continual_principled_belief_smoke_r1.yaml deleted file mode 100644 index 0733350090..0000000000 --- a/scripts/configs/predicatorv3/continual_principled_belief_smoke_r1.yaml +++ /dev/null @@ -1,28 +0,0 @@ -# Smoke runs of the principled joint belief (principled-belief branch, -# Sept 25, 2026): EMPIRIC on the five benchmark settings, one seed each, -# before the matched comparison. The joint belief is on through -# continual_common.yaml (belief_joint_draws 16). Launch from a frozen -# worktree with --partition mit_preemptable --requeue --accounts a,b,c,d. -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 1 -FLAGS: - continual_skill_preflight: false - continual_validation_audit: false -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - mb_opus_joint_smoke_r1: - EXTENDS: mb_opus diff --git a/scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml b/scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml deleted file mode 100644 index 7b514acdf0..0000000000 --- a/scripts/configs/predicatorv3/continual_real_to_sim_benchmark_r1.yaml +++ /dev/null @@ -1,29 +0,0 @@ -# Agentic real-to-sim baseline on the five benchmark settings (Sept 18, -# 2026): the agent builds its own PyBullet scene from the engine, the -# scene manifest and the asset files (agent_continual_real_to_sim), on -# Opus, three seeds each: 5 domains x 3 seeds = 15 runs. Envs and the -# arm come from the menus (envs/continual.yaml, approaches/continual.yaml); -# the round key follows the eight-agent sweep, so the logs land under -# logs/agent_continual_real_to_sim/-real_to_sim_opus_benchmark_r1/seed. -# -# python scripts/engaging/launch.py -c predicatorv3/continual_real_to_sim_benchmark_r1.yaml --partition mit_preemptable,mit_normal --requeue --accounts a,b,c,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - real_to_sim_opus_benchmark_r1: - EXTENDS: real_to_sim_opus diff --git a/scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml b/scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml deleted file mode 100644 index fe4ccc7453..0000000000 --- a/scripts/configs/predicatorv3/continual_scene_only_benchmark_r1.yaml +++ /dev/null @@ -1,30 +0,0 @@ -# Scene-only ablation on the five benchmark settings (Sept 18, 2026): -# the exact scene twin with corrected base calibration, no mechanism -# code, frozen for the run (agent_continual_scene_only, formerly -# "oracle scene"), on Opus, three seeds each: 5 domains x 3 seeds = 15 -# runs. Envs and the arm come from the menus (envs/continual.yaml, -# approaches/continual.yaml); the round key matches the eight-agent -# sweep so the logs land under logs/agent_continual_scene_only/ -# -scene_only_opus_benchmark_r1/seed either way. -# -# python scripts/engaging/launch.py -c predicatorv3/continual_scene_only_benchmark_r1.yaml --partition mit_preemptable,mit_normal --requeue --accounts c,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 3 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - scene_only_opus_benchmark_r1: - EXTENDS: scene_only_opus diff --git a/scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml b/scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml deleted file mode 100644 index f934a3a5f4..0000000000 --- a/scripts/configs/predicatorv3/continual_standalone_no_uncertainty_r2.yaml +++ /dev/null @@ -1,37 +0,0 @@ -# Relaunch of two comparison arms on the five benchmark settings, seed 0 -# only (Sept 18, 2026); more seeds follow once these look right. -# -# - Standalone sim, closer to WorldCoder: the run_python probe scores the -# agent's world_model.py on the recorded data and rolls a plan through -# it once; plan search (sim.refine), repeated-trial rollouts, predicate -# scoring and engine renders of predicted states are withheld. -# - No explicit uncertainty, with the observation noise undeclared: no -# noise section, no [noise] line, and the harness fit models none of it. -# -# The r1 rounds of these arms had the fuller probe and the declared noise; -# their logs stay under the _benchmark_r1 and _raw_obs_opus_r1 keys. -# -# Launched from a frozen worktree: -# python scripts/engaging/launch.py -c predicatorv3/continual_standalone_no_uncertainty_r2.yaml --partition mit_preemptable,mit_normal --requeue --accounts a,b,c,d -includes: - - continual_common.yaml - - envs/continual.yaml - - approaches/continual.yaml -START_SEED: 0 -NUM_SEEDS: 1 -ENVS: - balloons: - SKIP: false - bridge: - SKIP: false - boil: - SKIP: false - fan: - SKIP: false - domino_high_friction_turn: - SKIP: false -APPROACHES: - standalone_opus_benchmark_r2: - EXTENDS: standalone_opus - no_uncertainty_opus_benchmark_r2: - EXTENDS: no_uncertainty_opus diff --git a/scripts/configs/predicatorv3/envs/all.yaml b/scripts/configs/predicatorv3/envs/all.yaml deleted file mode 100644 index 3ad8e0dcac..0000000000 --- a/scripts/configs/predicatorv3/envs/all.yaml +++ /dev/null @@ -1,352 +0,0 @@ -ENVS: - domino: - NAME: "pybullet_domino" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 - pybullet_birrt_path_subsample_ratio: 2 - domino_turns: - NAME: "pybullet_domino" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 - pybullet_birrt_path_subsample_ratio: 2 - domino_test_turn_ratio: 1.0 - # Min-block friction-sysID tasks (reward = toppled - block_cost * blues); one - # block per mismatch direction. Forward: true friction below the planner's. - domino_low_friction: - NAME: "pybullet_domino" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 # raise to avoid placement collisions - pybullet_birrt_path_subsample_ratio: 2 - domino_min_block_tasks: True - domino_true_friction: 0.1 - domino_planning_friction: 0.5 - domino_min_block_span_lo: 0.13 - domino_min_block_span_hi: 0.30 - domino_min_block_num_blues: 4 - # Reverse: true friction above the planner's, so it over-builds and pays - # per-block cost. Geometry retuned 2026-07-12 (probe_min_block_bands.py). - domino_high_friction: - NAME: "pybullet_domino" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 # raise to avoid placement collisions - pybullet_birrt_path_subsample_ratio: 2 - domino_min_block_tasks: True - # Spans / turn legs: true K* 1 (2 on turns) vs believed 2 (3); four - # staged blues give the believed build a spare. - domino_true_friction: 0.5 - domino_planning_friction: 0.1 - domino_min_block_span_lo: 0.29 - domino_min_block_span_hi: 0.31 - domino_min_block_turn_entry_lo: 0.21 - domino_min_block_turn_entry_hi: 0.24 - domino_min_block_turn_exit_lo: 0.17 - domino_min_block_turn_exit_hi: 0.20 - domino_min_block_num_blues: 4 - domino_block_cost: 0.1 # doubled so one extra blue shows as a 0.1 gap - online_learning_early_stopping_ignore_reward_bar: True - # domino_high_friction with all-turn test tasks; exp_domino.yaml's env. - domino_high_friction_turn: - NAME: "pybullet_domino" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 # raise to avoid placement collisions - pybullet_birrt_path_subsample_ratio: 2 - domino_min_block_tasks: True - domino_true_friction: 0.5 - domino_planning_friction: 0.1 - domino_min_block_span_lo: 0.29 - domino_min_block_span_hi: 0.31 - domino_min_block_turn_entry_lo: 0.21 - domino_min_block_turn_entry_hi: 0.24 - domino_min_block_turn_exit_lo: 0.17 - domino_min_block_turn_exit_hi: 0.20 - domino_min_block_num_blues: 4 - domino_block_cost: 0.1 - domino_test_turn_ratio: 1.0 - online_learning_early_stopping_ignore_reward_bar: True - # domino_high_friction_turn on the real scene's setup (Panda + table tile); - # only the robot and table change, so differences are attributable to them. - domino_high_friction_turn_real: - NAME: "pybullet_domino_real_geometry" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 - pybullet_birrt_path_subsample_ratio: 2 - domino_min_block_tasks: True - domino_true_friction: 0.5 - domino_planning_friction: 0.1 - domino_min_block_span_lo: 0.29 - domino_min_block_span_hi: 0.31 - # Turn legs are the Fetch values and do NOT certify on the Panda (test - # split empty); recalibration is parked, see CALIBRATION_NOTES.md. - domino_min_block_turn_entry_lo: 0.21 - domino_min_block_turn_entry_hi: 0.24 - domino_min_block_turn_exit_lo: 0.17 - domino_min_block_turn_exit_hi: 0.20 - domino_min_block_num_blues: 4 - domino_block_cost: 0.1 - domino_test_turn_ratio: 1.0 - online_learning_early_stopping_ignore_reward_bar: True - pybullet_robot: "panda" - # Required for the Panda: rests the closed fingers on the 0.015 m - # domino's faces (grasp detection tolerance 0.0005). - pybullet_closed_fingers: 0.008 - # Real-world domino: learn and explore in sim on the reconstructed scene, - # test on the real Franka. The scene JSON sizes the env; set no counts here. - domino_real: - NAME: "pybullet_domino_real" - SKIP: True - # store_true flags go in ARGS (bare --flag), never in FLAGS. - ARGS: - - "make_test_images" - - "make_failure_images" - FLAGS: - num_test_tasks: 1 # one reconstructed scene - max_initial_demos: 0 # the grid oracle cannot solve the real scene - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 400 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 - pybullet_birrt_path_subsample_ratio: 2 - option_model_use_gui: False - agent_bilevel_log_state: False - agent_sim_learn_oracle_sim_program: False - agent_sim_learn_oracle_sim_params: False - code_sim_learning_num_mcmc_steps: 0 - pybullet_robot: "panda" - domino_use_skill_factories: True - domino_real_scene: "/home/amberli/babyrobot/BabyRobotPredicator/scenes/domino_straight.json" - # Roles by id (green start, purple target); ignored if the scene has 'role'. - domino_real_start_id: 6 - domino_real_target_id: 5 - pybullet_closed_fingers: 0.015 # fingers on the 0.029 m real domino's faces - real_robot_execute: False - # Heavy-block (immovable obstacle) tasks, mass-only mismatch; exp_domino_heavy.yaml. - domino_heavy: - NAME: "pybullet_domino" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "loc,rot,angle,direction" - horizon: 500 - domino_initialize_at_finished_state: False - domino_use_domino_blocks_as_target: True - domino_use_continuous_place: True - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: False - keep_failed_demos: True - predicate_invent_invent_derived_predicates: True - pybullet_birrt_extend_num_interp: 20 # raise to avoid placement collisions - pybullet_birrt_path_subsample_ratio: 2 - domino_heavy_block_tasks: True - domino_min_block_num_blues: 4 - domino_test_turn_ratio: 1.0 - # Flags from the boil oracle test; boil has a PO simulator. exp_boil_sweep.yaml. - boil: - NAME: "pybullet_boil" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "switch" - max_num_steps_option_rollout: 100 - horizon: 500 - boil_goal: "simple" - boil_require_jug_full_to_heatup: True - script_option_file_name: "boil.txt" - boil_water_fill_speed: 0.0015 - pybullet_birrt_path_subsample_ratio: 2 - boil_num_jugs_train: [1] - boil_num_jugs_test: [2] - boil_num_burner_train: [1] - boil_num_burner_test: [1] - fan: - NAME: "pybullet_fan" - SKIP: True - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: "switch" - terminate_on_goal_reached: True - process_planning_heuristic_weight: 10.0 - horizon: 500 - pybullet_birrt_path_subsample_ratio: 2 - process_planning_max_execution_replans: 3 # safety net for learned models - # Test tasks: maze layouts on the full 10 x 9 arena grid (see - # fan_test_task_generation in settings.py). - fan_train_num_pos_x: 3 - fan_train_num_pos_y: 3 - fan_test_num_pos_x: 10 - fan_test_num_pos_y: 9 - fan_train_num_walls_per_task: [1] - fan_test_num_walls_per_task: [16, 20, 24] - fan_train_task_generation: "uniform" - fan_test_task_generation: "maze" - fan_maze_min_segments: 4 - fan_maze_min_path_len: 10 - fan_maze_max_segment_len: 4 - bridge: - NAME: "pybullet_bridge" - SKIP: True - FLAGS: - max_initial_demos: 0 - horizon: 3000 - # Real episodes take 600-950 steps; 2000 leaves headroom for a 20-25 - # option plan while bounding a looping policy's data. - max_num_steps_interaction_request: 2000 - # Press 3 N into the support before release (post-release drift 4.5 -> 0.2 mm). - skill_place_settle_preload_force: 3.0 - process_planning_heuristic_weight: 10.0 # unweighted skeleton search is slow - # Each Wait ends on the first atom change, so multi-cure tails need a - # cheap replan-from-current-state. - process_planning_max_execution_replans: 3 - # The packed staging grid grazes by 2-3 mm; the default 1 mm margin - # makes those unrecoverable BiRRT rejections. - pybullet_birrt_contact_margin: -0.005 - # Kinematic pin of welded assemblies while held (carried bonded beams - # drift ~9-12 deg unpinned, 0 pinned). - pybullet_pin_held_weld_assemblies: True - busyboard: - NAME: "pybullet_busyboard" - # Parked by default; exp_busyboard.yaml un-skips it. - SKIP: True - FLAGS: - max_initial_demos: 0 - # A press is ~22 low-level steps and a lamp needs ~48 driven ones - # to light, so a three-press plan runs past the common 500 cap and - # every refinement would be rejected on the horizon check. - horizon: 2000 - # REQUIRED here, for the same reason pybullet_bridge needs it. A - # lamp's lighting is a delayed effect of a drive condition, so the - # tick it lands on depends on how many low-level steps the - # surrounding options happen to take; the symbolic delay places it - # on one tick and physics may deliver it on the neighbouring one, - # and the per-step atom check then rejects a plan that reaches the - # goal. Measured 2026-09-01 on oracle_process_planning over 10 test - # tasks: 8/10 with the check on, 10/10 with it off, at either - # candidate delay value. - sesame_check_expected_atoms: False - # Every button press that lights nothing is free on this board, so - # nothing in the goal stops a plan from latching extra buttons. The - # necessity gate refuses a capture that still reaches the goal with - # a step removed (run_20260902_152811 pressed three of four buttons - # for a two-press goal). - agent_plan_validation_necessity: True - icerink: - NAME: "pybullet_icerink" - # Parked by default; a launcher built on continual_common.yaml un-skips it. - SKIP: True - FLAGS: - max_initial_demos: 0 - # A push is ~60 low-level steps and a slide settles within ~40 - # more; a three-push plan plus a wait runs past the common 500 cap. - horizon: 2000 - # A slide lands on its target a few ticks after the push option - # ends, so the per-step atom check would reject a plan that - # reaches the goal (same reason as pybullet_busyboard). - sesame_check_expected_atoms: False - pybullet_birrt_path_subsample_ratio: 2 - launcher: - NAME: "pybullet_launcher" - SKIP: True - FLAGS: - max_initial_demos: 0 - horizon: 1500 - # The top block topples a few ticks after the launch option ends. - sesame_check_expected_atoms: False - pybullet_birrt_path_subsample_ratio: 2 - magnets: - NAME: "pybullet_magnets" - SKIP: True - FLAGS: - max_initial_demos: 0 - horizon: 1500 - # A pulled piece settles under the tip a few ticks after the move. - sesame_check_expected_atoms: False - pybullet_birrt_path_subsample_ratio: 2 - balloons: - NAME: "pybullet_balloons" - SKIP: True - FLAGS: - max_initial_demos: 0 - horizon: 1500 - # The box rises to the band a few ticks after the last tie. - sesame_check_expected_atoms: False - pybullet_birrt_path_subsample_ratio: 2 - crane: - NAME: "pybullet_crane" - SKIP: True - FLAGS: - max_initial_demos: 0 - horizon: 1500 - # The crate settles on the bin some ticks after the pull option ends. - sesame_check_expected_atoms: False - pybullet_birrt_path_subsample_ratio: 2 diff --git a/scripts/configs/predicatorv3/envs/continual.yaml b/scripts/configs/predicatorv3/envs/continual.yaml deleted file mode 100644 index 9a44c0d52d..0000000000 --- a/scripts/configs/predicatorv3/envs/continual.yaml +++ /dev/null @@ -1,266 +0,0 @@ -# Continual-protocol environment menu. Every entry is parked (SKIP); a -# launcher un-parks the one it runs (see continual_common.yaml). Env FLAGS -# override the approach's and the common ones, so per-domain noise levels -# and level budgets live here. Keys are the env half of the experiment id. -# -# The benchmark settings (Sept 18, 2026) are balloons (composition test -# levels), bridge (four-span test row), boil (two-jug test level), fan -# (maze test level) and domino_high_friction_turn; the other entries are -# the earlier or alternative splits they replaced. -ENVS: - # Balloons: the main tree's benchmark, composition test levels (every test - # level composes lifts the - # agent measured on training racks; the weakest-first release bursts) and - # the 25-step goal dwell. Round 1 (Sept 16, 2026) ran with dwell 1, which - # let a swinging box win at a turning point; not comparable. - balloons: - NAME: pybullet_balloons - SKIP: true - FLAGS: - max_initial_demos: 0 - horizon: 1500 - sesame_check_expected_atoms: false - pybullet_birrt_path_subsample_ratio: 2 - balloons_require_jam_decoy: false - balloons_goal_dwell_steps: 25 - continual_obs_noise_position: 0.01 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 2 - # Balloons, bundle test levels (Sept 17, 2026): the test rack ties its - # balloons into bundles of two, one clip per bundle, so no cut is a - # small trim and the arithmetic answer (the bundle resting in band) - # bursts on its first-cut overshoot; see the pybullet_balloons module - # doc. Training racks are the usual singles. - balloons_bundles: - NAME: pybullet_balloons - SKIP: true - FLAGS: - max_initial_demos: 0 - horizon: 1500 - sesame_check_expected_atoms: false - pybullet_birrt_path_subsample_ratio: 2 - balloons_require_jam_decoy: false - balloons_goal_dwell_steps: 25 - balloons_test_bundle_sizes: [2, 2, 2, 2] - balloons_max_sampling_attempts: 60 - continual_obs_noise_position: 0.01 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 2 - # Bridge, three spans on both splits (the Sept 14 comparison setting). - bridge_three_span: - NAME: pybullet_bridge - SKIP: true - FLAGS: &bridge_three_span_flags - max_initial_demos: 0 - horizon: 3000 - max_num_steps_interaction_request: 2000 - skill_place_settle_preload_force: 3.0 - process_planning_heuristic_weight: 10.0 - process_planning_max_execution_replans: 3 - wait_option_max_steps: 120 - pybullet_birrt_contact_margin: -0.005 - pybullet_pin_held_weld_assemblies: true - continual_obs_noise_position: 0.005 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 10000 - num_train_tasks: 1 - bridge_train_span_blocks: 3 - bridge_test_span_blocks: 3 - # Bridge, the benchmark setting: a three-span training row, a four-span - # test row, with the Sept 16 four-span repairs (rigid grasp, lift-first - # transit, certificate that waits for the robot to withdraw). Not - # comparable with the span_transfer_r1 cohort, which ran without the - # repairs. Opus MB 3/3 at 3852 steps against MF 3/3 at 6134 (Sept 17-18). - bridge: - NAME: pybullet_bridge - SKIP: true - FLAGS: &bridge_flags - <<: *bridge_three_span_flags - bridge_train_span_blocks: 3 - bridge_test_span_blocks: 4 - pybullet_grasp_max_force: 10000.0 - bridge_lift_before_transit: true - bridge_goal_robot_clearance: 0.01 - # The same setting under the key the Sept 17 span-transfer rounds ran - # under (their logs live at bridge_span_transfer-). - bridge_span_transfer: - NAME: pybullet_bridge - SKIP: true - FLAGS: *bridge_flags - # Boil: one jug in training, two jugs on the test level, the 5 cm faucet - # tolerance (the code default since d3ac34f02, pinned here; the old 10 cm - # rows were dropped on Sept 17). - boil: - NAME: pybullet_boil - SKIP: true - FLAGS: &boil_flags - max_initial_demos: 0 - excluded_objects_in_state_str: switch - max_num_steps_option_rollout: 100 - horizon: 500 - boil_goal: simple - boil_require_jug_full_to_heatup: true - script_option_file_name: boil.txt - boil_water_fill_speed: 0.0015 - pybullet_birrt_path_subsample_ratio: 2 - boil_num_jugs_train: [1] - boil_num_jugs_test: [2] - boil_num_burner_train: [1] - boil_num_burner_test: [1] - boil_faucet_align_threshold: 0.05 - continual_obs_noise_position: 0.0125 - continual_obs_noise_orientation: 0.05 - continual_obs_noise_scalar: 0.07 - continual_steps_per_level: 5000 - num_train_tasks: 1 - # Historical maze setting, retained for reproducible development configs. - fan_maze: - NAME: pybullet_fan - SKIP: true - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: switch - terminate_on_goal_reached: true - process_planning_heuristic_weight: 10.0 - horizon: 500 - pybullet_birrt_path_subsample_ratio: 2 - process_planning_max_execution_replans: 3 - fan_train_num_pos_x: 3 - fan_train_num_pos_y: 3 - fan_test_num_pos_x: 10 - fan_test_num_pos_y: 9 - fan_train_num_walls_per_task: [1] - fan_test_num_walls_per_task: [16, 20, 24] - fan_train_task_generation: uniform - fan_test_task_generation: maze - fan_maze_min_segments: 4 - fan_maze_min_path_len: 10 - fan_maze_max_segment_len: 4 - continual_obs_noise_position: 0.005 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 1 - # Default benchmark Fan: reviewed 3 mm ramp with longer landing. - fan: - NAME: pybullet_fan - SKIP: true - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: switch - terminate_on_goal_reached: true - process_planning_heuristic_weight: 10.0 - horizon: 500 - pybullet_birrt_path_subsample_ratio: 2 - process_planning_max_execution_replans: 3 - fan_train_num_pos_x: 3 - fan_train_num_pos_y: 3 - fan_exposed_transfer: true - fan_inertial_transfer: true - fan_ramp_transfer: true - fan_ramp_rise: 0.003 - fan_ramp_landing_extension: 0.10 - fan_train_num_walls_per_task: [0] - fan_test_num_walls_per_task: [0] - fan_test_num_pos_x: 3 - fan_test_num_pos_y: 3 - fan_train_task_generation: uniform - fan_test_task_generation: uniform - continual_obs_noise_position: 0.005 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 1 - # Domino: min-block friction tasks with a turn, true friction above the - # planner's belief. - domino_high_friction_turn: - NAME: pybullet_domino - SKIP: true - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: loc,rot,angle,direction - horizon: 500 - domino_initialize_at_finished_state: false - domino_use_domino_blocks_as_target: true - domino_use_continuous_place: true - process_planning_heuristic_weight: 2.0 - domino_has_glued_dominos: false - keep_failed_demos: true - predicate_invent_invent_derived_predicates: true - pybullet_birrt_extend_num_interp: 20 - pybullet_birrt_path_subsample_ratio: 2 - domino_min_block_tasks: true - domino_true_friction: 0.5 - domino_planning_friction: 0.1 - domino_min_block_span_lo: 0.29 - domino_min_block_span_hi: 0.31 - domino_min_block_turn_entry_lo: 0.21 - domino_min_block_turn_entry_hi: 0.24 - domino_min_block_turn_exit_lo: 0.17 - domino_min_block_turn_exit_hi: 0.2 - domino_min_block_num_blues: 4 - domino_block_cost: 0.1 - domino_test_turn_ratio: 1.0 - online_learning_early_stopping_ignore_reward_bar: true - continual_obs_noise_position: 0.01 - continual_obs_noise_orientation: 0.04 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 1 - # Fan, the original test split: uniform 6 x 6 test positions with two or - # three walls (the Sept 16 Fan row; those were the runtime defaults then, - # pinned here because the maze entry above overrides them). - fan_uniform: - NAME: pybullet_fan - SKIP: true - FLAGS: - max_initial_demos: 0 - excluded_objects_in_state_str: switch - terminate_on_goal_reached: true - process_planning_heuristic_weight: 10.0 - horizon: 500 - pybullet_birrt_path_subsample_ratio: 2 - process_planning_max_execution_replans: 3 - fan_test_num_pos_x: 6 - fan_test_num_pos_y: 6 - fan_test_num_walls_per_task: [2, 3] - fan_test_task_generation: uniform - continual_obs_noise_position: 0.005 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 1 - # Boil, one test jug (Sonnet MB 2/2 in 1240 vs MF 2/2 in 3379 on Sept 16). - boil_one_jug: - NAME: pybullet_boil - SKIP: true - FLAGS: - <<: *boil_flags - boil_num_jugs_test: [1] - # Balloons, the original Sept 16 pilot setting (the table's Balloons row): - # chute scene, original task distribution, jam decoy, dwell 1. Its - # balloons_scene and balloons_task_generation flags come from the - # continual-comparisons line (frozen worktrees), not the main tree, until - # that merge lands. - balloons_original: - NAME: pybullet_balloons - SKIP: true - FLAGS: - max_initial_demos: 0 - horizon: 1500 - sesame_check_expected_atoms: false - pybullet_birrt_path_subsample_ratio: 2 - balloons_scene: chute - balloons_task_generation: original - balloons_require_jam_decoy: true - balloons_goal_dwell_steps: 1 - continual_obs_noise_position: 0.01 - continual_obs_noise_orientation: 0.02 - continual_obs_noise_scalar: 0.0 - continual_steps_per_level: 5000 - num_train_tasks: 2 diff --git a/scripts/configs/predicatorv3/exp_boil_sweep.yaml b/scripts/configs/predicatorv3/exp_boil_sweep.yaml deleted file mode 100644 index ff29763cbb..0000000000 --- a/scripts/configs/predicatorv3/exp_boil_sweep.yaml +++ /dev/null @@ -1,119 +0,0 @@ -# BOIL sweep over every paper arm (C7 dropped), 3 seeds each. -# Usage: python scripts/engaging/launch.py -c predicatorv3/exp_boil_sweep.yaml ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -NUM_SEEDS: 3 -ENVS: - boil: - SKIP: False - FLAGS: - num_test_tasks: 5 -APPROACHES: -# Every arm is tested after every learning cycle. Learning arms skip the pre-loop -# test; U1 and A2 (no cycles) are evaluated by it, and C1 keeps it as A1 (no learning). - # C1 - sim_predicator: - SKIP: False - ARGS: - - auto_resume - # C1 in policy mode (parked; the paper's C1 is plan mode). - sim_predicator_policy: - SKIP: True - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C2 - agent_model_free_planning: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C3 - nl_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C4 - code_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C5 - gnn_dynamics: - SKIP: False - FLAGS: - skip_initial_test: True - # C6 - operator_learning: - SKIP: False - FLAGS: - skip_initial_test: True - # C8 - maple_q: - SKIP: False - FLAGS: - skip_initial_test: True - # U1 (PO and open-loop, like C1) - agent_oracle_hybrid_sim: - SKIP: False - FLAGS: - partially_observable: True - agent_bilevel_max_execution_replans: 0 - ARGS: - - auto_resume - # A2 - sim_predicator_zero_shot: - SKIP: False - ARGS: - - auto_resume - # A3 - sim_predicator_undirected_explore: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A4 (the rule-param margin samples the declared intervals here) - sim_predicator_no_param_fit: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A5 - sim_predicator_no_validation: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A6 - sim_predicator_explore_no_disagreement: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A7 - sim_predicator_validation_no_uncertainty: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A8 - sim_predicator_no_predicates: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume diff --git a/scripts/configs/predicatorv3/exp_bridge.yaml b/scripts/configs/predicatorv3/exp_bridge.yaml deleted file mode 100644 index e981ab3fdf..0000000000 --- a/scripts/configs/predicatorv3/exp_bridge.yaml +++ /dev/null @@ -1,33 +0,0 @@ -# Bridge launcher: un-skips the bridge env and the OURS arm(s). -# Usage: python scripts/local/launch_simp.py -c predicatorv3/exp_bridge.yaml --parallel ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -# Seed-3 rerun on the 2026-09-02 execution fixes (stall completion -# 621190a1b, retreat lift-off 7ac637ddd, partial-open descent 4dcc7121e, -# clearance-aware certification); prior checkpoints parked in -# saved_approaches/backup_pre_clearance_relaunch_20260902/ so auto_resume -# starts fresh. Top-level scalars override common.yaml here only. -START_SEED: 3 -NUM_SEEDS: 1 -ENVS: - bridge: - SKIP: False -APPROACHES: - agent_oracle_hybrid_sim: - SKIP: True # regression passed 2026-08-24 (1/1) - sim_predicator: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume # resume the latest per-cycle checkpoint on requeue - # Same learner; the solve deliverable is a per-task closed-loop policy.py. - sim_predicator_policy: - SKIP: True # plan arm only for the seed-3 rerun - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume diff --git a/scripts/configs/predicatorv3/exp_bridge_sweep.yaml b/scripts/configs/predicatorv3/exp_bridge_sweep.yaml deleted file mode 100644 index 4054e755cb..0000000000 --- a/scripts/configs/predicatorv3/exp_bridge_sweep.yaml +++ /dev/null @@ -1,119 +0,0 @@ -# BRIDGE sweep over every paper arm (C7 dropped), 3 seeds each. -# Usage: python scripts/engaging/launch.py -c predicatorv3/exp_bridge_sweep.yaml ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -NUM_SEEDS: 3 -ENVS: - bridge: - SKIP: False - FLAGS: - num_test_tasks: 5 -APPROACHES: -# Every arm is tested after every learning cycle. Learning arms skip the pre-loop -# test; U1 and A2 (no cycles) are evaluated by it, and C1 keeps it as A1 (no learning). - # C1 - sim_predicator: - SKIP: False - ARGS: - - auto_resume - # C1 in policy mode (parked; the paper's C1 is plan mode). - sim_predicator_policy: - SKIP: True - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C2 (same per-attempt solve budget as C1) - agent_model_free_planning: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C3 - nl_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C4 - code_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C5 - gnn_dynamics: - SKIP: False - FLAGS: - skip_initial_test: True - # C6 - operator_learning: - SKIP: False - FLAGS: - skip_initial_test: True - # C8 (same interaction budget as the agent arms) - maple_q: - SKIP: False - FLAGS: - skip_initial_test: True - # U1 (PO and open-loop, like C1) - agent_oracle_hybrid_sim: - SKIP: False - FLAGS: - partially_observable: True - agent_bilevel_max_execution_replans: 0 - ARGS: - - auto_resume - # A2 - sim_predicator_zero_shot: - SKIP: False - ARGS: - - auto_resume - # A3 - sim_predicator_undirected_explore: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A4 (the rule-param margin samples the declared intervals here) - sim_predicator_no_param_fit: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A5 - sim_predicator_no_validation: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A6 - sim_predicator_explore_no_disagreement: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A7 - sim_predicator_validation_no_uncertainty: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A8 - sim_predicator_no_predicates: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume diff --git a/scripts/configs/predicatorv3/exp_busyboard.yaml b/scripts/configs/predicatorv3/exp_busyboard.yaml deleted file mode 100644 index 0c02090b53..0000000000 --- a/scripts/configs/predicatorv3/exp_busyboard.yaml +++ /dev/null @@ -1,30 +0,0 @@ -# Thin launcher: run the busyboard experiment (one arm at a time). -# Usage: python scripts/local/launch_simp.py -c predicatorv3/exp_busyboard.yaml --parallel -# Env definitions live in envs/all.yaml and approach definitions in -# approaches/all.yaml (all parked by default); this file only un-skips the -# env + arm(s) it runs. To run a baseline sweep, flip additional arms' -# SKIP to False here. -# -# The env block carries --sesame_check_expected_atoms False, which this -# domain requires; see envs/all.yaml for why. The demonstrator -# (oracle_process_planning) solves 10/10 test tasks with it. ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -ENVS: - busyboard: - SKIP: False - # busyboard has no agent-specific excluded_predicates. The wiring - # helper predicates (SoleDriver / JointDrivers) are injected for the - # oracle only, so they are not in env.predicates and an agent already - # runs over the physical vocabulary without them. -APPROACHES: - # Oracle hybrid-sim arm: solved 1/1 with the minimal plan (job 21840975). - agent_oracle_hybrid_sim: - SKIP: True - sim_predicator: - SKIP: False - ARGS: - - auto_resume # resume the latest per-cycle checkpoint on requeue diff --git a/scripts/configs/predicatorv3/exp_domino.yaml b/scripts/configs/predicatorv3/exp_domino.yaml deleted file mode 100644 index 99c51ce1c2..0000000000 --- a/scripts/configs/predicatorv3/exp_domino.yaml +++ /dev/null @@ -1,34 +0,0 @@ -# Thin launcher: run the domino friction-sysID experiment ("ours" arm). -# Usage: python scripts/local/launch_simp.py -c predicatorv3/exp_domino.yaml --parallel -# Env definitions live in envs/all.yaml and approach definitions in -# approaches/all.yaml (all parked by default); this file only un-skips the -# env + arm(s) it runs and sets agent-specific ENVS overrides. ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -ENVS: - # NOTE: launch_simp runs the ENVS x APPROACHES cross-product - park - # arms here before launching if the full matrix is not wanted. - domino: - SKIP: True - domino_turns: - SKIP: True - # Forward (over-reach) arm: true friction 0.1, planner believes 0.5. - domino_low_friction: - SKIP: True - # Reverse (under-reach) arm: true friction 0.5, planner believes 0.1. - domino_high_friction: - SKIP: True - domino_high_friction_turn: - SKIP: False -APPROACHES: - agent_oracle_hybrid_sim: - SKIP: True - FLAGS: - agent_sim_learn_kept_predicates_names: ["Holding", "HandEmpty"] - sim_predicator: - SKIP: False - FLAGS: - skip_initial_test: True diff --git a/scripts/configs/predicatorv3/exp_domino_heavy.yaml b/scripts/configs/predicatorv3/exp_domino_heavy.yaml deleted file mode 100644 index 90d715629c..0000000000 --- a/scripts/configs/predicatorv3/exp_domino_heavy.yaml +++ /dev/null @@ -1,20 +0,0 @@ -# Thin launcher: validate the heavy-block (mass-only mismatch) domino tasks -# with the oracle-sim upper-bound arm (GT hybrid sim + GT physical params, -# i.e. the planner knows the gray block's true heavy mass). -# Usage: python scripts/local/launch_simp.py -c predicatorv3/exp_domino_heavy.yaml --parallel -# Env definitions live in envs/all.yaml and approach definitions in -# approaches/all.yaml (all parked by default); this file only un-skips the -# env + arm(s) it runs. ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -ENVS: - domino_heavy: - SKIP: False -APPROACHES: - agent_oracle_hybrid_sim: - SKIP: False - FLAGS: - agent_sim_learn_kept_predicates_names: ["Holding", "HandEmpty"] diff --git a/scripts/configs/predicatorv3/exp_domino_real.yaml b/scripts/configs/predicatorv3/exp_domino_real.yaml deleted file mode 100644 index ae214c2028..0000000000 --- a/scripts/configs/predicatorv3/exp_domino_real.yaml +++ /dev/null @@ -1,196 +0,0 @@ -# Thin launcher: real-world domino testing through the Predicators -# pipeline. -# Usage: python scripts/local/launch_simp.py -c predicatorv3/exp_domino_real.yaml ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -FLAGS: - # One cycle for the first integration run: the point is to get one episode - # all the way through record -> post-process -> fit. Raise it once that - # path is known to work; with human_reset off there is no prompt between - # episodes, so a multi-cycle run would start the second one on whatever the - # first left behind. - num_online_learning_cycles: 1 - wait_option_max_steps: 200 - # One exploration episode per cycle. common.yaml asks for 2, but the - # explorer replays the same fixed plan every time, so the second episode - # costs a full run of the hardware (and a human scene reset) to collect a - # near-duplicate of the first. - online_nsrt_learning_requests_per_cycle: 1 -NUM_SEEDS: 1 -# Agent-only env overrides: deep-merged on top of envs/all.yaml. These -# excluded_predicates are dropped for "ours" runs (matches exp_domino.yaml; -# oracle.yaml would keep them in). -ENVS: - domino_real: - SKIP: False - FLAGS: - excluded_predicates: "InitialBlock,MovableBlock,Tilting,Upright,InFront" - real_robot_execute: True - # LIVE: the arm moves and the cameras look. Learning a friction from - # real observations needs a real cascade to observe, so neither half - # can be faked -- a dry arm leaves the scene untouched and a blind run - # has nothing to report. (For the no-motion rung instead, set - # real_robot_dry True with perception "none" and the two flags below - # False; "none" with the look on raises at executor construction.) - real_robot_dry: False - # -- open-loop execution, recorded end to end -------------------------- - # The whole episode's motion ships in one batch once it has been - # simulated, so the arm runs the plan as one contiguous stroke instead - # of idling through the next option's motion planning. Mutually - # exclusive with the boundary look, which is asserted at construction: - # a look has to happen BETWEEN two options and batching leaves no such - # moment. - real_robot_open_loop_episode: True - real_robot_observe_at_option_boundary: False - # Start the take in front of Push rather than in front of the whole - # batch. Only the cascade is scored: on run_20260818_092302 the first - # onset was 107 s into a 131 s track, so the pick-and-place before it - # was ~80% of the video and none of the evidence. The arm still runs - # the bridge as one batch, then pauses once while the take opens. - # - # This also lines the track's frame 0 up with the arrangement the push - # acts on, which is what the id matching compares against -- with the - # take starting at the reset, that run could not match 2 of 4 dominoes - # and logged 122 warnings. - real_robot_record_from_option: "Push" - # The cameras record the whole execution; the poses come out of - # post-processing afterwards. NOT "zed": that is the marker pipeline, - # whose 20mm tags do not resolve at this camera distance (1 of ~7 on - # one camera, 0 on the other), and it would also fight the recorder - # for the cameras. "scene_file" replays the captured layout, which is - # what a fixed-plan replay wants -- the plan names specific objects and - # a rebuild could renumber them. - real_robot_perception: "scene_file" - real_robot_human_reset: False - real_robot_record_episodes: True - real_robot_process_takes: True - # Fit the poses from 30264679. Markerless is single-camera and the two - # are not interchangeable: on hand-measured ground truth this one is 6x - # better on orientation (1.03 deg median against 6.29), which is what - # the topple onsets are read off. The other tracks more frames (99.9% - # against 82%), so revisit if coverage turns out to matter more. - real_robot_track_camera: "30264679" - # Trim the still lead-in before SAM-2 sees it. The take starts at the - # reset and the twin then simulates every option with the arm parked: - # on run_20260817_162250 that was 152 s of a static scene out of 420 s - # recorded, ~6.5k of 18.1k frames. Needs BabyRobotPredicator's - # --trim-motion; an older driver ignores the request rather than - # failing, and test_the_driver_honours_the_trim_request_once_it_can - # says which of the two you have. - real_robot_trim_still_frames: True - # Stage 2 needs one box per domino. Draw them once at the start of the - # run, in a drag window, while a human is still at the bench -- rather - # than producing a boxes.json out of band beforehand, or having a window - # open mid-run. Valid here because the fixed plan trains and tests on - # one arrangement, so the boxes drawn on the scene as it stands are the - # right ones for every take. Needs a display (X forwarding over SSH). - # Set real_robot_snapshot_boxes_json to an earlier run's boxes.json - # instead to skip the window entirely. The 2026-08-17 capture shipped - # one, drawn on this very arrangement -- scenes/ - # domino_row_20260817_markerless/boxes.json, four boxes at 1280x720 on - # camera 30264679 -- so it is reusable only if the run records on the - # same camera at the same resolution. Check before trusting it: boxes - # from another camera land on empty table and stage 2 fits nothing. - real_robot_pick_boxes_at_start: True - real_robot_snapshot_boxes_json: "" - # -- post-processing speed --------------------------------------------- - # run_20260818_092302 took 1008 s to turn a 138 s take into a track and - # missed the fit's 900 s deadline by 108 s, so the episode it had just - # recorded was skipped. These two take roughly 5 minutes off that. - # - # 30 of this machine's 32 cores for stage 4, which ran 16.0 cores busy - # for its whole 364 s. Sized for THIS box -- lower it on a smaller one, - # and remember the pipeline runs in the background while the next - # episode drives the robot. - real_robot_track_jobs: 30 - # Skip masks_overlay.mp4: 163 s, rendered before stage 4 and so paid - # straight out of time-to-track. Turn it back on when the tracks look - # wrong -- it is how id swaps are spotted. - real_robot_track_viz: False - real_robot_divergence_atol: 0.02 - # No learning for now, just hybrid sim. - # agent_sim_learn_oracle_sim_program: True - # agent_sim_learn_oracle_sim_params: True - # -- the mismatch ------------------------------------------------------ - # Applied only to sims built with skip_process_dynamics=True (the - # approach's base env and option models), so the twin keeps the - # settings.py default 0.5 and goes on modelling the real table. - domino_planning_friction: 0.1 - # -- the scene, and the roles it does not carry ------------------------- - # The 2026-08-17 capture: four dominoes, all standing, on an arc rather - # than a row. Its records are in id order, so capture id N lands in slot - # N and is named domino_N. - domino_real_scene: "/home/amberli/babyrobot/BabyRobotPredicator/scenes/domino_row_20260817.json" - # This capture has no per-domino 'role' field, so the env reads the roles - # off these ids -- green (the one Push acts on) is capture id 3, purple - # (the goal) is capture id 0, and ids 1 and 2 are the movables the plan - # bridges with. envs/all.yaml's 6 / 5 are domino_straight.json's ids and - # appear nowhere in this scene; left in place the task has no target at - # all and _task_from_perceived asserts on it. - domino_real_start_id: 3 - domino_real_target_id: 0 - # -- learning the friction from the recording -------------------------- - # Score the free-running rollout against the markerless pose track - # instead of against every recorded state. Under open-loop nothing - # corrects the twin, so those states ARE the twin's own simulation and - # scoring them recovers the twin's friction by construction -- the - # defect this experiment exists to fix. With this off, turning the two - # flags above on makes the fit worse, not better. - code_sim_learning_rollout_score_observed_only: True - # The run manifest the recorder writes, naming each episode's track. - code_sim_learning_rollout_track_path: "logs/zed_tracks/tracks.json" - # Drop the commanded arm and the non-kinematic features from the scored - # scope. The arm reproduces at every candidate friction so it can only - # dilute -- and with it in scope nothing in the episode ever rests, so - # rest-point segmentation can never cut. - code_sim_learning_rollout_scope_types: ["domino"] - # The fit blocks this long for tracks the manifest promised. - # Post-processing runs about 3x the length of a take, and the loop fits - # as soon as an episode ends; without the wait the fit finds nothing and - # falls back to the per-step scoring above. - code_sim_learning_track_wait_s: 900.0 - # The markerless pipeline emits poses in the ROBOT BASE frame; a twin - # state is in the env's world frame. For this env the two differ by a - # quarter turn about z plus (0.75, 0.72) -- exactly - # pybullet_domino.real_geometry.base_to_world_transform, whose - # constants these mirror. Matching absorbs a translation by voting over - # candidate offsets, but not a rotation: unset, every track/twin pair - # lands 144-307 mm apart against a 40 mm tolerance and no domino is - # matched at all. - code_sim_learning_track_frame_yaw: 1.5707963267948966 - code_sim_learning_track_frame_xy: [0.75, 0.72] -APPROACHES: - agent_oracle_hybrid_sim: - SKIP: True - sim_predicator: - SKIP: False - FLAGS: - # Replay one fixed plan every episode instead of planning. What is - # being tested is whether perception feeds the learner well enough to - # move the friction belief, so the exploration half should be a - # constant: a run that goes wrong is then the loop's fault and not the - # planner's, and every episode is directly comparable to the last. - # (explorer is only ever set in approaches/all.yaml, never in - # envs/all.yaml, so this override is not shadowed by the env block.) - explorer: "fixed_plan" - # Matched to domino_real_scene above: the plan names specific objects and - # specific world coordinates, so the two move together or the replay - # places dominoes into empty table. The sketch's header carries the - # geometry it was derived from. One alternate for this same scene sits - # beside it -- _release055, the same plan with the placement drop cut - # from 29 mm to 9 mm, for if the real placements bounce or land tilted. - fixed_plan_explorer_path: "scripts/plan_sketches/domino_row_20260817_bridge2.txt" - # The agent's own plan-testing simulator is built with - # skip_process_dynamics=agent_planner_use_base_simulator, and only a - # skip_process_dynamics=True env picks up domino_planning_friction. - # Left False (the default) the agent tests its plans against the TRUE - # friction -- its belief is not mismatched at all, and there is - # nothing for the sysID to discover. True is what makes the agent - # actually believe 0.1. - agent_planner_use_base_simulator: True - # The pre-loop test is a second full episode on the hardware before - # anything is learned; the experiment is about the cycle. - skip_initial_test: True diff --git a/scripts/configs/predicatorv3/exp_domino_real_geometry.yaml b/scripts/configs/predicatorv3/exp_domino_real_geometry.yaml deleted file mode 100644 index 79184f8153..0000000000 --- a/scripts/configs/predicatorv3/exp_domino_real_geometry.yaml +++ /dev/null @@ -1,58 +0,0 @@ -# Thin launcher: the domino_high_friction_turn system-ID experiment, run on -# the real scene's physical setup -- Franka Panda on its short pedestal and -# the extended table tile -- instead of the Fetch on a flat base. -# -# Usage: -# python scripts/local/launch_simp.py -c predicatorv3/exp_domino_real_geometry.yaml -# -# This is exp_domino.yaml with one env swapped. Tasks, friction mismatch -# (true 0.5 / believed 0.1), span and turn-leg bands, turn ratio, blue budget -# and reward semantics are identical to the domino_high_friction_turn arm, and -# the dominoes are the same simulated blocks. The robot and the table are the -# only things that move, so the two runs answer "does this result survive the -# real robot's kinematics?" and nothing else. -# -# Because the blocks did not change, the 2026-07-12 band calibration carries -# over unchanged -- those bands are a property of the blocks and the friction -# pair. Re-probe only if you change one of those: -# python scripts/domino_debug/probe_min_block_bands.py reach \ -# --frictions 0.1 0.5 --env pybullet_domino_real_geometry --robot panda -# -# This is NOT exp_domino_real.yaml. That one runs the pybullet_domino_real -# env, which discards the generated tasks and rebuilds a single task from a -# perceived scene JSON; the min-block machinery never runs there. This one -# keeps every generated task and changes only the physical setup. ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -FLAGS: - num_online_learning_cycles: 3 -ENVS: - # launch_simp runs the ENVS x APPROACHES cross-product - park arms here - # before launching if the full matrix is not wanted. - domino: - SKIP: True - domino_turns: - SKIP: True - domino_low_friction: - SKIP: True - domino_high_friction: - SKIP: True - # The simulated-setup arm this one is paired against. Un-skip it to run - # both and compare; parked by default so a launch is the real-setup arm - # alone. - domino_high_friction_turn: - SKIP: True - domino_high_friction_turn_real: - SKIP: False -APPROACHES: - agent_oracle_hybrid_sim: - SKIP: True - FLAGS: - agent_sim_learn_kept_predicates_names: ["Holding", "HandEmpty"] - sim_predicator: - SKIP: False - FLAGS: - skip_initial_test: True diff --git a/scripts/configs/predicatorv3/exp_domino_real_replay.yaml b/scripts/configs/predicatorv3/exp_domino_real_replay.yaml deleted file mode 100644 index 0f3093b46a..0000000000 --- a/scripts/configs/predicatorv3/exp_domino_real_replay.yaml +++ /dev/null @@ -1,65 +0,0 @@ -# Replay: the domino friction fit, scored against an ALREADY-RECORDED track. -# Usage: -# python scripts/local/launch_simp.py -c predicatorv3/exp_domino_real_replay.yaml -# -# Same experiment as exp_domino_real.yaml, with the two slow halves removed: -# the arm does not move and the cameras do not look. Everything downstream of -# the track -- id matching, interval residuals, the parameter sweep, and the -# agent's decision about what to declare -- runs exactly as it does live. -# -# WHY THIS EXISTS. One live episode costs a scene reset, ~110 s of arm motion -# and ~3 min of markerless post-processing, and it produces the same track -# every time the layout is the same. The objective is what has been changing, -# not the data, so iterating on the objective against a known-good recording -# is the loop that matters. run_20260820_123606 is that recording: four -# dominoes cascaded 3 -> 2 -> 1 -> 0, the twin reproduced all four, and the -# ids match 5 of 5 offline. -# -# WHAT STILL RUNS. The twin simulates the fixed plan, which is where the -# recorded trajectory comes from -- that half is pure PyBullet and was never -# the slow part. The arm is dry, so the plan's motion is a no-op at the -# hardware boundary, and perception is the captured scene file rather than a -# camera. The trajectory this produces is the same computation the live run -# performed, because under open-loop nothing corrects the twin mid-episode. -# -# WHAT THIS CANNOT TELL YOU. Whether the real world would have cascaded -# differently at a different friction. The track is fixed, so the replay -# answers "what does the fit do with this evidence", never "is the evidence -# right". Re-record when the scene or the plan changes. ---- -includes: - - exp_domino_real.yaml -ENVS: - domino_real: - FLAGS: - # -- the arm and the cameras, both off --------------------------------- - # Dry: no arm is built and arm calls are no-ops, so the plan is - # simulated and then dropped at the hardware boundary. The executor is - # still attached (real_robot_execute stays True) so the same shipping - # and batching path runs -- it just ships into nothing. - real_robot_dry: True - # "scene_file" replays domino_real_scene: cameraless, and it reports the - # captured layout, which is what the twin has to start from for its - # trajectory to match the recorded run's. - real_robot_perception: "scene_file" - # No takes, so no ZED session, no SVOs and no markerless pipeline. This - # also stops tracks.json being rewritten, which is what makes the frozen - # manifest below safe to point at. - real_robot_record_episodes: False - real_robot_process_takes: False - # Nothing to reset between episodes when nothing moved, and nothing to - # draw boxes on when no camera looked. Both of these BLOCK on a human - # (a terminal prompt and an OpenCV drag window), which would defeat the - # point of a replay. - real_robot_human_reset: False - real_robot_pick_boxes_at_start: False - real_robot_snapshot_rebuild: False - # -- the evidence ------------------------------------------------------ - # The frozen copy, NOT logs/zed_tracks/tracks.json: that file is - # rewritten by every live run, so a replay pointed at it would silently - # start scoring whatever was recorded most recently. - code_sim_learning_rollout_track_path: "logs/zed_tracks/replay_20260820_124013.json" - # The track is already on disk and complete, so there is nothing to wait - # for. Left long enough to be a real error rather than a hang if the - # manifest ever points somewhere wrong. - code_sim_learning_track_wait_s: 30 diff --git a/scripts/configs/predicatorv3/exp_domino_sweep.yaml b/scripts/configs/predicatorv3/exp_domino_sweep.yaml deleted file mode 100644 index fae7b1b2ea..0000000000 --- a/scripts/configs/predicatorv3/exp_domino_sweep.yaml +++ /dev/null @@ -1,118 +0,0 @@ -# DOMINO sweep over every paper arm (C7 dropped), 3 seeds each. -# Usage: python scripts/engaging/launch.py -c predicatorv3/exp_domino_sweep.yaml ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -NUM_SEEDS: 3 -ENVS: - domino_high_friction_turn: # exp_domino.yaml's env - SKIP: False - FLAGS: - num_test_tasks: 5 -APPROACHES: -# Every arm is tested after every learning cycle. Learning arms skip the pre-loop -# test; U1 and A2 (no cycles) are evaluated by it, and C1 keeps it as A1 (no learning). - # C1 - sim_predicator: - SKIP: False - ARGS: - - auto_resume - # C1 in policy mode (parked; the paper's C1 is plan mode). - sim_predicator_policy: - SKIP: True - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C2 - agent_model_free_planning: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C3 - nl_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C4 - code_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C5 - gnn_dynamics: - SKIP: False - FLAGS: - skip_initial_test: True - # C6 - operator_learning: - SKIP: False - FLAGS: - skip_initial_test: True - # C8 - maple_q: - SKIP: False - FLAGS: - skip_initial_test: True - # U1 (keeps HandEmpty like exp_domino.yaml) - agent_oracle_hybrid_sim: - SKIP: False - FLAGS: - agent_sim_learn_kept_predicates_names: ["Holding", "HandEmpty"] - ARGS: - - auto_resume - # A2 - sim_predicator_zero_shot: - SKIP: False - ARGS: - - auto_resume - # A3 - sim_predicator_undirected_explore: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A4 (the rule-param margin samples the declared intervals here) - sim_predicator_no_param_fit: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A5 - sim_predicator_no_validation: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A6 - sim_predicator_explore_no_disagreement: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A7 - sim_predicator_validation_no_uncertainty: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # A8 - sim_predicator_no_predicates: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume diff --git a/scripts/configs/predicatorv3/exp_fan.yaml b/scripts/configs/predicatorv3/exp_fan.yaml deleted file mode 100644 index 9764670636..0000000000 --- a/scripts/configs/predicatorv3/exp_fan.yaml +++ /dev/null @@ -1,30 +0,0 @@ -# Thin launcher: run the fan experiment (agent_oracle_hybrid_sim arm). -# Usage: python scripts/local/launch_simp.py -c predicatorv3/exp_fan.yaml --parallel -# Env definitions live in envs/all.yaml and approach definitions in -# approaches/all.yaml (all parked by default); this file only un-skips the -# env + arm(s) it runs. To run a baseline sweep, flip additional arms' -# SKIP to False here. ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -ENVS: - fan: - SKIP: False - # fan has no agent-specific excluded_predicates. The grid predicates - # (BallAtLoc / ClearLoc / SideOf / FanFacingSide / OppositeFan) are - # helper-only (injected for the oracle), so they are not in env.predicates - # and the agent already runs grid-free over the physical vocabulary. -APPROACHES: - agent_oracle_hybrid_sim: - SKIP: True - sim_predicator: - SKIP: False - FLAGS: - skip_initial_test: True - # Ablation axis: surface the fan env's base-sim source - # (pybullet_fan_base.py + pybullet_env.py) in the agent sandbox's - # ./reference/base_sim/. The wind dynamics, task generation, and - # goal semantics live in pybullet_fan.py, which is never provided. - agent_sim_provide_base_sim_source: True diff --git a/scripts/configs/predicatorv3/exp_fan_sweep.yaml b/scripts/configs/predicatorv3/exp_fan_sweep.yaml deleted file mode 100644 index 180a958887..0000000000 --- a/scripts/configs/predicatorv3/exp_fan_sweep.yaml +++ /dev/null @@ -1,129 +0,0 @@ -# FAN sweep over every paper arm (C7 dropped), 3 seeds each. -# Usage: python scripts/engaging/launch.py -c predicatorv3/exp_fan_sweep.yaml ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -NUM_SEEDS: 3 -ENVS: - fan: - SKIP: False - FLAGS: - num_test_tasks: 5 -APPROACHES: -# Every arm is tested after every learning cycle. Learning arms skip the pre-loop -# test; U1 and A2 (no cycles) are evaluated by it, and C1 keeps it as A1 (no learning). - # C1 (base-sim source surfaced, as in exp_fan.yaml, on every arm with a base sim) - sim_predicator: - SKIP: False - FLAGS: - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # C1 in policy mode (parked; the paper's C1 is plan mode). - sim_predicator_policy: - SKIP: True - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # C2 - agent_model_free_planning: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C3 - nl_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C4 - code_world_model: - SKIP: False - FLAGS: - skip_initial_test: True - ARGS: - - auto_resume - # C5 - gnn_dynamics: - SKIP: False - FLAGS: - skip_initial_test: True - # C6 - operator_learning: - SKIP: False - FLAGS: - skip_initial_test: True - # C8 - maple_q: - SKIP: False - FLAGS: - skip_initial_test: True - # U1 - agent_oracle_hybrid_sim: - SKIP: False - FLAGS: - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A2 - sim_predicator_zero_shot: - SKIP: False - FLAGS: - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A3 - sim_predicator_undirected_explore: - SKIP: False - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A4 (the rule-param margin samples the declared intervals here) - sim_predicator_no_param_fit: - SKIP: False - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A5 - sim_predicator_no_validation: - SKIP: False - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A6 - sim_predicator_explore_no_disagreement: - SKIP: False - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A7 - sim_predicator_validation_no_uncertainty: - SKIP: False - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume - # A8 - sim_predicator_no_predicates: - SKIP: False - FLAGS: - skip_initial_test: True - agent_sim_provide_base_sim_source: True - ARGS: - - auto_resume diff --git a/scripts/configs/predicatorv3/oracle.yaml b/scripts/configs/predicatorv3/oracle.yaml deleted file mode 100644 index f6a2c94000..0000000000 --- a/scripts/configs/predicatorv3/oracle.yaml +++ /dev/null @@ -1,30 +0,0 @@ -# Thin launcher: run the oracle arm. -# Usage: python scripts/local/launch_simp.py -c predicatorv3/oracle.yaml -# Approach definitions live in approaches/all.yaml and env definitions in -# envs/all.yaml (all parked by default); this file only un-skips the env + -# arm(s) it runs and sets oracle-specific ENVS overrides. ---- -includes: - - common.yaml - - envs/all.yaml - - approaches/all.yaml -ENVS: - # Pick the env for the oracle run: fan is active; flip SKIPs to run - # domino instead. - fan: - SKIP: False - # Oracle keeps restricted push (target inferred from state). The agent - # configs rely on the codebase default (False) so the LLM can name the - # trigger domino explicitly via Push(robot, domino). - domino: - FLAGS: - domino_restricted_push: True - domino_low_friction: - FLAGS: - domino_restricted_push: True - domino_high_friction: - FLAGS: - domino_restricted_push: True -APPROACHES: - oracle: - SKIP: False diff --git a/scripts/domino_debug/__init__.py b/scripts/domino_debug/__init__.py deleted file mode 100644 index e69de29bb2..0000000000 diff --git a/scripts/domino_debug/count_turns.py b/scripts/domino_debug/count_turns.py deleted file mode 100644 index d05a91c4c1..0000000000 --- a/scripts/domino_debug/count_turns.py +++ /dev/null @@ -1,91 +0,0 @@ -"""Count % of generated domino test tasks that contain a turn. - -Uses whatever DominoTaskGenerator is currently installed on disk, so it -can be run across git versions of the generator by swapping the file in -place. -""" -import numpy as np - -from predicators import utils -from predicators.envs.pybullet_domino.components.domino_component import \ - DominoComponent -from predicators.envs.pybullet_domino.env import PyBulletDominoEnv -from predicators.envs.pybullet_domino.task_generators import \ - domino_task_generator as dtg -from predicators.settings import CFG -from predicators.structs import Task - -N_TASKS = 40 -SEED = 0 - - -def ang_diff(a: float, b: float) -> float: - """Return the smallest unsigned angle between a and b (mod pi).""" - d = (a - b) % np.pi - return min(d, np.pi - d) - - -def is_turn(task: Task, comp: DominoComponent) -> bool: - """Return True if the task's start/target dominoes differ in yaw.""" - st = task.init - sy = ty = None - for d in st.get_objects(comp.domino_type): - r, g, b = st.get(d, "r"), st.get(d, "g"), st.get(d, "b") - if abs(r - comp.start_domino_color[0]) < 1e-2 and \ - abs(g - comp.start_domino_color[1]) < 1e-2: - sy = st.get(d, "yaw") - elif abs(r - comp.target_domino_color[0]) < 1e-2 and \ - abs(b - comp.target_domino_color[2]) < 1e-2: - ty = st.get(d, "yaw") - return sy is not None and ty is not None and \ - ang_diff(sy, ty) > np.deg2rad(30) - - -def main() -> None: - """Print the percentage of generated test tasks with a turn.""" - utils.reset_config({ - "env": "pybullet_domino", - "seed": SEED, - "num_train_tasks": 0, - "num_test_tasks": N_TASKS, - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_has_glued_dominos": False, - "domino_test_num_dominos": [3], - "domino_test_num_targets": [1, 2], - "domino_test_num_pivots": [0], - }) - env = PyBulletDominoEnv(use_gui=False) - comp = env._domino_component # pylint: disable=protected-access - assert comp is not None - robot_init_state = { - "x": env.robot_init_x, - "y": env.robot_init_y, - "z": env.robot_init_z, - "fingers": env.open_fingers, - "roll": env.robot_init_roll, - "tilt": env.robot_init_tilt, - "wrist": env.robot_init_wrist, - } - gen = dtg.DominoTaskGenerator( - domino_component=comp, - robot=env._robot, # pylint: disable=protected-access - robot_init_state=robot_init_state, - additional_components=[]) - # Reproduce the env's per-seed test path: seeds 0-4, 5 tasks each. - turns = total = 0 - for seed in range(5): - rng = np.random.default_rng(seed + 10000) - tasks = gen.generate_tasks( - num_tasks=5, - rng=rng, - possible_num_dominos=CFG.domino_test_num_dominos, - possible_num_targets=CFG.domino_test_num_targets, - possible_num_pivots=CFG.domino_test_num_pivots) - turns += sum(is_turn(t.task, comp) for t in tasks) - total += len(tasks) - print(f"RESULT turns={turns}/{total} = {100.0 * turns / total:.1f}%") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/measure_turn_diversity.py b/scripts/domino_debug/measure_turn_diversity.py deleted file mode 100644 index 28311c1dec..0000000000 --- a/scripts/domino_debug/measure_turn_diversity.py +++ /dev/null @@ -1,185 +0,0 @@ -"""Render the first 5 test tasks for seeds 0-4 (the exact tasks the agent run -used) and measure turn % with the grasp-clearance staging check (commit -50d56e940) ON ("after") vs OFF ("before"). - -Reproduces the env's generation path exactly: for seed N the test rng is -np.random.default_rng(N + CFG.test_env_seed_offset), 5 tasks per seed. - -A task is a TURN if the purple target block's yaw differs from the green start -block's yaw by > 30 deg (a straight chain shares one yaw mod pi; a turn90 chain -ends ~90 deg rotated). -""" -from typing import Any, List, Tuple - -import numpy as np -from matplotlib import pyplot as plt -from matplotlib.patches import Rectangle -from matplotlib.transforms import Affine2D - -from predicators import utils -from predicators.envs.pybullet_domino.env import PyBulletDominoEnv -from predicators.envs.pybullet_domino.task_generators import \ - domino_task_generator as dtg -from predicators.settings import CFG -from predicators.structs import EnvironmentTask - -SEEDS = [0, 1, 2, 3, 4] -TASKS_PER_SEED = 5 -OFFSET = 10000 # CFG.test_env_seed_offset - -DominoEntry = Tuple[float, float, float, str] - - -def ang_diff(a: float, b: float) -> float: - """Smallest angular difference modulo pi (radians).""" - d = (a - b) % np.pi - return min(d, np.pi - d) - - -def classify(task: EnvironmentTask, - comp: Any) -> Tuple[bool, List[DominoEntry]]: - """Return (is_turn, dominoes) for a task using start/target yaw.""" - st = task.init - dominoes, sy, ty = [], None, None - for d in st.get_objects(comp.domino_type): - r, g, b = st.get(d, "r"), st.get(d, "g"), st.get(d, "b") - x, y, yaw = st.get(d, "x"), st.get(d, "y"), st.get(d, "yaw") - if abs(r - comp.start_domino_color[0]) < 1e-2 and \ - abs(g - comp.start_domino_color[1]) < 1e-2: - role, sy = "start", yaw - elif abs(r - comp.target_domino_color[0]) < 1e-2 and \ - abs(b - comp.target_domino_color[2]) < 1e-2: - role, ty = "target", yaw - else: - role = "movable" - dominoes.append((x, y, yaw, role)) - is_turn = (sy is not None and ty is not None - and ang_diff(sy, ty) > np.deg2rad(30)) - return is_turn, dominoes - - -def build_generator(env: PyBulletDominoEnv) -> dtg.DominoTaskGenerator: - """Build a DominoTaskGenerator matching the env's config.""" - ris = { - "x": env.robot_init_x, - "y": env.robot_init_y, - "z": env.robot_init_z, - "fingers": env.open_fingers, - "roll": env.robot_init_roll, - "tilt": env.robot_init_tilt, - "wrist": env.robot_init_wrist, - } - # pylint: disable=protected-access - return dtg.DominoTaskGenerator( - domino_component=env._domino_component, # type: ignore[arg-type] - robot=env._robot, - robot_init_state=ris, - additional_components=[]) - - -def _no_block(*_a: Any, **_k: Any) -> bool: - """Stub replacement that never blocks grasp clearance.""" - return False - - -def gen_all(env: PyBulletDominoEnv, - disable_grasp: bool) -> List[Tuple[int, int, EnvironmentTask]]: - """Generate all (seed, task_idx, task) entries, optionally with the grasp- - clearance staging check disabled.""" - gen = build_generator(env) - cls = dtg.DominoTaskGenerator - # pylint: disable=protected-access - orig = cls._grasp_clearance_blocked - if disable_grasp: - cls._grasp_clearance_blocked = _no_block # type: ignore[method-assign] - try: - out: List[Tuple[int, int, EnvironmentTask]] = [] - for seed in SEEDS: - rng = np.random.default_rng(seed + OFFSET) - tasks = gen.generate_tasks( - num_tasks=TASKS_PER_SEED, - rng=rng, - possible_num_dominos=CFG.domino_test_num_dominos, - possible_num_targets=CFG.domino_test_num_targets, - possible_num_pivots=CFG.domino_test_num_pivots) - for ti, t in enumerate(tasks): - out.append((seed, ti, t)) - finally: - cls._grasp_clearance_blocked = orig # type: ignore[method-assign] - return out - - -def render(entries: List[Tuple[int, int, EnvironmentTask]], comp: Any, - title: str, path: str) -> None: - """Render a grid of task scenes and report the turn percentage.""" - nrows, ncols = len(SEEDS), TASKS_PER_SEED - fig, axes = plt.subplots(nrows, ncols, figsize=(2.4 * ncols, 2.4 * nrows)) - w, dpth = comp.domino_width, comp.domino_depth - cmap = {"start": "#2ca02c", "movable": "#6699ff", "target": "#cc66cc"} - by_key = {(s, ti): t for (s, ti, t) in entries} - turns = 0 - for r, seed in enumerate(SEEDS): - for c in range(TASKS_PER_SEED): - ax = axes[r][c] - t = by_key.get((seed, c)) - if t is None: - ax.axis("off") - continue - is_turn, dominoes = classify(t, comp) - turns += is_turn - for (x, y, yaw, role) in dominoes: - rect = Rectangle((-w / 2, -dpth / 2), - w, - dpth, - color=cmap[role]) - rect.set_transform(Affine2D().rotate(yaw).translate(x, y) + - ax.transData) - ax.add_patch(rect) - ax.set_xlim(comp.domino_x_lb - 0.05, comp.domino_x_ub + 0.05) - ax.set_ylim(comp.domino_y_lb - 0.05, comp.domino_y_ub + 0.05) - ax.set_aspect("equal") - ax.set_xticks([]) - ax.set_yticks([]) - ax.set_title( - f"seed{seed} t{c}: {'TURN' if is_turn else 'straight'}", - fontsize=9, - color="red" if is_turn else "black") - n = len(entries) - fig.suptitle( - f"{title} — {turns}/{n} turns ({100.0*turns/n:.0f}%)\n" - "green=start blue=movable(staged) purple=target", - fontsize=13) - fig.tight_layout(rect=[0, 0, 1, 0.97]) - fig.savefig(path, dpi=95) - plt.close(fig) - print(f" {title}: {turns}/{n} turns = {100.0*turns/n:.1f}% -> {path}") - - -def main() -> None: - """Generate, render, and compare turn % before/after the staging check.""" - utils.reset_config({ - "env": "pybullet_domino", - "seed": 0, - "num_train_tasks": 0, - "num_test_tasks": TASKS_PER_SEED, - "test_env_seed_offset": OFFSET, - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_has_glued_dominos": False, - "domino_test_num_dominos": [3], - "domino_test_num_targets": [1, 2], - "domino_test_num_pivots": [0], - }) - env = PyBulletDominoEnv(use_gui=False) - comp = env._domino_component # pylint: disable=protected-access - base = "/Users/ycliang/Code/predicators/scripts/domino_debug/" - for label, disable, fn in [ - ("AFTER (grasp-clearance ON = commit 50d56e940)", False, "after"), - ("BEFORE (grasp-clearance OFF = parent 50d56e940~1)", True, "before"), - ]: - entries = gen_all(env, disable) - render(entries, comp, label, f"{base}turn_diversity_{fn}.png") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/probe_cascade.py b/scripts/domino_debug/probe_cascade.py deleted file mode 100644 index 26d7c0dff1..0000000000 --- a/scripts/domino_debug/probe_cascade.py +++ /dev/null @@ -1,113 +0,0 @@ -"""Locate where a domino cascade dies: run the recorded sketch -(Pick->Place->Push->Wait) through the REAL option model and log every domino's -roll (topple angle) after each step. - -Usage: - PYTHONPATH=. python scripts/domino_debug/probe_cascade.py \ - -""" -import logging -import sys -from typing import List - -import numpy as np - -from predicators import utils -from predicators.approaches import create_approach -from predicators.approaches.agent_sim_learning_approach import \ - AgentSimLearningApproach -from predicators.envs import get_or_create_env -from predicators.ground_truth_models import get_gt_options -from predicators.ground_truth_models.domino import processes as P -from predicators.structs import Object, State -from scripts.domino_debug.replay_domino_sketches import _FLAGS - -logging.disable(logging.CRITICAL) - - -def rolls(state: State, dominoes: List[Object]) -> str: - """Format each domino's roll angle for one-line logging.""" - return " ".join(f"{d.name}:r={state.get(d,'roll'):+.3f}" - for d in dominoes) - - -def main() -> None: - """Probe a recorded cascade step-by-step via the real option model.""" - seed, ti = int(sys.argv[1]), int(sys.argv[2]) - utils.reset_config(dict(_FLAGS, seed=seed)) - - env = get_or_create_env("pybullet_domino") - options = get_gt_options(env.get_name()) - preds, _ = utils.parse_config_excluded_predicates(env) - train_tasks = [t.task for t in env.get_train_tasks()] - approach = create_approach("agent_sim_learning", preds, options, env.types, - env.action_space, train_tasks) - assert isinstance(approach, AgentSimLearningApproach) - # pylint: disable=protected-access - approach._maybe_install_oracle_samplers() - om = approach._option_model - # pylint: enable=protected-access - assert om is not None - opt = {o.name: o for o in options} - InFront = [p for p in preds if p.name == "InFront"][0] - Toppled = [p for p in preds if p.name == "Toppled"][0] - - task = env.get_test_tasks()[ti].task - state = task.init - dominoes = sorted([o for o in state if o.type.name == "domino"], - key=lambda o: o.name) - robot = [o for o in state if o.type.name == "robot"][0] - fallen = env.fallen_threshold if hasattr(env, "fallen_threshold") else None - print(f"# seed{seed} env-task{ti}: {len(dominoes)} dominoes, " - f"fallen_threshold={fallen}") - for d in dominoes: - print(f" {d.name}: x={state.get(d,'x'):.4f} y={state.get(d,'y'):.4f} " - f"yaw={state.get(d,'yaw'):+.4f}") - print(f" goal: {sorted(str(a) for a in task.goal)}") - print(f"\ninit {rolls(state, dominoes)}") - - held = dominoes[1] - ref0, ref1 = dominoes[0], dominoes[3] - sub = { - utils.GroundAtom(InFront, [held, ref0]), - utils.GroundAtom(InFront, [ref1, held]) - } - - # pylint: disable=protected-access - # Pick(d1) - pick_params = P._pick_option_sampler(state, set(), - np.random.default_rng(0), - [robot, held]) - s = om.get_next_state_and_num_actions( - state, opt["Pick"].ground([robot, held], pick_params))[0] - # Place(d1) at generator-faithful pose for the two InFront subgoals - pp = P._place_option_sampler(s, sub, np.random.default_rng(3), [robot]) - s = om.get_next_state_and_num_actions(s, opt["Place"].ground([robot], - pp))[0] - print(f"after Place {rolls(s, dominoes)} " - f"(d1 placed at {pp[0]:.3f},{pp[1]:.3f},yaw={pp[3]:+.3f})") - print(" InFront subgoals holding: " + - str([str(a) for a in sub if a.holds(s)])) - - # Push - push = opt["Push"] - pparams = P._push_option_sampler(s, set(), np.random.default_rng(0), - [robot]) - # pylint: enable=protected-access - pg = push.ground([robot], pparams) if len(push.types) == 1 else \ - push.ground([robot, dominoes[0]], pparams) - s = om.get_next_state_and_num_actions(s, pg)[0] - print(f"after Push {rolls(s, dominoes)}") - - # Wait - wait = opt["Wait"] - wparams = np.zeros(wait.params_space.shape[0], dtype=np.float32) - s = om.get_next_state_and_num_actions(s, wait.ground([robot], wparams))[0] - print(f"after Wait {rolls(s, dominoes)}") - print("\nToppled after Wait: " + - str({d.name: Toppled.holds(s, [d]) - for d in dominoes})) - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/probe_infront_drift.py b/scripts/domino_debug/probe_infront_drift.py deleted file mode 100644 index 2bde9c28b9..0000000000 --- a/scripts/domino_debug/probe_infront_drift.py +++ /dev/null @@ -1,116 +0,0 @@ -"""Measure InFront settle-drift: place a domino at the oracle sampler's -generator-faithful nominal pose through the REAL option model (PyBullet forward -sim), then compare the settled pose to the nominal one and to the InFront -tolerance window. - -Usage: PYTHONPATH=. python \ - scripts/domino_debug/probe_infront_drift.py -""" -import logging -import sys - -import numpy as np - -from predicators import utils -from predicators.approaches import create_approach -from predicators.envs import get_or_create_env -from predicators.ground_truth_models import get_gt_options -from predicators.ground_truth_models.domino import processes as P - -logging.disable(logging.CRITICAL) - -# Same flags the replay uses so the option model + oracle samplers match. -# pylint: disable=wrong-import-position -from scripts.domino_debug.replay_domino_sketches import _FLAGS # noqa: E402 - - -def main() -> None: - """Probe InFront settle-drift for one seed/test-task index.""" - seed = int(sys.argv[1]) - ti = int(sys.argv[2]) - - utils.reset_config(dict(_FLAGS, seed=seed)) - - env = get_or_create_env("pybullet_domino") - options = get_gt_options(env.get_name()) - preds, _ = utils.parse_config_excluded_predicates(env) - train_tasks = [t.task for t in env.get_train_tasks()] - approach = create_approach("agent_sim_learning", preds, options, env.types, - env.action_space, train_tasks) - # pylint: disable=protected-access - approach._maybe_install_oracle_samplers() # type: ignore[attr-defined] - om = approach._option_model # type: ignore[attr-defined] - assert om is not None, "no option model" - - task = env.get_test_tasks()[ti].task - state = task.init - dominoes = sorted([o for o in state if o.type.name == "domino"], - key=lambda o: o.name) - robot = [o for o in state if o.type.name == "robot"][0] - print(f"# seed{seed} env-task{ti}: {len(dominoes)} dominoes") - for d in dominoes: - print(f" {d.name}: x={state.get(d,'x'):.4f} y={state.get(d,'y'):.4f} " - f"yaw={state.get(d,'yaw'):+.4f} roll={state.get(d,'roll'):+.4f}") - - InFront = [p for p in preds if p.name == "InFront"][0] - infront_holds = InFront.holds - pos_gap = 0.098 # PyBulletDominoEnv.pos_gap - pos_tol = pos_gap * 0.3 - print(f"\n# InFront window: pos_tol={pos_tol:.4f} m, ang_tol=15deg, " - f"pos_gap={pos_gap:.4f}") - - opt_by_name = {o.name: o for o in options} - Pick, Place = opt_by_name["Pick"], opt_by_name["Place"] - - # Place domino_1 to satisfy InFront(domino_1, domino_0). - held, ref = dominoes[1], dominoes[0] - sub = {utils.GroundAtom(InFront, [held, ref])} - - # 1) Pick the held domino. - # pylint: disable=protected-access - pick_params = P._pick_option_sampler(state, set(), - np.random.default_rng(0), - [robot, held]) - pick_opt = Pick.ground([robot, held], pick_params) - assert pick_opt.initiable(state), "pick not initiable" - s1, _ = om.get_next_state_and_num_actions(state, pick_opt) - print(f"\n# after Pick({held.name}): is_held={s1.get(held,'is_held'):.2f}") - - # 2) Sample the generator-faithful placement and run Place. - for trial in range(5): - rng = np.random.default_rng(100 + trial) - place_params = P._place_option_sampler(s1, sub, rng, [robot]) - nom_x, nom_y, _, nom_yaw = [float(v) for v in place_params] - place_opt = Place.ground([robot], place_params) - if not place_opt.initiable(s1): - print(f" trial{trial}: place not initiable") - continue - s2, _ = om.get_next_state_and_num_actions(s1, place_opt) - gx, gy, gyaw = (s2.get(held, "x"), s2.get(held, - "y"), s2.get(held, "yaw")) - groll = s2.get(held, "roll") - rx, ry, ryaw = (s2.get(ref, "x"), s2.get(ref, "y"), s2.get(ref, "yaw")) - infront = infront_holds(s2, [held, ref]) - # nominal InFront check (kinematic, roll=0, exactly at sampler pose) - snom = s2.copy() - snom.set(held, "x", nom_x) - snom.set(held, "y", nom_y) - snom.set(held, "yaw", nom_yaw) - snom.set(held, "roll", 0.0) - nom_infront = infront_holds(snom, [held, ref]) - print(f"\n trial{trial}: nominal place=({nom_x:.4f},{nom_y:.4f}," - f"yaw={nom_yaw:+.4f}) nominal_InFront={nom_infront}") - print(f" settled {held.name}=({gx:.4f},{gy:.4f},yaw={gyaw:+.4f}," - f"roll={groll:+.4f})") - print(f" drift dx={gx-nom_x:+.4f} dy={gy-nom_y:+.4f} " - f"dyaw={gyaw-nom_yaw:+.4f} (pos_tol={pos_tol:.4f})") - tol10 = np.sin(np.radians(10)) - is_cardinal = (abs(np.sin(ryaw)) < tol10 or abs(np.cos(ryaw)) < tol10) - cardinal = "yes" if is_cardinal else "NO" - print(f" {ref.name}=({rx:.4f},{ry:.4f},yaw={ryaw:+.4f}) " - f"cardinal={cardinal}") - print(f" => settled InFront({held.name},{ref.name}) = {infront}") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/probe_min_block_bands.py b/scripts/domino_debug/probe_min_block_bands.py deleted file mode 100644 index d8bd6478f6..0000000000 --- a/scripts/domino_debug/probe_min_block_bands.py +++ /dev/null @@ -1,180 +0,0 @@ -"""Anchor probes for calibrating min-block task bands (spans / turn legs). - -The min-block differentiation bands are friction-pair specific: straight -spans need a window where the true friction's chain count is below the -planning friction's, and turn legs need cells where a natural corner -tops at the true friction while the believed side needs an extra blue. -This script measures both at the canonical probe anchor with the SAME -machinery task generation uses (memoized straight probes, the labeled -turn-layout family search, real Push rollouts), so its numbers transfer -to the generator's certificates. - -Used for the 2026-07-12 domino_high_friction short-leg retune (see the -env block comments in scripts/configs/predicatorv3/envs/all.yaml). Run -it whenever domino_true_friction / domino_planning_friction change: - - python scripts/domino_debug/probe_min_block_bands.py reach \ - --frictions 0.1 0.5 --span-lo 0.12 --span-hi 0.64 - python scripts/domino_debug/probe_min_block_bands.py turn \ - --frictions 0.1 0.5 --cells 0.22,0.18 0.23,0.19 --reps 3 - -Reading the output: - * reach: pick a span window where k(true) < k(planning) is stable - across --reps (repeats clear the probe memo, sampling solver-history - variance; a span whose count flips between rounds is knife-edge). - * turn: per (entry, exit) cell and friction, the first k with a - toppling layout plus per-family topple counts. "corner" (single - natural-yaw corner blue) is the agent-buildable style - a cell whose - only topplers are "pair" (the legacy 45-degree pair) is - agent-intractable and must NOT ship (that was the pre-retune - high_friction failure). On the planning-friction side, prefer cells - whose k ties are impossible: believed k should exceed the true k in - EVERY rep (single-rep believed reads flicker ~1/3 on knife-edge - cells, which is why the generator re-runs its believed certificates - twice post-staging). -""" -import argparse -import json -import time -from typing import Any, Dict, List, Sequence, Tuple - -import numpy as np - -from predicators import utils - - -def _build_env() -> Any: - """Env with the min-block flags that affect probe physics.""" - utils.reset_config({ - "env": "pybullet_domino", - "approach": "oracle", - "seed": 0, - "use_gui": False, - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_use_continuous_place": True, - "domino_has_glued_dominos": False, - "domino_use_skill_factories": True, - "skill_phase_use_motion_planning": True, - "pybullet_ik_validate": False, - "pybullet_birrt_extend_num_interp": 20, - "pybullet_birrt_path_subsample_ratio": 2, - "domino_min_block_tasks": True, - "horizon": 500, - }) - # pylint: disable-next=import-outside-toplevel - from predicators.envs.pybullet_domino.env import PyBulletDominoEnv - return PyBulletDominoEnv(use_gui=False) - - -def _turn_poses(mbu: Any, entry: float, exit_: float) -> Tuple[Any, Any]: - """Left-turn L at the canonical anchor, entry along +x (like the straight - probes).""" - sx, sy = mbu._PROBE_ANCHOR # pylint: disable=protected-access - syaw = np.pi / 2 - u = np.array([np.sin(syaw), np.cos(syaw)]) - p = np.array([-u[1], u[0]]) - t = np.array([sx, sy]) + entry * u + exit_ * p - tyaw = float(np.arctan2(-p[0], p[1])) - return (sx, sy, syaw), (float(t[0]), float(t[1]), tyaw) - - -def probe_reach(env: Any, mbu: Any, frictions: Sequence[float], - spans: Sequence[float], budget: int, - reps: int) -> Dict[str, List[Any]]: - """k = straight_span_k_star per (friction, span), ``reps`` rounds with - the memo cleared between rounds.""" - results: Dict[str, List[Any]] = {} - for rep in range(reps): - mbu._span_probe_memo.clear() # pylint: disable=protected-access - for f in frictions: - env.set_domino_physical_params(lateral_friction=f) - for span in spans: - k = mbu.straight_span_k_star(env, span, budget=budget) - results.setdefault(f"{f}|{span:.2f}", []).append(k) - print(f"reach rep{rep} f={f} span={span:.2f} -> k={k}", - flush=True) - return results - - -def probe_turn(env: Any, mbu: Any, frictions: Sequence[float], - cells: Sequence[Tuple[float, ...]], budget: int, - reps: int) -> Dict[str, List[Any]]: - """First k with a toppler per (friction, cell), with per-family topple - counts from full k-layer scans, ``reps`` times each.""" - comp = env._domino_component # pylint: disable=protected-access - results: Dict[str, List[Any]] = {} - for rep in range(reps): - for f in frictions: - env.set_domino_physical_params(lateral_friction=f) - for entry, exit_ in cells: - sp, tp = _turn_poses(mbu, entry, exit_) - push_opt = mbu._get_push_option(env) # pylint: disable=protected-access - t0 = time.time() - first_k, layers = None, [] - for k in range(budget + 1): - fams: Dict[str, int] = {} - # pylint: disable-next=protected-access - for fam, od, s, t in mbu._candidate_turn_layouts_labeled( - comp, k, sp, tp): - if mbu._layout_topples(env, od, s, t, push_opt): # pylint: disable=protected-access - fams[fam] = fams.get(fam, 0) + 1 - layers.append({"k": k, "topples": fams}) - if fams: - first_k = k - break # the K* layer is fully scanned; stop - results.setdefault(f"{f}|{entry}|{exit_}", []).append({ - "k": - first_k, - "layers": - layers, - }) - print( - f"turn rep{rep} f={f} legs=({entry},{exit_}) -> " - f"k={first_k} ({time.time() - t0:.1f}s) " - f"families={layers[-1]['topples']}", - flush=True) - return results - - -def _main() -> None: - parser = argparse.ArgumentParser(description=__doc__) - sub = parser.add_subparsers(dest="mode", required=True) - common = argparse.ArgumentParser(add_help=False) - common.add_argument("--frictions", type=float, nargs="+", required=True) - common.add_argument("--budget", type=int, default=5) - common.add_argument("--reps", type=int, default=1) - common.add_argument("--out", type=str, default="") - reach = sub.add_parser("reach", parents=[common]) - reach.add_argument("--span-lo", type=float, default=0.12) - reach.add_argument("--span-hi", type=float, default=0.64) - reach.add_argument("--span-step", type=float, default=0.04) - turn = sub.add_parser("turn", parents=[common]) - turn.add_argument("--cells", - type=str, - nargs="+", - required=True, - help="entry,exit leg pairs, e.g. 0.22,0.18") - args = parser.parse_args() - - env = _build_env() - # pylint: disable-next=import-outside-toplevel - from predicators.envs.pybullet_domino.task_generators import \ - min_block_utils as mbu - if args.mode == "reach": - n = int(round((args.span_hi - args.span_lo) / args.span_step)) + 1 - spans = [round(args.span_lo + i * args.span_step, 2) for i in range(n)] - results = probe_reach(env, mbu, args.frictions, spans, args.budget, - args.reps) - else: - cells = [tuple(float(v) for v in c.split(",")) for c in args.cells] - results = probe_turn(env, mbu, args.frictions, cells, args.budget, - args.reps) - if args.out: - with open(args.out, "w", encoding="utf-8") as fh: - json.dump(results, fh, indent=1) - print(f"saved {args.out}") - - -if __name__ == "__main__": - _main() diff --git a/scripts/domino_debug/probe_real_scene.py b/scripts/domino_debug/probe_real_scene.py deleted file mode 100644 index 37ec17534e..0000000000 --- a/scripts/domino_debug/probe_real_scene.py +++ /dev/null @@ -1,287 +0,0 @@ -"""Execute domino SKILLS on the real-bench scene IN SIMULATION and save an MP4. - -This is a sim-only debug harness. It reproduces the stock config -(``predicatorv3/exp_domino_real.yaml``) so the env, bench_setup patches, -geometry and scene match a real run exactly, grounds a hand-specified skill -sketch (Pick / Place / Push / Wait) with the domino oracle samplers, then -rolls it out through the env via ``env.step`` -- the same path -``run_testing`` uses to render its test videos -- capturing one frame per -low-level action into an MP4. - -Use it to watch a skill's arm motion + physics on the real scene and iterate on -skill geometry (grasp height, push pose, ...) without the LLM planner. - -Usage (from the predicators repo root, robot-ml env; set PYTHONHASHSEED=0): - PYTHONPATH=. python scripts/domino_debug/probe_real_scene.py - # custom sketch + a different scene + live GUI: - PYTHONPATH=. python scripts/domino_debug/probe_real_scene.py \ - --sketch "Pick:1 Place:1@6 Push:start Wait" --gui - -Sketch grammar (space-separated ``Skill[:obj[@ref]]`` tokens): - Push[:obj] topple a domino (obj: a domino id / "start" / "target"; - default "start"). Non-restricted Push grounds [robot, obj]. - Pick:obj grasp a domino. - Place:obj[@ref] place the held domino; @ref adds an InFront(obj, ref) - subgoal so the placer aims for it (else an empty goal). - Wait let the physics settle (the cascade). -Object refs: a domino id from the scene JSON, or "start"/"target" (by role). -""" -import argparse -import json -import logging -import os -from typing import Any, Callable, Dict, List, Optional, Tuple - -import numpy as np - -from predicators import utils -from predicators.envs import get_or_create_env -from predicators.ground_truth_models import get_gt_options -from predicators.ground_truth_models.domino import processes as P -from predicators.settings import CFG -from predicators.structs import Object, Predicate, State, _Option -from scripts.cluster_utils import generate_run_configs - -# This is a debug harness that deliberately pokes env / component / sampler -# internals to drive skills directly, so protected access is expected. -# pylint: disable=protected-access - - -def _load_config(config: str, scene: str | None, seed: int) -> None: - """reset_config from the stock launcher config (single source of truth), - optionally overriding the scene JSON.""" - rc = list(generate_run_configs(config))[0] - flags = dict(rc.flags) - flags.update({"env": rc.env, "approach": rc.approach, "seed": seed}) - flags.pop("log", None) # launcher arg, not a CFG flag - if scene is not None: - flags["domino_real_scene"] = scene - utils.reset_config(flags) - - -def _build_resolver( - env: Any, state: State -) -> Tuple[Callable[[str], Object], Object, Object, List[Object]]: - """id / 'start' / 'target' -> the predicators domino Object. - - Dominoes are placed in scene order, so scene index i -> object - ``domino_i``. - """ - comp = env._domino_component - with open(CFG.domino_real_scene, encoding="utf-8") as f: - scene_ids = [d["id"] for d in json.load(f)["dominoes"]] - id_to_obj = {sid: comp.dominos[i] for i, sid in enumerate(scene_ids)} - dominoes = sorted([o for o in state if o.type.name == "domino"], - key=lambda o: o.name) - start = next(d for d in dominoes if comp._StartBlock_holds(state, [d])) - target = next(d for d in dominoes if comp._TargetDomino_holds(state, [d])) - - def resolve(ref: str) -> Object: - if ref == "start": - return start - if ref == "target": - return target - return id_to_obj[int(ref)] - - return resolve, start, target, dominoes - - -def _ground_token(tok: str, state: State, opt: Dict[str, Any], - in_front: Optional[Predicate], robot: Object, - resolve: Callable[[str], Object], - rng: np.random.Generator) -> _Option: - """Ground one sketch token, sampling its parameters on ``state``. - - ``state`` is the state this option will actually start from, not the - episode's initial state -- see ``_lazy_option_policy``. - """ - name, _, rest = tok.partition(":") - objref, _, ref = rest.partition("@") - if name == "Pick": - d = resolve(objref) - params = P._pick_option_sampler(state, set(), rng, [robot, d]) - return opt["Pick"].ground([robot, d], params) - if name == "Place": - place_d = resolve(objref) if objref else None - goal = set() - if ref and in_front is not None and place_d is not None: - goal = {utils.GroundAtom(in_front, [place_d, resolve(ref)])} - params = P._place_option_sampler(state, goal, rng, [robot]) - return opt["Place"].ground([robot], params) - if name == "Push": - d = resolve(objref) if objref else resolve("start") - params = P._push_option_sampler(state, set(), rng, [robot]) - push = opt["Push"] - return (push.ground([robot], params) - if len(push.types) == 1 else push.ground([robot, d], params)) - if name == "Wait": - wait = opt["Wait"] - params = np.zeros(wait.params_space.shape[0], dtype=np.float32) - return wait.ground([robot], params) - raise ValueError(f"unknown skill in sketch: {name!r}") - - -def _lazy_option_policy(sketch: str, env: Any, robot: Object, - resolve: Callable[[str], Object], seed: int, - recorded: List[_Option]) -> Callable[[State], _Option]: - """Ground the sketch one token at a time, each against the live state.""" - tokens = sketch.split() - opt = {o.name: o for o in get_gt_options(env.get_name())} - in_front = next((p for p in env.predicates if p.name == "InFront"), None) - rng = np.random.default_rng(seed) - index = 0 - - def _option_policy(state: State) -> _Option: - nonlocal index - if index >= len(tokens): - # The rollout ends here rather than erroring - raise utils.OptionExecutionFailure("sketch exhausted") - tok = tokens[index] - index += 1 - ground = _ground_token(tok, state, opt, in_front, robot, resolve, rng) - print(f"# ground : {tok} -> {ground.simple_str()}" - f"[{', '.join(f'{float(p):.4f}' for p in ground.params)}]") - # Append the grounded option to the plan - recorded.append(ground) - return ground - - return _option_policy - - -def _dump_plan(path: str, plan: List[_Option], header: List[str]) -> None: - """Write the grounded plan in ``replay_plan.py``'s text format. - - The continuous parameters here came out of the oracle samplers, so this - file is the only record of the exact numbers that were just watched - working. ``replay_plan`` re-grounds them verbatim, which is the whole - point: the plan that reaches the Franka is the plan that was verified in - simulation, not a fresh sample that merely came from the same sampler. - - ``_Option.simple_str`` is deliberately parameter-free, so the line format - is built here. It has to satisfy ``replay_plan._LINE``, i.e. - ``Name(objs)[nums]``. - """ - lines = [f"# {h}" for h in header] - for opt in plan: - objs = ", ".join(f"{o.name}:{o.type.name}" for o in opt.objects) - params = ", ".join(f"{float(p):.6f}" for p in opt.params) - lines.append(f"{opt.name}({objs})[{params}]") - os.makedirs(os.path.dirname(os.path.abspath(path)), exist_ok=True) - with open(path, "w", encoding="utf-8") as f: - f.write("\n".join(lines) + "\n") - - -def main() -> None: - """Parse args, roll out the sketch on the real scene, and save the MP4.""" - ap = argparse.ArgumentParser( - description=__doc__, - formatter_class=argparse.RawDescriptionHelpFormatter) - ap.add_argument("--config", - default="predicatorv3/exp_domino_real.yaml", - help="launcher config to reproduce (env + flags + scene).") - ap.add_argument("--scene", - default=None, - help="override CFG.domino_real_scene (a capture JSON).") - ap.add_argument("--sketch", - default="Push:start Wait", - help="skill sequence to execute (see module docstring).") - ap.add_argument("--seed", type=int, default=0) - ap.add_argument("--gui", - action="store_true", - help="also open a live PyBullet window while rendering.") - ap.add_argument("--max-steps", - type=int, - default=1500, - help="max low-level env steps (Wait can be long).") - ap.add_argument("--frame-stride", - type=int, - default=2, - help="keep every Nth rendered frame in the MP4.") - ap.add_argument("--out", - default=None, - help="output mp4 path (default: " - "logs/probe_real_scene/_.mp4).") - ap.add_argument("--dump-plan", - default=None, - help="also write the grounded plan (with the sampled " - "parameters) in replay_plan.py's format, so this exact " - "rollout can be shipped to the real arm.") - args = ap.parse_args() - logging.basicConfig(level=logging.INFO, format="%(message)s") - - _load_config(args.config, args.scene, args.seed) - if args.gui: - CFG.use_gui = True - - env = get_or_create_env(CFG.env) - preds, _ = utils.parse_config_excluded_predicates(env) - task = env.get_test_tasks()[0].task - state = task.init - robot = next(o for o in state if o.type.name == "robot") - resolve, start, target, dominoes = _build_resolver(env, state) - Toppled = next(p for p in preds if p.name == "Toppled") - - print(f"# scene : {os.path.basename(CFG.domino_real_scene)} " - f"({len(dominoes)} dominoes)") - print(f"# roles : start={start.name} target={target.name}") - print(f"# sketch : {args.sketch}") - print(f"# goal : {sorted(str(a) for a in task.goal)}") - - # Filled in as the rollout grounds each token; empty until then, which is - # why the grounded plan is printed per option above rather than up front. - plan: List[_Option] = [] - # Wait terminates on an abstract-atom change (CFG.wait_option_terminate_ - # on_atom_change), so the policy needs an abstract function over the preds. - policy = utils.option_policy_to_policy( - _lazy_option_policy(args.sketch, env, robot, resolve, args.seed, plan), - abstract_function=lambda s: utils.abstract(s, preds)) - monitor = utils.VideoMonitor(env.render) - traj, _ = utils.run_policy( - policy, - env, - "test", - 0, - termination_function=lambda s: False, - max_num_steps=args.max_steps, - exceptions_to_break_on={utils.OptionExecutionFailure}, - monitor=monitor) - - final = traj.states[-1] - toppled = {d.name: bool(Toppled.holds(final, [d])) for d in dominoes} - solved = bool(all(a.holds(final) for a in task.goal)) - print(f"# steps : {len(traj.actions)}") - print(f"# toppled : {toppled}") - print(f"# solved : {solved}") - - if args.dump_plan: - # The verdict rides along in the header so a plan file found later - # still says what it did in sim. replay_plan skips '#' lines. - _dump_plan(args.dump_plan, plan, [ - f"scene : {CFG.domino_real_scene}", - f"sketch : {args.sketch}", - f"seed : {args.seed}", - f"steps : {len(traj.actions)}", - f"toppled : {toppled}", - f"solved : {solved}", - "replay : python scripts/domino_debug/replay_plan.py " - f"--plan {args.dump_plan} --scene {CFG.domino_real_scene}", - ]) - print(f"# plan : {args.dump_plan}") - - video = monitor.get_video()[::max(1, args.frame_stride)] - # save_video writes to CFG.video_dir/ (default videos/) and only - # creates video_dir itself, so make the nested subdir first. An absolute - # --out still works (os.path.join ignores video_dir for an absolute path). - out = args.out or os.path.join( - "probe_real_scene", - f"{os.path.splitext(os.path.basename(CFG.domino_real_scene))[0]}" - f"__{args.sketch.replace(' ', '_').replace(':', '-')}.mp4") - resolved = os.path.join(CFG.video_dir, out) - os.makedirs(os.path.dirname(resolved), exist_ok=True) - utils.save_video(out, video) - print( - f"# saved : {resolved} ({len(video)} frames @ {CFG.video_fps} fps)") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/render_domino_initial_states.py b/scripts/domino_debug/render_domino_initial_states.py deleted file mode 100644 index 05fa39ee9d..0000000000 --- a/scripts/domino_debug/render_domino_initial_states.py +++ /dev/null @@ -1,83 +0,0 @@ -"""Render initial states of the domino test tasks for debugging. - -Reproduces the exact test tasks from a run (same env, seed, -test_env_seed_offset and domino flags) and saves a PNG of each test -task's initial state so failed tasks can be visualized. - -Usage: - PYTHONPATH=. python scripts/domino_debug/render_domino_initial_states.py -""" -import os - -import numpy as np -from PIL import Image - -from predicators import utils -from predicators.envs import create_new_env - -# Domino flags copied verbatim from the run namespace (info.log) so the -# generated test tasks match the run exactly. -_DOMINO_FLAGS = { - "env": "pybullet_domino", - "num_train_tasks": 1, - "num_test_tasks": 5, - "test_env_seed_offset": 10000, - "pybullet_camera_width": 1340, - "pybullet_camera_height": 720, - "domino_test_num_dominos": [3], - "domino_test_num_targets": [1, 2], - "domino_test_num_pivots": [0], - "domino_test_num_pos_x": 4, - "domino_test_num_pos_y": 3, - "domino_train_num_dominos": [2], - "domino_train_num_targets": [1], - "domino_train_num_pivots": [0], - "domino_train_num_pos_x": 3, - "domino_train_num_pos_y": 2, - "domino_use_continuous_place": True, - "domino_use_domino_blocks_as_target": True, - "domino_restricted_push": True, - "domino_only_straight_sequence_in_training": True, - "domino_use_skill_factories": True, - "domino_prune_actions": False, - "domino_has_glued_dominos": False, - "domino_some_dominoes_are_connected": False, - "domino_include_connected_predicate": False, - "domino_initialize_at_finished_state": False, - "domino_debug_layout": False, - "domino_domino_on_stairs": False, -} - -# Which 1-indexed test tasks failed in each seed (for labeling). -_FAILED = {0: {2, 3}, 2: {1, 2, 3, 5}} - -_OUT_DIR = ("logs/agent_sim_learning/" - "domino-agent_oracle_hybrid_sim_oracle_samplers/initial_states") - - -def main() -> None: - """Render and save initial states of the domino test tasks.""" - os.makedirs(_OUT_DIR, exist_ok=True) - for seed in (0, 2): - utils.reset_config({**_DOMINO_FLAGS, "seed": seed}) - # do_cache=False: a cached env keeps its seed-0 test tasks, so each - # seed must build a fresh env to regenerate its own test tasks. - env = create_new_env("pybullet_domino", do_cache=False) - tasks = env.get_test_tasks() - for idx in range(len(tasks)): - env.reset("test", idx) - rgb = np.asarray(env.render()[0], dtype=np.uint8) - task_num = idx + 1 # 1-indexed to match the run logs - status = "FAILED" if task_num in _FAILED.get(seed, set()) \ - else "solved" - fname = f"seed{seed}_task{task_num}_{status}.png" - path = os.path.join(_OUT_DIR, fname) - Image.fromarray(rgb).save( # type: ignore[no-untyped-call] - path) - goal = sorted(str(a) for a in tasks[idx].goal) - print(f"seed{seed} task{task_num} [{status}] -> {path}") - print(f" goal: {goal}") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/render_unsolved_domino_states.py b/scripts/domino_debug/render_unsolved_domino_states.py deleted file mode 100644 index 5194d65a2b..0000000000 --- a/scripts/domino_debug/render_unsolved_domino_states.py +++ /dev/null @@ -1,173 +0,0 @@ -"""Render init-state PNGs for the unsolved domino tasks (oracle-samplers runs). - -Uses the geometry-affecting flags from the experiment command line so the -regenerated test scenes match the runs exactly (verified: seed1 = [4,4,5,4,4] -dominoes, and the seed1.t2 grasp-infeasibility matches the run). Run ONE seed -per process (task-gen RNG is shared across seeds in one interpreter). - -Usage: PYTHONPATH=. python \ - scripts/domino_debug/render_unsolved_domino_states.py -""" -import os -import sys -from typing import Any, Dict, List, Optional, Sequence, Tuple - -import numpy as np -from numpy.typing import NDArray -from PIL import Image, ImageDraw, ImageFont - -from predicators import utils -from predicators.envs import create_new_env -from predicators.structs import State - - -def _project(xyz: Sequence[float], view_matrix: Sequence[float], - proj_matrix: Sequence[float], width: int, - height: int) -> Optional[Tuple[float, float]]: - """World (x,y,z) -> (u,v) pixel using pybullet's column-major matrices.""" - V = np.array(view_matrix).reshape((4, 4), order="F") - P = np.array(proj_matrix).reshape((4, 4), order="F") - clip = P @ (V @ np.array([xyz[0], xyz[1], xyz[2], 1.0])) - if clip[3] == 0: - return None - ndc = clip[:3] / clip[3] - return ((ndc[0] * 0.5 + 0.5) * width, - (1.0 - (ndc[1] * 0.5 + 0.5)) * height) - - -def _font(size: int) -> Any: - """Load a TrueType font at the given size, falling back to a default.""" - for path in ("/System/Library/Fonts/Supplemental/Arial Bold.ttf", - "/System/Library/Fonts/Helvetica.ttc"): - try: - return ImageFont.truetype( # type: ignore[no-untyped-call] - path, size) - except Exception: # pylint: disable=broad-except - pass - try: - return ImageFont.load_default(size=size) - except TypeError: - return ImageFont.load_default() - - -def _caption(rgb: NDArray[np.uint8], lines: List[str]) -> NDArray[np.uint8]: - """Draw a header banner (top-left) with the given text lines.""" - img = Image.fromarray(rgb) # type: ignore[no-untyped-call] - draw = ImageDraw.Draw(img, "RGBA") - font = _font(22) - pad, lh = 8, 26 - w = max( - draw.textlength( # type: ignore[no-untyped-call] - t, font=font) for t in lines) - draw.rectangle([0, 0, w + 2 * pad, lh * len(lines) + pad], - fill=(0, 0, 0, 170)) - for i, t in enumerate(lines): - draw.text((pad, pad + i * lh), t, fill=(255, 255, 255), font=font) - return np.asarray(img) - - -def _annotate(rgb: NDArray[np.uint8], init_state: State, - cam: Any) -> NDArray[np.uint8]: - """Label each domino with its index at its initial-state position.""" - img = Image.fromarray(rgb) # type: ignore[no-untyped-call] - draw = ImageDraw.Draw(img) - font = _font(26) - for o in sorted([o for o in init_state if o.type.name == "domino"], - key=lambda o: o.name): - x, y, z = (init_state.get(o, "x"), init_state.get(o, "y"), - init_state.get(o, "z")) - uv = _project((x, y, z + 0.13), *cam) - if uv is None: - continue - u, v = uv - idx = o.name.split("_")[-1] - col = (int(init_state.get(o, "r") * 255), - int(init_state.get(o, "g") * 255), - int(init_state.get(o, "b") * 255)) - r = 15 - draw.ellipse([u - r, v - r, u + r, v + r], - fill=(0, 0, 0), - outline=col, - width=3) - tb = draw.textbbox((0, 0), idx, font=font) - draw.text((u - (tb[2] - tb[0]) / 2, v - (tb[3] - tb[1]) / 2 - tb[1]), - idx, - fill=(255, 255, 255), - font=font) - return np.asarray(img) - - -# 1-indexed tasks unsolved in EITHER arm, with (arms, failure-mode) labels. -UNSOLVED: Dict[int, Dict[int, Tuple[str, str]]] = { - 0: { - 1: ("both", "push-dropped"), - 2: ("both", "place-MP+InFront"), - 3: ("no_demo", "pick+place-MP") - }, - 1: { - 1: ("demo", "exec-retreat-collision"), - 3: ("both", "pick+place-MP") - }, - 2: { - 1: ("no_demo", "pick+place-MP"), - 2: ("demo", "toppled-cascade"), - 4: ("both", "pick+place-MP"), - 5: ("both", "place-MP+toppled") - }, - 3: { - 5: ("demo", "holding+InFront+place-MP") - }, -} -FLAGS: Dict[str, Any] = { - "env": "pybullet_domino", - "num_train_tasks": 1, - "num_test_tasks": 5, - "pybullet_ik_validate": False, - "pybullet_camera_width": 900, - "pybullet_camera_height": 900, - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_use_continuous_place": True, - "domino_restricted_push": True, - "domino_has_glued_dominos": False, - "pybullet_birrt_extend_num_interp": 20, - "pybullet_birrt_path_subsample_ratio": 2, -} -OUT = "logs/agent_sim_learning/unsolved_init_states" - - -def main() -> None: - """Render annotated init-state PNGs for one seed's unsolved tasks.""" - seed = int(sys.argv[1]) - os.makedirs(OUT, exist_ok=True) - utils.reset_config(dict(FLAGS, seed=seed)) - env = create_new_env("pybullet_domino", do_cache=False) - tasks = env.get_test_tasks() - counts = [ - len([o for o in t.init if o.type.name == "domino"]) for t in tasks - ] - print(f"seed{seed} domino counts per task = {counts}") - # pylint: disable=protected-access - cam = env._get_camera_matrices() # type: ignore[attr-defined] - for t1, (arms, mode) in sorted(UNSOLVED.get(seed, {}).items()): - idx = t1 - 1 - env.reset("test", idx) - rgb = np.asarray(env.render()[0], dtype=np.uint8) - rgb = _annotate(rgb, tasks[idx].init, cam) - goal_ids = ",".join( - sorted( - str(a).rsplit("_", maxsplit=1)[-1].rstrip(":domino)") - for a in tasks[idx].goal)) - rgb = _caption(rgb, [ - f"seed {seed} task {t1} ({arms})", - f"goal: Toppled({goal_ids}) fail: {mode}" - ]) - fname = f"seed{seed}_task{t1}_{arms}_{mode}.png" - Image.fromarray( # type: ignore[no-untyped-call] - rgb).save(os.path.join(OUT, fname)) - goal = sorted(str(a) for a in tasks[idx].goal) - print(f" saved {fname} | {counts[idx]} dominoes | goal={goal}") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/replay_domino_sketches.py b/scripts/domino_debug/replay_domino_sketches.py deleted file mode 100644 index f1c83e81da..0000000000 --- a/scripts/domino_debug/replay_domino_sketches.py +++ /dev/null @@ -1,226 +0,0 @@ -"""Faithfully reproduce domino refinement failures by replaying the recorded -LLM sketches through the real bilevel refinement -- no LLM required. - -The agent's plan sketches were logged verbatim in each run's ``info.log`` -(``Sketch (attempt N):`` blocks). This script extracts them, regenerates the -deterministic test task, and runs the *exact same* ``refine_sketch`` the -pipeline uses (oracle option model + oracle samplers + subgoal checks, same -per-(sketch,refine) RNG seeding). The pass/fail outcome and the "stuck at step -K" reason therefore reproduce the run's solve-time failures deterministically. - -Run ONE seed per process (task-gen RNG is shared; see -reproduce_domino_failures). - -Usage: - PYTHONPATH=. python scripts/domino_debug/replay_domino_sketches.py \ - [--all] - --all replays every task; default replays only tasks the run - did not solve. -""" - -import logging -import re -import sys -from glob import glob -from typing import Dict, List, Optional, Tuple - -from predicators import utils -from predicators.agent_sdk import bilevel_sketch -from predicators.approaches import create_approach -from predicators.approaches.agent_sim_learning_approach import \ - AgentSimLearningApproach -from predicators.envs import get_or_create_env -from predicators.ground_truth_models import get_gt_options - -logging.disable(logging.CRITICAL) - -# Refine retries per sketch, matching the refine loop the audited runs -# ran with (their agent_bilevel_max_refine_retries setting); preserved -# here so old recordings keep replaying faithfully. -_REFINE_RETRIES = 5 - -ANSI = re.compile(r"\x1b\[[0-9;]*m") -STEP = re.compile( - r"^\s*\d+:\s*([A-Za-z]\w*)\((.*?)\)(?:\s*->\s*\{(.*)\})?\s*$") -SKETCH_HDR = re.compile(r"Sketch \(attempt (\d+)\)") -TASK_RES = re.compile( - r"\[main\.py\] Task (\d+) / \d+: (.*)|Task (\d+) / \d+: (SOLVED)") - -_FLAGS = { - "env": "pybullet_domino", - "approach": "agent_sim_learning", - "num_train_tasks": 1, - "num_test_tasks": 5, - "skill_phase_use_motion_planning": True, - "pybullet_ik_validate": False, - "demonstrator": "oracle_process_planning", - "bilevel_plan_without_sim": True, - "explorer": "agent_model_based", - "agent_sim_learn_oracle_sim_program": True, - "agent_sim_learn_oracle_sim_params": True, - "agent_sim_learn_parameterized_samplers": True, - "agent_sim_learn_oracle_samplers": True, - "execution_monitor": "subgoal_annotations", - "agent_bilevel_max_execution_replans": 2, - "horizon": 400, - "excluded_objects_in_state_str": "loc,rot,angle,direction", - "excluded_predicates": "InitialBlock,MovableBlock,Tilting,Upright", - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_use_continuous_place": True, - "domino_restricted_push": True, - "process_planning_heuristic_weight": 2.0, - "domino_has_glued_dominos": False, - "pybullet_birrt_extend_num_interp": 20, - "pybullet_birrt_path_subsample_ratio": 2, - "agent_sdk_use_local_sandbox": True, - "option_model_terminate_on_repeat": False, - "agent_planner_use_simulator": True, -} - - -def find_info_log(seed: int, arm: str) -> str: - """Return the newest matching run's info.log path for seed/arm.""" - exp = f"domino-agent_oracle_hybrid_sim_oracle_samplers_{arm}" - pat = f"logs/agent_sim_learning/{exp}/seed{seed}/run_*/info.log" - hits = sorted(glob(pat)) - if not hits: - raise SystemExit(f"no info.log at {pat}") - return hits[-1] - - -def extract_sketches(info_log: str) -> Dict[int, dict]: - """Return {task_idx (0-based): {"outcome": str, "sketches": ...}}. - - Each step is (option_name, [obj_names], raw_subgoal_str). - """ - tasks: Dict[int, dict] = {} - pending: List[List[Tuple[str, List[str], str]]] = [] - cur: Optional[List[Tuple[str, List[str], str]]] = None - with open(info_log, encoding="utf-8") as f: - for raw in f: - line = ANSI.sub("", raw.rstrip("\n")) - if SKETCH_HDR.search(line): - cur = [] - pending.append(cur) - continue - m = STEP.match(line) - if m and cur is not None: - opt, args, sg = m.group(1), m.group(2), m.group(3) or "" - objs = [ - a.split(":")[0].strip() for a in args.split(",") - if a.strip() - ] - cur.append((opt, objs, sg)) - continue - cur = None # any non-step line ends the current sketch block - tm = TASK_RES.search(line) - if tm: - ti = int(tm.group(1) or tm.group(3)) - 1 - outcome = (tm.group(2) or tm.group(4) or "").strip() - tasks[ti] = {"outcome": outcome, "sketches": pending} - pending = [] - return tasks - - -def typed_text(steps: List[Tuple[str, List[str], str]], - name_to_type: Dict[str, str]) -> str: - """Rebuild typed sketch text the option-plan parser expects.""" - lines = [] - for opt, objs, sg in steps: - typed = ", ".join(f"{o}:{name_to_type.get(o, 'object')}" for o in objs) - line = f"{opt}({typed})" - if sg: - line += f" -> {{{sg}}}" - lines.append(line) - return "\n".join(lines) - - -class _DeepestFail: - """Track the deepest (highest-index) step failure seen so far.""" - - def __init__(self) -> None: - self.idx: int = -1 - self.reason: str = "" - - def record(self, idx: int, _prefix: list, reason: str) -> None: - """on_step_fail callback: keep the deepest failure.""" - if idx > self.idx: - self.idx, self.reason = idx, reason - - -def main() -> None: - """Replay recorded sketches through the real refinement.""" - seed = int(sys.argv[1]) - arm = sys.argv[2] if len(sys.argv) > 2 else "no_demo" - replay_all = "--all" in sys.argv - - info_log = find_info_log(seed, arm) - tasks = extract_sketches(info_log) - - utils.reset_config(dict(_FLAGS, seed=seed)) - - env = get_or_create_env("pybullet_domino") - options = get_gt_options(env.get_name()) - preds, _ = utils.parse_config_excluded_predicates(env) - train_tasks = [t.task for t in env.get_train_tasks()] - approach = create_approach("agent_sim_learning", preds, options, env.types, - env.action_space, train_tasks) - assert isinstance(approach, AgentSimLearningApproach) - # pylint: disable=protected-access - approach._maybe_install_oracle_samplers() - # pylint: enable=protected-access - test_tasks = env.get_test_tasks() - name_to_type = {o.name: o.type.name for o in test_tasks[0].task.init} - - print(f"# seed{seed} {arm}: replaying recorded sketches through real " - f"refinement (oracle option-model + oracle samplers, no LLM)") - for ti in sorted(tasks): - rec = tasks[ti] - solved = rec["outcome"].upper().startswith("SOLVED") - if solved and not replay_all: - continue - task = test_tasks[ti].task - print(f"\n== task{ti} (run Task{ti+1}) | run outcome: " - f"{rec['outcome'][:60]}") - if not rec["sketches"]: - print(" (no sketches recorded)") - continue - for si, steps in enumerate(rec["sketches"]): - sketch = bilevel_sketch.parse_sketch_from_text( - typed_text(steps, name_to_type), - task, - predicates=preds, - options=set(options), - types=env.types) - if not sketch: - print(f" sketch{si}: unparseable") - continue - any_success = False - deepest_idx, deepest_reason = -1, "" - for r in range(_REFINE_RETRIES): - fail = _DeepestFail() - # attempt reproduces the audited runs' per-(sketch,refine) - # RNG seeding (rng = CFG.seed + attempt inside - # _refine_sketch), so recorded failures replay exactly. - # pylint: disable-next=protected-access - _, success = approach._refine_sketch( - task, - sketch, - timeout=600.0, - attempt=si * _REFINE_RETRIES + r, - on_step_fail=fail.record) - if success: - any_success = True - break - if fail.idx > deepest_idx: - deepest_idx, deepest_reason = fail.idx, fail.reason - verdict = "REFINED-OK" if any_success else \ - f"FAILED (stuck step {deepest_idx}: {deepest_reason[:60]})" - head = " -> ".join(f"{o}({','.join(a)})" for o, a, _ in steps) - print(f" sketch{si} [{len(steps)} steps]: {verdict}") - print(f" {head}") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/replay_ikval_sweep.py b/scripts/domino_debug/replay_ikval_sweep.py deleted file mode 100644 index babae9c7c1..0000000000 --- a/scripts/domino_debug/replay_ikval_sweep.py +++ /dev/null @@ -1,196 +0,0 @@ -"""Replay each task's recorded sketches with ik_validate=True and compare to -the recorded (ik_validate=False) run outcome — to separate pure IK-artifact -failures from genuine geometric ones, and to check for regressions on solved -tasks. - -Mirrors the pipeline solve loop: for each task, try recorded sketches in -order, up to N refine attempts each, stop at first success; enforce a -per-task wall budget. Run ONE seed+arm per process. Usage: PYTHONPATH=. -python scripts/domino_debug/replay_ikval_sweep.py -[budget_s] -""" -import logging -import re -import sys -import time -from glob import glob -from typing import Any, Callable, Dict, List, Tuple - -logging.disable(logging.CRITICAL) - -ANSI = re.compile(r"\x1b\[[0-9;]*m") -STEP = re.compile( - r"^\s*\d+:\s*([A-Za-z]\w*)\((.*?)\)(?:\s*->\s*\{(.*)\})?\s*$") -SKH = re.compile(r"Sketch \(attempt (\d+)\)") -TRES = re.compile( - r"\[main\.py\] Task (\d+) / \d+: (.*)|Task (\d+) / \d+: (SOLVED)") - - -def extract(info_log: str) -> Dict[int, dict]: - """Parse recorded sketches and outcomes per task from an info.log.""" - tasks: Dict[int, dict] = {} - pending: List[List[Tuple[str, List[str], str]]] = [] - cur: List[Tuple[str, List[str], str]] | None = None - for raw in open(info_log, encoding="utf-8"): - line = ANSI.sub("", raw.rstrip("\n")) - if SKH.search(line): - cur = [] - pending.append(cur) - continue - m = STEP.match(line) - if m and cur is not None: - cur.append((m.group(1), [ - a.split(":")[0].strip() for a in m.group(2).split(",") - if a.strip() - ], m.group(3) or "")) - continue - cur = None - tm = TRES.search(line) - if tm: - ti = int(tm.group(1) or tm.group(3)) - 1 - tasks[ti] = { - "outcome": (tm.group(2) or tm.group(4) or "").strip(), - "sketches": pending - } - pending = [] - return tasks - - -def main() -> None: - """Replay recorded sketches with ik_validate and report flips.""" - seed = int(sys.argv[1]) - arm = sys.argv[2] - budget = float(sys.argv[3]) if len(sys.argv) > 3 else 300.0 - ikv = (sys.argv[4].lower() == "true") if len(sys.argv) > 4 else True - exp = f"domino-agent_oracle_hybrid_sim_oracle_samplers_{arm}" - info_log = sorted( - glob(f"logs/agent_sim_learning/{exp}/seed{seed}/run_*/info.log"))[-1] - rec = extract(info_log) - FLAGS = { - "env": "pybullet_domino", - "approach": "agent_sim_learning", - "seed": seed, - "num_train_tasks": 1, - "num_test_tasks": 5, - "skill_phase_use_motion_planning": True, - "pybullet_ik_validate": ikv, # <-- the change under test - "demonstrator": "oracle_process_planning", - "bilevel_plan_without_sim": True, - "explorer": "agent_model_based", - "agent_sim_learn_oracle_sim_program": True, - "agent_sim_learn_oracle_sim_params": True, - "agent_sim_learn_parameterized_samplers": True, - "agent_sim_learn_oracle_samplers": True, - "execution_monitor": "subgoal_annotations", - "agent_bilevel_max_execution_replans": 2, - "horizon": 400, - "excluded_objects_in_state_str": "loc,rot,angle,direction", - "excluded_predicates": "InitialBlock,MovableBlock,Tilting,Upright", - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_use_continuous_place": True, - "domino_restricted_push": True, - "domino_has_glued_dominos": False, - "pybullet_birrt_extend_num_interp": 20, - "pybullet_birrt_path_subsample_ratio": 2, - "agent_sdk_use_local_sandbox": True, - "option_model_terminate_on_repeat": False, - "agent_planner_use_simulator": True - } - # pylint: disable=import-outside-toplevel - # Imports are deferred until after reset_config so module-level CFG - # reads in these modules observe the FLAGS set above. - from predicators import utils - utils.reset_config(FLAGS) - from predicators.agent_sdk import bilevel_sketch - from predicators.approaches import create_approach - from predicators.envs import get_or_create_env - from predicators.ground_truth_models import get_gt_options - env = get_or_create_env("pybullet_domino") - options = get_gt_options(env.get_name()) - preds, _ = utils.parse_config_excluded_predicates(env) - # Cast to Any: this script probes approach-specific protected members - # (sampler installers, option model) absent from the BaseApproach API. - ap: Any = create_approach("agent_sim_learning", preds, options, env.types, - env.action_space, - [t.task for t in env.get_train_tasks()]) - ap._maybe_install_oracle_samplers() # pylint: disable=protected-access - n2t = {o.name: o.type.name for o in env.get_test_tasks()[0].task.init} - - def typed(steps: List[Tuple[str, List[str], str]]) -> str: - """Render parsed sketch steps as typed operator lines.""" - lines = [] - for op, objs, sg in steps: - args = ", ".join(o + ":" + n2t.get(o, "object") for o in objs) - line = op + "(" + args + ")" - if sg: - line += " -> {" + sg + "}" - lines.append(line) - return "\n".join(lines) - - header = (f"# seed{seed} {arm} | ik_validate={ikv} | " - f"NEW task-gen | budget={budget}s") - print(header) - for ti in sorted(rec): - task = env.get_test_tasks()[ti].task - recout = "SOLVED" if rec[ti]["outcome"].upper().startswith( - "SOLVED") else "FAILED" - t0 = time.perf_counter() - solved_by = None - deepest: Tuple[int, str] = (-1, "") - for si, steps in enumerate(rec[ti]["sketches"]): - if time.perf_counter() - t0 > budget: - break - sk = bilevel_sketch.parse_sketch_from_text(typed(steps), - task, - predicates=preds, - options=set(options), - types=env.types) - if not sk: - continue - for r in range(2): - if time.perf_counter() - t0 > budget: - break - fail: Dict[str, object] = {"idx": -1, "reason": ""} - - def make_rc( - f: Dict[str, - object]) -> Callable[[int, object, str], None]: - """Build an on_step_fail recording the deepest fail.""" - - def rc(i: int, _p: object, reason: str) -> None: - if i > f["idx"]: # type: ignore[operator] - f["idx"], f["reason"] = i, reason - - return rc - - # attempt reproduces this script's historical RNG streams - # (rng = CFG.seed + attempt inside _refine_sketch). - # pylint: disable-next=protected-access - _, ok = ap._refine_sketch(task, - sk, - timeout=budget, - attempt=si * 5 + r, - on_step_fail=make_rc(fail)) - if ok: - solved_by = (si, r) - break - if fail["idx"] > deepest[0]: # type: ignore[operator] - deepest = (fail["idx"], fail["reason"]) # type: ignore - if solved_by: - break - dt = time.perf_counter() - t0 - verdict = "SOLVED" if solved_by else "FAILED" - flip = "" if verdict == recout else ( - " *** REGRESSION" if recout == "SOLVED" else " *** FIXED") - if solved_by: - extra = f"by sketch{solved_by[0]}" - else: - extra = f"deepest step{deepest[0]}: {deepest[1][:32]}" - line = (f" task{ti+1}: recorded(F)={recout:6s} -> " - f"new={verdict:6s} [{dt:5.0f}s] {extra}{flip}") - print(line) - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/replay_plan.py b/scripts/domino_debug/replay_plan.py deleted file mode 100644 index 70079128bb..0000000000 --- a/scripts/domino_debug/replay_plan.py +++ /dev/null @@ -1,218 +0,0 @@ -"""Replay an EXACT solved option plan on the real-scene domino env -(deterministic, no LLM). Grounds the plan's Pick/Place/Push/Wait options with -their exact parameters and rolls them through the env in TEST mode. With -``real_robot_execute=True`` a RealRobotExecutor is attached to the env and each -option's joint trajectory is shipped to the Franka as that option ends. The -scene is NOT re-perceived between options here: this tool replays a fixed plan, -and correcting the twin mid-replay would let the option policies see states the -recorded plan was never chosen against. The default dry-run stays pure sim -(optionally rendered to MP4). - -Plan-file format (one option per line; ``-> {...}`` subgoals optional/ignored): - Pick(robot:robot, domino_1:domino)[0.06] -> {Holding(robot, domino_1)} - Place(robot:robot)[0.70, 1.16, 0.55, 1.75] - Push(robot:robot, domino_0:domino)[0.03, 0.05] - Wait(robot:robot)[] - -Three rungs, in order. Take them all; each adds exactly one new source of -failure (from the predicators repo root, robot-ml; PYTHONHASHSEED=0): - - # 1. pure sim: no executor, no robot object at all - PYTHONPATH=. python scripts/domino_debug/replay_plan.py --plan plan.txt - - # 2. dry arm: the whole RealRobot minus the arm. Attachment, per-option - # chunking and the gripper split all run; nothing moves. Needs - # babyrobot importable, needs no hardware powered on. - PYTHONPATH=.:/path/to/BabyRobotPredicator \ - python scripts/domino_debug/replay_plan.py --plan plan.txt \ - --execute --dry - - # 3. MOVES THE ARM - PYTHONPATH=.:/path/to/BabyRobotPredicator \ - python scripts/domino_debug/replay_plan.py --plan plan.txt --execute -""" -import argparse -import logging -import os -import re -from typing import List, Tuple - -import numpy as np - -from predicators import utils -from predicators.approaches import create_approach -from predicators.cogman import CogMan, run_episode_and_get_observations -from predicators.envs import get_or_create_env -from predicators.envs.pybullet_domino_real import PyBulletDominoRealEnv -from predicators.execution_monitoring import create_execution_monitor -from predicators.ground_truth_models import get_gt_options -from predicators.perception import create_perceiver -from predicators.pybullet_helpers.real_robot_executor import attach_real_robot -from predicators.settings import CFG -from scripts.cluster_utils import SingleSeedRunConfig, generate_run_configs - -# pylint: disable=protected-access -_LINE = re.compile(r"^\s*(\w+)\s*\(([^)]*)\)\s*\[([^\]]*)\]") - - -def _parse_plan(text: str) -> List[Tuple[str, List[str], List[float]]]: - """[(option_name, [obj_names], [param_floats]), ...] from the plan text.""" - steps: List[Tuple[str, List[str], List[float]]] = [] - for raw in text.splitlines(): - line = raw.split("->", 1)[0].strip() - if not line or line.startswith("#"): - continue - m = _LINE.match(line) - if not m: - continue - name, args, params = m.group(1), m.group(2), m.group(3) - objs = [ - a.split(":", 1)[0].strip() for a in args.split(",") if a.strip() - ] - floats = [float(v) for v in params.split(",") if v.strip()] - steps.append((name, objs, floats)) - return steps - - -def main() -> None: - """Ground the plan's options and roll them through the real-scene env.""" - ap = argparse.ArgumentParser( - description=__doc__, - formatter_class=argparse.RawDescriptionHelpFormatter) - ap.add_argument("--plan", required=True, help="plan text file") - ap.add_argument("--config", default="predicatorv3/exp_domino_real.yaml") - ap.add_argument( - "--scene", - default=None, - help="override CFG.domino_real_scene (must match the plan)") - ap.add_argument("--execute", - action="store_true", - help="EXECUTE ON THE REAL FRANKA (needs the babyrobot " - "submodule installed). Default: dry-run (pure sim, no " - "motion).") - ap.add_argument("--dry", - action="store_true", - help="with --execute: build the whole RealRobot but with " - "NO arm attached. The executor still attaches, every " - "option still chunks and ships, the gripper split still " - "runs -- and nothing moves. This is the rung between pure " - "sim and real motion; take it before every new plan.") - ap.add_argument("--observe", - action="store_true", - help="look at the scene between options and correct the " - "twin from what was seen (bring-up Stage 5). Opens the " - "cameras. The replay stops being a pure replay -- that is " - "the point of the rung, not a side effect.") - ap.add_argument("--out", default=None, help="optional MP4 of the rollout") - args = ap.parse_args() - logging.basicConfig(level=logging.INFO, format="%(message)s") - - rc = list(generate_run_configs(args.config))[0] - assert isinstance(rc, SingleSeedRunConfig) - flags = dict(rc.flags) - flags.update({"env": rc.env, "approach": rc.approach, "seed": rc.seed}) - flags.pop("log", None) - if args.scene: - flags["domino_real_scene"] = args.scene - flags["real_robot_execute"] = bool(args.execute) - # Build the arm-less RealRobot: everything downstream of the executor runs - # for real, so this exercises chunking, the gripper split and the drift - # guard without a Franka in the room (and without one powered on). - flags["real_robot_dry"] = bool(args.dry) - # By default this tool replays an EXACT plan and never looks: re-syncing - # the twin mid-replay lets the option policies see states the recorded plan - # was never chosen against, and not looking keeps the tool usable with the - # cameras down. --observe opts into exactly that mid-replay correction, - # which is the closed-loop rung. Perception follows, because RealRobot - # opens its session at CONSTRUCTION -- left at the "zed" default a run that - # never looks would still hold both cameras open. - flags["real_robot_perception"] = "zed" if args.observe else "none" - flags["real_robot_observe_at_option_boundary"] = bool(args.observe) - # ...and no human reset either: that rebuilds the episode's task from a - # live look, which would rename and re-place the very objects the recorded - # plan refers to. The task must stay the captured scene the plan was - # written against. - flags["real_robot_human_reset"] = False - # ...which is exactly the case the stale-task guard exists for, so opt in - # explicitly: these poses are the ones the plan was written against. - flags["real_robot_allow_captured_scene_task"] = True - utils.reset_config(flags) - - env = get_or_create_env(CFG.env) - assert isinstance(env, PyBulletDominoRealEnv), \ - f"replay_plan drives the real-scene env; got {CFG.env}" - # Attaches the arm under --execute and is a no-op otherwise, so the env - # stays the same object either way. - attach_real_robot(env) - opts = {o.name: o for o in get_gt_options(env.get_name())} - env_task = env.get_test_tasks()[0] - task = env_task.task - by_name = {o.name: o for o in task.init} - - with open(args.plan, encoding="utf-8") as f: - steps = _parse_plan(f.read()) - assert steps, "no plan steps parsed" - - plan = [] - for name, obj_names, params in steps: - option = opts[name] - objs = [by_name[n] for n in obj_names] - plan.append(option.ground(objs, np.array(params, dtype=np.float32))) - print("# grounded plan:") - for g in plan: - print(" ", g.simple_str()) - # Say plainly whether metal is about to move: this banner is the last - # thing a human reads before deciding where their hand is. - if not CFG.real_robot_execute: - print("# SIM ONLY -- no executor attached, no arm, nothing moves") - elif CFG.real_robot_dry: - print("# DRY ARM -- RealRobot built without an arm; chunks ship, " - "nothing moves") - else: - print("# *** THE REAL FRANKA WILL MOVE *** (in-process RealRobot)") - - policy = utils.option_plan_to_policy( - plan, abstract_function=lambda s: utils.abstract(s, env.predicates)) - monitor = utils.VideoMonitor(env.render) if args.out else None - - # Roll out through CogMan's episode loop -- the same one main.py uses -- - # driven by an override policy, exactly as the online-learning path does - # (main.py sets the override, then resets). With an override in place - # CogMan never asks the approach to solve, so the plan being replayed is - # the plan that executes. - # - # The approach is therefore never consulted for control, and only has to - # *construct*. "random_options" is a plain BaseApproach that needs nothing - # but the option set. Not "oracle": that one builds ground-truth NSRTs in - # its constructor, and this env has none -- it is planned over processes, - # so get_gt_nsrts raises NotImplementedError for pybullet_domino_real and - # the replay dies before it renders a frame. - cogman = CogMan( - create_approach("random_options", env.predicates, - get_gt_options(env.get_name()), env.types, - env.action_space, [task]), - create_perceiver(CFG.perceiver), create_execution_monitor("trivial")) - cogman.set_override_policy(policy) - cogman.set_termination_function(lambda s: False) - cogman.reset(env_task) - (_, actions), _, _ = run_episode_and_get_observations( - cogman, - env, - "test", - 0, - max_num_steps=CFG.horizon, - terminate_on_goal_reached=False, - exceptions_to_break_on={utils.OptionExecutionFailure}, - monitor=monitor) - print(f"# steps={len(actions)} goal_reached={env.goal_reached()}") - - if args.out and monitor is not None: - os.makedirs(os.path.join(CFG.video_dir, - os.path.dirname(args.out) or "."), - exist_ok=True) - utils.save_video(args.out, monitor.get_video()) - print(f"# saved {os.path.join(CFG.video_dir, args.out)}") - - -if __name__ == "__main__": - main() diff --git a/scripts/domino_debug/reproduce_domino_failures.py b/scripts/domino_debug/reproduce_domino_failures.py deleted file mode 100644 index 563f3bfb40..0000000000 --- a/scripts/domino_debug/reproduce_domino_failures.py +++ /dev/null @@ -1,146 +0,0 @@ -"""Deterministic, LLM-free reproduction of the domino oracle-samplers failures. - -Reproduces the geometric / parsing root causes behind the unsolved tasks in -``domino-agent_oracle_hybrid_sim_oracle_samplers_{demo,no_demo}`` (seeds 0-4), -*without* invoking the LLM sketcher. The test-task scenes are deterministic -given the seed, so the BiRRT motion-planning infeasibilities and the option-plan -parser bug reproduce exactly. - -IMPORTANT: run ONE seed per process. Generating tasks for several seeds inside -one interpreter advances the shared RNG and changes the scenes (e.g. seed1 would -regenerate as [4,5,5,5,4] dominoes instead of the real [4,4,5,4,4]). The bash -wrapper at the bottom of the module docstring loops correctly. - -Usage: - # motion-planning reproduction for a single seed (fresh process each): - # for s in 0 1 2 3 4; do PYTHONPATH=. python \ - # scripts/domino_debug/reproduce_domino_failures.py mp $s; done - # option-plan parser (Push) bug: - # PYTHONPATH=. python \ - # scripts/domino_debug/reproduce_domino_failures.py push 0 -""" - -import logging -import sys -from typing import List, Optional, Tuple - -import numpy as np - -from predicators import utils -from predicators.envs import get_or_create_env -from predicators.envs.base_env import BaseEnv -from predicators.ground_truth_models import get_gt_options -from predicators.structs import ParameterizedOption, State, _Option - -logging.disable(logging.CRITICAL) - -# Geometry-affecting flags copied verbatim from the experiment command line. -_ARGS = { - "env": "pybullet_domino", - "approach": "oracle", - "num_train_tasks": 1, - "num_test_tasks": 5, - "pybullet_ik_validate": False, - "skill_phase_use_motion_planning": True, - "domino_initialize_at_finished_state": False, - "domino_use_domino_blocks_as_target": True, - "domino_use_continuous_place": True, - "domino_restricted_push": True, - "domino_has_glued_dominos": False, - "pybullet_birrt_extend_num_interp": 20, - "pybullet_birrt_path_subsample_ratio": 2, -} -_GRASP_Z_OFFSET = 0.0825 # value used by the oracle Pick sampler in the runs. -_POS_GAP = 0.098 # domino chain spacing (env.py: domino_width * 1.4). -_MAX_STEPS = 80 - - -def _setup(seed: int) -> Tuple[BaseEnv, List[ParameterizedOption]]: - """Reset the config and create the env + ground-truth options.""" - args = dict(_ARGS, seed=seed) - utils.reset_config(args) - env = get_or_create_env("pybullet_domino") - options = list(get_gt_options(env.get_name())) - return env, options - - -def _run_option(env: BaseEnv, opt: _Option, - state: State) -> Tuple[Optional[bool], str]: - """Drive a grounded option to termination; return (ok, failure_msg).""" - if not opt.initiable(state): - return None, "not-initiable" - s = state - for _ in range(_MAX_STEPS): - try: - a = opt.policy(s) - except utils.OptionExecutionFailure as e: - return False, str(e) - s = env.step(a) - if opt.terminal(s): - return True, "ok" - return True, "ran-max-steps" - - -def reproduce_mp(seed: int) -> None: - """Report grasp-infeasible dominoes and probe one Place into the gap.""" - env, options = _setup(seed) - Pick = next(o for o in options if o.name == "Pick") - tasks = env.get_test_tasks() - for ti in range(len(tasks)): - env.reset("test", ti) - st = env._current_state # pylint: disable=protected-access - dominoes = sorted([o for o in st if o.type.name == "domino"], - key=lambda o: o.name) - infeasible = [] - for d in dominoes: - env.reset("test", ti) - s = env._current_state # pylint: disable=protected-access - rb = next(o for o in s if o.type.name == "robot") - dd = next(o for o in s if o.name == d.name) - opt = Pick.ground([rb, dd], - np.array([_GRASP_Z_OFFSET], dtype=np.float32)) - ok, _ = _run_option(env, opt, s) - if ok is False: - infeasible.append(d.name) - print( - f"seed{seed} task{ti} (run Task{ti+1}): {len(dominoes)} dominoes " - f"| grasp-INFEASIBLE: {infeasible if infeasible else 'none'}") - - -def reproduce_push_bug(seed: int) -> None: - """Show the parser drops a Push line that names a target domino.""" - env, options = _setup(seed) - push = next(o for o in options if o.name == "Push") - print(f"Push option signature: types={[t.name for t in push.types]}") - state = env.get_test_tasks()[0].init - objects = list(state) - cases = { - "LLM-style 'Push(robot, domino_0)'": - "Pick(robot:robot, domino_1:domino)\n" - "Push(robot:robot, domino_0:domino)\nWait(robot:robot)", - "legal 'Push(robot)'": - "Pick(robot:robot, domino_1:domino)\n" - "Push(robot:robot)\nWait(robot:robot)", - } - for label, txt in cases.items(): - plan = utils.parse_model_output_into_option_plan( - txt, objects, env.types, options, parse_continuous_params=False) - names = [op.name for op, _, _ in plan] - flag = "PUSH DROPPED!" if "Push" not in names else "ok" - print(f" {label:42s} -> {names} ({flag})") - - -def _main() -> None: - """Dispatch to the requested reproduction mode.""" - mode = sys.argv[1] if len(sys.argv) > 1 else "mp" - seed = int(sys.argv[2]) if len(sys.argv) > 2 else 0 - if mode == "mp": - reproduce_mp(seed) - elif mode == "push": - reproduce_push_bug(seed) - else: - raise SystemExit(f"unknown mode {mode!r} (expected 'mp' or 'push')") - - -if __name__ == "__main__": - _main() diff --git a/scripts/dump_continual_arm_prompts.py b/scripts/dump_continual_arm_prompts.py index 4136d7ce5c..f20ca07414 100644 --- a/scripts/dump_continual_arm_prompts.py +++ b/scripts/dump_continual_arm_prompts.py @@ -9,7 +9,7 @@ is sent to a model. python -m scripts.dump_continual_arm_prompts \ - --config predicatorv3/continual_eight_agent_noisy_sweep.yaml \ + --config empiric/benchmark.yaml \ --domain balloons --out docs/prompt-review/2026-09-18-balloons """ import argparse diff --git a/scripts/engaging/launch.py b/scripts/engaging/launch.py index a56d31ea11..64b4107d0c 100644 --- a/scripts/engaging/launch.py +++ b/scripts/engaging/launch.py @@ -5,27 +5,33 @@ job, with one array task per seed, so all experiments run concurrently on compute nodes rather than in the current terminal/login node. -Usage example: +Usage example (continual configs name a round, which suffixes the run +folders so a new launch never resumes an earlier one): - python scripts/engaging/launch.py -c predicatorv3/exp_domino.yaml + python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 mit_normal is often saturated. To run on the much larger (but evictable) preemptable partition instead: - python scripts/engaging/launch.py -c predicatorv3/exp_domino.yaml \ + python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 \ --partition mit_preemptable +To launch a subset of the config's envs, approaches or seeds: + + python scripts/engaging/launch.py -c empiric/benchmark.yaml \ + --round fan_fix_r1 --envs fan --approaches mb_opus --seeds 2-4 + Agent runs draw on a Claude account's usage limit. To spread a launch's runs over several accounts (token files under ~/.claude-tokens, see claude_accounts.py): - python scripts/engaging/launch.py -c predicatorv3/exp_domino.yaml \ + python scripts/engaging/launch.py -c empiric/benchmark.yaml --round r2 \ --partition mit_preemptable --accounts a,b """ import argparse import sys from pathlib import Path -from typing import Optional +from typing import List, Optional, Tuple # Add project root to sys.path so `scripts` is importable without PYTHONPATH=. # parents[0] = scripts/engaging, parents[1] = scripts, parents[2] = repo root @@ -33,7 +39,7 @@ # pylint: disable=wrong-import-position from scripts.cluster_utils import BatchSeedRunConfig, config_to_cmd_flags, \ - config_to_logfile, generate_run_configs + config_to_logfile, generate_run_configs, parse_seed_range from scripts.engaging.claude_accounts import resolve_accounts from scripts.engaging.submit_engaging_job import submit_engaging_job @@ -70,21 +76,73 @@ def _main() -> None: "a token file ~/.claude-tokens/ or the reserved name 'login' " "(the CLI's stored login). Defaults to $PREDICATORS_CLAUDE_ACCOUNTS, " "else 'login'.") + parser.add_argument( + "--round", + type=str, + default=None, + help="Name of this launch's round, appended to every experiment id " + "(-_); overrides the config's ROUND. " + "Continual configs require one.") + parser.add_argument( + "--envs", + type=str, + default=None, + help="Comma-separated env keys of the config to launch, e.g. " + "fan,boil. Defaults to every env the config does not SKIP.") + parser.add_argument( + "--approaches", + type=str, + default=None, + help="Comma-separated approach keys of the config to launch, e.g. " + "mb_opus,mf_opus. Defaults to every approach the config does not " + "SKIP.") + parser.add_argument( + "--seeds", + type=str, + default=None, + help="Seeds to launch, N or N-M (e.g. 2-4); overrides the config's " + "START_SEED and NUM_SEEDS.") args = parser.parse_args() - _launch_experiments(args.config, args.partition, args.requeue, - args.accounts) + _launch_experiments( + args.config, + args.partition, + args.requeue, + args.accounts, + round_name=args.round, + envs=_keys(args.envs), + approaches=_keys(args.approaches), + seeds=parse_seed_range(args.seeds) if args.seeds is not None else None) + + +def _keys(text: Optional[str]) -> Optional[List[str]]: + """Split a comma-separated --envs or --approaches value.""" + if text is None: + return None + return [key.strip() for key in text.split(",") if key.strip()] def _launch_experiments(config_file: str, partition: Optional[str] = None, requeue: Optional[bool] = None, - accounts: Optional[str] = None) -> None: - # Validate the account list once, before anything is submitted. + accounts: Optional[str] = None, + round_name: Optional[str] = None, + envs: Optional[List[str]] = None, + approaches: Optional[List[str]] = None, + seeds: Optional[Tuple[int, int]] = None) -> None: + # Validate the account list and resolve every run once, before + # anything is submitted. account_names = resolve_accounts(accounts) + run_configs = list( + generate_run_configs(config_file, + batch_seeds=True, + round_name=round_name, + envs=envs, + approaches=approaches, + seeds=seeds, + require_round=True)) # Loop over run configs. The experiment's index staggers the account # round-robin across sibling experiments (claude_accounts.py). - for index, cfg in enumerate( - generate_run_configs(config_file, batch_seeds=True)): + for index, cfg in enumerate(run_configs): assert isinstance(cfg, BatchSeedRunConfig) cmd_flags = config_to_cmd_flags(cfg) log_dir = "logs" diff --git a/scripts/engaging/relaunch_on_timeout.py b/scripts/engaging/relaunch_on_timeout.py deleted file mode 100644 index 0113b2e600..0000000000 --- a/scripts/engaging/relaunch_on_timeout.py +++ /dev/null @@ -1,156 +0,0 @@ -"""Watch Slurm jobs and relaunch each one's experiment once on TIMEOUT. - -Preemption self-heals via ``sbatch --requeue``, but a job that hits its -wall-clock limit is simply killed. For auto_resume experiments the fix is -to resubmit the identical command so the run continues from its latest -checkpoint. This watcher does that, once per watched job, and exits when -every watched job has reached a terminal state. - -Each watched job maps to ONE experiment (approach x env arm) inside its -launch config, and only that experiment is resubmitted. Relaunching the -whole config would also resubmit sibling arms that may still be running -(jobs in one generation drift apart through preemption requeues), and a -duplicate arm races the live one on the same auto_resume checkpoints. - -Usage (detach with nohup; poll every 10 min): - - nohup python scripts/engaging/relaunch_on_timeout.py \ - -p mit_preemptable \ - 21338737:predicatorv3/exp_bridge_v2.yaml:bridge_v2-agent_pi_al \ - 21338740:predicatorv3/exp_bridge_v2.yaml:bridge_v2-agent_pi_al_pol \ - >> logs/auto_relaunch.log 2>&1 & - -Pass the launch's --accounts list too, so a relaunched experiment keeps -the Claude accounts its seeds started on (the assignment is a function -of the config, the account list and the experiment's index). -""" - -import argparse -import subprocess -import sys -import time -from dataclasses import dataclass -from pathlib import Path -from typing import List, Optional - -# Add project root to sys.path so `scripts` is importable without PYTHONPATH=. -sys.path.insert(0, str(Path(__file__).resolve().parents[2])) - -# pylint: disable=wrong-import-position -from scripts.cluster_utils import BatchSeedRunConfig, config_to_cmd_flags, \ - config_to_logfile, generate_run_configs -from scripts.engaging.claude_accounts import resolve_accounts -from scripts.engaging.submit_engaging_job import submit_engaging_job - - -@dataclass -class _WatchedJob: - job_id: str - config_file: str - experiment_id: str - resolved: bool = False - - -def _parse_spec(spec: str) -> _WatchedJob: - parts = spec.split(":") - if len(parts) != 3: - raise ValueError(f"Expected JOBID:CONFIG:EXPERIMENT_ID, got {spec!r}") - return _WatchedJob(job_id=parts[0], - config_file=parts[1], - experiment_id=parts[2]) - - -def _in_queue(job_id: str) -> bool: - out = subprocess.run(["squeue", "-h", "-j", job_id], - capture_output=True, - text=True, - check=False) - return bool(out.stdout.strip()) - - -def _terminal_state(job_id: str) -> Optional[str]: - """The job's sacct state, or None if sacct has nothing yet.""" - out = subprocess.run(["sacct", "-j", job_id, "-X", "-n", "-o", "State%30"], - capture_output=True, - text=True, - check=False) - lines = [ln.strip() for ln in out.stdout.splitlines() if ln.strip()] - return lines[0] if lines else None - - -def _relaunch_experiment(config_file: str, - experiment_id: str, - partition: str, - accounts: Optional[str] = None) -> None: - """Resubmit only ``experiment_id`` from ``config_file``.""" - account_names = resolve_accounts(accounts) - for index, cfg in enumerate( - generate_run_configs(config_file, batch_seeds=True)): - assert isinstance(cfg, BatchSeedRunConfig) - if cfg.experiment_id != experiment_id: - continue - cmd_flags = config_to_cmd_flags(cfg) - log_prefix = config_to_logfile(cfg, suffix="") - submit_engaging_job("main.py", cfg.experiment_id, "logs", log_prefix, - cmd_flags, cfg.start_seed, cfg.num_seeds, - cfg.use_gpu, cfg.use_mujoco, partition, True, - account_names, index) - return - print( - f"WARNING: experiment {experiment_id} not found in {config_file}; " - "nothing relaunched", - flush=True) - - -def _main() -> None: - parser = argparse.ArgumentParser() - parser.add_argument("specs", - nargs="+", - help="JOBID:CONFIG:EXPERIMENT_ID triples") - parser.add_argument("-p", "--partition", default="mit_preemptable") - parser.add_argument("--poll-seconds", type=int, default=600) - parser.add_argument("--accounts", - type=str, - default=None, - help="The launch's Claude account list (see " - "claude_accounts.py); defaults to " - "$PREDICATORS_CLAUDE_ACCOUNTS, else 'login'.") - args = parser.parse_args() - # Fail on a bad account list now, not at the first TIMEOUT. - resolve_accounts(args.accounts) - jobs: List[_WatchedJob] = [_parse_spec(s) for s in args.specs] - stamp = time.strftime("%m-%d %H:%M") - print( - f"{stamp} watching {len(jobs)} jobs: " - f"{', '.join(j.job_id for j in jobs)}", - flush=True) - while not all(j.resolved for j in jobs): - time.sleep(args.poll_seconds) - for job in jobs: - if job.resolved or _in_queue(job.job_id): - continue - state = _terminal_state(job.job_id) - if state is None: - continue - stamp = time.strftime("%m-%d %H:%M") - if state.startswith("TIMEOUT"): - print( - f"{stamp} job {job.job_id} TIMEOUT -> relaunching " - f"{job.experiment_id} from {job.config_file}", - flush=True) - _relaunch_experiment(job.config_file, job.experiment_id, - args.partition, args.accounts) - else: - print( - f"{stamp} job {job.job_id} ended {state}; " - "no relaunch", - flush=True) - job.resolved = True - print( - f"{time.strftime('%m-%d %H:%M')} all watched jobs resolved; " - "exiting", - flush=True) - - -if __name__ == "__main__": - _main() diff --git a/scripts/local/generate_random_action_gifs.py b/scripts/local/generate_random_action_gifs.py index 383e2fe63d..05f0db1fe4 100644 --- a/scripts/local/generate_random_action_gifs.py +++ b/scripts/local/generate_random_action_gifs.py @@ -10,7 +10,7 @@ Options: --skip-run Skip running the experiments (just convert existing MP4s) --config, -c Config file to use - (default: mara2/random_actions_pybullet.yaml) + (default: ExoPredicator/random_actions_pybullet.yaml) --video-dir Directory where MP4s are written (default: videos) --output-dir Directory for output GIFs (default: docs/envs/assets/random_action_gifs) @@ -114,7 +114,7 @@ def main() -> None: parser.add_argument( "-c", "--config", - default="mara2/random_actions_pybullet.yaml", + default="ExoPredicator/random_actions_pybullet.yaml", help="Config YAML file (relative to scripts/configs/).", ) parser.add_argument( diff --git a/scripts/local/render_init_state_gifs.py b/scripts/local/render_init_state_gifs.py index ffb230549c..0d71d54148 100644 --- a/scripts/local/render_init_state_gifs.py +++ b/scripts/local/render_init_state_gifs.py @@ -6,9 +6,11 @@ experiments use: - the five benchmark domains take their entry in - scripts/configs/predicatorv3/envs/continual.yaml, -- the other recent domains take their entry in envs/all.yaml, -- the older domains take their entry in random_actions_pybullet.yaml. + scripts/configs/empiric/envs.yaml, +- the older domains take their entry in + scripts/configs/ExoPredicator/random_actions_pybullet.yaml, +- the rest (busyboard, crane, icerink, launcher, magnets) render with their + defaults; no experiment changed their task distributions. Observation noise flags are irrelevant here (states are rendered, not observed). Run on a compute node, one env per call: @@ -27,32 +29,30 @@ from predicators import utils from predicators.envs import create_new_env -_CONFIG_DIR = "scripts/configs/predicatorv3" +_CONFIG_DIR = "scripts/configs" -# env name -> (menu file, menu key) +# env name -> (menu file, menu key) for the benchmark domains. _MENUS: Dict[str, Tuple[str, str]] = { - "pybullet_balloons": ("envs/continual.yaml", "balloons"), - "pybullet_bridge": ("envs/continual.yaml", "bridge"), - "pybullet_boil": ("envs/continual.yaml", "boil"), - "pybullet_fan": ("envs/continual.yaml", "fan"), - "pybullet_domino": ("envs/continual.yaml", "domino_high_friction_turn"), - "pybullet_busyboard": ("envs/all.yaml", "busyboard"), - "pybullet_crane": ("envs/all.yaml", "crane"), - "pybullet_icerink": ("envs/all.yaml", "icerink"), - "pybullet_launcher": ("envs/all.yaml", "launcher"), - "pybullet_magnets": ("envs/all.yaml", "magnets"), + "pybullet_balloons": ("empiric/envs.yaml", "balloons"), + "pybullet_bridge": ("empiric/envs.yaml", "bridge"), + "pybullet_boil": ("empiric/envs.yaml", "boil"), + "pybullet_fan": ("empiric/envs.yaml", "fan"), + "pybullet_domino": ("empiric/envs.yaml", "domino_high_friction_turn"), } -_LEGACY_MENU = "random_actions_pybullet.yaml" +_LEGACY_MENU = "ExoPredicator/random_actions_pybullet.yaml" def _menu_flags(env_name: str) -> Dict[str, Any]: - """Return the FLAGS of the menu entry that defines this env's tasks.""" + """Return the FLAGS of the menu entry that defines this env's tasks, or + none for an env that no menu lists.""" menu_file, key = _MENUS.get(env_name, (_LEGACY_MENU, "")) with open(os.path.join(_CONFIG_DIR, menu_file), encoding="utf-8") as f: config = yaml.safe_load(f) envs = config["ENVS"] if not key: - key = next(k for k, v in envs.items() if v["NAME"] == env_name) + key = next((k for k, v in envs.items() if v["NAME"] == env_name), "") + if not key: + return {} flags = dict(envs[key].get("FLAGS", {})) return {k: v for k, v in flags.items() if not k.startswith("continual_")} diff --git a/scripts/plotting/monitor_benchmark_arms.py b/scripts/plotting/monitor_benchmark_arms.py index e0a0ab0f8e..b6dbce8411 100644 --- a/scripts/plotting/monitor_benchmark_arms.py +++ b/scripts/plotting/monitor_benchmark_arms.py @@ -273,7 +273,8 @@ def report_text(original: str, rows: List[Row], plot: Dict[str, Any], "See the [illustrated task description]" "(../amps/fan-exposed-transfer.md) and " "[launch configuration]" - "(../../scripts/configs/predicatorv3/" + "(https://github.com/BasisResearch/predicators/blob/" + "iclr-empiric-submission/scripts/configs/predicatorv3/" "continual_fan_transfer_pilot_r1.yaml).\n" "This column is excluded from the paper figure and its " "data selection.\n\n") diff --git a/tests/agent_sdk/prompt_goldens/explore_query.md b/tests/agent_sdk/prompt_goldens/explore_query.md deleted file mode 100644 index 853eca4f3c..0000000000 --- a/tests/agent_sdk/prompt_goldens/explore_query.md +++ /dev/null @@ -1,60 +0,0 @@ -Design this episode's experiment for the task below. Reaching the goal is the most informative experiment available, so treat solving the task as part of information gathering. - -## Goal Description - -Put the thing on the fixture and switch it on. - -## Goal Atoms - -Active(fixture0:fixture) -At(thing0:thing, fixture0:fixture) - -## Experiment Guidance - -The learning phase left this ranked ledger of open questions: -1. Does Active require At? Experiment: MoveTo then Wait. - -## Initial State Atoms - -(none: no atom of the available predicates holds initially) - -## Initial State Features - - {'fixture0:fixture': {'x': 1.0000, 'y': 0.0000, 'is_on': 0.0000}, - 'thing0:thing': {'x': 0.0000, 'y': 0.0000}} - -## Objects - - fixture0: fixture - thing0: thing - -## Available Options - - MoveTo(thing, fixture) [params: dx, dy, range [-0.5, -0.5] to [0.5, 0.5]] - Wait() - -## Available Predicates (for subgoal annotations) - - Active(fixture) - At(thing, fixture): thing rests on fixture - -## Available Tools - - - submit_plan - - run_python - -## Plans Already Scheduled This Cycle - -The plan(s) below are already queued to run on this task before any learning happens, so their data will be collected regardless of what you propose now. - -Plan 1: - 0: MoveTo(thing0, fixture0)[0.2000, 0.0000] - NOTE: belief-certified; executes verbatim as a solve attempt. - -Propose a plan whose data is complementary rather than redundant: cover the open questions, mechanisms, or parameter regions the plan(s) above leave unmeasured. A goal-reaching plan is still preferred when it can carry that coverage; when it cannot, a designed experiment that settles what the scheduled plans will not is the better use of this episode. Only when the model is believed correct everywhere and no meaningfully different goal-reaching plan exists, repeat the best plan. - -If a scheduled plan is marked belief-certified, this episode is the second test of the belief model, and one success of one plan is weak evidence. In order of preference: (1) a STRUCTURALLY different goal-reaching plan (a different option sequence, order, grasp, or contact arrangement) validated through the same `submit_plan` gate; (2) when no structurally different plan exists for this goal, the same structure with materially different parameters (a different placement pose, offset, or timing, not a jitter), validated the same way; (3) only as a last resort, the certified plan resubmitted unchanged. State which of the three you chose and why. A certified plan that then fails for real is the most informative outcome this episode can produce, not a loss. - -## Instructions - -Inspect the environment with your tools, design the episode as the system prompt's Exploration Setting specifies, and output the plan lines as your final text. diff --git a/tests/agent_sdk/prompt_goldens/learn_message.md b/tests/agent_sdk/prompt_goldens/learn_message.md deleted file mode 100644 index 053bf139c6..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_message.md +++ /dev/null @@ -1,72 +0,0 @@ -Synthesize a residual dynamics simulator for this environment. There are 3 trajectories (42 step transitions) available: 1 oracle demonstration(s), which reached the goal by construction, and 2 interaction trajectory/ies collected during online learning, some of which may have failed to reach the goal. - -[0] demo, task 0 -[1] interaction, task 0 - -Each trajectory carries a `train_task_idx`. `is_goal_state(state, task_idx)` (equivalently `train_tasks[task_idx].goal_holds(state)`) checks a single state for the goal atoms. Reaching the goal atoms does not by itself mean an episode is solved; when a task objective is stated below, score full trajectories with `evaluate_trajectory`. Use `is_goal_state` to confirm which trajectories reached the goal atoms and to treat failed interaction trajectories as counterexamples: places where a predicate or rule said "this should work" and the environment disagreed. - -## Task objective (env ground-truth reward) - -reward = success - 0.1 x moves used - -The trajectory roster above shows each interaction episode's env-computed reward. In `run_python`, `evaluate_trajectory(states, actions=None, task_idx=0)` is the task's reward model: it scores any state sequence with the same rules, a collected trajectory's `states` and `actions`, or a rollout of your simulator (where a rule that replays physics runs on your belief simulator at its current fit, so the verdict is only as trustworthy as the simulator). It returns `{reward, solved, note}`; `solved` means the episode is scored as a success, a rollout can reach the goal atoms and still be `solved=False`, and `note` says what a replaying rule simulated and on what. Label transitions with `(option, objects, params)` (`None` for an unlabeled one) so such a rule replays your action rather than its canonical one. - -Prior cycle state: `./simulator.py` and `./predicates.py` already exist in the sandbox from a previous learning cycle. Read them first: they are the previous cycle's committed result and a reasonable starting point for incremental refinement, though a fresh rewrite is fine if the prior approach looks fundamentally wrong. Structural decisions are not binding across cycles: re-read the decision record at the top of `simulator.py` and re-decide the architecture itself (what the base sim carries and what the rules model, which features the rules own, the latent structure, whether disclosed base-sim parameters should be identified) rather than only tuning what exists. In particular, if the trajectory roster shows goal-reaching episodes scored `solved=0`, suspect a structural modeling error (for example mis-calibrated base physics that the rules only paper over near the fit data), not only parameter values. Earlier versions are in `./simulator_versions/` and `./predicates_versions/` (named `cycle_XXX_vers_YYY_*.py`); cross-reference the roster's provenance tags against those files to see which rules and predicates produced each failed plan. - -## Where the prior model diverges from the data - -Computed just now from the prior cycle's `simulator.py` with its parameters refit to all trajectories above, so the remaining mismatches need structural fixes, not tuning. Re-score any edit with `sim.residuals()` (same report, current file): - -fixture.is_on: 4 mismatches - -Data-structure source code is at: ./reference/structs.py - -The base simulator's own source code is available (read-only): - - - ./reference/base_sim/scene.py - -These files are byte-identical to the code your base-sim rollouts execute: scene geometry and constants, body construction, stepping, and state read/write. They deliberately omit the environment's hidden domain-specific step, the residual dynamics you are here to model, and its task generation and goal semantics. Use them to ground hypotheses (masses, damping, substeps per action, how switches toggle) instead of re-measuring those from data. - -A residual scan between the base simulator's prediction and the observed next state suggests that these features carry residual dynamics (a starting hint; it may include base-sim jitter, so refine it as you go): - -{'fixture': ['is_on']} - -## Available Predicates (for subgoal annotations) - - Holding(robot, thing) - -Subgoal annotations in plans for `sim.refine` / `sim.run` must reference these predicate names with matching arity and types. Any threshold or condition you bake into a rule must be consistent with what the predicate's classifier checks, or refinement rejects parameter samples that look correct on paper. - -## Object Types - -thing: x, y -fixture: x, y, is_on - -## Options - -Plans (for `sim.refine` / `sim.run`) and rules must match these typed signatures and parameter boxes exactly: - -MoveTo(thing, fixture)[dx, dy] - -## Available Tools - - - run_python - - Read - - Write - - Edit - -## This session - -Read the data-structures file first, then explore the trajectory data with `run_python`. Write your simulator to `./simulator.py`, exporting a `RESIDUAL_ENV` subclass with `AGENT_PARAM_SPECS` and `RESIDUAL_FEATURES`, and iterate with `Edit` and re-scoring. Pass `task_idx` explicitly to `sim.reset`; `sim.task(task_idx)` prints a task digest. Finish with the deliverables listed in the system prompt: a final `sim.fit()`, the GO/NO-GO check, the decision record, `./open_questions.md`, and `./strategy.md`. - -## Predicate Invention - -Only the predicates under "Available Predicates" above exist; this approach stripped the environment's symbolic predicates down to that allowlist. Invent every other subgoal predicate in `./predicates.py` as `LEARNED_PREDICATES`, following the system prompt's "Predicate Invention" section. - -Goal (natural language): switch the fixture on. - -Workflow: edit `predicates.py`, call `sim.predicates()` in `run_python`, then run `sim.refine` / `sim.run` with sketches that reference your invented names. Any predicate a sketch references must exist in `predicates.py` first. - -## Partial observability - -Some causally important quantities may be absent from the observation entirely (under no name), possibly several, possibly none. Inspect the trajectories first to judge whether any hidden process is at work and which observable features are your window into it; then, if latents are needed, declare subclass `MODEL_STATE_INIT` and implement `update_model_state`. diff --git a/tests/agent_sdk/prompt_goldens/learn_notes_message.md b/tests/agent_sdk/prompt_goldens/learn_notes_message.md deleted file mode 100644 index b93f415b47..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_notes_message.md +++ /dev/null @@ -1,35 +0,0 @@ -Write the world model document for this environment. There are 2 recorded trajectories (9 skill-level transitions) available: 0 oracle demonstration(s), which reached the goal by construction, and 2 interaction trajectory/ies collected during online learning, some of which may have failed to reach the goal. - - [0] interaction, task 0 - [1] interaction, task 0 - -Each trajectory carries a `train_task_idx`. `is_goal_state(state, task_idx)` (equivalently `train_tasks[task_idx].goal_holds(state)`) checks a single state for the goal atoms. Use it to confirm which trajectories reached the goal and to treat failed interaction trajectories as counterexamples: places where the environment disagreed with what a skill was expected to do. - -## Task goals (natural language) - -- Build the bridge. - -A `world_model.md` from an earlier cycle exists at `./world_model.md`. Read it first; this cycle's data may confirm, refine, or contradict what it says. Revise it in place. - -Data-structure source code is at: ./reference/structs.py - -## Available Predicates - -- Holding(robot:robot, block:block) - -## Object Types - -- robot: hand -- block: x, y, held - -## Options - -- Pick(robot:robot, block:block)[] - -## Available Tools - - - run_python - -## This session - -Read the data-structures file first, then explore the trajectory data with `run_python`. Write your world model to `./world_model.md` under the headings given in the system prompt, and finish with the deliverables listed there. diff --git a/tests/agent_sdk/prompt_goldens/learn_notes_system.md b/tests/agent_sdk/prompt_goldens/learn_notes_system.md deleted file mode 100644 index fe416ea3f5..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_notes_system.md +++ /dev/null @@ -1,31 +0,0 @@ -You are building a world model for a robotic manipulation environment as a natural-language document. No simulator will run from what you write: at planning time the same document is all the knowledge of the environment's dynamics the planner has, and it plans by reasoning over it, so what you write must let a careful reader predict what every skill does, when it works, and how the environment's own processes unfold over time. - -## What you produce - -One file, `world_model.md` (path given in the first message). Keep it organized under fixed headings so later cycles and the planner can find things: - -1. `# Mechanisms`: every process the environment runs on its own (delayed effects, gradual changes, propagation between objects, hidden state that changes what skills do), each with its trigger condition, its rate or duration in low-level steps, what it changes and by how much, and the evidence (trajectory and step) it comes from. -2. `# Skills`: for every skill, what it changes in the observed state when it succeeds (with the numbers: offsets, final poses, feature values as a function of the parameters), the conditions under which it fails and what the failure looks like, how many low-level steps it takes, and which of its continuous parameters matter and over what ranges. -3. `# Thresholds and geometry`: the quantitative gates the environment enforces (how close is close enough, which side of a fixture, what counts as supported), each bracketed by recorded attempts on both sides where the data allows. -4. `# Hidden state`: what the observation does not show, how it can be inferred from what it does show and from the history of skills executed, and how it evolves. -5. `# Recipes`: skill sequences, with parameter values, that the data shows reaching intermediate goals, and why they work. -6. `# Uncertainties and open questions`: what the data does not settle, phrased as the experiment that would settle it. - -Write for prediction, not description: a reader must be able to take a state and a skill call and write down the state after it. Prefer numbers over adjectives, and say where each number comes from. When you are unsure, say so and give the range. - -## Tools - -`run_python` is the one tool over the data: `trajectories` (`List[LowLevelTrajectory]`; each action's `get_option()` is the skill that produced it, so the skill-level transitions are the spans between skill changes), `describe_trajectory(i)`, `train_tasks`, `is_goal_state(state, task_idx)`, and `np`. Use `Read`, `Write` and `Edit` on `world_model.md`. - -## Deliverables of a learning session - -- The document, complete under the six headings above, with every mechanism the recorded episodes exercised reconciled against what you wrote before (earlier cycles' notes are yours to revise, not to append to). -- A decision record at the top: the key modeling commitments, the evidence behind each, and every hypothesis you kept without direct evidence, labelled as such. -- `./open_questions.md` with what the next exploration should collect first, and `./strategy.md` with how you would solve the train task given what you now know. - -## Workflow - -1. Explore the data with `run_python`: for each skill, which features change between its start and its end, under what conditions, and by how much; for each feature that changes while no skill touches it, what drives it. -2. `Write` or `Edit` `world_model.md`, one heading at a time, with the numbers and their evidence. -3. Check every claim against a transition it should predict: pick a recorded skill call, predict its outcome from your notes alone, compare. Fix the notes where the prediction is wrong. -4. Finish with the deliverables above. diff --git a/tests/agent_sdk/prompt_goldens/learn_program_message.md b/tests/agent_sdk/prompt_goldens/learn_program_message.md deleted file mode 100644 index 67e9c09b38..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_program_message.md +++ /dev/null @@ -1,34 +0,0 @@ -Synthesize a world model program for this environment. There are 0 recorded trajectories (0 skill-level transitions) available: 0 oracle demonstration(s), which reached the goal by construction, and 0 interaction trajectory/ies collected during online learning, some of which may have failed to reach the goal. - -Each trajectory carries a `train_task_idx`. `is_goal_state(state, task_idx)` (equivalently `train_tasks[task_idx].goal_holds(state)`) checks a single state for the goal atoms. Reaching the goal atoms does not by itself mean an episode is solved; when a task objective is stated below, score full trajectories with `evaluate_trajectory`. Use `is_goal_state` to confirm which trajectories reached the goal atoms and to treat failed interaction trajectories as counterexamples: places where the environment disagreed with what a skill was expected to do. - -Data-structure source code is at: ./reference/structs.py - -## Available Predicates (for subgoal annotations) - -- Holding(robot:robot, block:block) - -Subgoal annotations in plans for `sim.refine` / `sim.run` must reference these predicate names with matching arity and types. - -## Object Types - -- robot: hand -- block: x, y, held - -## Options - -Plans (for `sim.refine` / `sim.run`) and your `transition` must match these typed signatures and parameter boxes exactly: - -- Pick(robot:robot, block:block)[] - -## Available Tools - - - run_python - -## This session - -Read the data-structures file first, then explore the trajectory data with `run_python`. Write your world model to `./world_model.py`, defining `LATENT_FEATURES`, `initial_latent`, and `transition`, and iterate with `Edit` and `sim.score()`. Pass `task_idx` explicitly to `sim.reset`; `sim.task(task_idx)` prints a task digest. Finish with the deliverables listed in the system prompt: a final `sim.score()`, the GO/NO-GO check, the decision record, `./open_questions.md`, and `./strategy.md`. - -## Zero-shot synthesis - -No trajectory has been recorded and none will be before you finish: this session is the whole learning phase, and what you write here is what the planner uses on the test tasks. The trajectory counts above are zero for that reason, and `sim.score` has no data to score against. Build the world model from the task description, the object types and options, the scene (`sim.task`, `sim.reset`, `sim.render`) and your own knowledge of the mechanisms involved, and validate it with `sim.refine` / `sim.run` rollouts of a full plan. State each mechanism you commit to, and the evidence you would want for it, in the decision record. diff --git a/tests/agent_sdk/prompt_goldens/learn_program_system.md b/tests/agent_sdk/prompt_goldens/learn_program_system.md deleted file mode 100644 index 973379191f..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_program_system.md +++ /dev/null @@ -1,135 +0,0 @@ -You are synthesizing a world model for a robotic manipulation environment as a standalone program: given the observed state and the skill the robot executes, predict the observed state after the skill completes. There is no physics engine behind your program. Robot motion, grasping, contact, placement, and every process the environment runs (delayed effects, gradual changes, propagation between objects, hidden mechanisms) are yours to model, at the level of one skill call at a time. - -## What you produce - -One file, `world_model.py` (path given in the first message), defining three top-level names: - -```python -LATENT_FEATURES: Dict[str, List[str]] # {type_name: [hidden feature names]} your latent tracks - -def initial_latent(obs: State, rng: np.random.Generator) -> Dict[str, Any]: - """A draw of the hidden state consistent with the first observation.""" - -def transition(obs: State, latent: Dict[str, Any], option: _Option, - rng: np.random.Generator) -> Tuple[State, Dict[str, Any], int]: - """The observed state after `option` runs to completion from `obs`, - the updated hidden state, and the number of low-level steps used.""" -``` - -`obs` is a `State` over the environment's objects with only the OBSERVABLE features (`obs.get(obj, "x")`; `obs.set(obj, "x", v)` on the copy you return; `list(obs)` iterates the objects). `option` is a ground skill: `option.name`, `option.objects` (typed, in signature order), `option.params` (the continuous parameter vector, in the option's box), and for `Wait` the target atoms in `option.memory.get("wait_target_atoms")`. Return a new `State` with exactly the same objects (start from `obs.copy()`), your updated latent dict, and a positive step count (the environment's horizon is counted in low-level steps, so a skill that takes longer must cost more). - -The latent is yours: a plain dict of whatever the environment hides (process progress, attachments, cure state, per-object counters). Declare in `LATENT_FEATURES` what it tracks. `initial_latent` may be stochastic through `rng` - when the first observation leaves the hidden state genuinely undetermined, return a draw over the possibilities: the harness keeps a particle belief of several draws, scores the model with it, and re-validates every plan under every particle. A deterministic `initial_latent` is a belief with one particle. - -## Modeling guidance - -- Model a skill's effect on every feature it changes, not only on the ones the goal names. Gripper state, the held object's pose while it is carried, the poses of objects that move together, and the features a process advances are all read by the planner's predicates and by the next skill. -- Skills fail. When a skill's parameters put its target out of reach, into collision, or onto an unsupported spot, return an outcome the environment would produce (the object drops, stays put, the gripper closes on nothing), not the intended one; a model that always succeeds validates plans that fail. -- Processes take time. A hidden process that advances while the robot does other things advances in your latent on EVERY transition (including `Wait`), by an amount tied to the step count you return, so that a plan's timing is checked. Wait terminates when its target atoms hold or on the first observable change; model its duration accordingly. -- Ground every mechanism in the recorded data: find the transitions where a feature changes, characterize when it changes and by how much, and encode that. A mechanism you suspect but never observed is a hypothesis; record it in the decision record and, when the goal requires it, ship it as a labelled hypothesis with the experiment that would confirm it first in `./open_questions.md`. -- Thresholds and geometric gates (how close is close enough, which side of a fixture) come from the data too: find the recorded attempts on both sides of the boundary and place the gate between them. When in doubt, tighten toward the empirical boundary; a permissive model passes plans the environment rejects. -- `predicates.py` has no learned parameters in this arm: write thresholds as literals there, kept consistent with the ones in your transition, or read the hidden state through `state.latent` (a dict while the planner rolls your model; `None` on a raw observation, so a predicate the plan needs on the real robot must not depend on it). - -## Tools - -`run_python` is the one tool over the data, and it carries the `sim` probe over your CANDIDATE world model (reloaded whenever the file changes): - -- `sim.score()`: the model's score on the recorded trajectories - a particle-filter pseudo-likelihood over your hidden state (0 is a perfect model; each unit is one feature-std of mean error per transition), the per-feature error table, and the worst transitions. The inner-loop signal: re-score after every edit, and read the worst transitions to find WHICH mechanism is wrong. `sim.score(traj_idxs=[...])` restricts the data. -- `sim.refine(plan)`: backtracking parameter search on a plan sketch through your model. -- `sim.run(plan)`: forward rollout through your model with subgoal checking. -- `sim.reset(task_idx=..., mods={...})` and `sim.render(label, annotations=[...])`: stage a state and render it with overlays. -- `sim.predicates()`: score `predicates.py` on the recorded data. -- `trajectories`, `describe_trajectory(i)`, `train_tasks`, `is_goal_state`: the recorded evidence. Each action carries the skill that produced it (`action.get_option()`), so the option-level transitions are the spans between skill changes. - -Probe rollouts are candidate predictions; do not confuse them with the recorded `trajectories`. - -### Score vs. forward validation - -`sim.score` and refine-then-run test complementary things: pointwise accuracy on what was recorded versus goal reachability on what a plan needs. A model can score well on the data and still make a gate wide enough that refinement accepts a placement the environment rejects, or advance a process too fast so that a `Wait` looks sufficient in the model and is not on the robot. Use `sim.score` as the fast inner loop and refine-then-run as the slow, goal-relevant gate before declaring done. When `sim.refine` passes but `sim.run` reports a subgoal not reached, the model is more permissive than the environment's effective behavior: tighten the threshold toward the empirical boundary, never loosen it. - -## Predicate Invention (required for plan subgoals) - -You also invent the symbolic predicates the planner uses as subgoal atoms in plan sketches. Only `Holding` is provided as a primitive; placement, device-state, and process-completion predicates do not exist until you invent them. - -Goals are presented in natural language (see the first message) and goal achievement is checked externally by the environment through `is_goal_state(state, task_idx)` / `train_tasks[task_idx].goal_holds(state)`. You need not invent goal-named predicates or match environment predicate names: invented predicates exist for plan-sketch subgoals (gating `Wait`, `Place`, and similar steps) and can be named freely. - -Define them in `predicates.py` (path given in the first message): - -```python -LEARNED_PREDICATES: List[Predicate] -``` - -The exec namespace pre-injects `Predicate`, `np`, and a `_type` binding for each env type (for example `widget_type`, `fixture_type`). The names below are illustrative; use the types, features, and parameter names your digests and the trajectory data report. - -```python -# Placement: object xy within a learned distance of the fixture's -# functional point, NOT its recorded origin (see "Geometric gates"). -# The local-frame offset is declared as ParamSpecs in simulator.py -# and shared with the rule that gates the same physics. -def _widget_at_fixture(s, objs): - widget, fixture = objs - rot = s.get(fixture, "rot") - cos_r, sin_r = np.cos(rot), np.sin(rot) - rot_mat = np.array([[cos_r, -sin_r], [sin_r, cos_r]]) - local_offset = np.array([params["fixture_local_dx"], - params["fixture_local_dy"]]) - origin = np.array([s.get(fixture, "x"), s.get(fixture, "y")]) - anchor = origin + rot_mat @ local_offset # world-frame point - widget_xy = np.array([s.get(widget, "x"), s.get(widget, "y")]) - dist = np.linalg.norm(widget_xy - anchor) - return dist < params["widget_at_fixture_dist"] - -LEARNED_PREDICATES = [ - Predicate("WidgetAtFixture", [widget_type, fixture_type], - _widget_at_fixture), - # Device state: a feature exceeding a fixed cutoff (no learned param). - Predicate("FixtureActive", [fixture_type], - lambda s, objs: s.get(objs[0], "is_on") > 0.5), - # Process completion: a rule-driven feature reaches a learned threshold. - Predicate("WidgetReady", [widget_type], - lambda s, objs: s.get(objs[0], "progress") >= params["ready_threshold"]), -] -``` - -A pre-injected `params` view is in scope and always reads the current fitted values of every `ParamSpec` declared in `simulator.py`; after each refit, predicates reading `params["name"]` see the new values. Whenever one physical gate drives both a rule's firing condition and a predicate's "subgoal reached" check, declare its parameters (the distance threshold and the local-frame anchor offset it is measured from) once in `PARAM_SPECS` and reference `params["name"]` from both. That keeps the two anchored to the same point and gives the offset a fitting signal from the rule's step data. A parameter used only by predicates has no fitting signal and stays at its `init_value`, so choose those initial values carefully. - -What you typically need: - -- Placement predicates (object at a target location) for every open-ended option such as `Place`; without them refinement picks an arbitrary location. -- Device-state predicates (on/off) for every toggle option. -- Process-completion predicates over the features your rules drive, so `Wait` steps know when to terminate. Keep classifier thresholds consistent with the rules' saturation values; an inconsistency makes `sim.fit` look fine while `sim.refine` gets stuck on the `Wait` subgoal. -- Coverage: every option you expect in a sketch should have predicates that express its post-condition, so every sketch step can carry a subgoal annotation. Annotations are checked against the real state during execution to detect and replan diverged steps; a step with no annotatable effect is unmonitored. While drafting sketches, a step you cannot annotate with any invented predicate is a missing predicate. - -Verify every classifier against the scene and the data. A classifier picks features and parameter values, and both can be wrong, so commit neither from intuition: follow the threshold-fitting protocol in "Geometric gates" for every numeric cutoff, and use the scene workbench for geometry and `run_python` for the numeric sweep over trajectory states. - -`sim.predicates()` validates cheaply (first-flip step, monotonicity, coverage across all trajectories) and is also the loader: it updates the predicate set `sim.refine` uses, so call it after every edit to `predicates.py` and before re-running refinement. On goal-reaching trajectories (`reached_goal=True` in `describe_trajectory`) a milestone predicate should flip from false to true exactly once and stay true. On failed interaction trajectories (`reached_goal=False`) the same predicate may fire while the rest of the trajectory shows no goal completion; that is the signature of an over-loose threshold (the predicate fires, the downstream physics does not follow), so tighten it or share the gating parameter with the rule so they are fitted jointly. - -Predicates persist across online cycles: the file is preserved between synthesis sessions, and every successful `Write`/`Edit` (plus a final post-session check) is snapshotted to `predicates_versions/cycle_XXX_vers_YYY_predicates.py`. Each cycle re-runs synthesis with the full trajectory history, so failed past attempts remain visible. - -## Plan format for `sim.refine` / `sim.run` - -One option call per line, with every option argument supplied as a typed object reference (`obj:type`), matching the options digest in your prompt exactly. The parser is strict: an omitted argument is not auto-filled. Example: - -``` -PickWidget(robot:robot, widget0:widget) -Place(robot:robot) -> {WidgetAtFixture(widget0:widget, fixture0:fixture)} -ActivateFixture(robot:robot, fixture0:fixture) -Wait(robot:robot) -> {WidgetReady(widget0:widget)} -``` - -The names are illustrative; use the options, types, and predicates your prompt digests list. Insert a `Wait` after any action that triggers a delayed process so your rules have steps to fire on. - -Subgoal annotations (`-> {Atom(obj:type, ...)}` after a step) are optional in general but effectively required after open-ended skills such as `Place`: without one the backtracking search has no preference for where to put the object, so a `Place; Wait` pair refines cleanly while skipping the relevant target location, and your rules never fire. That looks like a rule bug but is a missing subgoal. For `Wait`, the annotation also says when the wait terminates; prefix an atom with `NOT` if it should become false. - -## Deliverables of a learning session - -- Decision record. Begin `world_model.py` with a short comment stating your key modeling choices and the evidence behind them: which mechanisms the data shows, which features each skill writes, what the latent tracks and how it is initialized, and every hypothesis shipped without direct evidence. Later cycles read this record before deciding what to keep. -- Completeness. Work through every mismatch the score's worst transitions reveal in this one session; each deferred mechanism costs a full explore-learn-test round trip. -- Final score and GO/NO-GO. Before ending, run `sim.score()` on the final file and record it in the decision record. Then refine a full solve of the train task in your model and validate it with several trials (`sim.refine`, then `sim.run(plan, trials=5)`). Record the verdict with the plan's weakest margin, the smallest distance from any step's operating point to a threshold your model enforces. NO-GO means the next test episode will likely fail: put exactly what is missing at the top of `./open_questions.md`, and keep `./strategy.md` current with what the next exploration should collect. - -## Workflow - -1. Explore the data with `run_python`: for each skill, which features change between its start and its end, and under what conditions. -2. `Write` `world_model.py`; `Edit` to iterate. -3. Score with `sim.score()` and read the worst transitions to find the mechanism to fix. Repeat until the remaining error is noise. -4. Propose an option-skeleton plan and validate it: `sim.reset(task_idx=i)`, `sim.refine(plan, require_goal=True)`, then a continuous `sim.run` of the refined plan from a fresh `sim.reset(task_idx=i)`. A stuck refine step means a gate is too tight or a mechanism is missing; a refine-pass whose `sim.run` diverges means the model is too permissive. Fix and re-validate; do not declare done until both pass. Step 4's sketches need subgoal predicates that do not exist until you invent them: before validating, write them to `predicates.py` and load them with `sim.predicates()` (see "Predicate Invention"). -5. Finish with the deliverables above: final `sim.score()`, GO/NO-GO, decision record, `./open_questions.md`, `./strategy.md`. diff --git a/tests/agent_sdk/prompt_goldens/learn_system.md b/tests/agent_sdk/prompt_goldens/learn_system.md deleted file mode 100644 index 2054c86e21..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_system.md +++ /dev/null @@ -1,81 +0,0 @@ -You are synthesizing a parameterized residual-dynamics simulator for a robotic manipulation environment. - -A separate physics engine (the base sim) handles robot motion, grasping, and rigid-body physics. Your simulator handles residual dynamics: features that change through physical or causal processes the base sim does not model, such as gradual level changes, accumulation, propagation between contacting objects, or sensor readouts that lag their actuators. - -## `simulator.py`: a simulator subclass - -Export `RESIDUAL_ENV`, a subclass of the supplied `BaseSimulator`, from `./simulator.py`. `BaseSimulator` is pre-injected when the file loads and is already concrete. It supplies this environment's visible physics, without inheriting hidden mechanism helpers, their constants, or task generators in the five benchmark domains. Mechanism readouts that the visible core cannot compute remain at their restored observed values until your model implements them. Override `_get_domain_specific_feature(self, obj, feature)` for such predicted readouts and `_set_domain_specific_state(self, state)` for their initialization, delegating other features and visible-state restoration to `super()`. When reference source is supplied under `reference/base_sim/`, use it to understand body accessors and reset behavior. Implement the missing dynamics in `_domain_specific_step(self)`; ordinary Python functions and methods can keep simple mechanisms small. Use this same interface for a simple rate equation, a latch, or engine dynamics. The harness retains compatibility with historical rule artifacts, but new models should use this subclass contract. - -Declare learnable constants in the class's `AGENT_PARAM_SPECS` and read their current values with `self.agent_param(name)`. Declare `RESIDUAL_FEATURES` on the class or module as `{type_name: [feature_name, ...]}` to select observed quantities for the fitting loss. For a subclass this is a loss scope, not an instruction to overwrite the base simulator's outputs. Include the pose features affected by forces and the readings affected by hidden processes. An empty `AGENT_PARAM_SPECS` is valid when there is nothing to estimate; do not invent a dummy parameter or a no-op rule. Export only `RESIDUAL_ENV` as the dynamics implementation. - -```python -# BaseSimulator is supplied by the loader. -from predicators.code_sim_learning.fit_space import ParamSpec - -class MyDynamics(BaseSimulator): - AGENT_PARAM_SPECS = [ParamSpec("rate", 0.03, lo=0.0, hi=0.1)] - RESIDUAL_FEATURES = {"widget": ["progress"]} - - def _domain_specific_step(self): - update_widgets(self, self.agent_param("rate")) - -RESIDUAL_ENV = MyDynamics -``` - -`widget` and `update_widgets` above illustrate the structure; use this environment's types and implement the helper from observed evidence. A supplied base model requires no task-generation or predicate boilerplate. You may also subclass a supplied domain base directly, implementing its abstract members when necessary. Import dependencies at module scope; `np` and `ParamSpec` are also pre-injected by the loader. - -## Step and restoration behavior - -Each primitive action advances the base physics, updates declared model memory, then calls `_domain_specific_step` once. Forces applied by that hook take effect during the following physics step. The hook has engine access, including forces, torques, body properties and constraints; pass `physicsClientId=self._physics_client_id` to PyBullet calls. Use real engine constraints for bodies that must move together. Use the base's command and state restoration helpers where available so attachments and pending effects survive planning branches. Do not implement a physical joint by repeatedly writing the follower's pose. - -Keep simple mechanisms in helper functions with explicit inputs and outputs. Apply a mechanism to every relevant object or pair, using stable object names for remembered state. Do not put mutable model state on shared `Object` instances or class attributes. Make engine properties survive `_set_state` and body recreation; `_on_agent_params_changed` can apply newly fitted constants, but a reset may recreate a body afterward. Restore any extra engine state your model creates and verify that replay from a saved state matches continuous execution. If inferred memory creates attachments or other persistent engine effects, implement `restore_model_state(self)` to realize them immediately after reset, before controller initiation and motion planning. The hook must be idempotent: do not step physics, advance counters, snap poses, or infer new joints there. For rigid links inferred by your own observation-driven model, call `self.restore_model_attachments([(name_a, name_b), ...])` from this hook and when the inferred links change during dynamics. This registers links for held-assembly collision checking and snapshot restoration; creating an unregistered engine constraint is insufficient. The helper does not supply attachment rules or infer links from the real environment. Run `sim.reset(current=True).check_restore()` after model edits and before trusting a held-assembly rehearsal. It checks pose and inferred-memory round trips in fresh worlds without physical steps; a pass does not establish that your inferred memory is correct. - -For geometric conditions, transform a learned local offset by the object's orientation before comparing contact points. Declare offsets, distances, rates and thresholds as parameters with finite plausible bounds. Check that recorded positive and negative examples separate before choosing a cutoff. Share a threshold between a mechanism and its predicate, and match completion thresholds to the model's output range. Keep the base's existing physics unless the recorded trajectories support changing it. - -## Supplied physical parameter menu - -The base simulator exposes these tunable quantities: - -- `lateral_friction` (built-in 0.5, fit box [0.05, 2], fitted in log-space): sliding friction of every body - -Declare parameters you want to estimate in `AGENT_PARAM_SPECS`. Names from this menu are connected to the supplied base's physical setters automatically. For additional constants, implement their effect in your own hook, read them with `self.agent_param(name)`, and ensure changes survive state restoration. Use full rollout validation to check that each fitted parameter affects the intended mechanism. - -## Fit and validate complete rollouts - -Edit `./simulator.py`, then explicitly call `sim.fit()` to estimate its declared constants from the recorded trajectories. Edits are loaded on the next probe call; a rollout does not implicitly fit parameters. Before fitting, the model uses its carried or declared values and is marked UNFITTED. If there are no learnable constants, skip fitting and call `sim.validate()`. - -`sim.validate()` replays every selected recording at the values currently deployed for planning, including recordings a robust fit rejected. `sim.residuals()` uses full simulator replay for subclass models to expose accumulated error; the report labels the parameter values it scores. `sim.fit(traj_idxs=[...])` and explicit validation parameter overrides are diagnostics and publish nothing. Compare candidates on the same recordings, inspect per-trajectory failures and preserve counterexamples. A low fitting error on a selected subset does not establish model fidelity or task solvability. - -Use `sim.refine(plan)` to search skill parameters, then run the resulting plan continuously with `sim.run(plan)` and check each annotated subgoal. Use `sim.reset(task_idx=..., mods=...)` and `sim.render(label, annotations=[...])` to inspect geometry. Evaluate trajectory success with the supplied evaluator when available; its verdict on a simulated trajectory depends on the model's fidelity. Prefer additional simulator checks over spending real steps on a prediction that disagrees with recorded evidence. Keep speculative mechanisms labeled as hypotheses and state what observation would distinguish competing explanations. - -## Plan format for `sim.refine` / `sim.run` - -One option call per line, with every option argument supplied as a typed object reference (`obj:type`), matching the options digest in your prompt exactly. The parser is strict: an omitted argument is not auto-filled. Example: - -``` -PickWidget(robot:robot, widget0:widget) -Place(robot:robot) -> {WidgetAtFixture(widget0:widget, fixture0:fixture)} -ActivateFixture(robot:robot, fixture0:fixture) -Wait(robot:robot) -> {WidgetReady(widget0:widget)} -``` - -The names are illustrative; use the options, types, and predicates your prompt digests list. Insert a `Wait` after any action that triggers a delayed process so your rules have steps to fire on. - -Subgoal annotations (`-> {Atom(obj:type, ...)}` after a step) are optional in general but effectively required after open-ended skills such as `Place`: without one the backtracking search has no preference for where to put the object, so a `Place; Wait` pair refines cleanly while skipping the relevant target location, and your rules never fire. That looks like a rule bug but is a missing subgoal. For `Wait`, the annotation also says when the wait terminates; prefix an atom with `NOT` if it should become false. - -## Deliverables of a learning session - -- Begin `simulator.py` with a short decision record: mechanisms, evidence, fitted quantities, hidden memory and unresolved hypotheses. -- Reconcile every mechanism exercised by the recordings with the model. Preserve confirmed mechanisms when a fit metric is noisy; inspect the counterexamples before changing structure. -- Ground physical changes in recorded behavior the base mispredicts. Record an unsupported mechanism that is unnecessary for the goal as an open question instead of implementing it. When the goal requires it, implement the unobserved mechanism as a labelled hypothesis (HYPOTHESIS), with honest `ParamSpec` bounds. Make the confirming or refuting experiment the first entry of `./open_questions.md`, naming the observation that distinguishes the alternatives. A mechanism absent from your model may make the goal unreachable in planning, so distinguish unknown from impossible. -- Declare uncertain constants as `ParamSpec`s with plausible ranges. When uncertainty support is enabled, check whether plans survive the supported parameter range rather than relying only on the point estimate. -- Run a final explicit `sim.fit()` if the model declares learnable constants, then `sim.validate()` on the full recordings. Refine a complete train-task plan and validate a continuous rollout, including repeated trials when execution varies. Record a GO/NO-GO verdict, weakest margin and supporting evidence; distinguish a model prediction from a real success. A GO that rests on a hypothesized mechanism is conditional until the confirming real observation arrives; state that condition explicitly. -- Write `./open_questions.md` as a ranked list of unresolved mechanisms or parameters. Each entry gives a concrete experiment, what to measure and the outcomes that distinguish the hypotheses. Remove questions the new evidence settles. -- Write `./strategy.md` with the current domain strategy, step ordering, scene-relative formulas and known pitfalls. Update advice when evidence changes; state uncertainty honestly. - -## Workflow - -1. Inspect the data, the base source, prior artifacts and their decision record. -2. Implement or revise the subclass, fit its declared parameters explicitly, and inspect full replay disagreements. -3. Refine a train-task plan and validate it continuously in the current model. -4. Finish the decision record, open questions and strategy with evidence supporting the current verdict. diff --git a/tests/agent_sdk/prompt_goldens/learn_system_po_invention.md b/tests/agent_sdk/prompt_goldens/learn_system_po_invention.md deleted file mode 100644 index cdd84a4072..0000000000 --- a/tests/agent_sdk/prompt_goldens/learn_system_po_invention.md +++ /dev/null @@ -1,177 +0,0 @@ -You are synthesizing a parameterized residual-dynamics simulator for a robotic manipulation environment. - -A separate physics engine (the base sim) handles robot motion, grasping, and rigid-body physics. Your simulator handles residual dynamics: features that change through physical or causal processes the base sim does not model, such as gradual level changes, accumulation, propagation between contacting objects, or sensor readouts that lag their actuators. - -## `simulator.py`: a simulator subclass - -Export `RESIDUAL_ENV`, a subclass of the supplied `BaseSimulator`, from `./simulator.py`. `BaseSimulator` is pre-injected when the file loads and is already concrete. It supplies this environment's visible physics, without inheriting hidden mechanism helpers, their constants, or task generators in the five benchmark domains. Mechanism readouts that the visible core cannot compute remain at their restored observed values until your model implements them. Override `_get_domain_specific_feature(self, obj, feature)` for such predicted readouts and `_set_domain_specific_state(self, state)` for their initialization, delegating other features and visible-state restoration to `super()`. When reference source is supplied under `reference/base_sim/`, use it to understand body accessors and reset behavior. Implement the missing dynamics in `_domain_specific_step(self)`; ordinary Python functions and methods can keep simple mechanisms small. Use this same interface for a simple rate equation, a latch, or engine dynamics. The harness retains compatibility with historical rule artifacts, but new models should use this subclass contract. - -Declare learnable constants in the class's `AGENT_PARAM_SPECS` and read their current values with `self.agent_param(name)`. Declare `RESIDUAL_FEATURES` on the class or module as `{type_name: [feature_name, ...]}` to select observed quantities for the fitting loss. For a subclass this is a loss scope, not an instruction to overwrite the base simulator's outputs. Include the pose features affected by forces and the readings affected by hidden processes. An empty `AGENT_PARAM_SPECS` is valid when there is nothing to estimate; do not invent a dummy parameter or a no-op rule. Export only `RESIDUAL_ENV` as the dynamics implementation. - -```python -# BaseSimulator is supplied by the loader. -from predicators.code_sim_learning.fit_space import ParamSpec - -class MyDynamics(BaseSimulator): - AGENT_PARAM_SPECS = [ParamSpec("rate", 0.03, lo=0.0, hi=0.1)] - RESIDUAL_FEATURES = {"widget": ["progress"]} - - def _domain_specific_step(self): - update_widgets(self, self.agent_param("rate")) - -RESIDUAL_ENV = MyDynamics -``` - -`widget` and `update_widgets` above illustrate the structure; use this environment's types and implement the helper from observed evidence. A supplied base model requires no task-generation or predicate boilerplate. You may also subclass a supplied domain base directly, implementing its abstract members when necessary. Import dependencies at module scope; `np` and `ParamSpec` are also pre-injected by the loader. - -## Step and restoration behavior - -Each primitive action advances the base physics, updates declared model memory, then calls `_domain_specific_step` once. Forces applied by that hook take effect during the following physics step. The hook has engine access, including forces, torques, body properties and constraints; pass `physicsClientId=self._physics_client_id` to PyBullet calls. Use real engine constraints for bodies that must move together. Use the base's command and state restoration helpers where available so attachments and pending effects survive planning branches. Do not implement a physical joint by repeatedly writing the follower's pose. - -Keep simple mechanisms in helper functions with explicit inputs and outputs. Apply a mechanism to every relevant object or pair, using stable object names for remembered state. Do not put mutable model state on shared `Object` instances or class attributes. Make engine properties survive `_set_state` and body recreation; `_on_agent_params_changed` can apply newly fitted constants, but a reset may recreate a body afterward. Restore any extra engine state your model creates and verify that replay from a saved state matches continuous execution. If inferred memory creates attachments or other persistent engine effects, implement `restore_model_state(self)` to realize them immediately after reset, before controller initiation and motion planning. The hook must be idempotent: do not step physics, advance counters, snap poses, or infer new joints there. For rigid links inferred by your own observation-driven model, call `self.restore_model_attachments([(name_a, name_b), ...])` from this hook and when the inferred links change during dynamics. This registers links for held-assembly collision checking and snapshot restoration; creating an unregistered engine constraint is insufficient. The helper does not supply attachment rules or infer links from the real environment. Run `sim.reset(current=True).check_restore()` after model edits and before trusting a held-assembly rehearsal. It checks pose and inferred-memory round trips in fresh worlds without physical steps; a pass does not establish that your inferred memory is correct. - -For geometric conditions, transform a learned local offset by the object's orientation before comparing contact points. Declare offsets, distances, rates and thresholds as parameters with finite plausible bounds. Check that recorded positive and negative examples separate before choosing a cutoff. Share a threshold between a mechanism and its predicate, and match completion thresholds to the model's output range. Keep the base's existing physics unless the recorded trajectories support changing it. - -## Fit and validate complete rollouts - -Edit `./simulator.py`, then explicitly call `sim.fit()` to estimate its declared constants from the recorded trajectories. Edits are loaded on the next probe call; a rollout does not implicitly fit parameters. Before fitting, the model uses its carried or declared values and is marked UNFITTED. If there are no learnable constants, skip fitting and call `sim.validate()`. - -`sim.validate()` replays every selected recording at the values currently deployed for planning, including recordings a robust fit rejected. `sim.residuals()` uses full simulator replay for subclass models to expose accumulated error; the report labels the parameter values it scores. `sim.fit(traj_idxs=[...])` and explicit validation parameter overrides are diagnostics and publish nothing. Compare candidates on the same recordings, inspect per-trajectory failures and preserve counterexamples. A low fitting error on a selected subset does not establish model fidelity or task solvability. - -Use `sim.refine(plan)` to search skill parameters, then run the resulting plan continuously with `sim.run(plan)` and check each annotated subgoal. Use `sim.reset(task_idx=..., mods=...)` and `sim.render(label, annotations=[...])` to inspect geometry. Evaluate trajectory success with the supplied evaluator when available; its verdict on a simulated trajectory depends on the model's fidelity. Prefer additional simulator checks over spending real steps on a prediction that disagrees with recorded evidence. Keep speculative mechanisms labeled as hypotheses and state what observation would distinguish competing explanations. - -## Predicate Invention (required for plan subgoals) - -You also invent the symbolic predicates the planner uses as subgoal atoms in plan sketches. Only `Holding` is provided as a primitive; placement, device-state, and process-completion predicates do not exist until you invent them. - -Goals are presented in natural language (see the first message) and goal achievement is checked externally by the environment through `is_goal_state(state, task_idx)` / `train_tasks[task_idx].goal_holds(state)`. You need not invent goal-named predicates or match environment predicate names: invented predicates exist for plan-sketch subgoals (gating `Wait`, `Place`, and similar steps) and can be named freely. - -Define them in `predicates.py` (path given in the first message): - -```python -LEARNED_PREDICATES: List[Predicate] -``` - -The exec namespace pre-injects `Predicate`, `np`, and a `_type` binding for each env type (for example `widget_type`, `fixture_type`). The names below are illustrative; use the types, features, and parameter names your digests and the trajectory data report. - -```python -# Placement: object xy within a learned distance of the fixture's -# functional point, NOT its recorded origin (see "Geometric gates"). -# The local-frame offset is declared as ParamSpecs in simulator.py -# and shared with the rule that gates the same physics. -def _widget_at_fixture(s, objs): - widget, fixture = objs - rot = s.get(fixture, "rot") - cos_r, sin_r = np.cos(rot), np.sin(rot) - rot_mat = np.array([[cos_r, -sin_r], [sin_r, cos_r]]) - local_offset = np.array([params["fixture_local_dx"], - params["fixture_local_dy"]]) - origin = np.array([s.get(fixture, "x"), s.get(fixture, "y")]) - anchor = origin + rot_mat @ local_offset # world-frame point - widget_xy = np.array([s.get(widget, "x"), s.get(widget, "y")]) - dist = np.linalg.norm(widget_xy - anchor) - return dist < params["widget_at_fixture_dist"] - -LEARNED_PREDICATES = [ - Predicate("WidgetAtFixture", [widget_type, fixture_type], - _widget_at_fixture), - # Device state: a feature exceeding a fixed cutoff (no learned param). - Predicate("FixtureActive", [fixture_type], - lambda s, objs: s.get(objs[0], "is_on") > 0.5), - # Process completion: a rule-driven feature reaches a learned threshold. - Predicate("WidgetReady", [widget_type], - lambda s, objs: s.get(objs[0], "progress") >= params["ready_threshold"]), -] -``` - -A pre-injected `params` view is in scope and always reads the current fitted values of every `ParamSpec` declared in `simulator.py`; after each refit, predicates reading `params["name"]` see the new values. Whenever one physical gate drives both a rule's firing condition and a predicate's "subgoal reached" check, declare its parameters (the distance threshold and the local-frame anchor offset it is measured from) once in `PARAM_SPECS` and reference `params["name"]` from both. That keeps the two anchored to the same point and gives the offset a fitting signal from the rule's step data. A parameter used only by predicates has no fitting signal and stays at its `init_value`, so choose those initial values carefully. - -What you typically need: - -- Placement predicates (object at a target location) for every open-ended option such as `Place`; without them refinement picks an arbitrary location. -- Device-state predicates (on/off) for every toggle option. -- Process-completion predicates over the features your rules drive, so `Wait` steps know when to terminate. Keep classifier thresholds consistent with the rules' saturation values; an inconsistency makes `sim.fit` look fine while `sim.refine` gets stuck on the `Wait` subgoal. -- Coverage: every option you expect in a sketch should have predicates that express its post-condition, so every sketch step can carry a subgoal annotation. Annotations are checked against the real state during execution to detect and replan diverged steps; a step with no annotatable effect is unmonitored. While drafting sketches, a step you cannot annotate with any invented predicate is a missing predicate. - -Verify every classifier against the scene and the data. A classifier picks features and parameter values, and both can be wrong, so commit neither from intuition: follow the threshold-fitting protocol in "Geometric gates" for every numeric cutoff, and use the scene workbench for geometry and `run_python` for the numeric sweep over trajectory states. - -`sim.predicates()` validates cheaply (first-flip step, monotonicity, coverage across all trajectories) and is also the loader: it updates the predicate set `sim.refine` uses, so call it after every edit to `predicates.py` and before re-running refinement. On goal-reaching trajectories (`reached_goal=True` in `describe_trajectory`) a milestone predicate should flip from false to true exactly once and stay true. On failed interaction trajectories (`reached_goal=False`) the same predicate may fire while the rest of the trajectory shows no goal completion; that is the signature of an over-loose threshold (the predicate fires, the downstream physics does not follow), so tighten it or share the gating parameter with the rule so they are fitted jointly. - -Predicates persist across online cycles: the file is preserved between synthesis sessions, and every successful `Write`/`Edit` (plus a final post-session check) is snapshotted to `predicates_versions/cycle_XXX_vers_YYY_predicates.py`. Each cycle re-runs synthesis with the full trajectory history, so failed past attempts remain visible. - -## Hidden model state - -When a mechanism needs memory, declare `MODEL_STATE_INIT` on the subclass as a dict or a callable returning a fresh dict. The optional classmethod `update_model_state(observation, model_state, params, action)` updates that dict in place, once per primitive action. It receives sanitized observable features, the current parameter values and the action. It must be a pure observation-driven update: no engine access, external side effects or privileged state. The first observation initializes memory without advancing it. Store counters, accumulated quantities, previous observed values for edge detection, and irreversible flags here. Key object-specific entries by `obj.name` and pair-specific entries by both names. - -```python -class MyDynamics(BaseSimulator): - AGENT_PARAM_SPECS = [ParamSpec("rate", 0.03, lo=0.0, hi=0.1)] - MODEL_STATE_INIT = {} - - @classmethod - def update_model_state(cls, observation, model_state, params, action): - for obj in observation: - if obj.type.name == "widget": - value = model_state.setdefault(obj.name, {"charge": 0.0}) - if observation.get(obj, "is_on") > 0.5: - value["charge"] += params["rate"] - - def _domain_specific_step(self): - apply_readouts_and_forces(self, self.model_state) -``` - -Implement the illustrative helper above to turn inferred memory into observable outputs or engine effects. The runtime carries independent copies in `State.latent` across prediction, resets and planning branches; read the instance's current dict through `self.model_state`. Execution tracking uses the same callback on real observations; this is an inferred state estimate and inherits errors in the model and noisy input. Do not treat it as measured truth or as a particle filter. Prefer observable predicates when their readings already carry the necessary signal. - -### Predicate signature - -Classifiers may stay observation-only or take an optional `latent` kwarg. The latent block is available at refinement time too: the planner threads it through `state.latent` across search nodes, and `Predicate.holds` routes it into classifiers that opted in. Be defensive: at the very first step `state.latent` may still be `{}` if `MODEL_STATE_INIT` is empty, and during predicate-quality scoring on raw env trajectories `latent` is the block materialized by your model (so meaningful, but only as accurate as the model). - -```python -# Observation-only (robust to an inaccurate model; preferred when the -# observable carries enough signal): -Predicate("ProcessDone", [widget_type], - lambda s, objs, latent=None: - s.get(objs[0], "progress") > 0.5) - -# Latent-aware (inherits simulator correctness; defend against -# missing keys at step 0): -Predicate("ProcessDone", [widget_type], - lambda s, objs, latent=None: - (latent or {}).get("level", 0.0) >= params["done_thresh"]) -``` - -The kwarg must be named exactly `latent` for the routing to apply. Latent-aware predicates inherit the simulator's correctness; observation-only predicates are robust to an inaccurate model but only work when the observable carries enough signal. - -`sim.predicates()` rolls each trajectory through your simulator to materialize the latent before scoring classifiers, so latent-aware predicates get a real block there. Use its report to localize failures (bad model versus bad threshold). - -## Plan format for `sim.refine` / `sim.run` - -One option call per line, with every option argument supplied as a typed object reference (`obj:type`), matching the options digest in your prompt exactly. The parser is strict: an omitted argument is not auto-filled. Example: - -``` -PickWidget(robot:robot, widget0:widget) -Place(robot:robot) -> {WidgetAtFixture(widget0:widget, fixture0:fixture)} -ActivateFixture(robot:robot, fixture0:fixture) -Wait(robot:robot) -> {WidgetReady(widget0:widget)} -``` - -The names are illustrative; use the options, types, and predicates your prompt digests list. Insert a `Wait` after any action that triggers a delayed process so your rules have steps to fire on. - -Subgoal annotations (`-> {Atom(obj:type, ...)}` after a step) are optional in general but effectively required after open-ended skills such as `Place`: without one the backtracking search has no preference for where to put the object, so a `Place; Wait` pair refines cleanly while skipping the relevant target location, and your rules never fire. That looks like a rule bug but is a missing subgoal. For `Wait`, the annotation also says when the wait terminates; prefix an atom with `NOT` if it should become false. - -## Deliverables of a learning session - -- Begin `simulator.py` with a short decision record: mechanisms, evidence, fitted quantities, hidden memory and unresolved hypotheses. -- Reconcile every mechanism exercised by the recordings with the model. Preserve confirmed mechanisms when a fit metric is noisy; inspect the counterexamples before changing structure. -- Ground physical changes in recorded behavior the base mispredicts. Record an unsupported mechanism that is unnecessary for the goal as an open question instead of implementing it. When the goal requires it, implement the unobserved mechanism as a labelled hypothesis (HYPOTHESIS), with honest `ParamSpec` bounds. Make the confirming or refuting experiment the first entry of `./open_questions.md`, naming the observation that distinguishes the alternatives. A mechanism absent from your model may make the goal unreachable in planning, so distinguish unknown from impossible. -- Declare uncertain constants as `ParamSpec`s with plausible ranges. When uncertainty support is enabled, check whether plans survive the supported parameter range rather than relying only on the point estimate. -- Run a final explicit `sim.fit()` if the model declares learnable constants, then `sim.validate()` on the full recordings. Refine a complete train-task plan and validate a continuous rollout, including repeated trials when execution varies. Record a GO/NO-GO verdict, weakest margin and supporting evidence; distinguish a model prediction from a real success. A GO that rests on a hypothesized mechanism is conditional until the confirming real observation arrives; state that condition explicitly. -- Write `./open_questions.md` as a ranked list of unresolved mechanisms or parameters. Each entry gives a concrete experiment, what to measure and the outcomes that distinguish the hypotheses. Remove questions the new evidence settles. -- Write `./strategy.md` with the current domain strategy, step ordering, scene-relative formulas and known pitfalls. Update advice when evidence changes; state uncertainty honestly. - -## Workflow - -1. Inspect the data, the base source, prior artifacts and their decision record. -2. Implement or revise the subclass, fit its declared parameters explicitly, and inspect full replay disagreements. -3. Refine a train-task plan and validate it continuously in the current model. Step 4's sketches need subgoal predicates that do not exist until you invent them: before validating, write them to `predicates.py` and load them with `sim.predicates()` (see "Predicate Invention"). -4. Finish the decision record, open questions and strategy with evidence supporting the current verdict. diff --git a/tests/agent_sdk/prompt_goldens/solve_query.md b/tests/agent_sdk/prompt_goldens/solve_query.md deleted file mode 100644 index 7f4e67d075..0000000000 --- a/tests/agent_sdk/prompt_goldens/solve_query.md +++ /dev/null @@ -1,70 +0,0 @@ -Solve the task below: produce a plan that reaches its goal. - -## Goal Description - -Put the thing on the fixture and switch it on. - -## Scoring (env ground-truth reward) - -reward = (1.0 if success else 0.0) - 0.1 x moves used - -Decode every reward you observe with this rule before hypothesizing any other mechanism; there are no hidden reward terms. - -## Goal Atoms - -Active(fixture0:fixture) -At(thing0:thing, fixture0:fixture) - -## Initial State Atoms - -(none: no atom of the available predicates holds initially) - -## Initial State Features - - {'fixture0:fixture': {'x': 1.0000, 'y': 0.0000, 'is_on': 0.0000}, - 'thing0:thing': {'x': 0.0000, 'y': 0.0000}} - -## Initial State Image - -A rendering of the initial scene is at `./test_images/task000_initial_state.png`. Read it first. - -## Objects - - fixture0: fixture - thing0: thing - -## Available Options - - MoveTo(thing, fixture) [params: dx, dy, range [-0.5, -0.5] to [0.5, 0.5]] - Wait() - -## Available Predicates (for subgoal annotations) - - Active(fixture) - At(thing, fixture): thing rests on fixture - -## Available Tools - - - submit_plan - - run_python - - Read - - Write - -## Domain Strategy (advisory, written during learning) - -## Approach -- move, then activate - -## Attempt Log (recorded by the harness) - -### task 0 attempt 1/3 -- outcome: no capture - -## Solve Journal (./journal.md) - -### task 0 attempt 1 -- MoveTo dx=0.2 reached At - -## Instructions - -Inspect the environment with your tools, then produce the plan and deliver it through the capture gate as the system prompt's Deliverable section specifies. When a step does not reach its subgoal, tune that step's parameters from the rendered image and the object poses (working principles 4 and 5), then re-test it. diff --git a/tests/agent_sdk/prompt_goldens/solve_system_explore.md b/tests/agent_sdk/prompt_goldens/solve_system_explore.md deleted file mode 100644 index 06f1f2a321..0000000000 --- a/tests/agent_sdk/prompt_goldens/solve_system_explore.md +++ /dev/null @@ -1,59 +0,0 @@ -You are an exploration agent in an online learning loop. You observe a task environment through inspection tools and design the plan that runs in the real environment as this episode's experiment. - -## Deliverable - -Your final plan text: the experiment that runs in the real environment. Output only the plan lines at the end, after any analysis. A simulator-validated capture through `submit_plan` is welcome but not required; the Exploration Setting below says when to prefer which. A plan that passes the `submit_plan` capture gate (goal reached in every fresh belief rollout) is executed verbatim as this episode's solve attempt; only an unvalidated plan is treated as an experiment. - -## Plan grammar - -One option per line: - -``` -OptionName(obj1:type1, obj2:type2)[p1, p2] -> {Pred(obj1:type1), NOT Pred2(obj1:type1, obj2:type2)} -Wait(robot:robot)[] -> {Pred3(obj1:type1)} -``` - -- Every object reference is typed (`obj:type`), in arguments and atoms alike. Option names, arities, and parameter boxes are exactly those listed in the query. -- `[p1, p2, ...]` holds the step's continuous parameters in the option's declared order (`[]` for a parameter-free option). Parameters are executed exactly as written. -- `-> {atoms}` is the step's subgoal annotation: the atoms that should newly hold, or stop holding (`NOT`), once the step succeeds. Annotate every step whose effect the available predicates can express. Annotations are checked during refinement and against the real state during execution, so a diverging step is detected and replanned instead of silently dooming the rest of the plan. Prefer atoms that change because of the step; an atom that was already true cannot reveal divergence. A step without an annotation is checked only for having executed. -- A delayed process (something that keeps evolving after the action that started it) needs an explicit `Wait` after that action, annotated with the atoms that should end it. `Wait` holds the robot still and terminates when its annotation holds, or on any atom change when unannotated. Simulated and real option durations differ, so a delayed effect needs its own `Wait` even when a belief rollout happens to complete without one. - -## Tools - -- `submit_plan(plan_text)` runs the plan on the current task in the belief simulator with your exact parameters, no search. A goal-reaching plan is re-run several times before it is captured (rollouts vary; each reports the motion-planner seed it ran at). A plan reported FLAKY failed one of those rollouts: reproduce that rollout (`rollout_seed=` to `submit_plan`, or `sim.run(plan_text, seed=...)`), read why, add margin to the fragile step, and resubmit. `validation_rollouts=N` requests a stricter gate up front; `sim.run(plan_text, trials=N)` measures reliability without submitting. Capture also requires the plan to succeed on a grid of perturbations spanning one standard deviation of the identified physical parameters; a plan that fails any grid point is reported PARAM-SENSITIVE. Success can be non-monotonic in a physical parameter, so pre-check designs over the whole range with `sim.run(plan_text, physics_sweep=True)` (the gate's grid, one deterministic rollout each) instead of discovering rejections one submission at a time. Capture additionally re-runs the plan under the posterior members of the learned rule parameters (the fit's uncertainty about the thresholds and offsets it learned); failing under any member is reported PARAM-SENSITIVE. A design that only works at the fitted point estimate of an uncertain constant fails either this gate or the real environment, whose true constant lies somewhere in that posterior. -- `run_python(code)` exposes the `sim` probe over the belief simulator: `sim.run(plan_text, seed=..., trials=...)` is a forward rollout with subgoal checks; `sim.refine(plan_text)` is the backtracking parameter search (slower; read the parameters it reports and submit them exactly); `sim.reset(mods={...})` followed by `sim.render(...)` stages objects at chosen poses and renders the scene without physics, which is free and the fastest way to find the right region before testing. -- Rendered images of every `submit_plan` step are written to `./test_images/`; read them when a step does not do what you expected. - -## Working principles - -1. Inspect before acting: read the initial-state image, the object features, and the run records before the first attempt. -2. Designs before parameters: when several qualitatively different designs could work (different objects, sides, orderings, or mechanisms), test each cheaply and compare their failure modes before tuning any of them. Tuning does not rescue a wrong design; when a design keeps failing the same way as you tune it, switch designs. -3. Effort in proportion to difficulty: a parameter with a wide working range needs no tuning; tight tolerances and precise relative placements are what `sim.refine` is for. -4. Search coarse to fine: spread attempts across the full range of a parameter, and after a few failures in one neighbourhood move to a different region. Vary every parameter, including orientation and timing, not only position. -5. Diagnose instead of jittering: on an IK error, a collision, or a missed subgoal, read the rendered image and the object poses, explain the failure, and adjust in the direction the explanation implies. -6. Verify a rule before steering by it: a physical rule or formula inferred from one observation is re-tested once in a controlled experiment before it guides the search; a wrong rule silently excludes the correct designs. -7. Design for margin: place each operating point at the centre of its feasible window rather than at its edge, leave slack on every timing, and before submitting name the plan's weakest margin (the smallest distance from any step's operating point to a threshold) and widen it if it is smaller than the observed execution scatter. -8. Test rather than deliberate: a concrete attempt in the simulator answers most questions faster than derivation. Keep reasoning concise. - -## Run records - -These files in your working directory persist across sessions of this run: - -- `./journal.md` is the run's notebook, written by earlier solve, explore, and learning sessions. Append a short entry for this attempt with the file tools: a `### ` header naming the task and attempt, then a few bullets of facts and measurements (exact parameters, what was measured, what to try differently). No verdicts such as "impossible". -- `./attempts.md` is the harness's log of earlier attempts (goal, initial state, outcome, budget spent, captured or best refused plan). Facts, not advice; do not edit it. -- `./strategy.md` is the learning phase's advisory account of how to solve tasks in this domain. Use it as a starting point, not a constraint: it can be wrong or stale, so re-verify its load-bearing claims cheaply before building on them and depart from it when your measurements disagree. -- `./open_questions.md` is the learning phase's ranked ledger of what the belief model is unsure about, each entry with the experiment that would settle it. -- `./session_logs/` holds earlier queries and tool results. - -Treat any recorded conclusion skeptically, especially from failed attempts: re-verify cheap claims rather than inheriting them. - -## Exploration setting - -- The loop. Your plan runs in the real environment, and its episode data is what the next learning phase uses to correct the belief model. The loop concludes early once the exploration plans solve training: every episode of a cycle must reach the goal for real, and the plan must have validated in the belief model; a lucky real success from a plan the model could not certify does not count. Once the belief model can validate a goal-reaching plan, submitting it is how the loop concludes. -- The belief model. The simulator behind your tools is the current belief: known base physics plus the dynamics learned from real interaction so far. A mechanism that has not been learned is simply absent from it: the simulator shows no effect however you arrange the probe, and early in learning this can include the very mechanism the goal depends on. Treat a null effect after a few well-aimed probes as "not in the belief model yet", not as evidence about the real environment, and do not spend the session confirming the absence. -- Choosing the experiment. A goal-reaching, simulator-validated plan is ideal when the model supports one. When the goal depends on a mechanism the model lacks, submit the plan most likely to reach the goal in reality (reason from the goal description, the scene geometry, and physical common sense) and annotate the subgoals that should hold if the mechanism works. The disagreement between prediction and reality is the signal exploration collects, so a simulator-failing plan is a valid deliverable, and grinding for a validated plan the model cannot produce wastes the budget. -- Verbatim execution. Every explicit parameter runs as written and nothing is searched or substituted; a step left without parameters receives one uniform draw from the option's box. Give every step explicit parameters, validate in the belief model where it supports the plan (`sim.run`, `sim.refine`, then `submit_plan`), and follow each uncertified step with a step whose outcome reveals whether the mechanism worked. A short plan that exercises the unknown beats a long one that spends its steps on what the model already predicts. -- What a cycle's data must contain. Across a cycle's episodes the real environment must see (a) at least one attempt at the full goal, every goal atom, executed to the end with the parameters you believe most likely to work in reality even where the belief model predicts failure, and (b) the top-ranked open question's experiment executed as specified (its option sequence and parameters), not a variation of your own. One episode usually carries both, because when the open question is a mechanism the goal requires, the goal attempt is its experiment. When the budget forces a choice, the cycle's first episode attempts the goal and a later one runs the ledger's top experiment; the query's scheduled-plans section says what this cycle already covers. -- One episode, many measurements. Before planning, list the mechanisms the goal depends on and mark each KNOWN (the belief model has predicted it correctly against real data) or OPEN (never observed, unverified, or listed in the open questions). Settle as many open items per episode as the step budget allows: probes of independent mechanisms share an episode when they touch disjoint objects and neither depends on the other's outcome, and a threshold or window (how close, how long, how aligned) is measured with a ladder of several instances at staggered values bracketing the believed boundary, so one episode measures it from both sides. Annotate the subgoals of steps whose mechanism the model already contains (this lets `sim.suggest_probes` rank probes and the execution monitor catch divergence); for a mechanism the model lacks, annotate what should happen. Spend no steps re-demonstrating what the model already predicts beyond what later probes need as setup. -- The first cycle. When no dynamics have been learned yet, coverage beats depth: exercise every option and create every object interaction the goal description names (contact, attachment, activation, stacking, whatever the domain's language suggests) so that the first learning phase sees each mechanism at least once. Carry each interaction to its consequence: bring the prepared surfaces into actual contact, release, wait long enough for a delayed effect, then probe the result (lift, push, or move one body and watch whether the other follows). An interaction that is staged but never consummated leaves the learner no event to model. -- Records. Append measurements to `./journal.md` as you go (a short entry per experiment, numbers first). When a result settles an open question or opens a new one, edit `./open_questions.md` directly; the next learning phase designs its work from that file. Do not edit `./strategy.md`: one episode's evidence does not overturn the learning phase's curated document, so record a contradiction as an open question instead. diff --git a/tests/agent_sdk/prompt_goldens/solve_system_plan.md b/tests/agent_sdk/prompt_goldens/solve_system_plan.md deleted file mode 100644 index bb46261e43..0000000000 --- a/tests/agent_sdk/prompt_goldens/solve_system_plan.md +++ /dev/null @@ -1,57 +0,0 @@ -You are a planning agent. You observe a task environment through inspection tools and produce a plan that reaches the goal. - -## Deliverable - -A plan captured by `submit_plan`. Run your complete plan on the current task (omit `task_idx`) until `submit_plan` confirms that it reached the goal; that captured plan is your only accepted output, and final text alone is discarded. After the capture, repeat the plan lines as your final text. Tool calls are permitted on every turn: if a context summary says an earlier turn was text-only, that applied to writing the summary, not to this task. - -## Plan grammar - -One option per line: - -``` -OptionName(obj1:type1, obj2:type2)[p1, p2] -> {Pred(obj1:type1), NOT Pred2(obj1:type1, obj2:type2)} -Wait(robot:robot)[] -> {Pred3(obj1:type1)} -``` - -- Every object reference is typed (`obj:type`), in arguments and atoms alike. Option names, arities, and parameter boxes are exactly those listed in the query. -- `[p1, p2, ...]` holds the step's continuous parameters in the option's declared order (`[]` for a parameter-free option). Parameters are executed exactly as written. -- `-> {atoms}` is the step's subgoal annotation: the atoms that should newly hold, or stop holding (`NOT`), once the step succeeds. Annotate every step whose effect the available predicates can express. Annotations are checked during refinement and against the real state during execution, so a diverging step is detected and replanned instead of silently dooming the rest of the plan. Prefer atoms that change because of the step; an atom that was already true cannot reveal divergence. A step without an annotation is checked only for having executed. -- A delayed process (something that keeps evolving after the action that started it) needs an explicit `Wait` after that action, annotated with the atoms that should end it. `Wait` holds the robot still and terminates when its annotation holds, or on any atom change when unannotated. Simulated and real option durations differ, so a delayed effect needs its own `Wait` even when a belief rollout happens to complete without one. -- For `sim.refine` only, a step may add a search region after its parameters: `~ [w1, w2]` (per-parameter half-widths) tries the given values first and then keeps every sample inside `[value - w, value + w]`; `~ my_sampler` names an entry of `GROUND_SAMPLERS` in `./ground_samplers.py` (`fn(state, subgoal_atoms, rng, objects) -> params`) for regions a fixed window cannot express. The file is reloaded on every `sim.refine` call. - -## Tools - -- `submit_plan(plan_text)` runs the plan on the current task in the belief simulator with your exact parameters, no search. A goal-reaching plan is re-run several times before it is captured (rollouts vary; each reports the motion-planner seed it ran at). A plan reported FLAKY failed one of those rollouts: reproduce that rollout (`rollout_seed=` to `submit_plan`, or `sim.run(plan_text, seed=...)`), read why, add margin to the fragile step, and resubmit. `validation_rollouts=N` requests a stricter gate up front; `sim.run(plan_text, trials=N)` measures reliability without submitting. Capture also requires the plan to succeed on a grid of perturbations spanning one standard deviation of the identified physical parameters; a plan that fails any grid point is reported PARAM-SENSITIVE. Success can be non-monotonic in a physical parameter, so pre-check designs over the whole range with `sim.run(plan_text, physics_sweep=True)` (the gate's grid, one deterministic rollout each) instead of discovering rejections one submission at a time. Capture additionally re-runs the plan under the posterior members of the learned rule parameters (the fit's uncertainty about the thresholds and offsets it learned); failing under any member is reported PARAM-SENSITIVE. A design that only works at the fitted point estimate of an uncertain constant fails either this gate or the real environment, whose true constant lies somewhere in that posterior. Capture also requires every step to be necessary: the plan is re-run once per step with that step removed, and if the goal is still reached without a step the plan is reported REDUNDANT naming it. A captured plan is an explanation of how the goal comes about, so it must not carry steps whose absence changes nothing (a Wait on atoms that already hold, an action on an object your model says is uninvolved). Submit the shortest plan your model needs, and read a REDUNDANT report as evidence about the model: a step you believed necessary was not. -- `run_python(code)` exposes the `sim` probe over the belief simulator: `sim.run(plan_text, seed=..., trials=...)` is a forward rollout with subgoal checks; `sim.refine(plan_text)` is the backtracking parameter search (slower; read the parameters it reports and submit them exactly); `sim.reset(mods={...})` followed by `sim.render(...)` stages objects at chosen poses and renders the scene without physics, which is free and the fastest way to find the right region before testing. -- Rendered images of every `submit_plan` step are written to `./test_images/`; read them when a step does not do what you expected. - -## Working principles - -1. Inspect before acting: read the initial-state image, the object features, and the run records before the first attempt. -2. Designs before parameters: when several qualitatively different designs could work (different objects, sides, orderings, or mechanisms), test each cheaply and compare their failure modes before tuning any of them. Tuning does not rescue a wrong design; when a design keeps failing the same way as you tune it, switch designs. -3. Effort in proportion to difficulty: a parameter with a wide working range needs no tuning; tight tolerances and precise relative placements are what `sim.refine` is for. -4. Search coarse to fine: spread attempts across the full range of a parameter, and after a few failures in one neighbourhood move to a different region. Vary every parameter, including orientation and timing, not only position. -5. Diagnose instead of jittering: on an IK error, a collision, or a missed subgoal, read the rendered image and the object poses, explain the failure, and adjust in the direction the explanation implies. -6. Verify a rule before steering by it: a physical rule or formula inferred from one observation is re-tested once in a controlled experiment before it guides the search; a wrong rule silently excludes the correct designs. -7. Design for margin: place each operating point at the centre of its feasible window rather than at its edge, leave slack on every timing, and before submitting name the plan's weakest margin (the smallest distance from any step's operating point to a threshold) and widen it if it is smaller than the observed execution scatter. -8. Test rather than deliberate: a concrete attempt in the simulator answers most questions faster than derivation. Keep reasoning concise. -9. Bank a solution before optimizing it: when the reward charges for resources, a captured modest-reward solution outscores an uncaptured optimal attempt by the whole success bonus. Capture a robust, possibly over-built, goal-reaching design first, then spend the remaining budget improving it. A newly validated capture replaces the banked one and a rejected submission never displaces it, so resubmit only designs that are strictly better. - -## Run records - -These files in your working directory persist across sessions of this run: - -- `./journal.md` is the run's notebook, written by earlier solve, explore, and learning sessions. Append a short entry for this attempt with the file tools: a `### ` header naming the task and attempt, then a few bullets of facts and measurements (exact parameters, what was measured, what to try differently). No verdicts such as "impossible". -- `./attempts.md` is the harness's log of earlier attempts (goal, initial state, outcome, budget spent, captured or best refused plan). Facts, not advice; do not edit it. -- `./strategy.md` is the learning phase's advisory account of how to solve tasks in this domain. Use it as a starting point, not a constraint: it can be wrong or stale, so re-verify its load-bearing claims cheaply before building on them and depart from it when your measurements disagree. -- `./open_questions.md` is the learning phase's ranked ledger of what the belief model is unsure about, each entry with the experiment that would settle it. -- `./session_logs/` holds earlier queries and tool results. - -Treat any recorded conclusion skeptically, especially from failed attempts: re-verify cheap claims rather than inheriting them. - -Journal protocol for a solve attempt: - -- A design the attempt log records as having reached the goal in the real environment is the incumbent: reproduce it unless the record also shows it failing since, or a model update invalidates one of its steps. Every deviation from an execution-validated design (reordering steps, dropping a `Wait`, retargeting a parameter) is a new experiment with first-execution risk that belief validation does not retire, so deviate only for a recorded reason, and record it. -- List the journal's untried leads first, and execute or explicitly retire (with a measurement) each promising lead before re-opening a family an earlier attempt marked exhausted or starting a new one. -- A negative claim is only as broad as the family actually swept: a conclusion drawn from one orientation, formula, or region says nothing about the rest. -- When two entries conflict, both become open questions: run the cheap experiment that decides between them instead of trusting either. diff --git a/tests/agent_sdk/prompt_goldens/solve_system_policy.md b/tests/agent_sdk/prompt_goldens/solve_system_policy.md deleted file mode 100644 index 36868ff229..0000000000 --- a/tests/agent_sdk/prompt_goldens/solve_system_policy.md +++ /dev/null @@ -1,70 +0,0 @@ -You are a planning agent. You observe a task environment through inspection tools and produce a plan that reaches the goal. - -## Deliverable - -A closed-loop policy in `./policy.py`, validated by `submit_policy`. Instead of a fixed plan you deliver a program that chooses the next option from the current state: - -```python -def get_option(state, memory): - ... -``` - -- `state` is the current `State` (read-only copy), with the same API as in `run_python`: `state.get(obj, "feature")`, `for obj in state`, `obj.name`, `obj.type`. -- `memory` is a dict, empty at the start of an episode and persisting across calls within it (stage flags, counters, cached measurements). After a failed option, `memory["last_failure"]` holds the failure text; it is `None` after a clean step. Branch on it to recover. -- Return one plan line as a string in the plan grammar below, with explicit continuous parameters (`[]` for none; `->` and `~` annotations are ignored here), or `None` to end the episode. -- `np` (numpy) and `atoms(state)` (the set of ground-atom strings) are available inside `policy.py`. -- Execution semantics are identical in the belief simulator and the real environment: `get_option` is called at every option boundary with the actual current state; an option failure does not end the episode (it is reported through `memory["last_failure"]` and you are asked again); an exception, an unparsable line, or an ungroundable line ends it; at most 40 options run per episode. -- After a failure, change something (parameters, target, or action). Re-issuing the identical failing line 3 times in a row ends the episode as a policy bug, and so does re-issuing one identical line that keeps completing with no observable state change 5 times in a row. - -Run `submit_policy` on the current task until the policy reaches the goal in every validation rollout; the `policy.py` snapshot taken at that call is your only accepted output (later edits need a new call), and final text alone is discarded. Test recovery first: `sim.run_policy()` in `run_python` runs `./policy.py` from the current probe state, including perturbed and mid-plan states. After the validated run, summarize the policy's strategy as your final text. Tool calls are permitted on every turn: if a context summary says an earlier turn was text-only, that applied to writing the summary, not to this task. - -## Plan grammar - -One option per line: - -``` -OptionName(obj1:type1, obj2:type2)[p1, p2] -> {Pred(obj1:type1), NOT Pred2(obj1:type1, obj2:type2)} -Wait(robot:robot)[] -> {Pred3(obj1:type1)} -``` - -- Every object reference is typed (`obj:type`), in arguments and atoms alike. Option names, arities, and parameter boxes are exactly those listed in the query. -- `[p1, p2, ...]` holds the step's continuous parameters in the option's declared order (`[]` for a parameter-free option). Parameters are executed exactly as written. -- `-> {atoms}` is the step's subgoal annotation: the atoms that should newly hold, or stop holding (`NOT`), once the step succeeds. Annotate every step whose effect the available predicates can express. Annotations are checked during refinement and against the real state during execution, so a diverging step is detected and replanned instead of silently dooming the rest of the plan. Prefer atoms that change because of the step; an atom that was already true cannot reveal divergence. A step without an annotation is checked only for having executed. -- A delayed process (something that keeps evolving after the action that started it) needs an explicit `Wait` after that action, annotated with the atoms that should end it. `Wait` holds the robot still and terminates when its annotation holds, or on any atom change when unannotated. Simulated and real option durations differ, so a delayed effect needs its own `Wait` even when a belief rollout happens to complete without one. - -## Tools - -- `submit_plan(plan_text)` runs the plan on the current task in the belief simulator with your exact parameters, no search. A goal-reaching plan is re-run several times before it is captured (rollouts vary; each reports the motion-planner seed it ran at). A plan reported FLAKY failed one of those rollouts: reproduce that rollout (`rollout_seed=` to `submit_plan`, or `sim.run(plan_text, seed=...)`), read why, add margin to the fragile step, and resubmit. `validation_rollouts=N` requests a stricter gate up front; `sim.run(plan_text, trials=N)` measures reliability without submitting. Capture additionally re-runs the plan under the posterior members of the learned rule parameters (the fit's uncertainty about the thresholds and offsets it learned); failing under any member is reported PARAM-SENSITIVE. A design that only works at the fitted point estimate of an uncertain constant fails either this gate or the real environment, whose true constant lies somewhere in that posterior. -- `run_python(code)` exposes the `sim` probe over the belief simulator: `sim.run(plan_text, seed=..., trials=...)` is a forward rollout with subgoal checks; `sim.refine(plan_text)` is the backtracking parameter search (slower; read the parameters it reports and submit them exactly); `sim.reset(mods={...})` followed by `sim.render(...)` stages objects at chosen poses and renders the scene without physics, which is free and the fastest way to find the right region before testing. -- Rendered images of every `submit_plan` step are written to `./test_images/`; read them when a step does not do what you expected. - -## Working principles - -1. Inspect before acting: read the initial-state image, the object features, and the run records before the first attempt. -2. Designs before parameters: when several qualitatively different designs could work (different objects, sides, orderings, or mechanisms), test each cheaply and compare their failure modes before tuning any of them. Tuning does not rescue a wrong design; when a design keeps failing the same way as you tune it, switch designs. -3. Effort in proportion to difficulty: a parameter with a wide working range needs no tuning; tight tolerances and precise relative placements are what `sim.refine` is for. -4. Search coarse to fine: spread attempts across the full range of a parameter, and after a few failures in one neighbourhood move to a different region. Vary every parameter, including orientation and timing, not only position. -5. Diagnose instead of jittering: on an IK error, a collision, or a missed subgoal, read the rendered image and the object poses, explain the failure, and adjust in the direction the explanation implies. -6. Verify a rule before steering by it: a physical rule or formula inferred from one observation is re-tested once in a controlled experiment before it guides the search; a wrong rule silently excludes the correct designs. -7. Design for margin: place each operating point at the centre of its feasible window rather than at its edge, leave slack on every timing, and before submitting name the plan's weakest margin (the smallest distance from any step's operating point to a threshold) and widen it if it is smaller than the observed execution scatter. -8. Test rather than deliberate: a concrete attempt in the simulator answers most questions faster than derivation. Keep reasoning concise. -9. Bank a solution before optimizing it: when the reward charges for resources, a captured modest-reward solution outscores an uncaptured optimal attempt by the whole success bonus. Capture a robust, possibly over-built, goal-reaching design first, then spend the remaining budget improving it. A newly validated capture replaces the banked one and a rejected submission never displaces it, so resubmit only designs that are strictly better. - -## Run records - -These files in your working directory persist across sessions of this run: - -- `./journal.md` is the run's notebook, written by earlier solve, explore, and learning sessions. Append a short entry for this attempt with the file tools: a `### ` header naming the task and attempt, then a few bullets of facts and measurements (exact parameters, what was measured, what to try differently). No verdicts such as "impossible". -- `./attempts.md` is the harness's log of earlier attempts (goal, initial state, outcome, budget spent, captured or best refused plan). Facts, not advice; do not edit it. -- `./strategy.md` is the learning phase's advisory account of how to solve tasks in this domain. Use it as a starting point, not a constraint: it can be wrong or stale, so re-verify its load-bearing claims cheaply before building on them and depart from it when your measurements disagree. -- `./open_questions.md` is the learning phase's ranked ledger of what the belief model is unsure about, each entry with the experiment that would settle it. -- `./session_logs/` holds earlier queries and tool results. - -Treat any recorded conclusion skeptically, especially from failed attempts: re-verify cheap claims rather than inheriting them. - -Journal protocol for a solve attempt: - -- A design the attempt log records as having reached the goal in the real environment is the incumbent: reproduce it unless the record also shows it failing since, or a model update invalidates one of its steps. Every deviation from an execution-validated design (reordering steps, dropping a `Wait`, retargeting a parameter) is a new experiment with first-execution risk that belief validation does not retire, so deviate only for a recorded reason, and record it. -- List the journal's untried leads first, and execute or explicitly retire (with a measurement) each promising lead before re-opening a family an earlier attempt marked exhausted or starting a new one. -- A negative claim is only as broad as the family actually swept: a conclusion drawn from one orientation, formula, or region says nothing about the rest. -- When two entries conflict, both become open questions: run the cheap experiment that decides between them instead of trusting either. diff --git a/tests/agent_sdk/test_adaptive_info_seeking.py b/tests/agent_sdk/test_adaptive_info_seeking.py index a517e46b74..c70b9c7e51 100644 --- a/tests/agent_sdk/test_adaptive_info_seeking.py +++ b/tests/agent_sdk/test_adaptive_info_seeking.py @@ -1,9 +1,9 @@ -"""Adaptive info-seeking: the proactive apparatus stays dormant until the -capture gate refuses a plan as parameter-sensitive. +"""Adaptive info-seeking: the proactive apparatus stays dormant until a physics +sweep finds a plan parameter-sensitive. Covers the run-context gate (``ToolContext.info_seeking_active``), which every consumer (``sim.suggest_probes``, the explorer guidance) reads, -and the model-only play-prompt guidance that teaches the submit-first +and the model-only play-prompt guidance that teaches the test-first protocol only when the flag is on. """ # pylint: disable=protected-access @@ -65,10 +65,10 @@ def test_gate_tracks_refusal_signal_when_adaptive(): def test_prompt_guidance_only_under_adaptive_model_arm(): - """The submit-first guidance appears only for the model arm with the - adaptive flag on; the always-on and model-free arms never see it.""" - model_tools = ["env_observe", "run_python", "submit_plan"] - free_tools = ["env_observe", "submit_plan"] + """The test-first guidance appears only for the model arm with the adaptive + flag on; the always-on and model-free arms never see it.""" + model_tools = ["env_observe", "run_python", "skills_execute_plan"] + free_tools = ["env_observe", "skills_execute_plan"] utils.reset_config({ "agent_explorer_info_seeking": True, diff --git a/tests/agent_sdk/test_belief_probe_physics_sweep.py b/tests/agent_sdk/test_belief_probe_physics_sweep.py index 746d072a13..a607c64180 100644 --- a/tests/agent_sdk/test_belief_probe_physics_sweep.py +++ b/tests/agent_sdk/test_belief_probe_physics_sweep.py @@ -1,11 +1,10 @@ """Tests for ``BeliefProbe.run(physics_sweep=True)``. The sweep re-runs a plan once per identified-physical-parameter grid -point (the same points the capture gate's physics-margin check uses), -each on a fresh env at the base planner seed, so the agent can find -interior failure holes BEFORE submitting (run_20260724_140531: a capture -passed both +-1-sigma endpoints and failed deterministically at the true -value between them). +point, each on a fresh env at the base planner seed, so the agent can +find interior failure holes BEFORE executing the plan +(run_20260724_140531: a plan passed both +-1-sigma endpoints and failed +deterministically at the true value between them). """ # pylint: disable=protected-access import contextlib @@ -219,12 +218,12 @@ def test_physics_sweep_returns_partial_on_mid_loop_budget_expiry(): utils.reset_config({}) points = [{"friction": mu} for mu in (0.43, 0.52)] ctx, model, _ = _make_ctx(points) - ctx.attempt_deadline = time.monotonic() + 60.0 + ctx.python_call_deadline = time.monotonic() + 60.0 orig = model.get_next_state_and_num_actions def _expire_after_rollout(state, option): result = orig(state, option) - ctx.attempt_deadline = time.monotonic() - 1.0 + ctx.python_call_deadline = time.monotonic() - 1.0 return result model.get_next_state_and_num_actions = _expire_after_rollout @@ -238,7 +237,7 @@ def _expire_after_rollout(state, option): ctx3, _, _ = _make_ctx(points) sim3 = BeliefProbe(ctx3) sim3.reset() - ctx3.attempt_deadline = time.monotonic() - 1.0 + ctx3.python_call_deadline = time.monotonic() - 1.0 with pytest.raises(ProbeBudgetExceeded): sim3.run("Move(block0:block)[0.95]", render=False, physics_sweep=True) diff --git a/tests/agent_sdk/test_bilevel_sketch_regions.py b/tests/agent_sdk/test_bilevel_sketch_regions.py index 4c88070930..e2094cc2ba 100644 --- a/tests/agent_sdk/test_bilevel_sketch_regions.py +++ b/tests/agent_sdk/test_bilevel_sketch_regions.py @@ -4,13 +4,9 @@ A region annotation gives a step's LLM-proposed params per-dimension half-widths: the exact center is tried once, then every later draw for the step is uniform inside ``clip([center - w, center + w], box)`` -instead of the full option box, taking precedence over any per-skill -sampler (region > sampler > uniform). +instead of the full option box. """ -import asyncio -from typing import Any - import numpy as np import pytest from gym.spaces import Box @@ -21,7 +17,7 @@ parse_sketch_from_text, strip_region_annotations from predicators.agent_sdk.sketch_refinement import refine_sketch from predicators.agent_sdk.sketch_types import GroundSampler, SketchStep -from predicators.agent_sdk.tools import ToolContext, create_mcp_tools +from predicators.agent_sdk.tools import ToolContext from predicators.structs import Action, GroundAtom, Object, \ ParameterizedOption, Predicate, State, Task, Type @@ -334,19 +330,6 @@ def test_region_draws_confined_to_window(): assert 0.9 <= float(plan[0].params[0]) <= 0.95 -def test_region_takes_precedence_over_sampler(): - """A registered per-skill sampler is never consulted for a region step.""" - - def sampler(*_args): - raise AssertionError("sampler called despite region annotation") - - plan, success, _ = _refine(_region_step(0.5, 0.5), - max_samples_per_step=200, - parameterized_samplers={"Move": sampler}) - assert success - assert float(plan[0].params[0]) >= 0.9 - - def test_region_window_clipped_to_box(): """An oversized width clips to the option box (draws stay in-box).""" plan, success, _ = _refine(_region_step(0.95, 10.0), @@ -381,24 +364,6 @@ def test_region_applies_on_info_seeking_path(): assert float(plan[0].params[0]) >= 0.9 -def test_region_step_not_capped_by_deterministic_sampler(): - """A deterministic-flagged sampler must not collapse a region step to a - single attempt: the region bypasses the sampler entirely.""" - - def sampler(*_args): - return np.array([0.95], dtype=np.float32) - - sampler.deterministic = True - plan, success, total = _refine(_region_step(0.5, 0.5), - max_samples_per_step=200, - parameterized_samplers={"Move": sampler}) - assert success - # The failing center consumed the first attempt; regional draws (not a - # single deterministic try) then found a passing value. - assert total > 1 - assert float(plan[0].params[0]) >= 0.9 - - def _named_step(fn, name="named"): return SketchStep(option=_Move, objects=[_block], @@ -455,7 +420,6 @@ def fn(*_args): def _tool_ctx(ground_samplers=True, sandbox_dir=None): utils.reset_config({ - "agent_bilevel_use_llm_initial_params": True, "agent_bilevel_max_samples_per_step": 200, "agent_bilevel_ground_samplers": ground_samplers, }) @@ -473,21 +437,6 @@ def _tool_ctx(ground_samplers=True, sandbox_dir=None): ) -def _run_tool(tool_name, args, ground_samplers=True, sandbox_dir=None): - ctx = _tool_ctx(ground_samplers=ground_samplers, sandbox_dir=sandbox_dir) - tools = { - t.name: t.handler - for t in create_mcp_tools(ctx, tool_names=[tool_name]) - } - try: - loop = asyncio.get_event_loop() - except RuntimeError: - loop = asyncio.new_event_loop() - asyncio.set_event_loop(loop) - result: Any = loop.run_until_complete(tools[tool_name](args)) - return result["content"][0]["text"] - - def _probe_refine(plan, ground_samplers=True, sandbox_dir=None): """``sim.refine`` on the fake model - the agent-facing refinement surface (same parser and search core as the explorer's refinement).""" @@ -515,23 +464,6 @@ def test_probe_refine_rejects_bad_region(): _probe_refine("Move(block0:block)[0.85] ~ [0.1, 0.2]") -def test_submit_plan_ignores_region(): - """submit_plan runs the exact center; the region is inert.""" - text = _run_tool( - "submit_plan", { - "plan": ("Move(block0:block)[0.95] ~ [0.05] -> " - "{ReachedHi(block0:block)}"), - "include_states": - False, - "include_atoms": - False, - }) - # Goal achieved proves the exact center 0.95 ran (only x >= 0.9 passes); - # a searched/perturbed value could not be distinguished, so also check - # the report is a plain execution (no refinement verdict lines). - assert "Goal achieved: True" in text - - def test_probe_refine_ignores_region_when_disabled(): """With agent_bilevel_ground_samplers off, the annotation is a no-op. diff --git a/tests/agent_sdk/test_bilevel_sketch_samplers.py b/tests/agent_sdk/test_bilevel_sketch_samplers.py index 93c58a8596..5007c30cea 100644 --- a/tests/agent_sdk/test_bilevel_sketch_samplers.py +++ b/tests/agent_sdk/test_bilevel_sketch_samplers.py @@ -1,12 +1,8 @@ -"""Tests for per-skill synthesized samplers in ``sketch_refinement``. - -Verifies that a sampler registered under an option name in -``parameterized_samplers`` is consulted (with the step's subgoal + -objects + the option's params box) to draw that option's continuous -params during refinement — on both the plain and info-seeking paths — -and that a missing / misbehaving sampler falls back to uniform sampling -so refinement is byte-for-byte unchanged when no usable sampler is -supplied. +"""Tests for how ``sketch_refinement`` draws a step's continuous params: + +uniformly by default, the agent's proposed params first, and on the +info-seeking path a pool of the proposal and sampled draws; plus the +sketch parsing and forward-execution helpers. """ # pylint: disable=unused-import @@ -23,7 +19,7 @@ parse_sketch_from_text, strip_subgoal_annotations from predicators.agent_sdk.sketch_refinement import \ refine_and_validate_report, refine_sketch, sample_params -from predicators.agent_sdk.sketch_types import SketchStep +from predicators.agent_sdk.sketch_types import GroundSampler, SketchStep from predicators.structs import Action, GroundAtom, Object, \ ParameterizedOption, Predicate, State, Task, Type @@ -126,107 +122,8 @@ def _easy_task_and_sketch(): return task, sketch -def test_registered_sampler_is_used(): - """A targeted sampler lands the hard subgoal on the first sample.""" - calls = [] - - def sampler(state, subgoal_atoms, rng, objects): - del state, rng - calls.append((objects, subgoal_atoms)) - return np.array([0.95], dtype=np.float32) - - model = _FakeOptionModel() - plan, success, total = refine_sketch( - _task_hi(), - _sketch_hi(), - model, - predicates={_ReachedHi}, - timeout=10.0, - rng=np.random.default_rng(0), - max_samples_per_step=50, - check_subgoals=True, - check_final_goal=False, - parameterized_samplers={"Move": sampler}) - assert success - assert np.isclose(float(plan[0].params[0]), 0.95) - # Feasible on the very first attempt — none of the uniform churn. - assert total == 1 - assert model.num_calls == 1 - # The sampler saw the right subgoal and objects. - objs, subgoal = calls[0] - assert [o.name for o in objs] == ["block0"] - assert GroundAtom(_ReachedHi, [_block]) in subgoal - - -def test_missing_entry_falls_back_to_uniform(): - """A sampler keyed by another option leaves Move on the uniform path.""" - seed = 7 - first = float(sample_params(_Move, np.random.default_rng(seed))[0]) - task, sketch = _easy_task_and_sketch() - - def other(*_args): - raise AssertionError("sampler for a different option was called") - - plan, success, _ = refine_sketch( - task, - sketch, - _FakeOptionModel(), - predicates={_Reached}, - timeout=10.0, - rng=np.random.default_rng(seed), - max_samples_per_step=50, - check_subgoals=True, - check_final_goal=False, - parameterized_samplers={"OtherOption": other}) - assert success - # Identical to the no-sampler uniform draw. - assert float(plan[0].params[0]) == first - - -def test_bad_shape_falls_back_to_uniform(): - """A wrong-shaped return is rejected; uniform sampling still succeeds.""" - task, sketch = _easy_task_and_sketch() - - def bad(*_args): - return np.array([0.5, 0.5], dtype=np.float32) # shape (2,) != (1,) - - plan, success, _ = refine_sketch(task, - sketch, - _FakeOptionModel(), - predicates={_Reached}, - timeout=10.0, - rng=np.random.default_rng(0), - max_samples_per_step=50, - check_subgoals=True, - check_final_goal=False, - parameterized_samplers={"Move": bad}) - assert success - assert 0.0 <= float(plan[0].params[0]) <= 1.0 - - -def test_raising_sampler_falls_back_to_uniform(): - """A sampler that raises is caught and uniform sampling proceeds.""" - task, sketch = _easy_task_and_sketch() - - def boom(*_args): - raise ValueError("nope") - - _, success, _ = refine_sketch(task, - sketch, - _FakeOptionModel(), - predicates={_Reached}, - timeout=10.0, - rng=np.random.default_rng(0), - max_samples_per_step=50, - check_subgoals=True, - check_final_goal=False, - parameterized_samplers={"Move": boom}) - assert success - - -def test_none_samplers_unchanged(): - """parameterized_samplers=None reproduces the plain first-uniform-draw - param.""" +def test_plain_step_draws_uniformly(): + """A step with no ground sampler takes the plain first uniform draw.""" seed = 7 first = float(sample_params(_Move, np.random.default_rng(seed))[0]) task, sketch = _easy_task_and_sketch() @@ -238,38 +135,11 @@ def test_none_samplers_unchanged(): rng=np.random.default_rng(seed), max_samples_per_step=50, check_subgoals=True, - check_final_goal=False, - parameterized_samplers=None) + check_final_goal=False) assert success assert float(plan[0].params[0]) == first -def test_sampler_used_on_info_seeking_path(): - """The info-seeking draw loop also routes through the sampler.""" - - def sampler(_s, _a, rng, _o): - # Jitter so candidates differ but all clear the x>=0.9 subgoal. - return np.array([0.9 + 0.05 * rng.random()], dtype=np.float32) - - model = _FakeOptionModel() - plan, success, _ = refine_sketch( - _task_hi(), - _sketch_hi(), - model, - predicates={_ReachedHi}, - timeout=10.0, - rng=np.random.default_rng(0), - max_samples_per_step=50, - check_subgoals=True, - check_final_goal=False, - info_scorer=lambda s, _a: s.get(_block, "x"), - info_n_feasible_target=4, - parameterized_samplers={"Move": sampler}) - assert success - # Every pooled candidate came from the sampler => satisfies x >= 0.9. - assert float(plan[0].params[0]) >= 0.9 - - # --------------------------------------------------------------------------- # # LLM-proposed initial_params (tried first, with sampling fallback). # --------------------------------------------------------------------------- # @@ -289,8 +159,7 @@ def test_initial_params_tried_first_without_sampler(): rng=np.random.default_rng(0), max_samples_per_step=50, check_subgoals=True, - check_final_goal=False, - parameterized_samplers=None) + check_final_goal=False) assert success # The proposal satisfied the hard subgoal on the very first attempt. assert np.isclose(float(plan[0].params[0]), 0.95) @@ -312,8 +181,7 @@ def test_initial_params_fall_back_to_uniform_on_failure(): rng=np.random.default_rng(0), max_samples_per_step=200, check_subgoals=True, - check_final_goal=False, - parameterized_samplers=None) + check_final_goal=False) assert success # The failed proposal was the first sample; uniform then found x >= 0.9. assert total > 1 @@ -333,14 +201,18 @@ def test_initial_params_clipped_to_box(): rng=np.random.default_rng(0), max_samples_per_step=50, check_subgoals=True, - check_final_goal=False, - parameterized_samplers=None) + check_final_goal=False) assert success # 5.0 clipped to the option's high bound (1.0), which clears x >= 0.9. assert np.isclose(float(plan[0].params[0]), 1.0) assert total == 1 +def _near_hi(_s, _a, rng, _o): + """Sampled candidates that clear x >= 0.9 but stay below 1.0.""" + return np.array([0.9 + 0.05 * rng.random()], dtype=np.float32) + + def test_initial_params_seeded_and_win_on_disagreement(): """LLM params are pooled with sampled draws; the argmax (most informative) is chosen. @@ -350,13 +222,11 @@ def test_initial_params_seeded_and_win_on_disagreement(): step = SketchStep(option=_Move, objects=[_block], subgoal_atoms={GroundAtom(_ReachedHi, [_block])}, - initial_params=np.array([1.0], dtype=np.float32)) + initial_params=np.array([1.0], dtype=np.float32), + ground_sampler=GroundSampler(fn=_near_hi, + name="near_hi")) model = _FakeOptionModel() - # Sampled candidates clear x >= 0.9 but stay below the guess's x = 1.0. - def sampler(_s, _a, rng, _o): - return np.array([0.9 + 0.05 * rng.random()], dtype=np.float32) - plan, success, _ = refine_sketch( _task_hi(), [step], model, @@ -367,8 +237,7 @@ def sampler(_s, _a, rng, _o): check_subgoals=True, check_final_goal=False, info_scorer=lambda s, _a: float(s.get(_block, "x")), - info_n_feasible_target=4, - parameterized_samplers={"Move": sampler}) + info_n_feasible_target=4) assert success # The guess had the highest disagreement (x = 1.0) => argmax picked it. assert np.isclose(float(plan[0].params[0]), 1.0) @@ -383,7 +252,10 @@ def test_initial_params_lose_to_more_informative_draw(): step = SketchStep(option=_Move, objects=[_block], subgoal_atoms={GroundAtom(_ReachedHi, [_block])}, - initial_params=np.array([0.9], dtype=np.float32)) + initial_params=np.array([0.9], dtype=np.float32), + ground_sampler=GroundSampler( + fn=lambda *_a: np.array([0.99], dtype=np.float32), + name="at_0_99")) plan, success, _ = refine_sketch( _task_hi(), [step], _FakeOptionModel(), @@ -394,10 +266,7 @@ def test_initial_params_lose_to_more_informative_draw(): check_subgoals=True, check_final_goal=False, info_scorer=lambda s, _a: float(s.get(_block, "x")), - info_n_feasible_target=4, - parameterized_samplers={ - "Move": lambda *_a: np.array([0.99], dtype=np.float32) - }) + info_n_feasible_target=4) assert success # The seeded guess (x = 0.9) was beaten by the more informative draw # (0.99) — proving it is pooled, not accepted just for being first. @@ -409,7 +278,9 @@ def test_initial_params_infeasible_seed_info_seeking_recovers(): step = SketchStep(option=_Move, objects=[_block], subgoal_atoms={GroundAtom(_ReachedHi, [_block])}, - initial_params=np.array([0.0], dtype=np.float32)) + initial_params=np.array([0.0], dtype=np.float32), + ground_sampler=GroundSampler(fn=_near_hi, + name="near_hi")) plan, success, _ = refine_sketch( _task_hi(), [step], _FakeOptionModel(), @@ -420,12 +291,7 @@ def test_initial_params_infeasible_seed_info_seeking_recovers(): check_subgoals=True, check_final_goal=False, info_scorer=lambda s, _a: float(s.get(_block, "x")), - info_n_feasible_target=4, - parameterized_samplers={ - "Move": - lambda _s, _a, rng, _o: np.array([0.9 + 0.05 * rng.random()], - dtype=np.float32) - }) + info_n_feasible_target=4) assert success # The infeasible guess (x = 0) wasn't pooled; sampled candidates won. assert float(plan[0].params[0]) >= 0.9 @@ -638,7 +504,7 @@ def test_execute_plan_forward_continues_past_zero_action_failure(): def test_execute_plan_forward_stop_on_failure_aborts(): - """With stop_on_failure (the submit_plan path), a 0-action step aborts + """With stop_on_failure (the probe's trials path), a 0-action step aborts execution like the real executor: later steps don't run and the goal is not reached.""" plan = [ diff --git a/tests/agent_sdk/test_capture_decision.py b/tests/agent_sdk/test_capture_decision.py deleted file mode 100644 index f6a18f2431..0000000000 --- a/tests/agent_sdk/test_capture_decision.py +++ /dev/null @@ -1,263 +0,0 @@ -"""Direct unit tests for the pure ``_decide_capture`` function. - -The e2e harness (``test_submit_plan_capture.py``) drives the -same policy through the real tool handler; these tests pin the decision -table itself, one test per :class:`CaptureDecision` case, including -guard combinations the e2e tests do not reach: - -* capture disabled (``capture_goal_reaching_plans=False``) is silent for - every otherwise-triggering combination; -* an empty grounded plan never captures (unreachable via the handler, - whose parser rejects empty plans; the guard is preserved); -* best-effort mode with an existing validated capture ("best-effort - never displaces validated") falls through to the loud refusals or to - silence; -* a non-terminated evaluator rejection (illegitimate verdict whose - ``terminated`` is False while the env goal-check passed) does not - block a validated capture; -* flaky + evaluator-rejected cannot co-occur in the handler (validation - repeats are skipped on a rejection) but the ``not - evaluator_rejected`` guard on the FLAKY refusal is preserved; -* a wrong-task run without a reached goal is silent. -""" - -from typing import Any, Dict - -from predicators.agent_sdk.tools.capture import BestEffortReason, \ - CaptureDecision, CaptureOutcome, _decide_capture - - -def _decide(**overrides: Any) -> CaptureOutcome: - """Call ``_decide_capture`` on the validated-solve baseline with - overrides.""" - kwargs: Dict[str, Any] = dict(capture_enabled=True, - is_current_task=True, - have_plan=True, - goal_achieved=True, - evaluator_rejected=False, - reward_hack=False, - flaky=False, - best_effort_mode=False, - have_validated_capture=False) - kwargs.update(overrides) - return _decide_capture(**kwargs) - - -def test_validated_capture(): - """A clean goal-reaching plan on the current task is a validated solve. - - e2e: test_robust_plan_is_captured_with_validation_note. - """ - outcome = _decide() - assert outcome.decision is CaptureDecision.VALIDATED_CAPTURE - assert outcome.best_effort_reason is None - assert outcome.captured - - -def test_flaky_no_capture(): - """A flaky validation repeat refuses the capture. - - e2e: test_flaky_plan_is_not_captured. - """ - outcome = _decide(flaky=True) - assert outcome.decision is CaptureDecision.FLAKY_NO_CAPTURE - assert outcome.best_effort_reason is None - assert not outcome.captured - - -def test_param_sensitive_no_capture(): - """A physics-margin rollout failure refuses the capture. - - e2e: test_param_sensitive_plan_is_not_captured. - """ - outcome = _decide(param_sensitive=True) - assert outcome.decision is CaptureDecision.PARAM_SENSITIVE_NO_CAPTURE - assert outcome.best_effort_reason is None - assert not outcome.captured - - -def test_redundant_no_capture(): - """A plan that reaches the goal without one of its steps is refused. - - e2e: test_redundant_step_is_not_captured. - """ - outcome = _decide(redundant=True) - assert outcome.decision is CaptureDecision.REDUNDANT_NO_CAPTURE - assert outcome.best_effort_reason is None - assert not outcome.captured - - -def test_best_effort_redundant(): - """Best-effort mode captures a redundant submission, flagged.""" - outcome = _decide(best_effort_mode=True, redundant=True) - assert outcome.decision is CaptureDecision.BEST_EFFORT_CAPTURE - assert outcome.best_effort_reason is BestEffortReason.REDUNDANT - - -def test_param_sensitive_outranks_redundant_reason(): - """Param-sensitivity is the reason when both are set in best-effort mode - (the handler skips the necessity sweep after a margin failure, but the - ordering is pinned).""" - outcome = _decide(best_effort_mode=True, - param_sensitive=True, - redundant=True) - assert outcome.best_effort_reason is BestEffortReason.PARAM_SENSITIVE - - -def test_best_effort_param_sensitive(): - """Best-effort mode captures a param-sensitive submission.""" - outcome = _decide(best_effort_mode=True, param_sensitive=True) - assert outcome.decision is CaptureDecision.BEST_EFFORT_CAPTURE - assert outcome.best_effort_reason is BestEffortReason.PARAM_SENSITIVE - - -def test_flaky_outranks_param_sensitive_reason(): - """When both gates fail in best-effort mode, FLAKY is the reason (the. - - handler never produces this combination - the margin loop is skipped - after a flaky repeat - but the reason ordering is pinned). - """ - outcome = _decide(best_effort_mode=True, flaky=True, param_sensitive=True) - assert outcome.decision is CaptureDecision.BEST_EFFORT_CAPTURE - assert outcome.best_effort_reason is BestEffortReason.FLAKY - - -def test_reward_hack_no_capture(): - """A goal-atoms-reaching but evaluator-rejected rollout is refused. - - e2e: test_illegitimate_plan_is_not_captured_and_skips_repeats. - """ - outcome = _decide(evaluator_rejected=True, reward_hack=True) - assert outcome.decision is CaptureDecision.REWARD_HACK_NO_CAPTURE - assert not outcome.captured - - -def test_wrong_task_note(): - """A success on a train task is flagged, never captured.""" - outcome = _decide(is_current_task=False) - assert outcome.decision is CaptureDecision.WRONG_TASK_NOTE - assert not outcome.captured - - -def test_no_capture_on_honest_failure(): - """A plan that simply misses the goal is silent - no capture, no flag.""" - outcome = _decide(goal_achieved=False) - assert outcome.decision is CaptureDecision.NO_CAPTURE - assert not outcome.captured - - -def test_best_effort_honest_shortfall(): - """Best-effort mode captures a goal-missing submission as a shortfall. - - e2e: test_best_effort_honest_shortfall_is_captured. - """ - outcome = _decide(best_effort_mode=True, goal_achieved=False) - assert outcome.decision is CaptureDecision.BEST_EFFORT_CAPTURE - assert outcome.best_effort_reason is BestEffortReason.GOAL_NOT_REACHED - assert outcome.captured - - -def test_best_effort_reward_hack(): - """Best-effort mode captures an evaluator-rejected submission. - - e2e: test_best_effort_certificate_rejected_is_captured. - """ - outcome = _decide(best_effort_mode=True, - evaluator_rejected=True, - reward_hack=True) - assert outcome.decision is CaptureDecision.BEST_EFFORT_CAPTURE - assert outcome.best_effort_reason is BestEffortReason.REWARD_HACK - - -def test_best_effort_flaky(): - """Best-effort mode captures a flaky submission. - - e2e: test_best_effort_flaky_plan_is_captured. - """ - outcome = _decide(best_effort_mode=True, flaky=True) - assert outcome.decision is CaptureDecision.BEST_EFFORT_CAPTURE - assert outcome.best_effort_reason is BestEffortReason.FLAKY - - -def test_validated_solve_wins_over_best_effort_mode(): - """Edge: best-effort mode does not demote a fully validated solve.""" - outcome = _decide(best_effort_mode=True) - assert outcome.decision is CaptureDecision.VALIDATED_CAPTURE - assert outcome.best_effort_reason is None - - -def test_best_effort_never_displaces_validated_capture(): - """Edge: with a validated capture already recorded, best-effort mode is - inert - the loud refusals still fire, everything else is silent.""" - # A flaky resubmission is refused (and would escalate the gate) - # instead of overwriting the validated capture. - outcome = _decide(best_effort_mode=True, - have_validated_capture=True, - flaky=True) - assert outcome.decision is CaptureDecision.FLAKY_NO_CAPTURE - # A reward-hack resubmission is likewise refused. - outcome = _decide(best_effort_mode=True, - have_validated_capture=True, - evaluator_rejected=True, - reward_hack=True) - assert outcome.decision is CaptureDecision.REWARD_HACK_NO_CAPTURE - # An honest shortfall is silent. - outcome = _decide(best_effort_mode=True, - have_validated_capture=True, - goal_achieved=False) - assert outcome.decision is CaptureDecision.NO_CAPTURE - - -def test_capture_disabled_is_always_silent(): - """Edge: without capture_goal_reaching_plans every combination is - NO_CAPTURE - the open-loop planner must never record captures.""" - for overrides in ( - {}, # would be a validated solve - { - "flaky": True - }, - { - "evaluator_rejected": True, - "reward_hack": True - }, - { - "is_current_task": False - }, - { - "best_effort_mode": True, - "goal_achieved": False - }, - ): - outcome = _decide(capture_enabled=False, **overrides) - assert outcome.decision is CaptureDecision.NO_CAPTURE - - -def test_empty_plan_never_captures(): - """Edge: no grounded steps means no capture, even for a would-be - validated solve or best-effort submission.""" - assert _decide(have_plan=False).decision is CaptureDecision.NO_CAPTURE - assert _decide(have_plan=False, best_effort_mode=True, - goal_achieved=False).decision is CaptureDecision.NO_CAPTURE - - -def test_non_terminated_evaluator_rejection_does_not_block(): - """Edge: an illegitimate verdict with terminated=False (env goal-check - passed, evaluator's own termination did not) is not a reward hack and - does not block the capture.""" - outcome = _decide(evaluator_rejected=True, reward_hack=False) - assert outcome.decision is CaptureDecision.VALIDATED_CAPTURE - - -def test_flaky_refusal_requires_no_evaluator_rejection(): - """Edge: the FLAKY refusal's ``not evaluator_rejected`` guard - with a - non-terminated rejection alongside flakiness the result is silence, not - a FLAKY refusal (the handler never produces this combination because a - rejection skips validation repeats).""" - outcome = _decide(flaky=True, evaluator_rejected=True) - assert outcome.decision is CaptureDecision.NO_CAPTURE - - -def test_wrong_task_note_requires_goal(): - """Edge: a train-task run that misses the goal is silent, not flagged.""" - outcome = _decide(is_current_task=False, goal_achieved=False) - assert outcome.decision is CaptureDecision.NO_CAPTURE diff --git a/tests/agent_sdk/test_clearance_probe.py b/tests/agent_sdk/test_clearance_probe.py deleted file mode 100644 index 0caa102988..0000000000 --- a/tests/agent_sdk/test_clearance_probe.py +++ /dev/null @@ -1,202 +0,0 @@ -"""Tests for the capture gate's robot-clearance probe (tools/clearance.py). - -Regression for the 2026-09-02 bridge seed3 rerun: two belief-certified -explore plans (8/8 validation rollouts each) died on real contacts of -6.5 mm and 9.9 mm against a block every rollout had cleared - the gate -certified against the belief's own variability but never measured how -close the robot came to a bystander, so plans tighter than the -executor's realization slop passed by the luck of the draw. -""" -from collections import namedtuple -from types import SimpleNamespace -from typing import List, cast - -import numpy as np -import pybullet as p - -from predicators import utils -from predicators.agent_sdk.tools.clearance import RobotClearanceProbe, \ - clearance_lines, phase_skill_of -from predicators.structs import _Option - - -class _FakeSkill: - """A stand-in with only the config the verdict reads.""" - - def __init__(self, tol: float) -> None: - self._config = SimpleNamespace(move_to_pose_tol=tol, simulator=None) - - -def test_verdict_uses_executor_pose_slop() -> None: - """The bar is sqrt(move_to_pose_tol): 1e-4 -> 10 mm.""" - probe = RobotClearanceProbe(_FakeSkill(1e-4)) - assert abs(probe.threshold - 0.01) < 1e-12 - probe.num_probes = 5 - probe.min_dist = 0.0042 - probe.where = "rollout 2, step PickBlock(span2) vs span1" - ok, summary, detail = probe.verdict() - assert not ok - assert "4.2 mm" in summary and "10 mm" in summary - assert "span1" in detail and "inside the executor's 10 mm" in detail - probe.min_dist = 0.0123 - ok, summary, detail = probe.verdict() - assert ok and detail == "" and "12.3 mm" in summary - - -def test_verdict_without_probes_or_approach_is_ok() -> None: - """No probes, or nothing within the query distance, is a pass.""" - probe = RobotClearanceProbe(_FakeSkill(1e-4)) - assert probe.verdict() == (True, "", "") - assert not clearance_lines(probe) - assert not clearance_lines(None) - probe.num_probes = 3 # nothing came within the query distance - ok, summary, detail = probe.verdict() - assert ok and detail == "" and summary.startswith("min robot clearance: >") - - -def test_phase_skill_of_requires_a_planning_simulator() -> None: - """Only a skill-factory option with a planning simulator qualifies.""" - no_sim = cast( - _Option, - SimpleNamespace(parent=SimpleNamespace(policy=SimpleNamespace( - __self__=_FakeSkill(1e-4))))) - assert phase_skill_of([no_sim]) is None - plain = cast( - _Option, - SimpleNamespace(parent=SimpleNamespace(policy=lambda s: None))) - assert phase_skill_of([plain]) is None - - -_Obj = namedtuple("_Obj", ["name"]) # hashable stand-in for an Object - - -class _RecordingProbe(RobotClearanceProbe): - """Records the exempt set handed to each clearance query.""" - - def __init__(self, skill, held: str = "") -> None: - super().__init__(skill, stride=1) - self.exempts: List[set] = [] - self._held = held - - def _min_robot_clearance(self, state, exempt): - self.exempts.append(set(exempt)) - return 0.02, "wall" - - def _held_object_name(self, state): - return self._held or None - - -def _grounded(name: str, objects, skill=None) -> _Option: - parent = SimpleNamespace(policy=SimpleNamespace(__self__=skill)) - return cast(_Option, - SimpleNamespace(name=name, parent=parent, objects=objects)) - - -def test_observe_exempts_declared_contacts_and_the_held_object() -> None: - """A push's switch (skill-declared contact) and the object held when the - option starts join the option's arguments in the exempt set; a skill - without the hook, or a failing hook, exempts only the arguments and the - held object.""" - faucet = _Obj("faucet") - switch = _Obj("faucet_switch") - push_skill = SimpleNamespace( - _config=SimpleNamespace(move_to_pose_tol=1e-4, simulator=None), - contact_objects=lambda state, objects: {switch}) - probe = _RecordingProbe(_FakeSkill(1e-4)) - probe.observe("rollout 1", _grounded("SwitchOn", [faucet], push_skill), - ["s0", "s1"]) - assert probe.exempts == [{"faucet", "faucet_switch"}] * 2 - # A Place holds domino_5 when it starts: the released domino is - # exempt for the whole option (its retreat passes it by design). - probe = _RecordingProbe(_FakeSkill(1e-4), held="domino_5") - probe.observe("rollout 1", _grounded("Place", [], _FakeSkill(1e-4)), - ["s0"]) - assert probe.exempts == [{"domino_5"}] - # A failing contact hook is best-effort: only the arguments remain. - bad_skill = SimpleNamespace(contact_objects=lambda s, o: 1 / 0) - probe = _RecordingProbe(_FakeSkill(1e-4)) - probe.observe("rollout 1", _grounded("Push", [faucet], bad_skill), ["s0"]) - assert probe.exempts == [{"faucet"}] - assert probe.min_dist == 0.02 and "vs wall" in probe.where - - -def test_push_contact_object_is_the_body_at_the_target_pose() -> None: - """object_at_pose picks the posed object nearest the push target within the - contact radius, skipping the grounding's own objects and pose-less - objects.""" - # pylint: disable=import-outside-toplevel - from predicators.ground_truth_models.skill_factories.push import \ - object_at_pose - from predicators.structs import Object, State, Type - posed = Type("posed", ["x", "y", "z"]) - plain = Type("plain", ["is_on"]) - robot = Object("robot", posed) - faucet = Object("faucet", plain) - switch = Object("faucet_switch", posed) - other = Object("burner_switch0", posed) - state = State({ - robot: np.array([1.0, 1.45, 0.7]), - faucet: np.array([0.0]), - switch: np.array([1.0, 1.45, 0.65]), - other: np.array([0.6, 1.3, 0.65]), - }) - target = (1.0, 1.452, 0.65) - assert object_at_pose(state, target, {robot, faucet}) == switch - # The robot is the nearest posed body but is excluded. - assert object_at_pose(state, (1.0, 1.45, 0.66), {robot, faucet}) \ - == switch - # Nothing within the radius: no contact target. - assert object_at_pose(state, (2.0, 2.0, 0.65), {robot, faucet}) is None - - -def test_bridge_probe_measures_and_exempts(tmp_path) -> None: - """On the bridge env's own planning simulator: a block moved under the - gripper reads as a small clearance, the option's argument objects are - exempt, and a far scene reads as beyond the query distance.""" - del tmp_path - utils.reset_config({ - "env": "pybullet_bridge", - "seed": 0, - "num_train_tasks": 1, - "num_test_tasks": 0, - "skill_phase_use_motion_planning": True, - }) - # pylint: disable=import-outside-toplevel - from predicators.envs.pybullet_bridge import PyBulletBridgeEnv - from predicators.ground_truth_models import get_gt_options - env = PyBulletBridgeEnv(use_gui=False) - try: - task = env._generate_train_tasks()[0] # pylint: disable=protected-access - options = {o.name: o for o in get_gt_options(env.get_name())} - robot = env._robot # pylint: disable=protected-access - span1 = next(b for b in env._blocks if b.name == "span1") # pylint: disable=protected-access - pick = options["PickBlock"].ground([robot, span1], - np.array([0.0], dtype=np.float32)) - skill = phase_skill_of([pick]) - assert skill is not None - probe = RobotClearanceProbe(skill, stride=1) - init = task.init - # Far scene: nothing within the query distance of the home pose. - far, _ = probe._min_robot_clearance(init, set()) # pylint: disable=protected-access - assert far >= 0.05 - # Park span1 right under the gripper: a real, small clearance - # (the robot's z is the fingertip point, the block sits 9 cm - # below it, i.e. ~3 cm from the finger geometry). - near = init.copy() - for feat, val in (("x", init.get(robot, - "x")), ("y", init.get(robot, "y")), - ("z", init.get(robot, "z") - 0.09)): - near.set(span1, feat, val) - dist, body = probe._min_robot_clearance(near, set()) # pylint: disable=protected-access - assert body == "span1" and dist < 0.05 - # The pick's own target is exempt: the measurement moves on. - dist_exempt, body_exempt = probe._min_robot_clearance( # pylint: disable=protected-access - near, {"span1"}) - assert body_exempt != "span1" and dist_exempt >= dist - # observe() on the pick's trajectory records the closest approach - # against non-argument bodies only. - probe.observe("rollout 1", pick, [near]) - assert probe.num_probes == 1 - assert not probe.where.endswith("vs span1") - finally: - p.disconnect(env._physics_client_id) # pylint: disable=protected-access diff --git a/tests/agent_sdk/test_docker_agent_runner.py b/tests/agent_sdk/test_docker_agent_runner.py deleted file mode 100644 index 0cdac4e306..0000000000 --- a/tests/agent_sdk/test_docker_agent_runner.py +++ /dev/null @@ -1,165 +0,0 @@ -"""Tests for stale Object hash repair after cross-process unpickling. - -Characterizes ``_rehash_objects_after_unpickle``: Object caches -``_hash = hash(str(self))`` in a cached_property, and PYTHONHASHSEED -randomization makes those cached values stale in a new process. The -tests simulate that by planting a bogus ``_hash`` before building the -``State.data`` dict, so the dict's stored entry hashes are stale, which -is exactly what a cross-seed dill roundtrip produces (verified E2E). - -Characterized behavior (current, not aspirational): - -- The rehash clears every reachable Object's cached ``_hash``/``_str``, - so stale and freshly created equal Objects hash identically again. -- The ``state.data = dict(state.data)`` rebuild does NOT re-key stale - entries: CPython's ``dict(d)`` copies each entry's stored hash without - recomputing it, so lookups into pre-existing dicts remain broken (a - dict comprehension would repair them). Tests below pin this down. - -Only ``_rehash_objects_after_unpickle`` is under test; ``main()`` and -``_run_query`` need Docker and the SDK. -""" -# pylint: disable=protected-access -from types import SimpleNamespace - -import numpy as np -import pytest - -from predicators.agent_sdk.docker_agent_runner import \ - _rehash_objects_after_unpickle -from predicators.structs import Action, GroundAtom, LowLevelTrajectory, \ - Object, Predicate, State, Task, Type - -_block_type = Type("block", ["x"]) -_OnTable = Predicate("OnTable", [_block_type], lambda s, o: True) - - -def _make_state(obj, value=0.0): - return State({obj: np.array([value], dtype=np.float32)}) - - -def _corrupt_hash(obj): - """Plant a stale cached hash, as if pickled under another hash seed.""" - true_hash = hash(obj) # Populate the cached_property. - # +1 guarantees a different bucket index modulo any power-of-two - # table size, so corrupted lookups miss deterministically. - obj.__dict__["_hash"] = true_hash + 1 - return true_hash - - -def _stale_state(name="block0"): - """Build a state whose dict entries are stored under a stale hash. - - Corrupting before insertion mirrors unpickling: dill inserts keys - while their cached (old-process) ``_hash`` is in effect. - """ - obj = Object(name, _block_type) - _corrupt_hash(obj) - state = _make_state(obj) - return obj, state - - -def test_stale_hash_breaks_fresh_object_lookup(): - """Precondition: a fresh equal Object cannot find the stale key.""" - stale_obj, state = _stale_state() - fresh_obj = Object("block0", _block_type) - assert fresh_obj == stale_obj - assert hash(fresh_obj) != hash(stale_obj) - with pytest.raises(KeyError): - _ = state.data[fresh_obj] - # The stale instance itself still works: its cached hash matches - # the hash stored at insertion. - assert state.data[stale_obj] is not None - - -def test_rehash_clears_hash_caches_on_all_ctx_surfaces(): - """All reachable Objects re-hash equal to fresh ones after rehash.""" - train_obj, train_state = _stale_state("train_block") - train_task = Task(train_state, {GroundAtom(_OnTable, [train_obj])}) - - cur_obj, cur_state = _stale_state("cur_block") - cur_task = Task(cur_state, {GroundAtom(_OnTable, [cur_obj])}) - - example_obj, example_state = _stale_state("example_block") - - traj_obj, traj_state0 = _stale_state("traj_block") - # traj_obj still carries the stale cache, so this second dict is - # keyed under it too, like a second unpickled state. - traj_state1 = _make_state(traj_obj, value=1.0) - traj = LowLevelTrajectory([traj_state0, traj_state1], - [Action(np.zeros(1, dtype=np.float32))]) - - ctx = SimpleNamespace(train_tasks=[train_task], - current_task=cur_task, - example_state=example_state, - offline_trajectories=[traj], - online_trajectories=[]) - _rehash_objects_after_unpickle(ctx) - - for name, obj in (("train_block", train_obj), ("cur_block", cur_obj), - ("example_block", example_obj), ("traj_block", - traj_obj)): - fresh = Object(name, _block_type) - assert hash(obj) == hash(fresh), f"cache not cleared for {name}" - # Goal-atom objects were cleared too (same instances here, but the - # atom path is walked independently of the state path). - goal_obj = next(iter(cur_task.goal)).objects[0] - assert hash(goal_obj) == hash(Object("cur_block", _block_type)) - - -def test_rehash_rekeys_stale_entries(): - """The rebuild re-keys entries under the repaired hashes. - - The comprehension in ``_process_state`` is load-bearing: ``dict(d)`` - would copy each entry's stored hash without calling ``__hash__``, - leaving lookups broken even after the caches are cleared. - """ - stale_obj, state = _stale_state() - ctx = SimpleNamespace(train_tasks=[Task(state, set())], current_task=None) - _rehash_objects_after_unpickle(ctx) - - fresh_obj = Object("block0", _block_type) - assert hash(stale_obj) == hash(fresh_obj) # Caches repaired... - assert fresh_obj in state.data # ...and the table re-keyed. - assert state.data[fresh_obj][0] == pytest.approx(0.0) - assert stale_obj in state.data - - -def test_rehash_makes_newly_built_dicts_consistent(): - """Dicts re-keyed by hand after the rehash serve fresh lookups.""" - stale_obj, state = _stale_state() - ctx = SimpleNamespace(train_tasks=[Task(state, set())], current_task=None) - _rehash_objects_after_unpickle(ctx) - # Re-inserting under the repaired hashes (what a comprehension or - # any post-rehash construction does) restores fresh-object lookups. - # pylint: disable-next=unnecessary-comprehension - rekeyed = {obj: val for obj, val in state.data.items()} - fresh_obj = Object("block0", _block_type) - assert fresh_obj in rekeyed - assert rekeyed[fresh_obj][0] == pytest.approx(0.0) - assert stale_obj in rekeyed - - -def test_rehash_processes_environment_task_init_obs(): - """Task-likes exposing init_obs (EnvironmentTask shape) are walked.""" - stale_obj, state = _stale_state("obs_block") - env_task = SimpleNamespace(init_obs=state, goal_description=None) - ctx = SimpleNamespace(train_tasks=[env_task], current_task=None) - _rehash_objects_after_unpickle(ctx) - assert hash(stale_obj) == hash(Object("obs_block", _block_type)) - - -def test_rehash_handles_minimal_ctx_and_none_current_task(): - """A ctx with no tasks, states, or trajectories is a no-op.""" - ctx = SimpleNamespace(current_task=None) - _rehash_objects_after_unpickle(ctx) # Should not raise. - - -def test_rehash_is_idempotent_on_hash_caches(): - """Running the rehash twice keeps hashes consistent.""" - stale_obj, state = _stale_state() - task = Task(state, set()) - ctx = SimpleNamespace(train_tasks=[task], current_task=task) - _rehash_objects_after_unpickle(ctx) - _rehash_objects_after_unpickle(ctx) - assert hash(stale_obj) == hash(Object("block0", _block_type)) diff --git a/tests/agent_sdk/test_ground_sampler_loader.py b/tests/agent_sdk/test_ground_sampler_loader.py new file mode 100644 index 0000000000..7031fba846 --- /dev/null +++ b/tests/agent_sdk/test_ground_sampler_loader.py @@ -0,0 +1,61 @@ +"""Tests for loading the agent's ``GROUND_SAMPLERS``, the named samplers a +sketch step references with ``~ name``.""" + +import numpy as np +from gym.spaces import Box + +from predicators.agent_sdk.proposal_exec import build_exec_context, \ + load_ground_samplers +from predicators.structs import Action, Object, ParameterizedOption, \ + Predicate, Type + +_block_type = Type("block", ["x"]) +_block = Object("block0", _block_type) + +_Reached = Predicate("Reached", [_block_type], lambda s, o: True) + +_Move = ParameterizedOption( + "Move", + types=[_block_type], + params_space=Box(low=np.array([0.0], dtype=np.float32), + high=np.array([1.0], dtype=np.float32)), + policy=lambda _s, _m, _o, _p: Action(np.zeros(1, dtype=np.float32)), + initiable=lambda _s, _m, _o, _p: True, + terminal=lambda _s, _m, _o, _p: False, +) + + +def test_load_ground_samplers_happy_and_bad_entries(): + """GROUND_SAMPLERS loads callables; bad keys/values warn and drop.""" + ctx = build_exec_context(types={_block_type}, + predicates={_Reached}, + options={_Move}) + code = """\ +def _fn(state, subgoal_atoms, rng, objects): + del state, subgoal_atoms, objects + return np.array([0.5], dtype=np.float32) + +GROUND_SAMPLERS = {"hi_band": _fn, "not-an-identifier": _fn, "seven": 7} +""" + fns, warnings, err = load_ground_samplers(code, ctx) + assert err is None + assert set(fns) == {"hi_band"} + assert len(warnings) == 2 + assert any("identifiers" in w for w in warnings) + assert any("not callable" in w for w in warnings) + + +def test_load_ground_samplers_errors(): + """Exec failures and non-dict bindings load nothing, with an error.""" + ctx = build_exec_context(types={_block_type}, + predicates={_Reached}, + options={_Move}) + fns, _, err = load_ground_samplers("raise RuntimeError('boom')", ctx) + assert not fns + assert err is not None and "boom" in err + ctx = build_exec_context(types={_block_type}, + predicates={_Reached}, + options={_Move}) + fns, _, err = load_ground_samplers("GROUND_SAMPLERS = [1]", ctx) + assert not fns + assert err is not None and "must be a dict" in err diff --git a/tests/agent_sdk/test_probe_synthesis.py b/tests/agent_sdk/test_probe_synthesis.py index b30af7d3e8..3e55023a52 100644 --- a/tests/agent_sdk/test_probe_synthesis.py +++ b/tests/agent_sdk/test_probe_synthesis.py @@ -20,7 +20,6 @@ create_mcp_tools from predicators.approaches.agent_sim_learning_approach import \ AgentSimLearningApproach -from predicators.code_sim_learning.fit_space import FitResult from predicators.option_model import _OptionModelBase from predicators.structs import Object, State, Task, Type @@ -83,7 +82,7 @@ def test_probe_reset_requires_task_idx_during_synthesis() -> None: assert sim._state is not None -def test_candidate_probe_model_provider_glue(tmp_path, monkeypatch) -> None: +def test_candidate_probe_model_provider_glue(tmp_path) -> None: """The provider gates on a loadable simulator.py, caches by content hash, rebuilds on change, and NEVER fits: the candidate runs at carried-over. @@ -92,36 +91,19 @@ def test_candidate_probe_model_provider_glue(tmp_path, monkeypatch) -> None: runs at the fitted values (status fitted). Exercises the real ``_make_candidate_probe_model_provider`` and the - real file loader; only the fit/build layer below - ``build_candidate_option_model`` is stubbed (its body is the shared - ``evaluate_plan_refinement`` path). + real file loader; only the option-model build layer below + ``build_candidate_option_model`` is stubbed. """ approach = object.__new__(AgentSimLearningApproach) approach._fitted_params = {} approach._latent_init = None approach._tool_context = ToolContext() - fit_calls = {"n": 0} - - def _fake_fit(rules, specs, triples, features): - del rules, triples, features - fit_calls["n"] += 1 - names = [s.name for s in specs] - return FitResult(names=names, - samples=np.array([[s.init_value for s in specs]]), - log_probs=np.array([0.0])), 0.0 - - monkeypatch.setattr( - "predicators.approaches.synthesis_validation.fit_rule_parameters", - _fake_fit) setattr(approach, "_build_combined_simulator", lambda learned: learned) setattr(approach, "_build_option_model", lambda sim: ("model", sim)) simulator_file = str(tmp_path / "simulator.py") - provider = approach._make_candidate_probe_model_provider( - simulator_file, - trajectories=[], - base_pred_triples=[], - inferred_hint={"thing": ["x"]}) + provider = approach._make_candidate_probe_model_provider(simulator_file, + trajectories=[]) # No file yet: hard error, never a fallback model. with pytest.raises(RuntimeError, match="no candidate simulator yet"): @@ -142,7 +124,6 @@ def _fake_fit(rules, specs, triples, features): f.write(valid) model = provider() # No implicit fit: declared init value, and the status says so. - assert fit_calls["n"] == 0 assert approach._fitted_params == {"k": 1.0} status = approach._tool_context.probe_param_status assert status is not None and status.startswith("UNFITTED") @@ -159,7 +140,6 @@ def _fake_fit(rules, specs, triples, features): assert approach._fitted_params == {"k": 1.7} assert approach._tool_context.probe_param_status == \ "fitted (cycle_000_vers_002)" - assert fit_calls["n"] == 0 # A rejected fit remains unvalidated through rebuild and cache reuse. approach._publish_probe_fit({"k": 1.0}, @@ -179,7 +159,6 @@ def _fake_fit(rules, specs, triples, features): provider() assert approach._tool_context.probe_param_status.startswith("PARTIAL FIT") assert "2/3" in approach._tool_context.probe_param_status - assert fit_calls["n"] == 0 # Changed content: rebuilt UNFITTED, carrying the last fit's value # for a param that still exists inside its box. @@ -198,10 +177,9 @@ def _fake_fit(rules, specs, triples, features): def test_probe_descriptions_follow_phase() -> None: - """The probe surface follows the session: the solve-phase run_python - carries the belief-simulator + submit_plan wording, while the synthesis - run_python's description carries the candidate-simulator + - evaluate_plan_refinement wording.""" + """The probe surface follows the session: the exploration run_python + carries the belief-simulator wording, the synthesis run_python the + candidate-simulator wording, and neither names the removed submit tools.""" utils.reset_config({}) def _desc(ctx: ToolContext) -> str: @@ -213,7 +191,8 @@ def _desc(ctx: ToolContext) -> str: solve_desc = _desc(ToolContext()) assert "belief simulator" in solve_desc - assert "submit_plan" in solve_desc + assert "skills_execute_plan" in solve_desc + assert "submit_plan" not in solve_desc # The solve namespace also carries the recorded real trajectories. assert "trajectories" in solve_desc assert "describe_trajectory" in solve_desc @@ -240,6 +219,7 @@ def _run_python_desc() -> str: # The fit/refine/forward-run protocol replaced the old validation # tool, and the probe is unconditional in synthesis sessions. assert "evaluate_plan_refinement" not in synth_desc + assert "submit_plan" not in synth_desc def test_probe_namespace_contract() -> None: diff --git a/tests/agent_sdk/test_prompt_goldens.py b/tests/agent_sdk/test_prompt_goldens.py index 1823ce87b9..f61ca84d5c 100644 --- a/tests/agent_sdk/test_prompt_goldens.py +++ b/tests/agent_sdk/test_prompt_goldens.py @@ -16,82 +16,19 @@ import os import re -import numpy as np import pytest -from gym.spaces import Box from predicators import utils -from predicators.agent_sdk import learn_prompts, play_prompts +from predicators.agent_sdk import play_prompts from predicators.agent_sdk.prompt_templates import _PROMPTS_DIR, \ load_sections, placeholders, render from predicators.agent_sdk.sandbox_prompts import build_claude_md -from predicators.agent_sdk.sketch_prompts import build_early_stop_note, \ - build_solve_prompt, build_solve_system_prompt from predicators.agent_sdk.tools.continual_tools import CONTINUAL_TOOL_NAMES from predicators.approaches.agent_continual_frozen_approach import \ AgentContinualOracleDynamicsApproach -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, TaskEvaluator, Type _GOLDEN_DIR = os.path.join(os.path.dirname(__file__), "prompt_goldens") -# -- fixtures ---------------------------------------------------------------- - -_THING = Type("thing", ["x", "y"]) -_FIXTURE = Type("fixture", ["x", "y", "is_on"]) - - -def _noop_policy(_s, _m, _o, _p): - return Action(np.zeros(1, dtype=np.float32)) - - -_MOVE = ParameterizedOption( - "MoveTo", - types=[_THING, _FIXTURE], - params_space=Box(low=np.array([-0.5, -0.5], dtype=np.float32), - high=np.array([0.5, 0.5], dtype=np.float32)), - policy=_noop_policy, - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: True, - params_description=("dx", "dy"), -) -_WAIT = ParameterizedOption( - "Wait", - types=[], - params_space=Box(low=np.zeros(0, dtype=np.float32), - high=np.zeros(0, dtype=np.float32)), - policy=_noop_policy, - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: True, -) -_AT = Predicate( - "At", [_THING, _FIXTURE], - lambda s, o: abs(s.get(o[0], "x") - s.get(o[1], "x")) < 0.1, - natural_language_assertion=lambda names: f"{names[0]} rests on {names[1]}") -_ON = Predicate("Active", [_FIXTURE], lambda s, o: s.get(o[0], "is_on") > 0.5) - - -class _Evaluator(TaskEvaluator): - """Evaluator whose objective statement reaches the prompt.""" - - def objective_description(self) -> str: - return "reward = (1.0 if success else 0.0) - 0.1 x moves used" - - -def _make_task(with_evaluator: bool) -> Task: - thing = Object("thing0", _THING) - fixture = Object("fixture0", _FIXTURE) - state = State({ - thing: np.array([0.0, 0.0]), - fixture: np.array([1.0, 0.0, 0.0]) - }) - goal = {GroundAtom(_AT, [thing, fixture]), GroundAtom(_ON, [fixture])} - return Task(state, - goal, - goal_nl="Put the thing on the fixture and switch it on.", - evaluator=_Evaluator(goal) if with_evaluator else None) - - # -- golden comparison ------------------------------------------------------- @@ -148,113 +85,22 @@ def test_render_rejects_missing_and_unused_placeholders() -> None: """A template placeholder without a value, or a value without a placeholder, fails loudly instead of shipping literal text.""" with pytest.raises(AssertionError): - render("solve_query", "goal_nl") + render("subclass_model", "physical_params") with pytest.raises(AssertionError): - render("solve_query", "goal_nl", goal_nl="g", extra="x") + render("subclass_model", "physical_params", param_list="p", extra="x") assert placeholders("a __B__ c __B__ __D__") == ["B", "D"] def test_rendered_prompts_have_no_leftover_placeholders() -> None: """No rendered prompt carries an unsubstituted ``__NAME__``.""" marker = re.compile(r"__[A-Z][A-Z0-9_]*__") - for text in (build_solve_system_prompt(explore=False), - build_solve_system_prompt(explore=True), build_claude_md()): + contract = play_prompts.build_model_contract(partially_observable=True) + tools = ["run_python"] + list(CONTINUAL_TOOL_NAMES) + for text in (play_prompts.build_play_system_prompt( + tools, model_contract=contract), build_claude_md()): assert not marker.search(text), marker.search(text) -# -- solve / explore --------------------------------------------------------- - - -def test_golden_solve_system_plan() -> None: - """Solve-phase system prompt, captured-plan deliverable, all gates.""" - _check_golden( - "solve_system_plan", - build_solve_system_prompt(explore=False, - propose_params=True, - ground_samplers=True, - physics_margin=True, - rule_param_margin=True, - necessity=True, - use_journal=True)) - - -def test_golden_solve_system_policy() -> None: - """Solve-phase system prompt, closed-loop policy deliverable.""" - _check_golden( - "solve_system_policy", - build_solve_system_prompt(explore=False, - policy_mode=True, - rule_param_margin=True, - policy_max_options=40, - policy_max_repeated_failures=3, - policy_max_repeated_noops=5)) - - -def test_golden_solve_system_explore() -> None: - """Explore-phase system prompt with the train-driven early-stop note.""" - utils.reset_config({ - "seed": - 0, - "online_learning_early_stopping": - True, - "online_learning_early_stopping_require_all_attempts": - True, - }) - _check_golden( - "solve_system_explore", - build_solve_system_prompt(explore=True, - physics_margin=True, - rule_param_margin=True, - early_stop_note=build_early_stop_note())) - - -def test_golden_solve_query() -> None: - """Solve query with scoring, run records, and the capture instructions.""" - utils.reset_config({"seed": 0}) - _check_golden( - "solve_query", - build_solve_prompt( - _make_task(with_evaluator=True), - all_predicates={_AT, _ON}, - all_options={_MOVE, _WAIT}, - tool_names=["submit_plan", "run_python", "Read", "Write"], - initial_image_section=( - "## Initial State Image\n\nA rendering of the initial " - "scene is at `./test_images/task000_initial_state.png`. " - "Read it first."), - propose_params=True, - require_tool_validation=True, - journal="### task 0 attempt 1\n- MoveTo dx=0.2 reached At", - strategy="## Approach\n- move, then activate", - attempts="### task 0 attempt 1/3\n- outcome: no capture", - )) - - -def test_golden_explore_query() -> None: - """Explore query with open questions and a belief-certified scheduled - plan.""" - utils.reset_config({"seed": 0}) - _check_golden( - "explore_query", - build_solve_prompt( - _make_task(with_evaluator=False), - all_predicates={_AT, _ON}, - all_options={_MOVE, _WAIT}, - tool_names=["submit_plan", "run_python"], - experiment_guidance=("The learning phase left this ranked " - "ledger of open questions:\n" - "1. Does Active require At? Experiment: " - "MoveTo then Wait."), - scheduled_plans=[ - " 0: MoveTo(thing0, fixture0)[0.2000, 0.0000]\n" - " NOTE: belief-certified; executes verbatim as a solve " - "attempt." - ], - propose_params=True, - explore_mode=True, - )) - - def test_golden_sandbox_claude_md() -> None: """The sandbox CLAUDE.md.""" _check_golden("sandbox_claude_md", build_claude_md()) @@ -263,144 +109,6 @@ def test_golden_sandbox_claude_md() -> None: # -- learn ------------------------------------------------------------------- -def test_golden_learn_system() -> None: - """Fully observable learn system prompt with a physical-parameter menu.""" - physical = learn_prompts.render_physical_params_section({ - "lateral_friction": { - "default": 0.5, - "lo": 0.05, - "hi": 2.0, - "scale": "log", - "description": "sliding friction of every body" - } - }) - _check_golden( - "learn_system", - learn_prompts.build_learn_system_prompt( - partially_observable=False, - residual_rule_signature="def residual_rule(state, updates, " - "params):", - scene_viz_hint="stage and render the scene", - physical_params_section=physical)) - - -def test_golden_learn_system_po_invention() -> None: - """Partially observable learn system prompt with predicate invention.""" - _check_golden( - "learn_system_po_invention", - learn_prompts.build_learn_system_prompt( - partially_observable=True, - residual_rule_signature="def residual_rule(observation, latent, " - "history, updates, params):", - scene_viz_hint="stage and render the scene", - extra_sections=[ - learn_prompts.render_predicate_invention_section( - "the scene workbench"), - ], - latent_extra_sections=[ - learn_prompts.render_predicate_latent_section(), - ], - workflow_extra=learn_prompts.render_predicate_workflow_extra())) - - -def test_golden_learn_message() -> None: - """Learn first message with a prior model, an objective, base-sim source, - and the predicate-invention and partial-observability additions.""" - _check_golden( - "learn_message", - learn_prompts.build_learn_message( - n_trajs=3, - n_transitions=42, - n_demos=1, - n_interaction=2, - trajectory_listing="[0] demo, task 0\n[1] interaction, task 0", - structs_ref="./reference/structs.py", - inferred_hint="{'fixture': ['is_on']}", - predicate_listing=" Holding(robot, thing)", - types_digest="thing: x, y\nfixture: x, y, is_on", - options_digest="MoveTo(thing, fixture)[dx, dy]", - simulator_file="./simulator.py", - objective_block=learn_prompts.render_objective_block( - "reward = success - 0.1 x moves used"), - prior_state_block=learn_prompts.render_prior_state_block( - ["`./simulator.py`", "`./predicates.py`"]), - divergence_block=learn_prompts.render_divergence_block( - "fixture.is_on: 4 mismatches", has_prior_model=True), - base_sim_block=learn_prompts.render_base_sim_block( - ["./reference/base_sim/scene.py"]), - tools_block=learn_prompts.render_tools_block( - ["run_python", "Read", "Write", "Edit"]), - extra_messages=[ - learn_prompts.render_predicate_invention_message( - "./predicates.py", - "Goal (natural language): switch the fixture on."), - learn_prompts.render_partial_observability_message(), - ])) - - -def test_golden_learn_program_system() -> None: - """The program-world-model learn system prompt with predicate invention - (paper arm C4).""" - _check_golden( - "learn_program_system", - learn_prompts.build_program_learn_system_prompt( - scene_viz_hint="stage and render the scene", - extra_sections=[ - learn_prompts.render_predicate_invention_section( - "the scene workbench"), - ], - workflow_extra=learn_prompts.render_predicate_workflow_extra())) - - -def test_golden_learn_program_message() -> None: - """The program-world-model learn first message, zero-shot variant.""" - _check_golden( - "learn_program_message", - learn_prompts.build_program_learn_message( - n_trajs=0, - n_transitions=0, - n_demos=0, - n_interaction=0, - trajectory_listing="", - structs_ref="./reference/structs.py", - predicate_listing="- Holding(robot:robot, block:block)", - types_digest="- robot: hand\n- block: x, y, held", - options_digest="- Pick(robot:robot, block:block)[]", - world_model_file="./world_model.py", - tools_block=learn_prompts.render_tools_block(["run_python"]), - extra_messages=[ - learn_prompts.render_program_zero_shot_message(), - ])) - - -def test_golden_learn_notes_system() -> None: - """The natural-language world-model learn system prompt (paper arm C3).""" - _check_golden("learn_notes_system", - learn_prompts.build_notes_learn_system_prompt()) - - -def test_golden_learn_notes_message() -> None: - """The natural-language world-model learn first message with prior notes - and a goal.""" - _check_golden( - "learn_notes_message", - learn_prompts.build_notes_learn_message( - n_trajs=2, - n_transitions=9, - n_demos=0, - n_interaction=2, - trajectory_listing=" [0] interaction, task 0\n" - " [1] interaction, task 0", - structs_ref="./reference/structs.py", - predicate_listing="- Holding(robot:robot, block:block)", - types_digest="- robot: hand\n- block: x, y, held", - options_digest="- Pick(robot:robot, block:block)[]", - notes_file="./world_model.md", - goal_nls=["Build the bridge.", "Build the bridge."], - has_prior_notes=True, - tools_block=learn_prompts.render_tools_block(["run_python"]))) - - @pytest.mark.parametrize("model_based,noise,repair", [ (True, False, False), (True, True, True), @@ -549,7 +257,8 @@ def test_golden_continual_system_ablation(arm): if arm == "no_uncertainty": flags["continual_obs_noise_declared"] = False # As the experiments run them: every arm but No uncertainty carries - # the joint belief (continual_common.yaml, approaches/continual.yaml). + # the joint belief (scripts/configs/empiric/common.yaml and + # approaches.yaml). flags["belief_joint_draws"] = 0 if arm == "no_uncertainty" else 16 utils.reset_config(flags) tools = ["run_python"] + list(CONTINUAL_TOOL_NAMES) diff --git a/tests/agent_sdk/test_refine_evaluator_gate.py b/tests/agent_sdk/test_refine_evaluator_gate.py index 53da53e2b2..9917f23314 100644 --- a/tests/agent_sdk/test_refine_evaluator_gate.py +++ b/tests/agent_sdk/test_refine_evaluator_gate.py @@ -88,7 +88,6 @@ def _make_probe(evaluator, model=None): # (p = 0.1 each) a 1e-9 event, so the certified/rejected paths # below are exercised at every probe seed, not just lucky ones. "agent_bilevel_max_samples_per_step": 200, - "agent_bilevel_use_llm_initial_params": False, }) init = State({_block: np.array([0.0], dtype=np.float32)}) goal = {GroundAtom(_ReachedHi, [_block])} diff --git a/tests/agent_sdk/test_sampler_synthesis_tools.py b/tests/agent_sdk/test_sampler_synthesis_tools.py deleted file mode 100644 index 9ab8f50ded..0000000000 --- a/tests/agent_sdk/test_sampler_synthesis_tools.py +++ /dev/null @@ -1,241 +0,0 @@ -"""Tests for the ``sim.samplers()`` loader (make_sampler_loader). - -Drives the real tool handler against a stub approach: loading -``LEARNED_SAMPLERS`` from ``samplers.py``, installing the validated dict -onto the approach, skip warnings for bad entries, error paths, snapshot -versioning, and the sanity check's empty-subgoal-set contract. -""" - -from typing import Any, Dict, Set - -import numpy as np -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk.belief_probe import BeliefProbe -from predicators.agent_sdk.proposal_exec import build_exec_context, \ - load_ground_samplers -from predicators.agent_sdk.tools import ToolContext, make_sampler_loader -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, Type - -_block_type = Type("block", ["x"]) -_block = Object("block0", _block_type) - -_Reached = Predicate("Reached", [_block_type], lambda s, o: True) - -_Move = ParameterizedOption( - "Move", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=lambda _s, _m, _o, _p: Action(np.zeros(1, dtype=np.float32)), - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: False, -) - - -class _StubApproach: - """The minimal approach surface make_sampler_loader uses.""" - - def __init__(self): - init = State({_block: np.array([0.0], dtype=np.float32)}) - self._types = {_block_type} - self._train_tasks = [Task(init, {GroundAtom(_Reached, [_block])})] - self._fitted_params: Dict[str, float] = {} - self._synthesized_samplers: Dict[str, Any] = {} - - def _get_all_predicates(self) -> Set[Predicate]: - return {_Reached} - - def _get_all_options(self) -> Set[ParameterizedOption]: - return {_Move} - - -def _run_sampler_loader(tmp_path, code=None): - utils.reset_config({"seed": 0}) - samplers_file = str(tmp_path / "samplers.py") - if code is not None: - with open(samplers_file, "w", encoding="utf-8") as f: - f.write(code) - approach = _StubApproach() - loader = make_sampler_loader( - samplers_file=samplers_file, - samplers_versions_dir=str(tmp_path / "samplers_versions"), - approach=approach, - cycle_index_provider=lambda: 1, - ) - # Through the probe, exactly as the agent reaches it. - ctx = ToolContext() - ctx.probe_artifact_loaders["samplers"] = loader - return BeliefProbe(ctx).samplers(), approach - - -_GOOD = """\ -def _move_sampler(state, subgoal_atoms, rng, objects): - del state, subgoal_atoms, objects - return np.array([0.25 + 0.01 * rng.random()], dtype=np.float32) - -LEARNED_SAMPLERS = {"Move": _move_sampler} -""" - - -def test_evaluate_sampler_installs_valid_samplers(tmp_path): - """A valid samplers.py is installed onto the approach and passes the sanity - check.""" - text, approach = _run_sampler_loader(tmp_path, _GOOD) - assert "1 per-skill sampler(s) installed" in text - assert "Move: OK" in text - assert "3/3 within the params box" in text - assert set(approach._synthesized_samplers) == {"Move"} # pylint: disable=protected-access - assert callable(approach._synthesized_samplers["Move"]) # pylint: disable=protected-access - - -def test_evaluate_sampler_warns_unknown_option(tmp_path): - """An entry keyed by a non-option name is skipped with a warning.""" - code = _GOOD + "\nLEARNED_SAMPLERS['Teleport'] = _move_sampler\n" - text, approach = _run_sampler_loader(tmp_path, code) - assert "Skipped 'Teleport' (not a known option name" in text - assert set(approach._synthesized_samplers) == {"Move"} # pylint: disable=protected-access - - -def test_evaluate_sampler_warns_non_callable(tmp_path): - """A non-callable value is skipped with a warning.""" - code = _GOOD + "\nLEARNED_SAMPLERS['Move'] = 3.0\n" - text, approach = _run_sampler_loader(tmp_path, code) - assert "Skipped 'Move' (value is not callable" in text - assert not approach._synthesized_samplers # pylint: disable=protected-access - - -def test_evaluate_sampler_reports_exec_error(tmp_path): - """A samplers.py that raises at import time reports the traceback.""" - text, approach = _run_sampler_loader(tmp_path, - "raise RuntimeError('boom')") - assert "Error executing" in text - assert "boom" in text - assert not approach._synthesized_samplers # pylint: disable=protected-access - - -def test_evaluate_sampler_reports_missing_symbol(tmp_path): - """A file without LEARNED_SAMPLERS names the missing symbol.""" - text, _ = _run_sampler_loader(tmp_path, "x = 1\n") - assert "LEARNED_SAMPLERS" in text - - -def test_evaluate_sampler_missing_file_hint(tmp_path): - """A missing samplers.py returns the Write hint, not a crash.""" - text, _ = _run_sampler_loader(tmp_path, code=None) - assert "Use Write to create it" in text - - -def test_evaluate_sampler_empty_dict_message(tmp_path): - """An empty LEARNED_SAMPLERS asks for entries instead of sanity lines.""" - text, _ = _run_sampler_loader(tmp_path, "LEARNED_SAMPLERS = {}\n") - assert "LEARNED_SAMPLERS is empty" in text - assert "Sanity check" not in text - - -def test_evaluate_sampler_version_tag_bumps_on_edit(tmp_path): - """Within one tool instance (one synthesis session), an edited samplers.py - gets a fresh version tag; an identical reload keeps it.""" - utils.reset_config({"seed": 0}) - samplers_file = tmp_path / "samplers.py" - samplers_file.write_text(_GOOD, encoding="utf-8") - loader = make_sampler_loader( - samplers_file=str(samplers_file), - samplers_versions_dir=str(tmp_path / "samplers_versions"), - approach=_StubApproach(), - cycle_index_provider=lambda: 1, - ) - - def _tag(): - return loader().split("]")[0].lstrip("[") - - tag1 = _tag() - tag2 = _tag() # unchanged file: same tag (snapshot deduped) - samplers_file.write_text(_GOOD.replace("0.25", "0.75"), encoding="utf-8") - tag3 = _tag() - assert tag1 == "cycle_001_vers_001" - assert tag2 == tag1 - assert tag3 == "cycle_001_vers_002" - - -def test_sanity_check_raising_sampler_mentions_empty_subgoal_contract( - tmp_path): - """A sampler that assumes a non-empty subgoal set gets the contract spelled - out: refinement calls samplers with subgoal_atoms=set() at steps with no - annotation, and the sanity check does the same.""" - code = """\ -def _needs_subgoal(state, subgoal_atoms, rng, objects): - atom = next(iter(subgoal_atoms)) - del state, rng, objects, atom - return np.array([0.5], dtype=np.float32) - -LEARNED_SAMPLERS = {"Move": _needs_subgoal} -""" - text, _ = _run_sampler_loader(tmp_path, code) - assert "ERROR" in text - assert "subgoal_atoms=set()" in text - assert "must not crash on an empty set" in text - - -def test_load_ground_samplers_happy_and_bad_entries(): - """GROUND_SAMPLERS loads callables; bad keys/values warn and drop.""" - ctx = build_exec_context(types={_block_type}, - predicates={_Reached}, - options={_Move}) - code = """\ -def _fn(state, subgoal_atoms, rng, objects): - del state, subgoal_atoms, objects - return np.array([0.5], dtype=np.float32) - -GROUND_SAMPLERS = {"hi_band": _fn, "not-an-identifier": _fn, "seven": 7} -""" - fns, warnings, err = load_ground_samplers(code, ctx) - assert err is None - assert set(fns) == {"hi_band"} - assert len(warnings) == 2 - assert any("identifiers" in w for w in warnings) - assert any("not callable" in w for w in warnings) - - -def test_load_ground_samplers_errors(): - """Exec failures and non-dict bindings load nothing, with an error.""" - ctx = build_exec_context(types={_block_type}, - predicates={_Reached}, - options={_Move}) - fns, _, err = load_ground_samplers("raise RuntimeError('boom')", ctx) - assert not fns - assert err is not None and "boom" in err - ctx = build_exec_context(types={_block_type}, - predicates={_Reached}, - options={_Move}) - fns, _, err = load_ground_samplers("GROUND_SAMPLERS = [1]", ctx) - assert not fns - assert err is not None and "must be a dict" in err - - -def test_sanity_check_wrong_shape_reports_error(tmp_path): - """A wrong-shaped return is reported with got/expected shapes.""" - code = """\ -def _bad_shape(state, subgoal_atoms, rng, objects): - del state, subgoal_atoms, rng, objects - return np.array([0.5, 0.5], dtype=np.float32) - -LEARNED_SAMPLERS = {"Move": _bad_shape} -""" - text, _ = _run_sampler_loader(tmp_path, code) - assert "ERROR" in text - assert "returned shape (2,), expected (1,)" in text - - -def test_probe_samplers_unavailable_without_a_loader(): - """Outside a sampler-synthesis session the probe has no samplers.py surface - and says so instead of silently doing nothing.""" - probe = BeliefProbe(ToolContext()) - try: - probe.samplers() - except RuntimeError as e: - assert "sim.samplers is unavailable" in str(e) - else: - raise AssertionError("sim.samplers() must raise without a loader") diff --git a/tests/agent_sdk/test_session_fatal.py b/tests/agent_sdk/test_session_fatal.py index ffa01c604b..680f1866cf 100644 --- a/tests/agent_sdk/test_session_fatal.py +++ b/tests/agent_sdk/test_session_fatal.py @@ -463,7 +463,6 @@ async def _fake_sleep(secs): class _Ctx: attempt_start = 100.0 - attempt_deadline = 2800.0 python_call_deadline = None paused = 0.0 @@ -471,7 +470,6 @@ def pause_attempt_clock(self, seconds): """Shift the armed marks like the real ToolContext does.""" self.paused += seconds self.attempt_start += seconds - self.attempt_deadline += seconds ctx = _Ctx() monkeypatch.setattr(sb, "stream_agent_response", _fake_stream) @@ -485,7 +483,7 @@ def pause_attempt_clock(self, seconds): assert collected == healthy_resp assert sleeps == [sb._LIMIT_POLL_SECS] * 3 assert ctx.paused == sum(sleeps) - assert ctx.attempt_deadline == 2800.0 + sum(sleeps) + assert ctx.attempt_start == 100.0 + sum(sleeps) # Past the total cap the limited response is handed back as-is. monkeypatch.setattr(sb, "_LIMIT_MAX_TOTAL_WAIT_SECS", 1.0) responses[:] = [list(limit_resp), list(healthy_resp)] diff --git a/tests/agent_sdk/test_solve_prompt_strategy.py b/tests/agent_sdk/test_solve_prompt_strategy.py deleted file mode 100644 index 10353c78d2..0000000000 --- a/tests/agent_sdk/test_solve_prompt_strategy.py +++ /dev/null @@ -1,190 +0,0 @@ -"""Behavioral checks on the solve/explore prompts. - -Each test pins one rule the prompts must carry (the goldens in -``test_prompt_goldens.py`` pin the full text): the reward form reaches -the solver, solutions are banked before being optimized, the run records -are used skeptically, and exploration is framed as experiment design -against a belief model. -""" -import numpy as np - -from predicators import utils -from predicators.agent_sdk.sketch_prompts import build_early_stop_note, \ - build_solve_prompt, build_solve_system_prompt -from predicators.structs import Object, State, Task, TaskEvaluator, Type - -_DOM = Type("thing", ["x"]) - - -class _StubEvaluator(TaskEvaluator): - """Evaluator whose objective statement must reach the prompt.""" - - def objective_description(self) -> str: - return "reward = (1.0 if success else 0.0) - 0.2 x widgets used" - - -def _make_task(evaluator=None) -> Task: - obj = Object("thing0", _DOM) - return Task(State({obj: np.array([0.0])}), - set(), - goal_nl="Do the thing.", - evaluator=evaluator) - - -def _render(task: Task, journal: str = "", strategy: str = "") -> str: - return build_solve_prompt(task, - all_predicates=set(), - all_options=set(), - propose_params=True, - require_tool_validation=True, - journal=journal, - strategy=strategy) - - -def _render_explore(task: Task, propose_params: bool = True) -> str: - return build_solve_prompt(task, - all_predicates=set(), - all_options=set(), - propose_params=propose_params, - require_tool_validation=False, - explore_mode=True) - - -def test_scoring_section_from_evaluator() -> None: - """The evaluator's public reward form renders as a Scoring section; without - an evaluator the section is absent.""" - utils.reset_config({"seed": 0}) - prompt = _render(_make_task(_StubEvaluator(set()))) - assert "## Scoring (env ground-truth reward)" in prompt - assert "0.2 x widgets used" in prompt - assert "Decode every reward you observe" in prompt - assert "## Scoring" not in _render(_make_task(None)) - - -def test_solve_system_prompt_banks_before_optimizing() -> None: - """The solve system prompt states the banking semantics (a validated - capture replaces the banked one, a rejection never does); the explore - system prompt, which has no capture deliverable, omits them.""" - prompt = build_solve_system_prompt(explore=False) - assert "Bank a solution before optimizing it" in prompt - assert "a rejected submission never displaces it" in prompt - assert "strictly better" in prompt - assert "Bank a solution" not in build_solve_system_prompt(explore=True) - - -def test_solve_system_prompt_journal_protocol() -> None: - """The run-record protocol (incumbent, untried leads, scoped negatives, - conflicting entries) lives in the solve system prompt and only there.""" - prompt = build_solve_system_prompt(explore=False, use_journal=True) - assert "is the incumbent" in prompt - assert "untried leads first" in prompt - assert "only as broad as the family actually swept" in prompt - assert "both become open questions" in prompt - assert "Journal protocol" not in build_solve_system_prompt( - explore=False, use_journal=False) - assert "Journal protocol" not in build_solve_system_prompt(explore=True) - - -def test_solve_system_prompt_verifies_rules_before_steering() -> None: - """A rule inferred from one observation is re-tested before it guides the - search.""" - prompt = build_solve_system_prompt(explore=False) - assert "Verify a rule before steering by it" in prompt - - -def test_query_carries_run_records_not_their_protocol() -> None: - """The query renders the journal and strategy contents under their headers; - the rules for using them are not repeated there.""" - utils.reset_config({"seed": 0}) - prompt = _render(_make_task(None), - journal="- notes", - strategy="## Glue first\n- dab twice") - assert "## Solve Journal (./journal.md)" in prompt - assert "- notes" in prompt - assert "## Domain Strategy (advisory, written during learning)" in prompt - assert "- dab twice" in prompt - assert "incumbent" not in prompt - bare = _render(_make_task(None)) - assert "## Solve Journal" not in bare - assert "## Domain Strategy" not in bare - - -def test_explore_system_prompt_states_the_setting() -> None: - """The explore system prompt discloses the belief model, accepts a - simulator-failing plan, states the cycle data contract and first-cycle - coverage, and executes parameters verbatim; the solve prompt has none.""" - prompt = build_solve_system_prompt(explore=True) - assert "## Exploration setting" in prompt - assert "not in the belief model yet" in prompt - assert "a simulator-failing plan is a valid deliverable" in prompt - assert "at least one attempt at the full goal" in prompt - assert ("top-ranked open question's experiment executed as specified" - in prompt) - assert "Carry each interaction to its consequence" in prompt - assert "nothing is searched or substituted" in prompt - solve = build_solve_system_prompt(explore=False) - assert "## Exploration setting" not in solve - assert "belief model yet" not in solve - - -def test_explore_deliverable_and_certified_note() -> None: - """The explore deliverable is the final plan text; the certified-plan note - follows the execute-certified-plan flag.""" - with_note = build_solve_system_prompt(explore=True, - execute_certified_plan=True) - assert "Your final plan text: the experiment" in with_note - assert "executed verbatim as this episode's solve attempt" in with_note - without = build_solve_system_prompt(explore=True, - execute_certified_plan=False) - assert "executed verbatim" not in without - - -def test_early_stop_note_follows_config() -> None: - """The early-stop note credits certified exploration plans that solve for - real (train-driven) or perfect test phases (test-driven), and is absent - when early stopping is off.""" - utils.reset_config({ - "seed": - 0, - "online_learning_early_stopping": - True, - "online_learning_early_stopping_require_all_attempts": - True, - }) - note = build_early_stop_note() - assert note.startswith("The loop concludes early once the exploration") - assert "every episode of a cycle" in note - prompt = build_solve_system_prompt(explore=True, early_stop_note=note) - assert "exploration plans solve training" in prompt - utils.reset_config({ - "seed": - 0, - "online_learning_early_stopping": - True, - "online_learning_early_stopping_by_test_solve_rate": - True, - "online_learning_early_stopping_consecutive_perfect_tests": - 2, - }) - assert "2 consecutive test phases" in build_early_stop_note() - utils.reset_config({"seed": 0, "online_learning_early_stopping": False}) - assert build_early_stop_note() == "" - - -def test_query_openings_by_phase() -> None: - """Explore queries open with the information-gathering framing; solve - queries with the solving framing. - - Both end in the instructions. - """ - utils.reset_config({"seed": 0}) - explore = _render_explore(_make_task(None)) - assert explore.startswith("Design this episode's experiment") - assert "part of information gathering" in explore - assert "output the plan lines as your final text" in explore - solve = _render(_make_task(None)) - assert solve.startswith("Solve the task below") - assert "deliver it through the capture gate" in solve - # Param-free sketch mode keeps the same delivery semantics. - sketch = _render_explore(_make_task(None), propose_params=False) - assert "output the plan lines as your final text" in sketch diff --git a/tests/agent_sdk/test_solve_restart_journal.py b/tests/agent_sdk/test_solve_restart_journal.py index ca151bedc3..8254a4833b 100644 --- a/tests/agent_sdk/test_solve_restart_journal.py +++ b/tests/agent_sdk/test_solve_restart_journal.py @@ -1,10 +1,8 @@ -"""Tests for the solve journal and the wall-clock exploration budgets. +"""Tests for the run journal and ``run_python``'s budgets. Covers the journal module (entry caps, prompt-injection trimming, the -harness-owned attempt log), the cooperative probe deadline -(:class:`ProbeBudgetExceeded`), and ``run_python``'s budget handling -(refusal after the attempt deadline, per-call timeout with partial -output, ``[budget]`` footer). +harness-owned round log) and ``run_python``'s budget handling (per-call +timeout with partial output, ``[budget]`` footer). """ # pylint: disable=protected-access import asyncio @@ -17,7 +15,7 @@ from predicators import utils from predicators.agent_sdk import journal as journal_mod -from predicators.agent_sdk.belief_probe import BeliefProbe, ProbeBudgetExceeded +from predicators.agent_sdk.belief_probe import BeliefProbe from predicators.agent_sdk.tools import ToolContext, create_mcp_tools from predicators.structs import Action, GroundAtom, LowLevelTrajectory, \ Object, ParameterizedOption, Predicate, State, Task, Type @@ -107,19 +105,17 @@ def test_journal_append_and_read(tmp_path): """Entries append under headers and read back verbatim.""" sandbox = str(tmp_path) assert journal_mod.read_journal(sandbox) == "" - assert journal_mod.append_entry(sandbox, "task 0 attempt 1 (auto)", - "- outcome: no capture") is None + journal_mod.append_entry(sandbox, "Round 1", "- no environment action") content = journal_mod.read_journal(sandbox, filename=journal_mod.ATTEMPTS_FILENAME) - assert "### task 0 attempt 1 (auto)" in content - assert "- outcome: no capture" in content + assert "### Round 1" in content + assert "- no environment action" in content def test_journal_entry_truncated_at_cap(tmp_path): - """Oversize entries are truncated with a notice.""" + """Oversize entries are truncated with a marker.""" sandbox = str(tmp_path) - note = journal_mod.append_entry(sandbox, "big", "x" * 10000) - assert note is not None and "truncated" in note + journal_mod.append_entry(sandbox, "big", "x" * 10000) content = journal_mod.read_journal(sandbox, max_chars=10**6, filename=_A) assert "[entry truncated at the per-entry size cap]" in content assert len(content) < 5000 @@ -145,26 +141,6 @@ def test_journal_read_no_sandbox(): assert journal_mod.read_journal(None) == "" -def test_journal_read_raw_and_restore(tmp_path): - """read_raw snapshots faithfully and restore rolls entries back.""" - sandbox = str(tmp_path) - assert journal_mod.read_raw(None) is None - assert journal_mod.read_raw(sandbox, filename=_A) is None - journal_mod.append_entry(sandbox, "Agent notes (pre-test phase)", - "- learning fact") - snapshot = journal_mod.read_raw(sandbox, filename=_A) - assert snapshot is not None and "- learning fact" in snapshot - journal_mod.append_entry(sandbox, "Agent notes (test task 0)", - "- test-phase fact") - journal_mod.restore(sandbox, snapshot, filename=_A) - assert journal_mod.read_raw(sandbox, filename=_A) == snapshot - # A None snapshot means no journal file existed: restore deletes. - journal_mod.restore(sandbox, None, filename=_A) - assert journal_mod.read_raw(sandbox, filename=_A) is None - # Deleting an already-absent journal is a no-op, not an error. - journal_mod.restore(sandbox, None, filename=_A) - - # --------------------------------------------------------------------------- # attempt log (harness-owned file next to the agent's journal) # --------------------------------------------------------------------------- @@ -172,99 +148,27 @@ def test_journal_read_raw_and_restore(tmp_path): def test_attempt_log_is_a_separate_file(tmp_path): """Harness entries land in attempts.md; the agent's journal.md is a plain - file the harness never writes, and each is read, snapshotted and restored - on its own.""" + file the harness never writes, and each is read on its own.""" sandbox = str(tmp_path) - assert journal_mod.append_entry(sandbox, "task 0 attempt 1/1 (auto)", - "- outcome: no capture") is None - assert not os.path.isfile(journal_mod.journal_path(sandbox)) - assert os.path.isfile(journal_mod.attempts_path(sandbox)) + journal_path = os.path.join(sandbox, journal_mod.JOURNAL_FILENAME) + journal_mod.append_entry(sandbox, "Round 1", "- no environment action") + assert not os.path.isfile(journal_path) + assert os.path.isfile(os.path.join(sandbox, _A)) assert journal_mod.read_journal(sandbox) == "" - attempts = journal_mod.read_journal(sandbox, - filename=journal_mod.ATTEMPTS_FILENAME) - assert "### task 0 attempt 1/1 (auto)" in attempts + assert "### Round 1" in journal_mod.read_journal(sandbox, filename=_A) # The agent writes its journal with the file tools. - with open(journal_mod.journal_path(sandbox), "w", encoding="utf-8") as f: - f.write("### task 0 attempt 1\n- tried x=0.5: stopped 3 cm short\n") - assert "stopped 3 cm short" in journal_mod.read_journal(sandbox) - snapshot = journal_mod.read_raw(sandbox, - filename=journal_mod.ATTEMPTS_FILENAME) - journal_mod.append_entry(sandbox, "task 1 attempt 1/1 (auto)", - "- outcome: captured") - journal_mod.restore(sandbox, - snapshot, - filename=journal_mod.ATTEMPTS_FILENAME) - assert "task 1" not in journal_mod.read_journal( - sandbox, filename=journal_mod.ATTEMPTS_FILENAME) + with open(journal_path, "w", encoding="utf-8") as f: + f.write("### Level 1\n- tried x=0.5: stopped 3 cm short\n") assert "stopped 3 cm short" in journal_mod.read_journal(sandbox) + assert "stopped 3 cm short" not in journal_mod.read_journal(sandbox, + filename=_A) # --------------------------------------------------------------------------- -# probe deadline +# probe rollout metering # --------------------------------------------------------------------------- -def test_probe_raises_after_attempt_deadline(): - """Past the attempt deadline every probe sim call raises.""" - utils.reset_config({}) - ctx = _make_ctx() - ctx.attempt_deadline = time.monotonic() - 1.0 - sim = BeliefProbe(ctx) - try: - sim.reset() - assert False, "expected ProbeBudgetExceeded" - except ProbeBudgetExceeded as e: - assert "submit your single best plan" in str(e) - - -def test_probe_deadline_skipped_during_best_effort_nudge(): - """The final-submission nudge is never blocked by the spent budget.""" - utils.reset_config({}) - ctx = _make_ctx() - ctx.attempt_deadline = time.monotonic() - 1.0 - ctx.capture_best_effort_plan = True - sim = BeliefProbe(ctx) - sim.reset() # must not raise - - -def test_probe_trials_returns_partial_on_mid_loop_budget_expiry(): - """A budget stop mid-trials returns the completed trials (they are minutes - of sim time living in the return value, not stdout) instead of discarding - them.""" - utils.reset_config({}) - ctx = _make_ctx() - ctx.attempt_deadline = time.monotonic() + 60.0 - model = ctx.option_model - orig = model.get_next_state_and_num_actions - - def _expire_after_rollout(state, option): - result = orig(state, option) - ctx.attempt_deadline = time.monotonic() - 1.0 - return result - - model.get_next_state_and_num_actions = _expire_after_rollout - sim = BeliefProbe(ctx) - sim.reset() - res = sim.run("Move(block0:block)[0.95]", render=False, trials=5) - assert len(res.trials) == 1 - assert res.successes == 1 - assert any("time budget expired after 1/5 trials" in n for n in res.notes) - - -def test_probe_trials_reraises_when_nothing_completed(): - """With zero completed trials there is nothing to salvage.""" - utils.reset_config({}) - ctx = _make_ctx() - sim = BeliefProbe(ctx) - sim.reset() - ctx.attempt_deadline = time.monotonic() - 1.0 - try: - sim.run("Move(block0:block)[0.95]", render=False, trials=3) - assert False, "expected ProbeBudgetExceeded" - except ProbeBudgetExceeded: - pass - - def test_probe_counts_rollouts(): """run() meters full-plan rollouts (single and trials).""" utils.reset_config({}) @@ -283,19 +187,6 @@ def test_probe_counts_rollouts(): # --------------------------------------------------------------------------- -def test_run_python_refuses_after_attempt_deadline(tmp_path): - """A call arriving past the attempt deadline is refused unrun.""" - utils.reset_config({}) - ctx = _make_ctx(sandbox_dir=str(tmp_path)) - ctx.attempt_start = time.monotonic() - 10.0 - ctx.attempt_deadline = time.monotonic() - 1.0 - text = _call(_get_tool(ctx, "run_python"), - {"code": "print('should not run')"}) - assert "wall-clock exploration budget" in text - assert "should not run" not in text - assert "[budget]" in text - - def test_python_call_timeout_returns_partial_output(tmp_path): """A per-call timeout stops the sweep and returns printed output.""" utils.reset_config({ @@ -317,12 +208,10 @@ def test_run_python_budget_footer(tmp_path): }) ctx = _make_ctx(sandbox_dir=str(tmp_path)) ctx.attempt_start = time.monotonic() - ctx.attempt_deadline = ctx.attempt_start + 2700 code = "sim.reset(); print(sim.run('Move(block0:block)[0.95]', " \ "render=False).goal_reached)" text = _call(_get_tool(ctx, "run_python"), {"code": code}) assert "[budget] attempt time" in text - assert "/45 min" in text assert "sim rollouts this attempt: 1 (+1 this call)" in text @@ -379,55 +268,6 @@ def test_run_python_no_footer_outside_attempt(tmp_path): assert "[budget]" not in text -# --------------------------------------------------------------------------- -# prompt injection -# --------------------------------------------------------------------------- - - -def test_solve_prompt_includes_journal_section(): - """build_solve_prompt renders the journal and attempt log contents.""" - # pylint: disable-next=import-outside-toplevel - from predicators.agent_sdk.sketch_prompts import build_solve_prompt - utils.reset_config({}) - ctx = _make_ctx() - task = ctx.train_tasks[0] - journal_text = ("### task 0 attempt 1/3 (auto)\n" - "- outcome: no capture") - prompt = build_solve_prompt(task, - all_predicates={_ReachedHi}, - all_options={_Move}, - journal="### notes\n- tried x=0.5", - attempts=journal_text) - assert "## Attempt Log" in prompt - assert "- outcome: no capture" in prompt - assert "## Solve Journal" in prompt - assert "- tried x=0.5" in prompt - assert "./journal.md" in prompt - assert "record_journal" not in prompt - # Without journal content the section is absent entirely. - prompt_no_journal = build_solve_prompt(task, - all_predicates={_ReachedHi}, - all_options={_Move}) - assert "## Solve Journal" not in prompt_no_journal - assert "## Attempt Log" not in prompt_no_journal - - -def test_read_strategy_absent_present_and_truncated(tmp_path): - """read_strategy: "" when absent, verbatim when small, head-kept cap.""" - sandbox = str(tmp_path) - assert journal_mod.read_strategy(sandbox) == "" - assert journal_mod.read_strategy(None) == "" - with open(journal_mod.strategy_path(sandbox), "w", encoding="utf-8") as f: - f.write("## Approach\n- glue both faces\n") - assert "- glue both faces" in journal_mod.read_strategy(sandbox) - with open(journal_mod.strategy_path(sandbox), "w", encoding="utf-8") as f: - f.write("HEADLINE\n" + "x" * 10000) - content = journal_mod.read_strategy(sandbox) - assert content.startswith("HEADLINE") - assert "[strategy truncated at the prompt cap" in content - assert len(content) < 4300 - - # --------------------------------------------------------------------------- # run_python path argument # --------------------------------------------------------------------------- diff --git a/tests/agent_sdk/test_submit_plan_capture.py b/tests/agent_sdk/test_submit_plan_capture.py deleted file mode 100644 index 12fa29558e..0000000000 --- a/tests/agent_sdk/test_submit_plan_capture.py +++ /dev/null @@ -1,1146 +0,0 @@ -"""Capture-gating tests for the ``submit_plan`` tool. - -Drives the real MCP tool handler with a fake option model (no PyBullet), -covering the two gates in front of ``ctx.solved_plan``: - -* multi-rollout validation - the shared sim env is nondeterministic - across repeats, so a goal-reaching plan is captured only after every - one of ``CFG.agent_plan_validation_rollouts`` rollouts succeeds; a - flaky plan is reported to the agent instead of captured; -* task-evaluator legitimacy - a goal-reaching but ``legitimate=False`` - rollout is refused as a reward hack (the real evaluator applies the - same certificate). The refusal is internal: the agent-facing report - speaks only in (terminated, reward) terms and never leaks the - certificate's legitimacy bool or reason string. - -Under ``capture_best_effort_plan`` (the final-submission nudge) neither -gate refuses: the submission is captured regardless - honest shortfall, -certificate-rejected rollout, or flaky repeat - but is marked as NOT a -validated solve (``solved_plan_reached_goal=False``), so it executes for -its honest reward without counting as a solve. -""" - -import asyncio -import contextlib -from typing import Any - -import numpy as np -import pytest -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk.tools import ToolContext, create_mcp_tools -from predicators.structs import Action, GroundAtom, LowLevelTrajectory, \ - Object, ParameterizedOption, Predicate, State, Task, TaskEvaluator, Type - -_block_type = Type("block", ["x"]) -_block = Object("block0", _block_type) - -_ReachedHi = Predicate("ReachedHi", [_block_type], - lambda s, o: s.get(o[0], "x") >= 0.9) - - -def _noop_policy(_s, _m, _o, _p): - return Action(np.zeros(1, dtype=np.float32)) - - -_Move = ParameterizedOption( - "Move", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: False, -) - -_PLAN_TEXT = "Move(block0:block)[0.95] -> {ReachedHi(block0:block)}" -# A plan that lands short of the goal (x=0.5 < 0.9), so ReachedHi never -# holds: an honest shortfall, not a reward hack. -_SHORTFALL_PLAN_TEXT = "Move(block0:block)[0.5]" - - -class _Model: - """Fake option model: Move sets block.x to its parameter value. - - ``succeed_first_n`` bounds how many calls apply the parameter; later - calls leave the state unchanged, emulating a flaky plan whose repeat - rollout misses the goal. Exposes ``last_trajectory`` so evaluator - verdicts are NON-coarse (a coarse verdict never blocks capture). - """ - - last_execution_failure = None - - def __init__(self, succeed_first_n=10**9): - self.num_calls = 0 - self._succeed_first_n = succeed_first_n - self.last_trajectory = None - - def get_next_state_and_num_actions(self, state, option): - """Roll the option forward one step, counting the call.""" - self.num_calls += 1 - nxt = state.copy() - if self.num_calls <= self._succeed_first_n and len(option.params): - nxt.set(_block, "x", float(option.params[0])) - self.last_trajectory = LowLevelTrajectory( - [state, nxt], [Action(np.zeros(1, dtype=np.float32))]) - return nxt, 1 - - -class _StubEvaluator(TaskEvaluator): - """Deterministic legitimacy verdict.""" - - def __init__(self, goal, legit): - super().__init__(goal) - self._legit = legit - - def _certify(self, states, step_options, sim_env=None): - if self._legit: - return True, "" - return False, "stub: the cascade was staged, not pushed" - - -def _make_ctx(model, evaluator=None, best_effort=False, goal_nl=None): - init = State({_block: np.array([0.0], dtype=np.float32)}) - goal = {GroundAtom(_ReachedHi, [_block])} - task = Task(init, goal, evaluator=evaluator, goal_nl=goal_nl) - ctx = ToolContext( - types={_block_type}, - predicates={_ReachedHi}, - processes=set(), - options={_Move}, - train_tasks=[task], - example_state=init, - option_model=model, - current_task=task, - ) - ctx.capture_goal_reaching_plans = True - ctx.capture_best_effort_plan = best_effort - return ctx - - -def _call_tool(ctx, plan_text=_PLAN_TEXT, extra_args=None): - """Invoke the real tool handler once against ``ctx``.""" - tools = { - t.name: t.handler - for t in create_mcp_tools(ctx, tool_names=["submit_plan"]) - } - try: - loop = asyncio.get_event_loop() - except RuntimeError: - loop = asyncio.new_event_loop() - asyncio.set_event_loop(loop) - call_args = {"plan": plan_text} - if extra_args: - call_args.update(extra_args) - result: Any = loop.run_until_complete(tools["submit_plan"](call_args)) - return result["content"][0]["text"] - - -def _run_tool(model, - evaluator=None, - rollouts=3, - plan_text=_PLAN_TEXT, - best_effort=False, - goal_nl=None, - extra_args=None): - utils.reset_config({"agent_plan_validation_rollouts": rollouts}) - ctx = _make_ctx(model, - evaluator=evaluator, - best_effort=best_effort, - goal_nl=goal_nl) - return _call_tool(ctx, plan_text, extra_args=extra_args), ctx - - -def test_robust_plan_is_captured_with_validation_note(): - """All rollouts succeed: captured, with the K/K validation note.""" - model = _Model() - text, ctx = _run_tool(model, rollouts=3) - assert "Captured as the current answer" in text - assert "Validated 3/3 rollouts" in text - assert ctx.solved_plan is not None - # One reported rollout + two validation repeats. - assert model.num_calls == 3 - - -def test_flaky_plan_is_not_captured(): - """A repeat rollout that misses the goal blocks capture, loudly.""" - model = _Model(succeed_first_n=1) - text, ctx = _run_tool(model, rollouts=3) - assert "FLAKY (plan NOT captured)" in text - assert "rollout 2/3 (planner seed" in text - assert "goal not reached" in text - assert "ReachedHi" in text - assert "Captured as the current answer" not in text - assert ctx.solved_plan is None - # The first rollout DID reach the goal - that is what makes it flaky. - assert "Goal achieved: True" in text - - -def test_single_rollout_config_disables_repeats(): - """``agent_plan_validation_rollouts=1`` restores single-rollout capture.""" - model = _Model() - text, ctx = _run_tool(model, rollouts=1) - assert "Captured as the current answer" in text - assert "Validated" not in text - assert ctx.solved_plan is not None - - -def test_flaky_message_reports_all_rollout_outcomes(): - """The FLAKY report lists EVERY rollout's outcome and an estimated. - - reliability, instead of stopping at the first failure - the - per-rollout list is what distinguishes failure modes. - """ - model = _Model(succeed_first_n=1) - text, _ = _run_tool(model, rollouts=3) - assert "estimated reliability 1/3" in text - assert "rollout 1 (planner seed" in text - assert "): goal reached" in text - assert "rollout 2 (planner seed" in text - assert "rollout 3 (planner seed" in text - assert text.count("FAILED -") >= 2 - # All three rollouts actually ran (no early break). - assert model.num_calls == 3 - - -def test_flaky_verdict_line_labeled_as_rollout_1(): - """When the submission is rejected as FLAKY, rollout 1's evaluator. - - verdict is labeled as such - unlabeled it read as a second, - contradictory verdict in the same message. - """ - model = _Model(succeed_first_n=1) - evaluator = _StubEvaluator({GroundAtom(_ReachedHi, [_block])}, True) - text, _ = _run_tool(model, evaluator=evaluator, rollouts=3) - assert "FLAKY (plan NOT captured)" in text - assert "[rollout 1 only - NOT the operative outcome" in text - - -def test_missing_goal_atoms_printed_even_with_goal_nl(): - """A goal-nl task still names the missing goal atoms on a shortfall - 'Goal - achieved: False' alone left agents unable to tell a near-miss from a non- - starter.""" - model = _Model() - text, _ = _run_tool(model, - plan_text=_SHORTFALL_PLAN_TEXT, - goal_nl="topple the target") - assert "Goal achieved: False" in text - assert "Missing goal atoms" in text - assert "ReachedHi" in text - assert model.num_calls == 1 - - -def test_illegitimate_plan_is_not_captured_and_skips_repeats(): - """A non-coarse ``legitimate=False`` verdict refuses capture before any - validation repeats are spent, reported in reward terms only: the - certificate's reason string never reaches the agent.""" - model = _Model() - goal = {GroundAtom(_ReachedHi, [_block])} - text, ctx = _run_tool(model, - evaluator=_StubEvaluator(goal, legit=False), - rollouts=3) - assert "NOT CAPTURED" in text - assert "scores it as a non-solve" in text - assert "stub: the cascade was staged" not in text - assert "legitimate" not in text - assert "Captured as the current answer" not in text - assert ctx.solved_plan is None - assert model.num_calls == 1 - - -def test_legitimate_plan_passes_both_gates(): - """A legitimate goal-reaching plan validates and captures normally.""" - model = _Model() - goal = {GroundAtom(_ReachedHi, [_block])} - text, ctx = _run_tool(model, - evaluator=_StubEvaluator(goal, legit=True), - rollouts=3) - assert "Captured as the current answer" in text - assert "Validated 3/3 rollouts" in text - assert "Task evaluator" in text and "reward=" in text - assert "solved=True" in text - assert "legitimate" not in text - assert ctx.solved_plan is not None - assert model.num_calls == 3 - - -def test_best_effort_honest_shortfall_is_captured(): - """An honest best-effort shortfall is captured, not refused. - - The plan does not reach the goal (terminated=False), so the - evaluator marks it legitimate=False - there is no genuine cascade to - certify. But it is not a reward hack, so under a best-effort - submission it is captured and executes for its honest reward instead - of being forfeited. - """ - model = _Model() - goal = {GroundAtom(_ReachedHi, [_block])} - text, ctx = _run_tool(model, - evaluator=_StubEvaluator(goal, legit=False), - rollouts=3, - plan_text=_SHORTFALL_PLAN_TEXT, - best_effort=True) - assert "Captured as the current answer" in text - assert "best-effort: goal NOT reached" in text - assert "will not count as a solve" in text - assert "NOT CAPTURED" not in text - assert ctx.solved_plan is not None - assert ctx.solved_plan_reached_goal is False - assert "Goal achieved: False" in text - - -def test_best_effort_certificate_rejected_is_captured(): - """A best-effort submission captures even a certificate-rejected rollout. - - The plan reaches the goal atoms (terminated=True) but the evaluator - scores the route as a non-solve. Outside best-effort mode that is - refused as a reward hack, but at final submission the budget is - spent: the plan is captured to execute for its honest reward, marked - as NOT a validated solve (run_20260714_145053 task 4: this refusal - forfeited the task entirely). - """ - model = _Model() - goal = {GroundAtom(_ReachedHi, [_block])} - text, ctx = _run_tool(model, - evaluator=_StubEvaluator(goal, legit=False), - rollouts=3, - best_effort=True) - assert "Captured as the current answer" in text - assert "best-effort" in text - assert "will not count as a solve" in text - assert "stub: the cascade was staged" not in text - assert "legitimate" not in text - assert "NOT CAPTURED" not in text - assert ctx.solved_plan is not None - assert ctx.solved_plan_reached_goal is False - # Certificate rejection skips the validation repeats. - assert model.num_calls == 1 - - -def test_flaky_rejection_escalates_later_captures(): - """A FLAKY rejection escalates the gate for later captures on the task. - - A flaky submission is evidence the agent is tuning in a marginal - parameter region, where a lucky streak can pass the base 3-rollout - gate and die on the single real episode (run_20260717_182321: a - 20/20-swept relay placement validated 3/3, then missed the target - for real). Resubmissions must therefore clear the escalated - ``agent_plan_validation_rollouts_after_flaky`` gate. - """ - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_rollouts_after_flaky": 6, - }) - flaky_model = _Model(succeed_first_n=1) - ctx = _make_ctx(flaky_model) - text = _call_tool(ctx) - assert "FLAKY (plan NOT captured)" in text - assert "captures require 6/6 successful rollouts" in text - # The resubmission (robust this time) faces the 6-rollout gate. - robust_model = _Model() - ctx.option_model = robust_model - text2 = _call_tool(ctx) - assert "Captured as the current answer" in text2 - assert "Validated 6/6 rollouts" in text2 - assert ctx.solved_plan is not None - assert robust_model.num_calls == 6 - - -def test_flaky_escalation_is_per_task(): - """Escalation is keyed to the task: a different test task keeps the base - gate.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_rollouts_after_flaky": 6, - }) - flaky_model = _Model(succeed_first_n=1) - ctx = _make_ctx(flaky_model) - ctx.test_task_idx = 0 - text = _call_tool(ctx) - assert "FLAKY (plan NOT captured)" in text - robust_model = _Model() - ctx.option_model = robust_model - ctx.test_task_idx = 1 - text2 = _call_tool(ctx) - assert "Validated 3/3 rollouts" in text2 - assert robust_model.num_calls == 3 - - -def test_validation_rollouts_enter_fresh_env_scope(): - """Every gate rollout - the reported first one included - runs inside - ``ctx.validation_env_scope``, so the whole gate shares one substrate - (reproducible in-session via ``sim.run(plan, trials=N)``).""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - }) - entered = [] - - @contextlib.contextmanager - def _scope(): - entered.append(True) - yield - - model = _Model() - ctx = _make_ctx(model) - ctx.validation_env_scope = _scope - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "freshly constructed simulator" in text - # All 3 rollouts, the reported first one included. - assert len(entered) == 3 - - -def test_fresh_env_scope_disabled_by_config(): - """``agent_plan_validation_fresh_env=False`` keeps repeats on the shared - env even when a scope is installed.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": False, - }) - entered = [] - - @contextlib.contextmanager - def _scope(): - entered.append(True) - yield - - model = _Model() - ctx = _make_ctx(model) - ctx.validation_env_scope = _scope - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "freshly constructed simulator" not in text - assert not entered - - -class _PhysicsAwareModel(_Model): - """Fake model whose success depends on a physics 'parameter'. - - Move only applies its parameter while ``friction >= 0.5``, emulating - a plan whose success band excludes part of the fit posterior. The - physics-margin scope perturbs ``friction`` the way the real scope - perturbs the fresh env's physical params. - """ - - def __init__(self): - super().__init__() - self.friction = 0.53 - - def get_next_state_and_num_actions(self, state, option): - if self.friction >= 0.5: - return super().get_next_state_and_num_actions(state, option) - self.num_calls += 1 - nxt = state.copy() - self.last_trajectory = LowLevelTrajectory( - [state, nxt], [Action(np.zeros(1, dtype=np.float32))]) - return nxt, 1 - - -def _physics_scope_ctx(model, points): - """A ctx whose fresh-env scope applies physics overrides to ``model``.""" - ctx = _make_ctx(model) - scope_overrides = [] - - @contextlib.contextmanager - def _scope(physical_overrides=None): - scope_overrides.append(physical_overrides) - prev = model.friction - if physical_overrides: - model.friction = physical_overrides["lateral_friction"] - try: - yield - finally: - model.friction = prev - - ctx.validation_env_scope = _scope - ctx.physics_margin_provider = lambda: list(points) - return ctx, scope_overrides - - -class _AdditiveModel(_PhysicsAwareModel): - """Fake model where each Move ADDS its parameter to block.x. - - Two Move(0.5) steps are both necessary to reach x >= 0.9, while a - Move(0.95) makes any other Move padding - the two shapes the - necessity gate has to tell apart. - """ - - def get_next_state_and_num_actions(self, state, option): - self.num_calls += 1 - nxt = state.copy() - nxt.set(_block, "x", - float(state.get(_block, "x")) + float(option.params[0])) - self.last_trajectory = LowLevelTrajectory( - [state, nxt], [Action(np.zeros(1, dtype=np.float32))]) - return nxt, 1 - - -_REDUNDANT_PLAN_TEXT = ( - "Move(block0:block)[0.95]\n" - "Move(block0:block)[0.95] -> {ReachedHi(block0:block)}") -_TWO_STEP_PLAN_TEXT = ("Move(block0:block)[0.5]\n" - "Move(block0:block)[0.5] -> {ReachedHi(block0:block)}") - - -def test_redundant_step_is_not_captured(): - """A plan that still reaches the goal with a step removed is refused. - - Regression for run_20260902_152811: a validated capture pressed - three of four buttons and released one that was never on, for a goal - its own model reached with two presses and a Wait. Every gate ran at - the full plan, so none could see the padding. - """ - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - "agent_plan_validation_necessity": True, - }) - ctx, _ = _physics_scope_ctx(_AdditiveModel(), []) - text = _call_tool(ctx, plan_text=_REDUNDANT_PLAN_TEXT) - assert "REDUNDANT (plan NOT captured)" in text - assert "without step 0 (Move(block0)): goal STILL reached" in text - assert ctx.solved_plan is None - # The best refused submission is stashed for the best-effort nudge. - assert ctx.best_uncaptured_plan_lines is not None - - -def test_necessary_steps_are_captured_with_note(): - """A plan whose every step is needed captures with the check's note.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - "agent_plan_validation_necessity": True, - }) - ctx, _ = _physics_scope_ctx(_AdditiveModel(), []) - text = _call_tool(ctx, plan_text=_TWO_STEP_PLAN_TEXT) - assert "Captured as the current answer" in text - assert "Necessity check passed" in text - assert ctx.solved_plan is not None - - -def test_necessity_gate_disabled_by_config(): - """The default-off flag captures the padded plan without ablations.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - "agent_plan_validation_necessity": False, - }) - ctx, scope_overrides = _physics_scope_ctx(_AdditiveModel(), []) - text = _call_tool(ctx, plan_text=_REDUNDANT_PLAN_TEXT) - assert "Captured as the current answer" in text - assert "REDUNDANT" not in text - # Main rollout + 2 execution repeats, no ablation rollouts. - assert scope_overrides == [None, None, None] - - -def test_param_sensitive_plan_is_not_captured(): - """A plan that fails at a -1-sigma physics point is refused. - - Regression for run_20260723_091108: a capture validated 8/8 at the - fitted lateral_friction 0.5319 failed deterministically at true 0.5 - - execution repeats at the fitted values cannot see zero margin to - the fit's parameter error. - """ - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": True, - }) - model = _PhysicsAwareModel() - ctx, scope_overrides = _physics_scope_ctx(model, [{ - "lateral_friction": 0.48 - }, { - "lateral_friction": 0.59 - }]) - text = _call_tool(ctx) - assert "PARAM-SENSITIVE (plan NOT captured)" in text - assert "lateral_friction=0.48" in text - assert ctx.solved_plan is None - # Main rollout + 2 execution repeats (no overrides) + 2 physics - # points. - assert scope_overrides == [ - None, None, None, { - "lateral_friction": 0.48 - }, { - "lateral_friction": 0.59 - } - ] - - -def test_param_sensitive_refusal_names_the_straddle(): - """Under the interval belief the refusal carries the certified fraction, - the passing and failing ranges and the probe cue.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": True, - "code_sim_learning_interval_belief": True, - }) - model = _PhysicsAwareModel() - ctx, _ = _physics_scope_ctx(model, [{ - "lateral_friction": 0.48 - }, { - "lateral_friction": 0.59 - }]) - text = _call_tool(ctx) - assert "PARAM-SENSITIVE (plan NOT captured)" in text - assert ("(1/2 belief-interval points passed; lateral_friction: fails at " - "0.48, passes at 0.59)") in text - assert "straddles this plan's success boundary" in text - assert "sim.suggest_probes" in text - assert ctx.param_sensitive_refusal_pending - assert ctx.solved_plan is None - - -def test_physics_margin_pass_is_captured_with_note(): - """Margin points inside the success band capture with the margin note.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": True, - }) - model = _PhysicsAwareModel() - ctx, _ = _physics_scope_ctx(model, [{ - "lateral_friction": 0.51 - }, { - "lateral_friction": 0.59 - }]) - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "Physics-margin check passed" in text - - -def test_physics_margin_disabled_by_config(): - """The default-off flag skips the margin rollouts entirely.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - }) - model = _PhysicsAwareModel() - ctx, scope_overrides = _physics_scope_ctx(model, [{ - "lateral_friction": 0.48 - }]) - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "PARAM-SENSITIVE" not in text - assert scope_overrides == [None, None, None] - - -def test_physics_margin_vacuous_without_points(): - """An empty provider (no fit / degenerate posterior) adds no note.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": True, - }) - model = _PhysicsAwareModel() - ctx, scope_overrides = _physics_scope_ctx(model, []) - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "Physics-margin check passed" not in text - assert scope_overrides == [None, None, None] - - -def _rule_param_scope_ctx(model, members): - """A ctx whose rule-param override scope applies members to ``model``.""" - ctx = _make_ctx(model) - applied = [] - - @contextlib.contextmanager - def _fresh_scope(physical_overrides=None): - del physical_overrides - yield - - @contextlib.contextmanager - def _override(point): - applied.append(point) - prev = model.friction - model.friction = point["dab_tol"] - try: - yield - finally: - model.friction = prev - - ctx.validation_env_scope = _fresh_scope - ctx.rule_param_margin_provider = lambda: list(members) - ctx.rule_param_override_scope = _override - return ctx, applied - - -def test_rule_param_sensitive_plan_is_not_captured(): - """A plan that fails under a calibrated rule-param ensemble member is - refused as PARAM-SENSITIVE. - - Regression for the bridge cycles 5-7: plans centered on marginal - operating points of uncertain learned constants (a glue dab at 16 mm - of the true 20 mm radius) passed 5/5 nominal validation rollouts and - the (empty) physics sweep, then failed in the real environment. - """ - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - "agent_plan_validation_rule_param_margin": True, - }) - model = _PhysicsAwareModel() - ctx, applied = _rule_param_scope_ctx(model, [{ - "dab_tol": 0.59 - }, { - "dab_tol": 0.48 - }]) - text = _call_tool(ctx) - assert "PARAM-SENSITIVE (plan NOT captured)" in text - assert "rule-param ensemble member 2/2" in text - assert "dab_tol=0.48" in text - assert ctx.solved_plan is None - assert applied == [{"dab_tol": 0.59}, {"dab_tol": 0.48}] - - -def test_rule_param_margin_pass_is_captured_with_note(): - """Members inside the success band capture with the margin note.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - "agent_plan_validation_rule_param_margin": True, - }) - model = _PhysicsAwareModel() - ctx, _ = _rule_param_scope_ctx(model, [{ - "dab_tol": 0.51 - }, { - "dab_tol": 0.59 - }]) - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "Rule-parameter margin check passed" in text - - -def test_rule_param_margin_disabled_by_config(): - """The default-off flag skips the rule-param sweep entirely.""" - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_plan_validation_fresh_env": True, - "agent_plan_validation_physics_margin": False, - "agent_plan_validation_rule_param_margin": False, - }) - model = _PhysicsAwareModel() - ctx, applied = _rule_param_scope_ctx(model, [{"dab_tol": 0.48}]) - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "PARAM-SENSITIVE" not in text - assert not applied - - -def test_best_effort_flaky_plan_is_captured(): - """A best-effort submission captures a flaky plan instead of refusing. - - Rollout 1 solves, rollout 2 misses. Outside best-effort mode that is - refused as FLAKY (the agent can add margin and resubmit), but at - final submission there is no budget left, so the plan is captured - with the flaky detail in the note and marked as NOT a validated - solve. - """ - model = _Model(succeed_first_n=1) - text, ctx = _run_tool(model, rollouts=3, best_effort=True) - assert "Captured as the current answer" in text - assert "best-effort" in text - assert "rollout 2/3 (planner seed" in text - assert "FLAKY (plan NOT captured)" not in text - assert ctx.solved_plan is not None - assert ctx.solved_plan_reached_goal is False - - -def test_capture_stashes_evaluator_reward(): - """A capture records the evaluator verdict's reward for the restart loop's - cross-attempt ranking.""" - model = _Model() - goal = {GroundAtom(_ReachedHi, [_block])} - _, ctx = _run_tool(model, - evaluator=_StubEvaluator(goal, legit=True), - rollouts=1) - assert ctx.solved_plan is not None - assert isinstance(ctx.solved_plan_eval_reward, float) - - -def test_capture_without_evaluator_has_no_reward(): - """No evaluator: the reward stash stays None (ranked below any rewarded - capture, above no capture).""" - model = _Model() - _, ctx = _run_tool(model, rollouts=1) - assert ctx.solved_plan is not None - assert ctx.solved_plan_eval_reward is None - - -def test_budget_footer_during_attempt(): - """The footer reports THIS call's rollout delta, not the attempt's - cumulative total masquerading as one.""" - import time as _time # pylint: disable=import-outside-toplevel - utils.reset_config({"agent_plan_validation_rollouts": 1}) - model = _Model() - ctx = _make_ctx(model) - ctx.attempt_start = _time.monotonic() - # Simulate a prior explore sweep this attempt. - ctx.attempt_rollout_count = 50 - text = _call_tool(ctx, _PLAN_TEXT) - assert "[budget] attempt time" in text - assert "sim rollouts this attempt: 51 (+1 this call)" in text - - -def test_no_budget_footer_outside_attempt(): - """No attempt in flight: no footer noise.""" - model = _Model() - text, _ctx = _run_tool(model, rollouts=1) - assert "[budget]" not in text - - -def test_validation_repeats_use_decorrelated_planner_seeds(): - """Each validation repeat rolls out under its own ``CFG.seed``. - - A fresh env per repeat is not enough for independent samples: the - skills' motion planning reads the constant ``CFG.seed`` at call - time, so identical-seed repeats are bit-identical replays and the - flaky gate detects nothing (run_20260722_204632: a 13/13-validated - capture was a coin flip on the real episode). The capture rollout - itself must keep the base seed; the repeats offset it; the base seed - must be restored afterward. - """ - - class _SeedRecordingModel(_Model): - """Records ``CFG.seed`` at each rollout step.""" - - def __init__(self): - super().__init__() - self.seeds = [] - - def get_next_state_and_num_actions(self, state, option): - from predicators.settings import \ - CFG # pylint: disable=import-outside-toplevel - self.seeds.append(CFG.seed) - return super().get_next_state_and_num_actions(state, option) - - model = _SeedRecordingModel() - _, ctx = _run_tool(model, rollouts=3) - from predicators.settings import \ - CFG # pylint: disable=import-outside-toplevel - base = CFG.seed - assert ctx.solved_plan is not None - # One capture rollout at the base seed, two decorrelated repeats. - assert model.seeds == [base, base + 1, base + 2] - - -class _SeedRecordingModel2(_Model): - """Records ``CFG.seed`` at each rollout step (module-level reuse).""" - - def __init__(self, succeed_first_n=10**9): - super().__init__(succeed_first_n=succeed_first_n) - self.seeds = [] - - def get_next_state_and_num_actions(self, state, option): - from predicators.settings import \ - CFG # pylint: disable=import-outside-toplevel - self.seeds.append(CFG.seed) - return super().get_next_state_and_num_actions(state, option) - - -def test_validation_rollouts_arg_raises_the_gate(): - """``validation_rollouts=N`` requests a stricter gate than configured.""" - model = _Model() - text, ctx = _run_tool(model, - rollouts=3, - extra_args={"validation_rollouts": 5}) - assert "Validated 5/5 rollouts" in text - assert ctx.solved_plan is not None - assert model.num_calls == 5 - - -def test_validation_rollouts_arg_cannot_lower_the_gate(): - """A request below the configured gate is ignored: the gate is a floor - - letting the agent lower it would let a lucky draw bypass validation.""" - model = _Model() - text, ctx = _run_tool(model, - rollouts=3, - extra_args={"validation_rollouts": 1}) - assert "Validated 3/3 rollouts" in text - assert ctx.solved_plan is not None - assert model.num_calls == 3 - - -def test_flaky_report_names_seeds_and_reproduction_path(): - """A FLAKY rejection names each rollout's planner seed and tells the agent - how to reproduce the failed rollout (``sim.run(plan, seed=S)``).""" - model = _Model(succeed_first_n=1) - text, _ = _run_tool(model, rollouts=3) - from predicators.settings import \ - CFG # pylint: disable=import-outside-toplevel - base = CFG.seed - assert f"rollout 1 (planner seed {base}): goal reached" in text - assert f"(planner seed {base + 1}): FAILED" in text - assert "sim.run(plan, seed=" in text - - -# --------------------------------------------------------------------------- -# annotation intersection over passing validation rollouts -# --------------------------------------------------------------------------- - -_MidHi = Predicate("MidHi", [_block_type], - lambda s, o: s.get(o[0], "x") >= 0.4) - -_TWO_STEP_PLAN = ("Move(block0:block)[0.5] -> {MidHi(block0:block)}\n" - "Move(block0:block)[0.95] -> {ReachedHi(block0:block)}") - -_NEG_STEP_PLAN = ("Move(block0:block)[0.5] -> {NOT ReachedHi(block0:block)}\n" - "Move(block0:block)[0.95] -> {ReachedHi(block0:block)}") - - -class _RepeatDriftModel(_Model): - """Two-step plans; validation repeats drift the FIRST step's landing. - - Rollout 1 applies each Move's parameter exactly. Later rollouts land - the first step of each pair at ``repeat_first_step_x`` instead, - while the second step still applies its parameter, so repeats reach - the goal (pass) with a different intermediate state. - """ - - def __init__(self, repeat_first_step_x): - super().__init__() - self._repeat_x = repeat_first_step_x - - def get_next_state_and_num_actions(self, state, option): - nxt, n = super().get_next_state_and_num_actions(state, option) - rollout_idx = (self.num_calls - 1) // 2 - step_in_rollout = (self.num_calls - 1) % 2 - if rollout_idx >= 1 and step_in_rollout == 0: - nxt.set(_block, "x", self._repeat_x) - return nxt, n - - -def test_annotation_pruned_when_absent_in_a_passing_repeat(): - """An atom that held in rollout 1 by luck is pruned by the repeats. - - The repeats pass (goal reached), but the intermediate MidHi does not - hold there, so the captured sketch drops it - keeping it would arm - the closed-loop monitor with a divergence the plan does not need. - """ - model = _RepeatDriftModel(repeat_first_step_x=0.3) - utils.reset_config({"agent_plan_validation_rollouts": 3}) - ctx = _make_ctx(model) - ctx.predicates.add(_MidHi) - text = _call_tool(ctx, _TWO_STEP_PLAN) - assert "Captured as the current answer" in text - sketch = ctx.solved_sketch - assert sketch is not None - assert not sketch[0].subgoal_atoms # MidHi pruned - assert {str(a) for a in sketch[1].subgoal_atoms} == \ - {"ReachedHi(block0:block)"} - assert ctx.solved_plan_validation_summary == \ - "validation: 3/3 rollouts ok" - - -def test_annotation_kept_when_only_a_failing_repeat_disagrees(): - """Failing rollouts contribute no evidence to the intersection. - - Repeats fail outright here (goal never reached), so under the best- - effort nudge the flaky capture falls back to the rollout-1 filter - and MidHi survives. - """ - model = _RepeatDriftModel(repeat_first_step_x=0.3) - # Make repeats FAIL: the second step of later rollouts also misses. - orig = _RepeatDriftModel.get_next_state_and_num_actions - - def _failing(self, state, option): - nxt, n = orig(self, state, option) - rollout_idx = (self.num_calls - 1) // 2 - if rollout_idx >= 1: - nxt.set(_block, "x", 0.3) - return nxt, n - - # pylint: disable-next=no-value-for-parameter - model.get_next_state_and_num_actions = _failing.__get__(model) - utils.reset_config({"agent_plan_validation_rollouts": 3}) - ctx = _make_ctx(model, best_effort=True) - ctx.predicates.add(_MidHi) - text = _call_tool(ctx, _TWO_STEP_PLAN) - assert "best-effort" in text - sketch = ctx.solved_sketch - assert sketch is not None - assert {str(a) for a in sketch[0].subgoal_atoms} == \ - {"MidHi(block0:block)"} - assert "first failure: rollout" in ctx.solved_plan_validation_summary - - -def test_negative_annotation_pruned_when_violated_in_a_passing_repeat(): - """The mirrored rule: a NOT atom must be absent in every passing repeat's - post-state to survive.""" - model = _RepeatDriftModel(repeat_first_step_x=0.95) - utils.reset_config({"agent_plan_validation_rollouts": 3}) - ctx = _make_ctx(model) - text = _call_tool(ctx, _NEG_STEP_PLAN) - assert "Captured as the current answer" in text - sketch = ctx.solved_sketch - assert sketch is not None - # NOT ReachedHi held after step 1 of rollout 1 (x=0.5) but is - # violated in the passing repeats (x=0.95), so it is pruned. - assert not sketch[0].subgoal_neg_atoms - - -def test_plan_capture_carries_validation_summary(): - """take_plan_capture surfaces the summary alongside the plan.""" - model = _Model() - _, ctx = _run_tool(model, rollouts=3) - capture = ctx.take_plan_capture() - assert capture.validation_summary == "validation: 3/3 rollouts ok" - assert ctx.solved_plan_validation_summary is None - - -def test_parallel_repeats_capture_robust_plan(): - """With agent_validation_parallel_workers set, the validation repeats run - in forked children: the verdict and note are identical to sequential mode, - and the parent-side model counter proves the repeats did NOT run in this - process (fork isolation).""" - from predicators.agent_sdk.parallel_rollouts import \ - parallel_rollouts_available # pylint: disable=import-outside-toplevel - if not parallel_rollouts_available(): - pytest.skip("fork not available on this platform") - model = _Model() - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_validation_parallel_workers": 2, - }) - ctx = _make_ctx(model) - text = _call_tool(ctx) - assert "Captured as the current answer" in text - assert "Validated 3/3 rollouts" in text - assert ctx.solved_plan is not None - # Only the reported rollout ran in the parent; both repeats ran in - # forked children whose counter increments never propagate back. - assert model.num_calls == 1 - # The parent-side rollout accounting still counts every repeat. - assert ctx.attempt_rollout_count >= 3 - - -def test_parallel_repeats_still_reject_flaky_plan(): - """Child-side failures propagate through the result queue.""" - from predicators.agent_sdk.parallel_rollouts import \ - parallel_rollouts_available # pylint: disable=import-outside-toplevel - if not parallel_rollouts_available(): - pytest.skip("fork not available on this platform") - model = _Model(succeed_first_n=1) - utils.reset_config({ - "agent_plan_validation_rollouts": 3, - "agent_validation_parallel_workers": 2, - }) - ctx = _make_ctx(model) - text = _call_tool(ctx) - assert "FLAKY (plan NOT captured)" in text - assert "Captured as the current answer" not in text - assert ctx.solved_plan is None - - -# ── Execution-verifiability probe ──────────────────────────────────── -# The monitor evaluates captured annotations on real observations, which -# carry no latent. An annotation whose truth in the certifying post-state -# depends on the belief latent (or whose classifier raises without one) -# is excluded from the monitored sketch at capture time. - - -class _LatentModel(_Model): - """Post-states carry a belief latent, as belief-sim rollouts do.""" - - def __init__(self, latent_value, **kwargs): - super().__init__(**kwargs) - self._latent_value = latent_value - - def get_next_state_and_num_actions(self, state, option): - nxt, n = super().get_next_state_and_num_actions(state, option) - nxt.latent = {"_bonds": self._latent_value} - return nxt, n - - -def _latent_only_classifier(s, o, latent=None): - del s, o - return bool((latent or {}).get("_bonds")) - - -_LatentBonded = Predicate("LatentBonded", [_block_type], - _latent_only_classifier) - - -def _raising_without_latent(s, o, latent=None): - del s, o - return bool(latent["_bonds"]) # TypeError when latent is None - - -_RaisingBond = Predicate("RaisingBond", [_block_type], _raising_without_latent) - -_LATENT_PLAN = ("Move(block0:block)[0.95] -> " - "{ReachedHi(block0:block), LatentBonded(block0:block)}") - -_NEG_RAISING_PLAN = ( - "Move(block0:block)[0.5] -> {NOT RaisingBond(block0:block)}\n" - "Move(block0:block)[0.95] -> {ReachedHi(block0:block)}") - - -def test_latent_only_annotation_excluded_from_monitoring(): - """A positive annotation that only holds through the belief latent is - dropped from the captured sketch (it would read false on every real - observation and abort a healthy episode), and the capture message says so; - the observable annotation survives.""" - model = _LatentModel({"a|b"}) - utils.reset_config({"agent_plan_validation_rollouts": 3}) - ctx = _make_ctx(model) - ctx.predicates.add(_LatentBonded) - text = _call_tool(ctx, _LATENT_PLAN) - assert "Captured as the current answer" in text - assert "cannot be verified from a real observation" in text - assert "LatentBonded" in text - sketch = ctx.solved_sketch - assert sketch is not None - assert {str(a) for a in sketch[0].subgoal_atoms} == \ - {"ReachedHi(block0:block)"} - - -def test_raising_negative_annotation_excluded_from_monitoring(): - """A negative annotation whose classifier RAISES without a latent is. - - dropped too - the monitor could not evaluate it on a real state. - """ - # Empty bond set: RaisingBond is False with the latent, so the NOT - # annotation holds in the belief rollout and survives to the probe. - model = _LatentModel(set()) - utils.reset_config({"agent_plan_validation_rollouts": 3}) - ctx = _make_ctx(model) - ctx.predicates.add(_RaisingBond) - # Belief rollouts attach an initial latent to the task init (the - # production path does this via _attach_initial_latent); without it - # the classifier would raise inside the rollout itself. - ctx.current_task.init.latent = {"_bonds": set()} - text = _call_tool(ctx, _NEG_RAISING_PLAN) - assert "Captured as the current answer" in text - assert "cannot be verified from a real observation" in text - assert "RaisingBond" in text - sketch = ctx.solved_sketch - assert sketch is not None - assert not sketch[0].subgoal_neg_atoms - - -def test_observation_backed_annotation_survives_probe(): - """An annotation that holds from observable features alone is kept: - - the probe only drops latent-dependent atoms. - """ - model = _LatentModel({"a|b"}) - utils.reset_config({"agent_plan_validation_rollouts": 3}) - ctx = _make_ctx(model) - text = _call_tool(ctx, _PLAN_TEXT) - assert "Captured as the current answer" in text - assert "cannot be verified from a real observation" not in text - sketch = ctx.solved_sketch - assert sketch is not None - assert {str(a) for a in sketch[0].subgoal_atoms} == \ - {"ReachedHi(block0:block)"} diff --git a/tests/agent_sdk/test_submit_policy_capture.py b/tests/agent_sdk/test_submit_policy_capture.py deleted file mode 100644 index 511e34683a..0000000000 --- a/tests/agent_sdk/test_submit_policy_capture.py +++ /dev/null @@ -1,219 +0,0 @@ -"""Capture-gating tests for the ``submit_policy`` tool (policy mode). - -Mirrors test_submit_plan_capture.py's fixtures: drives the real -MCP handler with a fake option model, covering the policy-mode gates in -front of ``ctx.solved_policy_source`` - multi-rollout validation with -fresh policy memory per rollout, source snapshotting, the -recovered-option-failure semantics, and the mode gate itself. -""" -import asyncio -import os -from typing import Any - -import numpy as np -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk.tools import ToolContext, create_mcp_tools -from predicators.structs import Action, GroundAtom, LowLevelTrajectory, \ - Object, ParameterizedOption, Predicate, State, Task, Type - -_block_type = Type("block", ["x"]) -_block = Object("block0", _block_type) -_ReachedHi = Predicate("ReachedHi", [_block_type], - lambda s, o: s.get(o[0], "x") >= 0.9) - - -def _noop_policy(_s, _m, _o, _p): - return Action(np.zeros(1, dtype=np.float32)) - - -_Move = ParameterizedOption( - "Move", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: False, -) - -_GOAL_POLICY = ''' -def get_option(state, memory): - for obj in state: - if state.get(obj, "x") >= 0.9: - return None - return "Move(block0:block)[0.95]" -''' - - -class _Model: - """Move sets block.x to its parameter; flaky after N calls.""" - - last_execution_failure = None - - def __init__(self, succeed_first_n=10**9): - self.num_calls = 0 - self._succeed_first_n = succeed_first_n - self.last_trajectory = None - - def get_next_state_and_num_actions(self, state, option): - """Roll the option forward one step, counting the call.""" - self.num_calls += 1 - nxt = state.copy() - if self.num_calls <= self._succeed_first_n and len(option.params): - nxt.set(_block, "x", float(option.params[0])) - self.last_trajectory = LowLevelTrajectory( - [state, nxt], [Action(np.zeros(1, dtype=np.float32))]) - return nxt, 1 - - -def _make_ctx(model, sandbox_dir, best_effort=False): - init = State({_block: np.array([0.0], dtype=np.float32)}) - goal = {GroundAtom(_ReachedHi, [_block])} - task = Task(init, goal) - ctx = ToolContext( - types={_block_type}, - predicates={_ReachedHi}, - processes=set(), - options={_Move}, - train_tasks=[task], - example_state=init, - option_model=model, - current_task=task, - sandbox_dir=sandbox_dir, - log_dir=sandbox_dir, - ) - ctx.capture_goal_reaching_plans = True - ctx.capture_best_effort_plan = best_effort - ctx.policy_capture_mode = True - return ctx - - -def _write_policy(sandbox_dir, source): - path = os.path.join(sandbox_dir, "policy.py") - with open(path, "w", encoding="utf-8") as f: - f.write(source) - return path - - -def _call_tool(ctx, extra_args=None): - tools = { - t.name: t.handler - for t in create_mcp_tools(ctx, tool_names=["submit_policy"]) - } - try: - loop = asyncio.get_event_loop() - except RuntimeError: - loop = asyncio.new_event_loop() - asyncio.set_event_loop(loop) - result: Any = loop.run_until_complete(tools["submit_policy"](extra_args - or {})) - return result["content"][0]["text"] - - -def _run_tool(model, - tmp_path, - source=_GOAL_POLICY, - rollouts=3, - best_effort=False, - extra_args=None): - # The local-sandbox path resolution reads /sandbox. - sandbox = os.path.join(str(tmp_path), "sandbox") - os.makedirs(sandbox, exist_ok=True) - utils.reset_config({ - "agent_plan_validation_rollouts": rollouts, - "agent_solve_policy_mode": True, - "agent_sdk_use_local_sandbox": True, - }) - _write_policy(sandbox, source) - ctx = _make_ctx(model, str(tmp_path), best_effort=best_effort) - return _call_tool(ctx, extra_args=extra_args), ctx, sandbox - - -def test_robust_policy_is_captured_with_validation_note(tmp_path): - """All rollouts succeed: source captured with the K/K note.""" - model = _Model() - text, ctx, _ = _run_tool(model, tmp_path, rollouts=3) - assert "Captured policy.py as the current answer" in text - assert "Validated 3/3 rollouts" in text - assert ctx.solved_policy_source is not None - assert "get_option" in ctx.solved_policy_source - assert ctx.solved_plan is None - assert ctx.solved_plan_reached_goal is True - assert ctx.solved_plan_validation_summary == \ - "validation: 3/3 rollouts ok" - - -def test_flaky_policy_is_not_captured(tmp_path): - """A failing validation repeat blocks the capture, loudly.""" - model = _Model(succeed_first_n=1) - text, ctx, _ = _run_tool(model, tmp_path, rollouts=3) - assert "FLAKY (policy NOT captured)" in text - assert ctx.solved_policy_source is None - - -def test_flaky_policy_best_effort_captured(tmp_path): - """Under the final nudge a flaky policy is captured as best-effort.""" - model = _Model(succeed_first_n=1) - text, ctx, _ = _run_tool(model, tmp_path, rollouts=3, best_effort=True) - assert "best-effort" in text - assert ctx.solved_policy_source is not None - assert ctx.solved_plan_reached_goal is False - - -def test_source_snapshot_not_rereading_file(tmp_path): - """Editing policy.py after the call cannot swap unvalidated code.""" - model = _Model() - _, ctx, sandbox = _run_tool(model, tmp_path, rollouts=1) - assert ctx.solved_policy_source is not None - _write_policy(sandbox, "def get_option(state, memory):\n return None\n") - assert "Move(block0:block)" in ctx.solved_policy_source - - -def test_recovered_option_failure_still_captures(tmp_path): - """Closed-loop: a surfaced-and-recovered failure does not disqualify.""" - model = _Model() - orig = _Model.get_next_state_and_num_actions - - def _first_call_fails(self, state, option): - if self.num_calls == 0: - self.num_calls += 1 - self.last_execution_failure = "simulated failure" - return state.copy(), 0 - return orig(self, state, option) - - # pylint: disable-next=no-value-for-parameter - model.get_next_state_and_num_actions = _first_call_fails.__get__(model) - text, ctx, _ = _run_tool(model, tmp_path, rollouts=1) - assert "OPTION FAILURE (surfaced to the policy" in text - assert "Captured policy.py as the current answer" in text - assert ctx.solved_policy_source is not None - - -def test_mode_gate_refuses_outside_policy_mode(tmp_path): - """The tool refuses when the attempt is not in policy mode.""" - model = _Model() - sandbox = os.path.join(str(tmp_path), "sandbox") - os.makedirs(sandbox, exist_ok=True) - utils.reset_config({ - "agent_solve_policy_mode": False, - "agent_sdk_use_local_sandbox": True, - }) - _write_policy(sandbox, _GOAL_POLICY) - ctx = _make_ctx(model, str(tmp_path)) - ctx.policy_capture_mode = False - text = _call_tool(ctx) - assert "only available in policy mode" in text - - -def test_missing_policy_file_is_instructive(tmp_path): - """A missing policy.py errors with writing instructions.""" - model = _Model() - utils.reset_config({ - "agent_solve_policy_mode": True, - "agent_sdk_use_local_sandbox": True, - }) - ctx = _make_ctx(model, str(tmp_path)) - text = _call_tool(ctx) - assert "No ./policy.py found" in text diff --git a/tests/agent_sdk/test_tool_registry.py b/tests/agent_sdk/test_tool_registry.py index 55c6963402..e581cca5db 100644 --- a/tests/agent_sdk/test_tool_registry.py +++ b/tests/agent_sdk/test_tool_registry.py @@ -119,12 +119,12 @@ def test_list_session_tool_names_filters_and_combines() -> None: """Filtered MCP names drop unknowns; ``extra_mcp_tools`` pass through.""" fake = SimpleNamespace(name="run_python") grouped = list_session_tool_names( - mcp_filter=["submit_plan", "not_a_tool", "run_python"], + mcp_filter=["not_a_tool", "run_python"], extra_mcp_tools=[fake], include_builtin=False, ) assert grouped == { - "mcp": ["submit_plan", "run_python"], + "mcp": ["run_python"], "extra": ["run_python"], } @@ -143,13 +143,13 @@ def test_solve_and_synthesis_tool_names_are_independent() -> None: class _Approach(AgentSessionMixin): def _get_solve_tool_names(self) -> Optional[List[str]]: - return ["run_python", "submit_plan"] + return ["run_python", "skills_execute_plan"] def _get_synthesis_tool_names(self) -> Optional[List[str]]: return ["run_python"] obj = _Approach() - assert obj._get_solve_tool_names() == ["run_python", "submit_plan"] + assert obj._get_solve_tool_names() == ["run_python", "skills_execute_plan"] assert obj._get_synthesis_tool_names() == ["run_python"] @@ -158,13 +158,11 @@ def test_get_allowed_tool_list_passes_dynamic_names_through() -> None: list is the single source of truth, with no silent filtering against ``ALL_TOOL_NAMES``.""" allowed = get_allowed_tool_list([ - "submit_plan", # static - "run_python", # dynamic synthesis tool + "run_python", # static, or a dynamic synthesis instance "my_dynamic_tool", # a dynamic tool the roster never lists ]) prefix = f"mcp__{MCP_SERVER_NAME}__" assert allowed == [ - f"{prefix}submit_plan", f"{prefix}run_python", f"{prefix}my_dynamic_tool", ] @@ -312,9 +310,8 @@ def test_agent_render_resolution() -> None: def test_synthesis_tool_names_run_python() -> None: - """Every session offers one ``run_python``: the synthesis roster carries - its own instance (fit data + the candidate-simulator probe in one - namespace), the solve roster the probe over the deployed belief model. + """The synthesis roster carries one ``run_python``, its own instance (fit + data + the candidate-simulator probe in one namespace). Fitting, residual reports, plan validation, and scene work are probe methods (``sim.fit`` / ``sim.residuals`` / ``sim.refine`` / @@ -328,9 +325,7 @@ def test_synthesis_tool_names_run_python() -> None: AgentSimPredicateInventionApproach sim_learn = object.__new__(AgentSimLearningApproach) - sim_learn._do_synthesize_samplers = False invention = object.__new__(AgentSimPredicateInventionApproach) - invention._do_synthesize_samplers = False utils.reset_config({}) names = _required_names(sim_learn._get_synthesis_tool_names()) @@ -339,28 +334,6 @@ def test_synthesis_tool_names_run_python() -> None: assert names.count("run_python") == 1 assert "evaluate_predicate_quality" not in names # sim.predicates() - # On the solve side every arm with a simulator gets the same surface: - # the probe (trajectories in its namespace, sim.task for the task - # digest) plus the submission tool. - utils.reset_config({ - "env": "cover", - "approach": "agent_sim_predicate_invention", - "agent_planner_use_simulator": True, - }) - names = _required_names(invention._get_solve_tool_names()) - assert names.count("run_python") == 1 - assert "submit_plan" in names - - # Without a simulator there is nothing to probe or validate against. - utils.reset_config({ - "env": "cover", - "approach": "agent_sim_predicate_invention", - "agent_planner_use_simulator": False, - }) - names = _required_names(invention._get_solve_tool_names()) - assert "run_python" not in names - assert "submit_plan" not in names - def test_attached_run_python_replaces_the_static_instance() -> None: """A session that attaches its own ``run_python`` (synthesis) gets. diff --git a/tests/agent_sdk/test_trajectory_summary.py b/tests/agent_sdk/test_trajectory_summary.py deleted file mode 100644 index 4d7b9bfbd3..0000000000 --- a/tests/agent_sdk/test_trajectory_summary.py +++ /dev/null @@ -1,110 +0,0 @@ -"""Tests for the ``## Trajectory Summary`` query section. - -The section is the only outcome feedback an agent without a learn phase -gets about its earlier episodes, so it must say what plan ran and how -the env judged it, not only how many steps it took. -""" -from typing import Sequence - -import numpy as np -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk.sketch_prompts import summarize_trajectories -from predicators.structs import Action, GroundAtom, LowLevelTrajectory, \ - Object, ParameterizedOption, Predicate, State, Task, Type - -_block_type = Type("block", ["x"]) -_block = Object("block0", _block_type) -_Far = Predicate("Far", [_block_type], lambda s, o: s.get(o[0], "x") > 0.5) -_Move = ParameterizedOption( - "Move", - types=[_block_type], - params_space=Box(0.0, 1.0, (1, )), - policy=lambda s, m, o, p: Action(np.zeros(1, dtype=np.float32)), - initiable=lambda s, m, o, p: True, - terminal=lambda s, m, o, p: True, -) -_Wait = ParameterizedOption( - "Wait", - types=[], - params_space=Box(0.0, 1.0, (0, )), - policy=lambda s, m, o, p: Action(np.zeros(1, dtype=np.float32)), - initiable=lambda s, m, o, p: True, - terminal=lambda s, m, o, p: True, -) - - -def _state(x: float) -> State: - return State({_block: np.array([x], dtype=np.float32)}) - - -def _traj(xs: Sequence[float], options: Sequence[str], - **kwargs) -> LowLevelTrajectory: - states = [_state(x) for x in xs] - actions = [] - for name in options: - act = Action(np.zeros(1, dtype=np.float32)) - if name == "Move": - act.set_option(_Move.ground([_block], np.array([0.5]))) - elif name == "Wait": - act.set_option(_Wait.ground([], np.array([]))) - actions.append(act) - return LowLevelTrajectory(states, actions, _train_task_idx=0, **kwargs) - - -def _setup() -> None: - utils.reset_config({"agent_sdk_max_trajectories_in_context": 2}) - - -def test_summary_reports_plan_verdict_and_stable_numbers() -> None: - """Each recent trajectory shows the option plan it executed (repeats - collapsed), the env's reward and goal verdict, and keeps its global index - so a number means the same episode in every session.""" - _setup() - trajs = [ - _traj([0.0, 0.0], ["Move"], _env_reward=0.0, _env_terminated=False), - _traj([0.0, 0.2, 0.9], ["Move", "Move"], - _env_reward=0.0, - _env_terminated=False), - _traj([0.0, 0.9, 0.9], ["Move", "Wait"], - _env_reward=1.0, - _env_terminated=True), - ] - text = summarize_trajectories(trajs, {_Far}) - assert "(3 total, showing last 2)" in text - assert "Trajectory 0:" not in text - assert "Trajectory 1: 2 steps" in text - assert "Executed: Move(block0)\n" in text - assert "Trajectory 2: 2 steps" in text - assert "Executed: Move(block0) -> Wait()" in text - assert "Outcome: env reward 0.00, goal NOT reached" in text - assert "Outcome: env reward 1.00, goal atoms held at the end" in text - assert "Gained: Far(block0:block)" in text - - -def test_summary_falls_back_to_the_task_goal_without_a_verdict() -> None: - """A trajectory the env never evaluated (no reward) still gets a goal - verdict from the train task's goal atoms when the tasks are given.""" - _setup() - task = Task(_state(0.0), {GroundAtom(_Far, [_block])}) - solved = _traj([0.0, 0.9], ["Move"]) - unsolved = _traj([0.0, 0.1], ["Move"]) - text = summarize_trajectories([solved, unsolved], {_Far}, - train_tasks=[task]) - assert text.count("Outcome: goal atoms held at the end") == 1 - assert text.count("Outcome: goal NOT reached") == 1 - # Without tasks there is nothing to judge against: no outcome line. - assert "Outcome" not in summarize_trajectories([solved], {_Far}) - - -def test_summary_without_option_tags_omits_the_plan_line() -> None: - """Raw-action trajectories (no option on the actions) keep the old step- - count and atom-delta lines and simply have no plan to show.""" - _setup() - traj = LowLevelTrajectory([_state(0.0), _state(0.9)], - [Action(np.zeros(1, dtype=np.float32))]) - text = summarize_trajectories([traj], {_Far}) - assert "Trajectory 0: 1 steps" in text - assert "Executed" not in text - assert "Gained: Far(block0:block)" in text diff --git a/tests/approaches/test_agent_continual_direct_scene_files.py b/tests/approaches/test_agent_continual_direct_scene_files.py index a806cd3951..ed0ca473d2 100644 --- a/tests/approaches/test_agent_continual_direct_scene_files.py +++ b/tests/approaches/test_agent_continual_direct_scene_files.py @@ -20,11 +20,12 @@ # pylint: disable=protected-access -CONFIG = "predicatorv3/continual_direct_scene_files_benchmark_r1.yaml" +CONFIG = "empiric/benchmark.yaml" +ARM = "mf_scene_package_opus" def _arm_flags(env_name: str) -> Dict[str, Any]: - cfg = next(c for c in generate_run_configs(CONFIG, False) + cfg = next(c for c in generate_run_configs(CONFIG, False, approaches=[ARM]) if c.env == env_name) # The launcher pins machine-specific output paths; tests keep their own. flags = { @@ -50,10 +51,10 @@ def _make(tmp_path: Any, **overrides: Any) -> Any: def test_config_is_the_direct_agent_with_the_scene_files() -> None: - """Five settings, three seeds, the MF arm with the package flag and none of + """Five settings, five seeds, the MF arm with the package flag and none of the model arm's flags.""" - runs = list(generate_run_configs(CONFIG, False)) - assert len(runs) == 15 + runs = list(generate_run_configs(CONFIG, False, approaches=[ARM])) + assert len(runs) == 25 for run in runs: assert run.approach == "agent_continual_model_free" assert run.flags["continual_provide_scene_package"] is True diff --git a/tests/approaches/test_agent_continual_real_to_sim_approach.py b/tests/approaches/test_agent_continual_real_to_sim_approach.py index 541db18424..2e53e009d4 100644 --- a/tests/approaches/test_agent_continual_real_to_sim_approach.py +++ b/tests/approaches/test_agent_continual_real_to_sim_approach.py @@ -20,7 +20,25 @@ # pylint: disable=protected-access -CONFIG = "predicatorv3/continual_real_to_sim_benchmark_r1.yaml" +# The agentic real-to-sim arm (no domain twin: the generic PyBulletEnv, a +# domain-agnostic SceneBase, the scene manifest and the asset files; the +# harness fits nothing and runs no uncertainty machinery) and the EMPIRIC +# from-assets arm it derives from. Neither is a benchmark arm; each is the +# benchmark's EMPIRIC run with these flags. +REAL_TO_SIM_FLAGS = { + "agent_sim_learn_declared_params_only": True, + "continual_uncertainty_decisions": False, + "agent_sim_learn_param_uncertainty": False, + "agent_explorer_info_seeking": False, + "agent_explorer_info_seeking_adaptive": False, + "agent_explorer_info_seeking_noise_aware": False, + "code_sim_learning_interval_belief": False, + "code_sim_learning_carry_posterior": False, +} +FROM_ASSETS_FLAGS = { + "agent_sim_learn_declared_params_only": False, + "continual_uncertainty_decisions": True, +} # The agent's simulator for the Boil training scene, written the way the # prompt describes: a SceneBase subclass whose initialize_pybullet loads # every body of the manifest under its observed name. @@ -115,16 +133,23 @@ def connect(*args: Any, **kwargs: Any) -> int: p.disconnect(client) -def _arm_flags() -> Dict[str, Any]: - cfg = next(c for c in generate_run_configs(CONFIG, False) - if c.env == "pybullet_boil") - # The launcher pins machine-specific output paths; tests keep their own. - flags = { +def _benchmark_flags(env_name: str) -> Dict[str, Any]: + """The benchmark EMPIRIC run's flags on ``env_name``, without the machine- + specific output paths (tests keep their own).""" + cfg = next(c for c in generate_run_configs( + "empiric/benchmark.yaml", False, approaches=["mb_opus"]) + if c.env == env_name) + return { k: v for k, v in cfg.flags.items() if k not in ("log", "continual_runs_dir") } - flags.update(approach=cfg.approach, - env=cfg.env, + + +def _arm_flags() -> Dict[str, Any]: + flags = _benchmark_flags("pybullet_boil") + flags.update(REAL_TO_SIM_FLAGS, + approach="agent_continual_real_to_sim", + env="pybullet_boil", continual_render=False, continual_make_video=False) return flags @@ -134,17 +159,10 @@ def _make(tmp_path: Any, from_assets: bool = False) -> Any: flags = _arm_flags() name = "agent_continual_real_to_sim" if from_assets: - cfg = next(c for c in generate_run_configs( - "predicatorv3/continual_from_assets_pilot_r1.yaml", False) - if c.env == "pybullet_bridge") - # The launcher pins machine-specific output paths; tests keep their own. - flags = { - k: v - for k, v in cfg.flags.items() - if k not in ("log", "continual_runs_dir") - } name = "agent_continual_from_assets" - flags.update(approach=name, + flags = _benchmark_flags("pybullet_bridge") + flags.update(FROM_ASSETS_FLAGS, + approach=name, env="pybullet_boil", continual_render=False, continual_make_video=False) @@ -158,17 +176,6 @@ def _make(tmp_path: Any, from_assets: bool = False) -> Any: return env, approach -def test_config_is_the_benchmark_arm() -> None: - """Five settings, three seeds, no fitting and no uncertainty flags.""" - runs = list(generate_run_configs(CONFIG, False)) - assert len(runs) == 15 - for run in runs: - assert run.approach == "agent_continual_real_to_sim" - assert run.flags["agent_sim_learn_declared_params_only"] is True - assert run.flags["continual_uncertainty_decisions"] is False - assert run.flags["continual_require_model_on_test"] is True - - def test_arm_refuses_uncertainty_machinery(tmp_path: Any) -> None: """A launcher that leaves an uncertainty switch on fails at construction.""" @@ -346,30 +353,3 @@ def query(message: str, *_args: Any, **_kwargs: Any) -> Any: fit = approach._last_fit_result assert fit is not None and fit.names == ["heat_rate"] assert fit.samples[0, 0] == pytest.approx(0.01) - - -def test_from_assets_pilot_retains_empiric_capabilities() -> None: - """Two seeds per domain, fitting and EMPIRIC's joint belief, no - preflight.""" - runs = list( - generate_run_configs( - "predicatorv3/continual_from_assets_pilot_r1.yaml", False)) - assert len(runs) == 10 - assert {r.env - for r in runs} == { - "pybullet_fan", "pybullet_bridge", "pybullet_domino", - "pybullet_balloons", "pybullet_boil" - } - for run in runs: - assert run.approach == "agent_continual_from_assets" - assert run.flags["agent_sim_learn_declared_params_only"] is False - assert run.flags["continual_uncertainty_decisions"] is True - assert run.flags["code_sim_learning_interval_belief"] is True - # The joint belief, with the prior at its declared centre. - assert run.flags["belief_joint_draws"] == 16 - assert run.flags["code_sim_learning_carry_posterior"] is False - assert run.flags["continual_skill_preflight"] is False - if run.env == "pybullet_fan": - assert run.flags["fan_ramp_transfer"] is True - assert run.flags["fan_ramp_rise"] == 0.003 - assert run.flags["fan_ramp_landing_extension"] == 0.10 diff --git a/tests/approaches/test_agent_continual_scene_package.py b/tests/approaches/test_agent_continual_scene_package.py index 05454418d4..803071c706 100644 --- a/tests/approaches/test_agent_continual_scene_package.py +++ b/tests/approaches/test_agent_continual_scene_package.py @@ -20,18 +20,27 @@ # pylint: disable=protected-access -CONFIG = "predicatorv3/continual_empiric_scene_package_benchmark_r1.yaml" +# EMPIRIC with what an agentic real-to-sim agent receives: the engine +# wrapper, the scene manifest and the URDF and mesh files, plus the twin's +# own core module where the domain declares one. Not a benchmark arm; it +# is the benchmark's EMPIRIC arm with these two flags. +SCENE_PACKAGE_FLAGS = { + "agent_sim_provide_base_sim_source": True, + "continual_provide_scene_package": True, +} def _arm_flags(env_name: str) -> Dict[str, Any]: - cfg = next(c for c in generate_run_configs(CONFIG, False) + cfg = next(c for c in generate_run_configs( + "empiric/benchmark.yaml", False, approaches=["mb_opus"]) if c.env == env_name) # The launcher pins machine-specific output paths; tests keep their own. flags = { k: v for k, v in cfg.flags.items() if k not in ("log", "continual_runs_dir") } - flags.update(approach=cfg.approach, + flags.update(SCENE_PACKAGE_FLAGS, + approach=cfg.approach, env=cfg.env, continual_render=False, continual_make_video=False) @@ -55,18 +64,6 @@ def _refs(sandbox: Path) -> List[str]: for q in (sandbox / "reference").rglob("*") if q.is_file()) -def test_config_is_empiric_with_the_scene_package() -> None: - """Five settings, three seeds, the MB arm with both reference flags.""" - runs = list(generate_run_configs(CONFIG, False)) - assert len(runs) == 15 - for run in runs: - assert run.approach == "agent_continual" - assert run.flags["continual_provide_scene_package"] is True - assert run.flags["agent_sim_provide_base_sim_source"] is True - assert run.flags["continual_require_model_on_test"] is True - assert "agent_sim_learn_declared_params_only" not in run.flags - - def test_fan_lists_the_twin_core_and_the_package(tmp_path: Any) -> None: """A domain with a declared core module gets it beside the package.""" _, approach = _make(tmp_path, "pybullet_fan") diff --git a/tests/approaches/test_agent_model_based_approach.py b/tests/approaches/test_agent_model_based_approach.py deleted file mode 100644 index d0c7683aaf..0000000000 --- a/tests/approaches/test_agent_model_based_approach.py +++ /dev/null @@ -1,1919 +0,0 @@ -"""Tests for AgentModelBasedApproach -- parsing and refinement logic.""" -# pylint: disable=protected-access,import-outside-toplevel -import os -from unittest.mock import MagicMock, patch - -import numpy as np -import pytest -from gym.spaces import Box - -from predicators import utils -from predicators.approaches.agent_model_based_approach import \ - AgentModelBasedApproach, _SketchStep -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, Type - -_TEST_DATA_DIR = os.path.join(os.path.dirname(__file__), "test_data") - -# --------------------------------------------------------------------------- -# Shared fixtures -# --------------------------------------------------------------------------- - -_block_type = Type("block", ["x", "y", "held"]) -_robot_type = Type("robot", ["x", "y"]) - -_block0 = Object("block0", _block_type) -_block1 = Object("block1", _block_type) -_robot = Object("robot0", _robot_type) - -_Holding = Predicate("Holding", [_block_type], - lambda s, o: s.get(o[0], "held") > 0.5) -_On = Predicate("On", [_block_type, _block_type], - lambda s, o: abs(s.get(o[0], "x") - s.get(o[1], "x")) < 0.1) -_HandEmpty = Predicate("HandEmpty", [_robot_type], lambda s, o: True) - -_ALL_PREDICATES = {_Holding, _On, _HandEmpty} -_ALL_OBJECTS = [_block0, _block1, _robot] - - -def _noop_policy(_s, _m, _o, _p): - return Action(np.zeros(1, dtype=np.float32)) - - -def _always_true(_s, _m, _o, _p): - return True - - -def _always_false(_s, _m, _o, _p): - return False - - -_Pick = ParameterizedOption( - "Pick", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_false, -) - -_Place = ParameterizedOption( - "Place", - types=[_block_type, _block_type], - params_space=Box(low=np.array([0.0, 0.0], dtype=np.float32), - high=np.array([1.0, 1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_false, -) - -_Wait = ParameterizedOption( - "Wait", - types=[_robot_type], - params_space=Box(low=np.array([], dtype=np.float32), - high=np.array([], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_false, -) - -_ALL_OPTIONS = {_Pick, _Place, _Wait} - - -def _make_state(overrides=None): - """Create a simple state with default feature values.""" - data = { - _block0: np.array([0.1, 0.2, 0.0], dtype=np.float32), - _block1: np.array([0.5, 0.6, 0.0], dtype=np.float32), - _robot: np.array([0.0, 0.0], dtype=np.float32), - } - if overrides: - for obj, vals in overrides.items(): - data[obj] = np.array(vals, dtype=np.float32) - return State(data) - - -def _make_approach(): - """Create an AgentModelBasedApproach with mock config and option model.""" - state = _make_state() - goal = {GroundAtom(_On, [_block0, _block1])} - task = Task(state, goal) - - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "option_model_name": "oracle", - "seed": 42, - "agent_bilevel_max_samples_per_step": 10, - "agent_bilevel_check_subgoals": True, - }) - - mock_option_model = MagicMock() - approach = AgentModelBasedApproach( - initial_predicates=_ALL_PREDICATES, - initial_options=_ALL_OPTIONS, - types={_block_type, _robot_type}, - action_space=Box(low=-1, high=1, shape=(1, )), - train_tasks=[task], - option_model=mock_option_model, - ) - return approach, mock_option_model, task - - -# --------------------------------------------------------------------------- -# Tests: _parse_subgoal_annotations -# --------------------------------------------------------------------------- - - -class TestParseSubgoalAnnotations: - """Tests for plan text subgoal parsing.""" - - def test_basic_subgoals(self): - """Test basic subgoals.""" - approach, _, _ = _make_approach() - text = ("Pick(block0:block) -> {Holding(block0:block)}\n" - "Place(block0:block, block1:block) -> " - "{On(block0:block, block1:block)}\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 2 - # First step: Holding(block0) - assert result[0] is not None - pos, neg = result[0] - assert GroundAtom(_Holding, [_block0]) in pos - assert len(neg) == 0 - # Second step: On(block0, block1) - assert result[1] is not None - pos2, neg2 = result[1] - assert GroundAtom(_On, [_block0, _block1]) in pos2 - assert len(neg2) == 0 - - def test_no_subgoals(self): - """Test no subgoals.""" - approach, _, _ = _make_approach() - text = ("Pick(block0:block)\n" - "Place(block0:block, block1:block)\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 2 - assert result[0] is None - assert result[1] is None - - def test_mixed_subgoals(self): - """Some lines have subgoals, some don't.""" - approach, _, _ = _make_approach() - text = ("Pick(block0:block) -> {Holding(block0:block)}\n" - "Wait(robot0:robot)\n" - "Place(block0:block, block1:block) -> " - "{On(block0:block, block1:block)}\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 3 - assert result[0] is not None - assert result[1] is None # Wait has no subgoal - assert result[2] is not None - - def test_multiple_atoms_in_subgoal(self): - """Test multiple atoms in subgoal.""" - approach, _, _ = _make_approach() - text = ( - "Place(block0:block, block1:block) " - "-> {On(block0:block, block1:block), HandEmpty(robot0:robot)}\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is not None - pos, neg = result[0] - assert len(pos) == 2 - assert len(neg) == 0 - assert GroundAtom(_On, [_block0, _block1]) in pos - assert GroundAtom(_HandEmpty, [_robot]) in pos - - def test_unknown_predicate_skipped(self): - """Test unknown predicate skipped.""" - approach, _, _ = _make_approach() - text = "Pick(block0:block) -> {FakePred(block0:block)}\n" - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is None # FakePred unrecognized, no valid atoms - - def test_unknown_object_skipped(self): - """Test unknown object skipped.""" - approach, _, _ = _make_approach() - text = "Pick(block0:block) -> {Holding(block99:block)}\n" - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is None # block99 doesn't exist - - def test_arity_mismatch_skipped(self): - """Test arity mismatch skipped.""" - approach, _, _ = _make_approach() - # Holding expects 1 arg, giving 2 - text = "Pick(block0:block) -> {Holding(block0:block, block1:block)}\n" - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is None - - def test_typed_object_refs_in_subgoals(self): - """Agent outputs obj:type in subgoal atoms — should still parse.""" - approach, _, _ = _make_approach() - text = ("Pick(block0:block) -> {Holding(block0:block)}\n" - "Place(block0:block, block1:block) " - "-> {On(block0:block, block1:block)}\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 2 - assert result[0] is not None - pos, _ = result[0] - assert GroundAtom(_Holding, [_block0]) in pos - assert result[1] is not None - pos2, _ = result[1] - assert GroundAtom(_On, [_block0, _block1]) in pos2 - - def test_numbered_prefix_subgoals(self): - """Agent numbers the lines (0:, 1:) — annotations must still align. - - Mirrors a real failure: the agent mirrored the numbered sketch - format shown in logs, embedding it between prose, and the - numbered prefix made every line parse as a non-option line so - the annotation list came back empty/misaligned. - """ - approach, _, _ = _make_approach() - text = ("Some analysis the agent wrote first.\n" - " 0: Pick(block0:block) -> {Holding(block0:block)}\n" - " 1: Place(block0:block, block1:block) " - "-> {On(block0:block, block1:block)}\n" - "Rationale: ...\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 2 - assert result[0] is not None - pos, _ = result[0] - assert GroundAtom(_Holding, [_block0]) in pos - assert result[1] is not None - pos2, _ = result[1] - assert GroundAtom(_On, [_block0, _block1]) in pos2 - - def test_preamble_ignored(self): - """Non-option lines should be ignored.""" - approach, _, _ = _make_approach() - text = ("Here is my analysis:\n" - "I think we should pick block0 first.\n" - "\n" - "Pick(block0:block) -> {Holding(block0:block)}\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is not None - - def test_whitespace_in_atoms(self): - """Spaces around commas in atom arguments.""" - approach, _, _ = _make_approach() - text = ("Place(block0:block, block1:block) -> " - "{ On( block0:block , block1:block ) }\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is not None - pos, _ = result[0] - assert GroundAtom(_On, [_block0, _block1]) in pos - - def test_not_atoms_in_subgoals(self): - """Test NOT prefix for negative target atoms.""" - approach, _, _ = _make_approach() - text = ( - "Wait(robot0:robot) -> " - "{Holding(block0:block), NOT On(block0:block, block1:block)}\n") - result = approach._parse_subgoal_annotations(text, _ALL_PREDICATES, - _ALL_OBJECTS) - - assert len(result) == 1 - assert result[0] is not None - pos, neg = result[0] - assert GroundAtom(_Holding, [_block0]) in pos - assert GroundAtom(_On, [_block0, _block1]) in neg - - -# --------------------------------------------------------------------------- -# Tests: check_wait_target_atoms -# --------------------------------------------------------------------------- - - -class TestCheckWaitTargetAtoms: - """Tests that Wait terminates on target atoms, not noisy changes.""" - - def test_no_targets_returns_none(self): - """No targets in memory -> returns None (fall back to any-change).""" - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - # No targets in memory - state = _make_state({_block0: [0.0, 0.0, 0.0]}) - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - result = utils.check_wait_target_atoms(opt, state, abstract_fn) - assert result is None - - def test_positive_target_met(self): - """Wait terminates when positive target atom holds.""" - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - target_atom = GroundAtom(_Holding, [_block0]) - opt.memory["wait_target_atoms"] = {target_atom} - - # State where Holding(block0) is true (held > 0.5) - state_held = _make_state({_block0: [0.0, 0.0, 1.0]}) - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - assert utils.check_wait_target_atoms(opt, state_held, abstract_fn) \ - is True - - def test_positive_target_not_met(self): - """Wait does NOT terminate when target atom doesn't hold yet.""" - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - target_atom = GroundAtom(_Holding, [_block0]) - opt.memory["wait_target_atoms"] = {target_atom} - - # State where Holding(block0) is false (held <= 0.5) - state_not_held = _make_state({_block0: [0.0, 0.0, 0.0]}) - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - assert utils.check_wait_target_atoms(opt, state_not_held, - abstract_fn) is False - - def test_noisy_atom_change_ignored_with_targets(self): - """Wait ignores noisy atom changes when specific targets are set. - - This is the key test: if the Wait is parameterized with a target - atom (e.g. Holding(block0)), it should NOT terminate when a - different atom changes (e.g. On(block0, block1)). - """ - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - # Only waiting for Holding(block0) - target_atom = GroundAtom(_Holding, [_block0]) - opt.memory["wait_target_atoms"] = {target_atom} - - # State where On(block0, block1) is true (noisy change) but - # Holding(block0) is still false - state_noisy = _make_state({ - _block0: [0.5, 0.0, 0.0], - _block1: [0.5, 0.0, 0.0] - }) - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - atoms = abstract_fn(state_noisy) - # On is true (positions are close), but Holding is false - assert GroundAtom(_On, [_block0, _block1]) in atoms - assert GroundAtom(_Holding, [_block0]) not in atoms - - # Wait should NOT terminate (target not met, despite On changing) - assert utils.check_wait_target_atoms(opt, state_noisy, - abstract_fn) is False - - def test_negative_target_met(self): - """Wait terminates when negative target atom is false.""" - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - neg_atom = GroundAtom(_On, [_block0, _block1]) - opt.memory["wait_target_neg_atoms"] = {neg_atom} - - # State where On(block0, block1) is false (positions far apart) - state = _make_state({ - _block0: [0.0, 0.0, 0.0], - _block1: [5.0, 0.0, 0.0] - }) - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - assert utils.check_wait_target_atoms(opt, state, abstract_fn) is True - - def test_negative_target_not_met(self): - """Wait does NOT terminate when negative target atom is still true.""" - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - neg_atom = GroundAtom(_On, [_block0, _block1]) - opt.memory["wait_target_neg_atoms"] = {neg_atom} - - # State where On(block0, block1) is true (positions close) - state = _make_state({ - _block0: [0.5, 0.0, 0.0], - _block1: [0.5, 0.0, 0.0] - }) - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - assert utils.check_wait_target_atoms(opt, state, abstract_fn) is False - - def test_mixed_positive_and_negative_targets(self): - """Both positive and negative targets must be satisfied.""" - opt = _Wait.ground([_robot], np.array([], dtype=np.float32)) - opt.memory["wait_target_atoms"] = {GroundAtom(_Holding, [_block0])} - opt.memory["wait_target_neg_atoms"] = { - GroundAtom(_On, [_block0, _block1]) - } - - abstract_fn = lambda s: utils.abstract(s, _ALL_PREDICATES) - - # Only positive met (Holding true, On still true) - state1 = _make_state({ - _block0: [0.5, 0.0, 1.0], - _block1: [0.5, 0.0, 0.0] - }) - assert utils.check_wait_target_atoms(opt, state1, abstract_fn) is False - - # Only negative met (On false, Holding false) - state2 = _make_state({ - _block0: [0.0, 0.0, 0.0], - _block1: [5.0, 0.0, 0.0] - }) - assert utils.check_wait_target_atoms(opt, state2, abstract_fn) is False - - # Both met (Holding true, On false) - state3 = _make_state({ - _block0: [0.0, 0.0, 1.0], - _block1: [5.0, 0.0, 0.0] - }) - assert utils.check_wait_target_atoms(opt, state3, abstract_fn) is True - - -# --------------------------------------------------------------------------- -# Tests: parse_wait_target_annotations and strip_wait_annotations -# --------------------------------------------------------------------------- - - -class TestWaitTargetParsing: - """Tests for parse_wait_target_annotations and strip_wait_annotations.""" - - def test_parse_positive_target(self): - """Parse a positive target atom.""" - line = "Wait(robot0:robot) -> {Holding(block0:block)}" - pos, neg = utils.parse_wait_target_annotations(line, _ALL_PREDICATES, - _ALL_OBJECTS) - assert GroundAtom(_Holding, [_block0]) in pos - assert len(neg) == 0 - - def test_parse_negative_target(self): - """Parse a NOT-prefixed target atom.""" - line = "Wait(robot0:robot) -> {NOT On(block0:block, block1:block)}" - pos, neg = utils.parse_wait_target_annotations(line, _ALL_PREDICATES, - _ALL_OBJECTS) - assert len(pos) == 0 - assert GroundAtom(_On, [_block0, _block1]) in neg - - def test_parse_mixed_targets(self): - """Parse both positive and negative target atoms.""" - line = ("Wait(robot0:robot) -> " - "{Holding(block0:block), NOT On(block0:block, block1:block)}") - pos, neg = utils.parse_wait_target_annotations(line, _ALL_PREDICATES, - _ALL_OBJECTS) - assert GroundAtom(_Holding, [_block0]) in pos - assert GroundAtom(_On, [_block0, _block1]) in neg - - def test_parse_no_annotation(self): - """Line without -> returns empty sets.""" - line = "Wait(robot0:robot)[]" - pos, neg = utils.parse_wait_target_annotations(line, _ALL_PREDICATES, - _ALL_OBJECTS) - assert len(pos) == 0 - assert len(neg) == 0 - - def test_strip_annotations(self): - """strip_wait_annotations removes -> {...} suffixes.""" - text = ("Pick(block0:block)[0.5]\n" - "Wait(robot0:robot)[] -> {Holding(block0:block)}\n" - "Place(block0:block, block1:block)[0.1, 0.2]\n") - stripped = utils.strip_wait_annotations(text) - assert "-> {" not in stripped - assert "Pick(block0:block)[0.5]" in stripped - assert "Wait(robot0:robot)[]" in stripped - assert "Place(block0:block, block1:block)[0.1, 0.2]" in stripped - - -# --------------------------------------------------------------------------- -# Tests: _refine_sketch -# --------------------------------------------------------------------------- - - -class TestRefineSketch: - """Tests for backtracking refinement search.""" - - def test_empty_sketch(self): - """Test empty sketch.""" - approach, _, task = _make_approach() - plan, success = approach._refine_sketch(task, [], timeout=5.0) - assert plan == [] - assert success is False - - def test_single_step_no_params(self): - """Option with empty params_space — should succeed in 1 try.""" - approach, mock_om, task = _make_approach() - - # Option model: Wait always succeeds, goal holds after - goal_state = _make_state({_block0: [0.5, 0.6, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (goal_state, 5) - - sketch = [ - _SketchStep(option=_Wait, objects=[_robot], subgoal_atoms=None) - ] - plan, success = approach._refine_sketch(task, sketch, timeout=5.0) - - assert success is True - assert len(plan) == 1 - assert plan[0].name == "Wait" - - def test_single_step_with_params_success(self): - """Option with params — should find working params via sampling.""" - approach, mock_om, task = _make_approach() - - goal_state = _make_state({_block0: [0.5, 0.6, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (goal_state, 3) - - sketch = [ - _SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None) - ] - plan, success = approach._refine_sketch(task, sketch, timeout=5.0) - - assert success is True - assert len(plan) == 1 - - def test_subgoal_check_pass(self): - """Subgoal atoms hold after execution.""" - approach, mock_om, task = _make_approach() - - # After Pick, Holding(block0) should hold — set held=1 - held_state = _make_state({_block0: [0.1, 0.2, 1.0]}) - # After Place, On(block0, block1) — set x close - goal_state = _make_state({_block0: [0.5, 0.6, 0.0]}) - - mock_om.get_next_state_and_num_actions.side_effect = [ - (held_state, 3), - (goal_state, 3), - ] - - sketch = [ - _SketchStep(option=_Pick, - objects=[_block0], - subgoal_atoms={GroundAtom(_Holding, [_block0])}), - _SketchStep(option=_Place, - objects=[_block0, _block1], - subgoal_atoms={GroundAtom(_On, [_block0, _block1])}), - ] - plan, success = approach._refine_sketch(task, sketch, timeout=5.0) - - assert success is True - assert len(plan) == 2 - - def test_subgoal_check_fail_triggers_resample(self): - """Subgoal atoms don't hold — should resample params.""" - approach, mock_om, task = _make_approach() - - # Holding never holds (held=0) — subgoal always fails - bad_state = _make_state({_block0: [0.1, 0.2, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (bad_state, 3) - - sketch = [ - _SketchStep(option=_Pick, - objects=[_block0], - subgoal_atoms={GroundAtom(_Holding, [_block0])}), - ] - _plan, success = approach._refine_sketch(task, sketch, timeout=5.0) - - # Should exhaust all samples and fail - assert success is False - # Option model called max_samples times (10) - assert mock_om.get_next_state_and_num_actions.call_count == 10 - - def test_backtracking_across_steps(self): - """Step 2 fails, causing step 1 to be re-sampled.""" - approach, mock_om, task = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "seed": 42, - "agent_bilevel_max_samples_per_step": 3, - "agent_bilevel_check_subgoals": False, - }) - - call_count = 0 - goal_state = _make_state({_block0: [0.5, 0.6, 0.0]}) - noop_state = _make_state() - - def side_effect(_state, option): - nonlocal call_count - call_count += 1 - if option.name == "Pick": - return (noop_state, 3) # Pick always succeeds - # Place: succeed only on the last attempt - if call_count >= 8: - return (goal_state, 3) - return (noop_state, 0) # fail (noop) - - mock_om.get_next_state_and_num_actions.side_effect = side_effect - - sketch = [ - _SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None), - _SketchStep(option=_Place, - objects=[_block0, _block1], - subgoal_atoms=None), - ] - plan, success = approach._refine_sketch(task, sketch, timeout=10.0) - - # Should have backtracked and eventually succeeded - assert success is True - assert len(plan) == 2 - assert call_count >= 4 # at least one backtrack cycle - - def test_not_initiable_triggers_resample(self): - """Option not initiable in current state — resample.""" - approach, mock_om, task = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "seed": 42, - "agent_bilevel_max_samples_per_step": 3, - }) - - # Create an option that is never initiable - not_initiable = ParameterizedOption( - "Pick", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_false, - terminal=_always_false, - ) - - sketch = [ - _SketchStep(option=not_initiable, - objects=[_block0], - subgoal_atoms=None) - ] - _plan, success = approach._refine_sketch(task, sketch, timeout=5.0) - - assert success is False - # Option model never called since initiable is always False - mock_om.get_next_state_and_num_actions.assert_not_called() - - def test_goal_check_on_final_step(self): - """Final step must satisfy the task goal even without subgoals.""" - approach, mock_om, task = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "seed": 42, - "agent_bilevel_max_samples_per_step": 5, - "agent_bilevel_check_subgoals": False, - }) - - # State that doesn't satisfy goal On(block0, block1) - bad_state = _make_state({_block0: [0.9, 0.2, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (bad_state, 3) - - sketch = [ - _SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None) - ] - _plan, success = approach._refine_sketch(task, sketch, timeout=5.0) - - # Goal never holds → exhausts samples - assert success is False - - -# --------------------------------------------------------------------------- -# Tests: _query_agent_for_plan_sketch (with mocked agent) -# --------------------------------------------------------------------------- - - -class TestQueryAgentForPlanSketch: - """Tests for end-to-end sketch extraction from mock agent responses.""" - - def _mock_responses(self, plan_text): - """Build mock agent response list containing plan_text.""" - return [ - { - "type": "assistant", - "content": [{ - "type": "text", - "text": plan_text - }], - }, - ] - - def test_basic_sketch_extraction(self): - """Test basic sketch extraction.""" - approach, _, task = _make_approach() - - plan_text = ("Pick(block0:block) -> {Holding(block0:block)}\n" - "Place(block0:block, block1:block) -> " - "{On(block0:block, block1:block)}\n") - - with patch.object(approach, - '_query_agent_sync', - return_value=self._mock_responses(plan_text)): - sketch = approach._query_agent_for_plan_sketch(task) - - assert len(sketch) == 2 - assert sketch[0].option.name == "Pick" - assert list(sketch[0].objects) == [_block0] - assert sketch[0].subgoal_atoms is not None - assert GroundAtom(_Holding, [_block0]) in sketch[0].subgoal_atoms - - assert sketch[1].option.name == "Place" - assert list(sketch[1].objects) == [_block0, _block1] - assert sketch[1].subgoal_atoms is not None - - def test_sketch_without_subgoals(self): - """Test sketch without subgoals.""" - approach, _, task = _make_approach() - - plan_text = ("Pick(block0:block)\n" - "Place(block0:block, block1:block)\n") - - with patch.object(approach, - '_query_agent_sync', - return_value=self._mock_responses(plan_text)): - sketch = approach._query_agent_for_plan_sketch(task) - - assert len(sketch) == 2 - assert sketch[0].subgoal_atoms is None - assert sketch[1].subgoal_atoms is None - - def test_sketch_with_code_fences(self): - """Test sketch with code fences.""" - approach, _, task = _make_approach() - - plan_text = ("Here is the plan:\n" - "```\n" - "Pick(block0:block) -> {Holding(block0:block)}\n" - "Place(block0:block, block1:block)\n" - "```\n") - - with patch.object(approach, - '_query_agent_sync', - return_value=self._mock_responses(plan_text)): - sketch = approach._query_agent_for_plan_sketch(task) - - assert len(sketch) == 2 - - def test_sketch_with_preamble(self): - """Agent includes analysis text before the plan.""" - approach, _, task = _make_approach() - - plan_text = ( - "After inspecting the environment, I found block0 and block1.\n" - "The goal is to place block0 on block1.\n" - "\n" - "Pick(block0:block)\n" - "Place(block0:block, block1:block)\n") - - with patch.object(approach, - '_query_agent_sync', - return_value=self._mock_responses(plan_text)): - sketch = approach._query_agent_for_plan_sketch(task) - - assert len(sketch) == 2 - - def test_sketch_with_wait(self): - """Test sketch with wait.""" - approach, _, task = _make_approach() - - plan_text = ("Pick(block0:block) -> {Holding(block0:block)}\n" - "Wait(robot0:robot)\n" - "Place(block0:block, block1:block) -> " - "{On(block0:block, block1:block)}\n") - - with patch.object(approach, - '_query_agent_sync', - return_value=self._mock_responses(plan_text)): - sketch = approach._query_agent_for_plan_sketch(task) - - assert len(sketch) == 3 - assert sketch[0].option.name == "Pick" - assert sketch[1].option.name == "Wait" - assert sketch[1].subgoal_atoms is None - assert sketch[2].option.name == "Place" - - def test_empty_response_raises(self): - """Agent returns no text → ApproachFailure.""" - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - - with patch.object(approach, - '_query_agent_sync', - return_value=[{ - "type": "result", - "content": [] - }]): - with pytest.raises(ApproachFailure, match="empty plan text"): - approach._query_agent_for_plan_sketch(task) - - def test_no_valid_options_raises(self): - """Agent returns text with no valid option names → ApproachFailure.""" - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - - plan_text = "I don't know what to do.\nSorry!\n" - - with patch.object(approach, - '_query_agent_sync', - return_value=self._mock_responses(plan_text)): - with pytest.raises(ApproachFailure, match="Parsed empty"): - approach._query_agent_for_plan_sketch(task) - - def test_sketch_from_file(self): - """Load sketch from a saved text file via CFG option.""" - approach, _, task = _make_approach() - sketch_path = os.path.join(_TEST_DATA_DIR, "simple_plan_sketch.txt") - - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "seed": 42, - "agent_bilevel_plan_sketch_file": sketch_path, - }) - - sketch = approach._query_agent_for_plan_sketch(task) - - assert len(sketch) == 2 - assert sketch[0].option.name == "Pick" - assert list(sketch[0].objects) == [_block0] - assert sketch[0].subgoal_atoms is not None - assert GroundAtom(_Holding, [_block0]) in sketch[0].subgoal_atoms - assert sketch[1].option.name == "Place" - assert list(sketch[1].objects) == [_block0, _block1] - assert sketch[1].subgoal_atoms is not None - assert GroundAtom(_On, [_block0, _block1]) in sketch[1].subgoal_atoms - - -# --------------------------------------------------------------------------- -# Tests: _sample_params -# --------------------------------------------------------------------------- - - -class TestValidatePlanForward: - """Tests for ``plan_execution.validate_plan_forward``. - - Covers the test-time forward validator that's the entire reason the - synthesis tool can catch refinement-passes/validation-fails - regressions. - """ - - def _grounded(self, option, objects, params=None): - if params is None: - params = np.zeros(option.params_space.shape[0], dtype=np.float32) - return option.ground(list(objects), np.asarray(params, - dtype=np.float32)) - - def test_goal_reached_returns_success(self): - """Plan that reaches the goal — validator passes, no diagnosis.""" - from predicators.agent_sdk import plan_execution - _, mock_om, task = _make_approach() - # Final post-state satisfies the goal (On(block0, block1)). - goal_state = _make_state({_block0: [0.55, 0.6, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (goal_state, 3) - - plan = [self._grounded(_Pick, [_block0], [0.5])] - ok, reason = plan_execution.validate_plan_forward( - task, plan, mock_om, predicates=_ALL_PREDICATES) - assert ok is True - assert reason == "" - - def test_goal_not_reached_diagnosis_names_missing_atoms(self): - """Plan terminates but goal isn't satisfied — diagnosis names the - missing atom set, not a generic 'validation failed'.""" - from predicators.agent_sdk import plan_execution - _, mock_om, task = _make_approach() - # Post-state doesn't satisfy On(block0, block1). - bad_state = _make_state({_block0: [0.1, 0.2, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (bad_state, 3) - - plan = [self._grounded(_Pick, [_block0], [0.5])] - ok, reason = plan_execution.validate_plan_forward( - task, plan, mock_om, predicates=_ALL_PREDICATES) - assert ok is False - assert "goal not reached" in reason - assert "On(block0:block, block1:block)" in reason - - def test_subgoal_divergence_logged_when_sketch_provided(self, caplog): - """When the sketch is passed in, per-step subgoal divergence is logged - with the missing atom — this is the diagnostic the synthesis agent - needs to see *which* step's predicate is spurious.""" - import logging as _logging - - from predicators.agent_sdk import plan_execution - _, mock_om, task = _make_approach() - # Post-state never establishes Holding(block0). Goal is also - # missing — but the subgoal log should fire first. - bad_state = _make_state({_block0: [0.1, 0.2, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (bad_state, 3) - - plan = [self._grounded(_Pick, [_block0], [0.5])] - sketch = [ - _SketchStep(option=_Pick, - objects=[_block0], - subgoal_atoms={GroundAtom(_Holding, [_block0])}) - ] - with caplog.at_level(_logging.INFO): - ok, _ = plan_execution.validate_plan_forward( - task, - plan, - mock_om, - predicates=_ALL_PREDICATES, - sketch=sketch, - run_id="test_run", - ) - assert ok is False - # Subgoal divergence log mentions the missing atom and the step. - assert any("subgoal divergence at step 0" in r.message - and "Holding(block0:block)" in r.message - for r in caplog.records) - - def test_option_failure_diagnosis_names_step(self): - """When the option model returns 0 actions (option execution failed), - the diagnosis identifies the failing step and surfaces the option - model's last_execution_failure.""" - from predicators.agent_sdk import plan_execution - _, mock_om, task = _make_approach() - # Simulate option failure: 0 actions, with a diagnostic message - # recorded on the option model. - mock_om.get_next_state_and_num_actions.return_value = (_make_state(), - 0) - mock_om.last_execution_failure = "IK timed out at waypoint 3" - - plan = [self._grounded(_Pick, [_block0], [0.5])] - ok, reason = plan_execution.validate_plan_forward( - task, plan, mock_om, predicates=_ALL_PREDICATES) - assert ok is False - assert "option execution failed at step 0" in reason - assert "Pick(block0)" in reason - assert "IK timed out at waypoint 3" in reason - - def test_empty_plan_with_goal_already_satisfied(self): - """Empty plan + init satisfies goal → success.""" - from predicators.agent_sdk import plan_execution - - # Goal trivially holds when block0 is already on block1. - init = _make_state({_block0: [0.55, 0.6, 0.0]}) - task = Task(init, {GroundAtom(_On, [_block0, _block1])}) - mock_om = MagicMock() - ok, reason = plan_execution.validate_plan_forward( - task, [], mock_om, predicates=_ALL_PREDICATES) - assert ok is True - assert reason == "" - - def test_empty_plan_with_unmet_goal(self): - """Empty plan + init does NOT satisfy goal → failure with explanatory - diagnosis.""" - from predicators.agent_sdk import plan_execution - _, _, task = _make_approach() # init does not satisfy goal - mock_om = MagicMock() - ok, reason = plan_execution.validate_plan_forward( - task, [], mock_om, predicates=_ALL_PREDICATES) - assert ok is False - assert "init state does not satisfy goal" in reason - - def test_sketch_length_mismatch_ignored_gracefully(self): - """Mismatched sketch length — validator should warn and fall back to - goal-only checking rather than crash.""" - from predicators.agent_sdk import plan_execution - _, mock_om, task = _make_approach() - goal_state = _make_state({_block0: [0.55, 0.6, 0.0]}) - mock_om.get_next_state_and_num_actions.return_value = (goal_state, 3) - - plan = [self._grounded(_Pick, [_block0], [0.5])] - # Sketch length 2, plan length 1. - sketch = [ - _SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None), - _SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None), - ] - ok, _ = plan_execution.validate_plan_forward( - task, - plan, - mock_om, - predicates=_ALL_PREDICATES, - sketch=sketch, - ) - # Validation still runs to completion against the goal. - assert ok is True - - -class TestSampleParams: - """TestSampleParams class.""" - - def test_empty_params_space(self): - """Test empty params space.""" - approach, _, _ = _make_approach() - rng = np.random.default_rng(0) - params = approach._sample_params(_Wait, _make_state(), rng) - assert params.shape == (0, ) - assert params.dtype == np.float32 - - def test_params_within_bounds(self): - """Test params within bounds.""" - approach, _, _ = _make_approach() - rng = np.random.default_rng(0) - for _ in range(100): - params = approach._sample_params(_Place, _make_state(), rng) - assert params.shape == (2, ) - assert np.all(params >= 0.0) - assert np.all(params <= 1.0) - assert params.dtype == np.float32 - - -# --------------------------------------------------------------------------- -# Tests: class metadata -# --------------------------------------------------------------------------- - - -def test_get_name(): - """Test get name.""" - assert AgentModelBasedApproach.get_name() == "agent_model_based" - # The pre-rename CLI name still resolves to this approach. - # pylint: disable=import-outside-toplevel - from predicators.approaches import _get_approach_cls_from_name - assert _get_approach_cls_from_name( - "agent_bilevel") is AgentModelBasedApproach - - -# --------------------------------------------------------------------------- -# Tests: closed-loop execution replanning (subgoal_annotations monitor + -# _maybe_replan_from_divergence / _replan_suffix) -# --------------------------------------------------------------------------- - -_PickDone = ParameterizedOption( - "Pick", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_true, -) - -_PlaceDone = ParameterizedOption( - "Place", - types=[_block_type, _block_type], - params_space=Box(low=np.array([0.0, 0.0], dtype=np.float32), - high=np.array([1.0, 1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_true, -) - - -def _make_two_step_plan(first_subgoals): - """Plan [Pick, Place] whose first step is annotated with first_subgoals.""" - plan = [ - _PickDone.ground([_block0], np.array([0.5], dtype=np.float32)), - _PlaceDone.ground([_block0, _block1], - np.array([0.5, 0.5], dtype=np.float32)), - ] - sketch = [ - _SketchStep(_PickDone, [_block0], first_subgoals), - _SketchStep(_PlaceDone, [_block0, _block1], None), - ] - return plan, sketch - - -def _enable_replanning(approach, budget): - """Turn on closed-loop execution and start a fresh episode.""" - utils.update_config({ - "agent_bilevel_max_execution_replans": budget, - "execution_monitor": "subgoal_annotations", - }) - approach.reset_for_new_episode() - - -def _make_monitor(approach): - """Create the monitor and sync it with the approach, CogMan-style.""" - from predicators.execution_monitoring import create_execution_monitor - monitor = create_execution_monitor("subgoal_annotations") - monitor.update_approach_info(approach.get_execution_monitoring_info()) - return monitor - - -def _sync(monitor, approach): - """Mimic CogMan pushing fresh approach info to the monitor.""" - monitor.update_approach_info(approach.get_execution_monitoring_info()) - - -class TestExecutionReplanning: - """Tests for closed-loop execution through the cogman monitor flow.""" - - def test_open_loop_when_disabled(self): - """With the flag at 0 (default), no monitoring info is exported and - divergence is never flagged.""" - approach, _, _ = _make_approach() - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - policy = approach._plan_to_policy(plan, sketch=sketch) - assert not approach.get_execution_monitoring_info() - state = _make_state() # block0 not held: subgoal would fail - monitor = _make_monitor(approach) - assert not monitor.step(state) - policy(state) # starts Pick - policy(state) # Pick terminal -> starts Place without any check - - def test_monitor_silent_when_subgoals_hold(self): - """Subgoals satisfied at the boundary: no replan is suggested.""" - approach, _, _ = _make_approach() - _enable_replanning(approach, 2) - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state({_block0: [0.1, 0.2, 1.0]}) # held: subgoal ok - monitor = _make_monitor(approach) - # Before any option is initiated (e.g. right after a replan, - # cogman asserts the monitor does not immediately re-fire). - assert not monitor.step(state) - policy(state) # starts Pick - _sync(monitor, approach) - assert not monitor.step(state) # boundary, but annotation holds - policy(state) # advances to Place - - def test_monitor_silent_mid_option(self): - """A failing annotation is only checked at the option boundary.""" - approach, _, _ = _make_approach() - _enable_replanning(approach, 2) - holding = {GroundAtom(_Holding, [_block0])} - # _Pick never terminates, so execution stays mid-option. - plan = [_Pick.ground([_block0], np.array([0.5], dtype=np.float32))] - sketch = [_SketchStep(_Pick, [_block0], holding)] - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() # block0 not held: subgoal fails - policy(state) - monitor = _make_monitor(approach) - assert not monitor.step(state) - - def test_monitor_detects_divergence_at_boundary(self): - """An unsatisfied annotation at the boundary suggests a replan.""" - approach, _, _ = _make_approach() - _enable_replanning(approach, 2) - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() # block0 not held: subgoal diverges - policy(state) # starts Pick (terminal at every state) - monitor = _make_monitor(approach) - assert monitor.step(state) - - def test_suffix_replan_preferred_on_divergence(self): - """The monitor-triggered re-solve resumes via the suffix path; no agent - re-query.""" - approach, _, task = _make_approach() - _enable_replanning(approach, 2) - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() - policy(state) - monitor = _make_monitor(approach) - assert monitor.step(state) - - # CogMan now re-invokes solve() on the current state. - def sentinel_policy(s): - del s # unused - return Action(np.full(1, 0.25, dtype=np.float32)) - - approach._replan_suffix = MagicMock(return_value=sentinel_policy) - approach._query_agent_for_plan_sketch = MagicMock() - new_policy = approach._solve(Task(state, task.goal), timeout=10) - assert new_policy is sentinel_policy - approach._query_agent_for_plan_sketch.assert_not_called() - approach._replan_suffix.assert_called_once() - args = approach._replan_suffix.call_args.args - assert args[0] is state # replans from the real current state - assert args[3] == 0 # the failed step is the annotated first step - - def test_openloop_resume_when_no_suffix_validates(self): - """Suffix path exhausted: the remaining plan resumes open-loop. - - An annotation is the agent's prediction, not proof the goal is - out of reach, so by default the episode keeps executing and the - goal check decides. A fresh sketch query would re-open the agent - turn budget the attempt already spent, so it stays opt-in - (agent_bilevel_replan_agent_fallback). - """ - approach, _, task = _make_approach() - _enable_replanning(approach, 2) - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() - policy(state) - approach._replan_suffix = MagicMock(return_value=None) - approach._query_agent_for_plan_sketch = MagicMock() - new_policy = approach._solve(Task(state, task.goal), timeout=10) - approach._replan_suffix.assert_called_once() - approach._query_agent_for_plan_sketch.assert_not_called() - # The resumed policy executes the remaining step (Place), and - # monitoring re-arms over exactly that suffix. - new_policy(state) - status = approach.get_execution_monitoring_info()[0] - assert status.steps_initiated == 1 - assert status.current_option.name == "Place" - - def test_openloop_resume_at_last_step_ends_plan(self): - """Divergence at the final step leaves nothing to resume: the returned - policy ends through the normal plan-exhausted path (so a goal-reached - terminator still gets its chance), not a divergence abort.""" - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - _enable_replanning(approach, 2) - holding = {GroundAtom(_Holding, [_block0])} - plan, _ = _make_two_step_plan(holding) - # Annotate the LAST step instead of the first. - sketch = [ - _SketchStep(_PickDone, [_block0], None), - _SketchStep(_PlaceDone, [_block0, _block1], holding), - ] - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() # block0 not held: Place's subgoal fails - policy(state) # starts Pick - policy(state) # Pick terminal -> starts Place - monitor = _make_monitor(approach) - assert monitor.step(state) - approach._replan_suffix = MagicMock(return_value=None) - new_policy = approach._solve(Task(state, task.goal), timeout=10) - with pytest.raises(ApproachFailure, match="exhausted"): - new_policy(state) - - def test_full_resolve_when_no_suffix_validates_with_fallback(self): - """With agent_bilevel_replan_agent_fallback, a failed suffix replan - falls through to a fresh agent sketch.""" - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - _enable_replanning(approach, 2) - utils.update_config({"agent_bilevel_replan_agent_fallback": True}) - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() - policy(state) - approach._replan_suffix = MagicMock(return_value=None) - # Reaching the fresh-sketch body raises its distinctive failure - - # proof we fell through to a fresh agent query. - sketch_query = MagicMock(side_effect=ApproachFailure("no sketch")) - approach._query_agent_for_plan_sketch = sketch_query - with patch.object(approach, '_nudge_final_submission', - MagicMock(return_value=None)): - with pytest.raises(ApproachFailure, match="Bilevel solve failed"): - approach._solve(Task(state, task.goal), timeout=10) - sketch_query.assert_called_once() - approach._replan_suffix.assert_called_once() - - def test_budget_shared_across_chained_replans(self): - """Chained replans share one per-episode budget; once it is exhausted, - a divergence resumes the remaining plan open-loop without paying for - further refinement.""" - approach, _, task = _make_approach() - _enable_replanning(approach, 1) - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - - def _suffix_replan(s, tsk, steps, k, t): - del s, tsk, steps, k, t # unused - new_plan, new_sketch = _make_two_step_plan(holding) - return approach._plan_to_policy(new_plan, sketch=new_sketch) - - approach._replan_suffix = MagicMock(side_effect=_suffix_replan) - approach._query_agent_for_plan_sketch = MagicMock() - policy = approach._plan_to_policy(plan, sketch=sketch) - state = _make_state() - policy(state) - monitor = _make_monitor(approach) - assert monitor.step(state) - # First divergence: budget 1 -> 0, replanned policy starts. - new_policy = approach._solve(Task(state, task.goal), timeout=10) - new_policy(state) - _sync(monitor, approach) - assert monitor.step(state) - # Second divergence: no budget left - the remaining plan resumes - # open-loop, with no further refinement attempt. - resumed = approach._solve(Task(state, task.goal), timeout=10) - approach._replan_suffix.assert_called_once() - approach._query_agent_for_plan_sketch.assert_not_called() - resumed(state) - status = approach.get_execution_monitoring_info()[0] - assert status.current_option.name == "Place" - - def test_reset_for_new_episode_clears_state(self): - """A new episode refreshes the budget and clears the live status.""" - approach, _, _ = _make_approach() - _enable_replanning(approach, 2) - assert approach._exec_replans_left == 2 - holding = {GroundAtom(_Holding, [_block0])} - plan, sketch = _make_two_step_plan(holding) - approach._plan_to_policy(plan, sketch=sketch) - assert approach.get_execution_monitoring_info() - approach._exec_replans_left = 0 - approach.reset_for_new_episode() - assert not approach.get_execution_monitoring_info() - assert approach._exec_replans_left == 2 - - def test_init_requires_subgoal_annotations_monitor(self): - """Enabling the budget without the monitor is a config error.""" - _, _, task = _make_approach() - utils.update_config({"agent_bilevel_max_execution_replans": 2}) - kwargs = dict( - initial_predicates=_ALL_PREDICATES, - initial_options=_ALL_OPTIONS, - types={_block_type, _robot_type}, - action_space=Box(low=-1, high=1, shape=(1, )), - train_tasks=[task], - option_model=MagicMock(), - ) - with pytest.raises(ValueError, match="subgoal_annotations"): - AgentModelBasedApproach(**kwargs) - utils.update_config({"execution_monitor": "subgoal_annotations"}) - AgentModelBasedApproach(**kwargs) - - def test_replan_suffix_walkback_and_validation(self): - """_replan_suffix tries the failed step first, walks back only to the - latest holding annotation, and forward-validates.""" - from predicators.agent_sdk import bilevel_sketch as bs - approach, _, task = _make_approach() - on_atom = {GroundAtom(_On, [_block0, _block1])} - holding = {GroundAtom(_Holding, [_block0])} - sketch = [ - _SketchStep(_PickDone, [_block0], on_atom), # holds (x close) - _SketchStep(_PickDone, [_block0], holding), # does not hold - _SketchStep(_PlaceDone, [_block0, _block1], holding), # failed - ] - # block0.x=0.5 == block1.x=0.5 so On holds; held=0 so Holding fails. - state = _make_state({_block0: [0.5, 0.2, 0.0]}) - tried = [] - - def _fake_refine(tsk, suffix, remaining, attempt=0): - del tsk, remaining, attempt # unused - tried.append(len(suffix)) - # Succeed only for the 2-step suffix (resume at step 1). - if len(suffix) == 2: - new_plan, _ = _make_two_step_plan(holding) - return new_plan, True - return [], False - - approach._refine_sketch = MagicMock(side_effect=_fake_refine) - with patch.object(bs, "validate_plan_forward", - return_value=(True, "")): - policy = approach._replan_suffix(state, task, sketch, 2, 10) - assert policy is not None - # Tried failed step (suffix len 1) first, then one step back - # (len 2); never walked past the holding annotation at step 0. - assert tried == [1, 2] - - -# --------------------------------------------------------------------------- -# Tests: scheduled-plans section in the solve/explore prompt -# --------------------------------------------------------------------------- - - -class TestScheduledPlansPromptSection: - """The explore prompt shows plans already generated this cycle so the next - request proposes a complementary plan instead of repeating the identical - one (run_20260707_112310 emitted the same 1-step plan for both of a cycle's - requests).""" - - @staticmethod - def _prompt(scheduled_plans): - from predicators.agent_sdk import sketch_prompts - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - }) - state = _make_state() - task = Task(state, {GroundAtom(_On, [_block0, _block1])}) - return sketch_prompts.build_solve_prompt( - task, - all_predicates=_ALL_PREDICATES, - all_options=_ALL_OPTIONS, - scheduled_plans=scheduled_plans, - propose_params=True, - ) - - def test_section_absent_without_scheduled_plans(self): - """No scheduled-plans section is emitted when none were scheduled.""" - for empty in (None, []): - prompt = self._prompt(empty) - assert "Plans Already Scheduled This Cycle" not in prompt - - def test_section_lists_plans_and_asks_for_different_one(self): - """Scheduled plans are listed so the agent proposes a different one.""" - plans = [ - " 0: Pick(block0)[0.5000]", - " 0: Place(block0, block1)[0.1000, 0.2000]", - ] - prompt = self._prompt(plans) - assert "## Plans Already Scheduled This Cycle" in prompt - assert "Plan 1:\n 0: Pick(block0)[0.5000]" in prompt - assert "Plan 2:\n 0: Place(block0, block1)[0.1000, 0.2000]" in prompt - assert "data is complementary rather than redundant" in prompt - # The instruction must keep the request goal-directed (this is what - # preserves the train-solve early-stopping semantics). - assert "repeat the best plan" in " ".join(prompt.split()) - - -# --------------------------------------------------------------------------- -# Tests: turn-cap exhaustion handling in _solve -# --------------------------------------------------------------------------- - - -class TestTurnCapHandling: - """Hitting agent_sdk_max_agent_turns_per_iteration ends the attempt with a - best-effort submission instead of burning the sketch retries.""" - - @staticmethod - def _cap_result(subtype=None, num_turns=None): - return { - "type": "result", - "subtype": subtype, - "num_turns": num_turns, - "total_cost_usd": 1.0, - } - - def test_responses_hit_turn_cap(self): - """Cap detection: subtype is authoritative, num_turns is fallback.""" - approach, _, _ = _make_approach() - cap = approach._responses_hit_turn_cap - assert cap([self._cap_result(subtype="error_max_turns")]) - max_turns = 50 - utils.update_config( - {"agent_sdk_max_agent_turns_per_iteration": max_turns}) - assert cap([self._cap_result(subtype="success", num_turns=max_turns)]) - assert not cap( - [self._cap_result(subtype="success", num_turns=max_turns - 1)]) - assert not cap([{"type": "assistant", "content": []}]) - assert not cap([]) - - def test_sketch_query_records_turn_cap(self): - """A capped session with no final text still marks the cap before the - empty-plan-text failure propagates.""" - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - responses = [self._cap_result(subtype="error_max_turns")] - with patch.object(approach, - '_query_agent_sync', - return_value=responses): - with pytest.raises(ApproachFailure, match="empty plan text"): - approach._query_agent_for_plan_sketch(task) - assert approach._last_sketch_query_hit_turn_cap - - def test_solve_one_query_per_attempt_on_turn_cap(self): - """A capped attempt takes the final nudge and stops.""" - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - query = MagicMock( - return_value=[self._cap_result(subtype="error_max_turns")]) - nudge = MagicMock(return_value=None) - with patch.object(approach, '_query_agent_sync', query), \ - patch.object(approach, '_nudge_final_submission', nudge): - with pytest.raises(ApproachFailure, match="Bilevel solve failed"): - approach._solve(task, timeout=10) - assert query.call_count == 1 # no re-query on the same context - nudge.assert_called_once_with() - - def test_solve_restarts_with_no_nudge_until_final_attempt(self): - """Non-final attempts restart directly with no nudge. - - The best-effort submission nudge fires only on the FINAL - attempt, as the ultimate fallback. - """ - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - utils.update_config({"agent_solve_max_attempts": 3}) - query = MagicMock( - return_value=[self._cap_result(subtype="error_max_turns")]) - nudge = MagicMock(return_value=None) - with patch.object(approach, '_query_agent_sync', query), \ - patch.object(approach, '_nudge_final_submission', nudge): - with pytest.raises(ApproachFailure, match="Bilevel solve failed"): - approach._solve(task, timeout=10) - assert query.call_count == 3 # one full query per attempt - nudge.assert_called_once_with() - - def test_solve_no_requery_on_non_cap_failure(self): - """A non-cap failure (e.g. unparseable output) ends the attempt too. - - The fresh-context restart is the ONLY retry: a query that merely - failed to submit does not buy a second full-price query on a - context that already contains whatever went wrong. - """ - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - utils.update_config({"agent_solve_max_attempts": 3}) - # Well under the cap, but no plan text: a real error, not budget end. - query = MagicMock( - return_value=[self._cap_result(subtype="success", num_turns=5)]) - nudge = MagicMock(return_value=None) - with patch.object(approach, '_query_agent_sync', query), \ - patch.object(approach, '_nudge_final_submission', nudge): - with pytest.raises(ApproachFailure, match="Bilevel solve failed"): - approach._solve(task, timeout=10) - assert query.call_count == 3 # one per attempt, not one per query - nudge.assert_called_once_with() - - def test_attempt_end_reason_labels_journal_outcome(self): - """A capture-less attempt records WHY it ended. - - That reason is the one fact the next fresh-context attempt - cannot rediscover from the transcript it no longer has. - """ - from predicators.approaches import ApproachFailure - approach, _, task = _make_approach() - nudge = MagicMock(return_value=None) - - for subtype, num_turns, reason in [ - ("error_max_turns", None, "turn cap"), - ("success", 5, "no submission"), - ]: - query = MagicMock( - return_value=[self._cap_result(subtype, num_turns)]) - with patch.object(approach, '_query_agent_sync', query), \ - patch.object(approach, '_nudge_final_submission', nudge): - with pytest.raises(ApproachFailure, match=reason): - approach._solve(task, timeout=10) - assert approach._last_attempt_end_reason == reason - assert approach._attempt_outcome_text( - None, None) == f"no capture ({reason})" - - def test_nudge_best_effort_flag_set_and_cleared(self): - """The nudge exposes best-effort capture to the tools only for the - duration of its own query.""" - approach, _, _ = _make_approach() - seen = {} - - def _fake_query(message, **kwargs): - del kwargs # unused - seen["flag"] = approach._tool_context.capture_best_effort_plan - seen["message"] = message - return [] - - with patch.object(approach, '_query_agent_sync', _fake_query): - policy = approach._nudge_final_submission() - assert policy is None - assert seen["flag"] is True - assert "even if it does not fully reach the goal" in seen["message"] - assert not approach._tool_context.capture_best_effort_plan - - def test_nudge_returns_captured_best_effort_plan(self): - """The nudge consumes a captured plan into a policy even when the - rollout did not reach the goal.""" - approach, _, _ = _make_approach() - plan = [_Pick.ground([_block0], np.array([0.5], dtype=np.float32))] - sketch = [_SketchStep(_Pick, [_block0], None)] - - def _fake_query(message, **kwargs): - del message, kwargs # unused - # Simulate submit_plan's best-effort capture. - assert approach._tool_context.capture_best_effort_plan - approach._tool_context.solved_plan = plan - approach._tool_context.solved_sketch = sketch - approach._tool_context.solved_plan_reached_goal = False - return [] - - with patch.object(approach, '_query_agent_sync', _fake_query): - policy = approach._nudge_final_submission() - assert policy is not None - assert approach._tool_context.solved_plan is None - assert approach._tool_context.solved_plan_reached_goal is None - - -# --------------------------------------------------------------------------- -# Policy mode (agent_solve_policy_mode) -# --------------------------------------------------------------------------- - -_POLICY_SOURCE = ''' -def get_option(state, memory): - if memory.get("issued"): - return None - memory["issued"] = True - return "Pick(block0:block)[0.5]" -''' - - -def test_policy_mode_constructor_checks(): - """Policy mode rejects replans>0 and sim-free configs.""" - state = _make_state() - task = Task(state, {GroundAtom(_On, [_block0, _block1])}) - base_kwargs = dict(initial_predicates=_ALL_PREDICATES, - initial_options=_ALL_OPTIONS, - types={_block_type, _robot_type}, - action_space=Box(low=-1, high=1, shape=(1, )), - train_tasks=[task]) - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 1, - "execution_monitor": "subgoal_annotations", - }) - with pytest.raises(ValueError, match="mutually exclusive"): - AgentModelBasedApproach(**base_kwargs) - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_planner_use_simulator": False, - }) - with pytest.raises(ValueError, match="use_simulator"): - AgentModelBasedApproach(**base_kwargs) - - -def test_consume_policy_capture_builds_executor(): - """A captured policy source composes and executes closed-loop.""" - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - }) - state = _make_state() - task = Task(state, {GroundAtom(_On, [_block0, _block1])}) - approach._tool_context.current_task = task - approach._tool_context.solved_policy_source = _POLICY_SOURCE - approach._tool_context.solved_plan_reached_goal = True - policy = approach._consume_validated_plan() - assert policy is not None - info = approach._last_capture_info - assert info is not None and info.validated - assert "policy.py sha=" in info.plan_lines[0] - # The composed executor runs the issued option (terminal is always - # False here, so the first call returns that option's action). - action = policy(state) - assert isinstance(action, Action) - # No sketch monitor is armed in policy mode. - assert approach._exec_status is None - - -def test_execution_policy_surfaces_option_failures(): - """A failed option is surfaced to the policy, not episode-fatal.""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 10, - }) - not_initiable = ParameterizedOption( - "Broken", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_false, - terminal=_always_false, - ) - seen = [] - - def option_fn(state, last_failure): - del state - seen.append(last_failure) - if last_failure is None and len(seen) == 1: - return not_initiable.ground([_block0], - np.array([0.5], dtype=np.float32)) - return None - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, match="DONE"): - policy(_make_state()) - assert seen[0] is None - assert seen[1] is not None # the failure was surfaced, not fatal - - -def test_execution_policy_budget_is_fatal(): - """The option cap converts an oscillating policy into a failure.""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 2, - }) - not_initiable = ParameterizedOption( - "Broken", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_false, - terminal=_always_false, - ) - - def option_fn(state, last_failure): - del state, last_failure - return not_initiable.ground([_block0], np.array([0.5], - dtype=np.float32)) - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, match="option budget"): - policy(_make_state()) - - -def _broken_option(): - return ParameterizedOption( - "Broken", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_false, - terminal=_always_false, - ) - - -def test_execution_policy_stuck_loop_is_fatal(): - """K consecutive failures of one identical command end the episode before - the option budget is burned (mirrors execute_policy_forward).""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 50, - }) - not_initiable = _broken_option() - issued = [] - - def option_fn(state, last_failure): - del state, last_failure - issued.append(1) - return not_initiable.ground([_block0], np.array([0.5], - dtype=np.float32)) - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, - match="re-issued the same failing option"): - policy(_make_state()) - assert len(issued) == 3 # the guard default, not the 50 cap - - -def test_execution_policy_stuck_loop_resets_on_changed_params(): - """Adapting the parameters after each failure avoids the guard: the episode - runs to the option budget instead.""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 6, - }) - not_initiable = _broken_option() - n_calls = [0] - - def option_fn(state, last_failure): - del state, last_failure - n_calls[0] += 1 - return not_initiable.ground([_block0], - np.array([0.1 * n_calls[0]], - dtype=np.float32)) - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, match="option budget"): - policy(_make_state()) - assert n_calls[0] == 6 - - -def test_execution_policy_stuck_loop_on_mid_execution_raise(): - """The guard also counts a skill that raises from inside its own policy - (e.g. a motion-planning refusal), whose exception carries no - last_failed_option of its own: option_policy_to_policy attributes it to the - executing option, so K identical re-issues end the episode instead of - burning the option budget (2026-08-25 policy-arm cycle-4 test).""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 50, - }) - - def _refusing_policy(_s, _m, _o, _p): - raise utils.OptionExecutionFailure( - "[Broken/MoveAbove] BiRRT collision: start configuration " - "in collision") - - refusing = ParameterizedOption( - "Broken", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_refusing_policy, - initiable=_always_true, - terminal=_always_false, - ) - issued = [] - - def option_fn(state, last_failure): - del state, last_failure - issued.append(1) - return refusing.ground([_block0], np.array([0.5], dtype=np.float32)) - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, - match="re-issued the same failing option"): - policy(_make_state()) - assert len(issued) == 3 # the guard default, not the 50 cap - - -def _instant_option(): - """Completes immediately (terminal always true) without failing.""" - return ParameterizedOption( - "Instant", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_true, - ) - - -def test_execution_policy_noop_livelock_is_fatal(): - """K consecutive clean completions of one identical command with no - observable state change end the episode as a livelock (mirrors - execute_policy_forward).""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 50, - }) - instant = _instant_option() - issued = [] - - def option_fn(state, last_failure): - del state, last_failure - issued.append(1) - return instant.ground([_block0], np.array([0.5], dtype=np.float32)) - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, match="no observable state change"): - for _ in range(10): - policy(_make_state()) - assert len(issued) == 3 # the guard default, not the 50 cap - - -def test_execution_policy_noop_livelock_resets_on_changed_params(): - """Varying the parameters makes each command a different one, so the - livelock guard never trips; the option budget ends the episode.""" - from predicators.approaches import ApproachFailure - approach, _, _ = _make_approach() - utils.reset_config({ - "env": "cover", - "approach": "agent_model_based", - "seed": 42, - "agent_solve_policy_mode": True, - "agent_bilevel_max_execution_replans": 0, - "agent_policy_max_options": 6, - }) - instant = _instant_option() - n_calls = [0] - - def option_fn(state, last_failure): - del state, last_failure - n_calls[0] += 1 - return instant.ground([_block0], - np.array([0.1 * n_calls[0]], dtype=np.float32)) - - policy = approach._policy_to_execution_policy(option_fn) - with pytest.raises(ApproachFailure, match="option budget"): - for _ in range(10): - policy(_make_state()) - assert n_calls[0] == 6 diff --git a/tests/approaches/test_agent_nl_world_model_approach.py b/tests/approaches/test_agent_nl_world_model_approach.py deleted file mode 100644 index 3ed76b8215..0000000000 --- a/tests/approaches/test_agent_nl_world_model_approach.py +++ /dev/null @@ -1,181 +0,0 @@ -"""Tests for the natural-language world model approach's harness glue (paper -arm C3).""" -# pylint: disable=protected-access -import os -from types import SimpleNamespace -from typing import Any, List - -from predicators import utils -from predicators.agent_sdk import learn_prompts -from predicators.agent_sdk.tools import ToolContext -from predicators.approaches import agent_nl_world_model_approach as anl -from predicators.envs import create_new_env -from predicators.explorers.agent_model_free_explorer import \ - AgentModelFreeExplorer -from predicators.ground_truth_models import get_gt_options -from predicators.settings import CFG -from predicators.structs import Task - -_NOTES = "# Mechanisms\n- the hand moves to the PickPlace parameter\n" - - -def _cover() -> Any: - utils.reset_config({ - "env": "cover", - "num_train_tasks": 2, - "num_test_tasks": 1, - "agent_sdk_use_local_sandbox": True, - "seed": 0, - }) - env = create_new_env("cover") - train_tasks = [ - Task(t.task.init, t.task.goal, goal_nl="Cover the target.") - for t in env.get_train_tasks() - ] - options = get_gt_options(env.get_name()) - return env, train_tasks, options - - -def _bare(env: Any, train_tasks: List[Task], options: Any, - log_dir: str) -> Any: - approach = anl.AgentNotesWorldModelApproach.__new__( - anl.AgentNotesWorldModelApproach) - approach._types = env.types - approach._initial_predicates = set(env.predicates) - approach._initial_options = options - approach._train_tasks = train_tasks - approach._tool_context = ToolContext(types=env.types, - predicates=set(env.predicates), - options=options, - train_tasks=train_tasks) - approach._agent_session = None - approach._notes = "" - approach._notes_version = None - approach._online_learning_cycle = 0 - approach._get_log_dir = lambda: log_dir # type: ignore[method-assign] - approach._get_all_options = lambda: options # type: ignore[method-assign] - approach._get_all_trajectories = lambda: [] # type: ignore[method-assign] - approach._offline_dataset = SimpleNamespace( # type: ignore[assignment] - trajectories=[]) - approach._online_trajectories = [] - approach._option_model = None - approach._synthesized_samplers = {} - return approach - - -def test_predicate_allowlist_and_paths(tmp_path, monkeypatch) -> None: - """The kept-predicate allowlist applies; sandbox paths mirror the code - arms' mapping.""" - env, train_tasks, options = _cover() - approach = _bare(env, train_tasks, options, str(tmp_path)) - monkeypatch.setattr(CFG, "agent_sim_learn_kept_predicates_names", []) - assert approach._get_all_predicates() == set(env.predicates) - monkeypatch.setattr(CFG, "agent_sim_learn_kept_predicates_names", - ["Holding"]) - assert {p.name for p in approach._get_all_predicates()} == {"Holding"} - paths = approach._notes_paths() - assert paths["notes_file"] == os.path.join(str(tmp_path), "sandbox", - "world_model.md") - assert paths["notes_file_for_agent"] == "./world_model.md" - assert approach._get_synthesis_tool_names() == ["run_python"] - - -def test_learn_notes_gating_and_checkpoint_round_trip(tmp_path, - monkeypatch) -> None: - """No data means no session unless zero-shot is on; the document survives a - checkpoint and is written back into the sandbox.""" - env, train_tasks, options = _cover() - approach = _bare(env, train_tasks, options, str(tmp_path)) - calls: List[Any] = [] - approach._run_notes_session = calls.append - monkeypatch.setattr(CFG, "agent_sim_learn_zero_shot", False) - approach._learn_notes() - assert not calls - monkeypatch.setattr(CFG, "agent_sim_learn_zero_shot", True) - approach._learn_notes() - assert calls == [[]] - # A session's document is loaded from the sandbox file. - paths = approach._notes_paths() - os.makedirs(paths["base"], exist_ok=True) - with open(paths["notes_file"], "w", encoding="utf-8") as f: - f.write(_NOTES) - approach._load_notes(paths) - assert approach._notes == _NOTES - assert approach._notes_version is not None - assert approach._tool_context.world_model_notes == _NOTES - assert approach._tool_context.world_model_notes_path == \ - "./world_model.md" - # Checkpoint round trip into a fresh instance with an empty sandbox. - saved = approach._extra_save_state() - assert saved["world_model_notes"] == _NOTES - other_dir = os.path.join(str(tmp_path), "other") - other = _bare(env, train_tasks, options, other_dir) - other._load_extra_save_state(saved) - assert other._notes == _NOTES - with open(other._notes_paths()["notes_file"], encoding="utf-8") as f: - assert f.read() == _NOTES - - -def test_notes_reach_the_solve_and_explore_prompts(tmp_path) -> None: - """The document is quoted into the solve prompt, the explore prompt, and - the system prompt names it; without notes nothing is quoted.""" - env, train_tasks, options = _cover() - approach = _bare(env, train_tasks, options, str(tmp_path)) - approach._initial_image_section = lambda: "" # type: ignore[method-assign] - assert approach._solve_prompt_extra_sections() == "" - approach._notes = _NOTES - approach._sync_tool_context() - extra = approach._solve_prompt_extra_sections() - assert "World model notes (./world_model.md)" in extra - assert "hand moves to the PickPlace parameter" in extra - prompt = approach._build_solve_prompt(train_tasks[0]) - assert "hand moves to the PickPlace parameter" in prompt - assert prompt.index("World model notes") < prompt.index("## Objects") - system = approach._get_agent_system_prompt() - assert "world_model.md" in system and "no simulator" in system - approach._learning_mode = True - assert "natural-language document" in approach._get_agent_system_prompt() - approach._learning_mode = False - explorer = AgentModelFreeExplorer(set(env.predicates), options, env.types, - env.action_space, train_tasks, 10, - approach._tool_context, - None) # type: ignore[arg-type] - explore_prompt = explorer._build_exploration_prompt(0) - assert "hand moves to the PickPlace parameter" in explore_prompt - assert explore_prompt.index("World model notes") < \ - explore_prompt.index("## Instructions") - - -def test_notes_learn_message_and_exec_namespace(tmp_path) -> None: - """The first message carries the data roster, goals, and the prior-notes - pointer; the exec namespace exposes the data helpers.""" - env, train_tasks, options = _cover() - approach = _bare(env, train_tasks, options, str(tmp_path)) - paths = approach._notes_paths() - message = approach._build_notes_learn_message([], paths) - assert "0 recorded trajectories" in message - assert "Cover the target." in message - assert message.count("Cover the target.") == 1 - assert "Read it first" not in message - assert "./reference/structs.py" in message - assert os.path.isfile( - os.path.join(paths["base"], "reference", "structs.py")) - os.makedirs(paths["base"], exist_ok=True) - with open(paths["notes_file"], "w", encoding="utf-8") as f: - f.write(_NOTES) - assert "Read it first" in approach._build_notes_learn_message([], paths) - ns = approach._build_notes_exec_ns([]) - assert set(ns) >= { - "trajectories", "train_tasks", "is_goal_state", "describe_trajectory", - "np" - } - assert ns["is_goal_state"](train_tasks[0].init, 0) is False - - -def test_notes_prompt_builders() -> None: - """Builders render without leftovers; the block is empty without notes.""" - system = learn_prompts.build_notes_learn_system_prompt() - assert "# Mechanisms" in system and "__" not in system - assert learn_prompts.render_world_model_notes_block("", "x") == "" - zero = learn_prompts.render_notes_zero_shot_message() - assert "No trajectory has been recorded" in zero diff --git a/tests/approaches/test_agent_program_world_model_approach.py b/tests/approaches/test_agent_program_world_model_approach.py index 9c7234b8e2..f136f68593 100644 --- a/tests/approaches/test_agent_program_world_model_approach.py +++ b/tests/approaches/test_agent_program_world_model_approach.py @@ -3,19 +3,15 @@ import os from typing import Any, List -import numpy as np - from predicators import utils -from predicators.agent_sdk import learn_prompts from predicators.agent_sdk.tools import ToolContext from predicators.approaches import agent_program_world_model_approach as apwm from predicators.approaches.agent_sim_learning_approach import _SynthesisPaths from predicators.code_sim_learning.program_world_model import \ - ProgramOptionModel, load_program_world_model + load_program_world_model from predicators.datasets import create_dataset from predicators.envs import create_new_env from predicators.ground_truth_models import get_gt_options -from predicators.settings import CFG _PROGRAM = ''' LATENT_FEATURES = {"robot": ["phase"]} @@ -64,42 +60,16 @@ def _bare(env: Any, train_tasks: List[Any], options: Any) -> Any: return approach -def test_belief_particles_and_override_scope() -> None: - """Particles, the nominal latent, the override scope, and the rolled - latents all come from the installed program.""" +def test_installed_program_rolls_latents() -> None: + """An installed program backs the option model, and materialise_latent + rolls it along a recorded trajectory.""" env, train_tasks, options = _cover() approach = _bare(env, train_tasks, options) - # No model yet: no particles, and the initial latent is left alone. - assert not approach._belief_particles() - assert approach._attach_initial_latent(train_tasks[0]) is train_tasks[0] program, err = load_program_world_model(_PROGRAM, env.types, env.predicates, options) assert err is None and program is not None approach._install_program(program) assert approach._option_model is approach._program_model - particles = approach._belief_particles() - # Distinct draws only: initial_latent has three outcomes. - assert 1 <= len(particles) <= 3 - assert len({p["phase"] for p in particles}) == len(particles) - # Deterministic across calls (seeded). - assert approach._belief_particles() == particles - # The current task drives the draw when one is set. - approach._tool_context.current_task = train_tasks[1] - assert approach._belief_particles() == particles - # The nominal latent is attached to the task the planner sees. - task = approach._attach_initial_latent(train_tasks[0]) - assert task.init.latent is not None and "phase" in task.init.latent - assert train_tasks[0].init.latent is None - # Under the scope every latent-less start rolls from the particle. - model: ProgramOptionModel = approach._program_model - (pick_place, ) = [o for o in options if o.name == "PickPlace"] - option = pick_place.ground([], np.array([0.4], dtype=np.float32)) - with approach._particle_override_scope({"phase": 20}): - nxt, _ = model.get_next_state_and_num_actions(train_tasks[0].init, - option) - assert nxt.latent == {"phase": 21} - assert model.initial_latent_override is None - # materialise_latent rolls the program along a recorded trajectory. dataset = create_dataset(env, train_tasks, options, env.predicates) traj = dataset.trajectories[0] latents = approach.materialise_latent(traj) @@ -108,32 +78,6 @@ def test_belief_particles_and_override_scope() -> None: assert approach._latent_tracking_available() is False -def test_learn_simulator_gating(monkeypatch) -> None: - """No data means no session unless zero-shot is on.""" - env, train_tasks, options = _cover() - approach = _bare(env, train_tasks, options) - approach._persist_fit_trajectories = lambda *a, **k: None - calls: List[Any] = [] - program, _ = load_program_world_model(_PROGRAM, env.types, env.predicates, - options) - - def _session(trajectories): - calls.append(list(trajectories)) - return program - - approach._run_program_synthesis_session = _session - monkeypatch.setattr(CFG, "agent_sim_learn_zero_shot", False) - approach._learn_simulator([]) - assert not calls and approach._program is None - monkeypatch.setattr(CFG, "agent_sim_learn_zero_shot", True) - approach._learn_simulator([]) - assert calls == [[]] and approach._program is program - dataset = create_dataset(env, train_tasks, options, env.predicates) - monkeypatch.setattr(CFG, "agent_sim_learn_zero_shot", False) - approach._learn_simulator(list(dataset.trajectories)) - assert len(calls) == 2 and len(calls[1]) == len(dataset.trajectories) - - def test_rehydrate_from_world_model_file(tmp_path, monkeypatch) -> None: """A checkpoint's world_model.py rebuilds the option model.""" env, train_tasks, options = _cover() @@ -162,31 +106,3 @@ def test_rehydrate_from_world_model_file(tmp_path, monkeypatch) -> None: assert program.latent_features == {"robot": ["phase"]} assert "world_model.py" in approach._CHECKPOINT_SANDBOX_FILES assert "world_model_versions" in approach._CHECKPOINT_SANDBOX_DIRS - - -def test_program_prompts_render() -> None: - """System prompt and first message render without leftovers.""" - system = learn_prompts.build_program_learn_system_prompt( - scene_viz_hint="x", - extra_sections=[ - learn_prompts.render_predicate_invention_section("workbench") - ], - workflow_extra=learn_prompts.render_predicate_workflow_extra()) - assert "world_model.py" in system and "sim.score" in system - assert "Plan format" in system and "Predicate Invention" in system - assert "__" not in system.replace("__init__", "") - message = learn_prompts.build_program_learn_message( - n_trajs=0, - n_transitions=0, - n_demos=0, - n_interaction=0, - trajectory_listing="", - structs_ref="./reference/structs.py", - predicate_listing="- Holding(robot, block)", - types_digest="types", - options_digest="options", - world_model_file="./world_model.py", - extra_messages=[learn_prompts.render_program_zero_shot_message()]) - assert "./world_model.py" in message - assert "No trajectory has been recorded" in message - assert "__" not in message diff --git a/tests/approaches/test_agent_sim_learning_ablations.py b/tests/approaches/test_agent_sim_learning_ablations.py index 8f9f9f1986..431812b326 100644 --- a/tests/approaches/test_agent_sim_learning_ablations.py +++ b/tests/approaches/test_agent_sim_learning_ablations.py @@ -5,8 +5,6 @@ ensemble. * ``agent_sim_learn_declared_params_only`` (A4): no estimation runs; the declaration is the estimate and its box the plausible interval. -* ``agent_sim_learn_zero_shot`` (A2): the synthesis session runs with - no recorded transitions. """ # pylint: disable=protected-access from typing import Any, Dict, List @@ -15,7 +13,6 @@ import pytest from predicators import utils -from predicators.agent_sdk import learn_prompts from predicators.agent_sdk.tools import create_synthesis_tools from predicators.approaches import agent_sim_learning_approach as asla from predicators.code_sim_learning.fit_space import ParamSpec, \ @@ -45,7 +42,6 @@ def _bare_approach() -> Any: approach._base_env = _RegistryEnv() approach._identified_physical_params = {} approach._identified_physical_sigma_points = [] - approach._cycle_applied_physical = {} approach._fitted_params = {} approach._param_ensemble = [] approach._param_specs = [] @@ -118,7 +114,6 @@ def test_deploy_declared_params_uses_the_declaration_as_the_estimate() -> None: # to the planning env. assert approach._fitted_params == {"k": 2.0} assert approach._base_env.applied == [{"lateral_friction": 0.5}] - assert approach._cycle_applied_physical == {"lateral_friction": 0.5} # Physics margin spans the declared box. frictions = [ p["lateral_friction"] @@ -140,13 +135,12 @@ def test_deploy_declared_params_uses_the_declaration_as_the_estimate() -> None: assert approach._fit_sse == 1.5 -def test_rule_param_margin_alone_builds_the_ensemble() -> None: - """A6 (info-seeking off, gate on) keeps the validation ensemble; with both - consumers off (A6+A7) none is built.""" +def test_ensemble_follows_info_seeking() -> None: + """Info-seeking, the ensemble's consumer, builds it; with info-seeking off + none is built.""" utils.reset_config({ "agent_sim_learn_declared_params_only": True, - "agent_explorer_info_seeking": False, - "agent_plan_validation_rule_param_margin": True, + "agent_explorer_info_seeking": True, "agent_explorer_info_ensemble_size": 5, }) approach = _bare_approach() @@ -157,7 +151,6 @@ def test_rule_param_margin_alone_builds_the_ensemble() -> None: utils.reset_config({ "agent_sim_learn_declared_params_only": True, "agent_explorer_info_seeking": False, - "agent_plan_validation_rule_param_margin": False, }) approach = _bare_approach() approach._physical_param_specs = list(_PHYS_SPECS) @@ -209,7 +202,6 @@ def test_no_data_seeding_applies_declared_physical_inits() -> None: which the zero-shot arm relies on.""" utils.reset_config({ "agent_sim_learn_declared_params_only": False, - "agent_sim_learn_oracle_sim_params": False, "agent_explorer_info_seeking": False, }) approach = _bare_approach() @@ -220,58 +212,6 @@ def test_no_data_seeding_applies_declared_physical_inits() -> None: assert approach._last_fit_result is None -def test_zero_shot_flag_gates_data_free_synthesis() -> None: - """With no transitions, _learn_simulator returns early unless the zero-shot - flag is set, in which case synthesis runs on empty data.""" - approach: Any = asla.AgentSimLearningApproach.__new__( - asla.AgentSimLearningApproach) - approach._explainability_cache = {} - approach._sysid_fit_cache = {} - approach._persist_fit_trajectories = lambda *a, **k: None - approach._maybe_install_oracle_samplers = lambda: None - approach._extract_obs_triples = lambda trajs: [] - approach._residual_rules = None - approach._learned_simulator = None - approach._fitted_params = {} - calls: List[Any] = [] - - def _synth(trajectories, obs_triples, base_pred_triples, inferred_hint): - calls.append( - (trajectories, obs_triples, base_pred_triples, inferred_hint)) - - approach._synthesize_with_agent = _synth - utils.reset_config({ - "agent_sim_learn_zero_shot": False, - "agent_sim_learn_oracle_sim_program": False, - }) - approach._learn_simulator([]) - assert not calls - utils.reset_config({ - "agent_sim_learn_zero_shot": True, - "agent_sim_learn_oracle_sim_program": False, - }) - approach._learn_simulator([]) - assert calls == [([], [], [], {})] - - -def test_declared_params_prompt_section_is_flag_gated() -> None: - """The no-estimation section renders only under the A3 flag.""" - kwargs: Dict[str, Any] = dict( - partially_observable=False, - residual_rule_signature="def rule(state, updates, params):", - scene_viz_hint="look", - ) - plain = learn_prompts.build_learn_system_prompt(**kwargs) - declared = learn_prompts.build_learn_system_prompt( - declared_params_only=True, **kwargs) - marker = "Harness parameter estimation is DISABLED" - assert marker not in plain - assert marker in declared - assert "__" not in declared.replace("__init__", "") - zero_shot = learn_prompts.render_zero_shot_message() - assert "No trajectory has been recorded" in zero_shot - - def test_estimation_surfaces_refuse_under_declared_params(tmp_path) -> None: """A4: sim.fit, fit_params and sweep_params refuse; the plain report runs.""" diff --git a/tests/approaches/test_agent_sim_learning_approach.py b/tests/approaches/test_agent_sim_learning_approach.py index 397bdee1e6..78b5845fff 100644 --- a/tests/approaches/test_agent_sim_learning_approach.py +++ b/tests/approaches/test_agent_sim_learning_approach.py @@ -6,24 +6,20 @@ that solve a pybullet_boil task. """ # pylint: disable=protected-access -import inspect import logging import os import re from types import SimpleNamespace from typing import List, Optional, Sequence, Set, Tuple, cast -import dill as pkl import numpy as np import pytest from predicators import utils +from predicators.agent_sdk.sketch_types import SketchStep as _SketchStep from predicators.approaches import agent_sim_learning_approach as asla -from predicators.approaches.agent_model_based_approach import _SketchStep from predicators.approaches.agent_sim_learning_approach import \ AgentSimLearningApproach -from predicators.code_sim_learning.fit_space import FitResult -from predicators.code_sim_learning.identifiability import Verdict from predicators.code_sim_learning.utils import LearnedSimulator, \ apply_rules, merge_updates from predicators.envs import create_new_env @@ -497,9 +493,8 @@ def test_fresh_validation_env_scope_disposes_crash_replacement(monkeypatch): def test_fresh_validation_env_scope_applies_physics_overrides(monkeypatch): - """``physical_overrides`` (the capture gate's physics-margin points) land - on the FRESH env on top of the identified params; the shared env is never - touched.""" + """``physical_overrides`` (physics-sweep points) land on the FRESH env on + top of the identified params; the shared env is never touched.""" class _MergingScopeEnv(_FakeScopeEnv): """Sticky per-param merge, matching the real override semantics.""" @@ -595,60 +590,6 @@ def test_rollout_fit_trajectories_subset() -> None: obj._rollout_fit_trajectories(traj_idxs=[3]) -def _cross_cycle_fit(value: float) -> Tuple[FitResult, dict]: - result = FitResult(names=["friction"], - samples=np.array([[value]]), - log_probs=np.zeros(1), - jacobian=None, - noise_sigma=0.05, - prior_sigma=np.array([0.75]), - scales=["log"]) - report = { - "friction": { - "posterior_std": 0.1, - "prior_std": 0.75, - "contraction": 0.13, - "verdict": Verdict.IDENTIFIED, - "note": "", - } - } - return result, report - - -def test_cross_cycle_inconsistent_holds_then_confirms() -> None: - """A many-sigma jump is held once, accepted on independent repeat. - - Regression for run_20260724_232411 seed2: cycle fits 0.3236 -> - 0.6267 (4.7 combined sigmas). The first jump must flag INCONSISTENT - (trusted history unchanged, both fits recorded as hull candidates); - a following cycle re-fitting near the new value confirms the jump - and the history moves - without confirmation the stale reference - would flag every future fit forever. - """ - obj = object.__new__(AgentSimLearningApproach) - obj._sysid_fit_history = {} - obj._sysid_pending_fit = {} - - result, report = _cross_cycle_fit(0.3236) - obj._check_cross_cycle_consistency(result, report, ["friction"]) - assert report["friction"]["verdict"] is Verdict.IDENTIFIED - assert obj._sysid_fit_history["friction"][0] == 0.3236 - - result, report = _cross_cycle_fit(0.6267) - obj._check_cross_cycle_consistency(result, report, ["friction"]) - assert report["friction"]["verdict"] is Verdict.INCONSISTENT - assert report["friction"]["candidate_values"] == [0.3236, 0.6267] - # Trusted history holds; the rejected fit waits as pending. - assert obj._sysid_fit_history["friction"][0] == 0.3236 - assert obj._sysid_pending_fit["friction"][0] == 0.6267 - - result, report = _cross_cycle_fit(0.63) - obj._check_cross_cycle_consistency(result, report, ["friction"]) - assert report["friction"]["verdict"] is Verdict.IDENTIFIED - assert obj._sysid_fit_history["friction"][0] == 0.63 - assert "friction" not in obj._sysid_pending_fit - - def test_make_probe_process_model_factory() -> None: """The certificate-probe factory mirrors the combined simulator. @@ -705,124 +646,6 @@ def latent_rule(state: State, latent: dict, history: list, updates: dict, assert factory()(state, noop).get(thing, "x") == 1.0 -def test_cross_cycle_arbitration_by_pooled_evidence() -> None: - """A flagged jump is accepted when pooled data decisively backs it. - - Regression for run_20260727_210827 seed1: the sharp-but-biased - 2-trajectory cycle-0 fit (0.9313, true 0.5) was held over the - 4-trajectory refit (0.4748) for the rest of the run even though the - refit explained the pooled data ~30x better. With a pooled-SSE probe - the arbitration must accept the new value immediately; an ambivalent - gap (or a failing probe) must keep the hold. - """ - obj = object.__new__(AgentSimLearningApproach) - obj._sysid_fit_history = {} - obj._sysid_pending_fit = {} - - result, report = _cross_cycle_fit(0.9313) - obj._check_cross_cycle_consistency(result, report, ["friction"]) - assert obj._sysid_fit_history["friction"][0] == 0.9313 - - def pooled_sse(theta: dict) -> float: - return 0.14 if abs(theta["friction"] - 0.4748) < 1e-9 else 4.4 - - result, report = _cross_cycle_fit(0.4748) - obj._check_cross_cycle_consistency(result, - report, ["friction"], - pooled_sse=pooled_sse) - assert report["friction"]["verdict"] is Verdict.IDENTIFIED - assert "candidate_values" not in report["friction"] - assert obj._sysid_fit_history["friction"][0] == 0.4748 - assert "friction" not in obj._sysid_pending_fit - - # Ambivalent pooled gap (below the decisive ratio): hold as before. - obj._sysid_fit_history = {"friction": (0.9313, 0.1, "log")} - obj._sysid_pending_fit = {} - result, report = _cross_cycle_fit(0.4748) - obj._check_cross_cycle_consistency(result, - report, ["friction"], - pooled_sse=lambda theta: 0.14) - assert report["friction"]["verdict"] is Verdict.INCONSISTENT - assert obj._sysid_fit_history["friction"][0] == 0.9313 - assert obj._sysid_pending_fit["friction"][0] == 0.4748 - - # A failing SSE probe must fall back to the hold, not crash. - obj._sysid_fit_history = {"friction": (0.9313, 0.1, "log")} - obj._sysid_pending_fit = {} - - def broken_sse(theta: dict) -> float: - raise RuntimeError("env died") - - result, report = _cross_cycle_fit(0.4748) - obj._check_cross_cycle_consistency(result, - report, ["friction"], - pooled_sse=broken_sse) - assert report["friction"]["verdict"] is Verdict.INCONSISTENT - assert obj._sysid_fit_history["friction"][0] == 0.9313 - - -def test_persist_fit_trajectories(tmp_path, monkeypatch) -> None: - """Fit data lands in /fit_data/, one numbered pickle per fit.""" - obj = object.__new__(AgentSimLearningApproach) - obj._fit_trajectories = cast(List[LowLevelTrajectory], - ["fake_traj_a", "fake_traj_b"]) - obj._physical_param_specs = [] - obj._identified_physical_params = {"friction": 0.5} - obj._get_log_dir = lambda: str(tmp_path) # type: ignore[method-assign] - - obj._persist_fit_trajectories() - obj._persist_fit_trajectories() - out_dir = tmp_path / "fit_data" - files = sorted(f.name for f in out_dir.glob("*.pkl")) - assert files == [ - "fit_trajectories_000_fitted.pkl", "fit_trajectories_001_fitted.pkl" - ] - with open(out_dir / files[0], "rb") as f: - payload = pkl.load(f) - assert payload["trajectories"] == ["fake_traj_a", "fake_traj_b"] - assert payload["identified_physical_params"] == {"friction": 0.5} - - monkeypatch.setattr(CFG, "code_sim_learning_persist_fit_data", False) - obj._persist_fit_trajectories() - assert len(list(out_dir.glob("*.pkl"))) == 2 - - -def test_fit_data_is_dumped_even_when_no_fit_runs(tmp_path) -> None: - """A cycle that declines to fit is exactly the one worth post-morteming. - - Persistence used to sit only inside the sysID fit, so the branch - that never ran was the branch whose data mattered: - run_20260817_171402 declined on a sweep returning one identical SSE - for every value of five parameters, and left nothing on disk to - explain it. Dumping where the data ARRIVES is what makes that - replayable. - """ - obj = object.__new__(AgentSimLearningApproach) - obj._physical_param_specs = [] - obj._identified_physical_params = {} - obj._get_log_dir = lambda: str(tmp_path) # type: ignore[method-assign] - obj._explainability_cache = {} - obj._sysid_fit_cache = {} - - # The one line of _learn_simulator this is about, with no fit after it. - obj._fit_trajectories = cast(List[LowLevelTrajectory], ["traj"]) - obj._persist_fit_trajectories("recorded") - - files = [f.name for f in (tmp_path / "fit_data").glob("*.pkl")] - assert files == ["fit_trajectories_000_recorded.pkl"] - with open(tmp_path / "fit_data" / files[0], "rb") as f: - assert pkl.load(f)["trajectories"] == ["traj"] - - # The wiring, not just the function: _learn_simulator runs on every - # cycle whether or not a fit follows, so the dump has to hang off it. - # Asserted on the source because calling _learn_simulator for real - # needs a whole synthesis session, and without this the test above - # passes with the call deleted. - source = inspect.getsource(AgentSimLearningApproach._learn_simulator) # pylint: disable=protected-access - assert "_persist_fit_trajectories(\"recorded\")" in source, \ - "the recorded-data dump is no longer wired into _learn_simulator" - - def test_base_sim_reference_provisioning() -> None: """Base-sim source rides the sandbox reference registry (so every session phase gets it) and the agent-visible paths map per backend.""" @@ -854,13 +677,11 @@ def test_base_sim_reference_provisioning() -> None: def test_synthesis_tool_names_are_run_python_only(): - """The learn session's only tool is run_python; the journal is a plain file - the agent edits, whatever the journal flag says.""" - stub = SimpleNamespace(_do_synthesize_samplers=False) - for use_journal in (True, False): - utils.reset_config({"agent_solve_use_journal": use_journal}) - names = AgentSimLearningApproach._get_synthesis_tool_names(stub) - assert names == ["run_python"] + """The synthesis session's only tool is run_python; the journal is a plain + file the agent edits.""" + names = AgentSimLearningApproach._get_synthesis_tool_names( + SimpleNamespace()) + assert names == ["run_python"] # --------------------------------------------------------------------------- @@ -900,11 +721,9 @@ def _make_checkpoint_stub(tmp_path, monkeypatch): obj._carried_physical_prior = {"lateral_friction": 0.5} obj._fit_evidence_history = {"vers_001": {"log_evidence": -1.0}} obj._identified_physical_sigma_points = [{"lateral_friction": 0.55}] - obj._sysid_fit_history = {} obj._residual_features = {"block": ["x"]} obj._current_simulator_version = "cycle_001_vers_003" obj._current_predicates_version = None - obj._current_samplers_version = None return obj, sandbox @@ -987,7 +806,6 @@ def test_rehydrate_rebuilds_simulator_from_restored_file( obj._learned_simulator = None obj._latent_init = None obj._fit_trajectories = [] - obj._synthesized_samplers = {} obj._base_env = SimpleNamespace(get_physical_param_info=lambda: {}) calls = [] monkeypatch.setattr( @@ -1008,8 +826,6 @@ def test_rehydrate_rebuilds_simulator_from_restored_file( monkeypatch.setattr(AgentSimLearningApproach, "_apply_identified_physical_params", lambda self, p: calls.append(("apply", dict(p)))) - monkeypatch.setattr(AgentSimLearningApproach, "_samplers_enabled", - staticmethod(lambda: False)) monkeypatch.setattr(AgentSimLearningApproach, "_rebuild_param_ensemble", lambda self: calls.append("ensemble")) # Checkpointed fitted params carry a stale name -> fall back to init. @@ -1040,39 +856,3 @@ def test_rehydrate_without_simulator_is_graceful(tmp_path, monkeypatch): lambda self, base: hooks.append(base)) obj._rehydrate_from_artifacts() assert hooks == [str(sandbox)] - - -def test_checkpoint_cycle_counter_semantics(tmp_path, monkeypatch): - """``_c`` files store c even when saved after the counter advanced, and - loading ``_None`` resumes at cycle 0 while ``_c`` resumes at c+1.""" - # pylint: disable=import-outside-toplevel - from predicators.approaches.agent_model_free_approach import \ - AgentModelFreeApproach - from predicators.structs import Dataset - utils.reset_config({ - "env": "cover", - "approach": "agent_model_free", - "seed": 0, - "approach_dir": str(tmp_path), - }) - obj = object.__new__(AgentModelFreeApproach) - obj._offline_dataset = Dataset([]) - obj._online_trajectories = [] - obj._run_id = "run" - obj._agent_session = None - monkeypatch.setattr(AgentModelFreeApproach, "_sync_tool_context", - lambda self: None) - # Post-offline checkpoint: counter 0, file _None. - obj._online_learning_cycle = 0 - obj.save(None) - # Cycle-3 checkpoint written AFTER the counter already advanced to 4 - # (the sim-learning subclass saves post-learn): the file must still - # denote cycle 3. - obj._online_learning_cycle = 4 - obj.save(3) - fresh = object.__new__(AgentModelFreeApproach) - fresh._agent_session = None - fresh.load(None) - assert fresh._online_learning_cycle == 0 - fresh.load(3) - assert fresh._online_learning_cycle == 4 diff --git a/tests/approaches/test_agent_sim_prompt_formatting.py b/tests/approaches/test_agent_sim_prompt_formatting.py deleted file mode 100644 index 80971abb8b..0000000000 --- a/tests/approaches/test_agent_sim_prompt_formatting.py +++ /dev/null @@ -1,471 +0,0 @@ -"""Tests for synthesis-prompt formatter helpers. - -These are pure-Python staticmethods (or `self`-less methods) on -``AgentSimLearningApproach`` and ``AgentSimPredicateInventionApproach`` -that render parts of the agent's first synthesis message. They were -added so the agent (a) knows the provenance of each interaction -trajectory and (b) gets reminded about prior-cycle files in the sandbox. -""" -# pylint: disable=protected-access,import-outside-toplevel,unused-import -from __future__ import annotations - -import numpy as np -import pytest - -# Bootstrap circular imports before pulling from predicators.approaches. -from predicators import utils -from predicators.structs import Action, LowLevelTrajectory, State, Task, Type - - -@pytest.fixture(name="approach_cls") -def _approach_cls(): - """Late-import the class so test collection is cheap.""" - from predicators.approaches.agent_sim_learning_approach import \ - AgentSimLearningApproach - return AgentSimLearningApproach - - -def _mk_traj(is_demo, - task_idx, - sim_v=None, - preds_v=None, - reward=None, - terminated=None): - """Build a 1-action trajectory with the given provenance tags.""" - cup_type = Type("cup_type", ["f"]) - cup = cup_type("cup") - states = [State({cup: [0.0]}), State({cup: [1.0]})] - actions = [Action(np.array([0.5]))] - return LowLevelTrajectory( - states, - actions, - _is_demo=is_demo, - _train_task_idx=task_idx, - _source_simulator_version=sim_v, - _source_predicates_version=preds_v, - _env_reward=reward, - _env_terminated=terminated, - ) - - -# ── _format_trajectory_listing ────────────────────────────────────── - - -def test_trajectory_listing_empty(approach_cls): - """Empty trajectory list short-circuits to an empty string.""" - assert approach_cls._format_trajectory_listing([]) == "" - - -def test_trajectory_listing_demo_has_no_provenance_tail(approach_cls): - """Demo trajectories never carry provenance - even if the tags are set, the - listing should still render them as plain demos for consistency with the - offline-data semantics.""" - trajs = [_mk_traj(is_demo=True, task_idx=0)] - out = approach_cls._format_trajectory_listing(trajs) - assert "[0] demo, task 0" in out - assert "generated using" not in out - - -def test_trajectory_listing_interaction_with_provenance(approach_cls): - """Interaction trajectories with provenance show the sim/preds tags.""" - trajs = [ - _mk_traj(is_demo=False, - task_idx=2, - sim_v="cycle_001_vers_004", - preds_v="cycle_001_vers_003"), - ] - out = approach_cls._format_trajectory_listing(trajs) - assert "[0] interaction, task 2" in out - assert "sim cycle_001_vers_004" in out - assert "predicates cycle_001_vers_003" in out - - -def test_trajectory_listing_supervisor_rejected(approach_cls): - """A rejected episode surfaces only through its (reward, terminated) - pair: terminated with solved=0 and no bonus in the reward. No - REJECTED flag or violation specifics reach the roster - the rules - live in the NL goal description, so the agent must infer the - violation from its own trajectory rather than be told it. - """ - trajs = [ - _mk_traj(is_demo=False, task_idx=0), - _mk_traj(is_demo=False, task_idx=3, reward=-0.05, terminated=True), - ] - out = approach_cls._format_trajectory_listing(trajs) - lines = [l for l in out.splitlines() if l.startswith(" [")] - assert "REJECTED" not in out - assert "env reward=-0.05 (solved=0)" in lines[1] - # No violation specifics leak into the roster line. - assert "domino" not in lines[1] - assert "push" not in lines[1].lower() - - -def test_trajectory_listing_env_reward(approach_cls): - """Evaluated episodes show the env reward with a success flag; a rejected - topple counts as solved=0 even though it terminated.""" - trajs = [ - _mk_traj(is_demo=False, task_idx=0, reward=0.85, terminated=True), - _mk_traj(is_demo=False, task_idx=1, reward=-0.05, terminated=True), - _mk_traj(is_demo=False, task_idx=2), # never evaluated - ] - out = approach_cls._format_trajectory_listing(trajs) - lines = [l for l in out.splitlines() if l.startswith(" [")] - assert "env reward=0.85 (solved=1)" in lines[0] - assert "env reward=-0.05 (solved=0)" in lines[1] - assert "REJECTED" not in lines[1] - assert "reward" not in lines[2] - - -def test_trajectory_listing_partial_provenance(approach_cls): - """Only ``source_simulator_version`` set: list just the sim tag. - - No stray ``, `` may appear from the missing predicate half of the - provenance pair. - """ - trajs = [_mk_traj(is_demo=False, task_idx=1, sim_v="cycle_001_vers_007")] - out = approach_cls._format_trajectory_listing(trajs) - line = [l for l in out.splitlines() if l.startswith(" [0]")][0] - assert "sim cycle_001_vers_007" in line - assert "predicates" not in line - - -# ── _format_prior_state_block ──────────────────────────────────────── - - -def test_prior_state_block_empty_when_no_files(approach_cls, tmp_path): - """Neither simulator.py nor predicates.py exists → empty block.""" - out = approach_cls._format_prior_state_block(None, str(tmp_path)) - assert out == "" - - -def test_prior_state_block_simulator_only(approach_cls, tmp_path): - """Only simulator.py exists → block mentions it and not predicates.py.""" - (tmp_path / "simulator.py").write_text("# sim") - out = approach_cls._format_prior_state_block(None, str(tmp_path)) - assert "`./simulator.py`" in out - assert "`./predicates.py`" not in out - # Always points at the versioned-snapshot dirs for cross-reference. - assert "./simulator_versions/" in out - - -def test_prior_state_block_both_files(approach_cls, tmp_path): - """Both files exist → block lists them joined with ' and '.""" - (tmp_path / "simulator.py").write_text("# sim") - (tmp_path / "predicates.py").write_text("LEARNED_PREDICATES = []") - out = approach_cls._format_prior_state_block(None, str(tmp_path)) - assert "`./simulator.py` and `./predicates.py`" in out - # Soft language so the agent isn't forbidden from a fresh rewrite. - assert "fresh rewrite is fine" in out - - -# ── _format_goal_nl_block (predicate-invention subclass) ──────────── - - -def test_goal_nl_block_empty_when_no_tasks_have_goal_nl(): - """No ``goal_nl`` populated → empty block (no header).""" - from predicators.approaches.agent_sim_predicate_invention_approach import \ - AgentSimPredicateInventionApproach - fake_self = type( - "_FakeApproach", - (), - { - "_train_tasks": [type("_T", (), {"goal_nl": None})()] * 2, - }, - )() - out = AgentSimPredicateInventionApproach._format_goal_nl_block(fake_self) - assert out == "" - - -def test_goal_nl_block_dedups_identical_goals(): - """Same NL goal across tasks shows up once, with the single-task header.""" - from predicators.approaches.agent_sim_predicate_invention_approach import \ - AgentSimPredicateInventionApproach - fake_task = type("_T", (), {"goal_nl": "boil the water"}) - fake_self = type( - "_FakeApproach", - (), - { - "_train_tasks": [fake_task() for _ in range(3)], - }, - )() - out = AgentSimPredicateInventionApproach._format_goal_nl_block(fake_self) - assert out.startswith("Goal (natural language): boil the water") - # Trailing blank line separates from the next paragraph in the prompt. - assert out.endswith("\n\n") - - -def test_goal_nl_block_multiple_distinct_goals(): - """Distinct goals across tasks render as a bulleted list.""" - from predicators.approaches.agent_sim_predicate_invention_approach import \ - AgentSimPredicateInventionApproach - tasks = [ - type("_T1", (), {"goal_nl": "boil the water"})(), - type("_T2", (), {"goal_nl": "stack the cups"})(), - ] - fake_self = type("_FakeApproach", (), {"_train_tasks": tasks})() - out = AgentSimPredicateInventionApproach._format_goal_nl_block(fake_self) - assert "Goals across train tasks (natural language):" in out - assert " - boil the water" in out - assert " - stack the cups" in out - - -# ── _build_synthesis_system_prompt (FO vs PO rule signature) ──────── -# These render the whole synthesis system prompt. The method only touches -# ``self`` through pure no-state helpers (``_rule_signature_section``, -# ``_residual_rule_signature``, ``_extra_synthesis_system_prompt``), so a -# bare instance via ``object.__new__`` is enough to render it. - - -def _render_prompt(cls): - from predicators.approaches.agent_sim_learning_approach import \ - AgentSimLearningApproach - return AgentSimLearningApproach._build_synthesis_system_prompt( - object.__new__(cls)) - - -def test_synthesis_prompt_no_leftover_placeholders(approach_cls): - """Every templated placeholder is substituted in the rendered prompt.""" - prompt = _render_prompt(approach_cls) - for placeholder in ("__RULE_SIGNATURE_SECTION__", - "__RESIDUAL_RULE_SIGNATURE__", - "__SYNTHESIS_PROMPT_EXTRA__"): - assert placeholder not in prompt - - -def test_synthesis_prompt_sections_not_duplicated(approach_cls): - """The system prompt has exactly one of each major section. - - Guards against the bad-merge artifact that duplicated the Tools / - Refinement / Plan-format blocks (and double-injected the extra). - """ - prompt = _render_prompt(approach_cls) - for header in ("## `simulator.py`: a simulator subclass", - "## Step and restoration behavior", "## Plan format", - "## Fit and validate complete rollouts"): - assert prompt.count(header) == 1, (header, prompt.count(header)) - - -def test_fo_prompt_uses_subclass_contract(approach_cls): - """Fully observable models use the same subclass contract as PO ones.""" - prompt = _render_prompt(approach_cls) - assert "class MyDynamics(BaseSimulator):" in prompt - assert "RESIDUAL_ENV = MyDynamics" in prompt - assert "def _domain_specific_step(self):" in prompt - assert "RESIDUAL_RULES" not in prompt - assert "def residual_rule(" not in prompt - assert "## Hidden model state" not in prompt - - -def test_po_prompt_uses_subclass_memory_contract(): - """Both PO approaches receive one canonical model-state callback. - - A competing rule signature must not reappear beside the subclass - contract, and only predicate invention adds classifier guidance. - """ - import re - - from predicators.approaches.agent_sim_learning_approach import \ - AgentSimLearningApproach - from predicators.approaches.agent_sim_predicate_invention_approach import \ - AgentSimPredicateInventionApproach - from predicators.settings import CFG - old_flag = CFG.partially_observable - CFG.partially_observable = True - try: - for cls in (AgentSimLearningApproach, - AgentSimPredicateInventionApproach): - prompt = _render_prompt(cls) - assert "class MyDynamics(BaseSimulator):" in prompt - assert ("def update_model_state(cls, observation, model_state, " - "params, action):" in prompt) - assert "RESIDUAL_RULES" not in prompt - assert "LATENT_INIT" not in prompt - assert "def residual_rule(" not in prompt - # Memory guidance is injected exactly once. - headers = re.findall(r"(?m)^## Hidden model state$", prompt) - assert len(headers) == 1, cls - # The predicate-side latent guidance is invention-only. - has_pred_section = "### Predicate signature" in prompt - assert has_pred_section == ( - cls is AgentSimPredicateInventionApproach), cls - finally: - CFG.partially_observable = old_flag - - -# ── _make_evaluate_trajectory_fn / _format_objective_block ────────── - - -def test_evaluate_trajectory_helper(approach_cls): - """The exec-ns evaluate_trajectory helper returns verdict dicts (never the - evaluator), labels Action inputs by their producing options, and rejects - bad task indices.""" - from types import SimpleNamespace - - from predicators.structs import TaskEvaluator - - class _RecordingEvaluator(TaskEvaluator): - """Rejecting evaluator that records the labels it saw.""" - - def __init__(self): - super().__init__(set()) # empty goal: terminated is True - self.seen_options = None - - def _certify(self, states, step_options, sim_env=None): - self.seen_options = step_options - return False, "nope" - - evaluator = _RecordingEvaluator() - - cup_type = Type("cup_type", ["f"]) - cup = cup_type("cup") - states = [State({cup: [0.0]}), State({cup: [1.0]})] - # _option_model mirrors the real approach attribute the helper reads - # for the certificate's sim_env (None => kinematics-only scoring). - stub = SimpleNamespace(_train_tasks=[ - Task(states[0], set(), evaluator=evaluator), - Task(states[0], set()), - ], - _option_model=None) - fn = approach_cls._make_evaluate_trajectory_fn(stub) - push = utils.SingletonParameterizedOption( - "Push", lambda s, m, o, p: Action(np.zeros(1, dtype=np.float32))) - act = Action(np.zeros(1, dtype=np.float32)) - act.set_option(push.ground([], np.zeros(0, dtype=np.float32))) - - verdict = fn(states, [act], task_idx=0) - # The public verdict includes a replay note, but excludes internal - # legitimacy details and goal-atom termination. This evaluator does - # not replay physics, so the note is empty. - assert verdict == { - "reward": 0.0, # bonus gated by the internal rejection - "solved": False, - "note": "", - } - # Labels are (name, objects, params) triples since plan-capture - # gating started matching on exact params. - assert evaluator.seen_options == [("Push", (), ())] - # Pre-built labels pass through unchanged. - fn(states, [("Push", ("robot", ))], task_idx=0) - assert evaluator.seen_options == [("Push", ("robot", ))] - with pytest.raises(ValueError, match="no task evaluator"): - fn(states, None, task_idx=1) - with pytest.raises(ValueError, match="out of range"): - fn(states, None, task_idx=2) - with pytest.raises(ValueError, match="non-empty"): - fn([], None, task_idx=0) - - -def test_evaluate_trajectory_physics_sweep(approach_cls): - """physics_sweep=True scores the sequence at every physics-margin point on - a fresh env at that physics and reports the fraction scored solved; with no - points to sweep it says so.""" - import contextlib - import functools - from types import SimpleNamespace - - from predicators.structs import TaskEvaluator - - physics = {"friction": 0.5} - - class _FrictionEvaluator(TaskEvaluator): - """Certifies only when the (swept) friction is at least 0.5.""" - - def __init__(self): - super().__init__(set()) - - def _certify(self, states, step_options, sim_env=None): - return physics["friction"] >= 0.5, "friction" - - cup_type = Type("cup_type", ["f"]) - cup = cup_type("cup") - states = [State({cup: [0.0]}), State({cup: [1.0]})] - seen = [] - - @contextlib.contextmanager - def _scope(physical_overrides=None): - seen.append(dict(physical_overrides or {})) - prev = dict(physics) - physics.update(physical_overrides or {}) - try: - yield - finally: - physics.clear() - physics.update(prev) - - stub = SimpleNamespace( - _train_tasks=[Task(states[0], set(), evaluator=_FrictionEvaluator())], - _option_model=None, - _identified_physical_sigma_points=[{ - "friction": 0.4 - }, { - "friction": 0.5 - }, { - "friction": 0.6 - }], - _fresh_validation_env_scope=_scope) - stub._sweep_evaluation = functools.partial(approach_cls._sweep_evaluation, - stub) - fn = approach_cls._make_evaluate_trajectory_fn(stub) - plain = fn(states, None, task_idx=0) - assert "sweep" not in plain and plain["solved"] is True - swept = fn(states, None, task_idx=0, physics_sweep=True) - assert swept["solved"] is True - sweep = swept["sweep"] - assert [p["solved"] for p in sweep["points"]] == [False, True, True] - assert sweep["solved_fraction"] == pytest.approx(2 / 3) - assert sweep["certified"] is False - assert seen == [{"friction": 0.4}, {"friction": 0.5}, {"friction": 0.6}] - assert physics == {"friction": 0.5} # the scope restored the physics - stub._identified_physical_sigma_points = [] - assert fn(states, None, task_idx=0, physics_sweep=True)["sweep"] is None - - -def test_format_objective_block(approach_cls): - """The objective block renders the first stated objective and is empty when - no evaluator states one.""" - from types import SimpleNamespace - - from predicators.structs import TaskEvaluator - - class _StatingEvaluator(TaskEvaluator): - """Evaluator with a public objective statement.""" - - def objective_description(self): - return "Topple the target legitimately; each blue costs 0.05." - - cup_type = Type("cup_type", ["f"]) - init = State({cup_type("cup"): [0.0]}) - - def _task(evaluator=None): - return Task(init, set(), evaluator=evaluator) - - fmt = approach_cls._format_objective_block - assert fmt(SimpleNamespace(_train_tasks=[])) == "" - assert fmt(SimpleNamespace(_train_tasks=[_task()])) == "" - assert fmt( - SimpleNamespace(_train_tasks=[_task(TaskEvaluator(set()))])) == "" - out = fmt( - SimpleNamespace( - _train_tasks=[_task(), _task(_StatingEvaluator(set()))])) - assert "## Task objective (env ground-truth reward)" in out - assert "each blue costs 0.05" in out - assert "evaluate_trajectory" in out - - -def test_learn_message_ships_goal_required_mechanisms_as_hypotheses(): - """The learn message distinguishes a hypothesis the goal can do without - (record, do not ship) from one the goal REQUIRES (ship as a labelled - hypothesis with declared ParamSpecs and a first-ranked confirming - experiment). - - The rule lives in the learn system prompt's deliverables section. - """ - from predicators.agent_sdk import learn_prompts - prompt = learn_prompts.build_learn_system_prompt( - partially_observable=False, - residual_rule_signature="def residual_rule(state, updates, params):", - scene_viz_hint="render the scene") - assert "When the goal requires it" in prompt - assert "labelled hypothesis" in prompt - assert "first entry of `./open_questions.md`" in prompt - assert "A GO that rests on a hypothesized mechanism" in prompt diff --git a/tests/approaches/test_agent_solve_restart.py b/tests/approaches/test_agent_solve_restart.py deleted file mode 100644 index 682853c753..0000000000 --- a/tests/approaches/test_agent_solve_restart.py +++ /dev/null @@ -1,459 +0,0 @@ -"""Tests for the time-boxed restart loop in AgentModelBasedApproach._solve. - -The loop runs up to ``agent_solve_max_attempts`` solve attempts, each on -a fresh conversation when ``agent_solve_fresh_context`` is set: a -validated (evaluator-solved) capture returns immediately, best-effort -captures are banked and ranked by evaluator reward, journal auto-entries -record every attempt, and total failure re-raises the last error. -""" -# pylint: disable=protected-access -import time - -import numpy as np -import pytest -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk import journal as journal_mod -from predicators.agent_sdk.session_base import AgentSessionFatalError -from predicators.approaches import ApproachFailure, ApproachTimeout -from predicators.approaches.agent_model_based_approach import \ - AgentModelBasedApproach, _CaptureInfo -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, Type - -_block_type = Type("block", ["x"]) -_block0 = Object("block0", _block_type) -_Reached = Predicate("Reached", [_block_type], - lambda s, o: s.get(o[0], "x") >= 0.9) - - -def _noop_policy(_s, _m, _o, _p): - return Action(np.zeros(1, dtype=np.float32)) - - -_Move = ParameterizedOption( - "Move", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: False, -) - - -def _make_approach(overrides=None, sandbox_dir=None): - state = State({_block0: np.array([0.0], dtype=np.float32)}) - task = Task(state, {GroundAtom(_Reached, [_block0])}) - config = { - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "seed": 42, - } - config.update(overrides or {}) - utils.reset_config(config) - approach = AgentModelBasedApproach( - initial_predicates={_Reached}, - initial_options={_Move}, - types={_block_type}, - action_space=Box(low=-1, high=1, shape=(1, )), - train_tasks=[task], - option_model=None, - ) - if sandbox_dir is not None: - approach._tool_context.sandbox_dir = sandbox_dir - return approach, task - - -class _AttemptScript: - """Scripted per-attempt outcomes for a stubbed _solve_attempt. - - Each script item is ``("validated", reward)``, ``("best_effort", - reward)``, ``("fail", message)`` (raises ApproachFailure), or - ``("error", exception)`` (raises that exception). The stub sets - ``_last_capture_info`` exactly as _consume_validated_plan would. - """ - - def __init__(self, approach, script): - self.approach = approach - self.script = list(script) - self.calls = 0 - self.policies = [] - - def __call__(self, _task): - kind, value = self.script[self.calls] - self.calls += 1 - if kind == "fail": - raise ApproachFailure(value) - if kind == "error": - raise value - self.approach._last_capture_info = _CaptureInfo( - validated=(kind == "validated"), - reward=value, - plan_lines=[f"Move(block0:block)[{0.9 + self.calls / 100.0}]"], - validation_summary="validation: 5/6 rollouts ok") - policy = lambda _s: Action(np.zeros(1, dtype=np.float32)) - self.policies.append(policy) - return policy - - -def test_validated_capture_returns_immediately(): - """A validated first attempt short-circuits the loop.""" - approach, task = _make_approach({ - "agent_solve_max_attempts": 3, - "agent_solve_fresh_context": True, - }) - closes = [] - approach._close_agent_session = lambda: closes.append(1) - script = _AttemptScript(approach, [("validated", 0.95)]) - approach._solve_attempt = script - policy = approach._solve(task, timeout=10) - assert policy is script.policies[0] - assert script.calls == 1 - assert len(closes) == 1 # fresh context for the (only) attempt - - -def test_best_effort_banked_and_best_reward_wins(): - """Best-effort captures are ranked by evaluator reward across attempts.""" - approach, task = _make_approach({"agent_solve_max_attempts": 3}) - script = _AttemptScript(approach, [ - ("best_effort", -0.10), - ("best_effort", 0.40), - ("fail", "attempt 3 found nothing"), - ]) - approach._solve_attempt = script - policy = approach._solve(task, timeout=10) - assert script.calls == 3 - assert policy is script.policies[1] # the reward-0.40 capture - - -def test_validated_on_later_attempt_beats_banked_best_effort(): - """A later validated capture wins over an earlier banked one.""" - approach, task = _make_approach({"agent_solve_max_attempts": 3}) - script = _AttemptScript(approach, [ - ("best_effort", 0.90), - ("validated", 0.95), - ]) - approach._solve_attempt = script - policy = approach._solve(task, timeout=10) - assert script.calls == 2 - assert policy is script.policies[1] - - -def test_all_attempts_fail_reraises(): - """With nothing captured anywhere, the last failure propagates.""" - approach, task = _make_approach({"agent_solve_max_attempts": 2}) - script = _AttemptScript(approach, [ - ("fail", "first"), - ("fail", "second"), - ]) - approach._solve_attempt = script - with pytest.raises(ApproachFailure, match="second"): - approach._solve(task, timeout=10) - assert script.calls == 2 - - -def test_no_fresh_context_keeps_session(): - """agent_solve_fresh_context=False never closes the session.""" - approach, task = _make_approach({ - "agent_solve_max_attempts": 2, - "agent_solve_fresh_context": False, - }) - closes = [] - approach._close_agent_session = lambda: closes.append(1) - script = _AttemptScript(approach, [("fail", "a"), ("validated", 0.9)]) - approach._solve_attempt = script - approach._solve(task, timeout=10) - assert not closes - - -def test_journal_auto_entries_record_each_attempt(tmp_path): - """Every attempt leaves an auto entry with outcome and plan.""" - approach, task = _make_approach( - { - "agent_solve_max_attempts": 2, - "agent_solve_use_journal": True, - }, - sandbox_dir=str(tmp_path)) - approach._tool_context.test_task_idx = 0 - script = _AttemptScript(approach, [ - ("fail", "nothing"), - ("best_effort", -0.05), - ]) - approach._solve_attempt = script - approach._solve(task, timeout=10) - content = journal_mod.read_journal(str(tmp_path), - filename=journal_mod.ATTEMPTS_FILENAME) - assert "### cycle 0 task 0 attempt 1/2 (auto)" in content - assert "- outcome: no capture" in content - assert "### cycle 0 task 0 attempt 2/2 (auto)" in content - assert "best-effort capture (evaluator reward -0.05)" in content - assert "Move(block0:block)" in content - # The capture-time validation record travels into the auto entry so a - # later fresh-context attempt sees how reliable the capture was. - assert "- validation: 5/6 rollouts ok" in content - # Task context (goal + init state dict, the prompt's own - # representation) is a dedicated entry written once, at the TOP of - # the task's section - before any attempt entry. - assert content.count("### cycle 0 task 0 goal + initial state (auto)") == 1 - assert content.count("- goal: Reached(block0:block)") == 1 - assert content.count("- initial state features:") == 1 - assert "'block0:block'" in content - assert "'x'" in content - assert content.index("goal + initial state") < content.index( - "### cycle 0 task 0 attempt 1/2") - - -def test_attempt_bookkeeping_reset_per_attempt(): - """attempt_index/rollout counter/deadline are set and cleared.""" - approach, task = _make_approach({ - "agent_solve_max_attempts": 2, - "agent_solve_attempt_wall_clock": 2700, - }) - ctx = approach._tool_context - seen = [] - - def _attempt(_task): - seen.append( - (ctx.attempt_index, ctx.attempt_rollout_count, ctx.attempt_deadline - is not None)) - ctx.attempt_rollout_count += 7 - raise ApproachFailure("no capture") - - approach._solve_attempt = _attempt - with pytest.raises(ApproachFailure): - approach._solve(task, timeout=10) - assert seen == [(1, 0, True), (2, 0, True)] - assert ctx.attempt_deadline is None - assert ctx.attempt_start is None - # attempt_index resets too: a stale index would mislabel journal - # entries recorded outside any attempt. - assert ctx.attempt_index == 0 - - -def test_attempt_wall_spent(): - """_attempt_wall_spent trips BEFORE the deadline, at the spent floor. - - Agents that watch the [budget] footer end their query shortly before - the deadline (run_20260718_125643 queries 002-003); an attempt whose - remaining tools would only refuse must be labelled spent, not "no - submission". - """ - approach, _task = _make_approach({"agent_solve_attempt_wall_clock": 2700}) - ctx = approach._tool_context - # No deadline set: never spent. - assert not approach._attempt_wall_spent() - ctx.attempt_deadline = time.monotonic() - 1.0 - assert approach._attempt_wall_spent() - # 2 min left of 45: below the 20% floor (540s), counts as spent. - ctx.attempt_deadline = time.monotonic() + 120.0 - assert approach._attempt_wall_spent() - # 20 min left: plenty of budget still on the clock. - ctx.attempt_deadline = time.monotonic() + 1200.0 - assert not approach._attempt_wall_spent() - - -def test_nudge_suspends_and_restores_attempt_deadline(): - """The nudge must not permanently disarm the wall clock. - - It suspends the deadline so its own submission is not refused, but - a helper that leaked the None would let any future mid-attempt - caller run unbounded - the exact runaway the time-box targets. - """ - approach, _task = _make_approach({}) - ctx = approach._tool_context - deadline = time.monotonic() + 1000.0 - ctx.attempt_deadline = deadline - approach._query_agent_sync = lambda *a, **k: [] - policy = approach._nudge_final_submission() - assert policy is None - assert ctx.attempt_deadline == deadline - assert ctx.capture_best_effort_plan is False - - -def test_unexpected_error_salvages_banked_capture(): - """A non-ApproachFailure error executes the banked best-effort capture - instead of forfeiting it, and stops further attempts.""" - approach, task = _make_approach({"agent_solve_max_attempts": 3}) - ctx = approach._tool_context - script = _AttemptScript(approach, [ - ("best_effort", 0.40), - ("error", RuntimeError("pybullet exploded")), - ]) - approach._solve_attempt = script - policy = approach._solve(task, timeout=10) - assert policy is script.policies[0] - assert script.calls == 2 # attempt 3 never runs - # Bookkeeping fully cleaned despite the unexpected error. - assert ctx.attempt_start is None - assert ctx.attempt_deadline is None - assert ctx.attempt_index == 0 - - -def test_fatal_session_error_reraises_even_with_banked_capture(): - """AgentSessionFatalError is never salvaged by a banked capture: the - session backend is unusable, so the run must terminate (bookkeeping still - cleaned by the finally).""" - approach, task = _make_approach({"agent_solve_max_attempts": 3}) - ctx = approach._tool_context - script = _AttemptScript(approach, [ - ("best_effort", 0.40), - ("error", AgentSessionFatalError("3 consecutive agent queries died")), - ]) - approach._solve_attempt = script - with pytest.raises(AgentSessionFatalError): - approach._solve(task, timeout=10) - assert script.calls == 2 # attempt 3 never runs - assert ctx.attempt_start is None - assert ctx.attempt_deadline is None - assert ctx.attempt_index == 0 - - -def test_unexpected_error_without_bank_reraises_after_cleanup(): - """ApproachTimeout (an ApproachFailure SIBLING) propagates, but only. - - after attempt bookkeeping is cleared - stale fields would pollute - later sessions sharing the ToolContext. - """ - approach, task = _make_approach({ - "agent_solve_max_attempts": 3, - "agent_solve_attempt_wall_clock": 2700, - }) - ctx = approach._tool_context - script = _AttemptScript(approach, [("error", ApproachTimeout("slow"))]) - approach._solve_attempt = script - with pytest.raises(ApproachTimeout): - approach._solve(task, timeout=10) - assert ctx.attempt_start is None - assert ctx.attempt_deadline is None - assert ctx.attempt_index == 0 - - -def test_journal_task_context_written_once_even_on_resolve(tmp_path): - """Re-entering _solve for the same task (mid-episode replan) must not - duplicate the goal + init-state entry.""" - approach, task = _make_approach( - { - "agent_solve_max_attempts": 1, - "agent_solve_use_journal": True, - }, - sandbox_dir=str(tmp_path)) - approach._tool_context.test_task_idx = 0 - script = _AttemptScript(approach, [("validated", 0.95), - ("validated", 0.96)]) - approach._solve_attempt = script - approach._solve(task, timeout=10) - approach._solve(task, timeout=10) - content = journal_mod.read_journal(str(tmp_path), - filename=journal_mod.ATTEMPTS_FILENAME) - assert content.count("### cycle 0 task 0 goal + initial state (auto)") == 1 - assert content.index("- goal:") < content.index("- outcome:") - - -def test_test_phase_journal_archived_and_rolled_back(tmp_path): - """Learning content persists across evaluations; each evaluation's own - additions (harness attempt-log entries and agent journal notes) are - archived outside the sandbox, then rolled back so the next evaluation - starts from learning knowledge only (no test-task leaks).""" - sandbox = tmp_path / "sandbox" - log_dir = tmp_path / "run_logs" - approach, task = _make_approach( - { - "agent_solve_max_attempts": 1, - "agent_solve_use_journal": True, - "log_file": str(log_dir), - }, - sandbox_dir=str(sandbox)) - ctx = approach._tool_context - attempts = journal_mod.ATTEMPTS_FILENAME - # A learning-phase note, written by the agent before any evaluation. - sandbox.mkdir(parents=True, exist_ok=True) - (sandbox / journal_mod.JOURNAL_FILENAME).write_text( - "### learn cycle notes\n- learning fact\n", encoding="utf-8") - # First evaluation: one test-task solve writes attempt-log entries - # and the agent adds a note. - approach.begin_test_phase() - ctx.test_task_idx = 0 - script = _AttemptScript(approach, [("validated", 0.95), - ("validated", 0.96)]) - approach._solve_attempt = script - approach._solve(task, timeout=10) - with open(sandbox / journal_mod.JOURNAL_FILENAME, "a", - encoding="utf-8") as f: - f.write("### task 0 attempt 1\n- eval-time note\n") - content = journal_mod.read_journal(str(sandbox), filename=attempts) - assert "### cycle 0 task 0 goal + initial state (auto)" in content - assert "- learning fact" in journal_mod.read_journal(str(sandbox)) - approach.end_test_phase() - # Rolled back: learning content survives, eval additions are gone. - assert "task 0" not in journal_mod.read_journal(str(sandbox), - filename=attempts) - notes = journal_mod.read_journal(str(sandbox)) - assert "- learning fact" in notes - assert "eval-time note" not in notes - # Both files were archived outside the sandbox first, one copy per - # evaluation phase. This first evaluation precedes any online - # learning, so it archives as the initial test. - archived = (log_dir / "attempts_eval_initial.md").read_text() - assert "### cycle 0 task 0 goal + initial state (auto)" in archived - archived_notes = (log_dir / "journal_eval_initial.md").read_text() - assert "- learning fact" in archived_notes - assert "eval-time note" in archived_notes - # Second evaluation on the same task, after a learning phase advanced - # the cycle: the context-entry dedup key was rolled back too, so the - # goal + init entry is re-written (else the attempt records would be - # uninterpretable). - approach._online_learning_cycle = 1 - approach.begin_test_phase() - ctx.test_task_idx = 0 - approach._solve(task, timeout=10) - content = journal_mod.read_journal(str(sandbox), filename=attempts) - assert content.count("### cycle 0 task 0 goal + initial state (auto)") == 1 - approach.end_test_phase() - # The second evaluation ran after cycle 0's learn advanced the - # counter to 1, so it archives under the 0-based cycle it evaluates. - assert sorted(p.name for p in log_dir.glob("attempts_eval*.md")) == [ - "attempts_eval_cycle0.md", "attempts_eval_initial.md" - ] - assert journal_mod.read_raw(str(sandbox)) is not None - assert "task 0" not in journal_mod.read_journal(str(sandbox), - filename=attempts) - - -def test_test_phase_journal_rollback_noop_without_journal(tmp_path): - """With the journal disabled the phase hooks touch nothing.""" - approach, _task = _make_approach({"agent_solve_use_journal": False}, - sandbox_dir=str(tmp_path)) - approach.begin_test_phase() - approach.end_test_phase() - assert journal_mod.read_raw(str(tmp_path)) is None - - -def test_journal_records_best_refused_submission(tmp_path): - """An attempt with no capture journals the best refused submission the - tools stashed, so the plan survives the fresh-context restart.""" - approach, task = _make_approach( - { - "agent_solve_max_attempts": 1, - "agent_solve_use_journal": True, - }, - sandbox_dir=str(tmp_path)) - ctx = approach._tool_context - ctx.test_task_idx = 0 - - def _attempt(_task): - ctx.best_uncaptured_plan_lines = ["Move(block0:block)[0.87]"] - ctx.best_uncaptured_reward = -0.05 - raise ApproachFailure("no capture") - - approach._solve_attempt = _attempt - with pytest.raises(ApproachFailure): - approach._solve(task, timeout=10) - content = journal_mod.read_journal(str(tmp_path), - filename=journal_mod.ATTEMPTS_FILENAME) - assert ("- best refused submission (evaluator reward -0.05, " - "not captured):") in content - assert "Move(block0:block)[0.87]" in content diff --git a/tests/approaches/test_continual_comparison_approach.py b/tests/approaches/test_continual_comparison_approach.py index e67665d07a..3db04992cb 100644 --- a/tests/approaches/test_continual_comparison_approach.py +++ b/tests/approaches/test_continual_comparison_approach.py @@ -1,4 +1,5 @@ """Continual comparison contracts exercised through real play tools.""" +import dataclasses import os import re import shlex @@ -20,15 +21,31 @@ from predicators.run.level_players import create_level_player from predicators.settings import CFG from predicators.structs import Dataset -from scripts.cluster_utils import config_to_cmd_flags, generate_run_configs +from scripts.cluster_utils import SingleSeedRunConfig, config_to_cmd_flags, \ + generate_run_configs from tests.approaches.test_agent_continual_approach import _call, _config, \ _result -# The benchmark sweep: eight arms on the five benchmark settings, three -# seeds each (Sept 18, 2026). -CONFIG = "predicatorv3/continual_eight_agent_noisy_sweep.yaml" -ARM_COUNT = 8 -SEEDS = {0, 1, 2} +# The benchmark: seven arms (six approach classes; the scene-package arm +# shares the model-free class) on the five settings, five seeds each. +CONFIG = "empiric/benchmark.yaml" +ARM_COUNT = 7 +CLASS_COUNT = 6 +SEEDS = set(range(5)) +# Comparison arms outside the benchmark. Their menu entries set only the +# model, like the standalone arm's, so they reuse its run config. +UNBENCHMARKED = ("agent_continual_scene_only", "agent_continual_zero_shot") + + +def _arm_config(domain: str, approach: str) -> SingleSeedRunConfig: + """The benchmark run config of ``approach`` on ``domain``.""" + if approach in UNBENCHMARKED: + base = _arm_config(domain, "agent_continual_program_world_model") + return dataclasses.replace(base, approach=approach) + cfg = next(c for c in generate_run_configs(CONFIG, False) + if c.env == f"pybullet_{domain}" and c.approach == approach) + assert isinstance(cfg, SingleSeedRunConfig) + return cfg @pytest.fixture(autouse=True) @@ -58,10 +75,10 @@ def connect(*args: Any, **kwargs: Any) -> int: def test_comparison_config_matches_existing_domains(monkeypatch: Any) -> None: - """Eight arms share each domain's settings and run paired seeds.""" + """The seven arms share each domain's settings and run paired seeds.""" new = list(generate_run_configs(CONFIG, False)) approaches = {c.approach for c in new} - assert len(approaches) == ARM_COUNT + assert len(approaches) == CLASS_COUNT assert len(new) == ARM_COUNT * 5 * len(SEEDS) seeds: Dict[Tuple[str, str], Set[int]] = {} for cfg in new: @@ -123,8 +140,7 @@ def test_no_uncertainty_rejects_smoothing(flag: str) -> None: def test_ablation_play_tools(tmp_path: Any, monkeypatch: Any, arm: str, domain: str) -> None: """No-uncertainty uses raw observations; tools enforce arm restrictions.""" - cfg = next(c for c in generate_run_configs(CONFIG, False) - if c.env == f"pybullet_{domain}" and c.approach.endswith(arm)) + cfg = _arm_config(domain, f"agent_continual_{arm}") _config( tmp_path, **{ **{k: v @@ -286,8 +302,7 @@ def _sandbox_listing(ctx: Any) -> str: def _configure_domain_comparison(tmp_path: Any, domain: str, approach: str) -> str: """Use the benchmark's real domain settings with local test outputs.""" - cfg = next(c for c in generate_run_configs(CONFIG, False) - if c.env == f"pybullet_{domain}" and c.approach == approach) + cfg = _arm_config(domain, approach) _config( tmp_path, **{ **{k: v diff --git a/tests/approaches/test_evaluate_trajectory_helper.py b/tests/approaches/test_evaluate_trajectory_helper.py new file mode 100644 index 0000000000..3b0209d49b --- /dev/null +++ b/tests/approaches/test_evaluate_trajectory_helper.py @@ -0,0 +1,159 @@ +"""Tests for the ``evaluate_trajectory`` helper the synthesis namespace +offers.""" +# pylint: disable=protected-access,import-outside-toplevel,unused-import +from __future__ import annotations + +import numpy as np +import pytest + +# Bootstrap circular imports before pulling from predicators.approaches. +from predicators import utils +from predicators.structs import Action, State, Task, Type + + +@pytest.fixture(name="approach_cls") +def _approach_cls(): + """Late-import the class so test collection is cheap.""" + from predicators.approaches.agent_sim_learning_approach import \ + AgentSimLearningApproach + return AgentSimLearningApproach + + +# ── _format_trajectory_listing ────────────────────────────────────── + +# ── _format_prior_state_block ──────────────────────────────────────── + +# ── _format_goal_nl_block (predicate-invention subclass) ──────────── + +# ── _build_synthesis_system_prompt (FO vs PO rule signature) ──────── +# These render the whole synthesis system prompt. The method only touches +# ``self`` through pure no-state helpers (``_rule_signature_section``, +# ``_residual_rule_signature``, ``_extra_synthesis_system_prompt``), so a +# bare instance via ``object.__new__`` is enough to render it. + +# ── _make_evaluate_trajectory_fn / _format_objective_block ────────── + + +def test_evaluate_trajectory_helper(approach_cls): + """The exec-ns evaluate_trajectory helper returns verdict dicts (never the + evaluator), labels Action inputs by their producing options, and rejects + bad task indices.""" + from types import SimpleNamespace + + from predicators.structs import TaskEvaluator + + class _RecordingEvaluator(TaskEvaluator): + """Rejecting evaluator that records the labels it saw.""" + + def __init__(self): + super().__init__(set()) # empty goal: terminated is True + self.seen_options = None + + def _certify(self, states, step_options, sim_env=None): + self.seen_options = step_options + return False, "nope" + + evaluator = _RecordingEvaluator() + + cup_type = Type("cup_type", ["f"]) + cup = cup_type("cup") + states = [State({cup: [0.0]}), State({cup: [1.0]})] + # _option_model mirrors the real approach attribute the helper reads + # for the certificate's sim_env (None => kinematics-only scoring). + stub = SimpleNamespace(_train_tasks=[ + Task(states[0], set(), evaluator=evaluator), + Task(states[0], set()), + ], + _option_model=None) + fn = approach_cls._make_evaluate_trajectory_fn(stub) + push = utils.SingletonParameterizedOption( + "Push", lambda s, m, o, p: Action(np.zeros(1, dtype=np.float32))) + act = Action(np.zeros(1, dtype=np.float32)) + act.set_option(push.ground([], np.zeros(0, dtype=np.float32))) + + verdict = fn(states, [act], task_idx=0) + # The public verdict includes a replay note, but excludes internal + # legitimacy details and goal-atom termination. This evaluator does + # not replay physics, so the note is empty. + assert verdict == { + "reward": 0.0, # bonus gated by the internal rejection + "solved": False, + "note": "", + } + # Labels are (name, objects, params) triples since plan-capture + # gating started matching on exact params. + assert evaluator.seen_options == [("Push", (), ())] + # Pre-built labels pass through unchanged. + fn(states, [("Push", ("robot", ))], task_idx=0) + assert evaluator.seen_options == [("Push", ("robot", ))] + with pytest.raises(ValueError, match="no task evaluator"): + fn(states, None, task_idx=1) + with pytest.raises(ValueError, match="out of range"): + fn(states, None, task_idx=2) + with pytest.raises(ValueError, match="non-empty"): + fn([], None, task_idx=0) + + +def test_evaluate_trajectory_physics_sweep(approach_cls): + """physics_sweep=True scores the sequence at every physics-margin point on + a fresh env at that physics and reports the fraction scored solved; with no + points to sweep it says so.""" + import contextlib + import functools + from types import SimpleNamespace + + from predicators.structs import TaskEvaluator + + physics = {"friction": 0.5} + + class _FrictionEvaluator(TaskEvaluator): + """Certifies only when the (swept) friction is at least 0.5.""" + + def __init__(self): + super().__init__(set()) + + def _certify(self, states, step_options, sim_env=None): + return physics["friction"] >= 0.5, "friction" + + cup_type = Type("cup_type", ["f"]) + cup = cup_type("cup") + states = [State({cup: [0.0]}), State({cup: [1.0]})] + seen = [] + + @contextlib.contextmanager + def _scope(physical_overrides=None): + seen.append(dict(physical_overrides or {})) + prev = dict(physics) + physics.update(physical_overrides or {}) + try: + yield + finally: + physics.clear() + physics.update(prev) + + stub = SimpleNamespace( + _train_tasks=[Task(states[0], set(), evaluator=_FrictionEvaluator())], + _option_model=None, + _identified_physical_sigma_points=[{ + "friction": 0.4 + }, { + "friction": 0.5 + }, { + "friction": 0.6 + }], + _fresh_validation_env_scope=_scope) + stub._sweep_evaluation = functools.partial(approach_cls._sweep_evaluation, + stub) + fn = approach_cls._make_evaluate_trajectory_fn(stub) + plain = fn(states, None, task_idx=0) + assert "sweep" not in plain and plain["solved"] is True + swept = fn(states, None, task_idx=0, physics_sweep=True) + assert swept["solved"] is True + sweep = swept["sweep"] + assert [p["solved"] for p in sweep["points"]] == [False, True, True] + assert sweep["solved_fraction"] == pytest.approx(2 / 3) + assert sweep["certified"] is False + assert seen == [{"friction": 0.4}, {"friction": 0.5}, {"friction": 0.6}] + assert physics == {"friction": 0.5} # the scope restored the physics + stub._identified_physical_sigma_points = [] + assert fn(states, None, task_idx=0, physics_sweep=True)["sweep"] is None diff --git a/tests/approaches/test_gnn_dynamics_approach.py b/tests/approaches/test_gnn_dynamics_approach.py deleted file mode 100644 index 70d7a00580..0000000000 --- a/tests/approaches/test_gnn_dynamics_approach.py +++ /dev/null @@ -1,129 +0,0 @@ -"""Tests for the GNN dynamics + shooting baseline (paper arm C5).""" -# pylint: disable=protected-access -import numpy as np -import pytest - -from predicators import utils -from predicators.approaches import ApproachFailure, ApproachTimeout, \ - create_approach -from predicators.datasets import create_dataset -from predicators.envs import create_new_env -from predicators.ground_truth_models import get_gt_options -from predicators.settings import CFG -from predicators.structs import InteractionResult - - -def _setup(env_name: str = "cover"): - utils.reset_config({ - "env": env_name, - "num_train_tasks": 3, - "num_test_tasks": 2, - "gnn_num_epochs": 20, - "gnn_use_validation_set": False, - "gnn_do_normalization": True, - "gnn_dynamics_history_len": 2, - "gnn_dynamics_shooting_max_tries": 5, - "gnn_dynamics_max_plan_length": 4, - "explorer": "random_options", - "online_nsrt_learning_requests_per_cycle": 1, - "max_num_steps_interaction_request": 5, - "horizon": 10, - "timeout": 5, - }) - env = create_new_env(env_name) - train_tasks = [t.task for t in env.get_train_tasks()] - options = get_gt_options(env.get_name()) - approach = create_approach("gnn_dynamics_shooting", env.predicates, - options, env.types, env.action_space, - train_tasks) - predicates, _ = utils.parse_config_excluded_predicates(env) - dataset = create_dataset(env, train_tasks, options, predicates) - return env, train_tasks, options, approach, dataset - - -def test_gnn_dynamics_learns_predicts_and_plans(): - """Learn from demos, predict a transition, shoot a plan, save and load.""" - env, train_tasks, options, approach, dataset = _setup() - assert approach.is_learning_based - task = env.get_test_tasks()[0].task - with pytest.raises(ApproachFailure): # nothing learned yet - approach.solve(task, timeout=CFG.timeout) - approach.learn_from_offline_dataset(dataset) - assert approach._gnn is not None - # Option-level examples came out of the demos with history attached. - examples = approach._generate_examples() - assert examples - assert all(len(h) <= 2 for h, *_ in examples) - history, state, option, next_state, num_actions = examples[-1] - pred_state, pred_steps = approach.predict_next_state( - history, state, option) - assert set(pred_state) == set(state) - assert pred_steps >= 1 - assert isinstance(num_actions, int) - for obj in state: - assert pred_state[obj].shape == next_state[obj].shape - # A shooting policy is returned; executing it either reaches the goal - # or fails honestly (the tiny model is not expected to be accurate). - try: - policy = approach.solve(task, timeout=CFG.timeout) - utils.run_policy_with_simulator(policy, - env.simulate, - task.init, - task.goal_holds, - max_num_steps=CFG.horizon, - exceptions_to_break_on={ - utils.OptionExecutionFailure, - ApproachFailure, - }) - except (ApproachFailure, ApproachTimeout): - pass - # Save / load round trip rebuilds the model. - approach2 = create_approach("gnn_dynamics_shooting", env.predicates, - options, env.types, env.action_space, - train_tasks) - approach2.load(online_learning_cycle=None) - assert approach2._gnn is not None - assert approach2._feat_to_index == approach._feat_to_index - s2, _ = approach2.predict_next_state(history, state, option) - for obj in state: - assert np.allclose(s2[obj], pred_state[obj]) - - -def test_gnn_dynamics_online_cycle(): - """Interaction requests come from the configured explorer and the results - extend the data the next model is trained on.""" - env, _, _, approach, dataset = _setup() - approach.learn_from_offline_dataset(dataset) - n_before = len(approach._trajectories) - requests = approach.get_interaction_requests() - assert len(requests) == 1 - request = requests[0] - task = approach._train_tasks[request.train_task_idx] - traj, _ = utils.run_policy(request.act_policy, - env, - "train", - request.train_task_idx, - request.termination_function, - max_num_steps=5, - exceptions_to_break_on={ - utils.RequestActPolicyFailure, - }) - del task - result = InteractionResult(traj.states, traj.actions, - [None] * len(traj.states)) - approach.learn_from_interaction_results([result]) - assert len(approach._trajectories) == n_before + 1 - assert approach._online_learning_cycle == 1 - # The interaction trajectory is not a demo and keeps its task index. - assert not approach._trajectories[-1].is_demo - assert approach._trajectories[-1].train_task_idx == request.train_task_idx - - -def test_gnn_dynamics_no_transitions_keeps_model(): - """With no option-bearing data the previous model is kept.""" - _, _, _, approach, dataset = _setup() - approach.learn_from_offline_dataset(dataset) - gnn = approach._gnn - approach._trajectories = [] - approach._learn_model() - assert approach._gnn is gnn diff --git a/tests/approaches/test_oracle_process_planning_boil.py b/tests/approaches/test_oracle_process_planning_boil.py index 03d808874b..f407e64db2 100644 --- a/tests/approaches/test_oracle_process_planning_boil.py +++ b/tests/approaches/test_oracle_process_planning_boil.py @@ -1,7 +1,7 @@ """End-to-end test: oracle_process_planning solves a boil task. -Mirrors the config from ``predicatorv3/oracle.yaml`` + -``predicatorv3/envs/all.yaml`` + ``predicatorv3/common.yaml`` so that a +Mirrors the retired phased config (``predicatorv3/oracle.yaml`` + +``envs/all.yaml`` + ``common.yaml`` at tag iclr-empiric-submission) so that a regression in either the approach (process planning + bilevel refinement) or the boil env's skill execution would surface here. @@ -30,7 +30,7 @@ def _oracle_boil_config() -> dict: - """Flags from predicatorv3/{common,envs/all,oracle}.yaml flattened. + """Flags from the retired predicatorv3/{common,envs/all,oracle}.yaml. Kept minimal: 1 train task and 1 test task, no online learning cycles (oracle approach is not learning-based), no LLM (oracle diff --git a/tests/approaches/test_oracle_process_planning_bridge.py b/tests/approaches/test_oracle_process_planning_bridge.py index bd8480aace..1737a21914 100644 --- a/tests/approaches/test_oracle_process_planning_bridge.py +++ b/tests/approaches/test_oracle_process_planning_bridge.py @@ -1,10 +1,10 @@ """End-to-end test: oracle_process_planning solves a bridge task. -Mirrors the config from ``predicatorv3/oracle.yaml`` + -``predicatorv3/envs/all.yaml`` (bridge entry) + ``predicatorv3/common.yaml`` -so that a regression in the approach (process planning + bilevel -refinement), the bridge env's glue/weld machinery, or the skill -factories would surface here. +Mirrors the retired phased config (``predicatorv3/oracle.yaml`` + +``envs/all.yaml`` (bridge entry) + ``common.yaml`` at tag +iclr-empiric-submission) so that a regression in the approach (process +planning + bilevel refinement), the bridge env's glue/weld machinery, or +the skill factories would surface here. Runs the smallest viable config (1 train task, 1 test task; 5 blocks, 2 glue joints) and asserts: @@ -34,7 +34,7 @@ def _oracle_bridge_config(n_spans: int = 3, pool: int = 0) -> dict: - """Flags from predicatorv3/{common,envs/all,oracle}.yaml flattened.""" + """Flags from the retired predicatorv3/{common,envs/all,oracle}.yaml.""" return { # --- env: bridge from envs/all.yaml --- "env": diff --git a/tests/approaches/test_published_fit_reuse.py b/tests/approaches/test_published_fit_reuse.py index 4db849f2c8..6ccba125e2 100644 --- a/tests/approaches/test_published_fit_reuse.py +++ b/tests/approaches/test_published_fit_reuse.py @@ -72,14 +72,12 @@ def test_published_fit_is_reused_for_the_fitted_file(tmp_path: Any) -> None: is None -def test_reused_physics_fit_restores_the_margin_gate_state( +def test_reused_physics_fit_restores_the_sweep_points( tmp_path: Any, monkeypatch: Any) -> None: """Deploying a published physics fit re-applies its physical values and - restores the cycle-applied snapshot and the physics-margin sigma points - (applying resets them), so the capture gate's margin sweep survives the - skip of the harness refit.""" + restores the physics-margin sigma points (applying resets them), so the + physics sweep survives the skip of the harness refit.""" utils.reset_config({ - "agent_sim_learn_oracle_sim_params": False, "agent_explorer_info_seeking": False, }) sim_file = tmp_path / "simulator.py" @@ -94,7 +92,6 @@ def test_reused_physics_fit_restores_the_margin_gate_state( # and make the identity assert below unreachable for mypy. setattr(approach, "_last_fit_result", None) approach._fit_sse = float("inf") - approach._cycle_applied_physical = {} approach._identified_physical_sigma_points = [] approach._rng = np.random.default_rng(0) applied_calls = [] @@ -127,7 +124,6 @@ def _paths() -> Any: assert approach._last_fit_result is fit assert approach._fit_sse == 0.5 assert applied_calls == [{"mu": 0.7}] - assert approach._cycle_applied_physical == {"mu": 0.7} assert approach._identified_physical_sigma_points == sigma @@ -147,7 +143,7 @@ def test_publish_without_a_fit_result_never_deploys(tmp_path: Any) -> None: def test_unfitted_deployment_carries_values_and_clears_evidence( tmp_path: Any, monkeypatch: Any, has_data: bool) -> None: """Edits retain compatible values, initialize new specs, and retire the old - model's posterior without calling any fitting backend.""" + model's posterior.""" utils.reset_config({"agent_sim_learn_param_uncertainty": False}) sim_file = tmp_path / "simulator.py" sim_file.write_text("before", encoding="utf-8") @@ -161,8 +157,6 @@ def test_unfitted_deployment_carries_values_and_clears_evidence( setattr(approach, "_last_fit_result", old_fit) approach._identified_physical_params = {"mu": .7, "removed": 9.} approach._identified_physical_sigma_points = [{"mu": .6}] - approach._cycle_applied_physical = dict( - approach._identified_physical_params) approach._physical_param_specs = [ParamSpec("mu", .5, 0., 1.)] specs = [ ParamSpec("k", 1., 0., 2.), @@ -179,20 +173,11 @@ def apply(params): approach._identified_physical_params = dict(params) approach._identified_physical_sigma_points = [] - def forbidden(*_args, **_kwargs): - pytest.fail("Deployment must not fit") - monkeypatch.setattr(approach, "_apply_identified_physical_params", apply) - monkeypatch.setattr(approach, "_fit_parameters_joint_rollout", forbidden) - monkeypatch.setattr(approach, "_fit_parameters_recurrent", forbidden) - monkeypatch.setattr( - "predicators.approaches.agent_sim_learning_approach" - ".fit_rule_parameters", forbidden) triples: Any = [(None, None, None)] if has_data else [] approach._fit_params_after_synthesis([], specs, triples, {}) assert approach._fitted_params == {"k": 1.5, "bounded": 2., "new": .25} assert applied == [{"mu": .7}] - assert approach._cycle_applied_physical == {"mu": .7} assert not approach._identified_physical_sigma_points assert approach._last_fit_result is None assert approach._fit_sse == float("inf") diff --git a/tests/approaches/test_sampler_learning_mixin.py b/tests/approaches/test_sampler_learning_mixin.py deleted file mode 100644 index 87c1d8829e..0000000000 --- a/tests/approaches/test_sampler_learning_mixin.py +++ /dev/null @@ -1,174 +0,0 @@ -"""Tests for SamplerLearningMixin's loader and oracle-install logic. - -Covers ``_load_samplers_from_module_file`` (missing file, exec error, -non-dict, bad entries, happy path) and -``_maybe_install_oracle_samplers`` (GT install, fallback to synthesis, -disabled no-op) on a minimal host. -""" - -from typing import Any, Dict, Set - -import numpy as np -from gym.spaces import Box - -from predicators import utils -from predicators.approaches import sampler_learning_mixin -from predicators.approaches.sampler_learning_mixin import SamplerLearningMixin -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, Type - -_block_type = Type("block", ["x"]) -_block = Object("block0", _block_type) - -_Reached = Predicate("Reached", [_block_type], lambda s, o: True) - -_Move = ParameterizedOption( - "Move", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=lambda _s, _m, _o, _p: Action(np.zeros(1, dtype=np.float32)), - initiable=lambda _s, _m, _o, _p: True, - terminal=lambda _s, _m, _o, _p: False, -) - - -class _Host(SamplerLearningMixin): # pylint: disable=abstract-method - """Minimal host supplying the mixin's contract surface.""" - - def __init__(self): - init = State({_block: np.array([0.0], dtype=np.float32)}) - self._types = {_block_type} - self._train_tasks = [Task(init, {GroundAtom(_Reached, [_block])})] - self._fitted_params: Dict[str, float] = {} - self._synthesized_samplers: Dict[str, Any] = {} - self._init_sampler_learning_state() - - def _get_all_predicates(self) -> Set[Predicate]: - return {_Reached} - - def _get_all_options(self) -> Set[ParameterizedOption]: - return {_Move} - - def _learning_cycle_index(self) -> int: - return 1 - - -def _host(**config): - utils.reset_config({"seed": 0, **config}) - return _Host() - - -# --------------------------------------------------------------------------- # -# _load_samplers_from_module_file. -# --------------------------------------------------------------------------- # - - -def test_load_samplers_missing_file_returns_empty(tmp_path): - """A missing samplers.py loads as the empty dict (samplers optional).""" - host = _host() - assert not host._load_samplers_from_module_file( # pylint: disable=protected-access - str(tmp_path / "samplers.py")) - - -def test_load_samplers_exec_error_returns_empty(tmp_path): - """A file that raises at exec time loads as the empty dict.""" - path = tmp_path / "samplers.py" - path.write_text("raise RuntimeError('boom')\n", encoding="utf-8") - host = _host() - assert not host._load_samplers_from_module_file( # pylint: disable=protected-access - str(path)) - - -def test_load_samplers_non_dict_returns_empty(tmp_path): - """LEARNED_SAMPLERS bound to a non-dict loads as the empty dict.""" - path = tmp_path / "samplers.py" - path.write_text("LEARNED_SAMPLERS = [1, 2]\n", encoding="utf-8") - host = _host() - assert not host._load_samplers_from_module_file( # pylint: disable=protected-access - str(path)) - - -def test_load_samplers_skips_unknown_and_non_callable_entries(tmp_path): - """Unknown option names and non-callables are dropped, the rest kept.""" - path = tmp_path / "samplers.py" - path.write_text("""\ -def _fn(state, subgoal_atoms, rng, objects): - del state, subgoal_atoms, objects - return np.array([0.5], dtype=np.float32) - -LEARNED_SAMPLERS = {"Move": _fn, "Teleport": _fn, "Reached": 7} -""", - encoding="utf-8") - host = _host() - loaded = host._load_samplers_from_module_file(str(path)) # pylint: disable=protected-access - assert set(loaded) == {"Move"} - - -def test_load_samplers_happy_path(tmp_path): - """A valid file loads a callable that draws correctly shaped params.""" - path = tmp_path / "samplers.py" - path.write_text("""\ -def _fn(state, subgoal_atoms, rng, objects): - del state, subgoal_atoms, objects - return np.array([0.25 + 0.01 * rng.random()], dtype=np.float32) - -LEARNED_SAMPLERS = {"Move": _fn} -""", - encoding="utf-8") - host = _host() - loaded = host._load_samplers_from_module_file(str(path)) # pylint: disable=protected-access - assert set(loaded) == {"Move"} - draw = loaded["Move"]( - host._train_tasks[0].init, # pylint: disable=protected-access - set(), - np.random.default_rng(0), - [_block]) - assert np.asarray(draw).shape == (1, ) - - -# --------------------------------------------------------------------------- # -# _maybe_install_oracle_samplers. -# --------------------------------------------------------------------------- # - - -def _gt_sampler(state, subgoal_atoms, rng, objects): - del state, subgoal_atoms, rng, objects - return np.array([0.5], dtype=np.float32) - - -def test_oracle_samplers_installed_when_available(monkeypatch): - """With oracle_samplers on and GT available: install, skip synthesis.""" - monkeypatch.setattr(sampler_learning_mixin, "get_gt_samplers", - lambda _env: {"Move": _gt_sampler}) - host = _host(agent_sim_learn_parameterized_samplers=True, - agent_sim_learn_oracle_samplers=True) - host._maybe_install_oracle_samplers() # pylint: disable=protected-access - assert host._synthesized_samplers == {"Move": _gt_sampler} # pylint: disable=protected-access - assert host._current_samplers_version == "oracle" # pylint: disable=protected-access - assert not host._do_synthesize_samplers # pylint: disable=protected-access - - -def test_oracle_samplers_fall_back_to_synthesis_when_none(monkeypatch): - """With oracle_samplers on but no GT for the env: synthesize instead.""" - monkeypatch.setattr(sampler_learning_mixin, "get_gt_samplers", - lambda _env: {}) - host = _host(agent_sim_learn_parameterized_samplers=True, - agent_sim_learn_oracle_samplers=True) - host._maybe_install_oracle_samplers() # pylint: disable=protected-access - assert not host._synthesized_samplers # pylint: disable=protected-access - assert host._do_synthesize_samplers # pylint: disable=protected-access - - -def test_samplers_disabled_no_synthesis_no_install(monkeypatch): - """With the master gate off nothing is installed or synthesized.""" - - def _boom(_env): - raise AssertionError("get_gt_samplers called with samplers disabled") - - monkeypatch.setattr(sampler_learning_mixin, "get_gt_samplers", _boom) - host = _host(agent_sim_learn_parameterized_samplers=False, - agent_sim_learn_oracle_samplers=False) - host._maybe_install_oracle_samplers() # pylint: disable=protected-access - assert not host._synthesized_samplers # pylint: disable=protected-access - assert not host._do_synthesize_samplers # pylint: disable=protected-access diff --git a/tests/approaches/test_sim_learning_info_seeking.py b/tests/approaches/test_sim_learning_info_seeking.py index 27ea816d0c..cc4dc2ee3c 100644 --- a/tests/approaches/test_sim_learning_info_seeking.py +++ b/tests/approaches/test_sim_learning_info_seeking.py @@ -156,37 +156,6 @@ def test_rebuild_param_ensemble_respects_flag(): assert approach._param_ensemble[0] == {"a": 1.0} # member 0 is anchor -def test_rebuild_param_ensemble_empty_under_oracle_params(): - """Oracle params carry no uncertainty, so no ensemble is built. - - Without this the uniform-jitter fallback would hand the capture gate - members that no plan can satisfy (a zero rate, a rewired lamp), and - the gate would refuse every plan on a model that is exactly right. - """ - from predicators.code_sim_learning.fit_space import ParamSpec - approach = object.__new__(AgentSimLearningApproach) - approach._fitted_params = {"a": 1.0} - approach._param_specs = [ParamSpec("a", 1.0, lo=0.0, hi=2.0)] - approach._param_ensemble = [{"a": 1.0}, {"a": 2.0}] - approach._last_fit_result = None - approach._rng = np.random.default_rng(0) - utils.reset_config({ - "agent_plan_validation_rule_param_margin": True, - "agent_explorer_info_ensemble_size": 5, - "agent_sim_learn_oracle_sim_params": True, - }) - approach._rebuild_param_ensemble() - assert approach._param_ensemble == [] - - utils.reset_config({ - "agent_plan_validation_rule_param_margin": True, - "agent_explorer_info_ensemble_size": 5, - "agent_sim_learn_oracle_sim_params": False, - }) - approach._rebuild_param_ensemble() - assert len(approach._param_ensemble) == 5 - - def _selector_approach(fit_result): from predicators.code_sim_learning.fit_space import ParamSpec approach = object.__new__(AgentSimLearningApproach) @@ -261,14 +230,9 @@ def test_select_ensemble_uniform_when_calibration_disabled(): assert method == "uniform-perturb" -def test_fit_params_no_data_seeds_declared_inits(monkeypatch): - """With no transitions, params seed from inits and no fit runs. - - This is the oracle-sim-program no-demos path: every demo failed, so - ``_learn_simulator`` reaches the fit with empty - ``base_pred_triples`` and must fall back to the declared init values - instead of fitting. - """ +def test_fit_params_no_data_seeds_declared_inits(): + """With no transitions and no published fit, the deployed params are the + declared init values.""" from predicators.code_sim_learning.fit_space import ParamSpec approach = object.__new__(AgentSimLearningApproach) @@ -280,16 +244,7 @@ def test_fit_params_no_data_seeds_declared_inits(monkeypatch): approach._last_fit_result = None approach._fit_sse = 0.0 approach._rng = np.random.default_rng(0) - - def _fail_fit(*args, **kwargs): - del args, kwargs - raise AssertionError("fit must not run with no data") - - monkeypatch.setattr( - "predicators.approaches.agent_sim_learning_approach" - ".fit_rule_parameters", _fail_fit) utils.reset_config({ - "agent_sim_learn_oracle_sim_params": False, "agent_explorer_info_seeking": False, }) specs = [ParamSpec("a", 1.5, lo=0.0, hi=5.0)] diff --git a/tests/code_sim_learning/test_bridge_transfer_oracle.py b/tests/code_sim_learning/test_bridge_transfer_oracle.py index 60d94b9230..7d6fa88f36 100644 --- a/tests/code_sim_learning/test_bridge_transfer_oracle.py +++ b/tests/code_sim_learning/test_bridge_transfer_oracle.py @@ -18,57 +18,17 @@ from tests.code_sim_learning.test_continual_oracle import _load -def test_oracle_repair_pilot_is_six_new_preflight_off_runs() -> None: - """The pilot cannot launch other domains or resume historical seeds.""" - runs = list( - generate_run_configs( - "predicatorv3/continual_oracle_validation_r2.yaml", False)) - assert len(runs) == 6 - assert {r.env for r in runs} == {"pybullet_bridge", "pybullet_domino"} - seeds = set() - for run in runs: - assert isinstance(run, SingleSeedRunConfig) - seeds.add(run.seed) - assert run.approach == "agent_continual_oracle_dynamics" - assert run.flags["continual_skill_preflight"] is False - assert "benchmark_r2" in run.experiment_id - assert seeds == {0, 1, 2} - - -def test_empiric_r2_is_ten_new_shadow_runs() -> None: - """The prospective round has two fresh seeds, no mandatory gate.""" - runs = list( - generate_run_configs( - "predicatorv3/continual_empiric_benchmark_r2.yaml", False)) - assert len(runs) == 10 - assert len({r.env for r in runs}) == 5 - for run in runs: - assert isinstance(run, SingleSeedRunConfig) - assert run.seed in (3, 4) - assert run.approach == "agent_continual" - assert run.flags["continual_skill_preflight"] is False - assert run.flags["continual_validation_audit"] is True - assert run.flags["continual_validation_audit_seconds"] == 600. - assert run.experiment_id.endswith("-mb_opus_benchmark_r2") - - def test_transfer_comparisons_preserve_arm_contracts(monkeypatch: Any) -> None: - """The current eight-arm benchmark uses the four-span task consistently. - - The old pilot launchers were removed when these settings became the - benchmark defaults; test the maintained menu-based launcher instead. - """ - transfer = [ - c for c in generate_run_configs( - "predicatorv3/continual_eight_agent_noisy_sweep.yaml", False) - if c.env == "pybullet_bridge" - ] - assert len(transfer) == 24 - assert len({cfg.approach for cfg in transfer}) == 8 + """Every arm of the benchmark plays the same four-span Bridge task, and the + launch command parses back to it.""" + transfer = list( + generate_run_configs("empiric/benchmark.yaml", False, envs=["bridge"])) + assert len(transfer) == 7 * 5 + assert len({cfg.approach for cfg in transfer}) == 6 for cfg in transfer: assert isinstance(cfg, SingleSeedRunConfig) assert cfg.env == "pybullet_bridge" - assert cfg.seed in (0, 1, 2) + assert cfg.seed in range(5) assert cfg.flags["bridge_train_span_blocks"] == 3 assert cfg.flags["bridge_test_span_blocks"] == 4 assert cfg.flags["continual_steps_per_level"] == 10000 diff --git a/tests/code_sim_learning/test_continual_oracle.py b/tests/code_sim_learning/test_continual_oracle.py index a8f58f4164..4137d9db37 100644 --- a/tests/code_sim_learning/test_continual_oracle.py +++ b/tests/code_sim_learning/test_continual_oracle.py @@ -353,10 +353,12 @@ def test_scene_only_physical_calibration(domain: str) -> None: AgentContinualSceneOnlyApproach # pylint: disable=import-outside-toplevel from scripts.cluster_utils import \ generate_run_configs # pylint: disable=import-outside-toplevel + + # The benchmark's env and common flags; the scene-only class needs no + # arm flags of its own. config = next(c for c in generate_run_configs( - "predicatorv3/continual_eight_agent_noisy_sweep.yaml", False) - if c.env == f"pybullet_{domain}" - and c.approach == "agent_continual_scene_only") + "empiric/benchmark.yaml", False, approaches=["mb_opus"]) + if c.env == f"pybullet_{domain}") utils.reset_config({ **{k: v for k, v in config.flags.items() if k != "log"}, "env": config.env, diff --git a/tests/code_sim_learning/test_param_fitting.py b/tests/code_sim_learning/test_param_fitting.py index 236b94e4ea..baf24ff616 100644 --- a/tests/code_sim_learning/test_param_fitting.py +++ b/tests/code_sim_learning/test_param_fitting.py @@ -13,7 +13,7 @@ import predicators.approaches # noqa: F401 # pylint: disable=unused-import from predicators import utils -from predicators.approaches.agent_model_based_approach import _SketchStep +from predicators.agent_sdk.sketch_types import SketchStep as _SketchStep from predicators.code_sim_learning.fit_space import ParamSpec from predicators.code_sim_learning.fitting import compute_sse, fit_params from predicators.envs import create_new_env diff --git a/tests/code_sim_learning/test_program_world_model.py b/tests/code_sim_learning/test_program_world_model.py index 0016cffc31..49557cab7c 100644 --- a/tests/code_sim_learning/test_program_world_model.py +++ b/tests/code_sim_learning/test_program_world_model.py @@ -91,8 +91,8 @@ def _ground_pick_place(options: Any, param: float) -> Any: def test_program_option_model_steps_and_carries_latent() -> None: - """Transitions write features, seed and advance the latent, honor the - override, and never mutate the input state.""" + """Transitions write features, seed and advance the latent, and never + mutate the input state.""" env, train_tasks, options, _ = _cover() model = ProgramOptionModel(_load(_HAND_PROGRAM, env, options), seed=0) init = train_tasks[0].init @@ -110,12 +110,6 @@ def test_program_option_model_steps_and_carries_latent() -> None: nxt2, _ = model.get_next_state_and_num_actions(nxt, option) assert nxt2.latent is not None assert nxt2.latent["count"] == nxt.latent["count"] + 1 - # The override pins every latent-less start. - model.initial_latent_override = {"count": 40} - nxt3, _ = model.get_next_state_and_num_actions(init, option) - assert nxt3.latent is not None - assert nxt3.latent["count"] == 41 - model.initial_latent_override = None assert model.last_execution_failure is None diff --git a/tests/envs/test_pybullet_fan_transfer.py b/tests/envs/test_pybullet_fan_transfer.py index 17826b2723..235cb58600 100644 --- a/tests/envs/test_pybullet_fan_transfer.py +++ b/tests/envs/test_pybullet_fan_transfer.py @@ -17,13 +17,14 @@ @pytest.mark.parametrize("seed", [0, 1]) @pytest.mark.parametrize("arm", [0, 1]) -@pytest.mark.parametrize("variant", ["transfer", "inertial", "ramp"]) -def test_launch_config_constructs_both_levels(seed, arm, variant): - """Resolve the actual launcher, including list overrides, before reset.""" +def test_launch_config_constructs_both_levels(seed, arm): + """Resolve the benchmark's Fan setting, including list flags, before + reset.""" configs = list( - generate_run_configs( - f"predicatorv3/continual_fan_{variant}_pilot_r1.yaml", - batch_seeds=True)) + generate_run_configs("empiric/benchmark.yaml", + batch_seeds=True, + envs=["fan"], + approaches=["mb_opus", "mf_opus"])) assert len(configs) == 2 config = configs[arm] flags = { diff --git a/tests/execution_monitoring/test_execution_monitoring.py b/tests/execution_monitoring/test_execution_monitoring.py index 422f91b5f1..6c7c5acfe2 100644 --- a/tests/execution_monitoring/test_execution_monitoring.py +++ b/tests/execution_monitoring/test_execution_monitoring.py @@ -1,20 +1,14 @@ """Tests for execution monitors.""" -import numpy as np import pytest -from gym.spaces import Box from predicators.execution_monitoring import create_execution_monitor from predicators.execution_monitoring.expected_atoms_monitor import \ ExpectedAtomsExecutionMonitor from predicators.execution_monitoring.mpc_execution_monitor import \ MpcExecutionMonitor -from predicators.execution_monitoring.subgoal_annotations_monitor import \ - SubgoalAnnotationsExecutionMonitor, SubgoalExecutionStatus from predicators.execution_monitoring.trivial_execution_monitor import \ TrivialExecutionMonitor -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Type def test_create_execution_monitor(): @@ -28,131 +22,6 @@ def test_create_execution_monitor(): exec_monitor = create_execution_monitor("expected_atoms") assert isinstance(exec_monitor, ExpectedAtomsExecutionMonitor) - exec_monitor = create_execution_monitor("subgoal_annotations") - assert isinstance(exec_monitor, SubgoalAnnotationsExecutionMonitor) - with pytest.raises(NotImplementedError) as e: create_execution_monitor("not a real monitor") assert "Unrecognized execution monitor" in str(e) - - -class _FakeSketchStep: - """Duck-typed sketch step (see agent_sdk.bilevel_sketch.SketchStep).""" - - def __init__(self, option, subgoal_atoms, subgoal_neg_atoms=None): - self.option = option - self.subgoal_atoms = subgoal_atoms - self.subgoal_neg_atoms = subgoal_neg_atoms - - -def test_subgoal_annotations_monitor(): - """Unit tests for SubgoalAnnotationsExecutionMonitor.step().""" - block_type = Type("block", ["held"]) - block = Object("block0", block_type) - held = Predicate("Held", [block_type], - lambda s, o: s.get(o[0], "held") > 0.5) - state_held = State({block: np.array([1.0], dtype=np.float32)}) - state_free = State({block: np.array([0.0], dtype=np.float32)}) - - def _make_option(terminal): - param_opt = ParameterizedOption( - "Pick", - types=[block_type], - params_space=Box(low=np.zeros(1, dtype=np.float32), - high=np.ones(1, dtype=np.float32)), - policy=lambda s, m, o, p: Action(np.zeros(1, dtype=np.float32)), - initiable=lambda s, m, o, p: True, - terminal=lambda s, m, o, p: terminal, - ) - return param_opt, param_opt.ground([block], - np.zeros(1, dtype=np.float32)) - - done_parent, done_option = _make_option(True) - _, running_option = _make_option(False) - held_atom = GroundAtom(held, [block]) - - monitor = create_execution_monitor("subgoal_annotations") - - # No approach info (e.g. exploration): never replan. - assert not monitor.step(state_free) - - # Info of an unexpected shape (another approach's export): ignore. - monitor.update_approach_info([{"something": "else"}]) - assert not monitor.step(state_free) - - def _status(option, steps_initiated, pos=None, neg=None): - step = _FakeSketchStep(done_parent, pos, neg) - return SubgoalExecutionStatus(sketch=[step], - steps_initiated=steps_initiated, - current_option=option) - - # No option initiated yet (fresh policy right after a replan). - monitor.update_approach_info([_status(None, 0, {held_atom})]) - assert not monitor.step(state_free) - - # Mid-option: the current option has not terminated. - monitor.update_approach_info([_status(running_option, 1, {held_atom})]) - assert not monitor.step(state_free) - - # Boundary, annotation holds: no replan. - monitor.update_approach_info([_status(done_option, 1, {held_atom})]) - assert not monitor.step(state_held) - - # Boundary, unannotated step: nothing to check. - monitor.update_approach_info([_status(done_option, 1, None)]) - assert not monitor.step(state_free) - - # Boundary, positive atom unsatisfied: replan. - monitor.update_approach_info([_status(done_option, 1, {held_atom})]) - assert monitor.step(state_free) - - # Boundary, negative atom violated: replan. - monitor.update_approach_info([_status(done_option, 1, None, {held_atom})]) - assert monitor.step(state_held) - - -def test_subgoal_annotations_monitor_skips_unverifiable_atoms(): - """A classifier that cannot evaluate on a bare observation (e.g. it indexes - a latent that real env states never carry) is skipped with a warning - instead of crashing the episode or counting as divergence.""" - block_type = Type("block", ["held"]) - block = Object("block0", block_type) - - def _latent_only(s, o, latent=None): - del s, o - return bool(latent["_bonds"]) # TypeError when latent is None - - latent_pred = Predicate("LatentBonded", [block_type], _latent_only) - latent_atom = GroundAtom(latent_pred, [block]) - held = Predicate("Held", [block_type], - lambda s, o: s.get(o[0], "held") > 0.5) - held_atom = GroundAtom(held, [block]) - state = State({block: np.array([0.0], dtype=np.float32)}) # no latent - - param_opt = ParameterizedOption( - "Pick", - types=[block_type], - params_space=Box(low=np.zeros(1, dtype=np.float32), - high=np.ones(1, dtype=np.float32)), - policy=lambda s, m, o, p: Action(np.zeros(1, dtype=np.float32)), - initiable=lambda s, m, o, p: True, - terminal=lambda s, m, o, p: True, - ) - done_option = param_opt.ground([block], np.zeros(1, dtype=np.float32)) - monitor = create_execution_monitor("subgoal_annotations") - - def _status(pos=None, neg=None): - step = _FakeSketchStep(param_opt, pos, neg) - return SubgoalExecutionStatus(sketch=[step], - steps_initiated=1, - current_option=done_option) - - # The unverifiable atom alone: skipped, no divergence, no crash. - monitor.update_approach_info([_status(pos={latent_atom})]) - assert not monitor.step(state) - # Same in the negative polarity. - monitor.update_approach_info([_status(neg={latent_atom})]) - assert not monitor.step(state) - # A genuinely failed observable atom alongside it still fires. - monitor.update_approach_info([_status(pos={latent_atom, held_atom})]) - assert monitor.step(state) diff --git a/tests/explorers/test_agent_model_based_explorer.py b/tests/explorers/test_agent_model_based_explorer.py deleted file mode 100644 index 3af97c9330..0000000000 --- a/tests/explorers/test_agent_model_based_explorer.py +++ /dev/null @@ -1,491 +0,0 @@ -"""Tests for AgentModelBasedExplorer.""" -# pylint: disable=protected-access - -from unittest.mock import AsyncMock, MagicMock - -import numpy as np -import pytest -from gym.spaces import Box - -from predicators import utils -from predicators.agent_sdk.sketch_types import SketchStep -from predicators.agent_sdk.tools import ToolContext -from predicators.explorers import create_explorer -from predicators.explorers.agent_model_based_explorer import \ - AgentModelBasedExplorer -from predicators.explorers.base_explorer import BaseExplorer -from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, Type - -# --------------------------------------------------------------------------- -# Fixtures (parallel the bilevel approach tests) -# --------------------------------------------------------------------------- - -_block_type = Type("block", ["x", "y", "held"]) -_robot_type = Type("robot", ["x", "y"]) - -_block0 = Object("block0", _block_type) -_block1 = Object("block1", _block_type) -_robot = Object("robot0", _robot_type) - -_Holding = Predicate("Holding", [_block_type], - lambda s, o: s.get(o[0], "held") > 0.5) -_On = Predicate("On", [_block_type, _block_type], - lambda s, o: abs(s.get(o[0], "x") - s.get(o[1], "x")) < 0.1) -_HandEmpty = Predicate("HandEmpty", [_robot_type], lambda s, o: True) - -_ALL_PREDICATES = {_Holding, _On, _HandEmpty} -_ALL_TYPES = {_block_type, _robot_type} - - -def _noop_policy(_s, _m, _o, _p): - return Action(np.zeros(1, dtype=np.float32)) - - -def _always_true(_s, _m, _o, _p): - return True - - -def _always_false(_s, _m, _o, _p): - return False - - -_Pick = ParameterizedOption( - "Pick", - types=[_block_type], - params_space=Box(low=np.array([0.0], dtype=np.float32), - high=np.array([1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_false, -) - -_Place = ParameterizedOption( - "Place", - types=[_block_type, _block_type], - params_space=Box(low=np.array([0.0, 0.0], dtype=np.float32), - high=np.array([1.0, 1.0], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_false, -) - -_Wait = ParameterizedOption( - "Wait", - types=[_robot_type], - params_space=Box(low=np.array([], dtype=np.float32), - high=np.array([], dtype=np.float32)), - policy=_noop_policy, - initiable=_always_true, - terminal=_always_false, -) - -_ALL_OPTIONS = {_Pick, _Place, _Wait} - - -def _make_state(overrides=None): - data = { - _block0: np.array([0.1, 0.2, 0.0], dtype=np.float32), - _block1: np.array([0.5, 0.6, 0.0], dtype=np.float32), - _robot: np.array([0.0, 0.0], dtype=np.float32), - } - if overrides: - for obj, vals in overrides.items(): - data[obj] = np.array(vals, dtype=np.float32) - return State(data) - - -def _make_task(): - state = _make_state() - goal = {GroundAtom(_On, [_block0, _block1])} - return Task(state, goal) - - -def _assistant_response(text: str): - return [{ - "type": "assistant", - "content": [{ - "type": "text", - "text": text - }], - }] - - -def _make_explorer(option_model, query_impl): - """Build an AgentModelBasedExplorer with stubbed session + tool_context.""" - tool_context = ToolContext( - types=_ALL_TYPES, - predicates=_ALL_PREDICATES, - options=_ALL_OPTIONS, - train_tasks=[_make_task()], - option_model=option_model, - ) - agent_session = MagicMock() - agent_session.query = query_impl - agent_session.tool_names = None - explorer = AgentModelBasedExplorer( - predicates=_ALL_PREDICATES, - options=_ALL_OPTIONS, - types=_ALL_TYPES, - action_space=Box(low=-1, high=1, shape=(1, )), - train_tasks=[_make_task()], - max_steps_before_termination=50, - tool_context=tool_context, - agent_session=agent_session, - ) - return explorer, tool_context - - -def _reset_config(**overrides): - base = { - "env": "cover", - "approach": "agent_model_based", - "num_train_tasks": 1, - "num_test_tasks": 1, - "seed": 42, - "agent_bilevel_max_samples_per_step": 5, - "agent_bilevel_check_subgoals": True, - "agent_bilevel_log_state": False, - "agent_explorer_fallback_to_random": True, - "agent_sdk_max_trajectories_in_context": 5, - } - base.update(overrides) - utils.reset_config(base) - - -# --------------------------------------------------------------------------- -# Tests -# --------------------------------------------------------------------------- - - -def test_factory_registration(): - """AgentModelBasedExplorer is reachable through create_explorer, under its - own name and under the deprecated agent_bilevel alias.""" - _reset_config() - tool_context = ToolContext( - types=_ALL_TYPES, - predicates=_ALL_PREDICATES, - options=_ALL_OPTIONS, - train_tasks=[_make_task()], - option_model=MagicMock(), - ) - agent_session = MagicMock() - explorer = create_explorer( - "agent_model_based", - _ALL_PREDICATES, - _ALL_OPTIONS, - _ALL_TYPES, - Box(low=-1, high=1, shape=(1, )), - [_make_task()], - tool_context=tool_context, - agent_session=agent_session, - ) - assert isinstance(explorer, BaseExplorer) - assert isinstance(explorer, AgentModelBasedExplorer) - legacy = create_explorer( - "agent_bilevel", - _ALL_PREDICATES, - _ALL_OPTIONS, - _ALL_TYPES, - Box(low=-1, high=1, shape=(1, )), - [_make_task()], - tool_context=tool_context, - agent_session=agent_session, - ) - assert isinstance(legacy, AgentModelBasedExplorer) - - -def test_happy_path_returns_policy_and_stashes_subgoals(): - """Canned sketch → refined plan → policy and stashed subgoals.""" - _reset_config() - - goal_state = _make_state({_block0: [0.5, 0.6, 0.0]}) - option_model = MagicMock() - option_model.get_next_state_and_num_actions.return_value = (goal_state, 3) - - plan_text = ("Pick(block0:block)\n" - "Place(block0:block, block1:block) -> " - "{On(block0:block, block1:block)}\n") - query = AsyncMock(return_value=_assistant_response(plan_text)) - - explorer, tool_context = _make_explorer(option_model, query) - policy, term_fn = explorer._get_exploration_strategy(0, timeout=5) - - assert callable(policy) - assert term_fn(_make_state()) is False - assert tool_context.last_sketch_subgoals is not None - assert len(tool_context.last_sketch_subgoals) == 2 - # Second step's positive subgoal should be {On(block0, block1)}. - pos2, _neg2 = tool_context.last_sketch_subgoals[1] - assert pos2 == {GroundAtom(_On, [_block0, _block1])} - assert tool_context.last_sketch_options == [ - ("Pick", ["block0"]), - ("Place", ["block0", "block1"]), - ] - assert query.await_count == 1 - - -def test_wait_memory_injection_on_grounding(): - """A Wait step's annotated subgoal rides on the grounded option as - ``wait_target_atoms`` so WaitOption terminates on the intended atoms.""" - _reset_config() - explorer, _ = _make_explorer(MagicMock(), None) - step = SketchStep(option=_Wait, - objects=[_robot], - subgoal_atoms={GroundAtom(_On, [_block0, _block1])}) - plan = explorer._ground_sketch_verbatim([step]) - assert len(plan) == 1 and plan[0].name == "Wait" - assert plan[0].memory["wait_target_atoms"] == { - GroundAtom(_On, [_block0, _block1]) - } - - -def test_sketch_executes_verbatim_without_belief_refinement(): - """The agent's explicit parameters execute exactly as written: the belief - model is never rolled, the verdict is not-certified, and the cycle record - shows the executed values.""" - _reset_config(agent_bilevel_use_llm_initial_params=True) - option_model = MagicMock() - plan_text = ("```\nPick(block0:block)[0.42] -> {Holding(block0:block)}\n" - "Place(block0:block, block1:block)[0.11, 0.22] -> " - "{On(block0:block, block1:block)}\n```") - query = AsyncMock(return_value=_assistant_response(plan_text)) - explorer, tool_context = _make_explorer(option_model, query) - policy, term_fn = explorer._get_exploration_strategy(0, timeout=5) - assert callable(policy) and term_fn(_make_state()) is False - assert not option_model.get_next_state_and_num_actions.called - assert tool_context.last_mental_model_solved is False - record = tool_context.cycle_scheduled_plans[-1] - assert "Pick(block0)[0.4200]" in record - assert "Place(block0, block1)[0.1100, 0.2200]" in record - assert "-> {On(block0:block, block1:block)}" in record - assert "without belief-model certification" in record - assert tool_context.last_sketch_options == [ - ("Pick", ["block0"]), - ("Place", ["block0", "block1"]), - ] - - -def test_missing_params_get_one_draw_from_the_box(): - """A step the agent left without parameters is grounded on one uniform. - - draw from the option's box - no search, and no crash. - """ - _reset_config(agent_bilevel_use_llm_initial_params=True) - explorer, _ = _make_explorer(MagicMock(), None) - steps = [ - SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None), - SketchStep(option=_Place, - objects=[_block0, _block1], - subgoal_atoms=None, - initial_params=np.array([0.5], dtype=np.float32)), - ] - plan = explorer._ground_sketch_verbatim(steps) - assert plan[0].params.shape == (1, ) and 0.0 <= plan[0].params[0] <= 1.0 - # Wrong arity counts as missing. - assert plan[1].params.shape == (2, ) - - -def _make_captured(pick_params, place_params): - """Build the (solved_plan, solved_sketch) a tool capture would stash.""" - grounded_plan = [ - _Pick.ground([_block0], np.array(pick_params, dtype=np.float32)), - _Place.ground([_block0, _block1], - np.array(place_params, dtype=np.float32)), - ] - captured_sketch = [ - SketchStep(option=_Pick, objects=[_block0], subgoal_atoms=None), - SketchStep(option=_Place, - objects=[_block0, _block1], - subgoal_atoms={GroundAtom(_On, [_block0, _block1])}), - ] - return grounded_plan, captured_sketch - - -def test_recovers_captured_plan_when_final_text_unparseable(): - """Agent validates a plan via submit_plan but ends in prose: - - explorer recovers the captured plan instead of falling back to - random and executes it at the captured continuous params. - """ - _reset_config() - option_model = MagicMock() - pick_params, place_params = [0.42], [0.11, 0.22] - grounded_plan, captured_sketch = _make_captured(pick_params, place_params) - explorer, tool_context = _make_explorer(option_model, None) - - async def query_impl(_msg, **_kw): - # Simulate the agent capturing a validated plan via the tool during - # the query (set AFTER the explorer's entry-time capture clear), then - # ending with prose that does NOT parse into a sketch. - tool_context.solved_plan = grounded_plan - tool_context.solved_sketch = captured_sketch - return _assistant_response("Solved it. Plan: 1. pick 2. place. Done.") - - explorer._agent_session.query = query_impl - policy, term_fn = explorer._get_exploration_strategy(0, timeout=5) - # Recovered (not random fallback): subgoals/options come from the capture. - assert callable(policy) - assert term_fn(_make_state()) is False - assert tool_context.last_sketch_options == [ - ("Pick", ["block0"]), - ("Place", ["block0", "block1"]), - ] - # The capture was consumed (cleared) so it can't leak into a later solve. - assert tool_context.solved_plan is None - assert tool_context.solved_sketch is None - # The captured params execute verbatim; the belief is not re-rolled. - assert not option_model.get_next_state_and_num_actions.called - record = tool_context.cycle_scheduled_plans[-1] - assert "Pick(block0)[0.4200]" in record - assert "Place(block0, block1)[0.1100, 0.2200]" in record - - -def test_fallback_when_query_fails_and_flag_on(): - """Agent raises → random options fallback when flag enabled.""" - _reset_config(agent_explorer_fallback_to_random=True) - - option_model = MagicMock() - - async def failing_query(_msg): - raise RuntimeError("boom") - - explorer, _ = _make_explorer(option_model, failing_query) - policy, term_fn = explorer._get_exploration_strategy(0, timeout=5) - assert callable(policy) - assert term_fn(_make_state()) is False - - -def test_fallback_disabled_raises(): - """Agent raises → RequestActPolicyFailure when fallback flag off.""" - _reset_config(agent_explorer_fallback_to_random=False) - - option_model = MagicMock() - - async def failing_query(_msg): - raise RuntimeError("boom") - - explorer, _ = _make_explorer(option_model, failing_query) - with pytest.raises(utils.RequestActPolicyFailure): - explorer._get_exploration_strategy(0, timeout=5) - - -def test_experiment_guidance_gated_by_info_seeking(): - """Experiment guidance appears iff info-seeking is on.""" - _reset_config(agent_explorer_info_seeking=True) - explorer, _ = _make_explorer(MagicMock(), MagicMock()) - guidance = explorer._build_experiment_guidance() # pylint: disable=protected-access - assert "sim.suggest_probes" in guidance - # Off => section absent entirely. - _reset_config(agent_explorer_info_seeking=False) - assert explorer._build_experiment_guidance() == "" # pylint: disable=protected-access - - -def test_experiment_guidance_injects_open_questions_ledger(tmp_path): - """The learn phase's open_questions.md ledger reaches the explore query - verbatim, independent of the info-seeking flag, and an oversized ledger - keeps its head (the ranking's top).""" - _reset_config(agent_explorer_info_seeking=False) - explorer, tool_context = _make_explorer(MagicMock(), MagicMock()) - # No sandbox / no file => no section (and no crash). - assert explorer._build_experiment_guidance() == "" # pylint: disable=protected-access - tool_context.sandbox_dir = str(tmp_path) - assert explorer._build_experiment_guidance() == "" # pylint: disable=protected-access - ledger = ("1. Bond window: place pairs at spacings 0.100/0.104/" - "0.110/0.114 and record which bond.") - (tmp_path / "open_questions.md").write_text(ledger, encoding="utf-8") - guidance = explorer._build_experiment_guidance() # pylint: disable=protected-access - assert ledger in guidance - assert "OPEN QUESTIONS" in guidance - assert "The TOP entry is mandatory" in guidance - # Info-seeking on: both the ledger and the boundary-probing note. - _reset_config(agent_explorer_info_seeking=True) - guidance = explorer._build_experiment_guidance() # pylint: disable=protected-access - assert ledger in guidance - assert "sim.suggest_probes" in guidance - # Oversized ledger: head survives, truncation is announced. - head = "TOP-RANKED ENTRY" - (tmp_path / "open_questions.md").write_text(head + "x" * 10000, - encoding="utf-8") - guidance = explorer._build_experiment_guidance() # pylint: disable=protected-access - assert head in guidance - assert "ledger truncated" in guidance - - -def _make_certified_capture(pick_params, place_params): - """A capture as the belief's validation gate leaves it: goal reached.""" - grounded_plan, captured_sketch = _make_captured(pick_params, place_params) - return grounded_plan, captured_sketch - - -def test_certified_capture_executes_verbatim_and_next_request_queries(): - """A plan the session validated through the capture gate (reached_goal - True) is executed verbatim as a solve attempt with a True mental-model - verdict; the cycle's next request on the task queries the agent again - (asking for a different certified plan) rather than replaying it.""" - _reset_config(agent_explorer_info_seeking=True) - option_model = MagicMock() - option_model.get_next_state_and_num_actions.return_value = (_make_state( - {_block0: [0.5, 0.6, 0.0]}), 3) - pick_params, place_params = [0.42], [0.11, 0.22] - grounded_plan, captured_sketch = _make_certified_capture( - pick_params, place_params) - explorer, tool_context = _make_explorer(option_model, None) - tool_context.atom_disagreement_fn = lambda _s, _atoms: 1.0 - queries = [] - - async def query_impl(msg, **_kw): - queries.append(msg) - tool_context.solved_plan = grounded_plan - tool_context.solved_sketch = captured_sketch - tool_context.solved_plan_reached_goal = True - tool_context.solved_plan_validation_summary = "5/5 rollouts ok" - return _assistant_response("Validated 5/5; submitting.") - - explorer._agent_session.query = query_impl - policy, term_fn = explorer._get_exploration_strategy(0, timeout=5) - assert callable(policy) and term_fn(_make_state()) is False - # Verbatim: no refinement rollouts, verdict True, capture consumed. - assert not option_model.get_next_state_and_num_actions.called - assert tool_context.last_mental_model_solved is True - assert tool_context.solved_plan is None - assert "belief-certified" in tool_context.cycle_scheduled_plans[-1] - assert "replayed" not in tool_context.cycle_scheduled_plans[-1] - assert tool_context.last_sketch_options == [("Pick", ["block0"]), - ("Place", ["block0", - "block1"])] - # The policy runs the captured options with their captured params. - act = policy(_make_state()) - assert isinstance(act, Action) - # Second request of the cycle on the same task: a new query that - # shows the certified plan as already scheduled. - tool_context.last_mental_model_solved = None - policy2, _ = explorer._get_exploration_strategy(0, timeout=5) - assert callable(policy2) - assert len(queries) == 2 - assert "belief-certified" in queries[1] - assert "STRUCTURALLY different" in queries[1] - assert tool_context.last_mental_model_solved is True - - -def test_uncertified_capture_executes_its_plan_verbatim(): - """A capture whose gate verdict is not True (best-effort, flaky) is not - certified: it executes at its captured params as an experiment, with a - False mental-model verdict.""" - _reset_config() - option_model = MagicMock() - grounded_plan, captured_sketch = _make_captured([0.42], [0.11, 0.22]) - explorer, tool_context = _make_explorer(option_model, None) - - async def query_impl(_msg, **_kw): - tool_context.solved_plan = grounded_plan - tool_context.solved_sketch = captured_sketch - tool_context.solved_plan_reached_goal = False - return _assistant_response("Best effort only, no sketch block.") - - explorer._agent_session.query = query_impl - policy, _ = explorer._get_exploration_strategy(0, timeout=5) - assert callable(policy) - assert not option_model.get_next_state_and_num_actions.called - assert "belief-certified" not in tool_context.cycle_scheduled_plans[-1] - assert tool_context.last_mental_model_solved is False diff --git a/tests/test_agent_harness_fixes.py b/tests/test_agent_harness_fixes.py index 5d66f73394..a7023f7d55 100644 --- a/tests/test_agent_harness_fixes.py +++ b/tests/test_agent_harness_fixes.py @@ -13,9 +13,8 @@ from predicators import utils from predicators.agent_sdk.session_base import max_session_log_number -from predicators.agent_sdk.tools.testing import _missing_goal_atoms from predicators.structs import Action, GroundAtom, Object, \ - ParameterizedOption, Predicate, State, Task, Type + ParameterizedOption, Predicate, State, Type def test_max_session_log_number(tmp_path: Path) -> None: @@ -29,20 +28,12 @@ def test_max_session_log_number(tmp_path: Path) -> None: assert max_session_log_number(str(tmp_path)) == 7 -def test_real_episode_step_budget_is_phase_aware() -> None: - """Explore episodes are capped by the interaction-request cap too.""" - utils.reset_config({ - "horizon": 3000, - "max_num_steps_interaction_request": 1000 - }) - assert utils.real_episode_step_budget("explore") == 1000 +def test_real_episode_step_budget_is_the_horizon() -> None: + """Every agent session phase gets the horizon.""" + utils.reset_config({"horizon": 3000}) assert utils.real_episode_step_budget("solve") == 3000 + assert utils.real_episode_step_budget("synthesis") == 3000 assert utils.real_episode_step_budget(None) == 3000 - utils.reset_config({ - "horizon": 3000, - "max_num_steps_interaction_request": 5000 - }) - assert utils.real_episode_step_budget("explore") == 3000 _block_type = Type("block", ["x"]) @@ -106,14 +97,3 @@ def test_strip_latent_wait_targets_keeps_observable_atoms() -> None: assert wait.memory["wait_target_atoms"] == {GroundAtom(far, [block])} assert "wait_target_neg_atoms" not in wait.memory assert other.memory["wait_target_atoms"] == {GroundAtom(far, [block])} - - -def test_missing_goal_atoms_uses_the_goal_classifiers() -> None: - """Atoms that hold are not reported missing even when the goal predicates - are absent from the agent's predicate set.""" - near, far_block = Object("near", _block_type), Object("far", _block_type) - state = State({near: np.array([0.0]), far_block: np.array([1.0])}) - far = _geometric_pred() - goal = {GroundAtom(far, [near]), GroundAtom(far, [far_block])} - task = Task(state, goal) - assert _missing_goal_atoms(task, state) == {GroundAtom(far, [near])} diff --git a/tests/test_agent_sdk_tools.py b/tests/test_agent_sdk_tools.py index 94a57f082d..0f874e005e 100644 --- a/tests/test_agent_sdk_tools.py +++ b/tests/test_agent_sdk_tools.py @@ -1,12 +1,10 @@ """Tests for agent SDK tool enhancements. Validates: -1. submit_plan always saves scene images -2. submit_plan shows "Missing goal atoms" when goal not achieved -3. submit_plan shows object poses on failure -4. format_object_poses helper -5. render_scene_image helper -6. _sync_tool_context sets ctx.env from option model +1. the run_python probe (reset, run, refine, render) +2. format_object_poses helper +3. render_scene_image helper +4. _sync_tool_context sets ctx.env from option model Usage: python tests/test_agent_sdk_tools.py @@ -163,63 +161,6 @@ def _get_valid_option_plan_step(ctx: Any) -> dict[str, Any] | None: return None -def _plan_to_text(plan: Any, ctx: Any) -> str: - """Render structured option-plan steps as the text grammar that submit_plan - now expects (typed object refs + params in []).""" - type_of = {o.name: o.type.name for o in ctx.current_task.init} - lines = [] - for step in plan: - objs = ", ".join(f"{n}:{type_of.get(n, 'object')}" - for n in step["object_names"]) - params = ", ".join(str(p) for p in step["params"]) - lines.append(f"{step['option_name']}({objs})[{params}]") - return "\n".join(lines) - - -def test_option_plan_missing_goal_atoms(ctx: Any) -> None: - """submit_plan reports missing goal atoms when goal not achieved.""" - tools = _make_tools(ctx, ["submit_plan"]) - - step = _get_valid_option_plan_step(ctx) - assert step is not None, "No valid option found for testing" - plan = [step] - - result = _run(tools["submit_plan"]({ - "plan": _plan_to_text(plan, ctx), - "include_atoms": True, - })) - text = result["content"][0]["text"] - - # Three possible outcomes: - if "Goal achieved: False" in text: - # Either the env exposes goal atoms (and we show "Missing goal - # atoms: ...") or it sets goal_nl (and we show that instead, - # to avoid leaking env predicate names to predicate-invention - # agents). - assert ("Missing goal atoms:" in text - or "Goal (natural language):" in text) - print(" PASS: submit_plan (failure diagnostic shown)") - elif "Goal achieved: True" in text: - assert "Missing goal atoms:" not in text - print(" PASS: submit_plan (goal achieved, no missing atoms)") - else: - # Plan failed early (grounding error, NOT INITIABLE, etc.) - assert ("NOT INITIABLE" in text or "FAILURE REASON:" in text - or "EXECUTION ERROR" in text or "Failed to ground" in text) - print(" PASS: submit_plan (plan failed early, " - "goal check not reached)") - - -def test_option_plan_description_submission_split(ctx: Any) -> None: - """submit_plan's description routes exploration to the probe and frames - this tool as the submission path.""" - from predicators.agent_sdk.tools import create_mcp_tools - tool_obj = create_mcp_tools(ctx, tool_names=["submit_plan"])[0] - desc = getattr(tool_obj, "description", "") - assert "run_python" in desc and "SUBMIT" in desc - print(" PASS: submit_plan (submission-split description)") - - def test_run_python_render_annotations(ctx: Any) -> None: """sim.render(annotations=...) overlays temporary geometry for one render (bodies removed after) and surfaces bad annotations as loud errors.""" @@ -272,11 +213,10 @@ def test_run_python_exec_and_persistence(ctx: Any) -> None: def test_run_python_probe_sim(ctx: Any) -> None: """BeliefProbe: reset with mods, full-precision state, run from the - modified state, snapshot/restore - and nothing is ever captured.""" + modified state, snapshot/restore.""" tools = _make_tools(ctx, ["run_python"]) domino = next(o for o in ctx.current_task.init if o.type.name == "domino") robot = next(o for o in ctx.current_task.init if o.type.name == "robot") - ctx.capture_goal_reaching_plans = True prior_dir = ctx.image_save_dir try: with tempfile.TemporaryDirectory() as tmpdir: @@ -298,16 +238,15 @@ def test_run_python_probe_sim(ctx: Any) -> None: result = _run(tools["run_python"]({"code": code})) saved = [f for f in os.listdir(tmpdir) if f.endswith(".png")] finally: - ctx.capture_goal_reaching_plans = False ctx.image_save_dir = prior_dir text = result["content"][0]["text"] assert "modx 0.95" in text assert "steps 1" in text assert "restx 0.95" in text assert "natoms" in text - # sim.run saves the same per-step audit images submit_plan - # does, and reports their paths on each step; render=False (for - # tight sweep loops) skips the render entirely. + # sim.run saves per-step audit images and reports their paths on + # each step; render=False (for tight sweep loops) skips the render + # entirely. assert "quietimg None" in text if saved: assert any("probe_step_0_" in f for f in saved) @@ -315,20 +254,16 @@ def test_run_python_probe_sim(ctx: Any) -> None: assert len(saved) == 1 else: print(" NOTE: rendering not available, image save not checked") - # The probe carries no scoring surface: nothing it ran was captured. - assert ctx.solved_plan is None - print(" PASS: run_python (BeliefProbe reset/run/snapshot, no capture)") + print(" PASS: run_python (BeliefProbe reset/run/snapshot)") def test_run_python_probe_refine(ctx: Any) -> None: - """BeliefProbe.refine searches params from the current state, reports per- - step samples and a refined plan line, and captures nothing.""" + """BeliefProbe.refine searches params from the current state and reports + per-step samples and a refined plan line.""" tools = _make_tools(ctx, ["run_python"]) domino = next(o for o in ctx.current_task.init if o.type.name == "domino") robot = next(o for o in ctx.current_task.init if o.type.name == "robot") - ctx.capture_goal_reaching_plans = True - try: - code = f""" + code = f""" sim.reset() res = sim.refine( "Pick({robot.name}:robot, {domino.name}:domino)[0.06] " @@ -338,15 +273,12 @@ def test_run_python_probe_refine(ctx: Any) -> None: print("samples", res.total_samples, res.step_samples) print("line", res.plan_lines[0]) """ - result = _run(tools["run_python"]({"code": code})) - finally: - ctx.capture_goal_reaching_plans = False + result = _run(tools["run_python"]({"code": code})) text = result["content"][0]["text"] assert "success True" in text assert "samples" in text assert "line Pick(" in text and "Holding(" in text - assert ctx.solved_plan is None - print(" PASS: run_python (BeliefProbe.refine, no capture)") + print(" PASS: run_python (BeliefProbe.refine)") def test_run_python_probe_refine_verdict_line(ctx: Any) -> None: @@ -526,112 +458,6 @@ def test_run_python_run_contacts(ctx: Any) -> None: print(" PASS: run_python (contact recording)") -def test_option_plan_not_initiable_shows_poses(ctx: Any) -> None: - """submit_plan shows object poses when option is NOT INITIABLE.""" - tools = _make_tools(ctx, ["submit_plan"]) - - # Find Place option and try it without Pick first - place_opt = None - for opt in ctx.options: - if opt.name == "Place": - place_opt = opt - break - - if place_opt is None: - print(" SKIP: submit_plan (no Place option)") - return - - # Build object names from types - state = ctx.current_task.init - obj_names = [] - for t in place_opt.types: - for obj in state: - if obj.type == t and obj.name not in obj_names: - obj_names.append(obj.name) - break - - low = place_opt.params_space.low - high = place_opt.params_space.high - params = ((low + high) / 2).tolist() - - plan = [{ - "option_name": "Place", - "object_names": obj_names, - "params": params, - }] - - result = _run(tools["submit_plan"]({ - "plan": _plan_to_text(plan, ctx), - })) - text = result["content"][0]["text"] - - if "NOT INITIABLE" in text: - assert "Object poses at failure:" in text - print(" PASS: submit_plan (NOT INITIABLE shows poses)") - elif "Failed to ground" in text: - print(" SKIP: submit_plan (Place could not be grounded)") - else: - print(" SKIP: submit_plan (Place was initiable, " - "can't test NOT INITIABLE path)") - - -def test_option_plan_saves_images(ctx: Any) -> None: - """submit_plan always saves scene images (never returns inline).""" - with tempfile.TemporaryDirectory() as tmpdir: - ctx.image_save_dir = tmpdir - - tools = _make_tools(ctx, ["submit_plan"]) - - step = _get_valid_option_plan_step(ctx) - assert step is not None, "No valid option found for testing" - plan = [step] - - result = _run(tools["submit_plan"]({ - "plan": _plan_to_text(plan, ctx), - })) - - content = result["content"] - # Should have text block only (no inline images) - assert any(b["type"] == "text" for b in content) - assert not any(b["type"] == "image" for b in content) - - # Check files were saved if env rendering works - saved = [f for f in os.listdir(tmpdir) if f.endswith(".png")] - if saved: - print(f" PASS: submit_plan ({len(saved)} images saved)") - else: - print(" SKIP: submit_plan (rendering not available)") - - ctx.image_save_dir = None - - -def test_option_plan_failure_shows_poses(ctx: Any) -> None: - """submit_plan shows object poses when option returns 0 actions.""" - tools = _make_tools(ctx, ["submit_plan"]) - - step = _get_valid_option_plan_step(ctx) - assert step is not None, "No valid option found for testing" - plan = [step] - - result = _run(tools["submit_plan"]({ - "plan": _plan_to_text(plan, ctx), - })) - text = result["content"][0]["text"] - - # Check the output is well-formed — it should have either step info - # or a grounding error - assert ("Step 0:" in text or "Failed to ground" in text - or "Testing option plan" in text) - if "FAILURE REASON:" in text: - assert "Object poses at failure:" in text - print(" PASS: submit_plan (failure shows poses)") - elif "NOT INITIABLE" in text: - assert "Object poses at failure:" in text - print(" PASS: submit_plan (NOT INITIABLE shows poses)") - else: - print(" PASS: submit_plan (no failures in output)") - - def testformat_object_poses(ctx: Any) -> None: """format_object_poses formats object positions correctly.""" from predicators.agent_sdk.tools import format_object_poses @@ -732,21 +558,14 @@ def main() -> None: print("=== Tool Enhancement Tests ===\n") - # submit_plan tests - print("1. submit_plan tests:") - test_option_plan_missing_goal_atoms(ctx) - test_option_plan_not_initiable_shows_poses(ctx) - test_option_plan_saves_images(ctx) - test_option_plan_failure_shows_poses(ctx) - # Helper function tests - print("\n2. Helper function tests:") + print("1. Helper function tests:") testformat_object_poses(ctx) testrender_scene_image(ctx) test_render_scene_no_env(ctx) # _sync_tool_context test (creates fresh env) - print("\n3. Context sync tests:") + print("\n2. Context sync tests:") test_sync_tool_context_sets_env() print("\n=== All tests passed! ===") diff --git a/tests/test_benchmark_plots.py b/tests/test_benchmark_plots.py index af53d8f498..fab8f18200 100644 --- a/tests/test_benchmark_plots.py +++ b/tests/test_benchmark_plots.py @@ -54,22 +54,18 @@ def test_current_fan_is_only_ramp() -> None: assert plot.DISPLAY_TITLE["Fan (ramp transfer)"] == "Fan" -def test_default_fan_config_preserves_archived_variants() -> None: - """The menu promotes the reviewed layout without altering old pilots.""" +def test_benchmark_fan_is_the_ramp_transfer() -> None: + """The benchmark's Fan setting is the reviewed ramp layout the paper + figures plot.""" from scripts.cluster_utils import \ parse_configs # pylint: disable=import-outside-toplevel - config = next(parse_configs("predicatorv3/envs/continual.yaml")) + config = next(parse_configs("empiric/envs.yaml")) fan = config["ENVS"]["fan"]["FLAGS"] assert fan["fan_ramp_transfer"] assert fan["fan_inertial_transfer"] assert fan["fan_exposed_transfer"] assert fan["fan_ramp_rise"] == 0.003 assert fan["fan_ramp_landing_extension"] == 0.10 - for variant in ("transfer", "inertial", "ramp"): - pilot = next( - parse_configs( - f"predicatorv3/continual_fan_{variant}_pilot_r1.yaml")) - assert pilot["ENVS"][f"fan_{variant}"]["EXTENDS"] == "fan_maze" def test_oracle_r2_layout_and_scope() -> None: diff --git a/tests/test_cluster_utils_configs.py b/tests/test_cluster_utils_configs.py index 7e26c28b92..3b93cc59d6 100644 --- a/tests/test_cluster_utils_configs.py +++ b/tests/test_cluster_utils_configs.py @@ -1,29 +1,132 @@ -"""The experiment config loader: includes, parked menu entries and EXTENDS.""" +"""The experiment config loader and the EMPIRIC benchmark config: includes, +parked entries, rounds and launch subsets.""" import os -from typing import Any, Dict +from typing import Any, Dict, List import pytest import yaml -from scripts.cluster_utils import _resolve_extends, generate_run_configs +from scripts.cluster_utils import SingleSeedRunConfig, generate_run_configs, \ + parse_seed_range +BENCHMARK = "empiric/benchmark.yaml" -def test_oracle_defaults_preserve_repaired_pilot_policy() -> None: - """Both Oracle presets retain the validated nonblocking policy.""" - with open("scripts/configs/predicatorv3/approaches/continual.yaml", - encoding="utf-8") as stream: - approaches = yaml.safe_load(stream)["APPROACHES"] - for name in ("oracle_dynamics_opus", "oracle_dynamics_sonnet"): - flags = approaches[name]["FLAGS"] - assert flags["continual_skill_preflight"] is False - assert flags["continual_validation_audit"] is False - runs = list( - generate_run_configs( - "predicatorv3/continual_oracle_validation_r2.yaml", False)) - assert len(runs) == 6 + +def _benchmark_runs(**kwargs: Any) -> List[SingleSeedRunConfig]: + runs = [] + for config in generate_run_configs(BENCHMARK, False, **kwargs): + assert isinstance(config, SingleSeedRunConfig) + runs.append(config) + return runs + + +def test_benchmark_is_seven_arms_on_five_settings() -> None: + """The default config runs the seven arms on the five settings, five seeds + each, with each arm's capability contract.""" + runs = _benchmark_runs(round_name="r9") + assert len(runs) == 7 * 5 * 5 + assert {r.seed for r in runs} == set(range(5)) + assert len({(r.experiment_id, r.seed) for r in runs}) == len(runs) + assert {r.env + for r in runs} == { + "pybullet_balloons", "pybullet_bridge", "pybullet_boil", + "pybullet_fan", "pybullet_domino" + } + assert {r.experiment_id.split("-", 1)[1] + for r in runs} == { + f"{arm}_opus_r9" + for arm in ("mb", "mf", "mf_scene_package", "standalone", + "oracle_dynamics", "no_fitting", "no_uncertainty") + } + # The scene-package arm shares the model-free class, not its flags. + assert len({r.approach for r in runs}) == 6 for run in runs: - assert run.flags["continual_skill_preflight"] is False - assert run.flags["continual_validation_audit"] is False + flags = run.flags + assert flags["agent_sdk_model_name"] == "claude-opus-5" + assert flags["experiment_protocol"] == "continual" + assert flags["partially_observable"] + assert not flags.get("continual_skill_preflight", False) + assert not flags.get("continual_validation_audit", False) + assert flags["continual_wall_clock_hours"] == 48.0 + assert "auto_resume" in run.args + # The principled joint belief is the default; only No uncertainty + # turns it off. + assert not flags["code_sim_learning_carry_posterior"] + assert flags["belief_joint_draws"] == ( + 0 if run.approach == "agent_continual_no_uncertainty" else 16) + if run.approach == "agent_continual_model_free": + assert not flags["agent_planner_use_simulator"] + assert not flags["continual_uncertainty_decisions"] + assert bool(flags.get("continual_provide_scene_package")) == ( + "mf_scene_package" in run.experiment_id) + elif run.approach == "agent_continual_no_fitting": + assert flags["agent_sim_learn_declared_params_only"] + assert flags["continual_uncertainty_decisions"] + assert flags["continual_require_model_on_test"] + elif run.approach == "agent_continual_no_uncertainty": + for name in ("continual_obs_noise_declared", + "continual_uncertainty_decisions", + "agent_sim_learn_param_uncertainty", + "code_sim_learning_interval_belief", + "code_sim_learning_rollout_noise_filter", + "continual_belief_frame"): + assert not flags[name] + assert flags["continual_require_model_on_test"] + elif run.approach == "agent_continual_oracle_dynamics": + # The repaired Oracle policy: no automatic execution gate or + # shadow audit. + assert flags["continual_skill_preflight"] is False + assert flags["continual_validation_audit"] is False + if run.env == "pybullet_bridge": + assert flags["bridge_train_span_blocks"] == 3 + assert flags["bridge_test_span_blocks"] == 4 + assert flags["continual_steps_per_level"] == 10000 + else: + assert flags["continual_steps_per_level"] == 5000 + if run.env == "pybullet_balloons": + assert flags["num_train_tasks"] == 2 + assert flags["balloons_goal_dwell_steps"] == 25 + if run.env == "pybullet_fan": + assert flags["fan_ramp_transfer"] + assert flags["fan_ramp_rise"] == 0.003 + if run.env == "pybullet_boil": + assert flags["boil_num_jugs_test"] == [2] + + +def test_benchmark_subsets_and_rounds() -> None: + """--envs, --approaches and --seeds pick a subset; --round names every + experiment id.""" + runs = _benchmark_runs(round_name="fan_fix_r1", + envs=["fan"], + approaches=["mb_opus", "mf_opus"], + seeds=parse_seed_range("2-4")) + assert sorted({r.experiment_id + for r in runs + }) == ["fan-mb_opus_fan_fix_r1", "fan-mf_opus_fan_fix_r1"] + assert sorted(r.seed for r in runs) == [2, 2, 3, 3, 4, 4] + with pytest.raises(ValueError, match="unknown envs"): + _benchmark_runs(round_name="r1", envs=["fan_maze"]) + with pytest.raises(ValueError, match="unknown approaches"): + _benchmark_runs(round_name="r1", approaches=["mb_sonnet"]) + with pytest.raises(ValueError, match="letters, digits"): + _benchmark_runs(round_name="r 1") + + +def test_continual_launch_requires_a_round() -> None: + """A launch that could resume an earlier one's run folders fails until it + names a round.""" + with pytest.raises(ValueError, match="name a round"): + _benchmark_runs(require_round=True) + assert _benchmark_runs(round_name="r2", require_round=True) + + +def test_parse_seed_range() -> None: + """Seed ranges are N or N-M, inclusive.""" + assert parse_seed_range("3") == (3, 1) + assert parse_seed_range("0-4") == (0, 5) + for bad in ("4-2", "a", "1-", "-3"): + with pytest.raises(ValueError): + parse_seed_range(bad) def _write(configs_dir: str, name: str, content: Dict[str, Any]) -> None: @@ -31,10 +134,11 @@ def _write(configs_dir: str, name: str, content: Dict[str, Any]) -> None: yaml.safe_dump(content, f) -def test_extends_gives_a_menu_entry_a_new_id(monkeypatch: Any, - tmp_path: Any) -> None: - """A launcher un-parks a menu env and derives round-specific arms from - parked menu arms with EXTENDS; ids, names and merged flags follow.""" +def test_includes_skip_round_and_precedence(monkeypatch: Any, + tmp_path: Any) -> None: + """A launcher includes menus and a common file; parked entries stay out + unless selected, the ROUND key names the ids, and env flags override arm + flags, which override the common ones.""" configs_dir = tmp_path / "configs" os.makedirs(configs_dir / "menu") _write( @@ -52,7 +156,6 @@ def test_extends_gives_a_menu_entry_a_new_id(monkeypatch: Any, "ENVS": { "balloons": { "NAME": "pybullet_balloons", - "SKIP": True, "FLAGS": { "over": "env" } @@ -68,7 +171,6 @@ def test_extends_gives_a_menu_entry_a_new_id(monkeypatch: Any, "APPROACHES": { "mb": { "NAME": "agent_continual", - "SKIP": True, "FLAGS": { "gate": True, "over": "arm" @@ -79,21 +181,12 @@ def test_extends_gives_a_menu_entry_a_new_id(monkeypatch: Any, _write( str(configs_dir), "launch.yaml", { "includes": ["common.yaml", "menu/envs.yaml", "menu/arms.yaml"], - "ENVS": { - "balloons": { - "SKIP": False - } - }, + "ROUND": "r2", "APPROACHES": { - "mb_r2": { - "EXTENDS": "mb", + "mb": { "FLAGS": { "round": 2 } - }, - "mb_parked": { - "EXTENDS": "mb", - "SKIP": True } } }) @@ -107,34 +200,12 @@ def test_extends_gives_a_menu_entry_a_new_id(monkeypatch: Any, assert run.flags["gate"] is True assert run.flags["round"] == 2 assert run.flags["shared"] == 1 - # Env flags override arm flags, which override the common ones. assert run.flags["over"] == "env" - - -def test_extends_rejects_unknown_and_chained_bases() -> None: - """EXTENDS names a menu entry of the same section, one level deep.""" - with pytest.raises(ValueError, match="unknown"): - _resolve_extends({"a": {"EXTENDS": "missing"}}) - with pytest.raises(ValueError, match="itself EXTENDS"): - _resolve_extends({ - "base": { - "NAME": "x" - }, - "mid": { - "EXTENDS": "base" - }, - "top": { - "EXTENDS": "mid" - } - }) - resolved = _resolve_extends({ - "base": { - "NAME": "x", - "SKIP": True - }, - "top": { - "EXTENDS": "base" - } - }) - assert resolved["base"]["SKIP"] is True - assert resolved["top"] == {"NAME": "x", "SKIP": False} + # A command-line round overrides the config's; naming a parked entry + # launches it. + runs = list( + generate_run_configs("launch.yaml", + batch_seeds=True, + round_name="r3", + envs=["bridge"])) + assert [r.experiment_id for r in runs] == ["bridge-mb_r3"] diff --git a/tests/test_docker_option_plan.py b/tests/test_docker_option_plan.py index 9b12d164b2..c94db7ea2a 100644 --- a/tests/test_docker_option_plan.py +++ b/tests/test_docker_option_plan.py @@ -1,4 +1,4 @@ -"""Test that submit_plan produces correct results. +"""Test that multi-step option plans execute correctly. Validates that multi-step option plans (Pick→Place→Pick→Place→Push) produce non-zero actions at every step, both in-process and in a subprocess that @@ -28,7 +28,7 @@ import predicators.utils as pred_utils from predicators.settings import CFG -# Config matching predicatorv3/predicator_v3.yaml (mf_agent approach) +# Flags of the retired phased mf_agent config. _CFG_OVERRIDES = { "env": "pybullet_domino", "approach": "agent_model_free", diff --git a/tests/test_five_seed_benchmark.py b/tests/test_five_seed_benchmark.py deleted file mode 100644 index 66a5abb951..0000000000 --- a/tests/test_five_seed_benchmark.py +++ /dev/null @@ -1,54 +0,0 @@ -"""The additional paper seeds preserve the intended capability contracts.""" -from scripts.cluster_utils import SingleSeedRunConfig, generate_run_configs - - -def test_five_seed_launch_contracts() -> None: - """Exactly six arms, five domains, and two fresh seeds with matched - costs.""" - runs = [] - for config in generate_run_configs( - "predicatorv3/continual_benchmark_five_seeds.yaml", False): - assert isinstance(config, SingleSeedRunConfig) - runs.append(config) - assert len(runs) == 60 - assert {r.seed for r in runs} == {3, 4} - assert len({(r.experiment_id, r.seed) for r in runs}) == 60 - assert len({r.env for r in runs}) == 5 - # Direct + scene shares the direct-agent class, not its capability flags. - assert len({r.approach for r in runs}) == 5 - assert len({r.experiment_id.split("-", 1)[1] for r in runs}) == 6 - for run in runs: - flags = run.flags - assert flags["agent_sdk_model_name"] == "claude-opus-5" - assert flags["partially_observable"] - assert not flags["continual_skill_preflight"] - assert not flags["continual_validation_audit"] - assert flags["continual_wall_clock_hours"] == 48.0 - assert "auto_resume" in run.args - if run.approach == "agent_continual_model_free": - assert not flags["agent_planner_use_simulator"] - assert not flags["continual_uncertainty_decisions"] - assert bool(flags.get("continual_provide_scene_package")) == ( - "mf_scene_package" in run.experiment_id) - elif run.approach == "agent_continual_no_fitting": - assert flags["agent_sim_learn_declared_params_only"] - assert flags["continual_uncertainty_decisions"] - assert flags["continual_require_model_on_test"] - elif run.approach == "agent_continual_no_uncertainty": - for name in ("continual_obs_noise_declared", - "continual_uncertainty_decisions", - "agent_sim_learn_param_uncertainty", - "code_sim_learning_interval_belief", - "code_sim_learning_carry_posterior", - "code_sim_learning_rollout_noise_filter", - "continual_belief_frame"): - assert not flags[name] - assert flags["continual_require_model_on_test"] - if run.env == "pybullet_bridge": - assert flags["bridge_test_span_blocks"] == 4 - assert flags["continual_steps_per_level"] == 10000 - else: - assert flags["continual_steps_per_level"] == 5000 - if run.env == "pybullet_balloons": - assert flags["num_train_tasks"] == 2 - assert flags["balloons_goal_dwell_steps"] == 25 diff --git a/tests/test_main.py b/tests/test_main.py index 65f2587eb6..fda9517ac9 100644 --- a/tests/test_main.py +++ b/tests/test_main.py @@ -13,8 +13,8 @@ from predicators import utils from predicators.approaches import ApproachFailure, ApproachTimeout, \ BaseApproach, create_approach -from predicators.approaches.agent_model_free_approach import \ - AgentModelFreeApproach +from predicators.approaches.pp_online_process_learning_approach import \ + OnlineProcessLearningAndPlanningApproach from predicators.cogman import CogMan from predicators.envs.cover import CoverEnv from predicators.execution_monitoring import create_execution_monitor @@ -465,9 +465,9 @@ def test_stash_resume_restores_request_bookkeeping(): get_interaction_requests, so the result->train-task pairing that learn_from_interaction_results needs must come from restore_interaction_requests (run_20260828_173451 asserted on it).""" - # The model-free family records the pairing in get_interaction_requests - # and asserts on it in learn_from_interaction_results. - approach = object.__new__(AgentModelFreeApproach) + # The online process learner records the pairing in + # get_interaction_requests and reads it in learn_from_interaction_results. + approach = object.__new__(OnlineProcessLearningAndPlanningApproach) approach._requests_train_task_idxs = None # pylint: disable=protected-access approach.restore_interaction_requests([0, 0]) assert approach._requests_train_task_idxs == [0, 0] # pylint: disable=protected-access