Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
62 changes: 62 additions & 0 deletions .github/workflows/flow-modes-verify.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
name: Managed and direct flow modes

on:
workflow_dispatch:
push:
paths:
- '.github/workflows/flow-modes-verify.yml'
- 'SKILL.md'
- 'flow/**'
- 'tools/**'
- 'scripts/**'
- 'tests/**'
- 'setup/**'
- 'toolchain.json'
pull_request:
paths:
- '.github/workflows/flow-modes-verify.yml'
- 'SKILL.md'
- 'flow/**'
- 'tools/**'
- 'scripts/**'
- 'tests/**'
- 'setup/**'
- 'toolchain.json'

permissions:
contents: read

jobs:
tool-modes:
runs-on: ubuntu-24.04
timeout-minutes: 20
env:
PYTHONPATH: ''
NAJAEDA_SRC: ''
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.13'
- name: Offline skills, mode routing and proof guards
run: python -m unittest discover -s tests -v
- name: Install or reuse packaged tools and check MCP discovery
run: |
mkdir -p .cache/mode-agent
python setup/mcp.py configure --client codex --project .cache/mode-agent \
--venv .cache/mode-python --apply
.cache/mode-python/bin/python setup/mcp.py check --venv .cache/mode-python
- name: Managed history inside a real Jupyter kernel
run: .cache/mode-python/bin/python scripts/versioned_session_regression.py --jupyter --work-dir runs/mode-managed
- name: Direct NajaEDA and MCP tools without the flow helper
run: .cache/mode-python/bin/python scripts/direct_tools_regression.py --work-dir runs/mode-direct
- name: Preserve actual mode evidence
if: always()
uses: actions/upload-artifact@v4
with:
name: flow-mode-regressions
path: |
runs/mode-managed/
runs/mode-direct/
if-no-files-found: warn
retention-days: 14
98 changes: 98 additions & 0 deletions .github/workflows/gcd-undo-verify.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
name: GCD packaged undo

on:
workflow_dispatch:
push:
paths:
- '.github/workflows/gcd-undo-verify.yml'
- 'scripts/**'
- 'setup/**'
- 'examples/backend/gcd/**'
- 'toolchain.json'
- 'tools/**'
- 'tests/**'
pull_request:
paths:
- '.github/workflows/gcd-undo-verify.yml'
- 'scripts/**'
- 'setup/**'
- 'examples/backend/gcd/**'
- 'toolchain.json'
- 'tools/**'
- 'tests/**'

permissions:
contents: read

jobs:
undo-replay:
runs-on: ubuntu-24.04
timeout-minutes: 45
defaults:
run:
shell: bash
env:
RUN: runs/gcd-undo
PYTHONPATH: ''
NAJAEDA_SRC: ''
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.13'
- name: Offline checks
run: python -m unittest discover -s tests -v
- name: Install Nix (reuse an existing installation)
uses: cachix/install-nix-action@v31
with:
install_url: https://releases.nixos.org/nix/nix-2.35.1/install
enable_kvm: false
extra_nix_config: |
experimental-features = nix-command flakes
max-jobs = 0
builders =
- name: Obtain cached OpenROAD (fail before downloading dependencies on a cache miss)
run: |
mkdir -p .cache/gcd-tools
df -h . /nix | tee .cache/gcd-tools/disk-before.txt
openroad=$(python -c 'import json; print(json.load(open("toolchain.json"))["openroad"]["installable"])')
cache=$(python -c 'import json; print(json.load(open("toolchain.json"))["openroad"]["cache"])')
bash tools/install-cached-package.sh openroad "$openroad" "$cache" .cache/gcd-tools
echo "$PWD/.cache/gcd-tools/openroad/bin" >> "$GITHUB_PATH"
- name: Install native Python wheels and pinned pure-Python Kepler MCP package
run: |
mkdir -p .cache/agent-config-codex .cache/agent-config-claude
python setup/mcp.py configure --client codex --project .cache/agent-config-codex \
--venv .cache/gcd-python --apply | tee .cache/gcd-tools/setup-codex.log
python setup/mcp.py configure --client claude-code --project .cache/agent-config-claude \
--venv .cache/gcd-python --apply | tee .cache/gcd-tools/setup-claude.log
.cache/gcd-python/bin/python -m pip freeze > .cache/gcd-tools/python-packages.txt
echo "$PWD/.cache/gcd-python/bin" >> "$GITHUB_PATH"
- name: Agent MCP discovery and existing live-session regressions
run: |
python setup/mcp.py check --venv .cache/gcd-python
python scripts/agent_mcp_regression.py --work-dir runs/agent-mcp
python scripts/live_session_regression.py --work-dir runs/live-session
python scripts/live_inspection_regression.py --work-dir runs/live-inspection
- name: Version history retention, failures, undo and continued editing
run: python scripts/versioned_session_regression.py --jupyter --work-dir runs/versioned-session
- name: GCD baseline, edit and undo with Scope, SEC and three fresh OpenROAD runs
run: |
openroad -version | tee .cache/gcd-tools/openroad-version.txt
python scripts/gcd_undo_regression.py --work-dir "$RUN"
- name: Preserve baseline, edited and restored observations and reports
if: always()
uses: actions/upload-artifact@v4
with:
name: gcd-packaged-undo
path: |
runs/gcd-undo/
runs/versioned-session/
runs/live-inspection/
runs/agent-mcp/
runs/live-session/
.cache/gcd-tools/*.json
.cache/gcd-tools/*.txt
.cache/gcd-tools/*.log
if-no-files-found: warn
retention-days: 14
3 changes: 2 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# Working In 22b

- Read [SKILL.md](SKILL.md), then the selected flow skill. Load tool references
- Read [SKILL.md](SKILL.md), then the selected flow and execution-mode skills.
Honor direct mode without importing the flow helper. Load tool references
only when their operation is needed.
- Keep reusable tool knowledge in `tools/`, application guidance in `flow/`,
and design-specific inputs and recipes in the matching `examples/` directory.
Expand Down
37 changes: 28 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@ measurement. There is no chatbot runtime or model dependency in this repository.

```text
SKILL.md Parent orchestration skill
flow/backend/ Physical-design analysis and improvement
flow/rtl/ RTL authoring and design changes
flow/backend/ Backend guidance, managed/ and direct/ skills
flow/rtl/ RTL guidance, managed/ and direct/ skills
examples/backend/gcd/ GCD design and independent model task
examples/rtl/ RTL example conventions
tools/ Shared tool skills and package installation guides
Expand All @@ -36,17 +36,33 @@ if the inline player is unavailable.
Kepler Formal runs through its
Python-backed MCP with native wheels; OpenROAD uses Nix. No source submodules
are required in 22b.
2. Choose the [backend](flow/backend/SKILL.md) or [RTL](flow/rtl/SKILL.md) flow.
2. Choose the [backend](flow/backend/SKILL.md) or [RTL](flow/rtl/SKILL.md) flow,
then its managed or direct execution flavor.
3. For a concrete backend attempt, give the model the [GCD task](examples/backend/gcd/task.md).

An agent can read these files directly. [AGENTS.md](AGENTS.md) points agents to
the same entry point; human users can follow the same procedures. Skills are
instructions, not a security sandbox. For iterative work, the optional
[persistent Python/Jupyter session](tools/live-session.md) validates each edit
and automatically runs SEC on the cumulative candidate against unchanged
golden, without design dumps or reloads between edits. The model stays in the
same kernel throughout. The pinned MCP includes its attached-session report API;
the existing file-based flow is unchanged.
instructions, not a security sandbox. Both flows offer two flavors:

| Flow | Helper-managed | Direct tools, no flow helper |
| --- | --- | --- |
| Backend | [Managed skill](flow/backend/managed/SKILL.md) | [Direct skill](flow/backend/direct/SKILL.md) |
| RTL | [Managed skill](flow/rtl/managed/SKILL.md) | [Direct skill](flow/rtl/direct/SKILL.md) |

Honor the requested flavor; managed is the default for iterative structural
work. Both share [session policies](flow/session-policy.md) and tool documentation.
Direct mode follows an explicit [file-based recipe](flow/direct-revisions.md),
with NajaEDA Python and direct Scope/Kepler MCP calls. Required checks are agent
responsibilities, not automatically enforced by instructions. Neither mode adds
an RTL elaboration frontend or changes tool input support.

Managed mode starts the [versioned Python/Jupyter session](tools/live-session.md).
It validates each edit,
runs SEC against unchanged golden and separately verifies numbered Verilog
checkpoints. It keeps ten recent edits plus baseline by default, supports undo,
and can retain a measured best result independently. The model stays in the same
kernel; only undo reloads the candidate. The original no-export helper and
existing file-based flow remain available without changing existing callers.

## Verification And Evidence

Expand All @@ -70,3 +86,6 @@ The separate [GCD reference workflow](.github/workflows/gcd-reference-verify.yml
tests package installation and real tool stages using the saved solution under
[reference/](examples/backend/gcd/reference/README.md). It uses no model, and its
success does not establish that a model can solve the independent task.
The [mode regression](.github/workflows/flow-modes-verify.yml) checks both skill
routes, real managed history and helper-free direct tool execution. It tests
structural fixtures, not a model's ability to follow skills or synthesize RTL.
22 changes: 12 additions & 10 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,10 @@ description: Coordinate open-source hardware design tools for backend optimizati

- For a mapped design and physical reports, read [backend](flow/backend/SKILL.md).
- For RTL creation or changes, read [RTL](flow/rtl/SKILL.md).
- Select that flow's **managed** or **direct** skill before editing. Honor an
explicitly requested mode; otherwise use managed for iterative structural
work. Read only the selected mode, not both. Switching requires a deliberate
handoff of inputs and evidence, not a silent fallback after a failure.
- Read [package setup](tools/README.md) only when a needed tool is absent or its
version does not match the experiment. Check existing installations first.
- If Kepler tools are absent from the agent's own tool list, use
Expand All @@ -32,18 +36,16 @@ or comparison afterward, not hints for independent discovery.
stale output files as evidence for a new run.
3. Inspect using reports and, when structural connectivity matters,
[Naja-Scope](tools/naja-scope/SKILL.md). Separate observations from hypotheses.
In a live editing session, refresh Scope only when the next decision needs
current connectivity; use its revision-labelled inspection checkpoint, not
a stale copy left from an earlier edit.
Refresh Scope when the next decision needs a different numbered revision;
do not reuse a stale or historical copy for a current-design question.
4. Use [NajaEDA](tools/najaeda/SKILL.md) for structural edits. Review and syntax
check generated code before running it with only the needed file access.
5. Run [Kepler Formal SEC through MCP](tools/kepler-formal/SKILL.md). For iterative
in-memory work, use the [persistent session](tools/live-session.md): keep one
unchanged golden and one cumulatively edited candidate, with automatic SEC
after every edit and no design dumps for verification. Optional inspection
copies never replace either live design. If a design is later
exported for another tool, verify that exported representation separately.
Preserve the structured outcome, logs and actual output coverage.
5. Run [Kepler Formal SEC through MCP](tools/kepler-formal/SKILL.md) and retain
actual outcomes and coverage. Managed mode enforces live SEC and separately
checks exported checkpoints through its helper. Direct mode calls the tools
explicitly; the skills require the checks but do not mechanically enforce
them. Follow the shared [session policy](flow/session-policy.md). Never
describe a live proof as proof of an exported file.
6. For backend tasks, rerun [OpenROAD](tools/openroad/SKILL.md) with the same
physical setup. Compare timing, area, estimated power, hold and routing checks.

Expand Down
11 changes: 11 additions & 0 deletions flow/backend/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,17 @@ description: Improve a synthesized design using OpenROAD physical reports, Naja-

# Backend Improvement

Select one execution flavor:

- [Managed](managed/SKILL.md): the Python session helper handles edits, SEC,
checkpoints and undo. Default for iterative structural work.
- [Direct](direct/SKILL.md): the agent coordinates tools and files without the
flow helper. Use when explicitly requested; do not silently switch to managed.

Both follow the [session policy](../session-policy.md). Measurements belong to
an exact numbered revision and unchanged physical setup, not just "the latest"
filename. Scope must inspect the restored revision after undo.

Read the [parent contract](../../SKILL.md). Start from mapped Verilog, Liberty,
LEF/technology data and an SDC; RTL synthesis is not implicit in this flow.

Expand Down
36 changes: 36 additions & 0 deletions flow/backend/direct/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
---
name: backend-direct
description: Run backend optimization through NajaEDA Python, Naja-Scope MCP, Kepler Formal MCP and OpenROAD directly, with agent-managed revision files and no 22b flow helper.
---

# Direct Backend

Follow the [backend objectives](../SKILL.md),
[shared session policy](../../session-policy.md) and
[direct revision recipe](../../direct-revisions.md). Do not instantiate a 22b
session helper or use its checkpoint, undo or Scope adapter. The agent performs
the documented steps; no Python flow controller is installed by this skill.

1. Preserve baseline design, libraries and constraints. Run baseline OpenROAD
and keep the reports before proposing an edit.
2. Select the numbered design file, load it in the separate
[Scope MCP](../../../tools/naja-scope/SKILL.md) server, and inspect connectivity.
3. Review and syntax-check the proposed script. Use the
[NajaEDA Python API](../../../tools/najaeda/SKILL.md) in a fresh candidate
process to load the selected file, edit it and export to a new staging path.
NajaEDA is a Python library here, not a separately configured editing MCP.
4. Call the agent's [Kepler MCP tools](../../../tools/kepler-formal/SKILL.md)
directly on golden versus the exported file, explicitly selecting SEC.
Preserve the structured result, reports, hashes and actual coverage.
5. Publish the revision only after interpreting the proof. Measure it with
[OpenROAD](../../../tools/openroad/SKILL.md) under unchanged settings; record
the revision and hashes with every report. Promote best only on comparable
measured improvement under the recorded objective and proof policy.
6. Restore and reverify the previous retained file when undo is requested.
Refresh Scope and external inputs before discarding the undone checkpoint.

Use the agent's registered MCP connections, not a hidden notebook flow client.
The [setup guide](../../../setup/README.md) registers Kepler; Scope has its own
installation/client instructions. Direct mode has no automatic edit validator,
mandatory-SEC gate, retention scheduler or crash recovery. The skill requires
those actions but cannot guarantee the agent performed them: retain evidence.
33 changes: 33 additions & 0 deletions flow/backend/managed/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
---
name: backend-managed
description: Run backend optimization with the 22b Python session helper for cumulative structural edits, checked checkpoints, undo and measured best results. Use for the managed execution flavor, not direct tool orchestration.
---

# Managed Backend

Follow the [backend objectives](../SKILL.md) and
[shared session policy](../../session-policy.md). Use this mode's instructions
only; do not load the direct-mode recipe unless the caller switches modes.

1. Follow [session startup](../../../tools/live-session.md) to create one
`VersionedDesignSession` in a dedicated Python/Jupyter kernel. Keep golden
unchanged. Configure retention (ten edited checkpoints by default) and the
measurement objective using [session history](../../../tools/session-history.md).
2. Review the NajaEDA script, then call `session.apply_edit(script)`. The helper
validates it, runs live SEC, exports a numbered candidate and runs file SEC.
Keep both proof outcomes and actual coverage; a warning is not full proof.
3. Resolve `session.checkpoint()` for current or an explicit revision for
history. Load that file and `session.libraries` into the separate Scope MCP
server. Use `session.use_checkpoint()` to pin inputs during external runs.
4. Run OpenROAD under the same physical setup. Record actual measurements and
reports with `session.record_measurement(...)`; let the configured objective
select best, never substitute estimated improvements.
5. For undo, use `session.undo()`, not manual database reset or file deletion.
After successful restoration, get `session.mcp_attachment()` again and
reattach external Kepler clients. Refresh Scope to the restored checkpoint.

A failed edit may leave an unsaved live candidate; repair it or undo to the last
saved checkpoint. Do not treat that checkpoint as the current live state while
the helper reports unsaved changes. Rejected proofs and tool errors stop progress.
The older no-export helper is an explicit compatibility option, not an automatic
fallback when checkpointing fails.
Loading
Loading