Discover a task in an application UI once, review the resulting capability, then
call it by name through deterministic replay. For example, an agent can request
member_savings_balance(member_id="10003") and receive balance 1411.21 without
asking a model to navigate the application again.
Discovery produces a draft. A reviewer describes and approves its exact content.
Replay uses no model and returns success, business_outcome, failure, or
escalated. A supported browser session can be handed to a person when automation
cannot safely continue. Writes also require invocation-specific consent; artifact
approval alone does not authorize a commit.
Goal → discovery → draft → describe / approve → registered capability
↓
CLI / HTTP / MCP / workflow → deterministic replay → typed result + evidence
↓
human takeover / resume
- Browser: Playwright Chromium, including frames and accessibility-based controls.
- Native desktop: Windows UI Automation in an interactive Windows session,
with the
windowsextra. Desktop screenshots, traces, vision and session handoff are unavailable; unsupported feature requests are refused before action. - Bundled examples are a synthetic legacy banking UI and Windows DeskCalc. Other applications need their own tenant, policy, family and reviewed capability.
Native macOS/Linux adapters, additional browsers and remote browser viewing are future extensions. See platform support and limitations.
Use Python 3.11 or newer and run these commands in a checkout. The demo currently
uses repository resources. Playwright's Chromium binary is installed separately
from the Python package. The short demo starts its own services and runs all seven
stages with three independent replays; no provider key or .env setup is needed.
PowerShell (Windows):
git clone https://github.com/aniaisec/inter-cua.git
Set-Location inter-cua
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e ".[dev]"
.\.venv\Scripts\python.exe -m playwright install chromium
.\.venv\Scripts\python.exe scripts/demo/run_demo.py --repetitions 3POSIX shell (Linux/macOS browser environment):
git clone https://github.com/aniaisec/inter-cua.git
cd inter-cua
python3 -m venv .venv
.venv/bin/python -m pip install -e ".[dev]"
.venv/bin/python -m playwright install chromium
.venv/bin/python scripts/demo/run_demo.py --repetitions 3On Debian/Ubuntu, install the matching Python venv package if python3 -m venv
reports that ensurepip is unavailable (for example sudo apt install python3-venv).
On Linux, missing system browser libraries may require
.venv/bin/python -m playwright install --with-deps chromium with permission to
install OS packages. For later CLI examples, activate with
.\.venv\Scripts\Activate.ps1 in PowerShell or . .venv/bin/activate in POSIX.
If PowerShell activation is restricted, use the environment's executable directly,
for example .\.venv\Scripts\python.exe -m cua.cli --help.
A successful run writes summary.md, summary.json, a masked command manifest and
evidence under a fresh demo/demo_<session-id>/. Expect three successful replays,
zero replay model calls, one approved commit and one blocked hostile-page attack.
Discovery, review and takeover are scripted in this demonstration; it validates
the mechanism, not live model planning or a person's effort.
See getting started, the complete demo guide
and troubleshooting.
An ordinary replay prints JSON. A successful lookup includes these fields (excerpt; the full result also carries timing, run identity and evidence):
{"kind": "success", "outputs": {"savings_balance": "1411.21"}, "side_effect": "none"}| Exit | Replay result | Meaning |
|---|---|---|
| 0 | success |
The operation completed with typed outputs |
| 2 | business_outcome |
A declared answer such as NOT_FOUND |
| 1 | failure |
Inspect the code, evidence and side effect |
| 3 | escalated |
A person or consent is needed; retain the resume token |
Usage/configuration errors generally exit 64; argparse errors exit 2. Decimal
outputs are strings. If side_effect is unknown, reconcile with the application
before repeating a write. Preserve an idempotency key across retries.
Normal discovery and replay evidence defaults to evidence/runs/<run-id>/.
Discovery records model calls; replay uses none. Preflight refusals may create no
run directory. See CLI contracts and operations
for result fields, handoff, consent and evidence handling.
Follow your first application to create a tenant
binding, restrict the policy, define the product's family template, discover a
read-only operation, review it and test a second input and failure. Configuration
is currently manual and relative to the working directory; there is no init or
doctor command yet. See configuration and secret precedence.
Live discovery needs an Anthropic or Gemini provider key and incurs charges. Scripted discovery needs a target-specific tool-call script. Replay, operator and routine tests need no provider key. A happy-path run cannot infer all business outcomes or recovery conditions.
| Entry point | Use it for | Guide |
|---|---|---|
| CLI / catalog | Shell commands and subprocess callers | CLI, caller contract |
| HTTP | Shared local service with authorization and polling | HTTP |
| MCP | Approved capabilities as stdio tools for an agent | MCP |
| Workflows | Compose approved operations with typed bindings | Workflow semantics |
| Operator / resume | Consent and takeover of the same browser session | Operations |
The documentation index links configuration, troubleshooting, architecture, extensions, operations and tests. CI gates describe quality, strict browser verification and benchmark/security artifacts.
In the paired synthetic evaluation, replay used zero model calls and lower latency with a similar exact-result rate; it stopped safely on renamed controls. The evaluation covers one mock application, excludes rerun provider failures, and does not establish performance on other applications. Measured results and qualifications, benchmark reproduction and curated evidence retain the complete tables and cost/human assumptions.
Architecture explains approval, tenant isolation, consent, idempotency, drift and recovery. See the threat model and security reporting policy. REPORT and phase history are historical design snapshots.
MIT.