diff --git a/.github/RELEASE.md b/.github/RELEASE.md new file mode 100644 index 0000000..a33a8a7 --- /dev/null +++ b/.github/RELEASE.md @@ -0,0 +1,59 @@ +# Release publication and catalog notification + +Publish only an accepted plugin package bound to its reviewed source/tag and exact archive digest. +Keep an existing public tag and archive immutable. A candidate, draft or notification receipt does +not establish authenticated vendor or installed-host acceptance. + +The Validate and package workflow produces candidate artifacts; it does not make a release public. After the accepted release becomes public, +`notify-catalog.yml` requests a complete catalog reconciliation. It also observes public edits, +channel promotion, unpublishing and deletion; those events never authorize catalog withdrawal by +themselves. The central publisher retains verified history and applies its reviewed withdrawal +policy. It verifies actual GitHub release sources rather than trusting an event payload. + +## Notification authority + +The workflow pins the website's central notification action to a reviewed full commit. Publish that +central commit before enabling a plugin workflow that references it. Review and update this pin +when adopting changes to the notification contract. The caller checks its immutable repository ID, +does not check out package code, and grants its own job token no repository permissions. + +Supply `CATALOG_DISPATCH_TOKEN` using existing reviewed authority with Actions write access to +`computer-mcp/computer-mcp.github.io` only. Website Contents write access is unnecessary. The action +can also receive an existing temporary token directly from a publishing job. Neither workflow +creates or persists credentials. Missing or rejected authority fails visibly; the +publisher's independent schedule still reconciles missed notifications. + +## Publication and retry + +A manual public release emits the release event. Publication performed with a repository's +`GITHUB_TOKEN` does not trigger ordinary release-event workflows. After that publication succeeds, +its automation must explicitly call this reusable workflow as a dependent job: + +```yaml +notify-catalog: + needs: publish + uses: ./.github/workflows/notify-catalog.yml + secrets: + CATALOG_DISPATCH_TOKEN: ${{ secrets.CATALOG_DISPATCH_TOKEN }} +``` + +Here `publish` is the job that actually makes the accepted release public, not the candidate-build +or draft-upload job. When using an existing short-lived token within that publishing job, invoke +the same pinned central action directly after publication instead. Keep token values out of command +arguments, printed output and release metadata. + +For an operator-driven publication or a missed/failed notification, explicitly dispatch: + +```sh +gh workflow run notify-catalog.yml --repo computer-mcp/plugin-claude --ref main +``` + +This schedules notification using its configured authority; it does not publish or rewrite a +release. Inspect the notification run and its returned central `run_url`. A successful dispatch +proves request acceptance only. Verify the central run completed successfully and the public index +contains the exact expected release identities and generation. If the release is already public and +notification fails, retry notification without changing or republishing the release. Complete +reconciliation is idempotent and repairs duplicate/missed events. + +See the central [catalog publication and notification contract](https://github.com/computer-mcp/computer-mcp.github.io/blob/main/docs/plugin-catalog.md) +for provenance, credentials, retry bounds and deployment semantics. diff --git a/.github/workflows/notify-catalog.yml b/.github/workflows/notify-catalog.yml new file mode 100644 index 0000000..c3ba8c3 --- /dev/null +++ b/.github/workflows/notify-catalog.yml @@ -0,0 +1,28 @@ +name: Notify official plugin catalog +on: + release: + types: [published, edited, released, unpublished, deleted] + workflow_dispatch: + workflow_call: + secrets: + CATALOG_DISPATCH_TOKEN: + description: Existing receiver-scoped Actions write authority + required: true + outputs: + run_url: + description: Accepted central run; verify its deployment separately + value: ${{ jobs.notify.outputs.run_url }} +permissions: {} +jobs: + notify: + if: github.repository_id == '1384746086' + runs-on: ubuntu-latest + timeout-minutes: 3 + outputs: + run_url: ${{ steps.catalog.outputs.run-url }} + steps: + - name: Request complete catalog reconciliation + id: catalog + uses: computer-mcp/computer-mcp.github.io/.github/actions/notify-catalog@fc27dd0f370d028a3e5021e3585274891f696578 + with: + token: ${{ secrets.CATALOG_DISPATCH_TOKEN }} diff --git a/Documentation/Architecture/Package.md b/Documentation/Architecture/Package.md index b7b63b7..b296e53 100644 --- a/Documentation/Architecture/Package.md +++ b/Documentation/Architecture/Package.md @@ -6,4 +6,11 @@ This repository owns the Claude Code CLI description and print/stream-json-to-MC `bin/plugin_runtime.py` is this package's private standard-library runtime for bounded MCP I/O, validation, process supervision and retention. It is shipped with the plugin, not loaded from another plugin or host implementation module. Python 3.13+ must be available on the launch PATH. A separate supervisor lifeline and positive cleanup receipt distinguish process exit from confirmed cleanup. These are internal process mechanisms, not a new host contribution type. +The ordinary MCP work resource projects these same run owners, including +pending startup, uncertain cleanup and retained completed results. Correlation +is bound at run creation and is not an authorization grant. Confirmed worker +and process cleanup are prerequisites for result eviction or explicit release. +Reading a result has no release side effect. The adapter bounds and versions +complete snapshots; the host owns configuration generations and routing. + The documented headless CLI is the selected upstream interface; the full Agent SDK's interactive callbacks and hosted Managed Agents are not substituted or claimed. Tests use deterministic peers and controlled process trees. Scripts validate real native version/help, deterministic archives and the unchanged host in isolated standalone mode. Authenticated backend execution and production activation require separate evidence. diff --git a/Documentation/Reference/Installation.md b/Documentation/Reference/Installation.md index 77faa13..8a0ee42 100644 --- a/Documentation/Reference/Installation.md +++ b/Documentation/Reference/Installation.md @@ -14,6 +14,8 @@ The example exposes this plugin's complete MCP tool catalog with no extra prefix On Computer MCP 1.2.2, plugin activation/selection changes require idle Gateway client admission. Finish or safely pause clients before production installation changes. No host binary replacement or host release is required. Do not restart the active development control connection merely to test installation. +Computer MCP 1.3.0 publishes plugin configuration changes to connected clients atomically. New calls use the current configuration; existing work keeps its owning runtime until release. Installation does not grant access, and later calls use current authorization. + ## Build ```sh @@ -28,16 +30,24 @@ CI runs deterministic fixture tests and packaging; it does not install a vendor ## Isolated host interoperability -After packaging, validate the exact ZIP with an unchanged installed Computer MCP 1.2.2 binary: +After packaging, validate the exact ZIP against the reviewed Computer MCP binary. Select its release version explicitly; a mismatch fails before package extraction: ```sh python3 Scripts/validate_host.py \ - --host "/Applications/Computer MCP.app/Contents/Resources/computer-mcp" \ + --host "/absolute/path/to/candidate/computer-mcp" \ + --expected-host-version 1.3.0 \ --archive /output/PLUGIN.zip \ --output /new/evidence/directory ``` -Replace `PLUGIN.zip` with the package's actual archive name. This uses a temporary directory and the installed host's archive worker and standalone MCP entrypoint. It does not connect to the production App's control socket or database. Vendor tool execution is replaced with inert fixtures; the native version/help check is a separate command. A new evidence directory is required to avoid overwriting an earlier run. +Replace `PLUGIN.zip` with the package's actual archive name. This uses a temporary directory and the selected host's archive worker and standalone MCP entrypoint. It does not connect to the production App's control socket or database. Vendor tool execution is replaced with inert fixtures; the native version/help check is a separate command. A new evidence directory is required to avoid overwriting an earlier run. + +For a separately built candidate host that supports the ordinary MCP work +resource, add `--require-work-ownership`. This also verifies ownership during +background execution, retained cancelled/completed results, and explicit release +on the same connection. The candidate host runs only with isolated configuration +and state. This option is not an authenticated-model or production activation +check. ## Result interpretation diff --git a/Documentation/Reference/Interface.md b/Documentation/Reference/Interface.md index ca2a3f3..8d62892 100644 --- a/Documentation/Reference/Interface.md +++ b/Documentation/Reference/Interface.md @@ -6,7 +6,17 @@ The native version assertion runs before each adapter execution and each host-projected CLI call. Native command availability does not establish account authentication or backend access. The CLI contribution returns a single native JSON result. The MCP adapter independently consumes print-mode stream-json with verbose and partial-message output; it does not parse the terminal UI or use hosted Managed Agents. -The adapter implements newline-delimited MCP JSON-RPC initialization, tools/list, tools/call, ping and cancellation. Supported MCP dates are 2024-11-05, 2025-03-26 and 2025-06-18. Unsupported proposals receive a supported date, not an unimplemented echo. Tool schemas describe accepted arguments. Tool results use `structuredContent.result`; `isError` indicates a failed operation, distinct from a JSON-RPC protocol error. +The adapter implements newline-delimited MCP JSON-RPC initialization, tools/list, tools/call, resources/list, resources/read, ping and cancellation. Supported MCP dates are 2024-11-05, 2025-03-26 and 2025-06-18. Unsupported proposals receive a supported date, not an unimplemented echo. Tool schemas describe accepted arguments. Tool results use `structuredContent.result`; `isError` indicates a failed operation, distinct from a JSON-RPC protocol error. + +## Host risk metadata + +Every MCP tool declares `_meta["io.github.computer-mcp/risk"]`. Model execution +and continuation declare `full-shell`: native permission defaults are not a +host-enforced sandbox. Catalog, result, event and pending-request inspection +declare `read-only`. Cancellation and owned-process retirement declare +`destructive`. The host applies these as minimum classifications, intersects +its own grants, and retains approval authority. Standard MCP annotations remain +hints rather than permissions. ## Tools @@ -18,8 +28,48 @@ The adapter implements newline-delimited MCP JSON-RPC initialization, tools/list | `claude.run.result` | Return current state or an explicitly completed final result | | `claude.run.events` | Read a bounded cursor page of native stream-json events | | `claude.run.cancel` | Cancel one adapter-owned run and wait through normal cleanup on subsequent result reads | - -The host may prefix these names; discover actual projections with tools/list. The adapter retains at most 32 total runs, of which at most four can be active. Run handles and event pages are memory-owned, not a second vendor conversation database. Replacing the adapter invalidates these handles; it does not delete native saved conversations. +| `claude.run.release` | Discard one completed result and its events after process and worker cleanup are confirmed | + +The host may prefix these names; discover actual projections with tools/list. +The adapter retains at most 32 total runs, of which at most four can have active +or unconfirmed cleanup. Starting a new run at capacity evicts the oldest completed +result whose process and worker cleanup are confirmed. Explicit `run.release` +discards the same retained result sooner. Active work returns `run_active`; +uncertain cleanup returns `cleanup_unconfirmed` and remains retained. Reading +results/events and requesting cancellation do not release the handle. + +Run handles and event pages are memory-owned, not a second vendor conversation +database. Replacing the adapter invalidates these handles; it does not delete +native saved conversations. Explicit release also leaves native persistence +unchanged. Save the native session ID before releasing a result if it is needed +for a later resume. + +## Ordinary MCP work resource + +Every tool declares `_meta["io.github.computer-mcp/work"]` with `format_version: 1` +and URI `computer-mcp://runtime/work/v1`. The same URI is available through +`resources/list` and `resources/read`; it follows the host's +[provider-work contract](https://github.com/computer-mcp/computer-mcp/blob/master/Documentation/Reference/MCPProtocol.md#downstream-provider-work). +It does not require private Host Services, change native permissions, or itself +enable host configuration changes. + +A host supplies `_meta["io.github.computer-mcp/work-invocation"]` as a UUID on +each tool call. Each newly reserved run retains its creation call's UUID. +Arguments cannot supply this identity. Subsequent result, event, cancel and +release calls do not rebind the original acquisition. Ordinary clients may +omit the metadata and still use every tool. While an unbound handle is retained, +work observation returns an error because it cannot prove a complete host-bound +snapshot. + +The resource returns a complete bounded snapshot with one instance UUID and a +monotonic revision, containing `kind: claude.run`, the adapter `run_id`, +`acquired_by`, and `state`. Pending version checks, process startup, streaming +execution and retained completed results are `active` owners. A cleanup failure +is `uncertain`, even when `completed` is true. Confirmed completion alone does +not remove retained result access: release or bounded capacity eviction ends +the handle's ownership. Snapshot errors preserve the last valid revision. +The declaration is connection-local; gateway reexports omit it from their wire +tool metadata. ## Session and permission semantics @@ -33,7 +83,7 @@ Model, effort, appended system prompt and optional budget are forwarded only whe Each native frame is bounded to 1 MiB before waiting for a newline. The event buffer retains at most 256 events/256 KiB; it reports overflow, omitted oversized events and stale cursors. Pages expose next_cursor, has_more and missed_events, never an unbounded whole history. The final native event is retained separately with a 128 KiB bound; oversized final data fails rather than masquerading as successful truncated output. -Success requires exactly one valid native result event, a non-error native outcome, a successful process exit, and confirmed process cleanup. A final event followed by a hung process still times out. Missing/duplicate/malformed final events, a nonzero exit, cancellation, timeout or cleanup uncertainty are tool errors. Native errors and the final event remain available in the bounded result. `completed: true` means this adapter run settled, not that its task succeeded; inspect `is_error` and state. +Success requires exactly one valid native result event, a non-error native outcome, a successful process exit, and confirmed process cleanup. A final event followed by a hung process still times out. Missing/duplicate/malformed final events, a nonzero exit, cancellation, timeout or cleanup uncertainty are tool errors. Native errors and the first valid final event remain available in the bounded result, including after later malformed output, timeout, cancellation or cleanup failure. `cleanup_error` records cleanup failure separately from an earlier execution `error`; `cleanup_confirmed` reports the process cleanup outcome. Successful final events require a `success` subtype, native session identity and textual result; unknown extension fields are retained. `completed: true` means this adapter run settled, not that its task succeeded; inspect `is_error` and state. The default execution timeout is 600 seconds, with a configurable maximum of 1800 seconds. Process setup/cleanup can add bounded latency. Synchronous MCP cancellation targets the corresponding run. For detached runs use run.cancel with the returned run_id. Cancellation kills only the owned vendor process group; it does not erase a persisted conversation or retry the prompt. @@ -46,3 +96,16 @@ A private supervisor observes parent/connection shutdown, pins the native group - The installed native `claude --version` and `claude --help` used to maintain the pinned CLI tree. Fixture protocol checks, native interface checks, exact-host interoperability and authenticated model execution are recorded separately. + +## Continuation binding + +Tools that accept an existing adapter handle declare +`_meta["io.github.computer-mcp/continuation"]` with format version 1. The selector +matches kind `claude.run` and primary resource `id` against argument +`run_id` using JSON Pointer `/run_id`. This identifies the actual +connection-owned lifetime; it does not rebind acquisition or grant permissions. +New work and unscoped listings do not claim an existing owner. The declaration +uses ordinary MCP metadata and requires no private Host Services. Hosts validate +and retain it on its originating connection; gateway reexports strip it. Runtime +generation selection remains host-owned, and this declaration alone does not +enable live configuration changes. diff --git a/README.md b/README.md index 4db0cde..3118448 100644 --- a/README.md +++ b/README.md @@ -6,7 +6,12 @@ The CLI contribution describes verified non-interactive commands. The native bas ## MCP capabilities -The adapter provides `claude.run`, `claude.run.start`, `claude.run.list`, `claude.run.result`, `claude.run.events` and `claude.run.cancel`. It consumes native print-mode stream-json, including partial messages, and retains the final result, native session ID and explicit error status. Discovery does not launch Claude Code. +The adapter provides `claude.run`, `claude.run.start`, `claude.run.list`, `claude.run.result`, `claude.run.events`, `claude.run.cancel` and `claude.run.release`. It consumes native print-mode stream-json, including partial messages, and retains the final result, native session ID and explicit error status. Release discards a completed result only after cleanup is confirmed. Discovery does not launch Claude Code. + +The ordinary MCP work resource reports active runs and retained result handles +with their original acquisition reference. Hosts can account for this work +after its creating tool returns, without private Host Services permission. +Cleanup uncertainty remains owned and cannot be released or evicted. `permission_mode` defaults to `dontAsk`; `plan` and `acceptEdits` are supported explicit choices. The noninteractive contract does not expose permission-bypass modes or fabricate interactive permission responses. Existing vendor configuration continues to apply. Resume/continue/new-session selection is explicit; cancelling a run does not delete its saved native conversation. diff --git a/Scripts/validate_host.py b/Scripts/validate_host.py index ce4ad19..e3ddedf 100644 --- a/Scripts/validate_host.py +++ b/Scripts/validate_host.py @@ -108,7 +108,7 @@ def configuration(package, workspace, cli_fixture, acp_fixture, vendor, readonly ''' -def validate(host, archive, output): +def validate(host, archive, output, expected_host_version, require_work_ownership=False): host = host.resolve(strict=True) archive = archive.resolve(strict=True) output.mkdir(parents=True, exist_ok=False) @@ -126,7 +126,8 @@ def validate(host, archive, output): environment = {'PATH': os.pathsep.join([str(Path(sys.executable).parent), '/usr/bin', '/bin', '/usr/sbin', '/sbin']), 'HOME': str(home), 'TMPDIR': str(work), 'PYTHONDONTWRITEBYTECODE': '1', 'LANG': 'en_US.UTF-8'} version = capture([str(host), '--version'], work, environment).decode().strip() - require(version.startswith('1.2.2 '), 'This validator targets Computer MCP 1.2.2; review another host before use') + require(bool(expected_host_version) and version.startswith(expected_host_version + ' '), + 'Host version differs from the explicitly selected acceptance version') inputs = work / 'inputs' inputs.mkdir() (inputs / f'{vendor}.zip').write_bytes(archive.read_bytes()) @@ -164,8 +165,21 @@ def validate(host, archive, output): projected = [tool for tool in catalog if tool.get('_meta', {}).get('cli', {}).get('command') == 'print'] require(len(projected) == 1, 'Expected one real host-projected print tool') native_names = [tool['name'] for tool in catalog if tool['name'].startswith(vendor + '.')] - require(len(native_names) == (12 if vendor == 'cursor' else 6), 'Adapter catalog missing required tools') + require(len(native_names) == (12 if vendor == 'cursor' else 7), 'Adapter catalog missing required tools') checks['catalog'] = {'status':'passed', 'cli_tools':len(projected), 'mcp_tools':len(native_names)} + if require_work_ownership: + require(vendor == 'claude', 'Work ownership acceptance requires the Claude run contract') + require(all(key not in tool.get('_meta',{}) for tool in catalog + for key in ('io.github.computer-mcp/work','io.github.computer-mcp/continuation')), + 'Gateway exports advertise downstream-only ownership metadata') + def work_status(count): + def observe(): + servers = checked(client,'mcp.servers.status',{'server':'fixture-adapter'})['servers'] + value = servers[0]['connection'].get('provider_work') + require(isinstance(value,dict), 'Candidate host does not expose provider-work observation') + return value + return wait_for(observe, lambda value:value['resource_count']==count + and value['unsettled_invocation_count']==0 and not value['observation_pending']) prompt = "--leading 'quotes' 中文\nnot-a-shell-command" arguments = {'prompt': prompt} if vendor == 'claude': @@ -197,10 +211,30 @@ def validate(host, archive, output): checked(client, 'cursor.acp.session.close', {'session':session}) require(not checked(client, 'cursor.acp.session.list')['sessions'], 'Closed ACP session remains live') else: + work_evidence = None + if require_work_ownership: + active = checked(client, 'claude.run.start', {'prompt':'slow'})['run_id'] + work_evidence = {'running':work_status(1)} + refusal = client.call('claude.run.release', {'run_id':active}) + require(Client.value(refusal).get('error',{}).get('code')=='run_active', 'Active run was released') + checked(client, 'claude.run.cancel', {'run_id':active}) + cancelled = wait_for(lambda:Client.value(client.call('claude.run.result', {'run_id':active})), lambda v:v.get('completed')) + require(cancelled['cleanup_confirmed'] and cancelled['state']=='cancelled', 'Cancellation cleanup was not confirmed') + work_evidence['cancelled_result_retained'] = work_status(1) + checked(client, 'claude.run.release', {'run_id':active}) + work_evidence['cancelled_result_released'] = work_status(0) run = checked(client, 'claude.run.start', {'prompt':'hello','permission_mode':'plan'})['run_id'] completed = wait_for(lambda: checked(client, 'claude.run.result', {'run_id':run}), lambda v:v.get('completed')) require(completed['result'] == 'hello', 'Native final result was not preserved') checked(client, 'claude.run.events', {'run_id':run,'max_bytes':2048}) + if work_evidence is not None: + work_evidence['completed_result_retained'] = work_status(1) + checked(client, 'claude.run.release', {'run_id':run}) + require(not checked(client, 'claude.run.list')['runs'], 'Released run result remains retained') + if work_evidence is not None: + work_evidence['released'] = work_status(0) + require(len({value['instance_id'] for value in work_evidence.values()})==1, 'Ownership crossed provider instances') + checks['provider_work'] = work_evidence invalid = client.call('claude.run', {'prompt':'not-executed','permission_mode':'bypassPermissions'}) require(invalid['result'].get('isError'), 'Bypass mode was admitted') checks['mcp_execution_and_events'] = 'passed' @@ -216,9 +250,30 @@ def validate(host, archive, output): finally: client.close() checks['read_only_profile'] = 'passed' + restricted = configuration(package, workspace, cli_fixture, ROOT / 'Tests/Fixtures/vendor.py', vendor) + restricted = restricted.replace('mode = "local-full-access"', 'mode = "workspace-operations"') + restricted = restricted.replace('full_shell_enabled = true', 'full_shell_enabled = false') + execution_tool = 'cursor.acp.prompt' if vendor == 'cursor' else 'claude.run' + # An explicit low host risk must not bypass the publisher's execution floor. + restricted += '\n[mcp.servers.tool_risks]\n' + json.dumps(execution_tool) + ' = "read-only"\n' + config.write_text(restricted) + client = Client(str(workspace), environment, [str(host), 'serve', 'stdio', '--config', str(config)]) + try: + tools = client.request('tools/list')['result']['tools'] + require(execution_tool not in {tool['name'] for tool in tools}, 'Restricted profile exposed arbitrary vendor execution') + inspection = 'cursor.acp.session.list' if vendor == 'cursor' else 'claude.run.list' + checked(client, inspection) + for name, arguments in [(execution_tool, {'prompt':'must-not-execute'}), + ('mcp.tools.call', {'server':'fixture-adapter','tool':execution_tool,'arguments':{'prompt':'must-not-execute'}})]: + denied = client.call(name, arguments) + require('error' in denied or denied['result'].get('isError'), 'Restricted profile admitted vendor execution') + finally: + client.close() + checks['publisher_floor_under_restricted_profile'] = 'passed' require(digest(host) == host_hash and digest(archive) == archive_hash, 'Host or archive changed during acceptance') report = {'status':'passed', 'observed_at':datetime.datetime.now(datetime.timezone.utc).isoformat(), 'plugin_id':vendor, 'plugin_version':manifest['version'], 'host_version':version, + 'expected_host_version':expected_host_version, 'host_sha256':host_hash, 'archive_sha256':archive_hash, 'checks':checks, 'scope':'host archive validation and ordinary standalone registrations with inert vendor fixtures', 'production_installation':False, 'authenticated_model_execution':False, @@ -230,10 +285,14 @@ def validate(host, archive, output): def main(): parser = argparse.ArgumentParser(description=__doc__) parser.add_argument('--host', required=True, type=Path) + parser.add_argument('--expected-host-version', required=True, + help='Exact reviewed host release version, for example 1.3.0') parser.add_argument('--archive', required=True, type=Path) parser.add_argument('--output', required=True, type=Path) + parser.add_argument('--require-work-ownership', action='store_true', + help='Require a candidate host to observe active and retained run ownership through release') args = parser.parse_args() - validate(args.host, args.archive, args.output) + validate(args.host, args.archive, args.output, args.expected_host_version, args.require_work_ownership) if __name__ == '__main__': diff --git a/Tests/support.py b/Tests/support.py index 63c2aae..2b16b5f 100644 --- a/Tests/support.py +++ b/Tests/support.py @@ -82,6 +82,8 @@ def wait_file(path, seconds=8): def process_alive(pid): r=subprocess.run(['/bin/ps','-p',str(pid),'-o','stat='],capture_output=True,text=True,timeout=1) + if r.stderr.strip() or r.returncode not in (0,1): + raise AssertionError('Cannot verify owned process state: '+r.stderr.strip()) return bool(r.stdout.strip()) and not r.stdout.strip().startswith('Z') diff --git a/Tests/test_adapter.py b/Tests/test_adapter.py index 382385e..4a3c89d 100644 --- a/Tests/test_adapter.py +++ b/Tests/test_adapter.py @@ -13,6 +13,14 @@ def value(self,response): self.assertFalse(response['result']['isError'],response) return Client.value(response) def call(self,name,args=None,timeout=10):return self.client.call('claude.'+name,args,timeout) + def test_catalog_declares_execution_and_inspection_risk(self): + expected = {'claude.run': 'full-shell', 'claude.run.start': 'full-shell', 'claude.run.list': 'read-only', 'claude.run.result': 'read-only', 'claude.run.events': 'read-only', 'claude.run.cancel': 'destructive', 'claude.run.release': 'destructive'} + tools = self.client.request('tools/list')['result']['tools'] + self.assertEqual({tool['name']:tool['_meta']['io.github.computer-mcp/risk'] for tool in tools},expected) + for tool in tools: + risk = expected[tool['name']] + self.assertEqual(tool['annotations']['readOnlyHint'],risk=='read-only') + self.assertEqual(tool['annotations']['destructiveHint'],risk in {'destructive','full-shell'}) def test_stream_json_through_mcp(self): r=self.value(self.call('run',{'prompt':'hello','permission_mode':'plan'})) self.assertEqual(r['result'],'hello');self.assertEqual(r['permission_mode'],'plan') diff --git a/Tests/test_lifecycle.py b/Tests/test_lifecycle.py index 9fb1f6b..b810aa6 100644 --- a/Tests/test_lifecycle.py +++ b/Tests/test_lifecycle.py @@ -31,6 +31,16 @@ def scenario(self,termination): def test_eof_retires_vendor_and_descendant(self):self.scenario('eof') def test_sigterm_retires_vendor_and_descendant(self):self.scenario(signal.SIGTERM) def test_sigkill_parent_loss_is_detected_by_supervisor(self):self.scenario(signal.SIGKILL) + def test_exited_leader_remains_owned_until_descendants_are_retired(self): + with tempfile.TemporaryDirectory() as temp: + marker=Path(temp)/'pids.json' + client=Client(temp,{'FIXTURE_MARKER':str(marker)}) + try: + response=client.call('claude.run',{'prompt':'earlychild'}) + self.assertFalse(response['result']['isError'],response) + wait_stopped(wait_file(marker)) + finally:client.close() + def test_host_context_is_not_forwarded_to_vendor(self): with tempfile.TemporaryDirectory() as temp: client=Client(temp,{'COMPUTER_MCP_TEST_PRIVATE':'must-not-inherit'}) diff --git a/Tests/test_protocol_admission.py b/Tests/test_protocol_admission.py new file mode 100644 index 0000000..29c7dc7 --- /dev/null +++ b/Tests/test_protocol_admission.py @@ -0,0 +1,76 @@ +"""Malformed peer input must not terminate unrelated MCP work.""" +import json +import sys +import threading +import unittest +from unittest.mock import patch +from support import ROOT +sys.path.insert(0,str(ROOT/'bin')) +import plugin_runtime as runtime + +class ProtocolAdmissionTests(unittest.TestCase): + def test_invalid_unicode_is_rejected_in_all_json_string_positions(self): + for raw in (b'{"id":"\\ud800"}',b'{"\\udfff":1}',b'["\\ud800"]'): + with self.subTest(raw=raw),self.assertRaises(ValueError): + runtime.decoded(raw) + self.assertEqual(runtime.decoded(b'"\\ud83d\\ude00"'),'\U0001f600') + def test_bad_id_does_not_end_active_tool_or_connection(self): + entered,release=threading.Event(),threading.Event() + def handler(_args,job): + entered.set() + if not release.wait(2): raise AssertionError('fixture not released') + job.check() + return {'finished':True} + server=runtime.MCPServer('fixture','1',[runtime.Tool('hold','fixture',runtime.schema({}),handler,risk='read-only')],lambda:None) + server.initialized=True + messages=[] + def write(_fd,data,*_): messages.append(json.loads(data)) + with patch.object(runtime,'write_bytes',side_effect=write): + server.dispatch({'jsonrpc':'2.0','id':1,'method':'tools/call','params':{'name':'hold'}}) + try: + self.assertTrue(entered.wait(1)) + server.dispatch({'jsonrpc':'2.0','id':'\ud800','method':'ping'}) + self.assertEqual(messages[-1]['error']['code'],-32600) + self.assertIsNone(messages[-1]['id']) + server.dispatch({'jsonrpc':'2.0','id':2,'method':'ping'}) + self.assertEqual(messages[-1]['result'],{}) + self.assertFalse(server.stop.is_set()) + finally: + with server.lock: threads=list(server.threads) + release.set() + for thread in threads: thread.join(2) + completion=next(m for m in messages if m['id']==1) + self.assertFalse(completion['result']['isError']) + def test_active_request_id_cannot_be_reused_by_ping(self): + server=runtime.MCPServer('fixture','1',[],lambda:None) + server.jobs[server.id_key(5)]=runtime.Job() + with patch.object(server,'emit') as emit: + server.dispatch({'jsonrpc':'2.0','id':5,'method':'ping'}) + self.assertIn('error',emit.call_args.args[0]) + self.assertIn(server.id_key(5),server.jobs) + + def test_denied_process_inspection_cannot_confirm_cleanup(self): + child=unittest.mock.Mock(pid=42,returncode=0) + writes=[] + report=unittest.mock.Mock(returncode=1,stdout='',stderr='permission denied') + with patch.object(runtime.subprocess,'Popen',return_value=child), patch.object(runtime.os,'waitid',return_value=object()), patch.object(runtime.os,'killpg',side_effect=PermissionError()), patch.object(runtime.subprocess,'run',return_value=report), patch.object(runtime.os,'write',side_effect=lambda fd,data:writes.append((fd,data))), patch.object(runtime.os,'close'), patch.object(runtime.signal,'signal'): + result=runtime.supervise(10,11,['/unused']) + self.assertEqual(result,70) + receipt=runtime.decoded(next(data for fd,data in writes if fd==11)) + self.assertFalse(receipt['cleanup_confirmed']) + + def test_protocol_negotiation_uses_supported_dates_and_typed_fields(self): + for date in (*runtime.SUPPORTED_MCP,'unsupported'): + server=runtime.MCPServer('fixture','1',[],lambda:None) + with patch.object(server,'emit') as emit: + server.dispatch({'jsonrpc':'2.0','id':1,'method':'initialize','params':{'protocolVersion':date,'capabilities':{},'clientInfo':{'name':'test','version':'1'}}}) + self.assertIn(emit.call_args.args[0]['result']['protocolVersion'],runtime.SUPPORTED_MCP) + for fields in ({'protocolVersion':True},{'protocolVersion':'2025-06-18','capabilities':[]}, + {'protocolVersion':'2025-06-18','capabilities':{},'clientInfo':{'name':4,'version':'1'}}): + server=runtime.MCPServer('fixture','1',[],lambda:None) + with patch.object(server,'emit') as emit: + server.dispatch({'jsonrpc':'2.0','id':1,'method':'initialize','params':fields}) + self.assertIn('error',emit.call_args.args[0]) + self.assertFalse(server.initialized) + +if __name__=='__main__': unittest.main() diff --git a/Tests/test_review_regressions.py b/Tests/test_review_regressions.py new file mode 100644 index 0000000..630fdb1 --- /dev/null +++ b/Tests/test_review_regressions.py @@ -0,0 +1,86 @@ +"""Final result ownership when later stream or cleanup work fails.""" +import unittest +from unittest.mock import patch +from test_completion import adapter +from plugin_runtime import Failure, Job, encoded + +FINAL = {'type':'result','subtype':'success','is_error':False, + 'session_id':'native-session','result':'kept','extension':{'unknown':True}} + +class Peer: + def __init__(self, frames, finish_error=None, cleanup_error=None): + self.frames=iter(frames) + self.finish_error=finish_error + self.cleanup_error=cleanup_error + def end_input(self): pass + def read(self, *_): + item=next(self.frames,None) + if isinstance(item,Exception): raise item + return item if isinstance(item,bytes) or item is None else encoded(item) + def finish(self, *_): + if self.finish_error: raise self.finish_error + return 0 + def close(self): + if self.cleanup_error: raise self.cleanup_error + +class ReviewRegressions(unittest.TestCase): + def execute(self, peer): + run=adapter.Run('owned-run',{'prompt':'fixture'}) + with patch.object(adapter,'verify_version'),patch.object(adapter,'Process',return_value=peer): + run.work('/unused') + return run.result() + def test_cancel_during_admission_never_reserves_native_work(self): + claude=adapter.Claude('/unused');job=Job() + def resolve(*_): + job.cancel() + return '/unused' + with patch.object(adapter,'resolve_executable',side_effect=resolve),patch.object(adapter.Run,'work') as work: + with self.assertRaises(Failure) as error:claude.start({'prompt':'fixture'},job) + self.assertEqual(error.exception.code,'cancelled') + self.assertEqual(claude.runs,{}) + work.assert_not_called() + + def test_final_survives_later_failure_and_event_eviction(self): + overflow=[{'type':'assistant','text':str(i)} for i in range(300)] + for error in ('timeout','cancelled'): + with self.subTest(error=error): + result=self.execute(Peer([FINAL,*overflow,Failure(error,'fixture')])) + self.assertTrue(result['completed']) + self.assertTrue(result['is_error']) + self.assertEqual(result['error']['code'],error) + self.assertEqual(result.get('final_event'),FINAL) + self.assertEqual(result.get('result'),'kept') + def test_final_survives_malformed_frame_duplicate_and_cleanup_failure(self): + peers=[Peer([FINAL,b'not-json']),Peer([FINAL,FINAL]), + Peer([FINAL],finish_error=Failure('timeout','fixture')), + Peer([FINAL],cleanup_error=Failure('cleanup_unconfirmed','fixture'))] + for peer in peers: + with self.subTest(peer=peer): + result=self.execute(peer) + self.assertTrue(result['is_error']) + self.assertEqual(result.get('final_event'),FINAL) + def test_invalid_final_cannot_claim_success(self): + invalid=[{'type':'result','is_error':False}, + dict(FINAL,session_id=None),dict(FINAL,subtype=4), + dict(FINAL,subtype='error_max_turns',is_error=False), + dict(FINAL,result={'not':'text'})] + for final in invalid: + with self.subTest(final=final): + result=self.execute(Peer([final])) + self.assertTrue(result['is_error']) + self.assertEqual(result['error']['code'],'invalid_vendor_response') + def test_native_failure_and_unknown_fields_are_retained(self): + final=dict(FINAL,subtype='error_during_execution',is_error=True, + errors=['native failure']) + final.pop('result') + result=self.execute(Peer([final])) + self.assertTrue(result['is_error']) + self.assertEqual(result['final_event'],final) + def test_cleanup_error_preserves_primary_execution_error(self): + result=self.execute(Peer([FINAL,Failure('timeout','fixture timeout')], + cleanup_error=Failure('cleanup_unconfirmed','fixture cleanup'))) + self.assertEqual(result['error']['code'],'timeout') + self.assertEqual(result.get('cleanup_error',{}).get('code'),'cleanup_unconfirmed') + self.assertEqual(result.get('final_event'),FINAL) + +if __name__=='__main__': unittest.main() diff --git a/Tests/test_work_resources.py b/Tests/test_work_resources.py new file mode 100644 index 0000000..9e1b82c --- /dev/null +++ b/Tests/test_work_resources.py @@ -0,0 +1,269 @@ +"""Run retention must not discard native cleanup ownership.""" +import json +import os +import tempfile +import threading +import time +import unittest +import uuid +from unittest.mock import Mock, patch +from test_completion import adapter +from test_review_regressions import Peer, FINAL +from support import Client +from plugin_runtime import CONTINUATION_METADATA, Failure, Job, MCPServer, Process, WORK_INVOCATION, WORK_METADATA, WORK_URI + + +class WorkResourceTests(unittest.TestCase): + def call(self,client,name,arguments=None,origin=None): + params = {'name':name,'arguments':arguments or {}} + if origin is not None: params['_meta'] = {WORK_INVOCATION:origin} + return client.request('tools/call',params) + + def snapshot(self,client): + result = client.request('resources/read',{'uri':WORK_URI}) + self.assertNotIn('error',result) + content = result['result']['contents'] + self.assertEqual(len(content),1) + self.assertEqual(content[0]['uri'],WORK_URI) + return json.loads(content[0]['text']) + + def settled(self,client,run): + deadline = time.monotonic()+5 + while time.monotonic()131072: raise Failure('result_too_large','Claude final result exceeds 128 KiB; no success was inferred from truncated data') - if type(event.get('is_error')) is not bool: - raise Failure('invalid_vendor_response','Claude result lacks a Boolean is_error field') - final = event + final = checked_final(event) + with self.lock: self.final_event = final code = process.finish(deadline,self.job) if final is None: raise Failure('vendor_failed',f'Claude exited with status {code} without a final result event') @@ -106,26 +143,29 @@ class Run: with self.lock: self.state = 'cancelled' if error.code=='cancelled' else 'failed' self.value = {'run_id':self.id,'session_id':self.session_id,'is_error':True,'error':{'code':error.code,'message':str(error)}} + if error.code=='cleanup_unconfirmed': self.record_cleanup_failure(error) except Exception: with self.lock: self.state = 'failed' self.value = {'run_id':self.id,'session_id':self.session_id,'is_error':True,'error':{'code':'internal_error','message':'Claude execution failed; no automatic retry was attempted'}} finally: - if process: - try: process.close() - except Failure as error: - with self.lock: - self.state = 'failed' - self.value = {'run_id':self.id,'session_id':self.session_id,'is_error':True,'error':{'code':error.code,'message':str(error)}} - # A completed result is visible only after owned process cleanup. - self.arguments = {} - self.done.set() + for owned in (process,self.version_process): + if owned is not None: + try: owned.close() + except Exception as error: self.record_cleanup_failure(error) + with self.lock: + self.cleanup_confirmed = self.cleanup_failure is None + self.arguments = {} + self.done.set() def result(self,limit=128): with self.lock: if not self.done.is_set(): return self.info() value = dict(self.value) + if self.final_event is not None: + value.update(final_event=self.final_event,result=self.final_event.get('result'), + subtype=self.final_event['subtype'],total_cost_usd=self.final_event.get('total_cost_usd')) page = self.events.page(0,limit) - return {**value,'state':self.state,'completed':True,'events':[row['event'] for row in page['events']],'event_page':{k:v for k,v in page.items() if k!='events'},'events_truncated':page['has_more'] or page['missed_events'] or page['omitted_events']>0} + return {**value,'state':self.state,'completed':True,'cleanup_confirmed':self.cleanup_confirmed,'events':[row['event'] for row in page['events']],'event_page':{k:v for k,v in page.items() if k!='events'},'events_truncated':page['has_more'] or page['missed_events'] or page['omitted_events']>0} class Claude: @@ -140,22 +180,26 @@ class Claude: validate(args,RUN_SCHEMA) executable = resolve_executable(self.executable,'CLAUDE_CODE_EXECUTABLE','claude') build_command(executable,args) - run = Run(uuid.uuid4().hex,dict(args)) + run = Run(uuid.uuid4().hex,dict(args),_job.work_invocation) with self.lock: - if self.stopping or sum(not x.done.is_set() for x in self.runs.values())>=4: + _job.check() + if self.stopping or sum(not x.execution_finished() for x in self.runs.values())>=4: raise Failure('capacity','Concurrent Claude run capacity reached or adapter is shutting down') while len(self.runs)>=32: - old = next((key for key,value in self.runs.items() if value.done.is_set()),None) + old = next((key for key,value in self.runs.items() if value.execution_finished()),None) if old is None: raise Failure('capacity','Run result capacity reached') del self.runs[old] self.runs[run.id] = run run.thread = threading.Thread(target=run.work,args=(executable,)) - run.thread.start() + try: run.thread.start() + except Exception: + if run.thread.ident is None: del self.runs[run.id] + raise Failure('startup_failed','Run worker could not start; no automatic retry was attempted') from None return run.info() def get(self,identifier): with self.lock: run = self.runs.get(identifier) if run is None: - raise Failure('unknown_run','Run belongs to another connection or its completed retention window has expired') + raise Failure('unknown_run','Run belongs to another connection or its result was released or evicted') return run def once(self,args,job): run = self.get(self.start(args,job)['run_id']) @@ -174,6 +218,19 @@ class Claude: def listing(self,_args,_job): with self.lock: runs = list(self.runs.values()) return {'runs':[r.info() for r in runs]} + def work_resources(self): + with self.lock: runs = list(self.runs.values()) + return [run.work_resource() for run in runs] + def release(self,args,_job): + run = self.get(args['run_id']) + if not run.done.is_set(): + raise Failure('run_active','Run is still active; cancel it and inspect its result before release') + if run.thread: run.thread.join(timeout=5) + if not run.execution_finished(): + raise Failure('cleanup_unconfirmed','Run cleanup has not been confirmed; its result remains retained') + with self.lock: + if self.runs.get(run.id) is run: del self.runs[run.id] + return {'run_id':run.id,'released':True} def shutdown(self): with self.lock: self.stopping = True @@ -181,15 +238,20 @@ class Claude: for run in runs: run.job.cancel() for run in runs: if run.thread: run.thread.join(timeout=5) + with self.lock: + self.runs = OrderedDict((key,run) for key,run in self.runs.items() if not run.execution_finished()) + if self.runs: + raise Failure('cleanup_unconfirmed','Claude work has not confirmed cleanup') def tools(self): identifier = text_schema(64) return [ - Tool('claude.run','Run Claude Code print mode to completion; preserve its final result, error flag, and bounded events.',RUN_SCHEMA,self.once), - Tool('claude.run.start','Start a Claude Code request; returns an adapter-local run ID immediately. Native session IDs remain separate.',RUN_SCHEMA,self.start), - Tool('claude.run.list','List running and retained completed requests belonging to this adapter connection.',schema({}),self.listing), - Tool('claude.run.result','Read completion status and result without restarting or replaying the native command.',schema({'run_id':identifier,'max_events':integer_schema(1,256,128)},['run_id']),self.result), - Tool('claude.run.events','Read byte-bounded native stream-json events using a run-local cursor.',schema({'run_id':identifier,'after_cursor':integer_schema(0,2**53-1,0),'limit':integer_schema(1,256,128),'max_bytes':integer_schema(1024,196608,196608)},['run_id']),self.events), - Tool('claude.run.cancel','Cancel the exact owned Claude run; subsequent result inspection confirms termination. No native session is deleted.',schema({'run_id':identifier},['run_id']),self.cancel), + Tool('claude.run','Run Claude Code print mode to completion; preserve its final result, error flag, and bounded events.',RUN_SCHEMA,self.once,risk='full-shell'), + Tool('claude.run.start','Start a Claude Code request; returns an adapter-local run ID immediately. Native session IDs remain separate.',RUN_SCHEMA,self.start,risk='full-shell'), + Tool('claude.run.list','List running and retained completed requests belonging to this adapter connection.',schema({}),self.listing,risk='read-only'), + Tool('claude.run.result','Read completion status and result without restarting or replaying the native command.',schema({'run_id':identifier,'max_events':integer_schema(1,256,128)},['run_id']),self.result,risk='read-only',continuation=('claude.run','run_id')), + Tool('claude.run.events','Read byte-bounded native stream-json events using a run-local cursor.',schema({'run_id':identifier,'after_cursor':integer_schema(0,2**53-1,0),'limit':integer_schema(1,256,128),'max_bytes':integer_schema(1024,196608,196608)},['run_id']),self.events,risk='read-only',continuation=('claude.run','run_id')), + Tool('claude.run.cancel','Cancel the exact owned Claude run; subsequent result inspection confirms termination. No native session is deleted.',schema({'run_id':identifier},['run_id']),self.cancel,risk='destructive',continuation=('claude.run','run_id')), + Tool('claude.run.release','Release the retained result and events of a completed, confirmed-cleaned run. Active or uncertain work is retained; native conversations are unchanged.',schema({'run_id':identifier},['run_id']),self.release,risk='destructive',continuation=('claude.run','run_id')), ] @@ -199,7 +261,7 @@ def main(): parser.add_argument('--executable',help='Absolute user-owned Claude Code path; otherwise CLAUDE_CODE_EXECUTABLE or PATH') args = parser.parse_args() claude = Claude(args.executable) - MCPServer('claude-mcp-adapter',VERSION,claude.tools(),claude.shutdown).serve() + MCPServer('claude-mcp-adapter',VERSION,claude.tools(),claude.shutdown,work=claude.work_resources).serve() if __name__=='__main__': main() diff --git a/bin/plugin_runtime.py b/bin/plugin_runtime.py index 0c64892..33e056e 100644 --- a/bin/plugin_runtime.py +++ b/bin/plugin_runtime.py @@ -16,10 +16,15 @@ import threading import time import tomllib +import uuid MAX_FRAME = 1_048_576 MAX_RESULT = 196_608 SUPPORTED_MCP = ('2024-11-05', '2025-03-26', '2025-06-18') +WORK_URI = 'computer-mcp://runtime/work/v1' +WORK_METADATA = 'io.github.computer-mcp/work' +CONTINUATION_METADATA = 'io.github.computer-mcp/continuation' +WORK_INVOCATION = 'io.github.computer-mcp/work-invocation' class Failure(Exception): @@ -47,7 +52,16 @@ def finite_float(text): if not math.isfinite(value): raise ValueError('JSON floating-point number is not finite') return value - return json.loads(raw.decode('utf-8'), object_pairs_hook=unique, parse_constant=invalid_number, parse_float=finite_float) + value = json.loads(raw.decode('utf-8'), object_pairs_hook=unique, parse_constant=invalid_number, parse_float=finite_float) + pending = [value] + while pending: + item = pending.pop() + if isinstance(item,str): item.encode('utf-8') + elif isinstance(item,dict): + pending.extend(item.keys()) + pending.extend(item.values()) + elif isinstance(item,list): pending.extend(item) + return value def schema(properties, required=()): @@ -107,8 +121,9 @@ def validate(value, spec, location='arguments', depth=0): class Job: - def __init__(self): + def __init__(self, work_invocation=None): self.cancelled = threading.Event() + self.work_invocation = work_invocation def cancel(self): self.cancelled.set() def check(self): @@ -237,7 +252,7 @@ def supervise(read_fd, receipt_fd, command): ['/bin/ps', '-g', str(child.pid), '-o', 'pid=,stat='], capture_output=True, text=True, timeout=1) rows = [line.split() for line in report.stdout.splitlines() if line.strip()] - if report.returncode not in (0, 1) or any( + if report.stderr.strip() or report.returncode not in (0, 1) or any( len(row) != 2 or not row[1].startswith('Z') for row in rows ): raise Failure('cleanup_unconfirmed', 'Owned group could not be retired') @@ -285,10 +300,15 @@ def __init__(self, command, cwd): self.stderr = bytearray() self.stderr_truncated = False self.stop_stderr = threading.Event() - os.set_blocking(self.child.stdin.fileno(), False) - self.output = LineReader(self.child.stdout) - self.stderr_thread = threading.Thread(target=self._stderr, daemon=True) - self.stderr_thread.start() + self.output, self.stderr_thread = None, None + try: + os.set_blocking(self.child.stdin.fileno(), False) + self.output = LineReader(self.child.stdout) + self.stderr_thread = threading.Thread(target=self._stderr, daemon=True) + self.stderr_thread.start() + except BaseException: + self.close() + raise def _stderr(self): try: with selectors.DefaultSelector() as selector: @@ -346,30 +366,33 @@ def close(self): raise self.cleanup_failure return self.closed = True - os.close(self.life) try: try: - self.child.wait(timeout=3) - except subprocess.TimeoutExpired: - self.child.terminate() + os.close(self.life) try: - self.child.wait(timeout=2) - except subprocess.TimeoutExpired as error: - self.child.kill() - self.child.wait(timeout=2) - raise Failure('cleanup_unconfirmed', 'Vendor supervisor did not confirm shutdown') from error - self.confirm_cleanup(self.receipt) - except Failure as error: - self.cleanup_failure = error - raise - finally: - os.close(self.receipt) - self.output.close() - self.stop_stderr.set() - self.stderr_thread.join(timeout=.5) - with self.writer_lock: - for stream in (self.child.stdin, self.child.stdout, self.child.stderr): - stream.close() + self.child.wait(timeout=3) + except subprocess.TimeoutExpired: + self.child.terminate() + try: + self.child.wait(timeout=2) + except subprocess.TimeoutExpired as error: + self.child.kill() + self.child.wait(timeout=2) + raise Failure('cleanup_unconfirmed', 'Vendor supervisor did not confirm shutdown') from error + self.confirm_cleanup(self.receipt) + finally: + os.close(self.receipt) + if self.output is not None: self.output.close() + self.stop_stderr.set() + if self.stderr_thread is not None and self.stderr_thread.ident is not None: + self.stderr_thread.join(timeout=.5) + with self.writer_lock: + for stream in (self.child.stdin, self.child.stdout, self.child.stderr): + stream.close() + except Exception as error: + self.cleanup_failure = error if isinstance(error,Failure) else Failure('cleanup_unconfirmed','Owned process cleanup could not be confirmed') + if self.cleanup_failure is error: raise + raise self.cleanup_failure from error def resolve_executable(explicit, environment_key, default): @@ -391,7 +414,7 @@ def load_package(entry): return manifest, tree -def verify_version(executable, tree, job, cwd): +def verify_version(executable, tree, job, cwd, retain_process=None): expected = next((x.get('stdout') for x in tree.get('executable_checks', []) if x.get('args') == ['--version']), None) if expected is None: raise Failure('invalid_configuration', 'Package lacks its native version assertion') @@ -399,6 +422,7 @@ def verify_version(executable, tree, job, cwd): deadline = time.monotonic()+5 output = bytearray() try: + if retain_process is not None: retain_process(process) process.end_input() while True: raw = process.read(deadline, job) @@ -459,12 +483,36 @@ class Tool: description: str input_schema: dict handler: object + risk: str + continuation: tuple[str,str] | None = None + def __post_init__(self): + if self.risk not in {'read-only','workspace-write','external-write','destructive','full-shell'}: + raise ValueError('Every tool requires a known publisher risk classification') + if self.continuation is not None: + kind, argument = self.continuation + if (not isinstance(kind,str) or not kind or len(kind.encode('utf-8'))>1024 + or any(ord(c)<32 or 127<=ord(c)<=159 for c in kind) + or not isinstance(argument,str) or not argument + or len(('/'+argument.replace('~','~0').replace('/','~1')).encode('utf-8'))>1024 + or any(ord(c)<32 or 127<=ord(c)<=159 for c in argument) + or argument not in self.input_schema.get('required',[]) + or self.input_schema.get('properties',{}).get(argument,{}).get('type') not in {'string','integer'}): + raise ValueError('Continuation requires a resource kind and required scalar handle') def definition(self): - return {'name':self.name,'description':self.description,'inputSchema':self.input_schema} + read_only = self.risk == 'read-only' + metadata = {'io.github.computer-mcp/risk':self.risk} + if self.continuation is not None: + kind, argument = self.continuation + pointer = '/' + argument.replace('~','~0').replace('/','~1') + metadata[CONTINUATION_METADATA] = {'format_version':1,'selectors':[{'kind':kind,'handles':{'id':pointer}}]} + return {'name':self.name,'description':self.description,'inputSchema':self.input_schema, + '_meta':metadata, + 'annotations':{'readOnlyHint':read_only,'destructiveHint':self.risk in {'destructive','full-shell'}, + 'idempotentHint':read_only,'openWorldHint':not read_only}} class MCPServer: - def __init__(self, name, version, tools, shutdown): + def __init__(self, name, version, tools, shutdown, work=None): self.name, self.version = name, version self.tools = {t.name:t for t in tools} self.shutdown = shutdown @@ -472,8 +520,54 @@ def __init__(self, name, version, tools, shutdown): self.jobs, self.threads = {}, set() self.lock, self.write_lock = threading.Lock(), threading.Lock() self.initialized = False + self.work = work + self.work_instance = str(uuid.uuid4()) + self.work_revision = 0 + self.work_last = None + def work_resource(self): + resources = self.work() + if not isinstance(resources,list) or len(resources)>1024: + raise Failure('work_unavailable','Complete work observation is unavailable') + keys = set() + rows = [] + def identifier(value): + return isinstance(value,str) and 0=2**63: + raise Failure('work_unavailable','Work observation revision is exhausted') + snapshot = {'format_version':1,'instance_id':self.work_instance,'revision':revision,'resources':rows} + body = encoded(snapshot) + if len(body)>524288: + raise Failure('work_unavailable','Complete work observation exceeds its byte bound') + self.work_revision, self.work_last = revision, rows + return {'contents':[{'uri':WORK_URI,'mimeType':'application/json','text':body.decode('utf-8')}]} def emit(self, message): - raw = encoded(message)+b'\n' + try: + raw = encoded(message)+b'\n' + except (ValueError,UnicodeError,RecursionError): + try: + self.id_key(message.get('id')) + identifier = message['id'] + except ValueError: + identifier = None + raw = encoded({'jsonrpc':'2.0','id':identifier,'error':{'code':-32603,'message':'Response cannot be encoded'}})+b'\n' if len(raw) > MAX_FRAME: raw = encoded({'jsonrpc':'2.0','id':message.get('id'),'error':{'code':-32603,'message':'Response exceeds the protocol bound'}})+b'\n' try: @@ -507,6 +601,7 @@ def call(self, request_id, key, tool, arguments, job): def id_key(value): if type(value) not in (int,str) or (isinstance(value,str) and (not value or len(value)>256)): raise ValueError('Invalid request ID') + if isinstance(value,str): value.encode('utf-8') return type(value).__name__, value def dispatch(self, message): if not isinstance(message, dict) or message.get('jsonrpc') != '2.0' or not isinstance(message.get('method'),str): @@ -524,16 +619,22 @@ def dispatch(self, message): try: key = self.id_key(request_id) except ValueError: self.rpc_error(None,-32600,'Invalid request ID'); return + with self.lock: active = key in self.jobs + if active: + self.rpc_error(request_id,-32600,'Request ID is already active'); return if not isinstance(params,dict): self.rpc_error(request_id,-32602,'Parameters must be an object'); return if method == 'initialize': if self.initialized: self.rpc_error(request_id,-32600,'Connection is already initialized'); return requested = params.get('protocolVersion') - if not isinstance(requested,str): - self.rpc_error(request_id,-32602,'protocolVersion is required'); return + client = params.get('clientInfo') + if not isinstance(requested,str) or not isinstance(params.get('capabilities'),dict) or not isinstance(client,dict) or not all(isinstance(client.get(k),str) and client[k] for k in ('name','version')): + self.rpc_error(request_id,-32602,'Initialize requires protocolVersion, capabilities and typed clientInfo'); return self.initialized = True - self.emit({'jsonrpc':'2.0','id':request_id,'result':{'protocolVersion':requested if requested in SUPPORTED_MCP else SUPPORTED_MCP[-1],'capabilities':{'tools':{'listChanged':False}},'serverInfo':{'name':self.name,'version':self.version}}}) + capabilities = {'tools':{'listChanged':False}} + if self.work is not None: capabilities['resources'] = {'subscribe':False,'listChanged':False} + self.emit({'jsonrpc':'2.0','id':request_id,'result':{'protocolVersion':requested if requested in SUPPORTED_MCP else SUPPORTED_MCP[-1],'capabilities':capabilities,'serverInfo':{'name':self.name,'version':self.version}}}) elif method == 'ping': self.emit({'jsonrpc':'2.0','id':request_id,'result':{}}) elif not self.initialized: @@ -541,18 +642,44 @@ def dispatch(self, message): elif method == 'tools/list': if params.get('cursor'): self.rpc_error(request_id,-32602,'This finite catalog has no continuation cursor'); return - self.emit({'jsonrpc':'2.0','id':request_id,'result':{'tools':[t.definition() for t in self.tools.values()]}}) + definitions = [t.definition() for t in self.tools.values()] + if self.work is not None: + for definition in definitions: + definition['_meta'][WORK_METADATA] = {'format_version':1,'uri':WORK_URI} + self.emit({'jsonrpc':'2.0','id':request_id,'result':{'tools':definitions}}) + elif method == 'resources/list' and self.work is not None: + if params.get('cursor'): + self.rpc_error(request_id,-32602,'This finite resource catalog has no continuation cursor'); return + self.emit({'jsonrpc':'2.0','id':request_id,'result':{'resources':[{'uri':WORK_URI,'name':'Runtime work','mimeType':'application/json','description':'Complete connection-owned live work and cleanup observation.'}]}}) + elif method == 'resources/read' and self.work is not None: + if params.get('uri') != WORK_URI: + self.rpc_error(request_id,-32602,'Unknown resource'); return + try: result = self.work_resource() + except Failure as error: + self.rpc_error(request_id,-32000,str(error)); return + except Exception: + self.rpc_error(request_id,-32000,'Complete work observation is unavailable'); return + self.emit({'jsonrpc':'2.0','id':request_id,'result':result}) elif method == 'tools/call': name = params.get('name') if not isinstance(name,str) or name not in self.tools: self.rpc_error(request_id,-32602,'Unknown tool'); return args = params.get('arguments', {}) + origin = None + if self.work is not None: + meta = params.get('_meta',{}) + if not isinstance(meta,dict): + self.rpc_error(request_id,-32602,'Tool metadata must be an object'); return + if WORK_INVOCATION in meta: + try: origin = str(uuid.UUID(meta[WORK_INVOCATION])) + except (ValueError,TypeError,AttributeError): + self.rpc_error(request_id,-32602,'Work invocation reference must be a UUID'); return with self.lock: if key in self.jobs: self.rpc_error(request_id,-32600,'Request ID is already active'); return if len(self.jobs)>=16: self.rpc_error(request_id,-32000,'Concurrent request capacity reached'); return - job = Job() + job = Job(work_invocation=origin) thread = threading.Thread(target=self.call,args=(request_id,key,self.tools[name],args,job)) self.jobs[key] = job self.threads.add(thread) diff --git a/computer-mcp-plugin.toml b/computer-mcp-plugin.toml index 72399ad..d27c6bf 100644 --- a/computer-mcp-plugin.toml +++ b/computer-mcp-plugin.toml @@ -1,6 +1,6 @@ id = "claude" name = "Claude Code" -version = "0.1.0" +version = "0.1.1" repository = "https://github.com/computer-mcp/plugin-claude" description = "Claude Code headless CLI automation plus a streaming CLI-to-MCP adapter." @@ -27,7 +27,7 @@ id = "stream" transport = "stdio" executable = { path = "bin/claude-mcp-adapter" } prefix = "claude" -capabilities = ["tools"] +capabilities = ["tools", "resources"] [[skills]] id = "claude-code" diff --git a/skills/claude-code/SKILL.md b/skills/claude-code/SKILL.md index b52ffbc..809c717 100644 --- a/skills/claude-code/SKILL.md +++ b/skills/claude-code/SKILL.md @@ -9,6 +9,12 @@ Discover actual host tool names/schemas and select an authorized workspace. Use For observable execution, use `claude.run.start`, retain run_id, page through `claude.run.events`, and inspect `claude.run.result`. Use `claude.run` only when waiting for the complete operation is appropriate. `claude.run.cancel` targets an exact run; inspect its subsequent result to confirm settlement. Cancelling the original start request after it returned does not cancel the detached run. +After saving the needed result, events and native session ID, use +`claude.run.release` to discard the retained adapter handle. Reading a result +does not release it. Release refuses active work or unconfirmed cleanup and +never deletes the native conversation. Unreleased completed results remain +available until bounded retention eviction or adapter shutdown. + Adapter run_id and native session_id are different. Resume with an explicit native resume_session_id, continue_previous, or create a session_id; do not combine them. The adapter's in-memory handles are invalid after process replacement, but native persisted conversations remain under vendor control. Unknown execution outcome does not authorize replaying a prompt. permission_mode defaults to dontAsk. Use plan or acceptEdits only for the requested native behavior. The adapter does not expose bypass modes or interactive permission brokerage. dontAsk is not a filesystem sandbox: existing native configuration still applies. Host Full Shell never silently changes Claude permissions.