Skip to content

feat: give Copilot publication and review bounded evidence tools #16

Description

@mingchuno

Problem Statement

A Copilot publication stage can receive a large captured change. To understand it, the agent may try to run a shell script over the diff. Publication is read-only, so the runner rejects shell permission requests. The agent can spend its stage deadline waiting or trying another route, leaving the operator without the requested commit and change-request text. The same navigation problem can affect Copilot review. Raising the timeout does not give the agent a reliable, bounded way to inspect the captured change.

Solution

Give Copilot publication and review invocations runner-owned tools to navigate the change evidence already captured for that stage. The agent can list changed paths, read a selected change in bounded parts, and search the captured text for a literal term. The runner restricts every result to the current invocation's verified evidence manifest. These tools are available by default; no client configuration changes. Copilot's built-in shell and write permissions remain denied for read-only stages. Codex keeps its current read-only access and evidence presentation.

User Stories

  1. As an operator, I want a Copilot publication stage to inspect a large change without asking for shell permission, so that it can finish drafting publication text within the configured deadline.
  2. As an operator, I want Copilot review to navigate the exact published change, so that its findings refer to the evidence captured for that review.
  3. As an operator, I want evidence tools available automatically, so that existing project configurations continue to work.
  4. As an operator, I want publication and review to remain read-only, so that inspecting changes cannot alter the managed checkout.
  5. As an operator, I want shell and write requests to remain denied, so that an evidence tool does not broaden agent permissions.
  6. As an operator, I want the evidence tools to serve only the current stage's captured change, so that another run's artifacts cannot be selected.
  7. As an operator, I want a clear error when a requested path, page, or chunk is absent, so that the agent can recover without guessing.
  8. As an operator, I want missing or modified evidence rejected, so that the agent cannot silently use stale or corrupted artifacts.
  9. As an operator, I want bounded tool responses and search results, so that large changes do not flood model context or stall a turn.
  10. As an operator, I want the agent to see which paths and portions remain unread, so that it can state material inspection limits honestly.
  11. As an operator, I want binary changes reported as metadata, so that the agent does not mistake absent text for an unchanged file.
  12. As an operator, I want tool calls and failures retained with the invocation events, so that I can diagnose an incomplete publication or review.
  13. As an operator, I want the existing output contract and format-correction rules preserved, so that evidence navigation cannot turn an incomplete or timed-out turn into a successful result.
  14. As a Copilot agent, I want a compact list of changed paths with change kinds and available chunks, so that I can choose what to inspect first.
  15. As a Copilot agent, I want to read one selected patch or untracked-content chunk at a time, so that I can inspect large changes incrementally.
  16. As a Copilot agent, I want to search captured text for a literal term and receive bounded locations, so that I can target relevant changes without a shell script.
  17. As a Codex user, I want existing publication and review behavior preserved, so that this Copilot improvement does not change my provider's tool access.

Implementation Decisions

  • Add a runner-owned evidence query layer over the current stage's captured manifest. It exposes changed-path listing with pagination, selected change reading in bounded chunks, and literal text search with bounded results and locations.
  • Use the existing evidence index and content chunks as the source of truth. Do not recapture a diff or read arbitrary checkout paths for these tools. Identify artifacts through manifest-backed path and chunk references, not caller-supplied filesystem paths.
  • Validate tool arguments, enforce page and response limits, and verify manifest membership and content hashes before returning text. Preserve the existing 64 KiB capture-chunk and 32 MiB stage-evidence limits.
  • Register the query operations as per-session Copilot custom tools only for publication and review invocations that have captured evidence. Tool descriptions and application-owned invocation instructions should direct the agent to use these tools for large changes and to report incomplete inspection.
  • Keep Copilot read-only permission handling unchanged: built-in shell, file-write, and other non-read requests remain denied for publication and review. Tool registration must not imply general shell access or a permission bypass.
  • Keep the runner's post-invocation checkout and evidence verification, output schema validation, deadline, and timeout behavior. Evidence tools do not make an incomplete turn eligible for JSON correction.
  • Preserve the stage prompt boundary: user-configurable prompts remain task instructions; evidence identity, tool availability, and permissions are application-owned.
  • Do not add a configuration field, migration, public SDK option, or client opt-in. Codex does not receive these custom tools in this change.
  • Make unavailable or corrupt evidence fail explicitly. A tool failure must not be represented as an empty successful search or empty patch.
  • Document the tools and Copilot read-only behavior in operator documentation without requiring users to change their configuration.

Testing Decisions

  • Prefer a high-level invocation test using a controlled Copilot session and captured evidence. Assert behavior visible to the stage: it can list, search, and read its own evidence; it cannot access another artifact or an arbitrary path; and shell/write requests remain denied. Extend the existing SDK contract and invocation test seams rather than creating a separate test-only architecture.
  • Cover pagination, response bounds, multi-chunk changes, literal-search matches and no-match results, binary metadata, invalid references, and hash mismatch with deterministic fixtures. Assert results and failures, not internal parsing steps.
  • Include a large synthetic change that would require multiple tool calls and verify the stage can produce a valid publication response without a shell approval. Keep this local fixture test distinct from a live-provider smoke test.
  • Verify that cancellation and stage timeout stop the invocation and that neither condition triggers format correction. Preserve existing tests for output validation and read-only checkout checks.
  • Run the repository's TypeScript, lint, and relevant test gates. A live Copilot run with the client's large diff is useful acceptance evidence but is separate from the deterministic gate.

Out of Scope

  • allowInspectionShell, arbitrary shell scripts, OS sandboxing, or broader publication/review write access.
  • Codex custom tools or changes to Codex sandbox and approval settings.
  • A new client-facing configuration interface or changes to stage timeout defaults.
  • Replacing the existing evidence capture format, lifting its size limits, or changing publication/review output contracts.
  • Granting arbitrary source-file access through the new evidence tools; existing provider read tools and stage policy remain responsible for source inspection outside the captured change.

Further Notes

The publication stage drafts commit and change-request text; commit and publication operations remain separate. Review inspects the exact published change. The runner already supplies indexed, hashed evidence and verifies it around read-only invocations. The new tools provide a bounded query interface over that evidence rather than a second source of truth. The original client timeout hypothesis should be checked against a live Copilot invocation after implementation; this spec does not claim the new tools alone will resolve every cause of stage timeout.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ready-for-agentSpecified and ready for agent implementation

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions