Skip to content

Add operator diagnostics, observability, and recovery controls #26

Description

@hyperpolymath

Outcome

Give operators a clear, local-first view of queue health, node capability, costs, failures, and recovery actions.

Scope

  • Implement just doctor/CLI diagnostics for toolchains, signing keys, GitHub access, writable state, worker reachability, capability drift, clock skew, and webhook configuration.
  • Add structured logs and metrics for queue depth, wait/run/report latency, retries, suspension reasons, node utilisation, and avoided hosted-runner cost.
  • Provide read-only commands to inspect a Baton, verify an envelope, replay reporting, and explain routing decisions.
  • Add explicit repair commands for safe retry/cancel/requeue with audit records.
  • Define alerts that work without a paid SaaS dependency.

Acceptance criteria

  • Common failures identify the broken layer and the exact safe next action.
  • Operators can distinguish “check failed”, “no capable node”, “attestation invalid”, and “GitHub report failed”.
  • Metrics and logs contain no secrets or unredacted command output by default.
  • Recovery actions are idempotent and ledgered.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    automationBots, schedulers, dispatch, self-healing, fan-outmeta:roadmapForward planning; not yet actionable workpriority:p1High - schedule next

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions