Skill invocation runner: install an Agent Skill in a fresh fixture, invoke an agent harness headlessly, and report what actually reached the model, with transcript evidence.
Skill authors publish for 25+ platforms that each load, present, and manage skills differently, and mostly invisibly. skillxp makes that behavior observable: it stages a skill on a real harness and answers "what did the platform actually do with it" from the session transcript, not from the model's self-reporting. It builds on agentsummons (headless invocation) and agentminutes (transcript parsing), and adds the third layer of lore: how each harness discovers, activates, and records skills.
skillxp renders no verdicts. It produces observations; graders consume
them. The first consumer is the
agent-skill-implementation
loading benchmark, whose checks and verdict logic live in that repo's
benchmark-runner/.
Early development. Supported harnesses: Antigravity CLI, Claude Code,
Codex CLI, and GitHub Copilot CLI. Run skillxp harnesses to see the harness versions your
installed release was validated against, and skillxp doctor to compare
them with what you have installed (newer harness releases usually keep
working; validation records coverage, not a compatibility bound).
Results reflect headless behavior, which may differ from
interactive use.
brew install agent-ecosystem/tap/skillxp
# or
npm install -g skillxp
# or
pip install skillxp
# or
go install github.com/agent-ecosystem/skillxp/cmd/skillxp@latestAs a Go library: go get github.com/agent-ecosystem/skillxp. The
harnesses you observe must be installed and authenticated. On Windows,
use pip or a release binary;
the npm package temporarily has no Windows support (a registry naming
issue is being worked out with npm).
# Where does each harness discover project-level skills?
skillxp harnesses
# Stage a skill, activate it, and trace how two phrases
# reached the model
skillxp observe -harness claude-code -install ./my-skill \
-prompt "Activate the my-skill skill and follow its instructions." \
-activation -trace "PHRASE-IN-BODY-1234,PHRASE-IN-REFERENCE-5678" \
-out out/observe writes a bundle: observation.json (run metadata and the
trace report), session.json (the normalized transcript), and the
archived native transcript(s) that evidence line numbers point into.
One rule about phrases: never put a phrase you plan to trace in the
prompt, or it contaminates every echo location in the transcript.
As a library:
obs, err := observe.Observe(ctx, observe.Config{ArchiveDir: dir},
agentsummons.ClaudeCode, observe.Spec{
SkillDirs: []string{"./my-skill"},
Prompt: "Activate the my-skill skill and follow its instructions.",
Activation: true,
})
if err != nil {
return err
}
occs := trace.Phrase(obs.Session, "PHRASE-IN-BODY-1234", obs.Profile.EchoSubtypes)Full documentation is available at skillxp.dev:
- Quickstart: install skillxp and observe your first skill invocation
- Use Cases: cross-platform skill CI, benchmarks, regression watching, and skill iteration
- CLI: the harnesses and observe commands, the observation bundle, and the trace report's classifications
- Go Library: Observe, multi-turn sessions, repeat runs and rates, and the three packages
- Sandboxing: isolated homes per run, user-scope installs, and per-harness auth setup
- Harness Lore: the empirically established per-harness behavior the profiles encode
MIT.