Controlled A/B testing for AI agent skills, .cursorrules, CLAUDE.md, and system prompts — measure whether agent instructions actually improve model performance or just waste context tokens.
Not a syntax linter or basic assertion check — a controlled scientific experiment that measures statistical effect size.
Most published SKILL.md files, .cursorrules, and agent prompts have never been tested against an unprompted baseline. You cannot tell whether they help your model, make it worse, or are purely decorative prompt bloat. Skillcheck replaces vibe checks with empirical AI evaluation: paired control vs. treatment trials, double-blind grading, and bootstrap confidence intervals.
Point it at any Markdown skill or rule file and it runs an automated A/B experiment: it generates fresh domain-specific tasks, has the model solve each task with and without the skill injected, grades both arms blind, and reports the measured effect with a 95% bootstrap confidence interval and a 0–100 satisfaction score.
Important
Empirical Findings from our 20-Skill Seed Corpus:
We benchmarked 20 popular community agent skills (mattpocock/skills, awesome-claude-md, etc.) across 600+ blind-graded tasks:
- 60% were statistical PLACEBOS — consuming hundreds of extra prompt tokens with
0.0 ppnet lift. - 15% actively HARMED accuracy — over-constraining the model and causing cognitive tunnel vision.
- Only 25% genuinely HELPED — delivering verified
+15to+35 pplifts on nuanced edge cases.
📊 View live benchmark results on the Skillcheck Leaderboard →
skillcheck-demo.mp4
$ skillcheck
✓ Evaluation tasks ready · 12s
✓ Trials complete (30/30) · 1m 38s
✓ Grading complete (30/30) · 41s
✓ Analysis complete · 2m 33s
╭────────────────────────────────────────────────────────╮
│ SKILLCHECK RESULT │
├────────────────────────────────────────────────────────┤
│ Skill API Documentation │
│ Run size 5 tasks × 3 trials │
│ │
│ Verdict HELPS │
│ The skill HELPED — model passed 80% of tasks with it │
│ vs 55% without. │
│ │
│ With skill 80.0% of tasks passed │
│ Without skill 55.0% of tasks passed │
│ Skill effect +25.0 pp change in pass rate │
│ Confidence +8.0 pp to +42.0 pp (95% range) │
│ Token cost +480 tokens to include the skill │
├────────────────────────────────────────────────────────┤
│ Satisfaction ██████████████░░░░ 75.0/100 GOOD │
╰────────────────────────────────────────────────────────╯
- Install
- Quick start
- Supported formats & agents
- Why skillcheck
- Skillcheck vs. traditional linters
- How it works
- Architecture
- Commands
- CI/CD integration
- Badges for your repository
- Reading the result
- Effort levels
- Terminal experience
- Configuration
- Model choice
- Self-hosting
- Development
- Star history
- License
npm install -g @sx4im/skillcheck@latest
# or run it without installing:
npx @sx4im/skillcheck@latestRequires Node.js 20+. Works on Linux, macOS, and Windows.
Skillcheck checks npm for a newer version about once a day and offers to update
(like the Codex and Gemini CLIs). Disable with SKILLCHECK_NO_UPDATE_CHECK=1.
Experience Skillcheck's interactive step tracker and satisfaction scorecard immediately:
npx @sx4im/skillcheck demoIf you already have any standard model API key set in your environment, Skillcheck runs out of the box with zero registration:
# Works instantly with OpenAI, Anthropic, Gemini, Groq, Mistral, OpenRouter, or NVIDIA NIM:
export OPENAI_API_KEY="sk-..." # or ANTHROPIC_API_KEY, GEMINI_API_KEY, etc.
npx @sx4im/skillcheck check ./SKILL.mdDon't have API keys handy? Run the interactive assistant:
skillcheckOn first run without an existing key, it offers to use Skillcheck Cloud (grab a free token from skillcheck.page — includes 10 free cloud evaluations). Key entry is masked, verified securely before saving, and launches a full-screen terminal file picker: navigate folders with the arrow keys, pick any .md file, select an effort level, and watch the live progress tracker until the result card lands.
Point it straight at a file or folder:
skillcheck check ./SKILL.md
skillcheck ./my-skill-folder # a folder containing a .md
skillcheck check ./SKILL.md --json # machine-readable output
skillcheck check ./SKILL.md --output result.jsonFully headless in CI / scripts:
export SKILLCHECK_TOKEN=chk_live_... # or OPENAI_API_KEY, etc.
skillcheck check ./SKILL.md --tasks 5 --trials 3 --jsonSkillcheck benchmarks any Markdown agent instruction, rule file, or prompt document:
| Tool / Framework | File / Path Convention | Example Command |
|---|---|---|
| Cursor | .cursorrules, .cursor/rules/*.mdc |
skillcheck check ./.cursorrules |
| Claude Code | SKILL.md, CLAUDE.md, .claude/skills/* |
skillcheck check ./SKILL.md |
| GitHub Copilot | .github/copilot-instructions.md |
skillcheck check ./.github/copilot-instructions.md |
| Windsurf & Cline | .windsurfrules, custom prompt files |
skillcheck check ./.windsurfrules |
| System & Agent Prompts | SYSTEM_PROMPT.md, prompts/*.md |
skillcheck check ./prompts/coding-agent.md |
| Skill Directories | Any folder with a supported file | skillcheck ./my-skill-directory/ |
- Ship skills with evidence — before you publish a Claude Code, Codex, or Cursor
SKILL.md, know whether it helps the model or just adds tokens. - Prompt A/B testing in one command — compare with-skill vs without-skill arms on the same tasks, with blind grading so the grader never sees which arm wrote the answer.
- CI-friendly LLM eval —
--json/--outputfor scripts and pipelines; hosted mode or bring-your-own-key across OpenAI, Anthropic, Gemini, Groq, Mistral, OpenRouter, and NVIDIA NIM. - Effect size, not vibes — bootstrap confidence intervals and a 0–100 satisfaction
score so you can tell a real lift from noise. Details in
METHODOLOGY.md.
Most prompt tools and eval frameworks test whether an output matches a static regex or schema. Skillcheck measures whether adding instructions produces statistically significant improvement over an unprompted model.
| Dimension | Vibe Check / Manual Testing | Assertion Linters (e.g. Promptfoo) | Skillcheck |
|---|---|---|---|
| Core Question | "Does the output look okay?" | "Does output pass hardcoded assertions?" | "Does this prompt actually improve model performance?" |
| Evaluation Design | Ad-hoc single prompts | Treatment-only (no control arm) | Controlled A/B Trial (With-skill vs. Without-skill) |
| Primary Metric | Subjective opinion | Boolean pass/fail rate |
Net Effect Size ( |
| Task Generation | Manual prompt typing | Manual YAML test authoring | Domain-Adaptive Synthesis (zero instruction leakage) |
| Grading Objectivity | High confirmation bias | Single-arm evaluation | Double-Blind Grading (grader never knows which arm produced output) |
| Outcome | Anecdotal | Test matrix table |
Statistical Verdict (HELPS / PLACEBO / HARMS) |
| Token Overhead | Unmeasured | Static token count | Marginal token cost measured against performance lift |
Skillcheck treats a skill like a drug trial treats a drug:
- Normalize — the skill file is parsed; its declared domain is read from
front matter (
domain:/description:) or the first heading. - Generate — a task generator sees only the domain, never the skill body, so the tasks can't leak the skill's instructions. It produces 2× candidate tasks; a seeded shuffle picks the final set.
- Run — every task runs
Ktrials in two arms: with the skill injected as a system prompt, and without it. Same model, same temperature. - Grade — a blind grader scores each output against the task's pass/fail criterion. It never knows which arm produced the output (outputs are shuffled), and grades at temperature 0 in JSON mode.
- Score — pass rates are compared pairwise and a 1000-iteration paired
bootstrap produces the effect size, a 95% confidence interval, and the verdict:
HELPS(CI fully above zero),HARMS(fully below), orPLACEBO(overlaps zero).
Every run is fresh: tasks and outputs are generated anew each time and check
stores nothing locally, so a repeated check is an independent measurement. Full
methodology in METHODOLOGY.md.
How a check flows from your terminal to the result card:
flowchart LR
subgraph CLI["skillcheck CLI (local)"]
direction TB
A["User input<br/>skillcheck check ./SKILL.md"] --> B{API key<br/>configured?}
B -- no --> C["Interactive setup<br/>masked key → verified → saved"]
B -- yes --> D
C --> D["Normalize skill<br/>name · domain · instructions"]
D --> E["Generate tasks<br/>domain only — never the skill body"]
E --> F["Run trials<br/>each task × K trials × 2 arms"]
F --> G["Blind grading<br/>shuffled outputs · temp 0 · JSON verdict"]
G --> H["Paired bootstrap<br/>1000 resamples → effect · 95% CI · verdict"]
H --> I{Output mode}
I -->|terminal| J["Result card<br/>+ animated satisfaction bar"]
I -->|"--json / --output"| K["JSON result<br/>task suite + transcript hashes"]
end
subgraph Cloud["Skillcheck Cloud (dashboard/, Vercel)"]
direction TB
P["Metered proxy<br/>/api/chat/completions<br/>authenticates chk_live key<br/>counts 1 run per check<br/>pins model · caps max_tokens"]
V["/api/key/verify"]
S["NVIDIA key<br/>stays server-side"]
P ~~~ V ~~~ S
end
subgraph NIM["NVIDIA NIM"]
direction TB
M["openai/gpt-oss-120b<br/>default for all three roles"]
end
CLI ==>|"model calls<br/>(generate · run · grade)"| Cloud
CLI -.->|"setup: key verify"| Cloud
Cloud ==>|server-side key| NIM
CLI -.->|"direct mode<br/>(NVIDIA_API_KEY)"| NIM
style CLI fill:#0b2942,stroke:#2d7dd2,color:#e8f0fe
style Cloud fill:#102a12,stroke:#3fa34d,color:#e8f5e9
style NIM fill:#2a2210,stroke:#d2a52d,color:#fdf6e3
Key properties:
- One metered run per check — every model call in a check shares a run id, so the hosted proxy counts the whole check as a single run.
- No provider key on your machine (hosted mode) — the CLI talks to the proxy; the NVIDIA key lives only on the server.
- Direct mode — set
NVIDIA_API_KEYto bypass the proxy entirely and call NVIDIA NIM with your own key.
skillcheck # interactive: pick a file, pick effort, run
skillcheck check <path> [--tasks N] [--trials K] [--output file.json] [--json] [--explain]
skillcheck setup # connect / change your API key
skillcheck logout # remove your saved API key
skillcheck eval <path> [--tasks N] [--trials K] [--output file.json] # raw JSON evaluator
skillcheck verify <result.json> [--sample n] # independently re-measure a published result
skillcheck corpus run --corpus corpus.json [--results dir] # batch-evaluate many skills
skillcheck rot [--results dir] [--output report.json] # detect skills that stopped helping
skillcheck --versionAccepted inputs: any Markdown (.md) file — SKILL.md, AGENTS.md, CLAUDE.md,
or any other .md — or a folder containing one. --tasks is capped at 50 and
--trials at 10; mistyped options are rejected rather than silently ignored.
Automate agent prompt and skill regression testing in GitHub Actions. Ensure no prompt edit or rule change silently degrades model performance:
name: Agent Skill Quality Gate
on:
pull_request:
paths:
- 'SKILL.md'
- '.cursorrules'
- '.github/copilot-instructions.md'
- 'prompts/**'
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- name: Run Skillcheck evaluation
run: npx @sx4im/skillcheck@latest check ./SKILL.md --tasks 5 --trials 3 --json --output skillcheck-result.json
env:
SKILLCHECK_TOKEN: ${{ secrets.SKILLCHECK_TOKEN }}Prove to users that your agent skill or .cursorrules delivers real, statistically verified lift:
[](https://github.com/sx4im/skillcheck)
[](https://github.com/sx4im/skillcheck)
[](https://github.com/sx4im/skillcheck)| Badge | Markdown Snippet |
|---|---|
| Tested with Skillcheck | [](https://github.com/sx4im/skillcheck) |
| Verdict: HELPS | [](https://github.com/sx4im/skillcheck) |
| Verdict: PLACEBO | [](https://github.com/sx4im/skillcheck) |
-
Verdict —
HELPS/PLACEBO/HARMS, decided by whether the 95% confidence interval clears zero.PLACEBOmeans no measurable difference, not necessarily a bad skill. -
Skill effect — the change in pass rate, in percentage points (pp).
-
Confidence — the 95% range for the true effect. A wide range means the run was inconclusive; re-run at a higher effort for a clearer signal.
-
Token cost — the prompt-token overhead of including the skill.
-
Satisfaction — a 0–100 quality score where 50 = no effect:
Score Band Score Band ≤10 Very bad 51–60 Decent 11–30 Bad 61–80 Good 31–50 Normal 81–100 Excellent
Each run is an independent experiment — tasks and model outputs are generated fresh every time, so results vary run to run. That variance is what the confidence interval quantifies.
Add --explain to see why a verdict landed where it did: a per-task breakdown
of the with/without pass rates, the change, and a contrasting example model output
from each arm — printed below the card, and included in --json output under
explain. It reuses the outputs the run already produced, so it costs nothing
extra.
skillcheck check ./SKILL.md --explain
skillcheck check ./SKILL.md --explain --json # breakdown under result.explainThe interactive run asks how thorough to be — more tasks/trials means a tighter confidence interval but a longer run:
| Level | Tasks × trials | Typical time |
|---|---|---|
| Quick | 2 × 1 | ~2–3 min |
| Standard | 3 × 3 | ~4–5 min |
| Thorough | 5 × 3 | ~6–7 min |
For scripted runs, set it explicitly: skillcheck check ./SKILL.md --tasks 5 --trials 3.
The CLI is built to feel like a first-class developer tool:
- Live step tracker — each phase persists as a receipt line
(
✓ Trials complete (30/30) · 1m 38s) while the active phase shows a spinner, a progress bar, and elapsed time. Progress renders on stderr, so piping stdout still gives you a clean result; piped stderr gets plain log lines instead of spinner frames. - Animated result card — the satisfaction bar sweeps to its score on interactive terminals; non-TTY output is the same card, static.
- Adaptive colour — truecolor gradients where supported, 256/16-colour
fallbacks elsewhere.
NO_COLOR(any non-empty value) disables colour entirely;FORCE_COLOR=1|2|3forces it on for piped output. - Quiet cancellation — backing out of a menu with
q/Ctrl+Cexits with code130and a one-line note, not an error dump. Run failures print a concise✗block on stderr. - Masked secrets — API-key entry never echoes; keys are stored at
~/.config/skillcheck/config.jsonwith0600permissions.
Credential precedence (highest wins):
| Setting | Mode | Effect |
|---|---|---|
<PROVIDER>_API_KEY |
direct | Call OpenAI, Anthropic, Gemini, Groq, Mistral, OpenRouter, or NVIDIA NIM with your own key |
SKILLCHECK_TOKEN |
hosted | Use a Skillcheck Cloud key without saving anything |
skillcheck setup |
interactive | Setup assistant: choose Hosted mode or Bring Your Own Key (BYOK) with live model selection |
Supported providers for Bring-Your-Own-Key (BYOK) direct mode:
- OpenAI (
OPENAI_API_KEY, models list viahttps://api.openai.com/v1/models) - Anthropic (
ANTHROPIC_API_KEY, models list viahttps://api.anthropic.com/v1/models) - Google Gemini (
GEMINI_API_KEY/GOOGLE_API_KEY, models list viahttps://generativelanguage.googleapis.com/v1beta/models) - Groq (
GROQ_API_KEY, models list viahttps://api.groq.com/openai/v1/models) - Mistral AI (
MISTRAL_API_KEY, models list viahttps://api.mistral.ai/v1/models) - OpenRouter (
OPENROUTER_API_KEY, models list viahttps://openrouter.ai/api/v1/models) - NVIDIA NIM (
NVIDIA_API_KEY, models list viahttps://integrate.api.nvidia.com/v1/models)
Optional environment variables:
| Variable | Default | Purpose |
|---|---|---|
SKILLCHECK_API_URL |
hosted cloud URL | Point at a self-hosted proxy deployment |
SKILLCHECK_MODEL |
provider default | Override the model for all three roles |
<PROVIDER>_GENERATOR_MODEL / <PROVIDER>_RUNNER_MODEL / <PROVIDER>_GRADER_MODEL |
— | Per-role model overrides for any provider (e.g. OPENAI_RUNNER_MODEL, ANTHROPIC_RUNNER_MODEL) |
OPENAI_BASE_URL / ANTHROPIC_BASE_URL / ... |
provider default | Custom API base URL per provider |
SKILLCHECK_TIMEOUT_MS / NVIDIA_TIMEOUT_MS |
120000 |
Per-request timeout |
SKILLCHECK_REQUEST_DELAY_MS / NVIDIA_REQUEST_DELAY_MS |
750 |
Minimum delay between requests (rate-limit safety) |
SKILLCHECK_MAX_ATTEMPTS |
8 |
Retry budget for retryable failures (429/5xx) |
SKILLCHECK_NO_UPDATE_CHECK |
— | 1 disables the daily update check |
SKILLCHECK_DEBUG |
— | 1 enables verbose per-call logging |
NO_COLOR |
— | Any non-empty value disables colour (spec) |
FORCE_COLOR |
— | 1/2/3 forces colour on, even when piped |
A .env file in the directory where you run skillcheck is loaded for convenience, but only an explicit allow-list is read from it: provider API keys (OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, GOOGLE_API_KEY, GROQ_API_KEY, MISTRAL_API_KEY, OPENROUTER_API_KEY, NVIDIA_API_KEY), SKILLCHECK_TOKEN / SKILLCHECK_API_KEY, model settings, and timeout settings. Everything else in the file is ignored, so a hostile .env in a cloned repo cannot redirect where your keys are sent or point the CLI at a hostile config file. Set anything else in your shell environment instead.
All three roles (task generator, runner, blind grader) default to the selected provider's default model (e.g., gpt-6-sol for OpenAI, claude-opus-5-5 for Anthropic, gemini-3.8-flash for Gemini, openai/gpt-oss-120b for NVIDIA NIM and Groq).
When running skillcheck setup, choosing Bring Your Own Key queries your provider's live /models endpoint, letting you pick any available model directly from your provider.
Role model overrides let you benchmark your specific production model while maintaining a strong generator and grader:
OPENAI_RUNNER_MODEL=gpt-6-luna skillcheck check ./SKILL.md or ANTHROPIC_RUNNER_MODEL=claude-haiku-4-5 skillcheck check ./SKILL.md.
Skillcheck's hosted tier runs behind a metered proxy so end users never need a provider key. The dashboard/ folder is a deployable Vercel app (Clerk sign-in, free-tier metering, optional Stripe upgrade) that issues chk_live_… keys and forwards completions to your server-side NVIDIA key. See the dashboard/README.md for deployment notes.
To bypass the hosted proxy, run skillcheck setup and select Bring Your Own Key, or set provider environment variables (e.g., OPENAI_API_KEY, ANTHROPIC_API_KEY, NVIDIA_API_KEY — see .env.example).
npm ci
npm run build # compile to dist/
npm test # vitest (280 tests)
npm run test:coverage # vitest + v8 coverage gate (84% statements/lines, 85% functions, 68% branches)
npm run lint # eslint (flat config, typescript-eslint)
npm run typecheck # strict TS, src + testsThe CLI lives in packages/cli (bin/skillcheck.ts → src/cli.ts).
packages/site is the Next.js leaderboard site; dashboard/ is the hosted cloud.
The suite runs fully offline: the model adapter is mocked, so an end-to-end test
drives the whole normalize → generate → run → grade → score pipeline (plus the
retry adapter, metering, and every command) without a single API call. The
interactive terminal shell is verified behaviourally rather than counted toward the
coverage percentage.
Every push and pull request runs ci.yml — lint,
typecheck, coverage, and a clean build on Node 20 and 22, a published-tarball
validation, and the dashboard's offline tests — and it makes no model calls, so it
runs on forks too. Tagging a release (npm version patch && git push --follow-tags)
triggers release.yml, which republishes to npm
with provenance. Separately, a scheduled rot workflow
re-runs the live corpus weekly and opens a PR when a skill's verdict regresses.
If Skillcheck saved you from shipping a placebo skill, a ⭐ helps other people find it.
Contributions are welcome! Check out our open Good First Issues to get started.
Please see CONTRIBUTING.md for local development setup and testing guidelines.