Skip to content

Repository files navigation

SKILLCHECK — is your skill actually helping the model?

npm version GitHub stars lifetime downloads CI Good First Issues License: MIT Node ≥20 Contributor Covenant

Controlled A/B testing for AI agent skills, .cursorrules, CLAUDE.md, and system prompts — measure whether agent instructions actually improve model performance or just waste context tokens.

Not a syntax linter or basic assertion check — a controlled scientific experiment that measures statistical effect size.

Most published SKILL.md files, .cursorrules, and agent prompts have never been tested against an unprompted baseline. You cannot tell whether they help your model, make it worse, or are purely decorative prompt bloat. Skillcheck replaces vibe checks with empirical AI evaluation: paired control vs. treatment trials, double-blind grading, and bootstrap confidence intervals.

Point it at any Markdown skill or rule file and it runs an automated A/B experiment: it generates fresh domain-specific tasks, has the model solve each task with and without the skill injected, grades both arms blind, and reports the measured effect with a 95% bootstrap confidence interval and a 0–100 satisfaction score.

Important

Empirical Findings from our 20-Skill Seed Corpus: We benchmarked 20 popular community agent skills (mattpocock/skills, awesome-claude-md, etc.) across 600+ blind-graded tasks:

  • 60% were statistical PLACEBOS — consuming hundreds of extra prompt tokens with 0.0 pp net lift.
  • 15% actively HARMED accuracy — over-constraining the model and causing cognitive tunnel vision.
  • Only 25% genuinely HELPED — delivering verified +15 to +35 pp lifts on nuanced edge cases.

📊 View live benchmark results on the Skillcheck Leaderboard →

skillcheck-demo.mp4
$ skillcheck

✓ Evaluation tasks ready · 12s
✓ Trials complete (30/30) · 1m 38s
✓ Grading complete (30/30) · 41s
✓ Analysis complete · 2m 33s

╭────────────────────────────────────────────────────────╮
│ SKILLCHECK RESULT                                      │
├────────────────────────────────────────────────────────┤
│ Skill         API Documentation                        │
│ Run size      5 tasks × 3 trials                       │
│                                                        │
│ Verdict       HELPS                                    │
│ The skill HELPED — model passed 80% of tasks with it   │
│ vs 55% without.                                        │
│                                                        │
│ With skill    80.0% of tasks passed                    │
│ Without skill 55.0% of tasks passed                    │
│ Skill effect  +25.0 pp change in pass rate             │
│ Confidence    +8.0 pp to +42.0 pp (95% range)          │
│ Token cost    +480 tokens to include the skill         │
├────────────────────────────────────────────────────────┤
│ Satisfaction  ██████████████░░░░  75.0/100  GOOD       │
╰────────────────────────────────────────────────────────╯

Table of contents

Install

npm install -g @sx4im/skillcheck@latest

# or run it without installing:
npx @sx4im/skillcheck@latest

Requires Node.js 20+. Works on Linux, macOS, and Windows.

Skillcheck checks npm for a newer version about once a day and offers to update (like the Codex and Gemini CLIs). Disable with SKILLCHECK_NO_UPDATE_CHECK=1.

Quick start

1. Instant 5-Second Demo (No API Keys Required)

Experience Skillcheck's interactive step tracker and satisfaction scorecard immediately:

npx @sx4im/skillcheck demo

2. Zero-Signup Evaluation (Bring Your Own Key)

If you already have any standard model API key set in your environment, Skillcheck runs out of the box with zero registration:

# Works instantly with OpenAI, Anthropic, Gemini, Groq, Mistral, OpenRouter, or NVIDIA NIM:
export OPENAI_API_KEY="sk-..."       # or ANTHROPIC_API_KEY, GEMINI_API_KEY, etc.
npx @sx4im/skillcheck check ./SKILL.md

3. Interactive Mode & Hosted Cloud (10 Free Checks Included)

Don't have API keys handy? Run the interactive assistant:

skillcheck

On first run without an existing key, it offers to use Skillcheck Cloud (grab a free token from skillcheck.page — includes 10 free cloud evaluations). Key entry is masked, verified securely before saving, and launches a full-screen terminal file picker: navigate folders with the arrow keys, pick any .md file, select an effort level, and watch the live progress tracker until the result card lands.

Point it straight at a file or folder:

skillcheck check ./SKILL.md
skillcheck ./my-skill-folder            # a folder containing a .md
skillcheck check ./SKILL.md --json      # machine-readable output
skillcheck check ./SKILL.md --output result.json

Fully headless in CI / scripts:

export SKILLCHECK_TOKEN=chk_live_...    # or OPENAI_API_KEY, etc.
skillcheck check ./SKILL.md --tasks 5 --trials 3 --json

Supported formats & agents

Skillcheck benchmarks any Markdown agent instruction, rule file, or prompt document:

Tool / Framework File / Path Convention Example Command
Cursor .cursorrules, .cursor/rules/*.mdc skillcheck check ./.cursorrules
Claude Code SKILL.md, CLAUDE.md, .claude/skills/* skillcheck check ./SKILL.md
GitHub Copilot .github/copilot-instructions.md skillcheck check ./.github/copilot-instructions.md
Windsurf & Cline .windsurfrules, custom prompt files skillcheck check ./.windsurfrules
System & Agent Prompts SYSTEM_PROMPT.md, prompts/*.md skillcheck check ./prompts/coding-agent.md
Skill Directories Any folder with a supported file skillcheck ./my-skill-directory/

Why skillcheck

  • Ship skills with evidence — before you publish a Claude Code, Codex, or Cursor SKILL.md, know whether it helps the model or just adds tokens.
  • Prompt A/B testing in one command — compare with-skill vs without-skill arms on the same tasks, with blind grading so the grader never sees which arm wrote the answer.
  • CI-friendly LLM eval — --json / --output for scripts and pipelines; hosted mode or bring-your-own-key across OpenAI, Anthropic, Gemini, Groq, Mistral, OpenRouter, and NVIDIA NIM.
  • Effect size, not vibes — bootstrap confidence intervals and a 0–100 satisfaction score so you can tell a real lift from noise. Details in METHODOLOGY.md.

Skillcheck vs. traditional linters

Most prompt tools and eval frameworks test whether an output matches a static regex or schema. Skillcheck measures whether adding instructions produces statistically significant improvement over an unprompted model.

Dimension Vibe Check / Manual Testing Assertion Linters (e.g. Promptfoo) Skillcheck
Core Question "Does the output look okay?" "Does output pass hardcoded assertions?" "Does this prompt actually improve model performance?"
Evaluation Design Ad-hoc single prompts Treatment-only (no control arm) Controlled A/B Trial (With-skill vs. Without-skill)
Primary Metric Subjective opinion Boolean pass/fail rate Net Effect Size ($\Delta$ Pass Rate) with 95% Bootstrap CI
Task Generation Manual prompt typing Manual YAML test authoring Domain-Adaptive Synthesis (zero instruction leakage)
Grading Objectivity High confirmation bias Single-arm evaluation Double-Blind Grading (grader never knows which arm produced output)
Outcome Anecdotal Test matrix table Statistical Verdict (HELPS / PLACEBO / HARMS)
Token Overhead Unmeasured Static token count Marginal token cost measured against performance lift

How it works

Skillcheck treats a skill like a drug trial treats a drug:

  1. Normalize — the skill file is parsed; its declared domain is read from front matter (domain:/description:) or the first heading.
  2. Generate — a task generator sees only the domain, never the skill body, so the tasks can't leak the skill's instructions. It produces 2× candidate tasks; a seeded shuffle picks the final set.
  3. Run — every task runs K trials in two arms: with the skill injected as a system prompt, and without it. Same model, same temperature.
  4. Grade — a blind grader scores each output against the task's pass/fail criterion. It never knows which arm produced the output (outputs are shuffled), and grades at temperature 0 in JSON mode.
  5. Score — pass rates are compared pairwise and a 1000-iteration paired bootstrap produces the effect size, a 95% confidence interval, and the verdict: HELPS (CI fully above zero), HARMS (fully below), or PLACEBO (overlaps zero).

Every run is fresh: tasks and outputs are generated anew each time and check stores nothing locally, so a repeated check is an independent measurement. Full methodology in METHODOLOGY.md.

Architecture

How a check flows from your terminal to the result card:

flowchart LR
    subgraph CLI["skillcheck CLI (local)"]
        direction TB
        A["User input<br/>skillcheck check ./SKILL.md"] --> B{API key<br/>configured?}
        B -- no --> C["Interactive setup<br/>masked key → verified → saved"]
        B -- yes --> D
        C --> D["Normalize skill<br/>name · domain · instructions"]
        D --> E["Generate tasks<br/>domain only — never the skill body"]
        E --> F["Run trials<br/>each task × K trials × 2 arms"]
        F --> G["Blind grading<br/>shuffled outputs · temp 0 · JSON verdict"]
        G --> H["Paired bootstrap<br/>1000 resamples → effect · 95% CI · verdict"]
        H --> I{Output mode}
        I -->|terminal| J["Result card<br/>+ animated satisfaction bar"]
        I -->|"--json / --output"| K["JSON result<br/>task suite + transcript hashes"]
    end

    subgraph Cloud["Skillcheck Cloud (dashboard/, Vercel)"]
        direction TB
        P["Metered proxy<br/>/api/chat/completions<br/>authenticates chk_live key<br/>counts 1 run per check<br/>pins model · caps max_tokens"]
        V["/api/key/verify"]
        S["NVIDIA key<br/>stays server-side"]
        P ~~~ V ~~~ S
    end

    subgraph NIM["NVIDIA NIM"]
        direction TB
        M["openai/gpt-oss-120b<br/>default for all three roles"]
    end

    CLI ==>|"model calls<br/>(generate · run · grade)"| Cloud
    CLI -.->|"setup: key verify"| Cloud
    Cloud ==>|server-side key| NIM
    CLI -.->|"direct mode<br/>(NVIDIA_API_KEY)"| NIM

    style CLI fill:#0b2942,stroke:#2d7dd2,color:#e8f0fe
    style Cloud fill:#102a12,stroke:#3fa34d,color:#e8f5e9
    style NIM fill:#2a2210,stroke:#d2a52d,color:#fdf6e3
Loading

Key properties:

  • One metered run per check — every model call in a check shares a run id, so the hosted proxy counts the whole check as a single run.
  • No provider key on your machine (hosted mode) — the CLI talks to the proxy; the NVIDIA key lives only on the server.
  • Direct mode — set NVIDIA_API_KEY to bypass the proxy entirely and call NVIDIA NIM with your own key.

Commands

skillcheck                                  # interactive: pick a file, pick effort, run
skillcheck check <path> [--tasks N] [--trials K] [--output file.json] [--json] [--explain]
skillcheck setup                            # connect / change your API key
skillcheck logout                           # remove your saved API key
skillcheck eval <path> [--tasks N] [--trials K] [--output file.json]   # raw JSON evaluator
skillcheck verify <result.json> [--sample n]  # independently re-measure a published result
skillcheck corpus run --corpus corpus.json [--results dir]             # batch-evaluate many skills
skillcheck rot [--results dir] [--output report.json]                  # detect skills that stopped helping
skillcheck --version

Accepted inputs: any Markdown (.md) file — SKILL.md, AGENTS.md, CLAUDE.md, or any other .md — or a folder containing one. --tasks is capped at 50 and --trials at 10; mistyped options are rejected rather than silently ignored.

CI/CD integration

Automate agent prompt and skill regression testing in GitHub Actions. Ensure no prompt edit or rule change silently degrades model performance:

name: Agent Skill Quality Gate
on:
  pull_request:
    paths:
      - 'SKILL.md'
      - '.cursorrules'
      - '.github/copilot-instructions.md'
      - 'prompts/**'

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22

      - name: Run Skillcheck evaluation
        run: npx @sx4im/skillcheck@latest check ./SKILL.md --tasks 5 --trials 3 --json --output skillcheck-result.json
        env:
          SKILLCHECK_TOKEN: ${{ secrets.SKILLCHECK_TOKEN }}

Badges for your repository

Prove to users that your agent skill or .cursorrules delivers real, statistically verified lift:

[![Tested with Skillcheck](https://img.shields.io/badge/tested%20with-skillcheck-0ea5e9?style=flat-square&logo=github)](https://github.com/sx4im/skillcheck)
[![Skillcheck: HELPS](https://img.shields.io/badge/skillcheck-HELPS%20%2B25%25-10b981?style=flat-square)](https://github.com/sx4im/skillcheck)
[![Skillcheck: PLACEBO](https://img.shields.io/badge/skillcheck-PLACEBO%20%C2%B10%25-64748b?style=flat-square)](https://github.com/sx4im/skillcheck)
Badge Markdown Snippet
Tested with Skillcheck [![Tested with Skillcheck](https://img.shields.io/badge/tested%20with-skillcheck-0ea5e9?style=flat-square&logo=github)](https://github.com/sx4im/skillcheck)
Verdict: HELPS [![Skillcheck: HELPS](https://img.shields.io/badge/skillcheck-HELPS%20%2B25%25-10b981?style=flat-square)](https://github.com/sx4im/skillcheck)
Verdict: PLACEBO [![Skillcheck: PLACEBO](https://img.shields.io/badge/skillcheck-PLACEBO%20%C2%B10%25-64748b?style=flat-square)](https://github.com/sx4im/skillcheck)

Reading the result

  • Verdict — HELPS / PLACEBO / HARMS, decided by whether the 95% confidence interval clears zero. PLACEBO means no measurable difference, not necessarily a bad skill.

  • Skill effect — the change in pass rate, in percentage points (pp).

  • Confidence — the 95% range for the true effect. A wide range means the run was inconclusive; re-run at a higher effort for a clearer signal.

  • Token cost — the prompt-token overhead of including the skill.

  • Satisfaction — a 0–100 quality score where 50 = no effect:

    Score Band Score Band
    ≤10 Very bad 51–60 Decent
    11–30 Bad 61–80 Good
    31–50 Normal 81–100 Excellent

Each run is an independent experiment — tasks and model outputs are generated fresh every time, so results vary run to run. That variance is what the confidence interval quantifies.

Add --explain to see why a verdict landed where it did: a per-task breakdown of the with/without pass rates, the change, and a contrasting example model output from each arm — printed below the card, and included in --json output under explain. It reuses the outputs the run already produced, so it costs nothing extra.

skillcheck check ./SKILL.md --explain
skillcheck check ./SKILL.md --explain --json    # breakdown under result.explain

Effort levels

The interactive run asks how thorough to be — more tasks/trials means a tighter confidence interval but a longer run:

Level Tasks × trials Typical time
Quick 2 × 1 ~2–3 min
Standard 3 × 3 ~4–5 min
Thorough 5 × 3 ~6–7 min

For scripted runs, set it explicitly: skillcheck check ./SKILL.md --tasks 5 --trials 3.

Terminal experience

The CLI is built to feel like a first-class developer tool:

  • Live step tracker — each phase persists as a receipt line (✓ Trials complete (30/30) · 1m 38s) while the active phase shows a spinner, a progress bar, and elapsed time. Progress renders on stderr, so piping stdout still gives you a clean result; piped stderr gets plain log lines instead of spinner frames.
  • Animated result card — the satisfaction bar sweeps to its score on interactive terminals; non-TTY output is the same card, static.
  • Adaptive colour — truecolor gradients where supported, 256/16-colour fallbacks elsewhere. NO_COLOR (any non-empty value) disables colour entirely; FORCE_COLOR=1|2|3 forces it on for piped output.
  • Quiet cancellation — backing out of a menu with q/Ctrl+C exits with code 130 and a one-line note, not an error dump. Run failures print a concise ✗ block on stderr.
  • Masked secrets — API-key entry never echoes; keys are stored at ~/.config/skillcheck/config.json with 0600 permissions.

Configuration

Credential precedence (highest wins):

Setting Mode Effect
<PROVIDER>_API_KEY direct Call OpenAI, Anthropic, Gemini, Groq, Mistral, OpenRouter, or NVIDIA NIM with your own key
SKILLCHECK_TOKEN hosted Use a Skillcheck Cloud key without saving anything
skillcheck setup interactive Setup assistant: choose Hosted mode or Bring Your Own Key (BYOK) with live model selection

Supported providers for Bring-Your-Own-Key (BYOK) direct mode:

  • OpenAI (OPENAI_API_KEY, models list via https://api.openai.com/v1/models)
  • Anthropic (ANTHROPIC_API_KEY, models list via https://api.anthropic.com/v1/models)
  • Google Gemini (GEMINI_API_KEY / GOOGLE_API_KEY, models list via https://generativelanguage.googleapis.com/v1beta/models)
  • Groq (GROQ_API_KEY, models list via https://api.groq.com/openai/v1/models)
  • Mistral AI (MISTRAL_API_KEY, models list via https://api.mistral.ai/v1/models)
  • OpenRouter (OPENROUTER_API_KEY, models list via https://openrouter.ai/api/v1/models)
  • NVIDIA NIM (NVIDIA_API_KEY, models list via https://integrate.api.nvidia.com/v1/models)

Optional environment variables:

Variable Default Purpose
SKILLCHECK_API_URL hosted cloud URL Point at a self-hosted proxy deployment
SKILLCHECK_MODEL provider default Override the model for all three roles
<PROVIDER>_GENERATOR_MODEL / <PROVIDER>_RUNNER_MODEL / <PROVIDER>_GRADER_MODEL — Per-role model overrides for any provider (e.g. OPENAI_RUNNER_MODEL, ANTHROPIC_RUNNER_MODEL)
OPENAI_BASE_URL / ANTHROPIC_BASE_URL / ... provider default Custom API base URL per provider
SKILLCHECK_TIMEOUT_MS / NVIDIA_TIMEOUT_MS 120000 Per-request timeout
SKILLCHECK_REQUEST_DELAY_MS / NVIDIA_REQUEST_DELAY_MS 750 Minimum delay between requests (rate-limit safety)
SKILLCHECK_MAX_ATTEMPTS 8 Retry budget for retryable failures (429/5xx)
SKILLCHECK_NO_UPDATE_CHECK — 1 disables the daily update check
SKILLCHECK_DEBUG — 1 enables verbose per-call logging
NO_COLOR — Any non-empty value disables colour (spec)
FORCE_COLOR — 1/2/3 forces colour on, even when piped

A .env file in the directory where you run skillcheck is loaded for convenience, but only an explicit allow-list is read from it: provider API keys (OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, GOOGLE_API_KEY, GROQ_API_KEY, MISTRAL_API_KEY, OPENROUTER_API_KEY, NVIDIA_API_KEY), SKILLCHECK_TOKEN / SKILLCHECK_API_KEY, model settings, and timeout settings. Everything else in the file is ignored, so a hostile .env in a cloned repo cannot redirect where your keys are sent or point the CLI at a hostile config file. Set anything else in your shell environment instead.

Model choice

All three roles (task generator, runner, blind grader) default to the selected provider's default model (e.g., gpt-6-sol for OpenAI, claude-opus-5-5 for Anthropic, gemini-3.8-flash for Gemini, openai/gpt-oss-120b for NVIDIA NIM and Groq).

When running skillcheck setup, choosing Bring Your Own Key queries your provider's live /models endpoint, letting you pick any available model directly from your provider.

Role model overrides let you benchmark your specific production model while maintaining a strong generator and grader: OPENAI_RUNNER_MODEL=gpt-6-luna skillcheck check ./SKILL.md or ANTHROPIC_RUNNER_MODEL=claude-haiku-4-5 skillcheck check ./SKILL.md.

Self-hosting

Skillcheck's hosted tier runs behind a metered proxy so end users never need a provider key. The dashboard/ folder is a deployable Vercel app (Clerk sign-in, free-tier metering, optional Stripe upgrade) that issues chk_live_… keys and forwards completions to your server-side NVIDIA key. See the dashboard/README.md for deployment notes.

To bypass the hosted proxy, run skillcheck setup and select Bring Your Own Key, or set provider environment variables (e.g., OPENAI_API_KEY, ANTHROPIC_API_KEY, NVIDIA_API_KEY — see .env.example).

Development

npm ci
npm run build          # compile to dist/
npm test               # vitest (280 tests)
npm run test:coverage  # vitest + v8 coverage gate (84% statements/lines, 85% functions, 68% branches)
npm run lint           # eslint (flat config, typescript-eslint)
npm run typecheck      # strict TS, src + tests

The CLI lives in packages/cli (bin/skillcheck.ts → src/cli.ts). packages/site is the Next.js leaderboard site; dashboard/ is the hosted cloud.

The suite runs fully offline: the model adapter is mocked, so an end-to-end test drives the whole normalize → generate → run → grade → score pipeline (plus the retry adapter, metering, and every command) without a single API call. The interactive terminal shell is verified behaviourally rather than counted toward the coverage percentage.

Every push and pull request runs ci.yml — lint, typecheck, coverage, and a clean build on Node 20 and 22, a published-tarball validation, and the dashboard's offline tests — and it makes no model calls, so it runs on forks too. Tagging a release (npm version patch && git push --follow-tags) triggers release.yml, which republishes to npm with provenance. Separately, a scheduled rot workflow re-runs the live corpus weekly and opens a PR when a skill's verdict regresses.

Star history

If Skillcheck saved you from shipping a placebo skill, a ⭐ helps other people find it.

Star history chart for sx4im/skillcheck

Contributing

Contributions are welcome! Check out our open Good First Issues to get started.

Please see CONTRIBUTING.md for local development setup and testing guidelines.

License

MIT

About

Blind A/B testing for AI agent skills, .cursorrules, CLAUDE.md, and system prompts. Detects placebo and harmful instructions.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

28 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages