Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,23 +10,23 @@ jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/checkout@v7

- uses: actions/setup-go@v6
- uses: actions/setup-go@v7
with:
go-version-file: go.mod

- uses: golangci/golangci-lint-action@v7
- uses: golangci/golangci-lint-action@v9

test:
strategy:
matrix:
os: [ubuntu-latest, windows-latest]
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v6
- uses: actions/checkout@v7

- uses: actions/setup-go@v6
- uses: actions/setup-go@v7
with:
go-version-file: go.mod

Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,15 +12,15 @@ jobs:
release:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/checkout@v7
with:
fetch-depth: 0

- uses: actions/setup-go@v6
- uses: actions/setup-go@v7
with:
go-version-file: go.mod

- uses: goreleaser/goreleaser-action@v6
- uses: goreleaser/goreleaser-action@v7
with:
version: "~> v2"
args: release --clean
Expand Down
49 changes: 49 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,55 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

### Added

- `analyze security` command and `security` check group (on by default in
`check`): scans skill files for prompt injection, credential access and
exfiltration, remote code execution (`curl … | sh`), disabled permission
checks, committed secrets, and invisible characters. The `security`
package is experimental
- `evals/evals.json` validation against the agentskills.io format, and an
informational note when a skill has no evals. Listing `evals` in
`--allow-dirs` skips the format check
- `evals/` and `agents/` (e.g. OpenAI Codex's `agents/openai.yaml`) are
accepted as conventional directories and excluded from token accounting
- Authoring checks from Anthropic's skill guidance: reference files linked
only from other references (one-level-deep rule), reference files over 100
lines without a table of contents, and backslash paths in SKILL.md
- Description checks: XML tags and the reserved words `anthropic`/`claude`
in names (rejected by the Claude API), first- or second-person wording, and
no statement of when to use the skill
- Client extension fields (`when_to_use`, `disable-model-invocation`,
`paths`, and other Claude Code and Grok Build fields) get a portability
note instead of an "unrecognized field" warning, and `description` plus
`when_to_use` over Claude Code's 1,536-character listing limit is flagged
- Content metrics `emphasis_markers`, `emphasis_ratio`, and
`rationale_markers`, with an informational note when all-caps emphasis is
dense

### Changed

- Skill names follow the `skills-ref` reference validator: NFKC-normalized
Unicode lowercase letters and digits are valid (with a portability warning
for non-ASCII names) instead of being rejected
- The LLM judge's Directive Precision rubric rewards unambiguous, gated
instructions that give their reasons, and no longer rewards emphatic
language; Novelty and Token Efficiency now count discoverable overviews and
rarely applicable instructions against a skill. Cached scores from the old
rubric are re-scored on the next run
- Default Anthropic judge model is now `claude-sonnet-5`; the default judge
content limit is 20,000 characters (up from 8,000), enough for a SKILL.md
at the spec's 5,000-token ceiling; the judge HTTP timeout is 120 seconds

### Fixed

- Judge content truncation counts characters, not bytes, so it no longer
splits multibyte characters or cuts CJK content to a third of the limit
- Cached judge scores are no longer served after the scored file changes;
the stored content hash is now checked, as the README described

## [1.6.2]

### Fixed
Expand Down
100 changes: 83 additions & 17 deletions README.md

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions cmd/analyze.go
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@ import (

var analyzeCmd = &cobra.Command{
Use: "analyze",
Short: "Analyze skill content or contamination",
Long: "Parent command for content and contamination analysis subcommands.",
Short: "Analyze skill content, contamination, or security",
Long: "Parent command for content, contamination, and security analysis subcommands.",
}

func init() {
Expand Down
53 changes: 53 additions & 0 deletions cmd/analyze_security.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
package cmd

import (
"github.com/spf13/cobra"

"github.com/agent-ecosystem/skill-validator/orchestrate"
"github.com/agent-ecosystem/skill-validator/types"
)

var strictSecurity bool

var analyzeSecurityCmd = &cobra.Command{
Use: "security <path>",
Short: "Scan for risky patterns (prompt injection, exfiltration, remote code, secrets)",
Long: `Scans every text file in the skill for patterns associated with the
vulnerability classes found in public skill marketplaces: prompt injection,
credential access and data exfiltration, remote code execution, disabled
safety controls, committed secrets, and invisible characters.

The rules are narrow signatures. A clean result is not proof of safety;
a finding is a prompt for human review.`,
Args: cobra.ExactArgs(1),
RunE: runAnalyzeSecurity,
}

func init() {
analyzeSecurityCmd.Flags().BoolVar(&strictSecurity, "strict", false, "treat warnings as errors (exit 1 instead of 2)")
analyzeCmd.AddCommand(analyzeSecurityCmd)
}

func runAnalyzeSecurity(cmd *cobra.Command, args []string) error {
_, mode, dirs, err := detectAndResolve(args)
if err != nil {
return err
}

eopts := exitOpts{strict: strictSecurity}
switch mode {
case types.SingleSkill:
r := orchestrate.RunSecurityAnalysis(dirs[0])
return outputReportWithExitOpts(r, false, eopts)
case types.MultiSkill:
mr := &types.MultiReport{}
for _, dir := range dirs {
r := orchestrate.RunSecurityAnalysis(dir)
mr.Skills = append(mr.Skills, r)
mr.Errors += r.Errors
mr.Warnings += r.Warnings
}
return outputMultiReportWithExitOpts(mr, false, eopts)
}
return nil
}
7 changes: 4 additions & 3 deletions cmd/check.go
Original file line number Diff line number Diff line change
Expand Up @@ -27,15 +27,15 @@ var (

var checkCmd = &cobra.Command{
Use: "check <path>",
Short: "Run all checks (structure + links + content + contamination)",
Short: "Run all checks (structure + links + content + contamination + security)",
Long: "Runs all validation and analysis checks. Use --only or --skip to select specific check groups.",
Args: cobra.ExactArgs(1),
RunE: runCheck,
}

func init() {
checkCmd.Flags().StringSliceVar(&checkOnly, "only", nil, "check groups to run: structure,links,content,contamination (comma-separated or repeatable)")
checkCmd.Flags().StringSliceVar(&checkSkip, "skip", nil, "check groups to skip: structure,links,content,contamination (comma-separated or repeatable)")
checkCmd.Flags().StringSliceVar(&checkOnly, "only", nil, "check groups to run: structure,links,content,contamination,security (comma-separated or repeatable)")
checkCmd.Flags().StringSliceVar(&checkSkip, "skip", nil, "check groups to skip: structure,links,content,contamination,security (comma-separated or repeatable)")
checkCmd.Flags().BoolVar(&perFileCheck, "per-file", false, "show per-file reference analysis")
checkCmd.Flags().BoolVar(&checkSkipOrphans, "skip-orphans", false,
"skip orphan file detection (unreferenced files in scripts/, references/, assets/)")
Expand All @@ -58,6 +58,7 @@ var validGroups = map[orchestrate.CheckGroup]bool{
orchestrate.GroupLinks: true,
orchestrate.GroupContent: true,
orchestrate.GroupContamination: true,
orchestrate.GroupSecurity: true,
}

func runCheck(cmd *cobra.Command, args []string) error {
Expand Down
7 changes: 4 additions & 3 deletions cmd/cmd_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -555,7 +555,8 @@ func TestValidateCommand_AllowedDirsSkill_WithoutFlag(t *testing.T) {
dir := fixtureDir(t, "allowed-dirs-skill")

r := structure.Validate(dir, structure.Options{})
// Without --allow-dirs, evals/ and testing/ should produce warnings
// Without --allow-dirs, testing/ should produce a warning; evals/ is a
// conventional directory (agentskills.io) and is accepted.
hasEvalsWarning := false
hasTestingWarning := false
for _, res := range r.Results {
Expand All @@ -566,8 +567,8 @@ func TestValidateCommand_AllowedDirsSkill_WithoutFlag(t *testing.T) {
hasTestingWarning = true
}
}
if !hasEvalsWarning {
t.Error("expected warning for evals/ without --allow-dirs")
if hasEvalsWarning {
t.Error("expected no warning for conventional evals/ directory")
}
if !hasTestingWarning {
t.Error("expected warning for testing/ without --allow-dirs")
Expand Down
4 changes: 2 additions & 2 deletions cmd/score_evaluate.go
Original file line number Diff line number Diff line change
Expand Up @@ -55,13 +55,13 @@ require an API key. This is useful when the CLI is already authenticated

func init() {
scoreEvaluateCmd.Flags().StringVar(&evalProvider, "provider", "anthropic", "LLM provider: anthropic, openai, or claude-cli")
scoreEvaluateCmd.Flags().StringVar(&evalModel, "model", "", "model name (default: claude-sonnet-4-5-20250929 for anthropic, gpt-5.2 for openai, sonnet for claude-cli)")
scoreEvaluateCmd.Flags().StringVar(&evalModel, "model", "", "model name (default: claude-sonnet-5 for anthropic, gpt-5.2 for openai, sonnet for claude-cli)")
scoreEvaluateCmd.Flags().StringVar(&evalBaseURL, "base-url", "", "API base URL (for openai-compatible endpoints)")
scoreEvaluateCmd.Flags().BoolVar(&evalRescore, "rescore", false, "re-score and overwrite cached results")
scoreEvaluateCmd.Flags().BoolVar(&evalSkillOnly, "skill-only", false, "score only SKILL.md, skip reference files")
scoreEvaluateCmd.Flags().BoolVar(&evalRefsOnly, "refs-only", false, "score only reference files, skip SKILL.md")
scoreEvaluateCmd.Flags().StringVar(&evalDisplay, "display", "aggregate", "reference score display: aggregate or files")
scoreEvaluateCmd.Flags().BoolVar(&evalFullContent, "full-content", false, "send full file content to LLM (default: truncate to 8,000 chars)")
scoreEvaluateCmd.Flags().BoolVar(&evalFullContent, "full-content", false, "send full file content to LLM (default: truncate to 20,000 chars)")
scoreEvaluateCmd.Flags().StringVar(&evalMaxTokensStyle, "max-tokens-style", "auto", "token parameter style: auto, max_tokens, or max_completion_tokens")
scoreCmd.AddCommand(scoreEvaluateCmd)
}
Expand Down
58 changes: 57 additions & 1 deletion content/content.go
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,10 @@ import (
)

// strongMarkerRes contains pre-compiled patterns for strong directive language
// markers (must, always, never, etc.) used to measure instruction specificity.
// markers (must, always, never, etc.) used to measure instruction specificity:
// the share of directive language that is strong rather than hedged. It is
// descriptive, not a quality score — a skill that explains when and why can
// be precise with few strong markers.
var strongMarkerRes = compilePatterns([]string{
`\bmust\b`, `\balways\b`, `\bnever\b`, `\bshall\b`,
`\brequired\b`, `\bdo not\b`, `\bdon't\b`, `\bensure\b`,
Expand All @@ -28,6 +31,26 @@ var weakMarkerRes = compilePatterns([]string{
`\bprefer\b`, `\btry to\b`, `\bif possible\b`,
})

// emphasisPattern matches all-caps emphasis (MUST, NEVER, CRITICAL, ...).
// Current Claude and GPT models follow instructions closely, so shouted
// directives cause overtriggering, and when many lines are emphasized none
// stands out. Matched case-sensitively: lowercase "must" is plain language.
var emphasisPattern = regexp.MustCompile(`\b(MUST|NEVER|ALWAYS|CRITICAL|IMPORTANT|MANDATORY|REQUIRED|SHALL|ESSENTIAL|DO NOT|DON'T)\b`)

// rationaleMarkerRes matches phrases that explain why an instruction exists.
// Instructions that give their reason generalize better than bare rules.
var rationaleMarkerRes = compilePatterns([]string{
`\bbecause\b`, `\bso that\b`, `\botherwise\b`, `\bto avoid\b`,
`\bto prevent\b`, `\bwhich means\b`, `\bthis ensures\b`, `\bthe reason\b`,
})

// Thresholds for the emphasis advisory: at least this many all-caps
// markers, appearing at this rate per sentence.
const (
emphasisAdvisoryMin = 5
emphasisAdvisoryRatio = 0.1
)

func compilePatterns(patterns []string) []*regexp.Regexp {
res := make([]*regexp.Regexp, len(patterns))
for i, p := range patterns {
Expand Down Expand Up @@ -274,6 +297,16 @@ func AnalyzeWithConfig(content string, cfg *ImperativeConfig) *types.ContentRepo
instructionSpecificity = float64(strongCount) / float64(totalMarkers)
}

// Emphasis and rationale, measured on prose only (code is not advice)
prose := util.CodeBlockStrip.ReplaceAllString(content, "")
prose = util.InlineCodeStrip.ReplaceAllString(prose, "")
emphasisCount := len(emphasisPattern.FindAllString(prose, -1))
emphasisRatio := 0.0
if sentenceCount > 0 {
emphasisRatio = float64(emphasisCount) / float64(sentenceCount)
}
rationaleCount := countMarkerMatches(prose, rationaleMarkerRes)

// Section count (H2+ headers)
sectionCount := len(sectionPattern.FindAllString(content, -1))

Expand All @@ -292,11 +325,34 @@ func AnalyzeWithConfig(content string, cfg *ImperativeConfig) *types.ContentRepo
StrongMarkers: strongCount,
WeakMarkers: weakCount,
InstructionSpecificity: util.RoundTo(instructionSpecificity, 4),
EmphasisMarkers: emphasisCount,
EmphasisRatio: util.RoundTo(emphasisRatio, 4),
RationaleMarkers: rationaleCount,
SectionCount: sectionCount,
ListItemCount: listItemCount,
}
}

// Advisories returns informational results for content patterns that
// current agent-vendor guidance advises against. file names the analyzed
// file in the results.
func Advisories(cr *types.ContentReport, file string) []types.Result {
if cr == nil {
return nil
}
ctx := types.ResultContext{Category: "Content", File: file}
var results []types.Result
if cr.EmphasisMarkers >= emphasisAdvisoryMin && cr.EmphasisRatio >= emphasisAdvisoryRatio {
results = append(results, ctx.Infof(
"%d all-caps emphasis markers (MUST, NEVER, CRITICAL, ...) across %d sentences — current models follow "+
"instructions closely and overtrigger on shouted directives, and when many lines are emphasized none "+
"stands out; use plain wording, explain why an instruction matters, and reserve emphasis for the one "+
"rule agents keep missing",
cr.EmphasisMarkers, cr.SentenceCount))
}
return results
}

func countImperativeSentencesWithDetector(sentences []string, d *imperativeDetector) int {
count := 0
for _, sentence := range sentences {
Expand Down
33 changes: 33 additions & 0 deletions content/content_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -462,3 +462,36 @@ func TestAnalyze_ChineseFullContent(t *testing.T) {
t.Errorf("expected at least 3 imperative sentences, got %d", r.ImperativeCount)
}
}

func TestAnalyze_EmphasisAndRationale(t *testing.T) {
text := "You MUST run the tests. NEVER skip linting. ALWAYS format code. " +
"This is CRITICAL. IMPORTANT: commit often. Use `MUST` in code freely.\n\n" +
"```\nMUST NEVER ALWAYS\n```\n\n" +
"Run migrations first because the schema changes. Pin versions so that builds repeat."
r := Analyze(text)
if r.EmphasisMarkers != 5 {
t.Errorf("EmphasisMarkers = %d, want 5 (code excluded)", r.EmphasisMarkers)
}
if r.RationaleMarkers != 2 {
t.Errorf("RationaleMarkers = %d, want 2", r.RationaleMarkers)
}
if r.EmphasisRatio <= 0 {
t.Errorf("EmphasisRatio = %v, want > 0", r.EmphasisRatio)
}
if n := len(Advisories(r, "SKILL.md")); n != 1 {
t.Errorf("expected 1 emphasis advisory, got %d", n)
}
}

func TestAdvisories_PlainWording(t *testing.T) {
r := Analyze("Run the tests before committing. You must pin versions because builds drift.")
if r.EmphasisMarkers != 0 {
t.Errorf("lowercase directives are not emphasis, got %d", r.EmphasisMarkers)
}
if got := Advisories(r, "SKILL.md"); len(got) != 0 {
t.Errorf("expected no advisories, got %v", got)
}
if got := Advisories(nil, "SKILL.md"); got != nil {
t.Errorf("expected nil for nil report, got %v", got)
}
}
1 change: 1 addition & 0 deletions doc.go
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@
// - [github.com/agent-ecosystem/skill-validator/structure] — directory layout, frontmatter, tokens, internal links
// - [github.com/agent-ecosystem/skill-validator/content] — content quality metrics (density, specificity, imperative ratio)
// - [github.com/agent-ecosystem/skill-validator/contamination] — cross-language contamination detection
// - [github.com/agent-ecosystem/skill-validator/security] — risky-pattern scanning (EXPERIMENTAL)
// - [github.com/agent-ecosystem/skill-validator/links] — external HTTP/HTTPS link validation
// - [github.com/agent-ecosystem/skill-validator/skill] — SKILL.md parsing (frontmatter + body)
// - [github.com/agent-ecosystem/skill-validator/skillcheck] — skill detection and reference file analysis
Expand Down
Loading
Loading