Skip to content

Align validation with current skill-authoring guidance - #97

Closed
aminmesbahi wants to merge 3 commits into
agent-ecosystem:mainfrom
aminmesbahi:best-practices-alignment
Closed

aminmesbahi wants to merge 3 commits into
agent-ecosystem:mainfrom
aminmesbahi:best-practices-alignment

Conversation

@aminmesbahi

@aminmesbahi aminmesbahi commented Sep 27, 2026 •

Copy link
Copy Markdown

TL;DR

Brings skill-validator in line with current skill-authoring guidance from the Agent Skills spec (skills-ref), Anthropic, OpenAI, and xAI, and with recent research (ETH Zurich's AGENTS.md study, SkillsBench, Agent Skills in the Wild):

  • Two bug fixes:
    • The LLM judge truncated input by bytes, which cut CJK text to about a third of the limit.
    • Cached judge scores were served even after a file changed.
  • Checks aligned with current guidance:
    • Names follow skills-ref (Unicode and NFKC).
    • Claude Code and Grok extension fields are no longer reported as "unrecognized".
    • evals/ and agents/ are accepted and evals/evals.json is validated.
    • New authoring checks: reference depth, tables of contents, backslash paths, and description wording.
  • The judge no longer rewards emphatic MUST/NEVER wording, which current prompting guidance says causes overtriggering. It now rewards clear, gated instructions that give their reasons.
  • New analyze security check for prompt injection, curl | sh, credential access, committed secrets, and invisible characters. About 26% of marketplace skills have at least one such pattern.
  • No --strict exit-code changes on existing skills: new heuristic checks report at info level.

What this PR does

Fixes

  • Judge truncation counts characters, not bytes. formatUserContent sliced content[:maxLen] even though DefaultMaxContentLen is documented in characters. That split multibyte characters and cut CJK content to a third of the limit, the same class of bug as Description length counts UTF-8 bytes but reports "characters" — multibyte descriptions rejected below 1024 chars #94.
  • Stale judge cache. Cache hits never compared the stored ContentHash, so an edited SKILL.md or reference file kept its old scores, contrary to the README. Hits now require a matching content hash and a matching RubricVersion. The rubric changes in this PR bump that version, so old scores are re-scored automatically.
  • CI actions updated

Spec and vendor alignment

  • Name validation matches skills-ref: names are NFKC-normalized, and Unicode lowercase letters and digits are allowed. Non-ASCII names were previously errors; they now get a portability warning, since the Claude API accepts only a-z0-9-. The directory-name comparison is NFKC-normalized too.
  • Warnings for Claude API rejections: the reserved words anthropic and claude in names, and XML tags in descriptions.
  • Client extension fields (when_to_use, disable-model-invocation, paths, context, effort, …) get an info-level portability note instead of an "unrecognized field" warning. Claude.ai uploads, the Skills API, and skills-ref reject these fields. description + when_to_use over Claude Code's 1,536-character listing limit is flagged.
  • Description wording (info): first- or second-person descriptions, and descriptions that never say when to use the skill.
  • evals/ and agents/ are conventional directories: no unknown-directory warning, and they are excluded from token accounting. agents/ covers OpenAI Codex's agents/openai.yaml. evals/evals.json is validated against the agentskills.io format, and an info note appears when a skill has no evals. --allow-dirs=evals opts out of the format check for custom eval formats.
  • Anthropic authoring rules:
    • Info note for reference files linked only from another reference, since references should be one level deep from SKILL.md.
    • Info note for reference files over 100 lines with no table of contents.
    • Warning for backslash paths in SKILL.md.

Content and scoring

  • New content metrics: emphasis_markers, emphasis_ratio, and rationale_markers, plus an info advisory when all-caps emphasis is dense. Existing JSON fields are unchanged.
  • Rubric changes:
    • Directive Precision now judges precision separately from intensity; pervasive CRITICAL/MUST counts against a skill.
    • Novelty counts content an agent could discover itself (directory overviews, restated READMEs) as common knowledge, following the ETH finding that overviews don't help.
    • Token Efficiency asks whether an agent would get the task wrong without each passage.
  • Judge defaults: claude-sonnet-5, a 20,000-character content limit (up from 8,000, so a SKILL.md at the spec's 5,000-token ceiling is scored whole), and a 120-second HTTP timeout.

Security (new, experimental)

  • New security package, analyze security command, and security check group, which check runs by default (--skip security turns it off). Findings include file and line.
  • Warnings for:
    • prompt-injection phrasing in markdown;
    • curl/wget piped into a shell or interpreter, and iwr … | iex;
    • encoded payloads that are decoded and executed;
    • access to credential stores;
    • environment variables sent over the network;
    • --dangerously-skip-permissions-style flags;
    • chmod 777;
    • invisible and text-direction characters.
  • Errors for committed private keys and well-known token formats.
  • Info when allowed-tools pre-approves unrestricted Bash.
  • evals/, hidden directories, and binary files are skipped.

Compatibility

  • New heuristic checks are info level, so --strict exit codes don't change for existing skills. On the repository's own fixtures and example skill, none of the new checks produced a warning or error.
  • Library behavior changes:
    • skill.UnrecognizedFields() keeps its meaning and is now sorted; the new skill.IsExtensionField and Skill.ExtensionFields do the filtering.
    • orchestrate.AllGroups() now includes security.
    • The security package is listed as experimental in the Stability section.
  • golang.org/x/text is pinned to v0.36.0 so the module still requires Go 1.25.5.

How to test

go test ./... -count=1
go run ./cmd/skill-validator check --skip links examples/review-skill
go run ./cmd/skill-validator analyze security testdata/valid-skill

New tests cover:

  • character-based truncation, including a CJK case;
  • the cache freshness helper, plus an end-to-end test that edited files are re-scored;
  • Unicode, NFKC, and reserved-word names;
  • description wording, extension fields, and the listing-length limit;
  • the evals schema, and conventional directories being excluded from token accounting;
  • tables of contents, backslash paths, and nested references;
  • emphasis and rationale metrics;
  • every security rule, including skipped directories and binary files.

Checklist

  • Tests pass locally (go test ./... -count=1). The one failure, TestDetectSkills/follows_symlinks, fails on main too: Windows needs extra privileges to create the symlink.
  • Tests pass with -race
  • Lint passes locally (golangci-lint run)
  • New functionality includes tests
  • Breaking changes are noted above (if any)

Bring the validator in line with the Agent Skills spec's reference
validator (skills-ref), Anthropic's and OpenAI's skill-authoring
guidance, and recent research on context files and skill security.

Fixes:
- Truncate judge input by characters, not bytes, so multibyte content
  is not split mid-rune or cut to a third of the limit
- Re-use cached judge scores only while the file content and rubric
  version match; edited files were previously served stale scores

Validation:
- Validate names like skills-ref: NFKC-normalized Unicode lowercase
  letters, digits, and hyphens; warn on non-ASCII names and on the
  reserved words "anthropic" and "claude" rejected by the Claude API
- Report client extension fields (when_to_use, paths, ...) as
  portability notes instead of unrecognized-field warnings; flag
  description + when_to_use over Claude Code's 1,536-character limit
- Check descriptions for XML tags, first/second-person wording, and a
  missing statement of when to use the skill
- Accept evals/ and agents/ as conventional directories, exclude them
  from token accounting, and validate evals/evals.json
- Flag nested reference chains, long reference files without a table
  of contents, and backslash paths in SKILL.md

Content and scoring:
- Add emphasis and rationale metrics with an advisory for dense
  all-caps emphasis
- Rework the Directive Precision rubric to reward unambiguous, gated
  instructions with reasons rather than emphatic language; count
  discoverable overviews and rarely applicable instructions against
  Novelty and Token Efficiency
- Default to claude-sonnet-5, a 20,000-character judge input limit, and
  a 120-second HTTP timeout

Security:
- Add an experimental security package, `analyze security` command, and
  `security` check group that flag prompt injection, remote code
  execution, credential access, disabled safety controls, committed
  secrets, and invisible characters

New heuristic checks report at info level so --strict pipelines keep
their current exit codes.
@aminmesbahi
aminmesbahi marked this pull request as draft September 27, 2026 12:46
staticcheck (ST1018) rejects string literals that contain Unicode
format characters. Write the zero-width and bidi-override characters
as escape sequences instead of literal characters; the matched set is
unchanged.
@aminmesbahi
aminmesbahi marked this pull request as ready for review September 27, 2026 13:01
@dacharyc

Copy link
Copy Markdown
Member

Hey @aminmesbahi - thank you for the PR, but unfortunately, I'm not able to accept this PR for a few reasons:

  • The scope of changes is far too wide for a single PR
  • Some of the changes you've proposed here aren't aligned with what harnesses are actually doing; for more details, you can take a look at https://agentskillimplementation.com , which is my companion research project. Specifically, I'm not trying to align with skills-ref, which Anthropic calls a "reference implementation" - I'm trying to align with whatever is required to get cross-harness portability. With naming, for example, I've already found harnesses diverge in practice on character type and length, and I've only got automated testing on 4 of the 25+ harnesses that implement support for skills; you can see more details here: https://agentskillimplementation.com/platforms/#validation-strictness
    • On this point, skill-validator is behind my research, and I plan to update it to reflect my findings here, but haven't gotten to that yet
  • Security is out-of-scope of this tool; there are already other tools working on this with far more resources, such as NVIDIA's Skill Spector: https://github.com/nvidia/skillspector

If you'd like to propose any new checks, feel free to open enhancement feature requests, one per check, describing the check behavior, why you think it's needed, and any evidence for it: https://github.com/agent-ecosystem/skill-validator/issues

If you'd like to propose changes to existing behavior, such as changes to judge behavior, I'd also propose one new enhancement feature request per behavioral change, so we can discuss each of them individually.

Thank you for your interest in contributing, and for taking the time to make the PR - apologies that I can't accept it.

@dacharyc dacharyc closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants