Skip to content

Add description-length-unit check and multibyte fixtures - #3

Merged
dacharyc merged 2 commits into
mainfrom
check/description-length-unit
Sep 20, 2026
Merged

dacharyc merged 2 commits into
mainfrom
check/description-length-unit

Conversation

@dacharyc

@dacharyc dacharyc commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Why

The spec says the description "must be 1-1024 characters" and never defines a character. Reading the public loaders (2026-09-19), implementations that enforce the limit split three ways:

Implementation Unit Behavior
skills-ref reference validator code points (Python len) rejects
Codex CLI code points (char_indices) truncates catalog entry at 1024 + ...
Cline UTF-16 code units (JS .length) rejects
Crush UTF-8 bytes (Go len) rejects
Claude Code, Antigravity, Gemini CLI, OpenCode no enforcement found loads

skill-validator counted bytes too until agent-ecosystem/skill-validator#94 (fix in #95). A 984-character Japanese description is 2952 bytes: compliant to the reference validator, rejected or truncated on a byte counter, with no error either way.

The existing probe-long-description fixture is ASCII, so bytes, UTF-16 units, and code points are all 1116 and it cannot tell the units apart.

What's added

  • Two fixtures, each with ASCII head and tail markers and a body canary, following the probe-long-description pattern:
    • probe-multibyte-description: Japanese prose. 848 code points, 848 UTF-16 units, 1822 bytes. Only a byte counter sees it as oversize.
    • probe-astral-description: emoji padding. 869 code points, 1319 UTF-16 units, 2219 bytes. UTF-16 and byte counters both see it as oversize.
  • description-length-unit check in Category 10 (checks.md), read together with the ASCII fixture. Four postures: no enforcement, code points, UTF-16 units, bytes.
  • Runner: three-session spec in checks/discovery.go grading each fixture's description as intact, truncated, rejected, or unobserved, then combining; verdict labels in report.go; spec-alignment judge in specalign.go. Check list version bumped to 0.3, with a changelog entry.
  • README, benchmark-skills README (inventory, mapping, structural validation, canary index), findings template, and site counts updated (41 checks, 35 fixtures).

Observed results (headless, 2026-09-19)

Harness Version ASCII (1116) Multibyte (848 cp / 1822 B) Astral (869 cp / 1319 u16) Verdict Confidence
Claude Code 2.1.267 intact intact intact no-length-enforcement transcript-direct
Antigravity agy intact intact intact no-length-enforcement behavioral-inference
Codex CLI 0.154.0 truncated at 1024 + ... intact intact counts-code-points transcript-direct

All three match the expectation set out before running. Codex's own answer for the ASCII fixture: "The catalog description ends with: 'If you can read every sentence of this description including the final...'", while both multibyte fixtures quoted their tail markers.

Reports and site platform pages are regenerated with these findings merged into the August batches (the runner's latest-finding-wins merge), which moves each page's test date to 2026-09-19 and adds the newer models to the observed list. Drop that commit's report changes if you would rather regenerate from a full batch later.

Dependency bump

The pinned agentminutes v0.3.1 failed both the Claude Code and Codex runs with format-drift errors (atis-latch in Claude Code 2.1.267, token_usage_record in Codex 0.154.0). agentminutes v0.5.0 and skillxp v0.1.3 (both 2026-09-13) recognize them, so the runner now pins those.

Notes

  • Both fixtures pass skill-validator validate structure on the #95 branch. On 1.6.1 and earlier they fail with "description exceeds 1024 characters" because those versions count bytes; the structural validation table says so.
  • Vale reports no new alerts on the changed prose.
  • report.go's verdict map was already unaligned per gofmt; I matched the surrounding style rather than reformat the whole map.

🤖 Generated with Claude Code

dacharyc and others added 2 commits September 19, 2026 20:55
The spec caps description at "1024 characters" without defining a
character, and enforcing implementations disagree: the skills-ref
reference validator and Codex count code points, Cline counts UTF-16
code units, and Crush counts UTF-8 bytes (skill-validator did too,
until issue #94). The existing ASCII oversize fixture cannot tell these
apart because its byte, code unit, and code point counts are equal.

Add two fixtures that stay under 1024 code points but exceed the limit
in other units: probe-multibyte-description (Japanese prose, over in
bytes only) and probe-astral-description (emoji, over in UTF-16 units
and bytes). Read with probe-long-description, the three outcomes
separate no enforcement, code points, UTF-16 units, and bytes.

Wire the check into the runner with a three-session spec, report
labels, and a spec-alignment judge; bump the check list to 0.3.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Results: Claude Code (2.1.267) and Antigravity load all three fixtures
intact (no-length-enforcement); Codex (0.154.0) truncates the ASCII
overrun at 1024 but delivers both fixtures under 1024 code points whole
(counts-code-points). All three match the expectations set out in the
check. Reports and site platform pages regenerated with the new
findings merged in.

Bump agentminutes to v0.5.0 and skillxp to v0.1.3 so the transcript
parser recognizes record types the current Claude Code and Codex
releases emit (atis-latch, token_usage_record); v0.3.1 failed both runs
with format-drift errors. Also set the finding's vehicle from the body
canary loads, as the other validation checks do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@dacharyc
dacharyc merged commit fe15b70 into main Sep 20, 2026
1 check passed
@dacharyc
dacharyc deleted the check/description-length-unit branch September 20, 2026 01:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant