Add description-length-unit check and multibyte fixtures - #3
Merged
Merged
Conversation
The spec caps description at "1024 characters" without defining a character, and enforcing implementations disagree: the skills-ref reference validator and Codex count code points, Cline counts UTF-16 code units, and Crush counts UTF-8 bytes (skill-validator did too, until issue #94). The existing ASCII oversize fixture cannot tell these apart because its byte, code unit, and code point counts are equal. Add two fixtures that stay under 1024 code points but exceed the limit in other units: probe-multibyte-description (Japanese prose, over in bytes only) and probe-astral-description (emoji, over in UTF-16 units and bytes). Read with probe-long-description, the three outcomes separate no enforcement, code points, UTF-16 units, and bytes. Wire the check into the runner with a three-session spec, report labels, and a spec-alignment judge; bump the check list to 0.3. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Results: Claude Code (2.1.267) and Antigravity load all three fixtures intact (no-length-enforcement); Codex (0.154.0) truncates the ASCII overrun at 1024 but delivers both fixtures under 1024 code points whole (counts-code-points). All three match the expectations set out in the check. Reports and site platform pages regenerated with the new findings merged in. Bump agentminutes to v0.5.0 and skillxp to v0.1.3 so the transcript parser recognizes record types the current Claude Code and Codex releases emit (atis-latch, token_usage_record); v0.3.1 failed both runs with format-drift errors. Also set the finding's vehicle from the body canary loads, as the other validation checks do. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The spec says the description "must be 1-1024 characters" and never defines a character. Reading the public loaders (2026-09-19), implementations that enforce the limit split three ways:
len)char_indices)....length)len)skill-validator counted bytes too until agent-ecosystem/skill-validator#94 (fix in #95). A 984-character Japanese description is 2952 bytes: compliant to the reference validator, rejected or truncated on a byte counter, with no error either way.
The existing
probe-long-descriptionfixture is ASCII, so bytes, UTF-16 units, and code points are all 1116 and it cannot tell the units apart.What's added
probe-long-descriptionpattern:probe-multibyte-description: Japanese prose. 848 code points, 848 UTF-16 units, 1822 bytes. Only a byte counter sees it as oversize.probe-astral-description: emoji padding. 869 code points, 1319 UTF-16 units, 2219 bytes. UTF-16 and byte counters both see it as oversize.description-length-unitcheck in Category 10 (checks.md), read together with the ASCII fixture. Four postures: no enforcement, code points, UTF-16 units, bytes.checks/discovery.gograding each fixture's description as intact, truncated, rejected, or unobserved, then combining; verdict labels inreport.go; spec-alignment judge inspecalign.go. Check list version bumped to 0.3, with a changelog entry.Observed results (headless, 2026-09-19)
no-length-enforcementno-length-enforcement...counts-code-pointsAll three match the expectation set out before running. Codex's own answer for the ASCII fixture: "The catalog description ends with: 'If you can read every sentence of this description including the final...'", while both multibyte fixtures quoted their tail markers.
Reports and site platform pages are regenerated with these findings merged into the August batches (the runner's latest-finding-wins merge), which moves each page's test date to 2026-09-19 and adds the newer models to the observed list. Drop that commit's report changes if you would rather regenerate from a full batch later.
Dependency bump
The pinned
agentminutesv0.3.1 failed both the Claude Code and Codex runs with format-drift errors (atis-latchin Claude Code 2.1.267,token_usage_recordin Codex 0.154.0).agentminutesv0.5.0 andskillxpv0.1.3 (both 2026-09-13) recognize them, so the runner now pins those.Notes
skill-validator validate structureon the #95 branch. On 1.6.1 and earlier they fail with "description exceeds 1024 characters" because those versions count bytes; the structural validation table says so.report.go's verdict map was already unaligned per gofmt; I matched the surrounding style rather than reformat the whole map.🤖 Generated with Claude Code