Skip to content

Count frontmatter field lengths in characters, not bytes - #95

Merged
dacharyc merged 1 commit into
mainfrom
fix/issue-94-length-in-characters
Sep 20, 2026
Merged

dacharyc merged 1 commit into
mainfrom
fix/issue-94-length-in-characters

Conversation

@dacharyc

Copy link
Copy Markdown
Member

Fixes #94.

What changed

  • structure/frontmatter.go: the name (64), description (1024), and compatibility (500) limits are now measured with utf8.RuneCountInString instead of len(). The counts printed in the error and pass messages use the same unit, and the limits are named constants.
  • judge/judge.go: sanitizeStringField caps at 1024 characters instead of 1024 bytes. Before this, a spec-compliant 1024-character CJK description (3072 bytes) was cut to about a third of its length before being sent to the judge.
  • Tests: CJK cases for description and compatibility on both sides of the limit, a multibyte name case, and rune-based truncation cases for the judge sanitizer.

Why code points

The spec says "characters" without defining the unit. The skills-ref reference validator uses Python len(), which is code points, and Codex CLI (the one major harness that enforces the description limit, by truncating its catalog entry at 1024) counts by char_indices. Counting code points matches both. Cline counts UTF-16 units and Crush counts bytes, so those two will still disagree on emoji and multibyte text respectively, but code points is the reference behavior.

The name pattern stays ASCII-only; the reference validator's Unicode-alphanumeric acceptance is a separate question.

Verification

$ skill-validator validate structure ./cjk-skill   # 984 CJK chars, 2952 bytes
before:  ✗ description exceeds 1024 characters (2952)
after:   ✓ description: (984 chars)

go test ./... and golangci-lint run are clean. No CHANGELOG entry, following the pattern of adding those in the release commit.

The name, description, and compatibility limits (64, 1024, 500) were
enforced with len(), which counts UTF-8 bytes, while the spec and the
error messages speak of characters. A 984-character CJK description is
2952 bytes and was rejected as "exceeds 1024 characters (2952)".

Count Unicode code points instead, the same unit the skills-ref
reference validator uses. The judge's sanitizeStringField cap moves to
1024 characters for the same reason: it was cutting spec-compliant
multibyte descriptions to a third of their length before scoring.

Fixes #94
@dacharyc
dacharyc merged commit 88289e4 into main Sep 20, 2026
3 checks passed
@dacharyc
dacharyc deleted the fix/issue-94-length-in-characters branch September 20, 2026 01:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Description length counts UTF-8 bytes but reports "characters" — multibyte descriptions rejected below 1024 chars

1 participant