An Agent Skill for skill maintainers. It hardens any other Agent Skill that ships templates or tests and runs automatic agent evals on its releases: score every run against a fixed fidelity rubric read from the agent's own log, fix the skill at the root cause of each deviation, cut a patch release, let the release re-run the evals, and repeat until Claude Code, Codex and Gemini CLI all score full.
An eval that passes its checks has not proved the agent used the skill as written. The deviations live in the agent's log, not in the result's checks, and each one traces back to a sentence in the skill that allowed it. Score the log, fix the sentence, and let the next release's eval prove the fix.
This skill was written by the maintainer who has run the loop on a published skill, from its first scored
release to a unanimous full score. What it carries holds by construction: a rubric fixed before round one and
the same for every agent, so scores compare across rounds; every agent's template edit reproduced against the
template before it is adopted; a template check that type-checks the templates and matches the
documented test count before any round ships; and a stop rule that is set in advance and never moved.
references/provenance.md has the record.
One command, via the skills.sh CLI, which installs the skill into every skills-compatible agent it detects, including Claude Code, Codex CLI and Gemini CLI:
npx skills add timerise-ai/skill-eval-loopName the agents instead with -a, for example npx skills add timerise-ai/skill-eval-loop -a claude-code -a codex.
Nothing here is Claude-specific: the skill is a plain Agent Skills folder,
SKILL.md plus markdown references with no file that calls a model, so cloning it into an agent's skills
directory is all an install is. For Claude Code:
git clone https://github.com/timerise-ai/skill-eval-loop.git ~/.claude/skills/skill-eval-loopTo scope it to a single project instead, clone it into that project's .claude/skills/ directory. For another
agent, clone into that agent's skills directory, or symlink the Claude Code copy so one git pull updates
every agent:
mkdir -p ~/.agents/skills
ln -s ~/.claude/skills/skill-eval-loop ~/.agents/skills/skill-eval-loopUpdate the skill with git pull in its directory. The current release is 0.1.2. See
CHANGELOG.md. The skills index lists the other
Timerise Skills and how to install them all at once.
The skill activates automatically when a maintainer asks to fix a skill from its evals, iterate it to a full
score, or score its eval runs. Invoke it explicitly with /skill-eval-loop in Claude Code, $skill-eval-loop
in Codex CLI, or from /skills in Gemini CLI, in the working directory of the skill to harden, or with its
path: /skill-eval-loop ../my-skill. /skill-eval-loop score scores the latest release's runs and
stops, with no fix and no release.
Each host matches a task against the description its own way, so invoke the skill explicitly on a first run
rather than assuming it fired. Only SKILL.md is read up front; the references/ files load on demand.
| File | Contents |
|---|---|
SKILL.md |
Entry point: the loop diagram, the seam, critical facts, hard rules, invocation, quick start and reference directory |
references/rubric.md |
The eight-item fidelity rubric, how to derive it for a target skill, what is not scored, how scores are recorded |
references/collecting.md |
Finding and waiting for a release's eval run, downloading the logs, and read_logs.py for all three agents |
references/fixing.md |
The root-cause table for deviations, reproducing an agent's template edit, and where each fix goes |
references/releasing.md |
The template check with extract_blocks.py, the patch release, and the dispatch round |
references/loop.md |
Rounds, the stop rule, autonomy, and the per-round and final reports |
references/provenance.md |
The engineering ledger: the recorded session and its scores per round, what was kept deliberately, and what was added here and never exercised |
README.md |
This file |
CHANGELOG.md |
One section per release, newest first |
CLAUDE.md |
The editing conventions, for an agent editing this repository |
LICENSE |
MIT |
evals/ |
The prompts a maintainer types after installing (prompts.md) and one file per agent eval: the skill installed into an empty Next.js app, one prompt that names no target skill, no help, then type-checked, built and tested |
.github/workflows/agent-eval.yml |
The caller of the index's reusable eval workflow, run on every published release and on a maintainer's dispatch |
The skill builds no code, so its evals score the agent's fidelity to the loop, not the app: the checks only
confirm the app was left intact, and the notes on each run carry the score. The seam with the target skill is
its own rules and recipes, read and never rewritten: rubric items 4 to 6 come from the target's
non-negotiables, and the template check and commit convention from the target's CLAUDE.md. The prompt, the
unattended note, the harness and the workflow belong to the index and stay outside it.
These are never optional. Each is stated as a hard rule in SKILL.md, in the same order:
- Result frontmatter is never edited and a run is never deleted. The frontmatter is what was measured, so
scores go in the result's body, in a
chore(evals)commit;git diffon the frontmatter stays empty. - The eval is never changed to make a skill pass. Prompt, unattended note, harness and workflow are the same for every skill, so a fix there proves nothing about this one; every fix lands in the skill.
- An agent's template edit is reproduced before it is adopted. An edit is a claim, and a probe against the shipped template settles it: a real defect gets a template fix and a test that fails on the old code; an improvisation gets a sentence that forbids it.
- No round ships without the template check. The templates type-check and the documented test count holds under every runner the skill names, because a round's fix can break the templates it did not touch.
- Evals never bump a version. A version marks a change to the skill; a round with nothing to fix re-runs by dispatch, and score commits never ride a release commit.
- The stop rule is fixed before round one and followed to the end. At least three rounds, unanimous through round five, two of three after; moving it mid-loop would let a round's scores choose its own finish line.
Everything else is the target skill's: its rubric items 4 to 6, its template-check recipe, its commit convention.
The gh CLI signed in with rights to push and publish releases on the target repository, and a target that
already runs its evals on every published release. Every round pushes, tags and publishes, which runs billed
agent sessions, so the skill confirms once before the first round and then runs unattended.
| Not this | Use instead |
|---|---|
| Writing a new skill, or turning an app's module into one | extract-skill or skill-creator |
| Cutting a release without scoring evals | bumpv |
| Checking an app's code against its docs | code-audit |
| Changing the eval harness, prompts or workflow | The skills index repository, by a maintainer |
Issues and pull requests are welcome here. Pure markdown, with no build step, but the two Python scripts in
the references are checked: read_logs.py is run against a downloaded eval run, and extract_blocks.py
against a skill whose references carry // file: blocks. Claims in this skill are meant to be verifiable: if
you change a factual claim about agent behaviour or a CLI, say how you verified it, whether against the run
whose log showed it, the gh or agent CLI's own help, or a reproduction.
Adding, removing or renaming a file in references/ means updating the quick start and the reference
directory table in SKILL.md, the file table above, and any relative cross-links. references/provenance.md
is the ledger that must stay truthful: its figures are measured, and anything the recorded session did not
exercise is marked as added; add an entry for anything you change. Commits follow Conventional Commits and
releases follow STANDARD.md in the index;
CLAUDE.md carries the full editing conventions.
This is one of the Timerise Skills: modules for Next.js App
Router apps written by our own senior engineers from the modules they have shipped, not synthetic, each
published as its own repository and indexed there. They share one layout, so an agent that has read one knows
how to read the next: a SKILL.md entry point, references/ loaded on demand, and a seam contract carrying
the module's non-negotiables. This one is for skill maintainers, of these or of any other skill: it turns
a skill's eval results into fixes.
Built and maintained by Timerise.
MIT. See LICENSE.