Browse this collection on skills.sh, including per-skill security audits.
bunx skills add https://github.com/will-wright-eng/skillsThis command uses the vercel-labs/skills CLI to implement skills in this repo.
bunx skills updateUpdates all installed skills; pass a name to update one (bunx skills update design-readiness). The CLI has no version pinning — update re-fetches whatever is at the head of each source repo — so review the diff after updating.
Every skill in this repo installs standalone. Cross-skill file references (../other-skill/DOC.md) are an antipattern: the skills CLI installs skills individually (--skill <name>, skills use, direct skill URLs) with no dependency resolution, so a link into a sibling skill directory dangles unless the whole repo happens to be installed. Shared docs are instead vendored into each skill that needs them, marked with a provenance comment (<!-- Vendored verbatim from ... -->). Vendored copies are updated by re-copying their source, never by editing in place.
Three sequential skills for adding a verifiable autonomous experiment loop to a git repository, generalized from karpathy/autoresearch.
After install, invoke them in order: autoresearch-method → autoresearch-verify → autoresearch-program. Once program.md is generated, hand it to a fresh agent session and the loop runs from there.
| Skill | Purpose |
|---|---|
autoresearch-method |
Explain the methodology and evaluate whether the current repo is a good fit. |
autoresearch-verify |
Build a repo-specific verifier script with light (per-candidate metric) and heavy (integrity matrix) modes. |
autoresearch-program |
Generate program.md at the repo root — the operating directive a fresh agent session uses to run the loop. |
Self-contained skills covering the design-doc lifecycle: verify a design before building, distill what was built into ADRs, and keep the domain glossary sharp. Each writes only on explicit confirmation.
| Skill | Purpose |
|---|---|
design-readiness |
Audit a design/implementation/proposal doc for consistency with the codebase, ADRs, and CONTEXT.md, plus completeness of definition, then interview through drift fixes and open design decisions — interview answers are the approval, so the revised doc is applied once findings are resolved. |
distill-adrs |
Distill existing implementation docs (plans, design docs, RFCs) into ADRs — extracts candidate decisions, verifies each against the code, and confirms them one at a time before writing. |
create-context |
Build or refine a CONTEXT.md glossary — explores the codebase for candidate domain terms, then confirms each term, relationship, and ambiguity one at a time before writing. |
| Skill | Purpose |
|---|---|
anneal |
Carve a god module into stable and volatile pieces along evidence from git history — hotspot ranking via the hc CLI (raw-git fallback when absent), a three-axis autopsy of the target file, a seam-by-seam interview, and a strangler-fig migration plan with a measurable baseline. |
prune-comments |
Aggressively delete comment narration that restates the code — including "pseudo-why" comments whose reason is already visible and step-heading comments over blocks — and condense verbose why-comments and doc comments to terse technical language; doubt resolves toward deletion. Scoped to the diff against a base branch by default (path sweep on request), optionally driven under a target comment ratio, never touching semantic comments (directives, pragmas, license headers) or genuine why-comments, and modifying comments only: never code, never markdown. |
anneal vendors copies of improve-codebase-architecture's LANGUAGE.md (architectural vocabulary) and grill-with-docs's ADR-FORMAT.md — see Skill Self-Containment.
Copied verbatim from their source repos — replicating third-party skills (after reading them) reduces prompt-injection risk versus installing from a remote source that can change underneath you.
Each replicated skill is mapped to its upstream path in scripts/replicated-skills.json. bash scripts/check_replicated_skills.sh diffs every copy against the upstream ref and exits non-zero on drift; the drift workflow runs the same check on pull requests (commenting the diff on the PR), weekly, and on demand. The lint fixers are excluded from these directories so the copies stay byte-for-byte. To resync a skill, re-copy it from the commit the report links to and review the diff — never hand-edit a copy.
From mattpocock/skills.
| Skill | Purpose |
|---|---|
improve-codebase-architecture |
Surface deepening opportunities — refactors that turn shallow modules into deep ones, using a fixed architectural vocabulary. |
grill-with-docs |
Interview-style session that stress-tests a plan against the project's domain language and updates CONTEXT.md / ADRs inline as decisions crystallise. |
Upstream, improve-codebase-architecture linked to grill-with-docs for its CONTEXT.md and ADR format docs; this repo vendors copies of those docs into the skill instead (see Skill Self-Containment). The two repointed link paths are the only local deviation from the replicated source.
From mattpocock/skills.
| Skill | Purpose |
|---|---|
grill-me |
Relentless, one-question-at-a-time interview that stress-tests a plan or design before you build, recommending an answer for each decision and exploring the codebase when it can answer a question itself. |
From JuliusBrussee/caveman.
| Skill | Purpose |
|---|---|
caveman |
Ultra-compressed response mode — cuts token usage ~75% by stripping articles, filler, and hedging while keeping full technical accuracy. Supports lite / full / ultra and 文言文 (wenyan-*) intensity levels. |
From DietrichGebert/ponytail. The upstream plugin also ships Node lifecycle hooks for always-on activation and a statusline badge; only the skill is replicated here.
| Skill | Purpose |
|---|---|
ponytail |
Lazy senior dev mode — climbs a ladder (YAGNI → reuse → stdlib → native platform → installed dep → one line → minimum) before writing code, never cutting validation, error handling, security, or accessibility. Supports lite / full / ultra intensity levels. Pairs with caveman: ponytail governs the code, caveman governs the prose. |
From multica-ai/andrej-karpathy-skills.
| Skill | Purpose |
|---|---|
karpathy-guidelines |
Behavioral guidelines to reduce common LLM coding mistakes, derived from Andrej Karpathy's observations — think before coding, simplicity first, surgical changes, goal-driven execution. |
Checked against Karpathy's 2026 public statements as of 2026-08-31: no conflicts. His Sequoia Ascent talk (Aug 2026) still criticizes agent output as "bloated, copy-pasted, awkwardly abstracted, brittle," and his "agentic engineering" framing (spec design, eval design, diff review) maps onto the skill's goal-driven-execution guideline. The skill's caution bias reads slightly conservative next to his shift toward agent autonomy (~80% agent-written code, AutoResearch), but his answer to autonomy is verifiability, which is that same guideline. The source tweet (Jan 2026) postdates his vibe-coding-to-agentic-engineering shift.
From max-sixty/worktrunk. The reference/ docs the skill reads at runtime are copied alongside it so the skill installs standalone (see Skill Self-Containment); upstream syncs them from worktrunk.dev, so refresh by re-copying the whole directory. Requires the wt CLI.
| Skill | Purpose |
|---|---|
worktrunk |
Guidance for the wt git-worktree CLI — which worktree a command acts on, user vs. project config (~/.config/worktrunk/config.toml vs. .config/wt.toml), lifecycle hooks, LLM commit-message generation, aliases, and escalating hook approvals to the user rather than passing --yes. |
Skills in this repo handle task-level behavior; global preferences live in ~/.claude/CLAUDE.md, which Claude Code applies to every project. That file is managed with gists3 (g3), an S3-inspired CLI that treats a GitHub gist as a bucket and its files as keys — the canonical copy lives in this gist. g3 link creates the local working copy, and g3 push / g3 pull sync edits with guards against overwriting unseen remote changes, giving the file free versioned storage without a full dotfiles repo.