Skip to content

Replace the two skills with seven smaller ones, add evals, and drop the downstream sync - #10

Merged
bcherry merged 13 commits into
mainfrom
bcherry/skills-refresh
Sep 24, 2026
Merged

bcherry merged 13 commits into
mainfrom
bcherry/skills-refresh

Conversation

@bcherry

@bcherry bcherry commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Replaces the two monolithic skills with seven smaller ones, one for each job in the life of an agent, adds tooling to check them, and removes the downstream sync.

Why

livekit-simulations was written for the pre-June CLI. Its run command (lk agent simulate --scenarios …) now prints help and exits 0, the bundled build_scenarios.py produced the old agent_description/metadata schema, and its beta notice was wrong on every point. livekit-agents told agents to write tests but never showed them how, and much of it was MCP setup boilerplate. Both are deleted here, not stubbed: a stub's description would sit in every user's context forever to say "don't use me".

The new skills

Stage Skill Owns
Know reading-livekit-docs Looking up any LiveKit fact via the Docs MCP server or lk docs; the others load it first
Build building-livekit-agents Architecture, voice-specific design, and keeping the model in charge of meaning while code owns state, approvals and effects
Try debugging-livekit-agents Live-testing during development with lk agent debugger; the default for a bare "test my agent"
Test testing-livekit-agents Turn-level pytest/Vitest tests; three kinds of evidence, kept distinct
Simulate writing-livekit-scenarios Authoring scenarios and the agent-side code that consumes them
Ship running-livekit-simulations Text vs audio, CI, triage
Operate operating-livekit-agents Deploying to LiveKit Cloud and running in production: worker model, prewarm, shutdown, upgrades, observability

They describe how to approach each job and send the agent to --help and the live docs for flags, schemas and API names, so they don't go stale when the CLI or SDK changes. Descriptions are third person, written as a set, and each names the sibling to use for adjacent work, with a negative clause where two skills border. Bodies are 900 lines across the seven, plus 420 in four reference files.

Two of them carry other people's work. The state-and-effects material in building (never classify intent with regex, omission-preserves mutation semantics, approval as a later real user message, commit-then-publish) is ported from @ShayneP's shaynep/skill-refresh, minus the benchmark framing, version-pinned SDK facts, and the 1,500-line teaching app. The production-operations half of operating (worker process model, prewarming, provider lifecycle, shutdown, upgrades, observability) is adapted from @rathoresids's #7, co-authored in the commit; deployment is new. Closes #6.

The sync is gone

Skill installation is moving into the LiveKit CLI, so this PR deletes trigger-skill-sync.yml. The receiving workflow is removed in livekit/internal-actions#16, and the stale livekit-agents copies it had placed downstream are removed in agent-starter-python#106, agent-starter-node#66, and cloud-api-server#2196. None of these depend on merge order. SYNC_DISPATCH_TOKEN and TARGET_REPO become unused.

Docs

  • README rewritten for users: who the skills are for, the lifecycle above, example requests that trigger each skill, and what to install.
  • AGENTS.md rewritten for agents maintaining the repo: how routing works, the authoring rules with their reasons, the order to verify facts in, how to run the evals, what counts as done, and how to remove a skill.
  • CONTRIBUTING and the PR template updated to match.

Evals

  • evals/validate.py — structural checks on every skill; exits non-zero. .github/workflows/validate.yml runs it on every PR and push to main, along with a compile check on the eval tooling — the repo's only CI, since the sync workflow is gone.
  • evals/trigger/ — does the right skill fire? Runs 50 labeled queries through claude -p against a credential-stripped clone of either starter template, and reports a confusion table.
  • evals/output/ — structural checks on generated scenario files, plus an LLM judge (no tools, evidence quoted) for the questions that need judgment.

Fixtures are git clones, not lk agent init, which resolves a Cloud project and writes its credentials. Every eval subprocess also gets dummy LIVEKIT_* variables, which lk honours over the configured default project, so an eval can't spend inference or upload code to a real project. The deciding model in every eval is the user's Claude.

Simulation content was verified against the in-progress simulations docs, since the live docs are behind the CLI.

Merge notes

Checklist

  • Content passes the "freeze forever" test: no flag rosters, schema listings, API surfaces, or version numbers; --help and the docs are referenced instead
  • No time-relative phrasing ("newer CLIs", "as of writing", "is landing")
  • Descriptions are third person, under 1,024 characters, carry real trigger phrases, and name sibling skills for adjacent jobs
  • python3 evals/validate.py passes
  • Every description change was followed by a trigger eval run with no misses (below)
  • Output eval run with three prompts on the scenario skill (below)

Eval results

All runs on Fable 5.1 with the whole set installed together.

Trigger — full set

39 queries at the time (6 should trigger nothing), 2 runs each, both fixtures, before and after a copy-edit pass that reworded every description:

fixture before copy-edit after
Python 78/78 78/78
Node 78/78 78/78

Identical confusion tables before and after. The only miss ever recorded was in the very first run — a bare "test my agent" going to running (76/78) — fixed by having debugging claim the phrase and running disclaim it.

Trigger — after each later description change

Each run is the affected group plus the neighbors it could steal from, 3 runs per query:

change queries × runs result
merged connecting into writing 16 × 3 48/48
running gains "regression test before deploying"; paired query for testing 13 × 3 39/39
building gains the state-and-effects phrases 7 × 3 21/21
new operating skill, checked against building/debugging/running 16 × 3 48/48
testing description loses its API identifiers 6 × 3 18/18

Query set is now 50, including 7 should-not-trigger near-misses.

Output

3 prompts, each run with and without the skill on the starter agent, graded by the committed judge:

run structural judged note
broad suite, with skill (22 scenarios) 6/6 6/6
guardrail focus, with skill (20) 6/6 6/6
guardrail focus, baseline (17) 6/6 5/6 fails expectations_are_decidable: "'Stays warm rather than preachy' — tone judgments a judge could score inconsistently"

Caveat: both starter templates already ship a scenarios.yaml in the style the skill teaches, so baseline runs come out nearly as good as skill runs. The judge's decidability finding was the one check that separated them.

The two original skills had drifted from the shipped surface: livekit-simulations
targeted the pre-June CLI (flat `lk agent simulate`, `agent_description`/`metadata`
schema) and its beta notice was stale on every point; livekit-agents spent much of
its length on MCP boilerplate and mandated tests it never taught.

Replace them with one skill per job in the development loop:

  reading-livekit-docs          fact lookup via Docs MCP / `lk docs`; loaded first by the rest
  building-livekit-agents       architecture: latency, context, handoffs and tasks
  debugging-livekit-agents      live-testing via `lk agent debugger`
  testing-livekit-agents        turn-level pytest/Vitest tests
  writing-livekit-scenarios     authoring scenarios, plus the agent-side wiring
                                (seeding userdata, session-scoped mocks, final-state grading)
  running-livekit-simulations   text vs audio, CI, triage

Skills are conceptual on purpose: they teach the shape of each job and send the
agent to `--help` and the live docs for flags, schemas and API names, so they stay
correct as the CLI and SDK move. Descriptions are third person, written as a set
with explicit hand-offs, and verified with a trigger-collision eval (76/78, then
100% after claiming a bare "test my agent" for debugging).

The old skills remain as DEPRECATED stubs pointing at their replacements; their
references/ and scripts/ (which emitted the superseded schema) are removed.
README: lead with who the skills are for and what changes when installed, frame
the six skills as a Know → Build → Try → Test → Simulate → Ship lifecycle, give
real trigger phrases per skill, state the requirements (`lk`, a Cloud project,
optionally the Docs MCP server), and explain why the skills are conceptual.

AGENTS.md: a working guide for agents maintaining this repo — how the set routes,
the authoring rules with their reasons, the fact-verification order (`--help`,
docs, the public repos' open PRs, examples), the three-tier evaluation process,
a definition of done, and the deprecation procedure with the downstream-sync
rationale. Names only public repos.

CONTRIBUTING: fix the stale reference to the deprecated skill, align naming and
frontmatter, defer detail to AGENTS.md. PR template: one file, a checklist that
can actually be checked, and a place for eval results.
evals/validate.py      exact checks on every SKILL.md — frontmatter, third-person
                       description under 1,024 chars, body under 500 lines,
                       references present with a TOC when long, cross-references
                       resolve, no time-relative phrasing. Exit code for CI.

evals/trigger/         does the right skill fire? Installs the non-deprecated set
                       into a throwaway fixture, runs each labeled query through
                       `claude -p`, reports a confusion table and misses. The
                       fixture is a shallow clone of agent-starter-python or
                       agent-starter-node — not `lk agent init`, which resolves a
                       Cloud project and writes its credentials. `.env*` files are
                       stripped and dummy LIVEKIT_* variables are injected, which
                       `lk` honours over the configured default project, so an eval
                       cannot spend or upload against the user's project. Tools are
                       read-only. 39 queries, parametrized per language.

evals/output/          does following the skill produce better work? A grader with
                       exact structural checks (schema against the CLI's scenario
                       struct) and an LLM judge for the judgment calls — coverage of
                       the agent's constraints, refusal-shaped expectations, invented
                       capabilities, rotting dates, decidable expectations — with
                       quoted evidence per verdict. Replaces keyword heuristics that
                       false-negatived on good work.

Eval results, logs and scratch fixtures are gitignored.
The stubs were insurance against a downstream sync whose deletion behavior
couldn't be verified. It has been now: the sync in livekit/internal-actions
copied a hardcoded list of two skill names, so it would neither remove the old
skills nor pick up the new ones. That workflow is being changed to mirror this
directory (discover skills from skills/, remove downstream copies of skills that
no longer exist here), which makes the stubs unnecessary — and a stub's
description would otherwise sit in every user's context on every request.

AGENTS.md's deprecation policy now says to delete, and why. README tells users
of the old skills to reinstall.
@bcherry bcherry changed the title Refresh skills: six atomic, conceptual skills plus eval tooling Replace the two skills with six smaller ones and add evals Sep 23, 2026
Comment thread skills/running-livekit-simulations/SKILL.md Outdated
Comment thread skills/running-livekit-simulations/SKILL.md Outdated
Comment thread skills/running-livekit-simulations/SKILL.md
Frame simulations positively: for regression-testing long-horizon behavior
before deploying to production, in one sentence, rather than leading with cost
in a way that made them never seem like the right option.

Drop the "what goes wrong in automation" list. If it goes wrong the coding
agent will find and fix it; the one design fact worth keeping (every committed
scenario must pass, so keep aspirational ones in a separate file) is now a
sentence in the automation paragraph.

Reframe reading results around the coding agent: the dashboard link is for the
human, export is the agent's path to a failing transcript.
Skill installation is moving into the LiveKit CLI, so skills will no longer be
pushed into the starter templates and fixtures on every merge. Delete the
dispatcher; the receiving workflow is removed in livekit/internal-actions#16.

AGENTS.md no longer reasons about skill removal in terms of the sync. The
skills/<name>/SKILL.md layout is the install contract installers depend on, and
nothing in this repo can reach a user's existing local copy.

The SYNC_DISPATCH_TOKEN secret and TARGET_REPO variable are now unused.
Comment thread skills/running-livekit-simulations/SKILL.md Outdated
Comment thread evals/output/grade_scenarios.py Outdated
Comment thread evals/output/grade_scenarios.py
running-livekit-simulations' body leads with regression-testing before
deployment, but its description never said so. Add "regression test my agent
before deploying" as a trigger, with two paired queries: that phrase routes to
running, "add a regression test for the bug where..." routes to testing.
Trigger eval on the affected groups, 13 queries x 3: 39/39.

The judge grader coerced verdicts with bool(), so a string "false" from the
model counted as a pass. Coerce only real booleans or unambiguous words;
anything else is recorded as unparseable and excluded from the majority rather
than counted either way. The example schema in the prompt showed the string
"true|false", which invited the problem; it now shows a JSON boolean. Verified
on a real judge run: six boolean verdicts, none unparseable.
Comment thread skills/testing-livekit-agents/SKILL.md Outdated

@Topherhindman Topherhindman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

one heads up before this merges: @ShayneP has been working in a nearby area on shaynep/skill-refresh. He rewrote livekit-agents around conversation and state contracts, and added example code and tests under assets/. this PR deletes livekit-agents, so there will be conflicts between the branches whenever one or the other merges. And possibly more importantly, some of the content overlaps. y'all may want to chat quickly

Port the engineering substance of Shayne Parker's livekit-agents rewrite
(shaynep/skill-refresh) into the skill that replaced it. Three new sections in
building-livekit-agents — never classify intent with code, make every change
mean exactly one thing, start with a failing complete-path test — and a
reference, state-and-effects.md, covering precise mutation semantics, review
and approval as separate runtime events, commit-then-publish, output ownership
and closing, input-mode hook differences, and repairing the first divergence.
testing-livekit-agents gains the three-kinds-of-evidence distinction and the
list of shortcuts that look like tests and aren't.

Deliberately not ported: the task-contract/benchmark framing, version-pinned
SDK facts and API names (kept conceptual per AGENTS.md), the 1,500-line
teaching application and its tests (a worked example belongs in livekit/agents
examples, versioned against the SDK), and the --experimental-auth CLI usage.

Description gains three trigger phrases for these failure modes. Trigger eval
on the building group plus two new paired queries, 7 x 3: 21/21.
bcherry and others added 2 commits September 23, 2026 17:08
The seventh job in the loop, after Ship: getting a version onto LiveKit Cloud
and keeping it healthy. Covers the deploy lifecycle conceptually (a directory
bound to an agent, container build or prebuilt image, versions, secrets as
environment, rollback, build vs deploy logs), the worker process model and
prewarming, safe async inside worker processes and how those bugs present,
provider timeouts and degradation, measure-first performance work and
endpointing tuning, drain on shutdown matched to the orchestrator, SDK
upgrades, observability, and how to change a codebase that is already live.

The worker-model, provider, performance, operations, and existing-codebase
material is adapted from Siddharth Rathod's livekit-agents-production skill
(#7, addressing #6), restructured to this set's conventions: gerund name,
third-person description with triggers disjoint from building and debugging,
no restated flags or SDK identifiers, the duplicated latency/context/listening
section dropped in favor of building-livekit-agents. Deployment is new. The
worker-model claims were checked against both SDKs: each exposes a prewarm
hook on a per-job process object.

README and AGENTS.md gain the seventh row; building, debugging and running
link to it. Trigger eval, 16 queries x 3 covering the new group and the
neighbors it could steal from: 48/48.

Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>
testing-livekit-agents' body had been made conceptual but its description still
named AgentSession, judge() and JudgeGroup. Replace them with the concepts
("the SDK's test session harness", "LLM judging of a reply against an
intent", "the built-in judges"). Trigger eval on the testing group, 6 x 3: 18/18.

The docs said "--judge-runs 2 or 3", but the grader's majority is a strict
greater-than, so a 1-1 split from two runs reads as a fail. Say 3 everywhere,
explain why in the argparse help, and warn when an even count is passed.
@bcherry bcherry changed the title Replace the two skills with six smaller ones and add evals Replace the two skills with seven smaller ones, add evals, and drop the downstream sync Sep 24, 2026
Run evals/validate.py on every pull request and on pushes to main, plus a
compile check on the eval tooling and a parse check on the trigger query set.
This is the structural half of AGENTS.md's definition of done, enforced where
the PR template already asks for it. The trigger and output evals stay local:
they need a logged-in `claude` and spend tokens.

It also clears the red check on this branch. The repo's CodeQL default setup
scans GitHub Actions, and removing the sync left the branch with no workflow
files, so the actions analysis failed with "CodeQL could not process any code
written in GitHub Actions".
Comment thread .github/workflows/validate.yml Fixed
CodeQL's actions/missing-workflow-permissions rule: a workflow that doesn't
declare permissions gets the repository's default GITHUB_TOKEN scope. This job
only reads the checkout, so grant contents: read and nothing else.
@bcherry
bcherry merged commit 4305938 into main Sep 24, 2026
5 checks passed
@bcherry
bcherry deleted the bcherry/skills-refresh branch September 24, 2026 05:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Separate skill for contributing to existing LiveKit agent codebases

4 participants