Replace the two skills with seven smaller ones, add evals, and drop the downstream sync - #10
Merged
Merged
Conversation
The two original skills had drifted from the shipped surface: livekit-simulations
targeted the pre-June CLI (flat `lk agent simulate`, `agent_description`/`metadata`
schema) and its beta notice was stale on every point; livekit-agents spent much of
its length on MCP boilerplate and mandated tests it never taught.
Replace them with one skill per job in the development loop:
reading-livekit-docs fact lookup via Docs MCP / `lk docs`; loaded first by the rest
building-livekit-agents architecture: latency, context, handoffs and tasks
debugging-livekit-agents live-testing via `lk agent debugger`
testing-livekit-agents turn-level pytest/Vitest tests
writing-livekit-scenarios authoring scenarios, plus the agent-side wiring
(seeding userdata, session-scoped mocks, final-state grading)
running-livekit-simulations text vs audio, CI, triage
Skills are conceptual on purpose: they teach the shape of each job and send the
agent to `--help` and the live docs for flags, schemas and API names, so they stay
correct as the CLI and SDK move. Descriptions are third person, written as a set
with explicit hand-offs, and verified with a trigger-collision eval (76/78, then
100% after claiming a bare "test my agent" for debugging).
The old skills remain as DEPRECATED stubs pointing at their replacements; their
references/ and scripts/ (which emitted the superseded schema) are removed.
README: lead with who the skills are for and what changes when installed, frame the six skills as a Know → Build → Try → Test → Simulate → Ship lifecycle, give real trigger phrases per skill, state the requirements (`lk`, a Cloud project, optionally the Docs MCP server), and explain why the skills are conceptual. AGENTS.md: a working guide for agents maintaining this repo — how the set routes, the authoring rules with their reasons, the fact-verification order (`--help`, docs, the public repos' open PRs, examples), the three-tier evaluation process, a definition of done, and the deprecation procedure with the downstream-sync rationale. Names only public repos. CONTRIBUTING: fix the stale reference to the deprecated skill, align naming and frontmatter, defer detail to AGENTS.md. PR template: one file, a checklist that can actually be checked, and a place for eval results.
evals/validate.py exact checks on every SKILL.md — frontmatter, third-person
description under 1,024 chars, body under 500 lines,
references present with a TOC when long, cross-references
resolve, no time-relative phrasing. Exit code for CI.
evals/trigger/ does the right skill fire? Installs the non-deprecated set
into a throwaway fixture, runs each labeled query through
`claude -p`, reports a confusion table and misses. The
fixture is a shallow clone of agent-starter-python or
agent-starter-node — not `lk agent init`, which resolves a
Cloud project and writes its credentials. `.env*` files are
stripped and dummy LIVEKIT_* variables are injected, which
`lk` honours over the configured default project, so an eval
cannot spend or upload against the user's project. Tools are
read-only. 39 queries, parametrized per language.
evals/output/ does following the skill produce better work? A grader with
exact structural checks (schema against the CLI's scenario
struct) and an LLM judge for the judgment calls — coverage of
the agent's constraints, refusal-shaped expectations, invented
capabilities, rotting dates, decidable expectations — with
quoted evidence per verdict. Replaces keyword heuristics that
false-negatived on good work.
Eval results, logs and scratch fixtures are gitignored.
The stubs were insurance against a downstream sync whose deletion behavior couldn't be verified. It has been now: the sync in livekit/internal-actions copied a hardcoded list of two skill names, so it would neither remove the old skills nor pick up the new ones. That workflow is being changed to mirror this directory (discover skills from skills/, remove downstream copies of skills that no longer exist here), which makes the stubs unnecessary — and a stub's description would otherwise sit in every user's context on every request. AGENTS.md's deprecation policy now says to delete, and why. README tells users of the old skills to reinstall.
u9g
approved these changes
Sep 23, 2026
Frame simulations positively: for regression-testing long-horizon behavior before deploying to production, in one sentence, rather than leading with cost in a way that made them never seem like the right option. Drop the "what goes wrong in automation" list. If it goes wrong the coding agent will find and fix it; the one design fact worth keeping (every committed scenario must pass, so keep aspirational ones in a separate file) is now a sentence in the automation paragraph. Reframe reading results around the coding agent: the dashboard link is for the human, export is the agent's path to a failing transcript.
Skill installation is moving into the LiveKit CLI, so skills will no longer be pushed into the starter templates and fixtures on every merge. Delete the dispatcher; the receiving workflow is removed in livekit/internal-actions#16. AGENTS.md no longer reasons about skill removal in terms of the sync. The skills/<name>/SKILL.md layout is the install contract installers depend on, and nothing in this repo can reach a user's existing local copy. The SYNC_DISPATCH_TOKEN secret and TARGET_REPO variable are now unused.
This was referenced Sep 23, 2026
running-livekit-simulations' body leads with regression-testing before deployment, but its description never said so. Add "regression test my agent before deploying" as a trigger, with two paired queries: that phrase routes to running, "add a regression test for the bug where..." routes to testing. Trigger eval on the affected groups, 13 queries x 3: 39/39. The judge grader coerced verdicts with bool(), so a string "false" from the model counted as a pass. Coerce only real booleans or unambiguous words; anything else is recorded as unparseable and excluded from the majority rather than counted either way. The example schema in the prompt showed the string "true|false", which invited the problem; it now shows a JSON boolean. Verified on a real judge run: six boolean verdicts, none unparseable.
Topherhindman
approved these changes
Sep 23, 2026
Topherhindman
left a comment
Collaborator
There was a problem hiding this comment.
one heads up before this merges: @ShayneP has been working in a nearby area on shaynep/skill-refresh. He rewrote livekit-agents around conversation and state contracts, and added example code and tests under assets/. this PR deletes livekit-agents, so there will be conflicts between the branches whenever one or the other merges. And possibly more importantly, some of the content overlaps. y'all may want to chat quickly
Port the engineering substance of Shayne Parker's livekit-agents rewrite (shaynep/skill-refresh) into the skill that replaced it. Three new sections in building-livekit-agents — never classify intent with code, make every change mean exactly one thing, start with a failing complete-path test — and a reference, state-and-effects.md, covering precise mutation semantics, review and approval as separate runtime events, commit-then-publish, output ownership and closing, input-mode hook differences, and repairing the first divergence. testing-livekit-agents gains the three-kinds-of-evidence distinction and the list of shortcuts that look like tests and aren't. Deliberately not ported: the task-contract/benchmark framing, version-pinned SDK facts and API names (kept conceptual per AGENTS.md), the 1,500-line teaching application and its tests (a worked example belongs in livekit/agents examples, versioned against the SDK), and the --experimental-auth CLI usage. Description gains three trigger phrases for these failure modes. Trigger eval on the building group plus two new paired queries, 7 x 3: 21/21.
4 of 5 tasks
The seventh job in the loop, after Ship: getting a version onto LiveKit Cloud and keeping it healthy. Covers the deploy lifecycle conceptually (a directory bound to an agent, container build or prebuilt image, versions, secrets as environment, rollback, build vs deploy logs), the worker process model and prewarming, safe async inside worker processes and how those bugs present, provider timeouts and degradation, measure-first performance work and endpointing tuning, drain on shutdown matched to the orchestrator, SDK upgrades, observability, and how to change a codebase that is already live. The worker-model, provider, performance, operations, and existing-codebase material is adapted from Siddharth Rathod's livekit-agents-production skill (#7, addressing #6), restructured to this set's conventions: gerund name, third-person description with triggers disjoint from building and debugging, no restated flags or SDK identifiers, the duplicated latency/context/listening section dropped in favor of building-livekit-agents. Deployment is new. The worker-model claims were checked against both SDKs: each exposes a prewarm hook on a per-job process object. README and AGENTS.md gain the seventh row; building, debugging and running link to it. Trigger eval, 16 queries x 3 covering the new group and the neighbors it could steal from: 48/48. Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>
testing-livekit-agents' body had been made conceptual but its description still
named AgentSession, judge() and JudgeGroup. Replace them with the concepts
("the SDK's test session harness", "LLM judging of a reply against an
intent", "the built-in judges"). Trigger eval on the testing group, 6 x 3: 18/18.
The docs said "--judge-runs 2 or 3", but the grader's majority is a strict
greater-than, so a 1-1 split from two runs reads as a fail. Say 3 everywhere,
explain why in the argparse help, and warn when an even count is passed.
This was referenced Sep 24, 2026
Run evals/validate.py on every pull request and on pushes to main, plus a compile check on the eval tooling and a parse check on the trigger query set. This is the structural half of AGENTS.md's definition of done, enforced where the PR template already asks for it. The trigger and output evals stay local: they need a logged-in `claude` and spend tokens. It also clears the red check on this branch. The repo's CodeQL default setup scans GitHub Actions, and removing the sync left the branch with no workflow files, so the actions analysis failed with "CodeQL could not process any code written in GitHub Actions".
CodeQL's actions/missing-workflow-permissions rule: a workflow that doesn't declare permissions gets the repository's default GITHUB_TOKEN scope. This job only reads the checkout, so grant contents: read and nothing else.
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Replaces the two monolithic skills with seven smaller ones, one for each job in the life of an agent, adds tooling to check them, and removes the downstream sync.
Why
livekit-simulationswas written for the pre-June CLI. Its run command (lk agent simulate --scenarios …) now prints help and exits 0, the bundledbuild_scenarios.pyproduced the oldagent_description/metadataschema, and its beta notice was wrong on every point.livekit-agentstold agents to write tests but never showed them how, and much of it was MCP setup boilerplate. Both are deleted here, not stubbed: a stub's description would sit in every user's context forever to say "don't use me".The new skills
reading-livekit-docslk docs; the others load it firstbuilding-livekit-agentsdebugging-livekit-agentslk agent debugger; the default for a bare "test my agent"testing-livekit-agentswriting-livekit-scenariosrunning-livekit-simulationsoperating-livekit-agentsThey describe how to approach each job and send the agent to
--helpand the live docs for flags, schemas and API names, so they don't go stale when the CLI or SDK changes. Descriptions are third person, written as a set, and each names the sibling to use for adjacent work, with a negative clause where two skills border. Bodies are 900 lines across the seven, plus 420 in four reference files.Two of them carry other people's work. The state-and-effects material in
building(never classify intent with regex, omission-preserves mutation semantics, approval as a later real user message, commit-then-publish) is ported from @ShayneP'sshaynep/skill-refresh, minus the benchmark framing, version-pinned SDK facts, and the 1,500-line teaching app. The production-operations half ofoperating(worker process model, prewarming, provider lifecycle, shutdown, upgrades, observability) is adapted from @rathoresids's #7, co-authored in the commit; deployment is new. Closes #6.The sync is gone
Skill installation is moving into the LiveKit CLI, so this PR deletes
trigger-skill-sync.yml. The receiving workflow is removed in livekit/internal-actions#16, and the stalelivekit-agentscopies it had placed downstream are removed in agent-starter-python#106, agent-starter-node#66, and cloud-api-server#2196. None of these depend on merge order.SYNC_DISPATCH_TOKENandTARGET_REPObecome unused.Docs
Evals
evals/validate.py— structural checks on every skill; exits non-zero..github/workflows/validate.ymlruns it on every PR and push tomain, along with a compile check on the eval tooling — the repo's only CI, since the sync workflow is gone.evals/trigger/— does the right skill fire? Runs 50 labeled queries throughclaude -pagainst a credential-stripped clone of either starter template, and reports a confusion table.evals/output/— structural checks on generated scenario files, plus an LLM judge (no tools, evidence quoted) for the questions that need judgment.Fixtures are
git clones, notlk agent init, which resolves a Cloud project and writes its credentials. Every eval subprocess also gets dummyLIVEKIT_*variables, whichlkhonours over the configured default project, so an eval can't spend inference or upload code to a real project. The deciding model in every eval is the user's Claude.Simulation content was verified against the in-progress simulations docs, since the live docs are behind the CLI.
Merge notes
7f31978,f526f65,dde2eb1; the two threads still marked open are the ones fixed in the last of those.operating-livekit-agentscontains substantial content from Add livekit-agents-production skill for existing codebases #7, so its CLA needs to be signed before this merges. Add livekit-agents-production skill for existing codebases #7 can close in favor of this.Analyze (actions)check went red when the sync workflow was deleted (no Actions code left to scan); the validate workflow gives it something to scan again.Checklist
--helpand the docs are referenced insteadpython3 evals/validate.pypassesdescriptionchange was followed by a trigger eval run with no misses (below)Eval results
All runs on Fable 5.1 with the whole set installed together.
Trigger — full set
39 queries at the time (6 should trigger nothing), 2 runs each, both fixtures, before and after a copy-edit pass that reworded every description:
Identical confusion tables before and after. The only miss ever recorded was in the very first run — a bare "test my agent" going to
running(76/78) — fixed by havingdebuggingclaim the phrase andrunningdisclaim it.Trigger — after each later description change
Each run is the affected group plus the neighbors it could steal from, 3 runs per query:
connectingintowritingrunninggains "regression test before deploying"; paired query fortestingbuildinggains the state-and-effects phrasesoperatingskill, checked against building/debugging/runningtestingdescription loses its API identifiersQuery set is now 50, including 7 should-not-trigger near-misses.
Output
3 prompts, each run with and without the skill on the starter agent, graded by the committed judge:
expectations_are_decidable: "'Stays warm rather than preachy' — tone judgments a judge could score inconsistently"Caveat: both starter templates already ship a
scenarios.yamlin the style the skill teaches, so baseline runs come out nearly as good as skill runs. The judge's decidability finding was the one check that separated them.