Add livekit-agents-production skill for existing codebases - #7
Closed
rathoresids wants to merge 1 commit into
Closed
rathoresids wants to merge 1 commit into
rathoresids wants to merge 1 commit into
Conversation
Adds a new skill focused on extending, debugging, and operating existing LiveKit agent codebases. Complements the existing livekit-agents skill (which covers building from scratch) with production-stage guidance: worker process model, safe async patterns, provider lifecycle, performance profiling, testing within established infrastructure, and production operations.
|
|
|
Any news on this? |
Contributor
|
Sorry for the long wait on this one but I'm incorporating a lot of the ideas here into new skills in #10 |
bcherry
added a commit
that referenced
this pull request
Sep 24, 2026
The seventh job in the loop, after Ship: getting a version onto LiveKit Cloud and keeping it healthy. Covers the deploy lifecycle conceptually (a directory bound to an agent, container build or prebuilt image, versions, secrets as environment, rollback, build vs deploy logs), the worker process model and prewarming, safe async inside worker processes and how those bugs present, provider timeouts and degradation, measure-first performance work and endpointing tuning, drain on shutdown matched to the orchestrator, SDK upgrades, observability, and how to change a codebase that is already live. The worker-model, provider, performance, operations, and existing-codebase material is adapted from Siddharth Rathod's livekit-agents-production skill (#7, addressing #6), restructured to this set's conventions: gerund name, third-person description with triggers disjoint from building and debugging, no restated flags or SDK identifiers, the duplicated latency/context/listening section dropped in favor of building-livekit-agents. Deployment is new. The worker-model claims were checked against both SDKs: each exposes a prewarm hook on a per-job process object. README and AGENTS.md gain the seventh row; building, debugging and running link to it. Trigger eval, 16 queries x 3 covering the new group and the neighbors it could steal from: 48/48. Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>
6 tasks done
bcherry
added a commit
that referenced
this pull request
Sep 24, 2026
…he downstream sync (#10) * Split monolithic skills into six atomic, conceptual skills The two original skills had drifted from the shipped surface: livekit-simulations targeted the pre-June CLI (flat `lk agent simulate`, `agent_description`/`metadata` schema) and its beta notice was stale on every point; livekit-agents spent much of its length on MCP boilerplate and mandated tests it never taught. Replace them with one skill per job in the development loop: reading-livekit-docs fact lookup via Docs MCP / `lk docs`; loaded first by the rest building-livekit-agents architecture: latency, context, handoffs and tasks debugging-livekit-agents live-testing via `lk agent debugger` testing-livekit-agents turn-level pytest/Vitest tests writing-livekit-scenarios authoring scenarios, plus the agent-side wiring (seeding userdata, session-scoped mocks, final-state grading) running-livekit-simulations text vs audio, CI, triage Skills are conceptual on purpose: they teach the shape of each job and send the agent to `--help` and the live docs for flags, schemas and API names, so they stay correct as the CLI and SDK move. Descriptions are third person, written as a set with explicit hand-offs, and verified with a trigger-collision eval (76/78, then 100% after claiming a bare "test my agent" for debugging). The old skills remain as DEPRECATED stubs pointing at their replacements; their references/ and scripts/ (which emitted the superseded schema) are removed. * Rewrite repo docs for users, contributors, and maintaining agents README: lead with who the skills are for and what changes when installed, frame the six skills as a Know → Build → Try → Test → Simulate → Ship lifecycle, give real trigger phrases per skill, state the requirements (`lk`, a Cloud project, optionally the Docs MCP server), and explain why the skills are conceptual. AGENTS.md: a working guide for agents maintaining this repo — how the set routes, the authoring rules with their reasons, the fact-verification order (`--help`, docs, the public repos' open PRs, examples), the three-tier evaluation process, a definition of done, and the deprecation procedure with the downstream-sync rationale. Names only public repos. CONTRIBUTING: fix the stale reference to the deprecated skill, align naming and frontmatter, defer detail to AGENTS.md. PR template: one file, a checklist that can actually be checked, and a place for eval results. * Add eval tooling: structure validator, trigger harness, output grader evals/validate.py exact checks on every SKILL.md — frontmatter, third-person description under 1,024 chars, body under 500 lines, references present with a TOC when long, cross-references resolve, no time-relative phrasing. Exit code for CI. evals/trigger/ does the right skill fire? Installs the non-deprecated set into a throwaway fixture, runs each labeled query through `claude -p`, reports a confusion table and misses. The fixture is a shallow clone of agent-starter-python or agent-starter-node — not `lk agent init`, which resolves a Cloud project and writes its credentials. `.env*` files are stripped and dummy LIVEKIT_* variables are injected, which `lk` honours over the configured default project, so an eval cannot spend or upload against the user's project. Tools are read-only. 39 queries, parametrized per language. evals/output/ does following the skill produce better work? A grader with exact structural checks (schema against the CLI's scenario struct) and an LLM judge for the judgment calls — coverage of the agent's constraints, refusal-shaped expectations, invented capabilities, rotting dates, decidable expectations — with quoted evidence per verdict. Replaces keyword heuristics that false-negatived on good work. Eval results, logs and scratch fixtures are gitignored. * Copy-edit pass across skills, docs, and eval tooling * Remove the deprecated stub skills The stubs were insurance against a downstream sync whose deletion behavior couldn't be verified. It has been now: the sync in livekit/internal-actions copied a hardcoded list of two skill names, so it would neither remove the old skills nor pick up the new ones. That workflow is being changed to mirror this directory (discover skills from skills/, remove downstream copies of skills that no longer exist here), which makes the stubs unnecessary — and a stub's description would otherwise sit in every user's context on every request. AGENTS.md's deprecation policy now says to delete, and why. README tells users of the old skills to reinstall. * running-livekit-simulations: address review feedback Frame simulations positively: for regression-testing long-horizon behavior before deploying to production, in one sentence, rather than leading with cost in a way that made them never seem like the right option. Drop the "what goes wrong in automation" list. If it goes wrong the coding agent will find and fix it; the one design fact worth keeping (every committed scenario must pass, so keep aspirational ones in a separate file) is now a sentence in the automation paragraph. Reframe reading results around the coding agent: the dashboard link is for the human, export is the agent's path to a failing transcript. * Remove the downstream skill sync Skill installation is moving into the LiveKit CLI, so skills will no longer be pushed into the starter templates and fixtures on every merge. Delete the dispatcher; the receiving workflow is removed in livekit/internal-actions#16. AGENTS.md no longer reasons about skill removal in terms of the sync. The skills/<name>/SKILL.md layout is the install contract installers depend on, and nothing in this repo can reach a user's existing local copy. The SYNC_DISPATCH_TOKEN secret and TARGET_REPO variable are now unused. * Address review: regression-test trigger phrase; judge verdict parsing running-livekit-simulations' body leads with regression-testing before deployment, but its description never said so. Add "regression test my agent before deploying" as a trigger, with two paired queries: that phrase routes to running, "add a regression test for the bug where..." routes to testing. Trigger eval on the affected groups, 13 queries x 3: 39/39. The judge grader coerced verdicts with bool(), so a string "false" from the model counted as a pass. Coerce only real booleans or unambiguous words; anything else is recorded as unparseable and excluded from the majority rather than counted either way. The example schema in the prompt showed the string "true|false", which invited the problem; it now shows a JSON boolean. Verified on a real judge run: six boolean verdicts, none unparseable. * building: the model interprets, code owns state and effects Port the engineering substance of Shayne Parker's livekit-agents rewrite (shaynep/skill-refresh) into the skill that replaced it. Three new sections in building-livekit-agents — never classify intent with code, make every change mean exactly one thing, start with a failing complete-path test — and a reference, state-and-effects.md, covering precise mutation semantics, review and approval as separate runtime events, commit-then-publish, output ownership and closing, input-mode hook differences, and repairing the first divergence. testing-livekit-agents gains the three-kinds-of-evidence distinction and the list of shortcuts that look like tests and aren't. Deliberately not ported: the task-contract/benchmark framing, version-pinned SDK facts and API names (kept conceptual per AGENTS.md), the 1,500-line teaching application and its tests (a worked example belongs in livekit/agents examples, versioned against the SDK), and the --experimental-auth CLI usage. Description gains three trigger phrases for these failure modes. Trigger eval on the building group plus two new paired queries, 7 x 3: 21/21. * Add operating-livekit-agents: deploy and run an agent in production The seventh job in the loop, after Ship: getting a version onto LiveKit Cloud and keeping it healthy. Covers the deploy lifecycle conceptually (a directory bound to an agent, container build or prebuilt image, versions, secrets as environment, rollback, build vs deploy logs), the worker process model and prewarming, safe async inside worker processes and how those bugs present, provider timeouts and degradation, measure-first performance work and endpointing tuning, drain on shutdown matched to the orchestrator, SDK upgrades, observability, and how to change a codebase that is already live. The worker-model, provider, performance, operations, and existing-codebase material is adapted from Siddharth Rathod's livekit-agents-production skill (#7, addressing #6), restructured to this set's conventions: gerund name, third-person description with triggers disjoint from building and debugging, no restated flags or SDK identifiers, the duplicated latency/context/listening section dropped in favor of building-livekit-agents. Deployment is new. The worker-model claims were checked against both SDKs: each exposes a prewarm hook on a per-job process object. README and AGENTS.md gain the seventh row; building, debugging and running link to it. Trigger eval, 16 queries x 3 covering the new group and the neighbors it could steal from: 48/48. Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com> * Address review: no identifiers in testing's description; odd judge-runs testing-livekit-agents' body had been made conceptual but its description still named AgentSession, judge() and JudgeGroup. Replace them with the concepts ("the SDK's test session harness", "LLM judging of a reply against an intent", "the built-in judges"). Trigger eval on the testing group, 6 x 3: 18/18. The docs said "--judge-runs 2 or 3", but the grader's majority is a strict greater-than, so a 1-1 split from two runs reads as a fail. Say 3 everywhere, explain why in the argparse help, and warn when an even count is passed. * Add a validate workflow; fixes the CodeQL actions check Run evals/validate.py on every pull request and on pushes to main, plus a compile check on the eval tooling and a parse check on the trigger query set. This is the structural half of AGENTS.md's definition of done, enforced where the PR template already asks for it. The trigger and output evals stay local: they need a logged-in `claude` and spend tokens. It also clears the red check on this branch. The repo's CodeQL default setup scans GitHub Actions, and removing the sync left the branch with no workflow files, so the actions analysis failed with "CodeQL could not process any code written in GitHub Actions". * validate workflow: least-privilege token permissions CodeQL's actions/missing-workflow-permissions rule: a workflow that doesn't declare permissions gets the repository's default GITHUB_TOKEN scope. This job only reads the checkout, so grant contents: read and nothing else. --------- Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Adds a new skill (
livekit-agents-production) focused on extending, debugging, and operating existing LiveKit agent codebases. This complements the existinglivekit-agentsskill (which covers building from scratch) with production-stage guidance.The new skill covers:
Addresses #6
Design decisions
livekit-agentsfor building,livekit-agents-productionfor extending and operatingChecklist