Skip to content

Add livekit-agents-production skill for existing codebases - #7

Closed
rathoresids wants to merge 1 commit into
livekit:mainfrom
rathoresids:improve-livekit-agents-skill
Closed

rathoresids wants to merge 1 commit into
livekit:mainfrom
rathoresids:improve-livekit-agents-skill

Conversation

@rathoresids

Copy link
Copy Markdown
Contributor

What does this PR do?

Adds a new skill (livekit-agents-production) focused on extending, debugging, and operating existing LiveKit agent codebases. This complements the existing livekit-agents skill (which covers building from scratch) with production-stage guidance.

The new skill covers:

  • Worker process model — parent-child architecture, shared state inheritance, and async context implications
  • Safe async patterns — avoiding the most common class of production bugs (event loop conflicts in worker processes)
  • Provider lifecycle — prewarming strategies, cold-start mitigation, failure handling
  • Testing within existing infrastructure — using established fixtures, mocking at boundaries, async test patterns
  • Performance profiling — what to measure (time-to-first-audio, tool execution, context growth) and common anti-patterns
  • Production operations — graceful shutdown, SDK upgrade strategy, observability, debugging common failures

Addresses #6

Design decisions

  • Separate skill, not a rewrite — keeps each skill focused: livekit-agents for building, livekit-agents-production for extending and operating
  • No code examples — all patterns described as behavioral principles per the "freeze forever" principle
  • Not Python-specific — worker model and async patterns framed universally, applicable to both Python and Node SDKs
  • 258 lines — well under the 500-line limit

Checklist

  • Content passes the "freeze forever" test — would it still be correct if never updated?
  • API specifics are directed to MCP/docs, not hardcoded in skill content
  • Skill stays under 500 lines
  • Trigger phrases in skill description are clear and representative
  • Tested with an AI coding agent to verify the skill produces correct behavior

   Adds a new skill focused on extending, debugging, and operating
   existing LiveKit agent codebases. Complements the existing
   livekit-agents skill (which covers building from scratch) with
   production-stage guidance: worker process model, safe async
   patterns, provider lifecycle, performance profiling, testing
   within established infrastructure, and production operations.
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@albertocubeddu

Copy link
Copy Markdown

Any news on this?

@bcherry

bcherry commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Sorry for the long wait on this one but I'm incorporating a lot of the ideas here into new skills in #10

bcherry added a commit that referenced this pull request Sep 24, 2026
The seventh job in the loop, after Ship: getting a version onto LiveKit Cloud
and keeping it healthy. Covers the deploy lifecycle conceptually (a directory
bound to an agent, container build or prebuilt image, versions, secrets as
environment, rollback, build vs deploy logs), the worker process model and
prewarming, safe async inside worker processes and how those bugs present,
provider timeouts and degradation, measure-first performance work and
endpointing tuning, drain on shutdown matched to the orchestrator, SDK
upgrades, observability, and how to change a codebase that is already live.

The worker-model, provider, performance, operations, and existing-codebase
material is adapted from Siddharth Rathod's livekit-agents-production skill
(#7, addressing #6), restructured to this set's conventions: gerund name,
third-person description with triggers disjoint from building and debugging,
no restated flags or SDK identifiers, the duplicated latency/context/listening
section dropped in favor of building-livekit-agents. Deployment is new. The
worker-model claims were checked against both SDKs: each exposes a prewarm
hook on a per-job process object.

README and AGENTS.md gain the seventh row; building, debugging and running
link to it. Trigger eval, 16 queries x 3 covering the new group and the
neighbors it could steal from: 48/48.

Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>
bcherry added a commit that referenced this pull request Sep 24, 2026
…he downstream sync (#10)

* Split monolithic skills into six atomic, conceptual skills

The two original skills had drifted from the shipped surface: livekit-simulations
targeted the pre-June CLI (flat `lk agent simulate`, `agent_description`/`metadata`
schema) and its beta notice was stale on every point; livekit-agents spent much of
its length on MCP boilerplate and mandated tests it never taught.

Replace them with one skill per job in the development loop:

  reading-livekit-docs          fact lookup via Docs MCP / `lk docs`; loaded first by the rest
  building-livekit-agents       architecture: latency, context, handoffs and tasks
  debugging-livekit-agents      live-testing via `lk agent debugger`
  testing-livekit-agents        turn-level pytest/Vitest tests
  writing-livekit-scenarios     authoring scenarios, plus the agent-side wiring
                                (seeding userdata, session-scoped mocks, final-state grading)
  running-livekit-simulations   text vs audio, CI, triage

Skills are conceptual on purpose: they teach the shape of each job and send the
agent to `--help` and the live docs for flags, schemas and API names, so they stay
correct as the CLI and SDK move. Descriptions are third person, written as a set
with explicit hand-offs, and verified with a trigger-collision eval (76/78, then
100% after claiming a bare "test my agent" for debugging).

The old skills remain as DEPRECATED stubs pointing at their replacements; their
references/ and scripts/ (which emitted the superseded schema) are removed.

* Rewrite repo docs for users, contributors, and maintaining agents

README: lead with who the skills are for and what changes when installed, frame
the six skills as a Know → Build → Try → Test → Simulate → Ship lifecycle, give
real trigger phrases per skill, state the requirements (`lk`, a Cloud project,
optionally the Docs MCP server), and explain why the skills are conceptual.

AGENTS.md: a working guide for agents maintaining this repo — how the set routes,
the authoring rules with their reasons, the fact-verification order (`--help`,
docs, the public repos' open PRs, examples), the three-tier evaluation process,
a definition of done, and the deprecation procedure with the downstream-sync
rationale. Names only public repos.

CONTRIBUTING: fix the stale reference to the deprecated skill, align naming and
frontmatter, defer detail to AGENTS.md. PR template: one file, a checklist that
can actually be checked, and a place for eval results.

* Add eval tooling: structure validator, trigger harness, output grader

evals/validate.py      exact checks on every SKILL.md — frontmatter, third-person
                       description under 1,024 chars, body under 500 lines,
                       references present with a TOC when long, cross-references
                       resolve, no time-relative phrasing. Exit code for CI.

evals/trigger/         does the right skill fire? Installs the non-deprecated set
                       into a throwaway fixture, runs each labeled query through
                       `claude -p`, reports a confusion table and misses. The
                       fixture is a shallow clone of agent-starter-python or
                       agent-starter-node — not `lk agent init`, which resolves a
                       Cloud project and writes its credentials. `.env*` files are
                       stripped and dummy LIVEKIT_* variables are injected, which
                       `lk` honours over the configured default project, so an eval
                       cannot spend or upload against the user's project. Tools are
                       read-only. 39 queries, parametrized per language.

evals/output/          does following the skill produce better work? A grader with
                       exact structural checks (schema against the CLI's scenario
                       struct) and an LLM judge for the judgment calls — coverage of
                       the agent's constraints, refusal-shaped expectations, invented
                       capabilities, rotting dates, decidable expectations — with
                       quoted evidence per verdict. Replaces keyword heuristics that
                       false-negatived on good work.

Eval results, logs and scratch fixtures are gitignored.

* Copy-edit pass across skills, docs, and eval tooling

* Remove the deprecated stub skills

The stubs were insurance against a downstream sync whose deletion behavior
couldn't be verified. It has been now: the sync in livekit/internal-actions
copied a hardcoded list of two skill names, so it would neither remove the old
skills nor pick up the new ones. That workflow is being changed to mirror this
directory (discover skills from skills/, remove downstream copies of skills that
no longer exist here), which makes the stubs unnecessary — and a stub's
description would otherwise sit in every user's context on every request.

AGENTS.md's deprecation policy now says to delete, and why. README tells users
of the old skills to reinstall.

* running-livekit-simulations: address review feedback

Frame simulations positively: for regression-testing long-horizon behavior
before deploying to production, in one sentence, rather than leading with cost
in a way that made them never seem like the right option.

Drop the "what goes wrong in automation" list. If it goes wrong the coding
agent will find and fix it; the one design fact worth keeping (every committed
scenario must pass, so keep aspirational ones in a separate file) is now a
sentence in the automation paragraph.

Reframe reading results around the coding agent: the dashboard link is for the
human, export is the agent's path to a failing transcript.

* Remove the downstream skill sync

Skill installation is moving into the LiveKit CLI, so skills will no longer be
pushed into the starter templates and fixtures on every merge. Delete the
dispatcher; the receiving workflow is removed in livekit/internal-actions#16.

AGENTS.md no longer reasons about skill removal in terms of the sync. The
skills/<name>/SKILL.md layout is the install contract installers depend on, and
nothing in this repo can reach a user's existing local copy.

The SYNC_DISPATCH_TOKEN secret and TARGET_REPO variable are now unused.

* Address review: regression-test trigger phrase; judge verdict parsing

running-livekit-simulations' body leads with regression-testing before
deployment, but its description never said so. Add "regression test my agent
before deploying" as a trigger, with two paired queries: that phrase routes to
running, "add a regression test for the bug where..." routes to testing.
Trigger eval on the affected groups, 13 queries x 3: 39/39.

The judge grader coerced verdicts with bool(), so a string "false" from the
model counted as a pass. Coerce only real booleans or unambiguous words;
anything else is recorded as unparseable and excluded from the majority rather
than counted either way. The example schema in the prompt showed the string
"true|false", which invited the problem; it now shows a JSON boolean. Verified
on a real judge run: six boolean verdicts, none unparseable.

* building: the model interprets, code owns state and effects

Port the engineering substance of Shayne Parker's livekit-agents rewrite
(shaynep/skill-refresh) into the skill that replaced it. Three new sections in
building-livekit-agents — never classify intent with code, make every change
mean exactly one thing, start with a failing complete-path test — and a
reference, state-and-effects.md, covering precise mutation semantics, review
and approval as separate runtime events, commit-then-publish, output ownership
and closing, input-mode hook differences, and repairing the first divergence.
testing-livekit-agents gains the three-kinds-of-evidence distinction and the
list of shortcuts that look like tests and aren't.

Deliberately not ported: the task-contract/benchmark framing, version-pinned
SDK facts and API names (kept conceptual per AGENTS.md), the 1,500-line
teaching application and its tests (a worked example belongs in livekit/agents
examples, versioned against the SDK), and the --experimental-auth CLI usage.

Description gains three trigger phrases for these failure modes. Trigger eval
on the building group plus two new paired queries, 7 x 3: 21/21.

* Add operating-livekit-agents: deploy and run an agent in production

The seventh job in the loop, after Ship: getting a version onto LiveKit Cloud
and keeping it healthy. Covers the deploy lifecycle conceptually (a directory
bound to an agent, container build or prebuilt image, versions, secrets as
environment, rollback, build vs deploy logs), the worker process model and
prewarming, safe async inside worker processes and how those bugs present,
provider timeouts and degradation, measure-first performance work and
endpointing tuning, drain on shutdown matched to the orchestrator, SDK
upgrades, observability, and how to change a codebase that is already live.

The worker-model, provider, performance, operations, and existing-codebase
material is adapted from Siddharth Rathod's livekit-agents-production skill
(#7, addressing #6), restructured to this set's conventions: gerund name,
third-person description with triggers disjoint from building and debugging,
no restated flags or SDK identifiers, the duplicated latency/context/listening
section dropped in favor of building-livekit-agents. Deployment is new. The
worker-model claims were checked against both SDKs: each exposes a prewarm
hook on a per-job process object.

README and AGENTS.md gain the seventh row; building, debugging and running
link to it. Trigger eval, 16 queries x 3 covering the new group and the
neighbors it could steal from: 48/48.

Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>

* Address review: no identifiers in testing's description; odd judge-runs

testing-livekit-agents' body had been made conceptual but its description still
named AgentSession, judge() and JudgeGroup. Replace them with the concepts
("the SDK's test session harness", "LLM judging of a reply against an
intent", "the built-in judges"). Trigger eval on the testing group, 6 x 3: 18/18.

The docs said "--judge-runs 2 or 3", but the grader's majority is a strict
greater-than, so a 1-1 split from two runs reads as a fail. Say 3 everywhere,
explain why in the argparse help, and warn when an even count is passed.

* Add a validate workflow; fixes the CodeQL actions check

Run evals/validate.py on every pull request and on pushes to main, plus a
compile check on the eval tooling and a parse check on the trigger query set.
This is the structural half of AGENTS.md's definition of done, enforced where
the PR template already asks for it. The trigger and output evals stay local:
they need a logged-in `claude` and spend tokens.

It also clears the red check on this branch. The repo's CodeQL default setup
scans GitHub Actions, and removing the sync left the branch with no workflow
files, so the actions analysis failed with "CodeQL could not process any code
written in GitHub Actions".

* validate workflow: least-privilege token permissions

CodeQL's actions/missing-workflow-permissions rule: a workflow that doesn't
declare permissions gets the repository's default GITHUB_TOKEN scope. This job
only reads the checkout, so grant contents: read and nothing else.

---------

Co-authored-by: Siddharth Rathod <109035755+rathoresids@users.noreply.github.com>
@bcherry bcherry closed this Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants