11 KiB
AGENTS.md
This is a LiveKit Agents project: a realtime voice AI mock interviewer — the AI plays the interviewer and conducts a two-stage interview (self-introduction, then past experience). LiveKit Agents is a Python SDK for building voice AI agents, used here with LiveKit Cloud. See @README.md for more about the rest of the LiveKit ecosystem.
The following is a guide for working with this project.
Goal
Build a demo that:
- Implements the self-introduction and past-experience stages using a LiveKit multi-agent design, with stable performance and clear interaction logic.
- Transitions between the two stages smoothly and naturally — well-defined switching logic, no interruptions, conflicts, or repeated prompts.
- Provides a time-based fallback so the interview always advances even when the normal (LLM-driven) switch does not fire.
Reference framework: https://github.com/livekit/agents
Environment
- WSL2 (Ubuntu); Python 3.11.
- Package manager:
uv(see Project structure). Dependencies are declared inpyproject.toml. - Secrets live in
.env.local(never commit). The starter uses LiveKit Inference for all models, so the only required keys areLIVEKIT_URL,LIVEKIT_API_KEY,LIVEKIT_API_SECRET. Provider keys (DEEPGRAM_API_KEY,OPENAI_API_KEY,CARTESIA_API_KEY) are needed only if you switch a model frominference.*to a direct plugin.
Project structure
This Python project uses the uv package manager. You should always use uv to install dependencies, run the agent, and run tests.
All app-level code is in the src/ directory. In general, simple agents can be constructed with a single agent.py file. Additional files can be added, but you must retain agent.py as the entrypoint (see the associated Dockerfile for how this is deployed).
Be sure to maintain code formatting. You can use the ruff formatter/linter as needed: uv run ruff format and uv run ruff check.
Architecture (target)
The starter's src/agent.py is the baseline. Build on its real structure:
- Models use LiveKit Inference (routed through LiveKit Cloud — no separate provider keys):
inference.LLM(model="openai/gpt-5.2-chat-latest"),inference.STT(model="deepgram/nova-3", language="multi"),inference.TTS(model="cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"). (gpt-4.1-mini is a cheaper/faster LLM swap if latency or cost bites.) silero.VADis prewarmed inprewarm(proc)and stored atproc.userdata["vad"](process-level), then read asctx.proc.userdata["vad"]. Also:turn_detection=MultilingualModel(),preemptive_generation=True, and noise cancellation viaroom_options=room_io.RoomOptions(audio_input=room_io.AudioInputOptions(noise_cancellation=ai_coustics.audio_enhancement(...))).- One
AgentSession[InterviewData]carries the shared STT/TTS/VAD/turn-detection/LLM defaults plususerdata=InterviewData()(session-level state). - Two stage agents —
SelfIntroAgent->PastExperienceAgent— areAgentsubclasses, each with its owninstructions,on_enter, and a completion@function_tool. They inherit the session's models (override the LLM per-agent only if a stage needs a different one, as LiveKit'smulti_agentexample does). Note: the starter sets the LLM on theAgent; for two agents, lift it onto the session so both stages share it. - Normal transition = the completion tool returns
(NextAgent(chat_ctx=self.chat_ctx.copy()), "<line>"). - Time-based fallback = a per-stage wall-clock watchdog that forces
session.update_agent(...). - Run pattern:
AgentServer()+server.setup_fnc = prewarm+@server.rtc_session(agent_name="interview-agent")+cli.run_app(server);await ctx.connect()comes aftersession.start(...). - New imports the refactor needs (not in the starter by default):
function_tool,RunContext,ChatContextfromlivekit.agents;TurnHandlingOptions,UserTurnLimitOptionsfor the fallback;dataclass,field.
Verify exact API names via the LiveKit docs (below) before coding. This follows LiveKit's handoffs-and-tasks model.
Handoffs and tasks ("workflows")
Voice AI agents are highly sensitive to excessive latency. For this reason, it's important to design complex agents in a structured manner that minimizes the amount of irrelevant context and unnecessary tools included in requests to the LLM. LiveKit Agents supports handoffs (one agent hands control to another) and tasks (tightly-scoped prompts to achieve a specific outcome) to support building reliable workflows. You should make use of these features, instead of writing long instruction prompts that cover multiple phases of a conversation. Refer to the documentation for more information.
Operating rules (IMPORTANT)
- Verify before coding. Confirm every LiveKit class / method / keyword argument against the LiveKit docs (CLI or MCP, see below) before writing it. This framework moves fast — do not rely on memory for API names.
- Stay in scope. Implement only what the current prompt asks. Do not scaffold ahead.
- Do NOT implement the time-based fallback (goal 3). The developer hand-writes that part. Touch it only if explicitly asked, and then only to verify API names — never to author the logic.
- Explain your work. After any substantive change, walk through how it behaves (especially the handoff and any timing) so the developer can verify it.
Known gotchas (verified against LiveKit issues / docs)
- Two different "userdata".
proc.userdata(a dict, process-level) holds the prewarmed VAD;session.userdata(ourInterviewDatadataclass, session-level) holds interview state. They are not the same object — don't putInterviewDataon the process, or look for the VAD on the session. - VAD is required, or
user_stateevents never fire (issue #2364). The starter already prewarms it. - New agents start with a fresh chat history; pass
chat_ctxon both tool-handoff andupdate_agent, or the candidate is forced to repeat themselves. - A handoff tool must return only
(next_agent, line)— no other reply-producing tool in the same turn, or the handoff races and repeats (issue #5150). - Forced transitions use
session.say(...)(deterministic), nevergenerate_reply(it wanders), and the fallback timer must be a standalone wall-clock task — not tied toagent_state_changed, which would make it cancel itself (issue #3148). - The starter sets
preemptive_generation=Trueas a flat session kwarg. If you addturn_handling=TurnHandlingOptions(...)for the fallback's user-turn limit, verify (don't guess) whetherpreemptive_generationmust move insideTurnHandlingOptions— some flat turn kwargs were deprecated in favor ofturn_handling. - The starter ships
ai_cousticsnoise cancellation. If startup errors on missing ai-coustics credentials in local dev, configure its key or temporarily comment out thenoise_cancellation=...line.
Testing
When possible, add tests for agent behavior. Read the documentation, and refer to existing tests in the tests/ directory. Run tests with uv run pytest.
Important: When modifying core agent behavior such as instructions, tool descriptions, and tasks/workflows/handoffs, never just guess what will work. Always use test-driven development (TDD) and begin by writing tests for the desired behavior. For instance, if you're planning to add a new tool, write one or more tests for the tool's behavior, then iterate on the tool until the tests pass correctly. This will ensure you are able to produce a working, reliable agent for the user.
For this project, the highest-priority test is the self-intro -> past-experience handoff: assert that intro_complete fires once, transfers control to PastExperienceAgent, and preserves context. Use LiveKit's test framework (result.expect, judge()).
LiveKit Documentation
LiveKit Agents is a fast-evolving project, and the documentation is updated frequently. You should always refer to the latest documentation when working with this project. For your convenience, LiveKit offers both a CLI and an MCP server that can be used to browse and search its documentation. If the developer has not yet installed the CLI, you should recommend that they install it.
LiveKit CLI
The LiveKit CLI lk docs subcommand gives full access to LiveKit documentation from the terminal. Requires CLI version 2.15.0+. Check with lk --version.
Install or update the CLI:
- macOS:
brew install livekit-cli(update:brew update && brew upgrade livekit-cli) - Linux:
curl -sSL https://get.livekit.io/cli | bash - Windows:
winget install LiveKit.LiveKitCLI(update:winget upgrade LiveKit.LiveKitCLI)
Run lk docs --help to see all available commands. Key commands: lk docs overview, lk docs search, lk docs get-page, lk docs code-search, lk docs changelog, lk docs submit-feedback. Run lk docs <command> --help before using a command for the first time, as the help text includes advice on usage and interpreting results. Prefer browsing (overview, get-page) over search, and prefer search over code-search, as docs pages are more up-to-date and provide better context than raw code.
LiveKit Docs MCP Server
For deeper IDE integration, LiveKit also offers an MCP server with the same documentation access. It uses Streamable HTTP transport at https://docs.livekit.io/mcp. Refer to your IDE's documentation for how to add it.
Submit docs feedback to LiveKit
If you use the LiveKit CLI or MCP Server to search or browse documentation, also submit constructive feedback to LiveKit with lk docs submit-feedback or the MCP submit_docs_feedback tool. Add this to your to-do list when planning work, to ensure you file feedback after completing your other tasks. Useful feedback includes incomplete documentation, broken examples, missing steps or any other documentation gap or issue.
LiveKit CLI
Beyond documentation access, the LiveKit CLI (lk) supports other tasks such as managing SIP trunks for telephony-based agents. Run lk --help to explore available commands.
Phasing — not yet in scope
The current phase is the foundational framework (goals 1-3 above). Do not build the following until the foundation is complete and the developer explicitly asks; if a request seems to call for one of these, confirm scope first: video avatar (the starter stubs anam), post-call scoring/report, recording (egress), resume/JD RAG, and any additional interview stages (behavioral, technical). When stages grow beyond two, reconsider the supervisor / task-group pattern instead of pairwise handoffs.