diff --git a/readme.md b/readme.md new file mode 100644 index 0000000..63699b8 --- /dev/null +++ b/readme.md @@ -0,0 +1,235 @@ +# AI Mock Interview Agent Walkthrough + +## Purpose + +This project turns the LiveKit Agents Python starter into a realtime voice AI mock interviewer. The agent plays the interviewer, named Jordan, and guides the candidate through two foundational interview stages: + +1. Self-introduction +2. Past experience + +The implementation is centered in `src/agent.py`. It uses LiveKit's multi-agent workflow model so each interview stage has its own focused agent, prompt, entry behavior, and transition rules. + +## Architecture at a Glance + +The runtime is one LiveKit `AgentSession[InterviewData]` with shared speech, voice, turn-taking, and LLM configuration. Stage-specific behavior lives in two `Agent` subclasses: + +- `SelfIntroAgent`: opens the interview and asks for the candidate's self-introduction. +- `PastExperienceAgent`: asks targeted questions about the candidate's prior work or projects. + +The shared session state is intentionally small: + +```python +@dataclass +class InterviewData: + """Session-level state for the mock interview.""" + + intro_was_given: bool = field(default=False) +``` + +The production session uses LiveKit Inference for models, prewarmed Silero VAD for speech detection, multilingual turn detection, adaptive interruption handling, preemptive generation, and AI-coustics noise cancellation: + +```python +session = AgentSession[InterviewData]( + llm=inference.LLM(model="openai/gpt-4.1-mini"), + stt=inference.STT(model="deepgram/nova-3", language="en"), + tts=inference.TTS( + model="cartesia/sonic-3", + voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc", + ), + vad=ctx.proc.userdata["vad"], + turn_handling=TurnHandlingOptions( + turn_detection=MultilingualModel(), + endpointing={"mode": "dynamic", "min_delay": 0.5, "max_delay": 3.0}, + interruption={ + "mode": "adaptive", + "min_words": 2, + "resume_false_interruption": True, + }, + preemptive_generation={"enabled": True, "preemptive_tts": True}, + ), + userdata=InterviewData(), +) +``` + +## Goal One: Two Interview Stages with Clear Logic + +The project avoids one large, general-purpose interviewer prompt. Instead, each stage has a dedicated agent with only the instructions needed for that phase. + +`SelfIntroAgent` is responsible for exactly one job: greet the candidate and collect a self-introduction. Its instructions explicitly tell the model to call `intro_complete` as soon as the candidate gives any substantive introduction: + +```python +class SelfIntroAgent(Agent): + def __init__(self, *, silence_budget: float = 20.0, stage_budget: float = 60.0) -> None: + super().__init__( + instructions=textwrap.dedent( + """\ + You are Jordan, a senior hiring manager... + You are conducting the opening stage of a mock job interview, and your + sole goal at this stage is to hear the candidate's self-introduction. + + ...Once they have shared a description of themselves + and their experience — even briefly — call `intro_complete` immediately + to advance the interview. Do not ask follow-up questions... + """ + ) + + "\n" + + _VOICE_RULES, + ) +``` + +When the self-introduction stage begins, the agent generates a short greeting and invitation, then starts both fallback timers: + +```python +async def on_enter(self) -> None: + await self.session.generate_reply( + instructions="Greet the candidate warmly and invite them to introduce themselves." + ) + self._silence_task = asyncio.create_task(self._silence_watchdog()) + self._stage_task = asyncio.create_task(self._stage_watchdog()) +``` + +`PastExperienceAgent` owns the second stage. It assumes the introduction has already happened and focuses on one meaningful experience at a time: + +```python +class PastExperienceAgent(Agent): + def __init__(self, chat_ctx: ChatContext | None = None) -> None: + super().__init__( + instructions=textwrap.dedent( + """\ + You are Jordan, the same senior hiring manager from the opening stage... + You are now conducting the past-experience stage of a mock job interview. + + Pick the most relevant experience the candidate mentioned and ask them + to walk you through it: what the challenge was, their specific + contribution, the outcome, and what they learned. + """ + ) + + "\n" + + _VOICE_RULES, + chat_ctx=chat_ctx, + ) +``` + +This split keeps each prompt short, reduces unnecessary context, and makes the conversation easier to test. + +## Goal Two: Smooth Stage Transition Without Repetition + +The normal transition is handled by the `intro_complete` function tool. It is called by `SelfIntroAgent` after the candidate gives a self-introduction. + +```python +@function_tool +async def intro_complete(self, context: RunContext[InterviewData]): + """Call this tool as soon as the candidate has described themselves and + their work experience.""" + context.userdata.intro_was_given = True + return ( + PastExperienceAgent(chat_ctx=self.chat_ctx.copy(exclude_instructions=True)), + "Thanks for that introduction. Now I'd like to explore your past experience.", + ) +``` + +There are three important details here: + +- `context.userdata.intro_was_given = True` records that the normal introduction happened. +- `PastExperienceAgent(chat_ctx=self.chat_ctx.copy(exclude_instructions=True))` preserves conversation history without carrying old system instructions into the next agent. +- The tool returns only the next agent and one transition line, avoiding competing replies during the handoff. + +When `PastExperienceAgent` enters, it checks the shared session state. If the self-introduction happened normally, it immediately asks about the most relevant experience without reintroducing Jordan or restating the transition: + +```python +async def on_enter(self) -> None: + if self.session.userdata.intro_was_given: + instructions = ( + "Ask the candidate to walk you through their most relevant past experience " + "or project. Do not re-introduce yourself or re-explain the transition — " + "go straight to the question." + ) + else: + instructions = ( + "Start directly: ask the candidate to walk you through their most recent " + "role or a project they've worked on. Do not reference any prior " + "introduction — go straight to the question." + ) + await self.session.generate_reply(instructions=instructions) +``` + +That fallback branch matters because forced transitions can happen even if the candidate never completed a normal introduction. + +## Goal Three: Time-Based Fallback for Continuous Progress + +The normal path depends on the LLM calling `intro_complete`. To guarantee the interview still moves forward, `SelfIntroAgent` starts two independent watchdogs after its opening prompt: + +- Silence watchdog: moves on after continuous candidate silence. +- Stage watchdog: moves on after the total stage time budget is reached. + +The default budgets are: + +```python +def __init__(self, *, silence_budget: float = 20.0, stage_budget: float = 60.0) -> None: +``` + +The silence watchdog resets while LiveKit reports that the user is speaking: + +```python +while True: + if self.session.user_state == "speaking": + last_speech_time = time.monotonic() + silence_elapsed = time.monotonic() - last_speech_time + if silence_elapsed >= self._silence_budget: + break + await asyncio.sleep(min(self._silence_budget - silence_elapsed, poll)) +``` + +When a fallback fires, it uses a deterministic spoken transition and then updates the active agent: + +```python +if not isinstance(self.session.current_agent, SelfIntroAgent): + return +if self._transitioning: + return +self._transitioning = True + +await self.session.say( + "Let's go ahead and move on to your background and experience." +) +self.session.update_agent( + PastExperienceAgent(chat_ctx=self.chat_ctx.copy(exclude_instructions=True)) +) +``` + +The fallback also protects against double transitions. Each watchdog checks whether `SelfIntroAgent` is still active, waits for any in-flight speech to finish, checks again, and then uses `_transitioning` as a final guard. + +On normal handoff or fallback handoff, `on_exit` cancels both watchdog tasks: + +```python +async def on_exit(self) -> None: + for task in (self._silence_task, self._stage_task): + if task is not None: + task.cancel() + self._silence_task = None + self._stage_task = None +``` + +This prevents the old stage from firing a second transition after control has already moved to `PastExperienceAgent`. + +## Options Considered but Not Used + +| Option | Why We Did Not Use It | +| --- | --- | +| One monolithic interviewer agent | It would mix stage instructions, increase prompt scope, and make transitions less reliable. | +| LLM-only transition logic | If the model fails to call the handoff tool, the interview can stall. The watchdog fallback guarantees progress. | +| `generate_reply` for forced transitions | Fallback movement needs deterministic wording. `session.say(...)` avoids wandering or repeated explanations. | +| Dropping chat history on handoff | The past-experience stage needs the candidate's introduction. Without copied `chat_ctx`, the candidate may have to repeat themselves. | +| Direct provider plugins for every model | LiveKit Inference keeps model routing under LiveKit Cloud credentials and avoids separate provider-key setup for the demo. | +| Extra interview stages now | The current roadmap limits scope to the foundational two-stage demo. More stages can be added later with the same handoff pattern. | + +## Verification + +The test suite in `tests/test_agent.py` covers the core behaviors: + +- `test_self_intro_greeting`: verifies the opening agent greets and asks for an introduction. +- `test_intro_complete_handoff`: verifies a complete introduction calls `intro_complete` and transfers to `PastExperienceAgent`. +- `test_no_premature_handoff`: verifies a bare greeting does not trigger the handoff. +- `test_watchdog_cancelled_on_normal_handoff`: verifies fallback timers are cancelled after normal transition. +- `test_fallback_handoff`: verifies the silence fallback moves to `PastExperienceAgent`. +- `test_silence_reset_on_speech`: verifies the silence timer resets while the candidate is speaking.