# Jev VAD — Client-side Turn Detection for Agora Conversational AI > **Copyable prompt for coding agents** > > Add client-side turn detection (Jev VAD) to this Agora Conversational AI app. Read this recipe fully. Start the agent with `turn_detection.config.start_of_speech.mode = "manual"`, `end_of_speech.mode = "manual"`, `interruption = {"enable": false, "disabled_config": {"strategy": "append"}}`, `advanced_features.enable_rtm = true` and `parameters.data_channel = "rtm"`. Copy `server/jev_vad.py` into the FastAPI backend and mount its router (`POST /judgeTurn`); add a `/api/judgeTurn` rewrite if the web app proxies `/api/*`. Copy `web/jev-vad.ts` into the web client and instantiate `JevVad` with the agent uid, local uid and the published microphone `MediaStreamTrack` once the agent is connected; call `stop()` on hang-up. Put `TYPESAFE_API_KEY` in the backend env and `pip install typesafe-sdk`. Do not let any other code call `manualSOS`, `manualEOS` or `interrupt`. Verify: first word is not clipped, "What's the weather like?" gets a reply within ~1 s of silence, "uh-huh" while the agent talks does not interrupt, "wait, stop" does. ## What you build Replace Agora's built-in start-of-speech / end-of-speech / interruption detection with your own: a cheap acoustic gate opens the turn, and a small language-model judge (TypeSafe's **Jev**) reads the live transcript to decide when the user is *actually* done and whether words spoken over the agent are a real interruption or just "uh-huh". This is a simplified, copy-in version of the full demo in this repository ([root README](../README.md)). It drops the side-talk detector, the "holding the floor" patience mode, the state-machine library, the tuning drawer and the telemetry — and keeps the two decisions that matter. ``` mic ──▶ acoustic gate ──▶ manualSOS ─────────────────────────────┐ ▼ Agora ASR ──▶ live transcript ──▶ POST /judgeTurn ──▶ Jev ──▶ turn_complete ≥ bar? ──▶ quiet ──▶ manualEOS (agent talking) taking_floor ≥ thr? ──▶ interrupt() ``` ## What you get | Behaviour | How | |---|---| | First word never clipped | `manualSOS` fires ~110 ms after the mic rises above the noise floor, before any words exist. | | Ends the turn when the *sentence* is done, not when the audio pauses | Jev scores `turn_complete` on the streaming text; a short answer to the agent's question closes immediately, "I have a question." does not. | | Long pauses don't stall | The bar Jev must clear relaxes from 0.72 to 0.55 over 3 s of measured mic silence; a hard 5 s cap backs it up. | | Coughs, doors, TV don't interrupt | The gate opens a turn, but the STT emits no words, so nothing is judged and nothing is sent. Empty turns never fire `manualEOS`. | | "Uh-huh" doesn't interrupt; "wait, stop" does | While the agent talks, the first transcribed words are judged with `taking_floor`; only a real floor-grab (or an answer to the agent's question) calls `interrupt()`. | Cost: one Jev call per transcript update, ~2 k input / 60 output tokens, 200–400 ms. ## Prerequisites - A working Agora ConvoAI app (agent starts, client joins, transcripts arrive over RTM). If not, start from the [Python quickstart](https://github.com/AgoraIO-Conversational-AI/agent-quickstart-python). - `agora-agent-client-toolkit` ≥ 2.10 on the web client (provides `manualSOS`, `manualEOS`, `interrupt`). - A TypeSafe API key (`TYPESAFE_API_KEY`) and `pip install typesafe-sdk`. - Agora App ID + App Certificate. ## Run the reference demo The full demo in this repository is the reference implementation. To see the behaviour before copying the two files into your own app: ```bash bun run setup # web deps + server venv agora login && agora project use # or edit server/.env by hand agora project env write server/.env # writes AGORA_APP_ID / AGORA_APP_CERTIFICATE echo 'TYPESAFE_API_KEY=' >> server/.env bun run dev # backend :8000 + web :3000 ``` Open http://localhost:3000 → **Start Conversation**. The in-call panel shows the machine state and Jev scores live; the tuning drawer changes every knob without a restart. ## Step 1 — Start the agent with manual turns and interruption off Everything else in this recipe assumes the server does **not** decide turns on its own. Python SDK (`agora-agent`): ```python agent = AgoraAgent( client=client, instructions=SYSTEM_PROMPT, greeting="Hi! How can I help?", turn_detection={ "mode": "default", "config": { "start_of_speech": {"mode": "manual"}, # client says when a turn opens "end_of_speech": {"mode": "manual"}, # client says when it closes }, }, interruption={ "enable": False, # voice alone never cuts the agent off "disabled_config": {"strategy": "append"}, # keep speech captured during agent talk }, advanced_features={"enable_rtm": True}, parameters={ "data_channel": "rtm", # transcripts + agent state to the client "enable_metrics": True, "silence_config": {"timeout_ms": 0}, # no "are you still there?" nags "audio_scenario": "chorus", # optional: low-latency profile for web }, ) ``` Raw `/join` equivalent: `properties.turn_detection.config.start_of_speech.mode = "manual"`, `…end_of_speech.mode = "manual"`, `properties.interruption = {"enable": false, "disabled_config": {"strategy": "append"}}`, `advanced_features.enable_rtm = true`, `parameters.data_channel = "rtm"`. Two server behaviours to know (measured, not documented): - `manualEOS` on a turn with **no ASR text** makes the LLM reply "your message didn't come through". Never close an empty turn. - An open manual turn with no words is **not** auto-closed by the server (tested to 120 s). False opens are harmless; leave them open and keep listening. ## Step 2 — Add the Jev judge route (server) Copy [`server/jev_vad.py`](server/jev_vad.py) next to your FastAPI app and mount it: ```python from jev_vad import router as jev_router app.include_router(jev_router) ``` `POST /judgeTurn` takes the current user transcript plus a little context and returns: ```json { "turn_complete": 0.83, "taking_floor": null, "should_end": true, "should_interrupt": false, "thresholds": { "turn_complete": 0.72, "taking_floor": 0.5 }, "latency_ms": 310 } ``` Two Jev *nouls* (yes/no judgments with a probability): - `turn_complete` — always. Prompted with the fact that it sees **text from a streaming recognizer, not audio**: punctuation is guessed, fillers are normal, coughs are invisible, and `silence_since_last_word_ms` is the only prosody it gets. - `taking_floor` — only when `agent_state` is `speaking`/`thinking`. Gets `assistant_current_speech` so it can tell an echo fragment or a backchannel from a real interruption, and treats an *answer* to a question the agent is asking as taking the floor. If your web app proxies `/api/*` to the backend (as the quickstart does), add a rewrite for `/api/judgeTurn`. ## Step 3 — Drop in the client controller (web) Copy [`web/jev-vad.ts`](web/jev-vad.ts) into your client and start it once the agent is connected and your microphone track is published: ```ts import { JevVad } from './jev-vad' const vad = new JevVad({ agentUid, // the agent's RTC uid (string) localUid, // your user's RTC uid (string) micTrack: localAudioTrack.getMediaStreamTrack(), judgeUrl: '/api/judgeTurn', onStatus: (s) => setVadStatus(s), // optional, for a status line }) vad.start() // on hang-up vad.stop() ``` It subscribes to the toolkit's `AGENT_STATE_CHANGED`, `TRANSCRIPT_UPDATED`, `USER_MANUAL_EOS` and `AGENT_MANUAL_EOS` events and owns all `manualSOS` / `manualEOS` / `interrupt` calls. Nothing else in your app should call those. ### What the controller does, in order 1. **Gate.** Every 20 ms: RMS of the mic vs. a slowly tracked noise floor. Above `noiseFloor × 3` (×5 while the agent speaks, to reject TTS echo) for 110 ms → `manualSOS`. Remembers the last user utterance so the previous turn's text is not judged again. 2. **Judge.** On each transcript update (debounced 150 ms; 0 ms during barge-in) → `POST /judgeTurn`. 3. **Barge-in.** If the turn opened while the agent was talking: `should_interrupt` → `interrupt()`, otherwise hold ("backchannel") and wait for more words. In-flight judgments are *not* cancelled when new partials arrive — a verdict on a prefix of the text is acted on as soon as it lands, which is what makes interruptions feel immediate. Once the agent goes quiet, held words are re-judged as a normal turn so they are not lost. 4. **End.** `turn_complete ≥ bar(quiet)` → wait until the mic has been quiet 400 ms → `manualEOS`. Otherwise re-judge after 900 ms of silence, passing `silence_ms` so Jev sees the pause; the bar relaxes toward 0.55 as silence grows; 5 s of silence closes regardless. 5. **Reset** on `manualEOS`, or when the server reports a manual-EoS result. ## Tuning | Knob | Default | Move it when… | |---|---|---| | `speechHoldMs` | 110 | Clipping → lower. Too many false opens → raise. | | `gateFactor` / `gateFactorAgentSpeaking` | 3 / 5 | Quiet speakers missed → lower. Room noise opens turns → raise. | | `eosThreshold` | 0.72 | Agent jumps in early → raise. Feels sluggish on clear questions → lower. | | `eosThresholdFloor` / `eosDecayMs` | 0.55 / 3000 | Complete-sounding sentences stall → lower floor or shorten decay. | | `confirmQuietMs` | 400 | Agent replies while you draw breath → raise. | | `maxSilenceMs` | 5000 | Hard cap. Rarely hit once the relaxing bar is in place. | | `JEV_BARGE_IN_THRESHOLD` (server) | 0.5 | Backchannels interrupt → raise. Real interruptions ignored → lower. | Measured on real sessions: fragments score 0.04–0.15, complete short questions 0.67–0.80, "so I have a few questions… pricing" 0.50–0.54 (Jev correctly hears more coming). Thresholds are where you express *your* preference for speed vs. patience. ## Pitfalls - **Don't let anything else open or close turns.** If the toolkit's default handlers or a push-to-talk button also call `manualSOS`/`manualEOS`, the server's turn state and the controller's will diverge. - **Don't judge the previous turn.** When a new turn opens, the transcript still holds the last utterance; the controller snapshots it as a baseline and ignores it. Skip this and the new turn closes instantly on stale text. - **Never `manualEOS` an empty turn** — see Step 1. - **Echo.** Without headphones, the agent's voice re-enters the mic. The stricter `gateFactorAgentSpeaking` handles most of it; a residual echo fragment that reaches the STT is caught by `taking_floor` (it mirrors `assistant_current_speech`). - **Threshold overrides ride with the request.** `eos_threshold` in the body beats the env default, so a UI slider can retune Jev mid-call with no server restart. ## Verification 1. Say a short question. The transcript's first word is intact (no clipping) and the agent replies within about a second of you stopping. 2. Say "So I have a question about…" and stop. The agent waits (Jev scores the sentence as incomplete); after ~3 s of silence the relaxed bar closes the turn anyway. 3. Cough or tap the desk. A turn may open, but no `manualEOS` fires and the agent stays quiet. 4. While the agent is talking, say "uh-huh". It keeps talking. Say "wait, stop" — it stops. 5. Let the agent ask you a question and answer over it with "pretty good". It stops and takes the answer (answers to the agent's own question count as taking the floor). 6. Backend log shows one `judgeTurn` line per transcript update with `complete=` and, during agent speech, `taking_floor=` scores. ## Non-goals - Replacing the STT. Jev judges text; the recognizer is still Agora's (any supported ASR vendor). - Server-side turn detection. This recipe deliberately moves the decision to the client so it can use mic energy, agent state and the live transcript together. - Side-talk / multi-party awareness and "holding the floor" patience — see below; they are in the full demo, not in the drop-in files. ## Going further (what the full demo adds) - `addressed_to_agent` — a third noul that holds the turn when the user is clearly talking to someone else in the room ("hey Mark, dinner's ready"). - `holding_floor` — patience mode: "hmm, let me think" extends the silence cap to 12 s and disables the relaxing bar, released by the latest words ("okay, go ahead"). - A pure, unit-tested state machine instead of instance fields; a live tuning drawer that persists to `localStorage`; and client-timeline mirroring into the server log for post-hoc analysis. See the [repository root README](../README.md) for the full architecture and the spike findings behind these numbers. Full-demo source: [`web/src/lib/turn-machine.ts`](../web/src/lib/turn-machine.ts), [`web/src/hooks/useTurnController.ts`](../web/src/hooks/useTurnController.ts), [`server/src/jev_eos.py`](../server/src/jev_eos.py).