Jev VAD: Client-side Turn Detection
Let a language-model judge decide when the user is done and when to barge in.
Use with Agora CLI
Clone the recipe and configure it with an Agora project.
agora init my-jev-vad --recipe jev-vadRecipe prompt
Paste into Cursor, Claude Code, v0, or your coding agentYou are implementing the "Jev VAD: Client-side Turn Detection" recipe in this project.
Read the recipe markdown first:
https://raw.githubusercontent.com/AgoraIO-Community/convoai-jev-vad/main/recipe/README.md
Use the source repository for cross-reference:
https://github.com/AgoraIO-Community/convoai-jev-vad
Build this recipe into the user's app using the markdown as the implementation guide. Inspect related source files through the repository links when the recipe points to them. Ask before installing new dependencies.Recipe
Rendered from the configured recipe markdown.
Jev VAD — Client-side Turn Detection for Agora Conversational AI
Copyable prompt for coding agents Add client-side turn detection (Jev VAD) to this Agora Conversational AI app. Read this recipe fully. Start the agent withturn_detection.config.start_of_speech.mode = "manual",end_of_speech.mode = "manual",interruption = {"enable": false, "disabled_config": {"strategy": "append"}},advanced_features.enable_rtm = trueandparameters.data_channel = "rtm". Copyserver/jev_vad.pyinto the FastAPI backend and mount its router (POST /judgeTurn); add a/api/judgeTurnrewrite if the web app proxies/api/*. Copyweb/jev-vad.tsinto the web client and instantiateJevVadwith the agent uid, local uid and the published microphoneMediaStreamTrackonce the agent is connected; callstop()on hang-up. PutTYPESAFE_API_KEYin the backend env andpip install typesafe-sdk. Do not let any other code callmanualSOS,manualEOSorinterrupt. Verify: first word is not clipped, "What's the weather like?" gets a reply within ~1 s of silence, "uh-huh" while the agent talks does not interrupt, "wait, stop" does.
What you build
Replace Agora's built-in start-of-speech / end-of-speech / interruption detection with your own: a cheap acoustic gate opens the turn, and a small language-model judge (TypeSafe's Jev) reads the live transcript to decide when the user is actually done and whether words spoken over the agent are a real interruption or just "uh-huh".
This is a simplified, copy-in version of the full demo in this repository (root README). It drops the side-talk detector, the "holding the floor" patience mode, the state-machine library, the tuning drawer and the telemetry — and keeps the two decisions that matter.
mic ──▶ acoustic gate ──▶ manualSOS ─────────────────────────────┐
▼
Agora ASR ──▶ live transcript ──▶ POST /judgeTurn ──▶ Jev ──▶ turn_complete ≥ bar? ──▶ quiet ──▶ manualEOS
(agent talking) taking_floor ≥ thr? ──▶ interrupt()What you get
| Behaviour | How |
|---|---|
| First word never clipped | manualSOS fires ~110 ms after the mic rises above the noise floor, before any words exist. |
| Ends the turn when the sentence is done, not when the audio pauses | Jev scores turn_complete on the streaming text; a short answer to the agent's question closes immediately, "I have a question." does not. |
| Long pauses don't stall | The bar Jev must clear relaxes from 0.72 to 0.55 over 3 s of measured mic silence; a hard 5 s cap backs it up. |
| Coughs, doors, TV don't interrupt | The gate opens a turn, but the STT emits no words, so nothing is judged and nothing is sent. Empty turns never fire manualEOS. |
| "Uh-huh" doesn't interrupt; "wait, stop" does | While the agent talks, the first transcribed words are judged with taking_floor; only a real floor-grab (or an answer to the agent's question) calls interrupt(). |
Cost: one Jev call per transcript update, ~2 k input / 60 output tokens, 200–400 ms.
Prerequisites
- A working Agora ConvoAI app (agent starts, client joins, transcripts arrive over RTM).
If not, start from the Python quickstart.
agora-agent-client-toolkit≥ 2.10 on the web client (providesmanualSOS,manualEOS,interrupt).- A TypeSafe API key (
TYPESAFE_API_KEY) andpip install typesafe-sdk. - Agora App ID + App Certificate.
Run the reference demo
The full demo in this repository is the reference implementation. To see the behaviour before copying the two files into your own app:
bun run setup # web deps + server venv
agora login && agora project use <project> # or edit server/.env by hand
agora project env write server/.env # writes AGORA_APP_ID / AGORA_APP_CERTIFICATE
echo 'TYPESAFE_API_KEY=<your-key>' >> server/.env
bun run dev # backend :8000 + web :3000Open http://localhost:3000 → Start Conversation. The in-call panel shows the machine state and Jev scores live; the tuning drawer changes every knob without a restart.
Step 1 — Start the agent with manual turns and interruption off
Everything else in this recipe assumes the server does not decide turns on its own. Python SDK (agora-agent):
agent = AgoraAgent(
client=client,
instructions=SYSTEM_PROMPT,
greeting="Hi! How can I help?",
turn_detection={
"mode": "default",
"config": {
"start_of_speech": {"mode": "manual"}, # client says when a turn opens
"end_of_speech": {"mode": "manual"}, # client says when it closes
},
},
interruption={
"enable": False, # voice alone never cuts the agent off
"disabled_config": {"strategy": "append"}, # keep speech captured during agent talk
},
advanced_features={"enable_rtm": True},
parameters={
"data_channel": "rtm", # transcripts + agent state to the client
"enable_metrics": True,
"silence_config": {"timeout_ms": 0}, # no "are you still there?" nags
"audio_scenario": "chorus", # optional: low-latency profile for web
},
)Raw /join equivalent: properties.turn_detection.config.start_of_speech.mode = "manual", …end_of_speech.mode = "manual", properties.interruption = {"enable": false, "disabled_config": {"strategy": "append"}}, advanced_features.enable_rtm = true, parameters.data_channel = "rtm".
Two server behaviours to know (measured, not documented):
manualEOSon a turn with no ASR text makes the LLM reply "your message didn't come
through". Never close an empty turn.
- An open manual turn with no words is not auto-closed by the server (tested to 120 s).
False opens are harmless; leave them open and keep listening.
Step 2 — Add the Jev judge route (server)
Copy `server/jev_vad.py` next to your FastAPI app and mount it:
from jev_vad import router as jev_router
app.include_router(jev_router)POST /judgeTurn takes the current user transcript plus a little context and returns:
{
"turn_complete": 0.83,
"taking_floor": null,
"should_end": true,
"should_interrupt": false,
"thresholds": { "turn_complete": 0.72, "taking_floor": 0.5 },
"latency_ms": 310
}Two Jev nouls (yes/no judgments with a probability):
turn_complete— always. Prompted with the fact that it sees **text from a streaming
recognizer, not audio**: punctuation is guessed, fillers are normal, coughs are invisible, and silence_since_last_word_ms is the only prosody it gets.
taking_floor— only whenagent_stateisspeaking/thinking. Gets
assistant_current_speech so it can tell an echo fragment or a backchannel from a real interruption, and treats an answer to a question the agent is asking as taking the floor.
If your web app proxies /api/* to the backend (as the quickstart does), add a rewrite for /api/judgeTurn.
Step 3 — Drop in the client controller (web)
Copy `web/jev-vad.ts` into your client and start it once the agent is connected and your microphone track is published:
import { JevVad } from './jev-vad'
const vad = new JevVad({
agentUid, // the agent's RTC uid (string)
localUid, // your user's RTC uid (string)
micTrack: localAudioTrack.getMediaStreamTrack(),
judgeUrl: '/api/judgeTurn',
onStatus: (s) => setVadStatus(s), // optional, for a status line
})
vad.start()
// on hang-up
vad.stop()It subscribes to the toolkit's AGENT_STATE_CHANGED, TRANSCRIPT_UPDATED, USER_MANUAL_EOS and AGENT_MANUAL_EOS events and owns all manualSOS / manualEOS / interrupt calls. Nothing else in your app should call those.
What the controller does, in order
- Gate. Every 20 ms: RMS of the mic vs. a slowly tracked noise floor. Above
noiseFloor × 3 (×5 while the agent speaks, to reject TTS echo) for 110 ms → manualSOS. Remembers the last user utterance so the previous turn's text is not judged again.
- Judge. On each transcript update (debounced 150 ms; 0 ms during barge-in) →
POST /judgeTurn.
- Barge-in. If the turn opened while the agent was talking:
should_interrupt→
interrupt(), otherwise hold ("backchannel") and wait for more words. In-flight judgments are not cancelled when new partials arrive — a verdict on a prefix of the text is acted on as soon as it lands, which is what makes interruptions feel immediate. Once the agent goes quiet, held words are re-judged as a normal turn so they are not lost.
- End.
turn_complete ≥ bar(quiet)→ wait until the mic has been quiet 400 ms →
manualEOS. Otherwise re-judge after 900 ms of silence, passing silence_ms so Jev sees the pause; the bar relaxes toward 0.55 as silence grows; 5 s of silence closes regardless.
- Reset on
manualEOS, or when the server reports a manual-EoS result.
Tuning
| Knob | Default | Move it when… |
|---|---|---|
speechHoldMs | 110 | Clipping → lower. Too many false opens → raise. |
gateFactor / gateFactorAgentSpeaking | 3 / 5 | Quiet speakers missed → lower. Room noise opens turns → raise. |
eosThreshold | 0.72 | Agent jumps in early → raise. Feels sluggish on clear questions → lower. |
eosThresholdFloor / eosDecayMs | 0.55 / 3000 | Complete-sounding sentences stall → lower floor or shorten decay. |
confirmQuietMs | 400 | Agent replies while you draw breath → raise. |
maxSilenceMs | 5000 | Hard cap. Rarely hit once the relaxing bar is in place. |
JEV_BARGE_IN_THRESHOLD (server) | 0.5 | Backchannels interrupt → raise. Real interruptions ignored → lower. |
Measured on real sessions: fragments score 0.04–0.15, complete short questions 0.67–0.80, "so I have a few questions… pricing" 0.50–0.54 (Jev correctly hears more coming). Thresholds are where you express your preference for speed vs. patience.
Pitfalls
- Don't let anything else open or close turns. If the toolkit's default handlers or a
push-to-talk button also call manualSOS/manualEOS, the server's turn state and the controller's will diverge.
- Don't judge the previous turn. When a new turn opens, the transcript still holds the
last utterance; the controller snapshots it as a baseline and ignores it. Skip this and the new turn closes instantly on stale text.
- Never `manualEOS` an empty turn — see Step 1.
- Echo. Without headphones, the agent's voice re-enters the mic. The stricter
gateFactorAgentSpeaking handles most of it; a residual echo fragment that reaches the STT is caught by taking_floor (it mirrors assistant_current_speech).
- Threshold overrides ride with the request.
eos_thresholdin the body beats the env
default, so a UI slider can retune Jev mid-call with no server restart.
Verification
- Say a short question. The transcript's first word is intact (no clipping) and the agent
replies within about a second of you stopping.
- Say "So I have a question about…" and stop. The agent waits (Jev scores the sentence as
incomplete); after ~3 s of silence the relaxed bar closes the turn anyway.
- Cough or tap the desk. A turn may open, but no
manualEOSfires and the agent stays quiet. - While the agent is talking, say "uh-huh". It keeps talking. Say "wait, stop" — it stops.
- Let the agent ask you a question and answer over it with "pretty good". It stops and
takes the answer (answers to the agent's own question count as taking the floor).
- Backend log shows one
judgeTurnline per transcript update withcomplete=and, during
agent speech, taking_floor= scores.
Non-goals
- Replacing the STT. Jev judges text; the recognizer is still Agora's (any supported ASR vendor).
- Server-side turn detection. This recipe deliberately moves the decision to the client so it
can use mic energy, agent state and the live transcript together.
- Side-talk / multi-party awareness and "holding the floor" patience — see below; they are in
the full demo, not in the drop-in files.
Going further (what the full demo adds)
addressed_to_agent— a third noul that holds the turn when the user is clearly talking
to someone else in the room ("hey Mark, dinner's ready").
holding_floor— patience mode: "hmm, let me think" extends the silence cap to 12 s and
disables the relaxing bar, released by the latest words ("okay, go ahead").
- A pure, unit-tested state machine instead of instance fields; a live tuning drawer that
persists to localStorage; and client-timeline mirroring into the server log for post-hoc analysis.
See the repository root README for the full architecture and the spike findings behind these numbers. Full-demo source: `web/src/lib/turn-machine.ts`, `web/src/hooks/useTurnController.ts`, `server/src/jev_eos.py`.