All recipes
PythonTypeScriptIntermediate

Jev VAD: Client-side Turn Detection

Let a language-model judge decide when the user is done and when to barge in.

Use with Agora CLI

Clone the recipe and configure it with an Agora project.

Agora CLI
agora init my-jev-vad --recipe jev-vad

Recipe prompt

Paste into Cursor, Claude Code, v0, or your coding agent
Use with your coding agent
You are implementing the "Jev VAD: Client-side Turn Detection" recipe in this project.

Read the recipe markdown first:
https://raw.githubusercontent.com/AgoraIO-Community/convoai-jev-vad/main/recipe/README.md

Use the source repository for cross-reference:
https://github.com/AgoraIO-Community/convoai-jev-vad

Build this recipe into the user's app using the markdown as the implementation guide. Inspect related source files through the repository links when the recipe points to them. Ask before installing new dependencies.

Recipe

Rendered from the configured recipe markdown.

Raw

Jev VAD — Client-side Turn Detection for Agora Conversational AI

Copyable prompt for coding agents Add client-side turn detection (Jev VAD) to this Agora Conversational AI app. Read this recipe fully. Start the agent with turn_detection.config.start_of_speech.mode = "manual", end_of_speech.mode = "manual", interruption = {"enable": false, "disabled_config": {"strategy": "append"}}, advanced_features.enable_rtm = true and parameters.data_channel = "rtm". Copy server/jev_vad.py into the FastAPI backend and mount its router (POST /judgeTurn); add a /api/judgeTurn rewrite if the web app proxies /api/*. Copy web/jev-vad.ts into the web client and instantiate JevVad with the agent uid, local uid and the published microphone MediaStreamTrack once the agent is connected; call stop() on hang-up. Put TYPESAFE_API_KEY in the backend env and pip install typesafe-sdk. Do not let any other code call manualSOS, manualEOS or interrupt. Verify: first word is not clipped, "What's the weather like?" gets a reply within ~1 s of silence, "uh-huh" while the agent talks does not interrupt, "wait, stop" does.

What you build

Replace Agora's built-in start-of-speech / end-of-speech / interruption detection with your own: a cheap acoustic gate opens the turn, and a small language-model judge (TypeSafe's Jev) reads the live transcript to decide when the user is actually done and whether words spoken over the agent are a real interruption or just "uh-huh".

This is a simplified, copy-in version of the full demo in this repository (root README). It drops the side-talk detector, the "holding the floor" patience mode, the state-machine library, the tuning drawer and the telemetry — and keeps the two decisions that matter.

mic ──▶ acoustic gate ──▶ manualSOS ─────────────────────────────┐
                                                                 ▼
Agora ASR ──▶ live transcript ──▶ POST /judgeTurn ──▶ Jev ──▶ turn_complete ≥ bar? ──▶ quiet ──▶ manualEOS
                    (agent talking)                          taking_floor ≥ thr?  ──▶ interrupt()

What you get

BehaviourHow
First word never clippedmanualSOS fires ~110 ms after the mic rises above the noise floor, before any words exist.
Ends the turn when the sentence is done, not when the audio pausesJev scores turn_complete on the streaming text; a short answer to the agent's question closes immediately, "I have a question." does not.
Long pauses don't stallThe bar Jev must clear relaxes from 0.72 to 0.55 over 3 s of measured mic silence; a hard 5 s cap backs it up.
Coughs, doors, TV don't interruptThe gate opens a turn, but the STT emits no words, so nothing is judged and nothing is sent. Empty turns never fire manualEOS.
"Uh-huh" doesn't interrupt; "wait, stop" doesWhile the agent talks, the first transcribed words are judged with taking_floor; only a real floor-grab (or an answer to the agent's question) calls interrupt().

Cost: one Jev call per transcript update, ~2 k input / 60 output tokens, 200–400 ms.

Prerequisites

  • A working Agora ConvoAI app (agent starts, client joins, transcripts arrive over RTM).

If not, start from the Python quickstart.

  • agora-agent-client-toolkit ≥ 2.10 on the web client (provides manualSOS, manualEOS, interrupt).
  • A TypeSafe API key (TYPESAFE_API_KEY) and pip install typesafe-sdk.
  • Agora App ID + App Certificate.

Run the reference demo

The full demo in this repository is the reference implementation. To see the behaviour before copying the two files into your own app:

bun run setup                              # web deps + server venv
agora login && agora project use <project> # or edit server/.env by hand
agora project env write server/.env        # writes AGORA_APP_ID / AGORA_APP_CERTIFICATE
echo 'TYPESAFE_API_KEY=<your-key>' >> server/.env
bun run dev                                # backend :8000 + web :3000

Open http://localhost:3000 → Start Conversation. The in-call panel shows the machine state and Jev scores live; the tuning drawer changes every knob without a restart.

Step 1 — Start the agent with manual turns and interruption off

Everything else in this recipe assumes the server does not decide turns on its own. Python SDK (agora-agent):

agent = AgoraAgent(
    client=client,
    instructions=SYSTEM_PROMPT,
    greeting="Hi! How can I help?",
    turn_detection={
        "mode": "default",
        "config": {
            "start_of_speech": {"mode": "manual"},   # client says when a turn opens
            "end_of_speech":   {"mode": "manual"},   # client says when it closes
        },
    },
    interruption={
        "enable": False,                              # voice alone never cuts the agent off
        "disabled_config": {"strategy": "append"},    # keep speech captured during agent talk
    },
    advanced_features={"enable_rtm": True},
    parameters={
        "data_channel": "rtm",                        # transcripts + agent state to the client
        "enable_metrics": True,
        "silence_config": {"timeout_ms": 0},          # no "are you still there?" nags
        "audio_scenario": "chorus",                   # optional: low-latency profile for web
    },
)

Raw /join equivalent: properties.turn_detection.config.start_of_speech.mode = "manual", …end_of_speech.mode = "manual", properties.interruption = {"enable": false, "disabled_config": {"strategy": "append"}}, advanced_features.enable_rtm = true, parameters.data_channel = "rtm".

Two server behaviours to know (measured, not documented):

  • manualEOS on a turn with no ASR text makes the LLM reply "your message didn't come

through". Never close an empty turn.

  • An open manual turn with no words is not auto-closed by the server (tested to 120 s).

False opens are harmless; leave them open and keep listening.

Step 2 — Add the Jev judge route (server)

Copy `server/jev_vad.py` next to your FastAPI app and mount it:

from jev_vad import router as jev_router
app.include_router(jev_router)

POST /judgeTurn takes the current user transcript plus a little context and returns:

{
  "turn_complete": 0.83,
  "taking_floor": null,
  "should_end": true,
  "should_interrupt": false,
  "thresholds": { "turn_complete": 0.72, "taking_floor": 0.5 },
  "latency_ms": 310
}

Two Jev nouls (yes/no judgments with a probability):

  • turn_complete — always. Prompted with the fact that it sees **text from a streaming

recognizer, not audio**: punctuation is guessed, fillers are normal, coughs are invisible, and silence_since_last_word_ms is the only prosody it gets.

  • taking_floor — only when agent_state is speaking/thinking. Gets

assistant_current_speech so it can tell an echo fragment or a backchannel from a real interruption, and treats an answer to a question the agent is asking as taking the floor.

If your web app proxies /api/* to the backend (as the quickstart does), add a rewrite for /api/judgeTurn.

Step 3 — Drop in the client controller (web)

Copy `web/jev-vad.ts` into your client and start it once the agent is connected and your microphone track is published:

import { JevVad } from './jev-vad'

const vad = new JevVad({
  agentUid,                                   // the agent's RTC uid (string)
  localUid,                                   // your user's RTC uid (string)
  micTrack: localAudioTrack.getMediaStreamTrack(),
  judgeUrl: '/api/judgeTurn',
  onStatus: (s) => setVadStatus(s),           // optional, for a status line
})
vad.start()

// on hang-up
vad.stop()

It subscribes to the toolkit's AGENT_STATE_CHANGED, TRANSCRIPT_UPDATED, USER_MANUAL_EOS and AGENT_MANUAL_EOS events and owns all manualSOS / manualEOS / interrupt calls. Nothing else in your app should call those.

What the controller does, in order

  1. Gate. Every 20 ms: RMS of the mic vs. a slowly tracked noise floor. Above

noiseFloor × 3 (×5 while the agent speaks, to reject TTS echo) for 110 ms → manualSOS. Remembers the last user utterance so the previous turn's text is not judged again.

  1. Judge. On each transcript update (debounced 150 ms; 0 ms during barge-in) →

POST /judgeTurn.

  1. Barge-in. If the turn opened while the agent was talking: should_interrupt →

interrupt(), otherwise hold ("backchannel") and wait for more words. In-flight judgments are not cancelled when new partials arrive — a verdict on a prefix of the text is acted on as soon as it lands, which is what makes interruptions feel immediate. Once the agent goes quiet, held words are re-judged as a normal turn so they are not lost.

  1. End. turn_complete ≥ bar(quiet) → wait until the mic has been quiet 400 ms →

manualEOS. Otherwise re-judge after 900 ms of silence, passing silence_ms so Jev sees the pause; the bar relaxes toward 0.55 as silence grows; 5 s of silence closes regardless.

  1. Reset on manualEOS, or when the server reports a manual-EoS result.

Tuning

KnobDefaultMove it when…
speechHoldMs110Clipping → lower. Too many false opens → raise.
gateFactor / gateFactorAgentSpeaking3 / 5Quiet speakers missed → lower. Room noise opens turns → raise.
eosThreshold0.72Agent jumps in early → raise. Feels sluggish on clear questions → lower.
eosThresholdFloor / eosDecayMs0.55 / 3000Complete-sounding sentences stall → lower floor or shorten decay.
confirmQuietMs400Agent replies while you draw breath → raise.
maxSilenceMs5000Hard cap. Rarely hit once the relaxing bar is in place.
JEV_BARGE_IN_THRESHOLD (server)0.5Backchannels interrupt → raise. Real interruptions ignored → lower.

Measured on real sessions: fragments score 0.04–0.15, complete short questions 0.67–0.80, "so I have a few questions… pricing" 0.50–0.54 (Jev correctly hears more coming). Thresholds are where you express your preference for speed vs. patience.

Pitfalls

  • Don't let anything else open or close turns. If the toolkit's default handlers or a

push-to-talk button also call manualSOS/manualEOS, the server's turn state and the controller's will diverge.

  • Don't judge the previous turn. When a new turn opens, the transcript still holds the

last utterance; the controller snapshots it as a baseline and ignores it. Skip this and the new turn closes instantly on stale text.

  • Never `manualEOS` an empty turn — see Step 1.
  • Echo. Without headphones, the agent's voice re-enters the mic. The stricter

gateFactorAgentSpeaking handles most of it; a residual echo fragment that reaches the STT is caught by taking_floor (it mirrors assistant_current_speech).

  • Threshold overrides ride with the request. eos_threshold in the body beats the env

default, so a UI slider can retune Jev mid-call with no server restart.

Verification

  1. Say a short question. The transcript's first word is intact (no clipping) and the agent

replies within about a second of you stopping.

  1. Say "So I have a question about…" and stop. The agent waits (Jev scores the sentence as

incomplete); after ~3 s of silence the relaxed bar closes the turn anyway.

  1. Cough or tap the desk. A turn may open, but no manualEOS fires and the agent stays quiet.
  2. While the agent is talking, say "uh-huh". It keeps talking. Say "wait, stop" — it stops.
  3. Let the agent ask you a question and answer over it with "pretty good". It stops and

takes the answer (answers to the agent's own question count as taking the floor).

  1. Backend log shows one judgeTurn line per transcript update with complete= and, during

agent speech, taking_floor= scores.

Non-goals

  • Replacing the STT. Jev judges text; the recognizer is still Agora's (any supported ASR vendor).
  • Server-side turn detection. This recipe deliberately moves the decision to the client so it

can use mic energy, agent state and the live transcript together.

  • Side-talk / multi-party awareness and "holding the floor" patience — see below; they are in

the full demo, not in the drop-in files.

Going further (what the full demo adds)

  • addressed_to_agent — a third noul that holds the turn when the user is clearly talking

to someone else in the room ("hey Mark, dinner's ready").

  • holding_floor — patience mode: "hmm, let me think" extends the silence cap to 12 s and

disables the relaxing bar, released by the latest words ("okay, go ahead").

  • A pure, unit-tested state machine instead of instance fields; a live tuning drawer that

persists to localStorage; and client-timeline mirroring into the server log for post-hoc analysis.

See the repository root README for the full architecture and the spike findings behind these numbers. Full-demo source: `web/src/lib/turn-machine.ts`, `web/src/hooks/useTurnController.ts`, `server/src/jev_eos.py`.