Voice AI
All recipes
PythonTypeScriptIntermediate

Speaker Lock

Lock onto the primary speaker and ignore other voices.

Recipe prompt

Paste into Cursor, Claude Code, v0, or your coding agent
Use with your coding agent
You are implementing the "Speaker Lock" recipe in this project.

Read the recipe markdown first:
https://raw.githubusercontent.com/AgoraIO-Conversational-AI/recipe-agent-voiceprint/main/README.md

Use the source repository for cross-reference:
https://github.com/AgoraIO-Conversational-AI/recipe-agent-voiceprint

Build this recipe into the user's app using the markdown as the implementation guide. Inspect related source files through the repository links when the recipe points to them. Ask before installing new dependencies.

Recipe

Rendered from the configured recipe markdown.

Raw

Agora Conversational AI — Voiceprint (Speaker Lock) Recipe (Python)

![License: MIT](./LICENSE) ![Python](https://www.python.org/) ![Bun](https://bun.sh/)

The voiceprint / speaker-lock recipe in the Agora Conversational AI recipes family. The agent auto-locks onto the primary speaker and suppresses other voices and background noise using Agora Speaker Lock (sal_mode: "locking"). Fully zero-key — no voiceprint enrollment required, and OpenAI is Agora-managed (no OPENAI_API_KEY needed unless you bring your own account).

Pipeline: DeepgramSTT(nova-3)OpenAI (plain assistant) → MiniMaxTTS

Speaker Lock (sal_mode: "locking") — the SDK auto-locks onto the first clear speaker detected in the channel and suppresses all other voices and background noise. No enrollment step, no voiceprint file required.

Optional named-speaker mode (not used here): The SDK also supports sal_mode: "recognition", which locks onto a specific named speaker. This mode requires a pre-hosted 16 kHz / 16-bit mono PCM voiceprint (≤ 2 MB) supplied via sample_urls. The SDK has no voiceprint-enrollment API — you must host the PCM file yourself. This recipe uses only "locking" (zero-key, no enrollment).

Prerequisites

Run It

# 1. Install web deps + create the Python venv
bun run setup

# 2. Add Agora credentials (CLI), or edit server/.env.local by hand
agora login
agora project use <your-project>          # select which project to use
agora project env write server/.env.local # writes App ID + Certificate

# 3. Run backend + web
bun run dev

Open http://localhost:3000Start Conversation → speak.

Working from a clone

If you cloned this repo (rather than scaffolding via the Agora CLI), the steps above are complete as written: bun run setup creates the Python venv and installs web dependencies, then bun run dev brings up both services. You still need Agora credentials in server/.env.local before a conversation can connect.

Services:

  • Frontend — http://localhost:3000
  • Backend — http://localhost:8000
  • Mock LLM — N/A (managed OpenAI, no local service)
  • API docs — http://localhost:8000/docs

Deploy

Deploy web (Next.js) and server (a reachable FastAPI backend). Set AGENT_BACKEND_URL in the web deployment so the Next rewrites reach the backend.

A backend-only Docker image is published to ghcr.io/AgoraIO-Conversational-AI/recipe-agent-voiceprint on v* tags. It exposes BACKEND-ONLY (:8000). No separate LLM container is needed — OpenAI is Agora-managed.

Environment variables

Backend env file: `server/.env.example`.

VariableRequiredDefaultNotes
AGORA_APP_IDAgora Console → Project → App ID
AGORA_APP_CERTIFICATEAgora Console → Project → App Certificate
OPENAI_MODELgpt-4o-miniOpenAI model
OPENAI_API_KEYOptional — Agora manages the OpenAI key by default (keyless). Set only if your account requires it.
TTS_VOICEEnglish_captivating_female1MiniMax TTS voice
AGENT_GREETINGbuilt-inOptional opening line override

Commands

bun run setup            # install web deps + create server/ venv
bun run dev              # run backend (:8000) + web (:3000)

bun run doctor           # prerequisite check (no creds needed)
bun run doctor:local     # + .env.local + credentials checks

bun run verify           # web-only gate (no Agora creds needed)
bun run verify:local     # full local gate: backend compile + smoke tests + web build
bun run clean            # remove venvs and build artifacts

Tests run standalone (no Agora cloud needed): pytest in server/, plus bun run verify in web/. CI runs them on Linux/macOS/Windows × Python 3.10 & 3.13.

Architecture

Browser (localhost:3000)
  │  fetch /api/*
  ▼
Next.js  ──rewrite──▶  Agent backend  (server/, localhost:8000)
                          │  starts agent session (managed OpenAI vendor)
                          │  sal={"sal_mode": "locking"}  ← Speaker Lock
                          ▼
                       Agora ConvoAI Cloud
                          │  Deepgram STT (managed, nova-3)
                          │  Speaker Lock — locks onto primary speaker, suppresses others
                          │  OpenAI assistant (Agora-managed, keyless)
                          │  MiniMax TTS (managed)
                          ▼
                       User hears agent focused on their voice only

No separate llm/ service — OpenAI is Agora-managed and requires no API key. See ARCHITECTURE.md.

What You Get

  • A Next.js web client (:3000) that drives the RTC/RTM lifecycle and only ever calls /api/*.
  • A FastAPI agent backend (:8000) that owns Agora token generation and the agent session lifecycle.
  • The /api/get_config · /api/startAgent · /api/stopAgent contract between the web client and the backend (Next rewrites, no Route Handlers).
  • Speaker Lock (sal_mode: "locking") wired via sal=build_sal() on AgoraAgent — no enrollment, no extra credentials.
  • Managed keyless OpenAI as a plain conversational assistant — Agora-managed, no OPENAI_API_KEY required.
  • Zero-key setup — the full pipeline runs with only Agora credentials.

How It Works

  1. The browser calls /api/get_config, which Next rewrites to the backend; the

backend mints an Agora token from AGORA_APP_ID + AGORA_APP_CERTIFICATE.

  1. The browser joins the RTC channel, then calls /api/startAgent; the backend

starts an agent session with sal={"sal_mode": "locking"} on AgoraAgent.

  1. Agora's Speaker Lock detects the first clear speaker in the channel and locks onto

that voice, suppressing other voices and background noise automatically.

  1. Deepgram STT transcribes the locked speaker's audio.
  2. Agora's managed OpenAI stage replies with a concise assistant response.
  3. MiniMax TTS speaks the response back into the channel.
  4. /api/stopAgent ends the session.

Repo Map

  • web/ — Next.js frontend (:3000); RTC/RTM lifecycle and UI.
  • server/ — FastAPI agent backend (:8000); Agora tokens + agent lifecycle.
  • server/src/sal_config.py — pure builder for the SAL (Speaker Lock) config dict.
  • ARCHITECTURE.md — system shape and component boundaries.
  • AGENTS.md — guide for coding agents working in this repo.

Troubleshooting

ProblemFix
Agent does not lock onto my voiceEnsure only one speaker is active at the start; Speaker Lock latches onto the first clear voice.
Local calls fail under a global proxy (Clash, etc.)Configure your proxy to send 127.0.0.1, localhost, and RFC-1918 ranges DIRECT.

More Docs

License

Released under the MIT License.