Gemini MLLM Voice Agent for Python
Build an end-to-end Gemini 3.8 Live MLLM voice agent with Python and Agora.
Use with Agora CLI
Clone the recipe and configure it with an Agora project.
agora init my-agora-gemini-mllm-python --recipe agora-gemini-mllm-pythonRecipe prompt
Paste into Cursor, Claude Code, v0, or your coding agentYou are implementing the "Gemini MLLM Voice Agent for Python" recipe in this project.
Read the recipe markdown first:
https://raw.githubusercontent.com/AgoraIO-Community/agora-gemini-mllm-python/main/docs/ai/RECIPE.md
Use the source repository for cross-reference:
https://github.com/AgoraIO-Community/agora-gemini-mllm-python
Build this recipe into the user's app using the markdown as the implementation guide. Inspect related source files through the repository links when the recipe points to them. Ask before installing new dependencies.Recipe
Rendered from the configured recipe markdown.
| Field | Value |
|---|---|
recipe_version | 1.0.0 |
recipe_status | experimental |
extension_points |
|
invariants |
|
stable_contracts |
|
Recipe Contract
This base recipe defines the reusable surface for a Python-backed Agora Conversational AI quickstart with a Next.js web client.
Recipe Role
- Role:
basequickstart recipe. - Target audience: developers bootstrapping a production-style Conversational AI app with a Python FastAPI backend and Next.js web client.
- Reuse model: keep this demo beside the local Python SDK checkout, configure three secrets, run, then customize the MLLM or UI.
Recipe Scope
This base recipe provides a copyable split-process starter with:
- Python FastAPI token generation and managed agent lifecycle.
- Next.js browser UI with RTC audio, RTM events, transcript, metrics, and connection status.
- Rewrite-only
/api/*browser facade that hides backend placement. - One Gemini 3.8 MLLM handles voice end to end; there is no separate STT, LLM, or TTS stage.
- Contract and local smoke verification that do not require live Agora calls.
Baseline Implementation Guidance
This Gemini demo is adapted from the Python-backed Agora quickstart. Use its source and progressive disclosure docs as the starting point for further customization.
Do not recreate Agora ConvoAI integration from memory. Provider schemas, SDK builder fields, token behavior, and RTM event details can drift. For a new baseline implementation, follow L1/L2/from_scratch_bootstrap.md while copying verified patterns from this repo.
Extension Points
| ID | Surface | How to extend | Required follow-up |
|---|---|---|---|
api.routes | server/src/server.py, web/next.config.ts, web/src/services/api.ts | Add FastAPI route, add rewrite, add browser fetch helper. | Extend web/scripts/verify-api-contracts.ts; add smoke coverage if the route belongs in local verification. |
agent.managed-config | server/src/agent.py | Change ADA_PROMPT, DEMO_GREETING, voice, or session options. The browser selects the public model and optional low/medium/high thinking level. | Run backend and local FastAPI checks; keep secrets in server/.env.example. |
web.conversation-ui | web/src/components/*, web/src/lib/conversation.ts | Customize pre-call, transcript, metrics, connection status, microphone, or visualizer UI. | Preserve RTC/RTM lifecycle ownership and transcript UID normalization. |
verification.contracts | web/scripts/*.ts, root package.json | Add contract checks for new browser/backend boundaries. | Keep checks runnable without live Agora credentials where possible. |
Invariants
- Browser code calls only
/api/get_config,/api/startAgent, and/api/stopAgentfor the default recipe flow. - Next.js owns the browser-facing
/api/*paths only through rewrites; do not addweb/app/api/**/route.tsfor agent or token logic. - FastAPI owns token generation,
AGORA_APP_CERTIFICATE, BYOK provider keys, and agent lifecycle. - The backend returns one RTC+RTM-capable token from
get_config; it must be issued for a concrete non-zero UID. LandingPage.tsxcoordinates config fetch, agent start, RTM login, token renewal, and call teardown.ConversationComponent.tsxowns RTC join, microphone publish,AgoraVoiceAIinitialization, transcript/state/metrics listeners, and explicit end-call media release.- The Google API key stays server-side; the Gemini MLLM and its prompt stay in the FastAPI agent configuration.
Stable Contracts
| Contract | Stable shape |
|---|---|
| Required backend env | AGORA_APP_ID, AGORA_APP_CERTIFICATE, GOOGLE_API_KEY |
| Backend port | Pass PORT only at launch; the default is 8000. |
| Required web deploy env | AGENT_BACKEND_URL |
GET /api/get_config | Query channel?, uid?; returns data.app_id, data.token, data.uid, data.channel_name, data.agent_uid. |
POST /api/startAgent | Body { channelName, rtcUid, userUid, model, thinkingLevel?, parameters? }; thinking is valid only for Extended Thinking; returns data.agent_id, data.channel_name, data.status. |
POST /api/stopAgent | Body { agentId }; returns { code: 0, msg: "success" }. |
| Success envelope | { "code": 0, "msg": "success", "data": ... } where route has data. |
| Verification entry points | bun run verify:web, bun run verify:backend, bun run verify:web:proxy, bun run verify:local:fastapi, bun run verify:local. |
Internal / Subject to Change
- Exact visual layout, component composition, Tailwind classes, and asset choices under
web/src/components/. - Gemini model IDs, thinking defaults, VAD timing, voice IDs, and prompt text; update docs and tests when changing them.
- In-memory
Agent._sessionsimplementation details; the stable behavior is start by channel/user and stop by returnedagent_id. - Verification implementation internals under
web/scripts/; the stable surface is the root script names and what they assert. agora-agentsSDK minor-version behavior; this recipe lower-bounds v2 but does not freeze every SDK field.
Related Progressive Disclosure Docs
L1/01_setup.md— setup, env, and command reference.L1/02_architecture.md— request flow and component topology.L1/05_workflows.md— common modification workflows.L1/06_interfaces.md— route, rewrite, env, and event contracts.L1/L2/from_scratch_bootstrap.md— implementation map for recreating the Python-backed quickstart recipe.L1/L2/managed_agent_config.md— full agent config detail.L1/L2/session_lifecycle.md— RTC/RTM/session orchestration.