
Voice AI · Proof of concept
Voice Calling Agent
Real-time streaming AI voice system
A proof-of-concept voice agent: the browser captures microphone audio and opens a WebRTC session with a FastAPI backend, which streams it to an AI provider (OpenAI or Gemini) and plays the voice reply back with live transcripts over WebSocket. A pluggable provider layer sits behind a unified interface, with mid-conversation tool calling (mocked in the repo) and handling of partial transcripts and interruptions. Per-call latency instrumentation logs each stage from WebRTC to AI response to tool call; the repo's tool calls run against a SQLite product catalogue.
Full-duplex · per-stage latency logging
Python / FastAPI / WebRTC / WebSockets / OpenAI / Gemini
