Problem Statement: A Conversational Loop Measured in Hundreds of Milliseconds
Frames the product as a latency-bound streaming pipeline, not a chat backend with a microphone attached.
Problem statement
Design a real-time voice assistant: the user speaks, the system streams audio into a speech recognizer, converts the transcript into intent with a large language model, optionally calls external skills or APIs, and answers with synthetic speech while preserving multi-turn context. The product is not transcription plus chat plus TTS bolted together; it is a single conversational loop whose quality is judged by how natural the silence feels. Humans perceive a response gap above roughly 800 ms to 1.2 s as sluggish, and interruptions (barge-in) must cancel in-flight synthesis within a couple of hundred milliseconds or the assistant talks over the user.
The defining architectural fact is that every stage is streaming. Audio arrives in 20-40 ms frames; the ASR emits partial hypotheses that revise themselves; the LLM emits tokens before it has finished thinking; the TTS emits phoneme chunks before the sentence is complete. A batch-oriented design (wait for full utterance, wait for full answer, then synthesize) adds one to three seconds of avoidable latency. Therefore the system is a chain of incremental producers and consumers with explicit backpressure, cancellation, and revision semantics.
Why the problem is distinctive
A text chatbot retries a slow token stream invisibly. A voice assistant cannot: silence is audible, overlap is rude, and a revised transcript may invalidate tokens already spoken. The design must therefore model three coupled state machines: the audio turn (who holds the floor), the semantic turn (what the user meant), and the response turn (what has been committed aloud). It must also decide where each function lives: client-side echo cancellation and VAD, edge media transport, regional GPU pools for ASR, LLM and TTS, and durable services for session memory, skills, billing, and safety.
Two product architectures exist in the wild. The cascaded pipeline (streaming ASR, then LLM, then streaming TTS) gives composability, inspectability, per-stage vendor choice, and text logs for safety review. The native speech-to-speech model (a single multimodal model mapping audio to audio, as OpenAI described for GPT-4o-class systems) removes two serialization points and preserves prosody, but weakens transcript-level control, tool-call auditability, and vendor substitution. This answer designs the cascaded pipeline as the default because an interview must show control over every latency budget, and treats native speech-to-speech as an explicit trade-off.
Public operating baseline versus design assumptions
Public evidence shows the category is real and measured. OpenAI published the Realtime API in 2024 with WebSocket and WebRTC transports, server-side VAD and turn detection, and GPT-4o-class audio-to-audio latency in the low hundreds of milliseconds in its launch material. Amazon disclosed Alexa-scale voice infrastructure with on-device wake word and cloud ASR/NLU/TTS, and publicly reported more than 100 million active Alexa users in 2019. Google demonstrated full-duplex telephone conversation with Duplex in 2018 and ships Gemini Live streaming multimodal assistants. NVIDIA Riva documents GPU-streaming ASR and TTS pipelines. These are cited context, not requirements.
For capacity planning this answer explicitly assumes a mature consumer platform: 30 million DAU, 6 million voice sessions per day, average session 6 minutes with 12 turns, 200,000 concurrent sessions at an 8x evening peak, and a 900 ms p50 voice-to-voice latency target. Unless a number is tied to a citation, it is a stated design assumption, budget, or target.
Key Highlights
- •The product is a single streaming loop: audio frames in, partial transcripts, tokens, phoneme chunks out; batch processing adds 1-3 s of audible silence.
- •Three coupled state machines: audio floor (who speaks), semantic turn (what was meant), response turn (what was committed aloud).
- •Cascaded ASR to LLM to TTS is the interview-default architecture; native speech-to-speech is a named trade-off, not the baseline.
- •Public anchors: OpenAI Realtime API (2024, WebSocket/WebRTC, server VAD), Alexa at 100M+ active users (2019), Google Duplex (2018) full-duplex demo.
- •Assumed scale: 30M DAU, 6M sessions/day, 200K concurrent at peak, 900 ms p50 voice-to-voice; every uncited number is labeled an assumption.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "Let me separate the conversational loop from the durable services before choosing any model or vendor."
- "I will treat silence and overlap as product defects with measurable budgets, not as cosmetic issues."