Background and Challenges
Voice interaction is evolving from a communication tool into an AI entry point. On one hand, public information from KOOK and YY Voice shows that voice rooms, cross-device connectivity, and low-latency co-hosting are already mature requirements. On the other hand, online TTS tools such as NowVoice and Airvoz are rapidly popularizing text-to-speech for dubbing, audio content, and accessibility scenarios. For developers, the real challenge is not building a standalone ASR or TTS demo, but combining voice input, large-model inference, and voice output into a stable, low-latency, interruptible real-time conversation pipeline.

Common pain points include: the system responds long after the user finishes speaking; the AI is mistakenly interrupted by background voices; ASR errors cause the model to answer irrelevantly; the TTS voice sounds natural but the first audio packet is too slow; and network jitter causes audio to break up. These problems must be addressed at the pipeline design level.
Core Concepts: A Real-Time Voice Conversation Pipeline
1. Audio Capture and Preprocessing
After the client captures PCM or Opus audio, it usually goes through echo cancellation, noise suppression, automatic gain control, and silence detection. If the AI is playing audio and the user suddenly speaks, echo cancellation directly affects whether ASR mistakenly recognizes the AI's own voice as user input.
2. VAD and Endpoint Detection
VAD determines whether the user is speaking, while endpoint detection decides when to send the speech segment to ASR. A real-time conversation should not wait until the user has paused for a long time before submitting, nor should it cut off too early during a normal pause.
3. Streaming ASR
ASR converts speech into text. Real-time pipelines usually upload audio via WebSocket or gRPC streaming and continuously receive intermediate and final results. In engineering, pay attention to sample rate, encoding format, hotwords, timestamps, and segmentation strategy.
4. LLM Conversation Orchestration
The LLM generates responses based on context. Voice scenarios are better suited to streaming output: first receive the initial token, then generate sentences progressively. If tool calls, knowledge retrieval, or safety review are involved, timeout and fallback strategies should be clearly defined in the orchestration layer.
5. Streaming TTS Synthesis
TTS converts text into speech. Public materials for products such as NowVoice and Airvoz often highlight capabilities like multiple voices, natural tone, and speed adjustment. However, in real-time conversations, first-packet latency, streaming return, text cleaning, and long-sentence splitting deserve even more attention.
6. Playback, Interruption, and State Machine
The conversation system needs clear states: idle, listening, thinking, speaking, and interrupted. When the user makes a valid interruption, playback should stop immediately, unplayed audio should be canceled, and the system should return to the listening state.
Practical Steps and Checklist
Step 1: Define the Latency Budget
For example, set the end-to-end target to 800 milliseconds to 1.5 seconds, and break it down across capture, network, ASR, LLM first token, TTS first packet, playback buffering, and other stages. Specific values vary by provider and deployment method; refer to official documentation.
Step 2: Make Everything Streaming
- ASR: Support recognition while speaking and return incremental text.
- LLM: Enable streaming generation to avoid waiting for a complete response.
- TTS: Support synthesis by sentence or segment, generating and playing at the same time.
Step 3: Design a TTS Chunking Strategy
Do not wait until the entire response is complete before synthesizing. Split by punctuation, semantics, and minimum length to avoid fragments that are too small and sound mechanical.
def split_for_tts(text):
parts = []
buf = ''
for ch in text:
buf += ch
if ch in '.!?;,':
if len(buf) >= 8:
parts.append(buf)
buf = ''
if buf:
parts.append(buf)
return parts
This example splits LLM output into fragments suitable for TTS. A production system can also consider complete semantics, maximum character count, and pause control.
Step 4: Implement Reliable Interruption Control
After detecting valid user speech, stop playback and cancel queued TTS tasks. To reduce false interruptions, add checks such as energy threshold, VAD confidence, and minimum speech duration.
Step 5: Establish Monitoring Metrics
- Time from when the user stops speaking to the final ASR text.
- LLM first-token latency.
- TTS first-packet audio latency.
- End-to-end response latency.
- False interruption rate, successful interruption rate, and reconnection rate.
Common Pitfalls and Recommendations
Pitfall 1: Focusing Only on Voice Quality While Ignoring Latency
High-quality voices are suitable for offline dubbing, but real-time conversations care more about whether the system responds promptly. Start with a low-latency voice, then improve audio quality based on user feedback.
Pitfall 2: Sending the Full Model Response to TTS at Once
This significantly increases waiting time. The right approach is: the model generates one sentence, TTS synthesizes one sentence, and the player buffers one sentence.
Pitfall 3: Overlooking Voice-Oriented Prompts
Voice responses should be short and conversational, avoiding long lists, code blocks, and complex tables. You can constrain output in the system prompt.
SYSTEM_PROMPT = (
'You are a voice assistant. Keep answers conversational and brief, '
'avoid lists and code. If unsure, say you need to confirm.'
)
This example helps the LLM generate responses that are more suitable for spoken playback.
Pitfall 4: Missing Text Normalization
Numbers, dates, English abbreviations, and emojis must be converted or filtered; otherwise TTS may read them incorrectly. For example, 10:30 should be read as a time, and AI can be pronounced as English letters or in Chinese depending on the product.
Pitfall 5: Lacking End-to-End Replay
It is recommended to record audio, ASR text, LLM responses, TTS audio, and state events. During review, this makes it possible to determine whether the issue lies in recognition, understanding, generation, or synthesis.
Directions for Further Reading
- Streaming ASR and VAD: Learn about public solutions such as WebRTC, Silero VAD, and RNNoise.
- Voice large-model architecture: Compare cascaded ASR + LLM + TTS with end-to-end speech models; refer to official documentation.
- TTS control: Pay attention to SSML, prosody, pauses, emotional voices, and multi-speaker synthesis.
- Real-time transport: WebSocket, WebRTC, Opus, jitter buffer, and packet loss recovery.
- Experience evaluation: WER, MOS, first-packet latency, task completion rate, and interruption naturalness.
Overall, combining voice with large models is not simply stitching together three APIs. It requires joint optimization across audio engineering, model serving, product interaction, and monitoring. First get the pipeline working, then continuously iterate around latency, interruption, voice quality, and context memory to build a truly usable real-time voice AI product.
References
Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.
Contact: chenxj.g@gmail.com
