Background and Challenges

Voice interaction is evolving from a communication tool into an AI entry point. On one hand, public information from KOOK and YY Voice shows that voice rooms, cross-device connectivity, and low-latency co-hosting are already mature requirements. On the other hand, online TTS tools such as NowVoice and Airvoz are rapidly popularizing text-to-speech for dubbing, audio content, and accessibility scenarios. For developers, the real challenge is not building a standalone ASR or TTS demo, but combining voice input, large-model inference, and voice output into a stable, low-latency, interruptible real-time conversation pipeline.

观山静思
观山静思

Common pain points include: the system responds long after the user finishes speaking; the AI is mistakenly interrupted by background voices; ASR errors cause the model to answer irrelevantly; the TTS voice sounds natural but the first audio packet is too slow; and network jitter causes audio to break up. These problems must be addressed at the pipeline design level.

Core Concepts: A Real-Time Voice Conversation Pipeline

1. Audio Capture and Preprocessing

After the client captures PCM or Opus audio, it usually goes through echo cancellation, noise suppression, automatic gain control, and silence detection. If the AI is playing audio and the user suddenly speaks, echo cancellation directly affects whether ASR mistakenly recognizes the AI's own voice as user input.

2. VAD and Endpoint Detection

VAD determines whether the user is speaking, while endpoint detection decides when to send the speech segment to ASR. A real-time conversation should not wait until the user has paused for a long time before submitting, nor should it cut off too early during a normal pause.

3. Streaming ASR

ASR converts speech into text. Real-time pipelines usually upload audio via WebSocket or gRPC streaming and continuously receive intermediate and final results. In engineering, pay attention to sample rate, encoding format, hotwords, timestamps, and segmentation strategy.

4. LLM Conversation Orchestration

The LLM generates responses based on context. Voice scenarios are better suited to streaming output: first receive the initial token, then generate sentences progressively. If tool calls, knowledge retrieval, or safety review are involved, timeout and fallback strategies should be clearly defined in the orchestration layer.

5. Streaming TTS Synthesis

TTS converts text into speech. Public materials for products such as NowVoice and Airvoz often highlight capabilities like multiple voices, natural tone, and speed adjustment. However, in real-time conversations, first-packet latency, streaming return, text cleaning, and long-sentence splitting deserve even more attention.

6. Playback, Interruption, and State Machine

The conversation system needs clear states: idle, listening, thinking, speaking, and interrupted. When the user makes a valid interruption, playback should stop immediately, unplayed audio should be canceled, and the system should return to the listening state.

Practical Steps and Checklist

Step 1: Define the Latency Budget

For example, set the end-to-end target to 800 milliseconds to 1.5 seconds, and break it down across capture, network, ASR, LLM first token, TTS first packet, playback buffering, and other stages. Specific values vary by provider and deployment method; refer to official documentation.

Step 2: Make Everything Streaming

  • ASR: Support recognition while speaking and return incremental text.
  • LLM: Enable streaming generation to avoid waiting for a complete response.
  • TTS: Support synthesis by sentence or segment, generating and playing at the same time.

Step 3: Design a TTS Chunking Strategy

Do not wait until the entire response is complete before synthesizing. Split by punctuation, semantics, and minimum length to avoid fragments that are too small and sound mechanical.

def split_for_tts(text):
    parts = []
    buf = ''
    for ch in text:
        buf += ch
        if ch in '.!?;,':
            if len(buf) >= 8:
                parts.append(buf)
                buf = ''
    if buf:
        parts.append(buf)
    return parts

This example splits LLM output into fragments suitable for TTS. A production system can also consider complete semantics, maximum character count, and pause control.

Step 4: Implement Reliable Interruption Control

After detecting valid user speech, stop playback and cancel queued TTS tasks. To reduce false interruptions, add checks such as energy threshold, VAD confidence, and minimum speech duration.

Step 5: Establish Monitoring Metrics

  • Time from when the user stops speaking to the final ASR text.
  • LLM first-token latency.
  • TTS first-packet audio latency.
  • End-to-end response latency.
  • False interruption rate, successful interruption rate, and reconnection rate.

Common Pitfalls and Recommendations

Pitfall 1: Focusing Only on Voice Quality While Ignoring Latency

High-quality voices are suitable for offline dubbing, but real-time conversations care more about whether the system responds promptly. Start with a low-latency voice, then improve audio quality based on user feedback.

Pitfall 2: Sending the Full Model Response to TTS at Once

This significantly increases waiting time. The right approach is: the model generates one sentence, TTS synthesizes one sentence, and the player buffers one sentence.

Pitfall 3: Overlooking Voice-Oriented Prompts

Voice responses should be short and conversational, avoiding long lists, code blocks, and complex tables. You can constrain output in the system prompt.

SYSTEM_PROMPT = (
    'You are a voice assistant. Keep answers conversational and brief, '
    'avoid lists and code. If unsure, say you need to confirm.'
)

This example helps the LLM generate responses that are more suitable for spoken playback.

Pitfall 4: Missing Text Normalization

Numbers, dates, English abbreviations, and emojis must be converted or filtered; otherwise TTS may read them incorrectly. For example, 10:30 should be read as a time, and AI can be pronounced as English letters or in Chinese depending on the product.

Pitfall 5: Lacking End-to-End Replay

It is recommended to record audio, ASR text, LLM responses, TTS audio, and state events. During review, this makes it possible to determine whether the issue lies in recognition, understanding, generation, or synthesis.

Directions for Further Reading

  • Streaming ASR and VAD: Learn about public solutions such as WebRTC, Silero VAD, and RNNoise.
  • Voice large-model architecture: Compare cascaded ASR + LLM + TTS with end-to-end speech models; refer to official documentation.
  • TTS control: Pay attention to SSML, prosody, pauses, emotional voices, and multi-speaker synthesis.
  • Real-time transport: WebSocket, WebRTC, Opus, jitter buffer, and packet loss recovery.
  • Experience evaluation: WER, MOS, first-packet latency, task completion rate, and interruption naturalness.

Overall, combining voice with large models is not simply stitching together three APIs. It requires joint optimization across audio engineering, model serving, product interaction, and monitoring. First get the pipeline working, then continuously iterate around latency, interruption, voice quality, and context memory to build a truly usable real-time voice AI product.

References

Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.

Contact: chenxj.g@gmail.com