What was announced
Alibaba's Qwen team released Qwen-Audio-3.1, a set of five audio models including Qwen-Audio-3.1-Realtime. This main model is a full-duplex speech system designed for voice agents that can think, act, and decide when to speak.
Model details
Qwen-Audio-3.1-Realtime supports a 262,000 token context window, with 245,000 tokens max input and 16,000 max output. It handles both text and audio as input and output. The model is available as a managed API on QwenCloud via WebSocket under the name qwen-audio-3.1-realtime-plus. No open weights were released.
Pricing is $6.4 per 1 million audio input tokens and $0.8 per 1 million text input tokens. Output tokens cost $24 per 1 million tokens, but text output tokens are free. The model supports function calling, web search, structured outputs, context caching, and fine-tuning.
Training and capabilities
The system runs two models sharing the same audio encoder and large language model design. One model decides whether to listen, speak, stop, or resume. The other converts speech to text. A voice renderer creates streaming speech based on conversation history and voice cues.
Training uses layers called Think, Act, and Speak and Coordinate. It applies on-policy distillation and reinforcement learning with human domain experts to improve empathy, intent, and acoustic scene understanding. The model also uses a tool pool and business policies to manage tasks.
Why it matters
The model reduces costs significantly and supports long context windows and real-time duplex interaction. It improves voice agent capabilities in managing conversations and calling tools.
What to do
Developers can access Qwen-Audio-3.1-Realtime via QwenCloud API for voice agent applications. Managers should consider the price cuts and new features when evaluating voice AI solutions.