Alibaba's Qwen team released Qwen3.8-LiveTranslate, a real-time interpretation model. It listens to live speech and optional video frames, then returns translated text and speech.
The main change is a new Interleave architecture. This reduces the average lag from 2.8 seconds to 2.3 seconds, an 18% decrease. The model improves faithfulness, fluency, and conciseness.
Qwen3.8-LiveTranslate is available as a hosted API. It runs on Alibaba Cloud Model Studio and QwenCloud, accessible via WebSocket. It supports 60 languages, speaking 29 and providing text for the rest.
Features include real-time speaker diarization, which distinguishes speakers and preserves voices. It also offers synchronized bilingual display and long-context disambiguation, keeping names and terms consistent.
Inputs can be audio and images, with visual cues helping in noisy environments. Developers can set hotwords to fix translations for specific terms. The API uses a default audio quality of 16 kHz input and 24 kHz output.
Pricing varies by location, with costs around $1.54 per hour for speech processing. The model handles a context window of over 53,000 tokens and limits requests to 10 per minute.
Qwen3.8-LiveTranslate improves translation speed and quality with new architecture and features. It supports multiple languages and is accessible via API for real-time applications.
Source: https://www.marktechpost.com/2026/09/19/alibaba-qwen-team-releases-qwen3-8-livetranslate/



