AWS has introduced a way to build real-time voice applications using vLLM-Omni on SageMaker AI. This setup allows models to stream speech as it is generated, reducing silent pauses.
The process uses the AWS vLLM-Omni Deep Learning Container (DLC) to deploy the Qwen3-TTS model. It streams text in and audio out over one WebSocket connection. This setup supports bidirectional communication, which is essential for interactive voice apps.
The vLLM-Omni project extends vLLM beyond just text. It can now process images, videos, and audio, making it suitable for multimodal applications. The DLC packages include routing middleware for SageMaker, simplifying deployment.
The example uses SageMaker instance pools with specific types like ml.g6.xlarge. It tries these types in order, based on availability and cost. You need enough quota for the selected instance types.
To try this setup, clone the sample repository, create a virtual environment, and install the required Python packages. The process is designed to be straightforward for developers.
This update improves the speed and responsiveness of voice applications by enabling real-time streaming. It helps developers build more natural and interactive voice tools.