AWS SageMaker AI now supports generating images and videos with vLLM-Omni models. It deploys two endpoints from the same container: one for image generation and one for video. The image endpoint runs in real-time, returning a PNG image directly. The video endpoint runs asynchronously, saving the MP4 file to Amazon S3.
The process starts with a text prompt sent to the image endpoint, which returns a base64 PNG. The image is resized and inserted into a video request. The video request is uploaded to Amazon S3, then processed by the video endpoint. Once finished, the MP4 appears in S3 for download.
Both models use the same container image, reducing complexity. The image model uses a real-time endpoint, while the video model uses an asynchronous endpoint. This setup matches the different response times needed for each task.
To try this, clone the sample repository, install dependencies, and deploy the endpoints with your AWS role. The sample code includes a command-line workflow and a Streamlit app for easy testing. You can choose instance types based on your workload and budget.
This update improves how models generate multimedia content, making it easier to build applications that create images and videos from text prompts. It also shows how to manage different workload types with separate endpoints and storage in Amazon S3.