Agents1 min read
SageMaker HyperPod: Model Caching Reduces Cold Starts
Amazon SageMaker HyperPod now offers model caching, drastically reducing inference cold starts from minutes to seconds by pre-loading model weights. This improves the performance of large models deployed on the HyperPod.
From AWS machine learning blog

