This post details how to run a physical AI model factory on a persistent Amazon SageMaker HyperPod cluster on Amazon EKS. The pipeline includes synthetic data generation, post-training processes, and closed-loop evaluation.
NVIDIA Cosmos 3 is used for the evaluation component, enabling continuous assessment of model performance. GPU goodput is emphasized as the primary metric to measure system efficiency.
Running such a pipeline on a resilient, scalable cluster supports ongoing model development and deployment in physical AI applications. This approach helps maintain system robustness and performance.
