Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod
Machine Learning Blog
This article demonstrates how to build a Physical AI model factory using NVIDIA Cosmos 3 on Amazon SageMaker HyperPod, enabling continuous pipelines of synthetic data generation, model post-training, and evaluation for robots and autonomous vehicles.
- Cosmos 3 uses a Mixture-of-Transformers design with per-layer joint attention and asymmetric training-versus-inference to unify world models, action labelers, and policies
- One persistent GPU cluster with shared storage replaces separate stage-specific clusters, improving GPU goodput across the entire flywheel
- Amazon SageMaker HyperPod on EKS provides health-checked auto-recovering capacity, pre-configured NCCL over EFA, and unified Kubernetes orchestration
- FSx for Lustre with EFA enables multi-terabyte dataset sharing across generation, post-training, and evaluation stages
- Distributed post-training demonstrated on robot-policy, vision SFT, and vision LoRA workloads with near-linear scaling efficiency (0.97–0.99) across 1–4 nodes
- Integrated observability dashboard unifies cosmos-framework metrics and GPU telemetry for measuring GPU goodput rather than peak throughput
- Resilient checkpointing with auto-resume recovers from node failures with minimal lost work, bounded by checkpoint interval
The reference architecture and reproducible methodology enable teams to operate Physical AI pipelines continuously and economically by treating the entire flywheel as one shared, persistent cluster rather than separate provisioning stages.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2025
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.