Home icon

Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod

Machine Learning Blog



This article demonstrates how to build a Physical AI model factory using NVIDIA Cosmos 3 on Amazon SageMaker HyperPod, enabling continuous pipelines of synthetic data generation, model post-training, and evaluation for robots and autonomous vehicles.

  • Cosmos 3 uses a Mixture-of-Transformers design with per-layer joint attention and asymmetric training-versus-inference to unify world models, action labelers, and policies
  • One persistent GPU cluster with shared storage replaces separate stage-specific clusters, improving GPU goodput across the entire flywheel
  • Amazon SageMaker HyperPod on EKS provides health-checked auto-recovering capacity, pre-configured NCCL over EFA, and unified Kubernetes orchestration
  • FSx for Lustre with EFA enables multi-terabyte dataset sharing across generation, post-training, and evaluation stages
  • Distributed post-training demonstrated on robot-policy, vision SFT, and vision LoRA workloads with near-linear scaling efficiency (0.97–0.99) across 1–4 nodes
  • Integrated observability dashboard unifies cosmos-framework metrics and GPU telemetry for measuring GPU goodput rather than peak throughput
  • Resilient checkpointing with auto-resume recovers from node failures with minimal lost work, bounded by checkpoint interval

The reference architecture and reproducible methodology enable teams to operate Physical AI pipelines continuously and economically by treating the entire flywheel as one shared, persistent cluster rather than separate provisioning stages.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Sep 17
2026
Amazon SageMaker AI now supports serverless model customization for NVIDIA Nemotron 3.5 Lightning
Feb 23
2026
Accelerating AI model production at Hexagon with Amazon SageMaker HyperPod
Sep 4
2026
Run agent-driven Amazon SageMaker HyperPod operations with InstantStart
Nov 24
2025
Amazon SageMaker HyperPod now supports NVIDIA Multi-Instance GPU (MIG) for generative AI tasks

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.