Introducing new Ray capabilities on SageMaker HyperPod
Machine Learning Blog
This article announces new Ray capabilities on Amazon SageMaker HyperPod that integrate Ray with purpose-built infrastructure for foundation model training and serving on Kubernetes.
- Create and manage Ray clusters directly from SageMaker Studio console without writing YAML manifests or kubectl commands
- Access Ray Dashboard and Amazon Managed Grafana observability dashboards with IAM-authenticated remote endpoints
- Attach JupyterLab or Code Editor workspaces to Ray clusters for interactive development with native Ray driver access
- Automatic node recovery and hung job detection for resilient training without code changes
- Tiered checkpointing for faster recovery by checking HyperPod Tiered Storage before S3
- Deploy models from SageMaker JumpStart catalog directly to Ray Serve endpoints
- Managed Tiered KV Cache reduces inference latency for long-context LLM serving
- Works with open-source KubeRay and standard Ray APIs, so existing scripts run without modification
SageMaker HyperPod now provides a complete Ray development experience on EKS, simplifying distributed ML workloads from cluster creation through production inference.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Aug 24
2026
2026
Amazon SageMaker HyperPod enhances support for Ray
Jul 10
2025
2025
Amazon SageMaker HyperPod announces new observability capability
Sep 11
2026
2026
Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts
Sep 4
2026
2026
Run agent-driven Amazon SageMaker HyperPod operations with InstantStart
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.