Home icon

Introducing Amazon SageMaker HyperPod Inference Gateway

Machine Learning Blog



This article announces Amazon SageMaker HyperPod Inference Gateway, a Kubernetes-native routing system that optimizes GPU utilization for large language model inference at scale.

  • GPU-aware routing uses real-time Prometheus metrics to place requests on best-suited pods, reducing first-token latency by up to 82%
  • Deploys as single EKS managed addon with zero application changes; uses Envoy Gateway, Body-Based Router, and Endpoint Picker components
  • Intelligently routes based on KV cache utilization, queue depth, LoRA adapter residency, and prefix cache hit rates
  • Supports multi-model routing and LoRA adapter affinity without application code modifications
  • Tier 2 Global Inference Router (coming soon) adds cross-cluster failover and cost-aware traffic shaping
  • Benchmarks show up to 98% latency reduction on mixed GPU generations and bursty traffic scenarios

The gateway eliminates GPU waste from naive round-robin routing by providing intelligent, metrics-driven request placement on Kubernetes clusters.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Sep 11
2026
Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts
May 20
2026
Amazon SageMaker HyperPod now supports data capture for inference workloads
Jul 9
2026
Enhancing enterprise inference on Amazon SageMaker HyperPod with data capture, Hugging Face, NVMe, and Route 53 integration
Sep 10
2026
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.