Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference
News
This article announces Amazon SageMaker HyperPod Inference Gateway, a Kubernetes-native routing system for scalable LLM inference on existing SageMaker HyperPod infrastructure.
- Reduces first-token latency by up to 82% and p99 TTFT by 97–98% versus round-robin load balancing
- Envoy Endpoint terminates HTTPS traffic and exposes single private endpoint per cluster
- Body-Based Router reads model name from requests and routes to correct GPU pool without client changes
- Endpoint Picker scores pods in real-time using KV cache utilization, queue depth, LoRA adapter residency, and other signals
- Compatible with OpenAI-compatible servers like vLLM and SGLang with no code changes
- Per-cluster routing available now; cross-cluster, cross-region, and global rate limiting coming soon
The gateway enables efficient multi-model serving from a single endpoint with intelligent request routing based on real-time inference metrics.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Sep 18
2026
2026
Introducing Amazon SageMaker HyperPod Inference Gateway
Sep 11
2026
2026
Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts
Sep 10
2026
2026
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
Apr 22
2025
2025
Supercharge your LLM performance with Amazon SageMaker Large Model Inference container v15
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.