Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6
Machine Learning Blog
This article benchmarks LLM inference performance on AWS SageMaker AI across G5, G6, and G7 GPU instances, demonstrating that G7 delivers superior throughput, latency, and cost-per-token for Mixture-of-Experts models.
- G7 instances with NVIDIA Blackwell GPUs achieve 60.8% higher throughput than G6 and 13% higher than G5 for Qwen3-Coder-30B
- G7 reduces average request latency by 37.6% versus G6 and P99 latency by 54.7% for the same workload
- Native NVFP4 4-bit quantization support on G7 provides structural advantage for Mixture-of-Experts model deployment
- Amazon SageMaker AI Generative AI Inference Recommendations automates benchmarking and configuration selection across GPU options
- G7.2xlarge achieves lowest cost per output token ($0.90/1M for chat, $2.29/1M for RAG) while G7.48xlarge maximizes throughput at 2,397 tokens/second
- Two complementary workflows: direct benchmarking with LMI for existing endpoints and automated recommendations for new deployments
G7 instances consistently outperform prior GPU generations for LLM inference, delivering measurable gains in throughput, latency, and cost efficiency across diverse workload profiles.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2025
2026
2025
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.