Introducing Gemma 4 models on Amazon Bedrock
Machine Learning Blog
This article announces the availability of Gemma 4 models on Amazon Bedrock, a family of open-weight foundation models from Google DeepMind with built-in reasoning and multimodal capabilities.
- Three instruction-tuned variants: Gemma 4 31B (dense), 26B-A4B (mixture-of-experts), and E2B (compact) for different cost/latency profiles
- All variants support 256K token context windows, native function calling, multimodal text/image input, and built-in reasoning mode
- Access via bedrock-mantle endpoint using OpenAI-compatible APIs with no infrastructure provisioning required
- Service tiers available: Standard (on-demand), Priority (25% better latency), and Flex (discounted pricing)
- Implicit prompt caching automatically enabled to reduce latency on repeated prefixes
- Handle scaling with exponential backoff for 503 errors and gradual traffic ramps to avoid capacity issues
- Available in four AWS Regions with per-token pricing varying by model and service tier
Gemma 4 on Amazon Bedrock enables organizations to deploy open-weight models with full data protection, regulatory compliance, and operational control without managing infrastructure.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2025
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.