Home icon

Amazon Bedrock now supports observability of First Token Latency and Quota Consumption

News



This article announces two new CloudWatch metrics for Amazon Bedrock to improve observability of AI inference performance and resource usage.

  • TimeToFirstToken measures streaming API latency from request to first token receipt
  • EstimatedTPMQuotaUsage tracks Tokens Per Minute quota consumption across all inference APIs
  • Both metrics enable CloudWatch alarms for performance monitoring and quota management
  • Available in all commercial Bedrock regions with no client-side instrumentation required
  • Metrics update every minute for successfully completed requests at no additional cost

These new metrics provide deeper visibility into Bedrock inference performance and quota consumption without requiring code changes or opt-in configuration.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

May 28
2026
Amazon Bedrock expands support for Service Quotas
Apr 9
2026
Amazon Bedrock now supports cost allocation by IAM user and role
May 21
2026
Amazon Bedrock expands support for request-level usage attribution
Dec 23
2024
Amazon Bedrock Agents, Flows, and Knowledge Bases now supports Latency Optimized Models

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.