Home icon

Optimizing cost and latency with Amazon Bedrock prompt caching

Machine Learning Blog



This article explains how to use Amazon Bedrock prompt caching to reduce input token costs by up to 90% when repeatedly sending the same context to foundation models.

  • Cache static content like documents, system prompts, and tool definitions using cachePoint markers in the Converse API
  • Achieve approximately 75% savings on input token costs for repeated context by paying 25% more on first write and 90% less on subsequent reads
  • Implement mixed TTL caching to assign different expiration times (1 hour for reference material, 5 minutes for session context) to different content tiers
  • Use SHA-256 tenant ID prefixes to isolate cached content per tenant in multi-tenant applications without separate AWS accounts
  • Integrate prompt caching with LangChain using ChatBedrockConverse.create_cache_point() for framework-based applications
  • Monitor cacheWriteInputTokens and cacheReadInputTokens metrics to track cache performance and optimize TTL settings

Prompt caching enables significant cost and latency improvements for RAG applications, agentic workflows, and persona-based assistants by eliminating redundant token processing.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Apr 7
2025
Effectively use prompt caching on Amazon Bedrock
Dec 4
2024
Reduce costs and latency with Amazon Bedrock Intelligent Prompt Routing and prompt caching (preview)
Dec 4
2024
Amazon Bedrock announces preview of prompt caching
Apr 7
2025
Amazon Bedrock announces general availability of prompt caching

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.