Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors
Machine Learning Blog
This article demonstrates how to implement Amazon S3 Vectors as a persistent memory backend for the NVIDIA NeMo Agent Toolkit, enabling multi-agent AI systems to share and retrieve knowledge across sessions.
- NAT's memory subsystem uses a MemoryEditor interface with add_items(), search(), and remove_items() methods for extensible backends
- Amazon S3 Vectors provides semantic retrieval, metadata filtering, strong consistency, and elastic scaling up to 2 billion vectors
- Implement a custom S3VectorsMemoryEditor plugin that generates embeddings with Amazon Titan and stores vectors with scoped metadata
- Configure auto_memory_agent workflow to automatically capture and retrieve memory without explicit LLM tool invocation
- Deploy agents on Amazon EKS using Kubernetes Deployments with IAM Roles for Service Accounts (IRSA) for S3 Vectors access
- Multi-agent coordination uses metadata filters (agent_id, team_id, ticker) to enable shared knowledge within the same index
- Memory consolidation distills episodic memories into semantic knowledge to maintain retrieval precision and efficiency
- Evaluate memory impact using NAT's evaluation harness to measure improvements in groundedness, token usage, and latency
This pattern enables production multi-agent systems to build cumulative knowledge across sessions, reducing redundant work and improving response quality through persistent, semantically queryable memory.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2025
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.