Teaching models to forget: Selective unlearning with Amazon Nova
Machine Learning Blog
This article introduces Reverse Direct Preference Optimization (rDPO), a novel unlearning technique powering Amazon Nova's Customizable Content Moderation Settings (CCMS) that allows organizations to selectively adjust AI safeguards for legitimate business use cases.
- rDPO reverses DPO preference pairs to simultaneously unlearn targeted behaviors while guiding models toward high-quality responses
- CCMS enables customization across four RAI pillars: Safety, Sensitive Content, Fairness, and Security
- LoRA adapters reduce deflection rates by up to 54 percentage points across policy categories
- Customized models retain near-baseline performance with less than 2 percentage point degradation on utility benchmarks
- Pre-trained adapters deploy via Amazon Bedrock with unique ARNs; no code changes required
- Amazon SageMaker AI supports DPO training for customers building custom unlearning experiments
rDPO enables organizations to overcome over-deflection in content moderation while maintaining core model capabilities and essential non-configurable safety controls.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.