Home icon

The state of AI for security: Measuring what matters most for building trust

Security Blog



This article introduces Deception Benchmark, a new evaluation framework measuring whether AI models can accurately distinguish real vulnerabilities from safe code that appears risky.

  • Deception Benchmark contains 14,822 code samples across 16 languages and 70+ CWE categories designed to test defensive precision
  • Evaluated 12 models from five providers; no model achieved both false positive and false negative rates below 10% production threshold
  • Direct prompting catches 95% of vulnerabilities but flags 41-99% of safe code; proof-of-exploit prompting reduces false positives but misses 7-44% of real vulnerabilities
  • Includes code-level challenges with subtle fixes and environment-gated challenges where infrastructure mitigations make exploits impossible
  • Labels generated through adversarial loops and multi-reviewer audit process to ensure quality and prevent memorization
  • Dataset and whitepaper released on GitHub for community evaluation of security AI tools

Deception Benchmark addresses the measurement gap in AI security by focusing on precision rather than vulnerability detection speed, helping teams evaluate whether AI-assisted security tools can be trusted in production.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

May 1
2026
Security posture improvement in the AI era
Aug 26
2026
Closing the AI agent trust gap with graduated autonomy
Dec 5
2024
Advancing AI trust with new responsible AI tools, capabilities, and resources
Nov 3
2025
New whitepaper available – AI for Security and Security for AI: Navigating Opportunities and Challenges

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.