Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload
Machine Learning Blog
This article benchmarks OpenAI models on Amazon Bedrock against cost-efficient OpenAI API baselines, demonstrating that true model selection should prioritize cost per successful outcome rather than token pricing alone.
- Measured cost per correct answer on AIME, GPQA Diamond, and MMLU-Pro benchmarks across five models
- Analyzed agent trajectory costs on DeepSearchQA, showing turn efficiency drives input-token volume quadratically
- Evaluated professional deliverables using GDPval rubric-graded tasks with real occupational standards
- GPT-5.6-Luna on Amazon Bedrock achieved lowest cost per passing outcome across benchmarks after July 2026 repricing
- Provided open-source benchmarking harness for customers to evaluate models on their own workloads
- Decision framework recommends Luna for high-volume tasks, Luna/Terra for agentic workloads, and Sol for accuracy-critical work
Organizations should measure full cost of successes and failures, including token efficiency and turn count, rather than relying on pricing pages alone when selecting models.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.