Home icon

Running GenAI Inference with AWS Graviton and Arcee AI Models

AWS Partner Network Blog



This article discusses running generative AI inference using AWS Graviton processors and Arcee AI's Small Language Models (SLMs).

  • AWS Graviton processors offer improved price performance for AI/ML workloads
  • Arcee AI provides Small Language Models tailored for business use cases
  • Key optimization techniques include:
    • Model quantization (reducing precision from 16-bit to 4-bit)
    • Using llama.cpp for model optimization
  • Experimental results show 4-bit quantized models can be:
  • 1.6x faster at 8-bit models
  • 2.5x faster than 16-bit models
  • With minimal performance degradation

The article encourages developers to experiment with Graviton instances and Arcee AI models to optimize generative AI inference performance.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Mar 19
2025
Run GenAI inference across environments with Amazon EKS Hybrid Nodes
Jun 12
2025
GenAI in Factor Modeling Data Pipelines: A Hedge Fund Workflow on AWS
Jun 5
2025
Run small language models cost-efficiently with AWS Graviton and Amazon SageMaker AI
May 14
2025
Cost-effective AI image generation with PixArt-Σ inference on AWS Trainium and AWS Inferentia

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.