Home icon

AWS adds support for NIXL with EFA to accelerate LLM inference at scale

News



This article announces AWS support for NVIDIA Inference Xfer Library (NIXL) with Elastic Fabric Adapter (EFA) to accelerate disaggregated LLM inference on Amazon EC2.

  • NIXL with EFA increases KV-cache throughput and reduces inter-token latency
  • Enables efficient KV-cache movement between storage layers
  • Compatible with all EFA-enabled EC2 instances across AWS regions
  • Integrates natively with NVIDIA Dynamo, SGLang, and vLLM frameworks
  • Available at no additional cost with NIXL 1.0.0+ and EFA installer 1.47.0+

NIXL with EFA provides flexible, performant disaggregated LLM inference at scale on EC2.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Mar 16
2026
Introducing Disaggregated Inference on AWS powered by llm-d
Mar 10
2026
Accelerate custom LLM deployment: Fine-tune with Oumi and deploy to Amazon Bedrock
Apr 15
2026
Accelerating decode-heavy LLM inference with speculative decoding on AWS Trainium and vLLM
Feb 24
2026
Announcing AWS Elemental Inference

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.