Home icon

Training Azerbaijani language models on Amazon SageMaker AI

Machine Learning Blog



This article describes how Azercell Telecom built an Azerbaijani language model on Amazon SageMaker AI through a three-stage framework combining custom tokenization, continued pre-training, and fine-tuning.

  • Custom tokenizer reduced tokens per word by 50%, doubling context window capacity
  • FSDP and Liger Kernel optimizations achieved 58% lower peak GPU memory usage
  • Training throughput increased 23% on ml.p5.48xlarge instances
  • Continued pre-training adapted Llama 3.2 1B to understand Azerbaijani morphology
  • LoRA fine-tuning with 2,000 question-answer pairs created conversational assistant
  • Framework scales from 1B to larger models with configuration changes only

The solution demonstrates efficient LLM adaptation for morphologically rich, low-resource languages using AWS infrastructure and open-source tools.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Jun 25
2026
Optimize model training on Amazon SageMaker AI with NVIDIA Blackwell
May 20
2026
Build real-time voice applications with Amazon SageMaker AI and vLLM
Jun 3
2026
Amazon SageMaker AI launches multi-turn reinforcement learning for AI agent model customization
May 4
2026
Amazon SageMaker AI launches AI agent experience for model customization

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.