Training Azerbaijani language models on Amazon SageMaker AI
Machine Learning Blog
This article describes how Azercell Telecom built an Azerbaijani language model on Amazon SageMaker AI through a three-stage framework combining custom tokenization, continued pre-training, and fine-tuning.
- Custom tokenizer reduced tokens per word by 50%, doubling context window capacity
- FSDP and Liger Kernel optimizations achieved 58% lower peak GPU memory usage
- Training throughput increased 23% on ml.p5.48xlarge instances
- Continued pre-training adapted Llama 3.2 1B to understand Azerbaijani morphology
- LoRA fine-tuning with 2,000 question-answer pairs created conversational assistant
- Framework scales from 1B to larger models with configuration changes only
The solution demonstrates efficient LLM adaptation for morphologically rich, low-resource languages using AWS infrastructure and open-source tools.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Jun 25
2026
2026
Optimize model training on Amazon SageMaker AI with NVIDIA Blackwell
May 20
2026
2026
Build real-time voice applications with Amazon SageMaker AI and vLLM
Jun 3
2026
2026
Amazon SageMaker AI launches multi-turn reinforcement learning for AI agent model customization
May 4
2026
2026
Amazon SageMaker AI launches AI agent experience for model customization
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.