Scaling Thomson Reuters’ language model research with Amazon SageMaker HyperPod
Machine Learning Blog
This article describes how Thomson Reuters scaled their research on training large language models (LLMs) using Amazon SageMaker HyperPod, a service that provides resilient and persistent clusters for distributed deep learning training.
Specifically, the article covers:
- The motivation for Thomson Reuters to train domain-adapted LLMs, including the limitations of public/commercial models like hallucinations, lack of domain knowledge, and speed/cost
- Thomson Reuters' research objectives around improving model performance on legal tasks using domain data and different training techniques like continuous pre-training and instruction fine-tuning
- Details on scaling model training using Amazon SageMaker HyperPod, including resilience features like deep health checks, automatic node replacement, and auto-resume
- Initial findings showing improvements in legal summarization and classification tasks with domain-adapted models compared to GPT-4
- Conclusion highlighting the benefits of Amazon SageMaker HyperPod for LLM training at scale
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Nov 21
2024
2024
Fine-tune large language models with Amazon SageMaker Autopilot
Sep 26
2024
2024
Scalable training platform with Amazon SageMaker HyperPod for innovation: a video generation case study
Aug 6
2024
2024
Large language models powered by Amazon Sagemaker Jumpstart available in Redshift ML
Oct 16
2024
2024
Scaling a Large Language Model with NVIDIA NIM on Amazon EKS with Karpenter
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.