Deploying Small Language Models at Scale with AWS IoT Greengrass and Strands Agents
Internet of Things Blog
This article demonstrates deploying Small Language Models (SLMs) at scale on edge devices using AWS IoT Greengrass and Strands Agents for real-time industrial automation.
- SLMs (3-15B parameters) run locally on constrained hardware for immediate responses without cloud dependency
- AWS IoT Greengrass deploys SLMs alongside Lambda functions directly to OPC-UA gateways
- Strands Agents SDK provides multi-agent orchestration with specialized agents for documentation and machine data
- S3FileDownloader component handles large model file distribution with automatic retry capabilities
- Ollama loads GGUF-format models and runs local inference on edge devices
- Hybrid architecture: edge handles real-time tasks, cloud manages fleet analytics and model retraining
- Solution includes OPC-UA simulator, maintenance runbooks, and MQTT-based query interface
- Applicable across manufacturing, automotive, energy, gaming, and education sectors
This approach enables secure, responsive AI-powered decision-making at the edge while maintaining cloud integration for complex analytics and model optimization.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Oct 30
2025
2025
Announcing an AI agent context pack for AWS IoT Greengrass developers
Nov 11
2025
2025
Deploying mission-critical edge applications with AWS IoT Greengrass in disconnected environments
Oct 27
2025
2025
Building large language models for the public sector on AWS
Jun 5
2025
2025
Run small language models cost-efficiently with AWS Graviton and Amazon SageMaker AI
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.