Your Models. Any Scale. Full Control.
Customize, train, and deploy your models into production faster on a service that scales from serverless simplicity to frontier scale.
Trusted by leading enterprises, ISVs and AI research labs
Customize models faster on your data to improve accuracy and lower costs
Your developers need access to latest models and all of the customization techniques to make your models understand your business. SageMaker AI offers the broadest set of serverless model customization techniques within a governed lifecycle, where you simply describe what you need and an coding agent handles the rest. Scale from serverless to thousands of accelerators when your workloads demand it. You own the weights and choose where they run, serverless on Bedrock or dedicated inference endpoints on SageMaker.
Deploy models to production in hours on dedicated infrastructure you control.
Model deployments shouldn't take weeks of benchmarking, a separate inference stack, or inefficient GPU spend. SageMaker benchmarks across hundreds of instance types, returning performance-tested configurations in hours. Pack thousands of models behind one endpoint, backed by capacity-aware pools that fail over automatically. Intelligent routing and prefix-aware caching cut time to first token for optimal price-performance, while blue/green and rolling deployments with auto-rollback keep releases safe. SageMaker provides inference you can trust to perform.
Maximize Every GPU Hour Building Domain and Specialized Intelligence
Training at scale means GPU failures, stranded capacity, and weeks of cluster setup. SageMaker HyperPod is the only managed service with both Slurm and Kubernetes on the same resilient cluster, so you train and serve without re-platforming. Recover from faults in under 2 minutes, run Ray reinforcement learning and inference on fully open-source Ray clusters, and get the fastest GPU storage in the cloud with FSx for Lustre built-in. Govern the full journey from training to deployment with MLflow, Grafana, Prometheus, and CloudWatch.
Amazon SageMaker Studio
One environment for the full AI lifecycle, from data prep to deployment.
Your IDE, your choice
JupyterLab, VS Code, Kiro, and more. Connect from your local machine or use managed notebooks.
AI Governance
End-to-end AI Governance from data to model management to model ops and observability.
Built-in AI Coding Agent
Work with the coding agent you prefer, including Kiro, Claude Code, and Codex.
One-Click to Production
From notebook to production inference endpoint in the service with observability built in.
More time building models, less time managing infrastructure.
Proven at scale. Chosen by leaders.
Figma used SageMaker AI to accelerate model development, enabling rapid iteration on custom models that power intelligent design features for millions of users.
Figma
Design Platform
Twelve Labs built pioneering AI video intelligence on AWS, leveraging SageMaker AI to train and serve models that understand video content at scale.
Twelve Labs
Video AISageMaker AI let us turn isolated GPU nodes into a high-performance fabric. Our training jobs now complete in a fraction of the time, and we never lose progress to node failures.
Salesforce
Enterprise CRM
SageMaker HyperPod cut our GPU downtime and boosted GenAI productivity by 35%. With the addition of EKS support, we now integrate HyperPod directly into our Kubernetes training pipelines and package that capability into our GenAI platform so our customers can run their own training and fine-tuning workloads at scale.
Articul8 AI
GenAI PlatformRocket Mortgage uses SageMaker AI Pipelines to rapidly evaluate new open-source LLMs through automated validation, cutting assessment time and freeing data science teams to focus on innovation instead of infrastructure.
Rocket Mortgage
Financial ServicesCommon questions
Questions
Open allML teams at enterprises, ISVs, and AI-native startups that need to customize, train, deploy, and govern AI models without building and managing the underlying infrastructure.
Yes. Bring any model, any framework (PyTorch, TensorFlow, JAX), and any container. SageMaker also offers 1,000+ pre-optimized models in JumpStart.
Use Bedrock when you want serverless access to frontier and open-weight models through a single API with pay-per-token pricing. Use SageMaker AI when you need full control over model weights, training infrastructure, serving runtimes, and GPU types, with the ability to scale from serverless fine-tuning to thousands of accelerators.
Serverless inference via Bedrock, managed endpoints with multi-model packing and scale-to-zero, and HyperPod inference with Kubernetes-native control and unified training-inference clusters.
Governance is built into every stage: MLflow tracking, Model Registry with IAM-enforced approval gates, Clarify for bias detection, Model Monitor for drift, and compliance certifications including SOC 2, HIPAA, FedRAMP, and EU AI Act.
Start building with SageMaker AI.
Your team was hired to build AI models, not to manage infrastructure. Give them the service that handles the rest, from a single GPU to tens of thousands of accelerators.
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages