One stop shop for running AI/ML on AWS
Docs · Available Images · Tutorials
AWS Deep Learning Containers (DLCs) are pre-built Docker images for running AI/ML workloads on AWS. Each image is tested and patched for security vulnerabilities. For more details, visit our documentation.
- [2026/10/01] vLLM Server ARM64 CPU v1.0 (AL2023) · EC2:
server-cpu-v1.0· SageMaker:server-sagemaker-cpu-v1.0· Initial release: vLLM0.30.0CPU backend on AWS Graviton (ARM64), no GPU required; bf16 kernels for Graviton 3 and later, RAM-based KV-cache defaults, and the OpenAI-compatible API on EC2 and SageMaker AI. - [2026/09/28] PyTorch v2.14.0 — EC2:
2.14-cu133-amzn2023· SageMaker:2.14-cu133-amzn2023-sagemaker· PyTorch 2.14.0 withtorchvision0.29.0; CUDA 13.3.1, EFA 1.50.0, TE 2.18.0, DeepSpeed 0.19.6. - [2026/09/26] vLLM Server v2.6 (AL2023) — EC2:
server-cuda-v2.6· SageMaker:server-sagemaker-cuda-v2.6· DeepEP v2 over EFA, built from theamazon-contributing/DeepEPfork; NCCL 2.31.2; EFA 1.50.0. vLLM stays at0.30.0(ec4a3a5). - [2026/09/26] vLLM-Omni v1.8 (AL2023) — EC2:
omni-cuda-v1.8· SageMaker:omni-sagemaker-cuda-v1.8· vLLM-Omni0.29.0rc1(up from 0.28.0); DeepEP v2 over EFA; FlashInfer 0.6.18, NCCL 2.31.2, EFA 1.50.0. - [2026/09/24] SGLang Server v1.4 (AL2023) — EC2:
server-cuda-v1.4· SageMaker:server-sagemaker-cuda-v1.4· SGLang0.5.19(up from 0.5.17); DeepEP v2 over EFA; PyTorch 2.13.0, NCCL 2.31.2, EFA 1.50.0. - [2026/09/23] vLLM v0.30.0 (Ubuntu) — EC2:
0.30.0-gpu-py312-ec2· SageMaker:0.30.0-gpu-py312· DeepSeek-V4.1-Flash, GLM-5.3-Flash, K2-Horizon; Fast Start GPU weight cache; Gumbel-max watermarking;transformerscapped below 5.17. - [2026/09/23] vLLM Server v2.5 (AL2023) — EC2:
server-cuda-v2.5· SageMaker:server-sagemaker-cuda-v2.5· vLLM0.30.0(up from 0.27.1), built fromec4a3a5for Transformers 5.17; FlashInfer 0.6.18.post1. - [2026/09/21] vLLM-Omni v1.7 (AL2023) — EC2:
omni-cuda-v1.7· SageMaker:omni-sagemaker-cuda-v1.7· vLLM-Omni0.28.0(up from 0.26.0); PEFT LoRA adapters validated on SageMaker — per-requestloraselection on/v1/images/generations; FlashInfer 0.6.16.post3, NCCL pinned 2.30.7, CUDA-13 mooncake wheel. - [2026/09/19] SGLang v0.5.20 (Ubuntu) — EC2:
0.5.20-gpu-py312-ec2· SageMaker:0.5.20-gpu-py312· GLM-5.3-Flash, Hy4-Preview, Qwen3.8-Flash-Next, K2 Horizon; RL sampling masks; unified radix tree; DeepSeek-V4 on Blackwell. - [2026/09/11] vLLM v0.29.0 (Ubuntu) — EC2:
0.29.0-gpu-py312-ec2· SageMaker:0.29.0-gpu-py312· New models: Hy4-preview (Tencent 770B/49B-active MoE with Gated DeepSeek Sparse Attention and native MTP), Qwen3.8-Flash-Next (BF16/FP8/NVFP4, MTP), GraniteSWA, GraniteMoeSWA, NemotronH_Omni_Reasoning_V3 (MTP), and Kimi K3 NVFP4 checkpoints. - [2026/09/07] SGLang v0.5.19 (Ubuntu) — EC2:
0.5.19-gpu-py312-ec2· SageMaker:0.5.19-gpu-py312· Qwen3.8, Ling-3.0, Spark2.5, Granite 4.2; beam search; DeepEP v2 MoE all-to-all. - [2026/09/01] Ray LLM v1.0 (2.58.0, AL2023) — EC2/EKS:
serve-llm-cuda-v1.0· Initial release: OpenAI-compatible LLM serving with Ray Serve and vLLM 0.26.0 on PyTorch 2.11.0 / CUDA 13.0.2 / Python 3.13;ray[llm]'sbuild_openai_appruns vLLM behind Ray Serve — single GPU on EC2, and multi-node serving on EKS via KubeRay. - [2026/09/01] Ray Train v1.1 (2.58.0, AL2023) — EC2/EKS:
train-ml-cuda-v1.1· EFA1.49.0(up from 1.47.0).
- [2026/04/28] We cannot guarantee security patching on Ubuntu-based vLLM and SGLang images due to the lack of Ubuntu Pro licensing. Customers may continue using these images at their own discretion and risk. We recommend migrating to our Amazon Linux-based images.
- [2026/02/10] Extended support for PyTorch 2.6 Inference containers until June 30, 2026
- PyTorch 2.6 Inference images will continue to receive security patches and updates through end of June 2026
- For complete framework support timelines, see our Support Policy
- Distributed Training on Amazon EKS - Configure and validate a distributed training cluster with DLCs on Amazon EKS.
- DLCs with Amazon SageMaker AI & MLflow - Use DLCs with SageMaker AI managed MLflow for experiment tracking and model management.
- LLM Serving on Amazon EKS with vLLM - Deploy and serve LLMs on Amazon EKS using vLLM DLCs.
- TorchServe to Ray Serve DLCs - Migrate TorchServe inference workloads to Ray Serve DLCs — AWS-managed, pre-tested containers without manual CUDA or version maintenance.
- Fine-tuning Meta Llama 3.2 Vision - Fine-tune and deploy Llama 3.2 Vision for web automation using DLCs, Amazon EKS, and Amazon Bedrock.
- DLCs with Amazon Q Developer and MCP - Streamline deep learning environments with Amazon Q Developer and Model Context Protocol.
- LLM Deployment on Amazon EKS - Deploy and optimize LLMs on Amazon EKS using vLLM DLCs. See also: Sample Code
This project is licensed under the Apache-2.0 License.