Skip to content

Repository files navigation

AWS Logo

AWS Deep Learning Containers

One stop shop for running AI/ML on AWS

Docs · Available Images · Tutorials

Auto Release - PyTorch 2.13 Auto Release - TensorFlow 2.21 Auto Release - vLLM Auto Release - vLLM-Omni Auto Release - SGLang Auto Release - Ray Auto Release - Base cu130 Auto Release - Base cu132


About

AWS Deep Learning Containers (DLCs) are pre-built Docker images for running AI/ML workloads on AWS. Each image is tested and patched for security vulnerabilities. For more details, visit our documentation.


🔥 What's New

🚀 Release Highlights

  • [2026/08/03] vLLM Server v2.2 (AL2023) — EC2: server-cuda-v2.2 · SageMaker: server-sagemaker-cuda-v2.2 · vLLM 0.26.0 (up from 0.24.0); FlashInfer 0.6.15.post1; DeepEP EPv2/GIN backend (NCCL pinned to 2.30.7); Inkling (piecewise CUDA graph, MTP speculative decoding, LoRA, NVFP4), Cosmos3 Edge Reasoner, TranslateGemma-12b-it, BertForMaskedLM.
  • [2026/07/31] Base cu132 (CUDA 13.2, AL2023) — EC2: devel-cu132-amzn2023 · runtime-cu132-amzn2023 · CUDA 13.2.1 with Python 3.13.12 (built from source) and uv pre-installed; devel and runtime variants; devel bundles the multi-node stack (GDRCopy, NCCL 2.29.7, EFA installer).
  • [2026/07/26] vLLM v0.26.0 (Ubuntu) — EC2: 0.26.0-gpu-py312-ec2 · SageMaker: 0.26.0-gpu-py312 · Inkling (piecewise CUDA graph, MTP speculative decoding, LoRA, NVFP4), Cosmos3 Edge Reasoner, TranslateGemma-12b-it, BertForMaskedLM; DeepSeek-V4 routing-kernel and fused_topk_bias speedups; fp32 lm_head via head_dtype; per-KV-cache-group attention backends.
  • [2026/07/26] SGLang v0.5.16 (Ubuntu) — EC2: 0.5.16-gpu-py312-ec2 · SageMaker: 0.5.16-gpu-py312 · Inkling day-0 support, LongCat 2.0 FP8, JetBrains Mellum v2, Pi0.5; DSpark speculative decoding (--speculative-algorithm DSPARK); UnifiedRadixTree now the default for SWA, Mamba, and DSA models.
  • [2026/07/20] PyTorch v2.13.0 — EC2: 2.13-cu133-amzn2023 · SageMaker: 2.13-cu133-amzn2023-sagemaker · Amazon Linux 2023 with EFA, flash-attn, and Transformer Engine; PyTorch 2.13.0 with CUDA 13.3.0, NCCL 2.30.7, TE 2.17.0, DeepSpeed 0.19.2.
  • [2026/07/15] vLLM v0.25.1 (Ubuntu) — EC2: 0.25.1-gpu-py312-ec2 · SageMaker: 0.25.1-gpu-py312 · Patch release: defer TorchCodec FFmpeg import error to runtime (unblocks startup without system FFmpeg); guard mixed-dtype allreduce RMSNorm quant fusions (fixes NVFP4 garbage output).
  • [2026/07/13] vLLM v0.25.0 (Ubuntu) — EC2: 0.25.0-gpu-py312-ec2 · SageMaker: 0.25.0-gpu-py312 · LLaVA-OneVision-2, Unlimited OCR, MOSS-Transcribe-Diarize, openai/privacy-filter, Hy3.
  • [2026/07/11] SGLang v0.5.15 (Ubuntu) — EC2: 0.5.15-gpu-py312-ec2 · SageMaker: 0.5.15-gpu-py312 · GLM 5.2 Tuned, Hy3, HRM-Text, LocateAnything-3B.
  • [2026/07/10] TensorFlow v2.21.0 (SageMaker training) — SageMaker CPU: 2.21.0-cpu-py312-amzn2023-sagemaker · SageMaker GPU: 2.21.0-gpu-py312-cu129-amzn2023-sagemaker · Amazon Linux 2023 with Python 3.12; GPU images ship CUDA 12.9.1.

📢 Support Updates

  • [2026/04/28] We cannot guarantee security patching on Ubuntu-based vLLM and SGLang images due to the lack of Ubuntu Pro licensing. Customers may continue using these images at their own discretion and risk. We recommend migrating to our Amazon Linux-based images.
  • [2026/02/10] Extended support for PyTorch 2.6 Inference containers until June 30, 2026
    • PyTorch 2.6 Inference images will continue to receive security patches and updates through end of June 2026
    • For complete framework support timelines, see our Support Policy

📝 Blog Posts

🎓 Workshop


License

This project is licensed under the Apache-2.0 License.

About

One stop shop for running AI/ML on AWS.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1.2k stars

Watchers

50 watching

Forks

Used by

Contributors

Languages