vLLM 0.29 makes Model Runner V2 the default and removes deprecated paths
The release changes the default execution path, removes ten deprecated model architectures and ships new admission-control and model-support features.
vLLM 0.29.0 changes the serving engine’s default execution path and removes several deprecated interfaces, making this more than a routine update for teams that package or operate vLLM-based inference services. The upstream release notes list 594 commits from 277 contributors.
What changed
Model Runner V2 is now the default for all models. The release adds CUDA-graph memory profiling for KV-cache sizing, batch-sharded sampling, prompt embeddings and additional speculative-decoding work to that execution path. Model Runner V1 remains in use for a limited set of ROCm models and features that V2 does not yet support.
The release also changes several defaults. FlashInfer all-reduce is enabled by default for tensor-parallel CUDA groups, while prefix-cache NONE_HASH behavior is deterministic by default. Operators gain --max-num-queued-reqs and --max-num-queued-tokens controls for admission management.
Model coverage expands to GraniteSWA and GraniteMoeSWA, alongside Hy4-preview, Qwen3.8-Flash-Next, NemotronH Omni Reasoning V3 and Kimi K3 NVFP4 checkpoints. The release includes additional work for Kimi K3, DeepSeek V4, speculative decoding, reinforcement-learning weight synchronization and Mamba prefix caching.
Who it affects
Platform teams should treat the update as a compatibility review rather than a drop-in patch. Ten deprecated model architectures have been removed. The PyAV video decoder backend is gone, and the python -m vllm.entrypoints.openai.api_server invocation is deprecated in favor of vllm serve. FlexOlmo, Olmo3 and Hunyuan V1/VL move to the Transformers modeling backend.
The default container and Python-package artifacts use CUDA 13.0. The project also publishes CUDA 12.9, ROCm, CPU and XPU artifacts, so image and accelerator choices need to remain explicit in deployment pipelines.
What to do
Before rollout, inventory workloads that rely on a removed architecture, PyAV decoding or the legacy Python module invocation. Test representative models against Model Runner V2, including memory sizing, speculative decoding and tensor-parallel behavior. CUDA operators that need the previous all-reduce path can opt out with VLLM_ALLREDUCE_USE_FLASHINFER=0.
Teams consuming vLLM through a downstream product should wait for that product’s validated build and support statement rather than substituting the upstream image directly. Teams operating upstream vLLM should pin the intended accelerator artifact and run compatibility and performance tests before promotion.
sources
- vLLM v0.29.0 release notesgithub.com
comments · 0