vLLM 0.30 adds faster restarts—and an upgrade checklist
The release introduces a persistent GPU weight cache, Gumbel-max watermarking and broad model support while removing several 0.29-era interfaces.
vLLM released version 0.30.0 on September 22 with a persistent GPU weight cache, built-in watermarking, new model support and a set of breaking changes that operators should review before upgrading.
What changed
The release’s new Fast Start path keeps post-quantized, tensor-parallel-sharded weights in GPU memory through a per-GPU daemon. Restarting engines can map those weights over CUDA IPC with --load-format ipc_cache instead of loading them again from disk. The release notes say the path supports FP4 checkpoints and multi-node tensor parallelism.
vLLM 0.30 also adds Gumbel-max generation watermarking and detection, including per-request controls and a dual-key mode compatible with speculative decoding. Serving changes include a stateless /v1/responses/render endpoint, reasoning-token accounting, additional Prometheus metrics and tighter validation around structured-output and scale-out requests.
Model and hardware work is extensive. The release adds support for DeepSeek-V4.1-Flash, DeepSeek-V4-Flash-Vision-Exp, GLM-5.3-Flash, K2-Horizon, Cohere Compass and Bailing V3 VL. It also expands optimization work across NVIDIA, AMD ROCm, Intel XPU and CPU backends.
Who it affects
The main operational impact falls on teams upgrading an existing serving deployment. Scale-out endpoints are no longer registered by plain vllm serve unless --enable-scale-out is passed, and the previous VLLM_ENABLE_SCALE_OUT_ENDPOINTS environment variable has been removed.
The release also removes GPTQ activation ordering through g_idx, removes environment variables deprecated for 0.29, changes YaRN handling to align with Transformers, and requires attention implementations to declare decode-context-parallel support explicitly. The default audio resampler moves from PyAV to torchaudio, while several hardware-specific kernel and communication defaults also change.
What to do
Before upgrading, operators should compare deployment manifests against the breaking-changes section of the 0.30.0 notes. In particular, check for the removed scale-out environment variable, deprecated 0.29 settings, GPTQ checkpoints that depend on activation ordering, YaRN-derived context lengths and custom attention backends used with decode context parallelism.
Teams evaluating Fast Start should treat it as a deployment change rather than a transparent speedup: it introduces a weight-cache daemon and a new load format. Watermarking is likewise opt-in functionality that should be tested with the serving stack’s speculative-decoding and request-control policies before production use.
sources
- vLLM v0.30.0 release notesgithub.com
comments · 0