vLLM 0.31 turns fast restarts and request trust into upgrade decisions
The release adds GPU-resident weight preloading and broader distributed-serving controls while changing multimodal request handling and several operator-facing flags.
vLLM 0.31.0 adds a GPU-resident restart path, expands large-scale serving controls, and changes several request and configuration defaults. The combination makes this more than a routine package refresh for platform teams operating shared inference services. The upstream release was published October 5, 2026.
What changed
The new vllm preload command starts a weight-cache daemon that keeps post-quantized model weights in GPU memory across engine restarts. The release extends that path to data parallelism and multi-token-prediction draft models, and adds health and readiness checks. An experimental snapshot workflow also uses CRIU to restore a fully initialized tensor-parallel-one engine.
Distributed-serving changes include a MoonEP all-to-all backend, prefill context parallelism with data parallelism, DeepEPv2 sequence parallelism, and back-pressure detection for KV-cache offloading. The scheduler also gains --max-num-active-seqs, separating active admission from the existing maximum sequence setting.
The security boundary for multimodal requests is stricter. Per-request mm_processor_kwargs and media_io_kwargs are rejected unless the server starts with --trust-request-mm-kwargs. The release also changes prefix-cache hashing to distinguish extra-key sources and include the LoRA path.
Who it affects
Operators that restart engines during model rollouts can evaluate preloading as a way to avoid reloading post-quantized weights. Teams running disaggregated or expert-parallel serving have new controls, but should benchmark them against their current topology rather than assume an automatic gain.
API owners need to review any client that sends per-request multimodal processor or media I/O options. Those requests now fail by default unless the deployment deliberately opts into trusting them.
What to do
Test 0.31.0 in staging with production request shapes and the same accelerator topology used in service. Audit clients for the newly gated multimodal fields before rollout.
The upgrade also requires a configuration review: tokenizer_mode="slow" is removed; --enable-mamba-fine-grained-prefix-cache is renamed to --enable-mamba-shared-prefix-checkpoint; XPU graphs are enabled by default; and several quantization and backend options changed or disappeared. Pin the release, compare startup and steady-state behavior, and update manifests before promoting it into a shared serving environment.
sources
- vLLM v0.31.0 release notesgithub.com
comments · 0