vLLM turns DeepSeek-V4.1-Flash’s day-zero path into a 5.3× throughput gain
SWA bounded replay, CUDA graphs and model-specific kernel fusion reshape the latency-throughput trade-off for long-horizon agent serving.
Three weeks after DeepSeek-V4.1-Flash arrived, the vLLM community says it has pushed the model to 1.9× its day-zero speed at low concurrency and 5.3× its throughput under a 150-tokens-per-second constraint. The result matters less as a single benchmark number than as a map of where agentic inference is spending time: sliding-window cache state, kernel launches and memory traffic rather than one dominant matrix multiplication.
What changed
The central change is sliding-window-attention bounded replay. DeepSeek-V4.1-Flash keeps a compact global KV cache alongside a larger, uncompressed sliding-window cache for the last 128 positions in each of 40 layers. Instead of storing the sliding-window state at every prefix-cache boundary, vLLM now rebuilds the last 128 tokens after a cache hit. On the decoder side, it runs the upper layers only on those final positions rather than across the whole prompt.
That replay is not bit-exact, but vLLM reports no meaningful accuracy difference on GSM8K and GPQA within about 1.5 standard errors. It is enabled by default for the model and can be disabled with --no-swa-bounded-replay.
The runtime also captures the trimmed upper layers in separate CUDA graphs. Without that step, launch overhead can erase the gain on short prompts; with it, vLLM measured a 30%–40% reduction in prefill computation time.
The rest is in the kernels
vLLM integrated DeepSeek’s MegaAttention, Mega-mHC, Mega-Gate and DeepSelect work, then fused more of the remaining path. The changes include scoring only fixed sparse-attention candidates, overlapping mHC coefficient work on a side CUDA stream, and keeping intermediate values on-chip in a fused low-latency kernel.
For the KV cache, MegaAttention reads an NVFP4 format that vLLM says is 45% smaller than its previous FP8 cache. At very long contexts, sparse MQA kernels are reported as 14×–23× faster per layer at 512K tokens than scoring the whole context and masking unused positions.
Who should care
Teams evaluating DeepSeek-V4.1-Flash for coding agents or other long-running tool loops should re-benchmark rather than carrying forward day-zero capacity estimates. vLLM’s low-latency setup used four-way tensor parallelism with FlashInfer attention; its throughput setup instead used data-parallel attention with experts split across GPUs to avoid duplicating the model’s shared KV latent.
Those are architecture choices, not drop-in promises for every cluster. The published figures come from the SemiAnalysis AgentX workload and selected NVIDIA Blackwell configurations. Operators should reproduce the latency-throughput curve with their own prompt lengths, concurrency and service-level target.
The practical takeaway is that DeepSeek-V4.1-Flash no longer has one obvious deployment shape in vLLM. Low-concurrency latency and high-concurrency throughput favor different attention and parallelism choices, while bounded replay changes both prefix-cache storage and prefill cost. Capacity planning should test both ends of that curve.
comments · 0