vLLM uses Mooncake to separate Kimi K3 serving from DSpark training
A new five-billion-parameter draft model and hidden-state connector turn speculative decoding into a capacity-planning choice across separate inference and training nodes.
The vLLM Project has extended its open-source Speculators training library to Kimi K3, a 2.8-trillion-parameter model, and released a five-billion-parameter DSpark draft model under the RedHatAI namespace. In the project’s September 15 engineering post, the team reports that the speculator raises single-stream generation on math workloads from about 110 to 435 tokens per second per user and delivers as much as 3.5 times more output throughput at matched interactivity under concurrent load.
What changed
DSpark uses a parallel draft-model backbone, then adds a Markov logit-bias head to restore local token dependencies and a confidence head to decide how much of a proposed block should be verified. The released Kimi K3 speculator proposes eight tokens per decoding step and reached a macro-average acceptance length of 4.11 tokens across nine evaluation domains, according to the vLLM results.
The more consequential engineering change is in training. Kimi K3 is too large for the project’s earlier single-node arrangement, even with four-bit weights. The team built a MooncakeHiddenStatesConnector that separates target-model serving from draft-model training and streams hidden states between vLLM and Speculators processes over RDMA or TCP.
The source supports a three-node working set: two GB300 nodes, each with four GPUs, serve the quantized Kimi K3 model while one four-GPU node trains the speculator. The team says that arrangement produced the best throughput among the configurations it tested. Mooncake lets operators scale the serving and training sides independently instead of forcing both into one fixed allocation.
Why platform teams should care
The work turns a research technique into a deployable artifact rather than only publishing benchmark numbers. vLLM provides a container recipe that loads RedHatAI/Kimi-K3-speculator.dspark directly through its speculative-decoding configuration. The same Speculators path has also been validated with draft models for Qwen3.6, Gemma 4 and GLM 5.2, the project says.
For teams operating large-model inference, the practical decision is whether the extra draft model and its training capacity are justified by workload shape. The gains are strongest in the project’s math and long-context tests, while acceptance length varies by domain. Operators should reproduce those measurements with their own prompts, concurrency and latency targets before planning capacity around the headline throughput increase.
The reusable pattern is the separation itself: vLLM serves the target model, Speculators trains the drafter, and Mooncake moves the required hidden states between them. That gives platform teams separate scaling controls for two workloads with different resource profiles. The published checkpoint offers a quick deployment path; the connector is the part that makes workload-specific retraining practical beyond one node.
comments · 0