vLLM’s TT plugin brings Tenstorrent accelerators behind the standard serving API
The out-of-tree backend preserves vLLM’s OpenAI-compatible surface while adapting scheduling, parallelism and sampling to Tenstorrent mesh hardware.
vLLM has introduced an out-of-tree platform plugin for Tenstorrent accelerators, adding another hardware backend without putting Tenstorrent-specific code into vLLM core. The plugin keeps vLLM’s existing OpenAI-compatible API and request format, while moving hardware-specific scheduling, model registration and execution into the separately installed TT package.
What changed
The plugin automatically registers Tenstorrent hardware when the ttnn runtime from TT-Metal is available. It currently maps a range of text and multimodal model families—including Llama, Qwen, Mistral, Gemma, DeepSeek V3 and GPT-OSS—to Tenstorrent implementations shipped with TT-Metal. Administrators can also register model bundles from an external directory without editing the plugin source.
That compatibility layer does not pretend Tenstorrent devices behave like GPUs. Tenstorrent systems compile and trace programs for a mesh of cores and chips, so the plugin replaces conventional tensor- and pipeline-parallel ranks with a mesh configuration selected through MESH_DEVICE. The published implementation rejects standard -tp and -pp settings rather than silently accepting options it cannot honor in the usual way.
The scheduler also separates prefill-only and decode-only steps. vLLM’s upstream scheduler can mix prefill and decode work within a token budget, but the Tenstorrent path favors homogeneous, shape-stable batches that can replay a captured trace. Long prompts can still use chunked prefill, with decode steps interleaved so active requests continue making progress.
Who it affects
The immediate audience is teams evaluating Tenstorrent hardware for vLLM-based serving. Existing clients can keep the familiar API surface, but operators need to account for different topology and scheduling behavior. On 32-chip Galaxy systems, some models use one compiled program spanning the mesh and expose multiple data-parallel KV-cache lanes inside a single engine process rather than assigning a process to each rank.
The plugin can sample tokens on-device when a request fits that path. Requests needing log probabilities, penalties, masks or custom logits processing fall back to vLLM’s host-side sampler automatically. An asynchronous decode path overlaps host readback with later scheduling, but the project describes it as a steady-state optimization rather than a universal asynchronous execution model.
What to do
Prospective users should start with the plugin’s supported-model table and a validated vLLM release, then choose a mesh configuration for the target model instead of carrying over GPU rank settings. Workloads that rely on mixed prefill/decode behavior, speculative decoding or custom sampling should be tested explicitly: the current plugin does not support speculative decoding, and some sampling features leave the device fast path.
The broader engineering signal is the extension boundary. Tenstorrent’s backend changes core execution assumptions—batch shape, parallel topology and where sampling runs—yet the integration remains outside vLLM core. That makes the plugin a useful test of whether vLLM’s hardware interfaces can accommodate architectures that are not merely GPU-compatible accelerators.
comments · 0