cloud-native · since 2025
llm-d
llm-d community
A Kubernetes-native distributed inference stack for large language models, combining scheduling, routing, cache-aware serving and accelerator-oriented deployment components.
AIdistributed-inferenceKubernetesvLLM
desk notes
Verified 2026-09-28 against the project site and GitHub API. Latest stable umbrella release: v0.9.0 (published 2026-08-17). Repository remains active.