vLLM Semantic Router 0.4 turns one model endpoint into a managed model system
Hermes adds virtual model names, bounded multi-model recipes, open routing models and route-level evaluation for teams operating mixed inference fleets.
vLLM Semantic Router 0.4, codenamed Hermes, changes the project’s center of gravity from choosing a backend for each request to operating a governed system of models behind one application-facing name. The release was published Sept. 27 after 982 commits from 130 contributors.
What changed
Hermes introduces a virtual-model abstraction: an application calls one stable model name, while an isolated Mixture-of-Models recipe decides which qualified backend—or bounded combination of backends—handles the request. The project’s release article describes three broad paths: select one eligible model, escalate through a controlled cascade, or run experimental panel-and-judge and role-based workflows.
The release also moves more policy into the route. Inputs can be classified by modality, metadata and learned signals; decisions can use explicit tiers and priorities; and algorithms filter candidates by context length, protocol and declared task capability before scoring them. New route-local plugins add context compression and sampled shadow dispatch, while cache compatibility checks now account for history, tools, response format, model and recipe.
Hermes ships alongside two open model families. Vela 1.0 contains 15 models for tasks including domain and modality classification, prompt-attack detection, safety, PII recognition, hallucination checks, embeddings and reranking. Decision 1.0 contributes six open-weight models, from 0.6 billion to 9 billion parameters, that score choices supplied by an application. The project says native Decision integration is still a next step, so those models should not be mistaken for the policy engine already running in the router.
Who should care
Platform teams running several model servers now have a clearer boundary between application contracts and inference-fleet changes. A stable virtual name can keep model selection, privacy boundaries, accelerator placement and fallback policy out of application code. That is especially relevant to Kubernetes and OpenShift operators managing heterogeneous GPU, edge, private-cloud and hosted endpoints; the repository describes the router as supporting heterogeneous compute and data-location boundaries.
The release does not make multi-model orchestration free. Panel, judge and workflow routes spend more calls and require explicit quorum, timeout and fallback choices. Hermes therefore pairs those paths with Replay explanations, per-attempt timing and token records, shadow traffic, and sr-bench comparisons so operators can test whether a more elaborate recipe actually improves quality, cost or latency.
What to do
Teams evaluating 0.4 should begin with one virtual model and a single-model recipe, then inspect route decisions before enabling cascades or collaboration. Pin the 0.4.0 package or release artifact, define failure behavior for unknown signals and undersized candidate pools, and use shadow dispatch plus a frozen benchmark set before changing a live route. The operational test is not whether the router can call more models; it is whether the added policy remains explainable when models, workloads and infrastructure change.
sources
- vLLM Semantic Router v0.4.0 releasegithub.com
- Hermes release article sourceraw.githubusercontent.com
- vLLM Semantic Router repositorygithub.com
comments · 0