live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

vLLM Semantic Router 0.4 adds stable virtual models and bounded multi-model routing

The Hermes release gives applications one model name while routing among qualified backends, and adds new evaluation, safety and operational controls.

Virtual model routing among backend models with safety and evaluation checks.
AI-generated illustration
By The News Desk· Sep 29, 2026the quick take — two AI hosts go live when you do

vLLM Semantic Router has released version 0.4, code-named Hermes, expanding the project from request classification into a more complete control plane for choosing, coordinating and evaluating model backends.

The central change is a virtual model: applications can address one stable model name while the router maps it to an isolated mixture-of-models recipe. That recipe can select one qualified backend, escalate through a bounded cascade, or coordinate multiple models. The release also filters candidates by context length, protocol and declared task capabilities before an algorithm runs, reducing the chance that a route silently chooses an incompatible endpoint.

Routing becomes a measurable system

Hermes adds experimental Fusion and Workflow algorithms. Fusion runs a panel-and-judge pattern with a usable-response quorum and configurable fallback. Workflows assign bounded roles and can resume tool-driven work. A separate experimental prompt selector can ask a helper model to choose among eligible candidates, while falling back if that selection call fails.

The release puts more evidence around those decisions. Preview and Replay now expose the strategy, tier and reason a route won. The sr-bench framework compares a routed experience against individual models on the same tasks, including quality, cost and latency. Shadow Dispatch can send a sampled copy of a request to a candidate without changing the live answer, providing a way to evaluate alternatives before promoting them.

New models and safer signals

The project also launched two open model families. Vela 1.0 contains 15 models for tasks including domain and modality classification, safety signals, PII detection, hallucination marking, embeddings and reranking. Decision 1.0 contains six open-weight models, ranging from 0.6 billion to 9 billion parameters, intended to score application-provided choices, conditions and ordered rubrics. Native Decision integration is described as a next step rather than part of the current router release.

For operators, Hermes adds bounded scanning of long inputs for guard and PII signals, response-stage jailbreak and hallucination observations, and clearer handling of classifier failures as unknown rather than safe. New route-local plugins include context compression and shadow dispatch, while response caching now checks compatibility across request history, tools, output format, model and recipe.

The release also adds Podman to the local installation choices, management-path permission checks, sensitive-field redaction, dashboard CSRF and origin checks, and a support matrix distinguishing maintained, supported and experimental surfaces.

Version 0.4.0 was published on Sept. 27. The project says it comprises 982 commits from 130 contributors since version 0.3.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.