live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
analysisAI

DiffusionGemma turns one vLLM pass into a typed decision engine

A merged vLLM change lets DiffusionGemma fill fixed answer slots and expose confidence, while Red Hat maps a preview path onto OpenShift AI.

By The News Desk· Sep 28, 2026the quick take — two AI hosts go live when you do

A merged vLLM change gives open-model teams a new serving primitive: ask DiffusionGemma several bounded questions, fill their answer slots in one denoising step and return probabilities that application code can act on. Red Hat has now published a deployment walkthrough and product roadmap for testing that pattern on Red Hat AI and OpenShift AI.

What changed in vLLM

The merged vLLM pull request adds the machinery for a “structured read.” A client seeds DiffusionGemma’s fixed-length token canvas with a response template, leaves only the answer positions unset and limits the request to one denoising step. vLLM then returns log probabilities for each allowed token; the client can select the top answer and derive uncertainty from the distribution.

The approach supports yes-or-no, scored and multiple-choice questions, and several questions can share one request. Choices must currently map to single tokens so that the fixed canvas does not shift. A client can translate longer labels such as moderation_spam to a token such as B and convert it back after inference.

This is not yet a standard vLLM endpoint. The pull request includes a prototype /v1/systemone server and low-level controls for seeded canvases, read-only requests and step limits. Red Hat’s walkthrough likewise labels the example server and custom runtime route as unsupported preview work.

The early performance signal

The pull-request author tested one DiffusionGemma deployment on a single NVIDIA DGX Spark. With a 32-token canvas and three decisions per request, it measured 8.7 requests per second at one-way concurrency and 54 requests per second at 32-way concurrency—about 162 decisions per second. Those are author-reported prototype results, not an independent benchmark, and the PR’s own task tests include misses on some classifications.

The operational trade-off is memory. Red Hat notes that diffusion state buffers grow with batch size, canvas length and the model vocabulary. Small decision canvases can sustain much higher concurrency than ordinary generation, but platform teams still need to size against their chosen model variant and schema.

Where Red Hat AI fits

Red Hat says DiffusionGemma 26B-A4B is already a validated model for its existing multimodal uses and that optimized FP8-dynamic and NVFP4 checkpoints are available in the RedHatAI collection. The company says the NVFP4 variant uses roughly one-third of the memory of BF16.

For now, teams can prototype the structured-decision mode with an upstream vLLM nightly and an unsupported custom serving runtime on Red Hat OpenShift AI, or use a Red Hat AI Inference preview image. Red Hat’s stated sequence is a preview after structured-read support reaches a stable vLLM release, followed by hardening based on feedback. Its post names an example-server Developer Preview for OpenShift AI 3.6 GA if vLLM 0.31 or later lands, while leaving the hardened endpoint’s timing and support level undecided.

The practical appeal is not another chatbot. It is a self-hosted component for routing, moderation, risk scoring and agent branching where applications need fixed answers and explicit confidence. The merged code establishes the primitive; supportability, interface stability and workload-specific accuracy are the next gates.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.