live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
analysisAI

vLLM brings day-zero serving support to NVIDIA Vera Rubin

The upstream project has daily CUDA 13.4 images, Rubin-tuned kernels and preliminary results showing as much as 7.84× the per-GPU AgentX throughput of GB200.

Chart of Rubin throughput gains versus GB200 and GB300.
Chart: figures from the story
By The News Desk· Oct 10, 2026the quick take — two AI hosts go live when you do

vLLM says its serving stack now runs on NVIDIA’s Vera Rubin NVL72 platform, with daily container builds, initial Rubin-tuned kernels and support for models including DeepSeek, Kimi, GLM and MiniMax. The work is an early hardware-enablement milestone rather than a finished performance release, but it gives inference teams a concrete path to start testing the next NVIDIA rack architecture.

What changed

The project’s first Rubin build is available as the vllm/vllm-openai:cu134-nightly image, using CUDA 13.4 and PyTorch 2.15. vLLM says Rubin remains compatible with many kernels built for NVIDIA’s Blackwell architecture family, while FlashInfer 0.7.0 adds Rubin-specific attention, GEMM and mixture-of-experts kernels.

A more structural optimization uses CUDA 13.4 locality domains to place shards of mixture-of-experts weights near the streaming multiprocessors that consume them. In vLLM’s preliminary MiniMax M3 layer tests, that placement produced about a 1.2× speedup for small-token forward passes. The project says the feature remains under active design and development.

The collaboration also gives Red Hat a direct role in the upstream enablement. vLLM credits NVIDIA and Red Hat with setting up the daily Rubin Docker builds, alongside contributions from Inferact, NVIDIA and the broader project community.

What the numbers mean

The headline result comes from SemiAnalysis AgentX: vLLM serving MiniMax M3 on Vera Rubin NVL72 reached up to 7.84× the throughput per GPU of NVIDIA GB200 at matched interactivity. Under a 150-tokens-per-second constraint, the reported advantage was 5.18×.

A separate MLPerf Inference 6.1 result used vLLM behind NVIDIA Dynamo to serve Qwen3-VL-235B-A22B. vLLM reports up to 3.7× the throughput of GB300 NVL72 across offline, server and interactive vision-language scenarios.

Those figures should be read as early, workload-specific results, not a blanket expectation for every model. vLLM explicitly calls the work an early look and lists substantial follow-on tasks, including fuller locality-domain support, additional FlashInfer integration and more Rubin-specific attention and MoE kernels.

Who should act

Teams planning Rubin evaluations can begin by reproducing their own model and latency targets with the CUDA 13.4 nightly image. The useful signal is not the largest multiplier by itself; it is that the upstream serving path, container pipeline and representative model coverage already exist before broader deployment.

Production teams should keep the nightly label in view. The sensible next step is controlled benchmarking against an existing GB200 or GB300 baseline, including interactivity constraints and the exact parallelism strategy, rather than treating the published peak as a capacity-planning constant.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.