Red Hat’s MLPerf 6.1 results make topology—not Kubernetes—the tuning problem
OpenShift led the four-GPU GB200 field, while CPU-only vLLM results exposed memory bandwidth and scheduler stability as the practical limits.
Red Hat’s MLPerf Inference 6.1 submissions put OpenShift at the top of the four-GPU NVIDIA GB200 field and show the same vLLM engine producing valid results on CPU-only Intel systems. The more useful finding is what made the difference: NUMA placement, memory bandwidth and scheduler isolation rather than a new model-serving abstraction.
MLCommons says Inference 6.1 drew a record 30 participating organizations and added end-to-end RAG and edge-agentic tests. Red Hat’s submission concentrated on existing datacenter workloads across gpt-oss-120b, Qwen3-VL, Llama 3.1 and Whisper.
OpenShift reaches the front of the GB200 group
A four-GPU Supermicro GB200 NVL4 system running OpenShift 4.21 and vLLM 0.24 produced 61,205.6 tokens per second in the offline gpt-oss-120b test and 51,095.2 tokens per second in the server scenario. Red Hat reports both as the highest results among four-GPU GB200 submissions, with its offline result 20% above the only other system at that scale.
Normalized per GPU, Red Hat calculates roughly 15,300 tokens per second—higher than the larger GB200 NVL72 submissions in this round. That supports a narrower conclusion than “Kubernetes has no overhead”: a tuned OpenShift system can lead this specific peer-reviewed configuration set.
The tuning detail is more transferable. On the Qwen3-VL workload, unpinned vLLM workers placed most memory on the remote socket and reduced host-to-device bandwidth from about 166 GB/s to about 50 GB/s. Pinning workers per socket restored bandwidth and improved end-to-end throughput by roughly 9%.
CPUs remain viable for bounded workloads
Red Hat and Intel also ran vLLM on a two-socket Xeon 6972P system with no accelerators. It delivered 1,507.21 tokens per second offline for Llama 3.1 8B and 1,966.11 samples per second for Whisper Large v3. Red Hat reports the highest per-core throughput among two-socket CPU-only submissions for both workloads.
The team reserved one CPU core per rank for serving and benchmark overhead, limited deep idle states and found memory bandwidth—not core count—to be the first-order limit. Those choices point to CPU inference for moderate-throughput, latency-tolerant work, not as a general substitute for accelerators.
What teams should test
Platform engineers should reproduce topology and scheduler settings before comparing hardware prices. A useful evaluation should record NUMA placement, memory bandwidth, reserved cores, idle-state policy and tail-latency behavior—not just aggregate token throughput.
The broader result is software portability: vLLM appeared in 22 submitters’ stacks, and Red Hat used it across both GB200 GPUs and Xeon CPUs. The benchmark does not prove every workload will move unchanged between them, but it does show that inference-engine investment can span substantially different hardware classes.
sources
comments · 0