live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
guideAI

A CPU-only path to benchmark Red Hat AI Inference before buying GPU capacity

Red Hat’s Emerging Technologies team packages vLLM serving, Podman and a repeatable cpueval workflow into an x86 evaluation pattern.

CPU inference benchmark with core-count comparisons and throughput results.
Chart: figures from the story
By The News Desk· Oct 7, 2026the quick take — two AI hosts go live when you do

Red Hat’s Emerging Technologies team has published a detailed path for running a language model on CPUs with Red Hat AI Inference 3.5, exposing it through an OpenAI-compatible API and measuring its behavior with a repeatable load test. The guide is useful less as a claim that CPUs replace accelerators than as a way to establish a baseline before committing scarce GPU capacity.

What the pattern assembles

The deployment starts with the Red Hat AI Inference CPU image and Podman on a RHEL host. Red Hat’s example serves Qwen2.5-7B-Instruct through vLLM, persists downloaded model weights in a host cache and exposes health, model-listing and chat-completion endpoints on port 8000.

The tested environment is specific: RHEL 10.2 on an AWS c8i.12xlarge instance with 48 virtual CPUs and 96 GB of memory. The guide calls for at least 16 physical cores and 32 GB of RAM, with 32 cores and 64 GB or more recommended for concurrent tests of 8-billion-parameter models. It also says the reference hardware uses AVX-512 or AMX instructions, so teams should not treat the procedure as a promise of comparable performance on arbitrary x86 hosts.

One operational detail matters immediately: root and non-root Podman sessions use separate credentials and storage. Red Hat warns users to keep the same privilege context across registry login, image pull and container launch rather than mixing sudo podman with rootless commands.

How the benchmark works

For repeatable testing, the guide introduces cpueval, an open-source wrapper around Ansible playbooks that provisions the container, runs the test and collects results. GuideLLM supplies the load, sweeping concurrency and reporting time to first token, time per output token, request rate and token throughput.

Placement is part of the experiment. On multi-NUMA systems, Red Hat recommends separating vLLM and the load generator across NUMA nodes. On a single-NUMA cloud VM, the example instead partitions cores between the server and GuideLLM. That distinction is important: without isolation, the benchmark client can compete with the model server and distort the baseline.

Red Hat reports roughly 9 tokens per second with eight cores for an 8B quantized model, about 15 tokens per second at 16 cores and roughly 283 tokens per second at 32 cores with concurrency set to 32. Those figures describe the published test conditions, not a general CPU sizing rule.

Who should try it

The pattern fits platform teams that need to verify model compatibility, tool calling and API behavior before reserving accelerators, as well as edge teams evaluating small models where CPU capacity is already available. It can also give procurement and capacity-planning discussions a locally reproducible starting point instead of a vendor benchmark detached from the intended workload.

There are clear boundaries. Registry credentials are required for the Red Hat image, gated models can require a Hugging Face token, and the post appears on Red Hat’s pre-product Emerging Technologies site. The site explicitly cautions that its how-tos are not supported products or roadmap commitments unless stated otherwise. Teams should therefore use the workflow as an evaluation harness, then confirm product support and production architecture separately.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.