live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
guideAI

Red Hat turns AMD GPU inference benchmarking into a reproducible Podman workflow

The new guide isolates host, container and workload variables, but its local single-node design is a baseline—not a production result.

AMD GPU benchmarking workflow in rootless Podman, showing baseline comparisons.
Side by side: what changed
By The News Desk· Sep 29, 2026the quick take — two AI hosts go live when you do

Red Hat’s latest Emerging Technologies guide offers something more useful than another unlabeled tokens-per-second chart: a reproducible path for testing a vLLM endpoint on AMD hardware with rootless Podman.

The walkthrough runs the Red Hat AI Inference 3.5 ROCm image, GuideLLM and an optional Open WebUI on one host. It is aimed at developers and performance engineers who need to separate GPU access failures from inference-server behavior before comparing results.

The permission chain is the first benchmark

The guide starts below vLLM. It checks that the host exposes /dev/kfd and /dev/dri, that the amdgpu module is loaded, and that ROCm reports the installed GPU architecture. It then verifies that the user belongs to the video and render groups and that rootless Podman uses crun, which is required by the guide’s --group-add=keep-groups configuration.

Only after those checks does it run PyTorch inside Red Hat’s vLLM container to count visible GPUs. That ordering is the strongest part of the method: a failed model launch can otherwise blur host-driver, device-permission, OCI-runtime and framework problems into one opaque container error.

The inference server binds its OpenAI-compatible API to 127.0.0.1:8000, mounts a persistent Hugging Face cache and passes the AMD devices into the rootless container. The example uses Qwen2.5-7B-Instruct, one-way tensor parallelism and Red Hat’s pinned 3.5 vLLM and GuideLLM image builds.

Three workloads answer different questions

The first GuideLLM run uses a balanced synthetic workload with 1,000 input and 1,000 requested output tokens across concurrency levels from one to 650. That is the saturation test: it reveals where throughput stops scaling and latency or errors begin to climb.

A second run uses a four-turn conversation so repeated prefixes can exercise vLLM’s automatic prefix cache. A third replays a Mooncake conversation trace containing request timing, token lengths and shared-prefix information. Red Hat has therefore separated raw concurrency, cache-sensitive conversation behavior and trace-driven traffic instead of collapsing them into one headline score.

The workflow also records the model, image names, GPU count, tensor-parallel setting and run ID beside JSON, CSV and HTML output. That metadata is what makes two runs comparable.

What the guide does not prove

The article publishes a method, not AMD performance results. All clients and the server share one machine and communicate over loopback, so the test excludes network, ingress, scheduling and multi-node effects. It also disables SELinux labeling for the GPU container, and the optional Open WebUI example uses an unpinned main image. Those choices make the lab easier to reproduce but should not quietly become production defaults.

For a platform team, the practical next step is to preserve this workload definition while changing one variable at a time: model, GPU, tensor parallelism or vLLM configuration. Then repeat it inside the actual OpenShift topology. Red Hat’s contribution here is the baseline needed to make those later comparisons credible.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.