Red Hat turns AMD GPU inference benchmarking into a reproducible Podman workflow
The new guide isolates host, container and workload variables, but its local single-node design is a baseline—not a production result.
Red Hat’s latest Emerging Technologies guide offers something more useful than another unlabeled tokens-per-second chart: a reproducible path for testing a vLLM endpoint on AMD hardware with rootless Podman.
The walkthrough runs the Red Hat AI Inference 3.5 ROCm image, GuideLLM and an optional Open WebUI on one host. It is aimed at developers and performance engineers who need to separate GPU access failures from inference-server behavior before comparing results.
The permission chain is the first benchmark
The guide starts below vLLM. It checks that the host exposes /dev/kfd and /dev/dri, that the amdgpu module is loaded, and that ROCm reports the installed GPU architecture. It then verifies that the user belongs to the video and render groups and that rootless Podman uses crun, which is required by the guide’s --group-add=keep-groups configuration.
Only after those checks does it run PyTorch inside Red Hat’s vLLM container to count visible GPUs. That ordering is the strongest part of the method: a failed model launch can otherwise blur host-driver, device-permission, OCI-runtime and framework problems into one opaque container error.
The inference server binds its OpenAI-compatible API to 127.0.0.1:8000, mounts a persistent Hugging Face cache and passes the AMD devices into the rootless container. The example uses Qwen2.5-7B-Instruct, one-way tensor parallelism and Red Hat’s pinned 3.5 vLLM and GuideLLM image builds.
Three workloads answer different questions
The first GuideLLM run uses a balanced synthetic workload with 1,000 input and 1,000 requested output tokens across concurrency levels from one to 650. That is the saturation test: it reveals where throughput stops scaling and latency or errors begin to climb.
A second run uses a four-turn conversation so repeated prefixes can exercise vLLM’s automatic prefix cache. A third replays a Mooncake conversation trace containing request timing, token lengths and shared-prefix information. Red Hat has therefore separated raw concurrency, cache-sensitive conversation behavior and trace-driven traffic instead of collapsing them into one headline score.
The workflow also records the model, image names, GPU count, tensor-parallel setting and run ID beside JSON, CSV and HTML output. That metadata is what makes two runs comparable.
What the guide does not prove
The article publishes a method, not AMD performance results. All clients and the server share one machine and communicate over loopback, so the test excludes network, ingress, scheduling and multi-node effects. It also disables SELinux labeling for the GPU container, and the optional Open WebUI example uses an unpinned main image. Those choices make the lab easier to reproduce but should not quietly become production defaults.
For a platform team, the practical next step is to preserve this workload definition while changing one variable at a time: model, GPU, tensor parallelism or vLLM configuration. Then repeat it inside the actual OpenShift topology. Red Hat’s contribution here is the baseline needed to make those later comparisons credible.
sources
comments · 0