live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
newsAI

Red Hat packages NVIDIA AI-Q as an OpenShift AI research-agent quickstart

The deployment guide combines multi-agent routing, vLLM model serving and an optional observability stack, with local-GPU and NVIDIA-hosted paths.

Local OpenShift AI deployment versus NVIDIA-hosted inference
Side by side: what changed
By The News Desk· Sep 11, 2026the quick take — two AI hosts go live when you do

Red Hat has published a hands-on quickstart for deploying a multi-agent research application on Red Hat AI Factory with NVIDIA. The guide adapts NVIDIA’s AI-Q Blueprint for Red Hat AI environments and adds deployment choices for locally served models or NVIDIA-hosted inference.

What the quickstart assembles

The application routes requests between simple responses, shallow tool-assisted research and deeper investigations that use planning, subagents and report generation. It can draw from web and academic search, uploaded files and an optional retrieval-augmented generation layer, while preserving citations in the output, according to the deployment documentation.

The underlying pattern combines NVIDIA NeMo Agent Toolkit with LangChain Deep Agents. For a local deployment, the guide serves three model roles through vLLM and KServe: RedHatAI/gpt-oss-120b as orchestrator, RedHatAI/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 for intent routing and research, and nvidia/Nemotron-Mini-4B-Instruct for summarization. Red Hat says the documented configuration was tested with OpenShift 4.20 and OpenShift AI 3.3.2.

Red Hat’s accompanying demonstration shows the workflow selecting a research path, collecting evidence and producing a report. It also shows operators tracking execution with OpenTelemetry and MLflow, watching model-serving metrics in Grafana and running evaluations from OpenShift AI workbenches.

Two deployment paths, with different costs

The local-vLLM option keeps model serving and data inside the cluster, but its standard configuration calls for three 80GB NVIDIA H100 or A100 GPUs. The guide describes a two-H100 minimum using Multi-Instance GPU partitioning. Teams can instead use NVIDIA NGC-hosted models without local GPUs, trading infrastructure control for a cloud API and pay-per-use inference.

The optional observability deployment installs OpenShift Logging with LokiStack, Grafana, OpenTelemetry Collector, user-workload monitoring and MLflow. The default application deployment uses a 10GB PostgreSQL persistent volume, while ChromaDB and application data use ephemeral storage unless an administrator adds another persistent volume.

The operational caveat

This is an engineering quickstart, not a blanket support statement. Red Hat explicitly says the material is authored by its experts but has not been tested on every supported configuration. The repository also uses prebuilt images based on NVIDIA AI-Q 2.1.0 with Red Hat-specific patches, and the documentation exposes those deployment files and patches for inspection.

For platform teams, the useful development is the amount of the agent stack that is made concrete: model endpoints, routing, retrieval, traces, metrics, evaluation and cleanup are all represented in deployable manifests. The caveat is equally concrete: the local path has a substantial GPU floor, and the faster hosted path sends inference to NVIDIA’s service.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.