live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
newsFIELD BUILDS

Red Hat quickstart turns embedding selection into an OpenShift AI pipeline

The proof of concept benchmarks embedding models, promotes the winner and deploys a CPU-based semantic-search service through KServe.

Pipeline diagram from benchmark to deployment.
AI-generated diagram
By The News Desk· Sep 18, 2026the quick take — two AI hosts go live when you do

A new Red Hat AI Quickstarts reference build packages embedding-model evaluation, selection and deployment into one workflow for Red Hat OpenShift AI. The proof of concept uses ZenML to benchmark candidate models against technical-support queries, selects the strongest result, rebuilds the retrieval index and deploys a searchable service through KServe.

The project is aimed at a practical problem in retrieval-augmented generation: the quality of the context supplied to an application or agent depends heavily on the embedding model, but model selection is often detached from deployment. This build makes that choice a repeatable pipeline step rather than a one-time experiment.

What the pipeline does

The documented workflow prepares a TechQA retrieval benchmark, evaluates configured embedding candidates in parallel and records results in MLflow. It then re-encodes the corpus with the selected model, stores a versioned FAISS bundle in MinIO and creates or updates a CPU-based KServe InferenceService.

The deployed FastAPI application exposes a browser search interface, ranked search results, an embedding endpoint, health status and OpenAPI documentation. ZenML run metadata links operators to the deployed interface and supporting endpoints. A smoke profile is included for a smaller end-to-end test before running the full benchmark.

The example targets OpenShift Container Platform 4.20 or later and OpenShift AI 3.4 or later; its authors say it was validated with OpenShift AI 3.4.3 and ZenML 0.96.2. It requires the OpenShift AI KServe and MLflow components, the integrated image registry and persistent storage, but no GPU. The bootstrap needs cluster-administrator privileges or an equivalent permission set.

A pattern, not a production search platform

The repository is unusually explicit about its production-readiness limits. It describes the deployment as a proof of concept for isolated development and demonstrations, not production use. The sample loads a static exact FAISS index in the application process and does not provide continuous ingestion, metadata filtering or access control.

Its supporting services are also deliberately small-scale: ZenML, MySQL, MinIO, MLflow and model serving are not configured for high availability; external secret management, network policies, autoscaling, monitoring and backup procedures are outside the build.

Those caveats matter because the useful part of the project is the engineering pattern, not the sample topology. It shows how an OpenShift AI team can connect an evidence-based retrieval benchmark to a controlled deployment step, while keeping the evaluation run and its artifacts visible. Teams adopting it would still need to replace the demonstration storage, security and serving choices with production-grade services and controls.

The repository includes Helm charts, deployment and cleanup commands, validation targets, architecture diagrams and a small search UI. It is a concrete starting point for teams that want retrieval quality to be measured before an updated component reaches a RAG application or AI agent.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.