live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

Open Data Hub 3.6 preview expands training, experiment tracking and llm-d serving

The early-access branch adds new modules and lifecycle changes that OpenShift AI operators should test as an integrated platform, not as isolated components.

Old stack vs new integrated AI platform modules.
Side by side: what changed
By The News Desk· Sep 4, 2026the quick take — two AI hosts go live when you do

Open Data Hub’s first 3.6 early-access build widens the platform in three places that usually span separate operational teams: distributed training, experiment tracking and large-model serving. The v3.6.0-ea.1 platform release bundles a Trainer module, MLflow Operator 1.1 and a larger llm-d component set alongside KServe, pipelines, model registry, Ray and TrustyAI.

This is an upstream early-access branch, not a Red Hat OpenShift AI general-availability announcement. Its value to OpenShift AI operators is as a preview of integration and upgrade boundaries worth testing before the branch matures.

What the new modules change

The Trainer release pulls upstream Kubeflow Trainer work into the Open Data Hub build and updates runtimes for Training Hub 0.9.2 and PyTorch 2.11.0. It also makes TrainJob immutable through admission handling to unblock Kueue unsuspension. At the platform layer, the older Training Operator v1 is deprecated and disabled while Trainer moves behind a module handler. That shifts validation from “does a training pod start?” to the full job lifecycle: admission, queue suspension, runtime selection, status progression and cleanup.

MLflow Operator 1.1 moves the managed service to MLflow 3.10.1 and adds namespace-specific artifact locations through MLflowConfig. Its release also includes custom CA bundle mounting, Prometheus metrics, status URLs, safer CORS defaults, revised RBAC, OpenShift-specific PostgreSQL and SeaweedFS manifests, and TLS support for database and S3-style test backends. Those are platform concerns, not merely a new tracking UI.

The llm-d side is a family rather than one router. The platform manifest now points to the llm-d router plus workload-variant autoscaling, a batch-gateway operator, asynchronous serving and a latency predictor. The router’s EA release is primarily a synchronization of upstream changes through late August, so operators should treat interoperability—not one headline router feature—as the test target.

What operators should validate

A useful 3.6 preview exercise should cover four seams:

  1. Upgrade and ownership: confirm the DataScienceCluster reaches Ready while out-of-tree modules install, upgrade and uninstall; specifically check the handoff from Training Operator v1 to Trainer and the platform controller’s module status.
  2. Disconnected deployment: mirror every related image by digest and verify Trainer, MLflow and the llm-d auxiliaries resolve without registry access. The platform changelog includes multiple fixes to image-digest synchronization and disconnected readiness.
  3. Identity, storage and certificates: test MLflow UI and SDK routes, namespace-scoped artifact paths, custom certificate authorities, RBAC boundaries and TLS to backing PostgreSQL and object storage.
  4. End-to-end AI flow: submit a queued training job, record artifacts and metrics in MLflow, register or serve the result, then exercise llm-d routing under mixed request sizes while watching autoscaling, latency and batch behavior.

The preview is most informative when those checks run together. The 3.6 branch is exposing where training, tracking and serving share certificates, storage, scheduling and lifecycle state—the integration work operators need to understand before any downstream release adopts the stack.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.