Red Hat AI 3.5 turns agent operations into a platform concern
The release makes model evaluation, workload priority, controlled rollouts and observability generally available while keeping several agent and cost controls in preview.
Red Hat has made Red Hat AI 3.5 available with a release centered less on adding another model endpoint and more on operating shared, agent-heavy AI infrastructure. The company says the release brings model and agent evaluation, workload priority, controlled model rollouts, distributed inference and observability into a more unified platform, while several newer agent-governance and cost-accounting functions remain technology previews.
What changed
EvalHub is now generally available for evaluating models and agents against risks including prompt injection and jailbreaks. Red Hat’s AI Model Catalog is also generally available with published safety, personally identifiable information exposure and toxicity results from Garak, alongside tool-calling validation and deployment configuration.
On shared infrastructure, priority-aware serving is generally available with admission control, fairness policies and starvation protection. Controlled deployments add canary traffic, graceful transitions for requests already in flight and rollback without service disruption. CPU and NVMe offload for vLLM’s key-value cache are generally available, as is agent tool calling in standard and distributed deployments.
The cloud footprint broadens too: llm-d distributed inference is generally available on CoreWeave Kubernetes Service and Microsoft Azure Kubernetes Service, while Amazon EKS support is a technology preview. Red Hat also says hosted control planes on OpenShift Virtualization are now supported for tenants that need stronger control-plane isolation.
Who it affects
The release is aimed at platform teams moving from isolated AI projects to shared services. The practical changes address recurring operational conflicts: latency-sensitive agents competing with batch workloads, model updates that need staged validation, long conversations exhausting accelerator memory, and application teams needing metrics without cluster-administrator access.
Developers get generally available Responses API and built-in retrieval-augmented generation through OGX, plus generally available agent templates and starter kits. AutoRAG and AutoML remain technology previews, as do gateway-level policy checks intended to block disallowed agent tool calls.
What to do
Teams evaluating 3.5 should separate generally available controls from previews in their rollout plans. The immediate production candidates are EvalHub, priority-aware serving, controlled deployments, KV-cache offload, Responses API support and the centralized observability framework.
MaaS token showback, visual agent tracing, AutoRAG, AutoML and gateway-level tool policy are still previews. Those functions may be useful in trials, but buyers should not treat them as production commitments. Platform owners should also verify the support status of their target Kubernetes service: CoreWeave and AKS are GA for distributed inference, whereas EKS is not yet at that level.
sources
comments · 0