OpenShift AI 3.5 AutoRAG turns RAG tuning into a measured search problem
Red Hat’s new architecture guidance separates indexing, retrieval and maintenance, then uses AutoRAG to compare pipeline configurations against accuracy, latency and cost constraints.
Red Hat is positioning AutoRAG in OpenShift AI 3.5 as an empirical alternative to the familiar cycle of enlarging chunks, retrieving more passages and moving to a bigger model whenever a retrieval-augmented generation system starts missing answers.
In new architecture guidance, the company argues that those changes can make a system both more expensive and less reliable: extra context consumes more KV-cache memory and increases time to first token, while loosely relevant passages can bury the facts a model needs. The proposed remedy is not one universal configuration, but automated evaluation against the organization’s own documents and questions.
Three pipelines instead of one
Red Hat’s reference design splits the workload into separate indexing, retrieval and maintenance pipelines. The indexing path parses and normalizes documents, creates chunks and embeddings, and populates a vector store. The real-time retrieval path combines dense and keyword search, reranking, access controls and generation. The maintenance path handles changes such as deletion requests, re-embedding and corpus-drift detection.
That separation matters operationally. Parsing choices can be tested without conflating them with model-serving behavior, while updates to the source corpus become a continuing process rather than an occasional rebuild. Red Hat maps the design to components including Docling for document conversion, vLLM and llm-d for inference, and MLflow for experiment and drift tracking.
AutoRAG searches the combinations
AutoRAG’s job is to test combinations of parsing, chunking, embeddings, hybrid search, reranking and generation settings. Teams supply representative documents and an evaluation set, select candidate models and define latency or cost limits. The system then presents configurations on a Pareto frontier rather than declaring that the most accurate result is automatically the right one.
The guidance emphasizes three measurements: context correctness, answer faithfulness and answer correctness. That distinction catches a subtle failure mode. A model can faithfully summarize the passages it receives even when retrieval supplied the wrong passages; measuring faithfulness alone would make that output look healthier than it is.
An earlier Red Hat Developer demonstration illustrates the tuning principle on a deliberately noisy sample corpus. Its selected setup kept the same measured context recall as the naive baseline while reducing average retrieved context from 1,367 words to 165. Red Hat is explicit about the limit of that result: the demo used deterministic retrieval scoring and did not make its one-billion-parameter model generally reliable. The full platform searches a broader configuration space and uses a fuller evaluation stack.
What platform teams should take from it
The practical decision is where to spend compute. Red Hat’s approach asks teams to remove irrelevant context before buying larger context windows or larger models, and to promote a configuration only after it is tested against ground-truth questions. Where retrieval is correct but generation remains weak, the guidance points to model customization such as retrieval-augmented fine-tuning rather than further widening the prompt.
The article is vendor guidance rather than an independent benchmark, and it does not publish comparative production costs for OpenShift AI 3.5. It does, however, make the platform’s intended operating model clearer: RAG configuration becomes a repeatable evaluation workflow with explicit accuracy, latency and budget boundaries, not a fixed set of tutorial defaults.
sources
- AutoRAG pipeline optimization in Red Hat OpenShift AIwww.redhat.com
- AutoRAG: Optimizing RAG for small modelsdevelopers.redhat.com
comments · 0