Red Hat AI Inference on IBM Cloud adds the retrieval half of RAG
A new OpenAI-compatible embeddings endpoint lets teams generate vectors, discover compatible models and keep retrieval and chat behind one managed project boundary.
Red Hat AI Inference on IBM Cloud now supports an OpenAI-compatible embeddings endpoint, extending the managed service beyond model discovery and chat completion into the vector-generation step needed for semantic search and retrieval-augmented generation (RAG), IBM Cloud announced Oct. 9.
What changed
Applications can submit one or more strings to /v1/embeddings and receive a vector for each input. The response follows the familiar OpenAI embeddings shape, including indexed vectors, a model identifier and token-usage fields, according to IBM’s walkthrough.
The same project-scoped service already exposes /v1/models for model discovery and /v1/chat/completions for generation. Model metadata now includes a supports_embeddings field so an application can identify models suitable for retrieval before it submits text, the post says.
Authentication remains tied to IBM Cloud Identity and Access Management. Requests can use an IBM Cloud API key or bearer token, subject to the caller’s IAM roles, and the example endpoint is scoped to a regional Red Hat AI Inference project, IBM documents.
Who it affects
The immediate audience is application teams building semantic search, knowledge assistants, document discovery, recommendations, classification or agent-memory workflows on the IBM-hosted Red Hat service. Platform teams can now offer retrieval and generation through one project context and authentication model rather than introducing a separate embedding-serving endpoint, according to IBM.
This does not supply a vector database. The documented RAG flow still expects teams to chunk source documents, generate embeddings, store vectors with the original text and metadata, retrieve the closest passages for a user query, and then send those passages to a chat-completion model, the walkthrough explains.
What to do
Teams evaluating the endpoint should first list the models available to their project and select one advertising supports_embeddings. They then need a paid IBM Cloud account, a Red Hat AI Inference project, the appropriate IAM role and an API key or bearer token before calling /v1/embeddings, IBM says.
The useful first test is small and measurable: embed a representative document set, store the vectors in the team’s chosen vector database, and compare retrieval quality for real queries before connecting the results to /v1/chat/completions. The launch makes that architecture easier to assemble; it does not remove the need to choose chunking, storage and relevance-testing strategies.
sources
comments · 0