Building a retrieval-augmented generation (RAG) pipeline is no longer a hurdle for many enterprise AI teams. The real issue is optimizing it once it's built. Between finding the right chunking strategy, tweaking embedding models, and aligning generation parameters, developers are often left to guess which combination of factors yields the highest model faithfulness.

AutoRAG reduces that guesswork by automating evaluation and hyperparameter tuning across RAG frameworks. Built on Kubeflow Pipelines and driven by the open source IBM ai4rag optimization engine, the upcoming advanced technical preview of AutoRAG in Red Hat OpenShift AI 3.5 bridges the gap between automated experimentation and scalable enterprise deployment.

The path to production

Finding the dream RAG pattern is one thing, but implementing it is a whole new beast. Historically, testing an optimized RAG pattern in a live environment meant manually rebuilding the pipeline architecture.

With AutoRAG, once you identify an optimal pattern based on your data and criteria, developers can chat with their RAG pipeline results instantly through an inline user interface (UI) chat. This shifts the timeline from "pipeline ready" to "first interactive test" from a lengthy deployment cycle down to a near-instant configuration load.

A user interface in Red Hat OpenShift AI displaying a completed AutoRAG experiment pipeline alongside an open inline chat window for testing the selected RAG pattern.

Figure 1: Image displaying the completed AutoRAG experiment pipeline and an open chat window with the selected pattern

When you select a winning pattern from AutoRAG’s leaderboard, the platform automatically generates production-ready artifacts for both halves of your RAG pipeline (the ingestion pipeline and the query endpoint), reducing the tedious work of manually extracting configurations from notebooks and rebuilding everything from scratch. 

AutoRAG compiles your optimized data ingestion workflow into a deployable Kubeflow Pipeline that embeds every detail of the winning pattern (document parsing, chunking strategy, embedding model configuration, and vector database indexing parameters), which you can deploy on OpenShift AI and re-run whenever new documents arrive. 

Simultaneously, it packages the retrieval and generation logic into a Responses application programming interface (API) configuration consumable via Open GenAI Framework (OGF) (formerly Llama Stack), generating a pattern.json with your complete setup, a pre-configured Python script for making API calls, and copy-paste-ready code snippets in Python, cURL, Go, and Node.js, all accessible through the dashboard's “View Code” tab. This bypasses traditional manual setup entirely, giving you a direct path to deployment, helping turn a complex RAG workflow into a scalable, developer-friendly endpoint.

Integrated multilingual support

The English language dominates global business, but local languages often still rule internal corporate data. For global enterprises, building a RAG pipeline for non-English speaking teams typically requires adding complex translation layers, which adds processing latency and severe risk of context loss during tokenization.

In OpenShift AI 3.5, AutoRAG includes native multilingual support. This means global enterprises can ingest documents in their native languages, eliminating context loss from translation and supporting higher localized accuracy.

To optimize compute efficiency across large search spaces, AutoRAG also includes a model preselector stage that automatically detects document languages using large language model (LLM)-driven detection (benchmarks embedding models and LLMs against your test data), removing poor performers before the full optimization matrix runs. 

This approach narrows the search space to models that demonstrate strong performance on your language, reducing unwanted computation on incompatible model combinations. Initial support for this feature includes German, Spanish, and Japanese, with future releases incorporating the majority of enterprise languages across Latin, Chinese, Japanese, and Korean (CJK), and Cyrillic scripts.

Democratizing the developer experience

Optimizing an AI pipeline shouldn't feel like navigating in the dark. Historically, understanding why AutoRAG selected one RAG pattern over another meant mentally reconstructing the pattern’s architecture from tabular leaderboard data, which chunking strategy fed into which embedding model, how retrieval methods stacked against re-ranking approaches, and where configurations failed.

AutoRAG in OpenShift AI 3.5 changes that with a visual pipeline representation that shows the end-to-end architecture of each tested RAG pattern at a glance.

Instead of drilling through complex configuration files and leaderboard tables, developers can see how data flows from ingestion through chunking, embedding, retrieval, and generation, with node-level details showing what happened at each stage.

This visual explainability makes it dramatically easier to compare patterns side-by-side, understand why certain patterns outperformed others, and debug failed experiments—turning AutoRAG from an optimization engine into a learning tool that builds your team's intuition about which RAG architectures work and why.

Combined with the interactive playground that lets you chat with pipeline results before committing to a full deployment, these quality-of-life updates give developers more visibility and control over their RAG workflows without ever leaving the OpenShift AI environment.

A visual pipeline graph in Red Hat OpenShift AI showing the end-to-end node architecture and data flow for a completed AutoRAG experiment.

Figure 2: Image of completed experiment pipeline after running AutoRAG

Meeting your infrastructure where it lives

Enterprise AI is all about optimization, and that includes infrastructure flexibility. Many large organizations have already standardized on PostgreSQL with pgvector for vector storage. They've already built operational runbooks, passed security reviews, trained database administrator (DBA) teams, and integrated it into their compliance controls. Forcing these teams to deploy and maintain a separate vector database to use AutoRAG creates adoption friction, increases total cost of ownership, and duplicates data management overhead.

OpenShift AI 3.5 lowers this barrier by adding pgvector support alongside the existing Milvus option, letting AutoRAG fit cleanly into your existing data architecture. If your infrastructure team already operates pgvector in production, you can now run AutoRAG optimization experiments against that same instance without provisioning new infrastructure or migrating data.

In the future, we plan to integrate Elasticsearch and Qdrant in later releases of OpenShift AI. This infrastructure flexibility means fewer roadblocks between evaluation and production deployment.

Bringing it all together on OpenShift AI

The features in AutoRAG 3.5 are built to reduce enterprise-scale friction. When combined with the scalable hybrid cloud infrastructure of Red Hat OpenShift AI, data science teams can move from a prototype RAG pipeline to a globally localizable, optimized, production-ready AI application in a fraction of the time.

For engineering teams, AutoRAG 3.5 reduces the manual overhead of database integration, evaluation file creation, and deployment scaffolding, allowing them to focus on what matters most—building intelligent applications that deliver real business value.

Want to get started?

Red Hat OpenShift AI 3.5 is already available, and AutoRAG 3.5 is currently available as a Technical Preview.

Product trial

Red Hat OpenShift AI on Developer Sandbox | 30-day self-serve trial of a Developer Sandbox for Red Hat OpenShift AI

Instant access to your own minimal developer cluster, hosted and managed by Red Hat for use with Red Hat OpenShift AI.

About the authors

Suhas Kashyap is a Product Manager on the Red Hat OpenShift AI team, where he focuses on AI/ML platform capabilities including model customization, RAG, and developer tooling. He brings over 22 years of software industry experience spanning development, architecture, and DevOps.
Before joining Red Hat, Suhas spent 9.5 years at IBM in AI Product Management, where he shipped the AI Toolkit for IBM Z and LinuxONE and worked extensively on model customization and advanced RAG capabilities within watsonx.ai.

Outside of work, Suhas is an avid cricketer, half-marathon runner, amateur photographer, and self-described lawn care enthusiast.

Isaac Tigges is a Developer Advocacy intern on Red Hat's AI business unit and a self-taught open source developer. He likes finding problems, common or niche, and solving them in the open, and he studied Business Management at Virginia Tech.

Anja is a Product Management Intern at Red Hat, focusing on AI infrastructure, system insights, and developer experience.

UI_Icon-Red_Hat-Close-A-Black-RGB

Browse by channel

automation icon

Automation

The latest on IT automation for tech, teams, and environments

AI icon

Artificial intelligence

Updates on the platforms that free customers to run AI workloads anywhere

open hybrid cloud icon

Open hybrid cloud

Explore how we build a more flexible future with hybrid cloud

security icon

Security

The latest on how we reduce risks across environments and technologies

edge icon

Edge computing

Updates on the platforms that simplify operations at the edge

Infrastructure icon

Infrastructure

The latest on the world’s leading enterprise Linux platform

application development icon

Applications

Inside our solutions to the toughest application challenges

Virtualization icon

Virtualization

The future of enterprise virtualization for your workloads on-premise or across clouds