Red Hat has published Sovereign AI Cloud as a Service, a joint reference architecture built with Rafay for telecommunications providers, sovereign cloud operators, and NeoClouds that want to turn distributed GPU infrastructure into a governed, self-service, revenue-generating AI cloud.

Built on Red Hat's foundational AI and multi-tenant cloud platform, the architecture integrates Rafay to add the self-service commercial workflows service providers need to resell that platform to their own customers. It is a validated blueprint, not a positioning document. It names the components, defines the capabilities at each layer, and traces a tenant request from SKU selection through to a metered GPU environment. You can read it in the Red Hat Architecture Center.

Addressing the GPU cloud service gap

Provisioning GPU compute as a cluster is not the same as delivering it as a cloud service.  Many organizations can stand up a GPU-enabled Kubernetes environment. The harder problem comes as AI adoption expands across teams, sites, and  infrastructure environments. Without a consistent operating model, each new deployment introduces its own provisioning process, policies, access model, and operational dependencies. What begins as a manageable infrastructure project becomes increasingly difficult to standardize and operate at scale.

Operators building AI clouds run into two distinct versions of this problem.

The first is a platform problem: environments have to come online repeatedly and identically. Each site needs servers provisioned, clusters built and joined to a fleet, network and storage carved into tenant-scoped allocations, policy applied consistently, and every layer kept patched and compliant over its lifetime. That is the problem Red Hat's platform solves, and it is the largest component of the work. 

The second is a commercial problem: Once the platform is running, the operator still has to sell what it produces. How does a customer request an environment without a ticket? What do we charge, and how do we prove what was consumed? Answering those questions requires SKU definition, rate cards, per-tenant entitlements, reservations, metering, and invoicing.

Both halves have to be automated. The effort to onboard a tenant is largely fixed regardless of what that tenant consumes, which is why manual onboarding breaks down, especially for the small and mid-size customers a self-service model depends on for volume. Without standardization across both platform operations and service delivery, onboarding time grows, GPU utilization falls, and operational overhead eats into service  margin.

Red Hat AI plus Red Hat OpenShift is the sovereign AI cloud engine

The reference architecture starts from the platform, because every property an operator ultimately sells - isolation, assurance, performance,  provenance - is created there.

A chart showing how Red Hat, Rafay, and NVIDIA support a sovereign AI cloud architecture.

OpenShift enforces the tenant boundary. Red Hat Advanced Cluster Management with Hosted Control Planes gives each tenant a dedicated Kubernetes API server, so control-plane separation is established before a workload ever schedules. Where containers alone are not sufficient, Red Hat OpenShift Virtualization delivers a full virtual machine (VM) boundary on the same platform and the same operating model, converging VM and container workloads rather than  splitting them across two stacks. Confidential containers with attestation-gated keys add a third tier, where the operator itself cannot read tenant memory. Three assurance levels, three price points, one platform - and the operator can sell all three without running three  infrastructures.

Red Hat Advanced Cluster Management drives the fleet, acting as the lifecycle and orchestration engine for the entire estate: 

  • It provisions and imports clusters
  • Creates and updates tenant namespaces and quotas
  • Instantiates Hosted Control Planes on demand
  • Applies governance policy across every cluster, and; 
  • Continuously reconciles configuration drift. 

Cluster lifecycle, security boundaries, and policy enforcement are owned end to end by Red Hat, at  fleet scale, through automated workflows that upstream systems invoke rather than replace.

Red Hat AI Enterprise runs the AI. From model development and validation through distributed inference and lifecycle management, Red Hat AI Enterprise delivers the full AI lifecycle under a single subscription, a single lifecycle, and a single escalation path. For service providers,  that is five separately priceable tenant services - models, notebooks, guardrails, evaluation, and red teaming - from one platform, with no second commercial agreement. Its llm-d inference layer performs KV-cache-aware routing and prefill/decode disaggregation, sustaining two to three times the throughput of round-robin distribution on prefill-bound traffic. Cost per token and the number of priceable services are the two  levers that determine whether an AI cloud has a margin; Red Hat moves both before anything is priced.

 Red Hat Enterprise Linux (RHEL) and RHEL CoreOS establish trust at the base through FIPS 140-2/3 validated cryptography, SELinux mandatory access control, signed immutable node images with secure boot, and GPU driver lifecycle management. Red Hat's hardware certification program validates RHEL and RHEL CoreOS across thousands of OEM, storage, and networking configurations, so operators can adopt next-generation hardware from any certified vendor without rebuilding or  re-qualifying the platform.

Sovereignty is a platform property, not a wrapper

This matters most for sovereign and regulated AI infrastructure, where the requirements land on the platform layer directly.

Sovereignty frameworks now grade operational control, technological autonomy, and assurance above residency. The 2026 question is the provenance and response posture of every layer between silicon and agent - supply chain is the dominant attack vector, weights load as unsigned binaries,and agents call tools holding real credentials. Against an in-country hyperscaler region, the operator's claim has to be control and inspectability: an open platform it can read and self-support, run by a qualified local party, with a signed chain it verifies itself.

Each layer of the Red Hat platform is developed upstream and published as open source, so the service remains operable independently of the  vendor relationship - the definition of technological autonomy, not merely data residency. Red Hat's update discipline backs it: a named owner for every CVE, fixes backported into the exact version the operator certified, certified versions of the NVIDIA operators, signed immutable node images, and participation in coordinated embargoed disclosure. Red Hat Advanced Cluster Management’s governance policies paired with Rafay's blueprints and drift detection then hold that baseline across every cluster in the fleet.

NVIDIA accelerated computing

NVIDIA provides the accelerated computing foundation the platform schedules against. The GPU Operator handles driver installation and MIG  partitioning. The Dynamic Accelerator Slicer enables fine-grained allocation per tenant workload, while NVIDIA Dynamo, CUDA-X, and DOCA/DPF supply the distributed inference and accelerated networking stack. Red Hat ships and supports certified versions of the NVIDIA operators as part of the platform lifecycle.

Where Rafay fits: the commercial and self-service layer

With the platform in place, Rafay provides an integrated commercial and self-service layer: turning platform capabilities into catalog items a customer can buy, and turning consumption into an invoice..

Rafay surfaces the platform's capabilities as governed, priceable services. SKU Studio lets operators define service offerings that combine GPU capacity with compute profiles, storage, networking, and policy controls, each with rate cards and tenant entitlements attached. The Self-Service Portal exposes them to tenants, white-labeled under the operator's own brand, domains, and UI. Tenant-aware identity access management (IAM) and roles-based access control (RBAC) federate enterprise identity providers into the offering and govern who can request what.

 Rafay's workflow engine handles the commercial path around a request - approvals, quota validation, capacity reservation, and metering - then calls the platform's automated workflows to fulfill it. Red Hat Advanced Cluster Management drives the deep lifecycle management and infrastructure provisioning across the  fleet; Rafay integrates those workflows into tenant-facing commercial catalogs and records what was consumed against the right rate card. The  same relationship holds for site build-out: Red Hat Advanced Cluster Management and OpenShift provision and manage the infrastructure, while Rafay tracks inventory and  sequences the tenant-facing service configuration that sits above it.

Two integration details are worth naming. The Rafay Cluster Controller installs on each OpenShift cluster as a Red Hat certified operator and creates an outbound-only zero-trust connection to the Rafay Control Plane, so sites behind NAT or in third-party facilities need no inbound firewall path. Then Rafay Token Factory adds a metered, multi-tenant billing and API layer on top of Red Hat AI's vLLM and llm-d capabilities, which provide a scalable inference service.  The result is a model in which infrastructure, platform software, and service operations remain distinct but work together as a unified AI cloud with clear ownership at every layer.

Get started

The publication of Sovereign AI Cloud as a Service gives operators a concrete, validated starting point for sovereign AI cloud delivery built on open source, certified hardware, and proven isolation mechanisms.  It also demonstrates how a validated technical architecture can support a larger operational objective: standardizing the platform, simplifying expansion, accelerating onboarding, and creating a foundation that can scale as enterprise AI adoption grows. Read “Sovereign AI Cloud as a Service” in the Red Hat Architecture Center for the full architecture, component detail, and provisioning flow.

If you are building an AI cloud on OpenShift and want to see the joint solution running against your own environment, get in touch and we will walk you through it.


About the author

UI_Icon-Red_Hat-Close-A-Black-RGB

Browse by channel

automation icon

Automation

The latest on IT automation for tech, teams, and environments

AI icon

Artificial intelligence

Updates on the platforms that free customers to run AI workloads anywhere

open hybrid cloud icon

Open hybrid cloud

Explore how we build a more flexible future with hybrid cloud

security icon

Security

The latest on how we reduce risks across environments and technologies

edge icon

Edge computing

Updates on the platforms that simplify operations at the edge

Infrastructure icon

Infrastructure

The latest on the world’s leading enterprise Linux platform

application development icon

Applications

Inside our solutions to the toughest application challenges

Virtualization icon

Virtualization

The future of enterprise virtualization for your workloads on-premise or across clouds