For the past 2 years, the GPU conversation was about supply. Could you get the hardware? How much? How fast? That conversation has shifted, and a growing number of operators now have GPUs in hand, or on order, and they're facing a different problem entirely. Operators are looking at how to turn racks of accelerators into a cloud service that customers can actually buy. This is the neocloud challenge, and the hardware was only the beginning.

This strategic importance is underscored by recent multibillion dollar agreements, such as Nebius’s $17.4 billion deal with Microsoft and CoreWeave’s infrastructure partnerships with Microsoft and OpenAI. These deals reveal a structural reliance where tech giants use neoclouds as dedicated third-party GPU infrastructure proxies to bridge capacity gaps, rather than adopting native platforms. These cloud services require multi-tenant, security-focused, and elastic infrastructure that can be consumed as an API-driven resource. Red Hat platforms are specifically designed to accelerate this transformation.

The platform is where it gets hard

Building an AI cloud requires much more than integrating accelerators into basic Kubernetes setups. It requires prioritizing operational details like availability, performance, security, and compliance. While this article focuses primarily on NVIDIA, neocloud operators can build with hardware solutions of their choice to power cloud operations. Today, more attention has been placed on networking in most distributed training scenarios because remote direct memory access (RDMA), NVIDIA® GPUDirect®, and high-bandwidth fabrics bypass the host CPU entirely to ensure the CPU does not become the bottleneck for training performance, which would be a fraction of what the silicon can deliver. Scheduling must be topology-aware because placing a multi-GPU job without understanding NVIDIA® NVLink™ connections and NVIDIA® NVSwitch™ fabric leaves performance on the table. And failures are expensive in ways that general-purpose clouds don't experience. For example, a multi-day training run that fails wastes GPU hours and erodes customer confidence.

Furthermore, managing multi-tenancy adds another layer of complexity. Tenant isolation has to extend to GPU memory, NVLink domains, and GPU health, not just Kubernetes namespaces. Standard namespace isolation is a starting point. Layer on the security expectations from enterprise and government buyers—whether it be domestic US standards like Federal Information Processing Standards (FIPS) and Federal Risk and Authorization Management Program (FedRAMP), or global benchmarks like ISO 27001, Bundesamt für Sicherheit in der Informationstechnik – Cloud Computing Compliance Criteria Catalogue (BSI C5), and General Data Protection Regulation (GDPR) compliance—and you're looking at a compliance investment that often takes far longer than operators plan for.

Additionally, none of this process is static, because GPU firmware needs updating across a live fleet. For example, new architectures call for driver and scheduler changes. Then, customers start asking for fractional GPUs, confidential computing, and Model-as-a-Service endpoints. This growth of the market and services means the platform is never done. Maintaining long-term momentum also requires an underlying platform that evolves dynamically and scales seamlessly without downtime for active services.

The build-vs-buy question is a time-to-revenue question

Every neocloud operator faces a build-or-buy decision based on resources available. Rather than balancing control against convenience, the true focus should be on maximizing time-to-revenue based on your team's existing capabilities and financial runway.

Building in-house works for operators with large fleets and the engineering organizations to match. Some established GPU cloud providers have invested heavily in custom Kubernetes distributions, proprietary schedulers, and bespoke control planes. However, regional GPU cloud providers, sovereign cloud operators, and enterprise data centers adding GPU-as-a-Service need to get from bare metal to revenue quickly. Every month of platform development is a month of GPUs generating cost but not income.

There's also a middle path that gets overlooked. Operators can adopt a platform for the infrastructure and orchestration layers, then build the differentiated services on top. The platform handles bare-metal lifecycle, Kubernetes, GPU scheduling, and multi-tenant isolation. At the same time, your team focuses on what creates competitive advantage, whether that's pricing models, specialized model hosting, industry-specific AI services, or edge deployments.

The platform is critical infrastructure, but isn't the product you're selling. The product is the GPU cloud service your customers buy. For example, NxtGen Cloud Technologies adopted a platform-based approach for their SpeedCloud offering to accelerate time-to-revenue. By using Red Hat as the foundation for infrastructure and orchestration, their engineering team bypassed the months of development typically required to build a proprietary control plane. This allowed them to prioritize delivering differentiated AI services directly to their customers, rather than managing the complexities of building a custom stack from bare metal. 

NVIDIA Cloud Partner reference guide

The NVIDIA Cloud Partner (NCP) Software Reference Guide establishes clear requirements without being prescriptive about specific products. The list of requirements includes infrastructure management, compute orchestration with GPU-aware scheduling and multi-instance GPU (MIG) partitioning, RDMA and GPUDirect networking, AI platform services like inference serving and model management, GPU-specific observability down to XID error tracking, and security and compliance for multi-tenant workloads. Meeting all of these requirements with a DIY platform is possible, but meeting them before your runway runs out might be a challenge.

Where Red Hat fits

Today, Red Hat is an NVIDIA AI Cloud Ready partner and an inaugural member of NVIDIA's Open Secure AI Alliance. As a software provider, Red Hat enables operators who need a production platform to run their AI cloud. How that plays out depends on where you are as an operator.

  

If you built your own Kubernetes, you might need a production-grade inference serving stack. Red Hat AI Inference, powered by llm-d, provides disaggregated inference with prefix-cache-aware routing, KV-cache offloading, and multi-backend support, all designed to improve GPU utilization and support multi-tenant inference without requiring you to rebuild your control plane. llm-d is a Cloud Native Computing Foundation (CNCF) sandbox project co-developed with CoreWeave, Google, IBM, and NVIDIA, with validated deployment blueprints available for CoreWeave CKS and Azure AKS.

If you need the full path from bare metal to revenue, the story is broader. Red Hat Enterprise Linux (RHEL) provides the enterprise host OS foundation, with NVIDIA® GPU driver and NVIDIA® CUDA® runtime support on RHEL, which means fleets come up as managed Linux, not one-off server builds. Red Hat OpenShift adds GPU-aware scheduling, MIG partitioning, topology-aware placement, GPU health monitoring, and bare-metal installer-provisioned infrastructure (IPI) for automated cluster deployment, cutting the time from racked accelerators to a production multi-tenant Kubernetes environment. Red Hat OpenShift AI builds on that platform for training, serving, and model operations when customers expect managed AI workflows.

Red Hat AI adds model serving, training pipelines, and model registry when you want to deliver managed AI services on top so your team ships the cloud service customers buy, not another platform rewrite.

For providers moving toward industrial-scale deployment, Red Hat AI Factory with NVIDIA delivers a co-engineered, unified architecture. This collaboration replaces fragmented components with a production-ready system that merges Red Hat’s enterprise orchestration with the NVIDIA accelerated computing stack, including NVIDIA BlueField. By using validated blueprints and rapid-deployment guides, operators can bypass significant architectural hurdles. This lets your engineering team prioritize delivering services that generate revenue, supported by a hardened, consistent environment optimized for rack-scale AI workloads. Additionally, this partnership provides day 0 support for new NVIDIA hardware and software releases, enabling operators to deploy the latest AI technologies immediately without waiting for long integration or validation cycles.

Security and compliance frequently becomes a major friction point for operators. For this, RHEL includes FIPS-validated cryptographic modules, Red Hat OpenShift supports FIPS mode deployment, and the stack provides the audit logging and access controls that Service Organization Control (SOC) 2 auditors and government procurement processes expect. For operators who can't afford a multi-year compliance build before their first enterprise deal, that's a practical differentiator.

A pattern worth watching: The edge neocloud

Zero Latency (0.lat) is building a distributed neocloud inference network across U.S. edge data centers using Red Hat AI Factory with NVIDIA. Zero Latency’s “Zerogrid” network has the GPU hardware and data center capacity, but needed a platform to operate it as a cloud service, and chose to adopt rather than build. Their deployment also illustrates an emerging pattern that more operators will need to solve, where a common platform unifies distributed GPU infrastructure across multiple edge locations with consistent security policies and centralized observability. As latency-sensitive inference pushes compute closer to end users, managing GPU infrastructure across many sites becomes a fundamentally different problem than running a single data center.

Why governance remains a top line priority

One dimension of the platform decision that deserves more attention is who controls the software your business relies on. GPU cloud operators are making multi-year infrastructure commitments, and the governance model of the platform you choose, who controls the roadmap, who can fork it, and who has access to the source directly affects your business risk. Platforms built on open source projects with community governance, like Kubernetes, KubeVirt, and llm-d under the CNCF, reduce single-vendor dependency. If the vendor changes direction or raises prices, you have options. Proprietary platforms don't offer the same flexibility. This isn't abstract. Licensing terms and governance matter as much as features when the software will underpin your business for years.

The window is open, but it won't stay open

Demand for AI compute still outruns available capacity in many markets, which is why new operators keep entering with hardware, searching for the platform that lets them compete. The operators who get the platform right will capture the next wave of demand. Those who spend too long building a custom platform risk finding that the window closed while they were still in development. Building a GPU cloud is no longer a question of if, but how to execute efficiently, while maintaining security and preserving economic viability.

Ready to go deeper?

Meet us at AI Infra Summit 2026

Ready to move beyond the complexities of building your own GPU platform? Join Red Hat at the AI Infra Summit 2026 in Santa Clara from September 15–17. Stop by Red Hat booth #222 for live demos and to discuss how to simplify your AI infrastructure and scale your cloud service efficiently. Learn more and register here.

Resource

The adaptable enterprise: Why AI readiness is disruption readiness

This e-book, written by Michael Ferris, Red Hat COO and CSO, navigates the pace of change and technological disruption with AI that faces IT leaders today.

About the author

Adam Wealand's experience includes marketing, social psychology, artificial intelligence, data visualization, and infusing the voice of the customer into products. Wealand joined Red Hat in July 2021 and previously worked at organizations ranging from small startups to large enterprises. He holds an MBA from Duke's Fuqua School of Business and enjoys mountain biking all around Northern California.

UI_Icon-Red_Hat-Close-A-Black-RGB

Browse by channel

automation icon

Automation

The latest on IT automation for tech, teams, and environments

AI icon

Artificial intelligence

Updates on the platforms that free customers to run AI workloads anywhere

open hybrid cloud icon

Open hybrid cloud

Explore how we build a more flexible future with hybrid cloud

security icon

Security

The latest on how we reduce risks across environments and technologies

edge icon

Edge computing

Updates on the platforms that simplify operations at the edge

Infrastructure icon

Infrastructure

The latest on the world’s leading enterprise Linux platform

application development icon

Applications

Inside our solutions to the toughest application challenges

Virtualization icon

Virtualization

The future of enterprise virtualization for your workloads on-premise or across clouds