AI & Machine Learning · Cloud · · 6 min read

Private AI Won't Be the Center of Your AI Universe. Hybrid Will.

In my last post I poked the bear. I said private cloud lost, and we’re going to see the same trend in private AI. If you’ve followed the blog or my content for any length of time, that shouldn’t surprise you. I’ve long said that private cloud is a fool’s errand from the position of trying to displace public cloud. That’s not to say private cloud doesn’t have its use cases. We identified them in that post.

Private AI is similar. But I need to be specific about what I mean by private AI. I mean serving models from your data center in an attempt to displace OpenAI, Anthropic, xAI, and similar services. That’s the thing I’m calling a fool’s errand. Not the hardware. Not the platform. The ambition.

Private cloud succeeded. It just didn’t win.

Look at the private cloud landscape: Nutanix, VMware, Red Hat, Kubernetes, OpenStack. There’s no shortage of ways to build a private cloud, and they succeed in the marketplace. VMware Cloud Foundation is doing fine. So where’s the gap between that commercial success and my assessment that private cloud doesn’t beat public cloud?

The gap is in how enterprises actually consume these platforms. They’re consumed like traditional virtualization and container management platforms. Steady-state applications run on them. But those applications are more and more integrated with public cloud services. What we’re seeing is a bifurcation of runtimes: the steady-state, static workloads run on premises on these platforms, and they lean on public cloud resources and services for everything that moves.

I’ve called this hybrid infrastructure for years. It’s the marriage of two realities. On-premises infrastructure and public cloud infrastructure will continue to coexist. I’ve been consistent on that, and I’m going to be consistent on what comes next.

This is what AI is going to look like

Hybrid AI will dominate the landscape. There won’t be a winner and a loser when it comes to where you run AI.

But here’s where I think the argument usually goes wrong, including in my own last post. The mistake is thinking hybrid AI is about where the model runs. The model is only one component of the AI application.

An AI outcome is produced by a set of services, and those services will be distributed across multiple execution environments. Retrieval may be local, because the data is. Identity is the enterprise’s, but how many of us still run our own identity service instead of pointing at Microsoft Entra ID? Policy may be enforced somewhere else entirely. Tool execution may have to happen inside the corporate network. Memory, observability, evaluation, and orchestration each have their own placement requirements. And the model may be OpenAI today and a local model tomorrow without any of the rest moving.

That’s not on-prem GPUs plus cloud GPUs. That’s a distributed application architecture. It’s the same bifurcation we saw with private cloud, where the steady-state runtime stayed on premises and leaned on public cloud services for everything around it. It’s also why private AI won’t win even if an enterprise eventually owns enough compute to satisfy most of its model requirements. It still won’t own every service surrounding the model.

What orchestrates the system?

Once you see the AI application that way, the question changes. It’s less about where the model will run and more about what orchestrates the entire system we just described. Something has to hold retrieval, identity, policy, tools, state, and the model together as one governed workload, and decide, per request, where each of those pieces is allowed to execute. I’ve been calling that the reasoning control plane, Layer 2C in my own framework. It isn’t a model picker. It’s the thing that makes a distributed AI application behave like one application.

Look at what Microsoft announced this week with NVIDIA. They doubled down on HydraFusion, the model orchestration they announced for GitHub Copilot last month, and extended it to Windows. Copilot will now decide whether a task runs on a local model, on the device in front of you, or on a cloud model.

That’s a model router, and it’s useful. It’s also the narrowest slice of the control plane the architecture needs. Model selection is one decision among many. Where the reasoning happens, where the data is retrieved from, which tools are permitted, where state is allowed to persist, which policies apply, and which execution environment is acceptable for this particular workload are the rest of them. Microsoft is showing you one slice. The rest has to come, because a distributed application can’t be governed one dependency at a time.

Placement, by comparison, is the more straightforward part. Plain and simple, there will be models that run in the data center. Many of you pushed back on me after that post and said as much, and you’re right. If I need a classifier, a basic 4 billion to 26 billion parameter model, it can run on CPU or GPU, and it will run in my data center. If it’s batch processing, same answer. We’ve proved in the lab that CPUs are more than good enough for batch work on models in the 26 billion parameter class. There’s no need to escalate that runtime into the public cloud. And when an application requires frontier-class reasoning, today that favors the hyperscalers and the model providers, on capability, on pace of improvement, and on economics. That can change. The control plane makes those calls, and remakes them when it does. The hard part is building the thing that makes them.

So where does the hardware go?

We’ll have a combination of runtimes, and some of them need to be close to the business activity. Think edge devices. That’s why AMD is investing in the Ryzen AI Max+ PRO 495 with 192 GB of unified memory. That’s why Microsoft is shipping Surface machines with NVIDIA silicon in them, and why NVIDIA is shipping DGX Stations. These devices will do real work. I own this class of hardware and I run real work on it.

The data center tier nobody has sized yet

What’s missing from that picture is the data center tier the reasoning control plane will place work on. There’s an undiscovered sweet spot there, and it’s bigger than the deskside story suggests. Anecdotally, I’ve seen more RTX PRO 6000s in production than I can count. Customers are running digital twins of entire manufacturing environments on them. That’s not a pilot. That’s steady-state, on-premises AI doing work that belongs close to the plant.

What we don’t know yet is what sits between an RTX PRO 6000 or DGX Station class system and a rack-scale AMD Helios cluster. That middle tier hasn’t shaken out, and the vendors haven’t shown us a shape for it.

Common enterprise architecture reasoning tells me how it ends. Like most higher-end compute, enterprises will rent the specialized capacity as needed rather than own it. It’s the exact same shape as yesterday’s high-performance computing (HPC). But the economics aren’t the point. The point is that rented rack-scale capacity becomes one more endpoint the reasoning control plane can place work on, alongside the workstation under the desk and the hyperscaler’s frontier model. Owned or rented, it’s an execution environment. The control plane decides when it’s the right one.

But to say the deskside and data center hardware will be the center of our AI universe in the enterprise? Generally speaking, that’s a fool’s errand. The center is the reasoning control plane. The hardware, and every other service in the stack, is wherever that plane places the work.