The VMware Reckoning: Broadcom, AI, and the Future of Enterprise Infrastructure

Enterprise IT didn’t plan for this problem—but now it has to solve it. In this episode of The CTO Advisor Podcast, Keith Townsend sits down with Lynn Comp, Head of Data Center Market Readiness at Intel, for a candid, unsugarcoated conversation about how Broadcom’s acquisition of VMware has fundamentally disrupted the enterprise infrastructure status quo. [...]

Transcript 3,398 words · about 23 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

All right, you're listening to another episode of the CTO Advisor Podcast, this one's sponsored by our friends at Intel. I can't believe that I've never had Lynn Kopp, head of data center market readiness at Intel on the podcast. First off, Lynn, welcome to the program. Thank you, Keith. It's good to be here and good to sit down and talk with you. I'm almost a little intimidated, my friend, because you are such an expert in this space. Around this topic for quite some time, so I don't know if intimidation is quite the right word.

Right. You're quite the heavy hitter in the industry. Awesome. Well, it'll be a great conversation. All right, so we are not going to sugarcoat this. We're not going to try and hide the big elephant in the room. VMware and Broadcom have completely disrupted the enterprise compute landscape. This is a problem that enterprise IT decision makers have to solve that they didn't know they were going to have to solve just two years ago. This enterprise workload placement data center was a solved problem.

In the data center, in the large enterprise, we use VMware. There might be a couple of exceptions to the rule, but for the most part, VMware was a solved solution and you built your AI and your cloud on top of that solution. So Lynn, let's start out the conversation. Where are you seeing this catalyst for change when it comes to virtualization workloads? You know, Keith, I think that I'm seeing it on multiple fronts. So first of all, you've got this situation where enterprises have been running on virtualization for many, many years, thought a license meant a license, and they thought a specific license meant that they could have it in perpetuity.

And as you know, they had planned disaster recovery, lifecycle management, all of these mandatory practices for running really mature business on credible infrastructure around this virtualization chassis. And the pressure points that are building are not just economic. We like to talk about the economic issues, but the reality is VMware has an incredible product. They've been developing it for many years. Many enterprises have been able to count on it. And as we've made shifts from pre-Sarbanes-Oxley to current world and environment of reporting and regulation, they've been there the whole time.

And that's not going to go away anytime soon. So as much as people complain about, oh, the costs went up, the reality is I'm not sure that they were factoring in virtualization as a core part of business operations, and we're bridging it more like a software license, quite frankly. Now they're looking at it going, oh my gosh, it's part of my operations construct, and it's not just a software license, and how am I going to blend between those two? And so we get a lot of questions around, are there options to move?

What are the options to move? So developers, developers, developers, if you're thinking about moving from one platform to another one, the number one constraint that I've seen is developers. You can't take a lamp stack or hundreds of lamp stacks and just point them to a target and say, I will migrate them. There's tooling to migrate them. We've tried this. We did a really nice project with Google Cloud to show what it takes to migrate a single app and then get to a development or a developer factory where you migrate these applications.

What we, I think, need is some heavy duty AI, some magentic coding to allow us to replace this constraint of developers and get us to the point where we can move from VMware vSphere to more of a natural OpenShift or Nutanix-based Kubernetes platform. You know, what I'm seeing more is, you know, there's the operational reality of VMware, but at the same time, when you're looking at things that are more cloud-native, that are more AI-native, it does give a real great opportunity for platforms and for considerations around OpenShift and Nutanix and some of these other platforms that are working to carry that into the future a little bit more.

And so the question really, though, is there's, how do you run AI in your infrastructure? And then there's, are we ever going to get to the point where AI can run your infrastructure? And there's a lot of debates on whether that's the case, you know. You could argue that AI makes mistakes and you have to do a lot of checking operationally. You could also argue that humans do the same thing. I think that we're going to be tiptoeing our way into that, and you're going to end up with things that are a little bit less of a profound dramatic shift where, you know, it's not any of the above of what we have right now.

We're going to be operating completely differently. Could you use AI to port things? I know people who have experimented. I know people who've experimented how to move their tech stack through AI capabilities and agents. I haven't seen much of it, to be honest. You know, you usually have developers that are higher level and they just don't care. And then you've got the hardware people who have to figure out how are they going to present something the developers want to use.

And it's really those choices that take you from the legacy infrastructure that we've had in a stable way towards the future operational states. By the time this publishes, I would have published the blog post, but how VMware is now in the oracle of IT infrastructure that it is so ingrained in not just your technology stack, but your entire operating model that simply moving away from it is, you know, kind of this aha moment that you've described that customers have had that, oh, wow, this isn't just a question of replacing the hypervisor with another hypervisor.

This is a question of I've, for some of us that went down the DevOps route, I've integrated literally VMware is the foundation building block of my infrastructure. And I've put OpenShift, Kubernetes, I've put that on top of it. As a matter of fact, when you look at OpenShift's playbook, I was just talking to an architect right before our call who talked about how she's looking at HGX, the reference architectures for NVIDIA on x86, how a lot of that is built on OpenShift that actually runs on VMware.

So it is absolutely integrated part of the operating model. And I think my second question, you've answered some of it, which is how are customers dealing with blending kind of that new world of DevOps, AIOps, MLOps, and this legacy? And I use legacy in air quotes, right? Because I'm running this new world on top of this legacy infrastructure of VMs. Yeah, it's an interesting question, because if you look at what people are realizing, if they try and build their own tech stack offline for private AI, they realize just how much of the software infrastructure is available to them that they really don't have to reconstruct in cloud.

And so I think it's a very similar thing where if you're trying to get off of virtualization, you start realizing, wow, VMware is so much of how I have run all my operations and run all my applications and software. Do I really want to rebuild that? And so it's a really interesting dichotomy. I have been hearing that customers are really looking at very tight definitions if they're trying to repatriate, if they're trying to get out of cloud, because you can't take on the world.

You can't replicate all of those capabilities that the cloud offers you. But with DevOps and containers and virtualization, because it's all been built on the virtualization foundations, I don't see them being able to get out of that easily. Because again, your business data is in a virtualized environment. Now, one way that you could potentially do things is treat it like the way virtualization came into the enterprise, which is there's legacy, and then there's virtualized, and never the twain shall meet. There's serious problems with that, to be honest, because you're thinking about there's blends of legacy storage and legacy data repositories into the new.

But how many duplications are you going to have of the same data in different vector searches and vector databases from your legacy storage to be able to go do an agentic system? They say that the efficiencies will make up for the sprawl, but I don't know that anybody's actually done that math. I think that what I've seen is there's judicious use of access to old data going into the new systems. At the same time, those new systems, I still believe, are being built on top of virtualization and VMware.

I don't think there's really an easy exit plan. Yeah, there's not an easy exit plan, and I've talked to several VMware customers. We are now in a dual phase, in my opinion, of repatriation. I've coined the term the fourth cloud. I don't know if I've coined it, but I've defined what fourth cloud means to me. This idea that the enterprise is running this quote-unquote fourth cloud where they have all of the missing services from the cloud providers that they can now repatriate some of these workloads that were thought to be cloud first just five years ago, and enterprises realize that the economics don't make sense.

I would love to hear where Intel is seeing customers' journeys when it comes to repatriation. Are you seeing a lot of that? Are you seeing customers ask you about it? What are some of the cost realities customers are encountering? The repatriation and going more native is an opportunity for vendors like Nutanix and Red Hat and OpenShift, if you look at it, because you're just going to have to rebuild and do something that's net new. Cloud makes things very operationally easy, especially for when you're hosting applications.

When you're repatriating, cloud has been observed to be 5x the cost of on-prem if you're looking at infrastructure and just renting versus buying. That doesn't factor into the question of, have we maintained the DevOps skills? Have we maintained the operational skills between the application layer and the hardware infrastructure? What are you going to be running between the servers and the applications that you have hosting in your enterprise? The folks at OpenShift just sent me this great video about a customer who had 30,000.

They got their Broadcom shock and they looked at their landscape and they had 9,000 VMs. Where is Intel seeing the opportunity for those customers that have a strong foot in that container, cloud-native world, but still have this large virtualization footprint for their legacy operations? Are they looking at continuing down the bimodal path or is there an opportunity to consolidate on cloud-native here? Yeah, honestly, I think that there's an opportunity to consolidate on cloud-native. Some of that has to do with the enterprise skills and where they're really going to spend their time.

Going more cloud-native is a great step towards going more AI-native if AI-native involves things like AI runs the infrastructure long, long-term. I think that it's a fantastic chance for OpenShift. It doesn't really end there. That's obvious numbers, but when you look at the efficiencies that you do get out of the containers, you don't necessarily get all of the ready-made infrastructure with the containers, but if you're already going down that journey and you're already going to be going more and more AI-native, it's like go all in.

Just go that direction. And then the other end of the spectrum is kind of the people who compete with VMware head-on. One of the number one questions that I get, I'm shocked by it, is Proxmox. Like Keith, what do you think about Proxmox in replacing my VMware landscape? You know what? If you're a small business and you have a handful of virtualization hosts, Proxmox probably is, depending on your complexity, a fine solution. When we get to enterprise scale, you have to understand how much time we spend on just pod design around a vSphere cluster or a VCF cluster.

Where are customers seeing the alternatives for Broadcom when it comes to just the low-level engineering of pods? Yeah. Yeah, I think that that's where the promise of HCI and converged infrastructure and things that are off the shelf, we've got Nutanix has an approach that they've been doing that really is seeing a great pickup in terms of interest in the market. Because the beautiful thing about this world that we find ourselves in and the forcing function back to the original question of Broadcom VMware, is that people are really realizing, hey, there's choices that I can make.

Based on my skill sets, based on my long-term plans with AI, home or not home in cloud or on-prem, I think that there's so many more good options on the table that are getting picked up. It's contextual. It's important to get the right context and not necessarily just say, oh, well, I'm just going to replace what was with this new thing. It's what makes sense in my infrastructure, given where I operate, how I'm going to be implementing the kind of DevOps talent that I actually have access to, and where my business is going in the future.

So Lynn, talk to me about the cost. Because I think where people tend to go is the TCO around the hardware and infrastructure. And I think we've talked about this in the past. I think customers will be surprised where Intel customers are having the conversation. Yeah. I think that there's multifaceted and it does come down to the skills that has been maintained. How much of your on-prem fleet are you maintaining that's consistent with what would be considered best practices and operations for cloud?

Because just about everybody's multi-cloud and hybrid cloud. And so I think that if it's been largely something where you've had very cordoned off infrastructure like HPC, things like that, then you're not going to end up having to keep up with the cloud DevOps skills. And then the question is really, what is replacing the cloud equivalent in your on-prem infrastructure of all the operational stacks, of all the services, of all of the storage APIs, of all of the things that go with building out an infrastructure service?

How are you going to replicate that? And do you still have the skills in-house to be able to develop that? And to what extent do things like OpenShift Nutanix or VMware give you a proxy that you can build on top of more easily? And there's not just the cost of the hardware and the TCO hardware. It's really the TCO of change management, the TCO of build versus buy. It's important to get this right because moving backwards is typically not an option for most organizations.

I just wrote this up that if you migrate from VMware and you start to go down the replatforming route, migrating back isn't as simple as hitting the undo button. This is not saying, oh, I went from Oracle to Postgres and now I want to move back from Postgres to Oracle. This is woven into your operation. So all of these talent factors you brought up is something really serious to think about. One, do you have the talent to move from one platform to the next?

And if you have to move an application back from another platform back to a platform is the problem. I'll quote the famous CEO and chairman of Microsoft, and I'll say developers, developers, developers are the key resource when making this back and forth flip. So let's get a little technical. There's been a technology that I've been looking at for years now that I know is near and dear, dear to your heart, and that's CXL. And specifically, I was really excited about kind of what started out with CXL Xeon 4, Xeon 5 made some improvements.

Where is Xeon 6 enabling kind of this adaptability of infrastructure for folks who have to make these types of decisions? Well, you know, I think that the main thing that's really exciting, and we've all been trying to get to it, there's basic CXL, which was a great advancement, standardized interfaces, you know, be able to get to larger memory pools. But I think people are really concerned or interested in application tiering as well as memory tiering, because it doesn't give you a chance to change the dynamics on storage and memory tiering between DRAM being close and the different kinds of storage, direct network, etc.

And so right now, what we're seeing is some real interest in CXL memory tiering, where you can, you have to do a little bit of homework to really understand is the additional latency of going through a CXL going to be offset by having these huge memory pools. And there are, you know, not every use case in the world is going to be able to do that and take advantage of it for you. But at the same time, there are a lot of very, very large and flat memory spaces that are absolutely going to benefit, like, you know, the potential for database and large database hosting or in-memory databases.

You're not going to be moving things around a lot, you're going to be doing a lot of work in that database. And so the ability to go into this large memory space and stay there and really do your analytics and you know, your ETL processes, if you will, updated maybe for the AI world is going to offset the additional latency that the CXL interface might be interjecting. So to bring this full circle, we talked about Broadcom and their disrupting the market. We talked about alternatives and we talked about AI.

What we haven't talked about, Lynn, is VCF and AI. Like, where's that story and where are you seeing that resonate with customers? Well, I've had a lot of customers tell me, you know, after they get over the grieving process with the disruption, they look at the AI features in VCF and think, wow, there's a lot here. And so, you know, for the right set of conditions in their infrastructure, you really do get a lot from the VMware Broadcom capabilities and it's included as part of VCF.

So, you know, very similar to other licensing deals where the OS comes with the hypervisor. This is a great opportunity to take a look at if you don't think that you're going to be making that radical shift of going completely AI native and 100% in the cloud or mostly in the cloud off-prem, then this is a fantastic way for you to step your way into having as part of just what you would get and the operational stability and the beauty of VCF. It's a great way to start augmenting and giving your employee base an opportunity to start becoming more and more AI native without having a complete disruption of learning totally different operating environments.

So I think it's a great thing. It's a both-and, and customers really do deserve choice in all of these changes that are hitting them. It's really just what is the best way for them to have those anchor points of what are my decision points. And if operational stability and racing to have your employees get better AI capabilities immediately, well, that might be the right choice for you is to just recognize it's an easy way to add in some AI capabilities in a less disruptive fashion.

So you're not paying for all that change management drama, but at the same time, you're not getting left behind by the AI revolution. Well, Lynn, thanks so much for joining the podcast. I'm looking forward to the second episode, which we will dive a little bit deeper into AI. Well, until then, thanks for joining. Awesome. Thank you so much, Keith.