AI in the Data Center: When Faster Isn’t Better

AI is changing the data center—but not always in the ways enterprises expect. In this episode, Keith Townsend is joined by Intel’s Lynn Comp for Part Two of their conversation, shifting the focus squarely to AI infrastructure realities. They explore why many AI workloads never justify GPUs, how CPU-based deployments often exceed real [...]

Transcript 4,127 words · about 28 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

All right, this is part two of a data center. I. is changing our data center. And the last episode we talked with Lincoln about the impact of the changes that Broadcom has me purchasing VMware to our infrastructure. I. portion of the the challenges that we're facing. Lynn, welcome back to the show. Thanks, Keith. I'm excited to have the next conversation here because, you know, we know that this conversation is evolving really fast in the industry. And so what was it a month ago that we spoke?

And lots of things have happened in between there. So lots to cover today. Yeah. I want to I think I want to set this one up. I. They do some pretty amazing things. Basically, they're looking to improve yields one to three percent every growing season. I. to do this. They have a cluster of Dell nodes running Intel processors and they added NVIDIA GPUs to the mix. I. processing, you know, a 10 X performance increase.

But what they discovered was that their SLA for getting the results back from the modeling data was 24 hours and GPUs reduced it to, you know, the performance down to maybe 30 minutes for processing. But what they discovered was their CPUs were already meeting the SLA by like three X, the requirements, like eight hours. And at the end of the day, they increased the complexity of their environment trying to manage. And I'll link it in the in the show notes. But they ended up with increased complexity so much so that they pulled the GPUs out of their cycle.

Are you seeing similar integration pain points with other Intel customers? Oh, yeah. I've seen at least three different situations. Can't mention the names, but, you know, one of them is is one of the largest, you know, data center owners that you would recognize as name brand value. And, you know, essentially they start out using CPUs because they can't get access to GPUs or they can't get the power or, you know, they have to wait a year and they can't wait on their deployment model.

And then what they find is once they get it deployed on on the CPUs, that it works for the SLAs they have. And I think that that's the main the the nuance here is it's easier to not think about SLAs and to run fast and just think I'll just fix it later. But the problem is if you don't understand your SLAs, then you're not going to understand how to have ROI in your model or how to calculate it. You'll just you know, you'll take the safe route, which is specific hardware.

But then what you didn't do is actually build in the view into what kind of work do we need to get done? And what is the definition of efficiency for this kind of situation? Yeah, I got to give my good friend Bobby Allen credit for this saying many companies ago that he worked for. He had the saying that you don't give unlimited capacity and capability to solutions that have a limited business value. So, right. This idea. And we've experienced this with cloud, right?

That, yes, we can architect something to go exceptionally fast. We can scale up to and this is the context in which he was talking about it. We can scale up to a thousand nodes to a process that doesn't have any business value beyond it running on one VM. And I'm seeing this pattern being repeated in AI. But. I think it's easy to make the argument and we talked about this a little bit before the start of the recording. I think the counter argument to that is that if these processes are going to eventually end up on AI or end up on GPU, doesn't it make more sense to start it out on GPU, even if, you know, you're going to be wasteful in resources and money?

I don't think many CEOs like that idea, but, you know, they don't have to replatform later. Intel has to be hearing that story in market land. Talk to me about this, this portability as a design goal. Well, I think that the portability isn't necessarily at the library level. That's not where you're really going to get it. You're really going to get it at the partner level. You know, so some of our partners, Kami Waza, Iterate, Iternal, there's there's a number of them and they can run on multiple different configurations.

NetApp is another one that is finding that there's some really good use cases. And not everybody needs the full blown, the highest end agentic. You can use a lot of small language models. In fact, NVIDIA just published in September that small language models are in many cases more effective. And we see a number of customers using vertical models. They're specific to their industry. And so the question is really is our CPUs and GPUs. You know, CPUs are more general purpose than GPUs.

GPUs are general purpose when they're running these large language models. And the question is really what gets you to the industry specific implementations. T. teams are facing is I need to get breakthrough usage models that can operate my infrastructure without spending any extra money. And I'm not thinking anything about change management. So I think that the industry is starting to realize we need to slow down to speed up a little bit. Now, you can tell me I'm wrong on that, but I find find so many instances where there's this.

I start with a CPU because of different assumption set. And then I find I don't really have to move. And if you do the portability correctly and the architecture correctly so that you can have a solution that's accelerated or not. And, you know, with with Intel Xeon 6, you could be using an accelerator on deck called AMX. That's really the important design criteria for having optionality as you move forward. Yeah, I don't think you're going to get a lot of pushback from me today.

I might have pushed back a couple of months ago when I built my CTO advisor, virtual CTO advisor, and then my CTO advisor stack builder that kind of implements a lot of our methodology. It was a really cool solution. I built it in about a month and I built it on Google Cloud. And then I got a little worried. I said, well, I don't want to be locked into Google models, Google pricing, et cetera. And I went to move it to an on prem solution.

I bought an NVIDIA GPU based system. And when I went to move it, I realized that I was locked in. I was indeed locked in. I didn't build to like CUDA or anything. I built to the API and I could move my Google vertex APIs on premises. So you're not going to get pushback on me on that, but this does raise the question of governance. And as architects, as CTOs are thinking through governance, what are some of like the smarter things you've seen customers do on the governance side?

Because this is not just a problem where you're talking about, do I want to run on X86? Do I want to run on all GPUs or even if I want to abstract it to the on the API layer? The governance question spans all three of those approaches. Yeah. You know, what's interesting with governance is since I ended up studying for and clearing the AI governance professional exam, there's a lot of things that have shifted. You know, you see some of the large cloud service providers saying, I'll give you a box.

You know, you can have a box with my stuff in it and that will be governed and it will cover the sovereignty and the governance. You've seen a lot of groups in other countries just saying, I'm just going to build data centers. And then you can ask yourself, you know, is that sustainable? You know, do they have the full production capacity for that? And so basically there's the orchestration and then there's where things are running and then there's where the data is residing.

And I think that we've gotten so cloud native, there's going to be some unpeeling that has to happen with where the data is placed and what they're giving access to and how they're doing things and taking advantage of things like confidential computing and, you know, TDX and all of that platform visibility and telemetry that you can get access to. And so it's so multifaceted for multinationals. I think it's almost simpler if you're a smaller company operating in a more limited region, you're not having to deal with things like the European Union AI Act and then Singapore has its own AI Act and Canada has its own AI Act.

And so does, you know, Brazil have its own policies. So I think this is going to be the next wave of AI. We got the technology there. There's a lot of great agentic solutions that can provision infrastructure, that can check in JIRA tickets, that can flag when things are happening, maybe even suggest fixes. But at the end of the day, who's signing off on it? So I'd be curious in your perspective on this. And, you know, one of the other governance questions that I've been asking is if using AI tools to write code means that you don't own the code, then what's the governance over your own innovation there?

So lots to unpack there. Yeah, if we could probably do a whole podcast series on this topic alone. I've wrote what I'm calling the reasoning layer. But whether you're automating this or doing it manually, you have to think through how you're managing governance. How do I ensure that data that my general model has access to in Europe isn't being leveraged for stuff in the US that European law doesn't allow you to do? And you have to really think through this entire challenge of data sovereignty, the resulting AI intelligence or AI data that comes as a result to it.

I've talked a lot about the unintended consequences of when you give someone access to something like a chat, GTP or copilot, and they have access to data that they have access to. But the insight, because you can now do data scientist level work with this $20 a month tool to get insights, a layer above your pay grade. How do you solve those problems? And I think this podcast, this single podcast about, you know, just the general problem of workload management by itself isn't enough time to go over it.

Well, it's interesting you mentioned that because like, you know, I deal with some of these tools internally. And what I've found is since the deployment of it, we've actually gotten stricter and stricter data silos. And so then you ask yourself, well, what's the $20 a month subscription for if things have gotten even tighter in terms of what we have access to? And so that's become a challenge as well, is you have an overreaction within some organizations where it becomes very difficult to be your own data analyst.

Yeah, I've talked to a lot of CAIOs, Chief AI Officers, who have simply just said no to a lot of AI tools. So part of RFPs is the question for SaaS solutions. Can you turn off AI until the business can catch up? And this is the business. This isn't an IT thing. This is the business is saying they want to slow down until they figure out the repercussions of AI. Let's shift the conversation back to technology. I can't help but recognize that AI is a memory bound workload.

Right. Intel has worked with CXL for several years now. Where are you seeing the impact of CXL on these enterprise AI workloads? Well, first it started out with memory attached. And now CXL is raising the bar so you can have accelerators attached. But that really comes down to higher memory utilization through pooling. Because if networking isn't your bottleneck memory, your storage is going to be your bottleneck or storage over a network. Essentially what CXL is envisioning of that was free up stranded memory, allocate it where the workloads need it most, do dynamic memory provisioning.

And so without having to reboot, you can basically reallocate flexibly and then can actually reassign AI accelerators without rigid coupling to the host. And so, you know, all of those are our workloads. And, you know, it's funny as I have from the Xeon desk series and I got some interesting comments on the one about CXL of, you know, it would be really, really interesting to sit down and talk through some of the challenges. But, again, there's so many opportunities with being able to dynamically adjust memory capacity available to a system.

It's going to be one of those areas that we keep posting and can't mention their name. But there's a very large, large company that has been able to do a memory mode using CXL attach that has really unlocked some of the data capabilities that are available. And so, again, it's a technology that we envisioned. It takes a little while to get through the enabling of the ecosystem, but the unlock that it can give customers is huge in terms of efficiency, performance, latency, and then not having to have restarts within their overall pipeline of getting things executed.

Yeah, I'm thinking through not just AI, but traditional analytics. What happens when, you know, I have my NetApp storage array that's aware of the CXL capabilities and I'm able to dynamically move or expand compute across several nodes to ingest data as a streaming in from this NetApp array or the ability to expand memory and do some really cool stuff like this? Stuff that I'm expecting Intel to figure out is this hit rate that you are working on when you were putting memory dams stored, basically storage memory dams inside of inside of servers.

And it's the same IO problem. How do I ensure that the compute that is needed to process the data and the memory that's needed is closest to the data? Data has gravity. And once CXL continues to mature, I can easily see parking back to our example earlier in the podcast of a company like Nature Fresh Farms. Again, extending the need to not have GPUs because they're they're they're they're troubling. They're solving the bottleneck until the bottleneck actually becomes real time latency.

And they just need to work through a batch process, you know, within this framework when their data grows. If CXL is there now, they don't again, they don't have to buy new equipment and make their environment more complex than needed. Yeah, I mean, I think it's it's it's one of those areas there's so many degrees of freedom. The entire data center architecture seems to be completely reopening in terms of, you know, what do efficiencies look like? How is it being managed?

And, you know, back to the integration pain point, I think that there's this huge question around how do you, you know, how do you have stability in the middle of all of these new architectures and deployments and where does that stability point come from? And I think that's that's one of the main integration pain points or integration challenges for every one that has deployment. Do you you mentioned you've got a bunch of CAIOs who basically are, was it CAIOs or CTOs that said, no, you don't get to put...

Yeah, they're CAIOs. These are not even, you know, they're not even IT people. Right. And they're saying no. And, you know, a lot of the AI capabilities that seem most natural are coming from vendors that are bolt ons into the into the properties that are already in the fleet. And so I think that there's so many considerations. I do not envy what some of the CTOs are dealing with these days. But when you start peeling it back to first principles, starting with governance, starting with architectural portability, so you've got optionality, use what you have in the fleet.

Look at how you can leverage memory and storage technologies more effectively. You end up with with at least a simplification in the middle of all of the AI change that that's constant. All right. So I'm going to get the last question, Lynn, it's going to be spicy. Intel is obviously going through a transformation right now, and you lead a entire team of folks at Intel, and I'm quite sure I've talked to them. You're asking them to do more with less resources than they've had in the past.

How are you and your team measuring success when it comes to AI projects so that not only is Intel as a technology company promoting the use of AI tools, but you're showing the way. How do you measure success with AI projects? It depends on what type of AI project it is. You know, an AI project metric of success in a design team is going to be different than in a go-to-market or marketing team like mine. In the go-to-market and marketing team, a lot of it is coming down to removing hours on non-productive things, like removing hours recrafting content, removing hours sitting in meetings when you can take a transcript, and then you can actually feed it into a pre-built agent, have that agent operate like a co-worker, and look at the transcript and then do the write-up.

Not just as a note, but potentially technical documentation as a result of sitting in that meeting. Right now, what I'm seeing as the ROI, at least for the collateral production, the messaging, the positioning, and planning, quite frankly, is do I have to write the document or can I have something else get the data access that's going to generate the document, and then I can work with that agent so that it is basically going to get 90% of the way there and I'm having to only correct a little bit more.

I know for a fact that there are software teams and hardware design teams that are measuring things slightly differently, which is how can I get to a design solution that's 80% or 90% of the way there without actually having to spend the weeks of analysis that we do. Can we feed all that data and all that history in and then basically have the agent learn from that and then tweak it at the end? I guess in some sense, that might be cost of design goes down because you're spending less time, and in my case in marketing, it's also time goes down, which is effectively cost, and so maybe you could argue that those are the same measures.

I'd be curious if you see those as the same as I described them to you. Yeah, I felt like you're sitting in my own internal meetings. I stretch a point across all of that. I use AI as a collaborator. I'm solo again, so I don't have the access to 10 other analyst brains to bounce ideas off of. I'll go use an LLM, give it an idea and say, you are this analyst with an opposing idea. Critique what I've presented to you.

At the same time, I'm developing code. Man, I haven't developed code in probably 15 years, and I can now develop code. I'm trying to figure out, and this is why I asked you the question. I'm trying to figure out how to codify this. For a solopreneur, obviously, this is 10x my operations, and when I talk to successful marketers, we're working with a great team at Articulate to talk about how they're codifying the same things. I'm looking for these patterns. I think it was a little bit unfair question for me to ask because the industry hasn't figured out how you actually quantify the thing.

Lynn, you're one of the smartest people I know in the industry, so I'm going to ask you the hard questions. I think the one thing that I would say that is the challenge in that, Teeth, is if you end up seeing some of the cuts that we're seeing in the industry and everybody's solopreneur, then how much space is there for that in the long run? I do think that there's some downside consequences that we tend to whistle past the graveyard about that I'll put back in your camp.

Where do you see things going? It is going to completely change the economy long term or maybe even short term. I'm not bullish on the inference flip happening in the very short term. I think enterprises have a lot of inertia that solopreneurs get to cut through. But when it is figured out, we're going to see a complete rethinking of organizational charts. I know you've given this a ton of thought as your team is probably a little bit more advanced and using these AI tools.

How are you going to mentor and how am I going to mentor the next generation of analysts and advisors? We're seeing it at PWC, how they're saying, you know what, we're hiring junior analysts out of college and we're not giving them the same assignments anymore. We're giving them assignments on managing AI. How do you take a Keith Townsend four plus one framework on AI infrastructure and give this now to a junior associate? And they can use my stack builder. They can use my virtual CTO advisor instance to now basically bring me into meetings.

That's, you know, that's a compelling from a productivity is compelling perspective. But what about the middle folks? What about the folks who did that? What about the pricing? What about the pricing too, right? This actually puts pricing pressure on me. If somebody virtual CTO advisors out there, if you go, if you ask it to write a white paper, it will. And it comes very close to my voice. And I open all of that. I allow people to do it and because I can't fight it.

So it's a threat to my own business model. It is an amazing, amazing and scary time. Yeah. So this is as always been an incredibly fun conversation. Folks at Intel are doing some pretty creative stuff. You folks are in a really interesting point in the industry. You're still the de facto X86 leaders. There's this question as to how much inference capacity do enterprises really need? There will be an inference flip. I'd be lying to you if I told you I knew if that inference flip is going to mean that people are going to buy less CPUs, more CPUs.

I think you'd argue more or if they're going to buy more GPUs or less GPUs. By the way, we are the most widely deployed host node. So from the standpoint, when I get asked what's Intel strategy, our strategy is to be the best solution in all the solution options. And I think there's a lot to be said. I'm working with your team on other research on the importance of CPUs and the entire AI data workflow. We'll get into that in future sessions.

But Lynn, I appreciate you stopping by. Awesome. You have a lot of work to do. I have a lot of work to do. If this comes out before the end of the year, Happy New Year to everyone. If it's at the beginning of the year, Happy New Year to everyone. I hope you enjoy a great holiday season. Lynn, I wish you a great holiday season. Thank you, Keith. It's always good to connect with you.