From CapEx Chaos to AI Scale: Kamiwaza CEO Luke Norris on the Future of Enterprise AI Infrastructure
📄 Episode Summary / Show Notes: The AI infrastructure game is changing—fast. In this episode of The CTO Advisor Podcast, Keith Townsend sits down with Luke Norris, CEO of Kamiwaza, just steps from #NvidiaGTC, to unpack what planned obsolescence really means in a world of $100K GPUs with 18-month lifespans. 🔥 In this episode, we [...]
Transcript
All right. You can't tell it, but we're just a couple of blocks from GTC. Literally. And Luke, I've had you on a few times. You're now Kamawaza. First, two congratulations. First, congrats on the funding. That was a big deal. That was a really big deal, actually. Colorado company, it's not the easiest, man. No, it's not. A Colorado AI company. A Colorado AI company. People are looking at me sideways, let me tell you.
The first day, I'm like, are you guys seriously in Colorado? It's the part of the fish tank, like, seriously, are you seriously in Colorado? That's kind of how it started every time. So, congrats on the funding. That is a big deal, especially even with AI. Everyone has really great ideas. I actually talked to a guy at the Biancuda event yesterday that absolutely is following in your trails, and his idea, and the challenges. So, I just saw Matt, and we'll probably post a picture somewhere here, with his Apple goggles on.
More than likely coding. He was absolutely coding, yes. More than likely coding. But the second piece of congratulations is plant obsolescence. You got scared, what, about a year or so ago, when you said NVIDIA is out just making sure that you cannot use the $100,000 GPUs you're buying today. You won't be able to use them a year, year and a half from now. They had to do this. Yes, thank you, by the way, on both of those. And if you just ran the math for the TAM that they're filling into, the fact that they were a $3 trillion company, meant so you weren't buying this on a normal depreciation cycle.
You couldn't buy an H100 and expect five years, six years, seven years, like you would try to eke out of a hard drive. To have that level of volume, to have that level of revenue, they have to force their customers to buy, typically, it looks like a 12-month cycle, 18 months on the long end, the most expensive CapEx ever in IT. And they're doing it, and everybody scared me. And I mean everybody. I think that one post on LinkedIn got to almost 400,000 impressions.
It was one of my widest pushed out ones ever. And it's because everybody's thinking it's just a hardware game. It's a hardware matched with software. ARM processors, unlike x86 processors, have a reductive set. So from one generation to the next, if you want a feature to be added into it, you have to have a net new chip to have the net new software they've written now to be able to access it. That means the previous version literally can't utilize that software for that new feature.
And they're marching that obsolescence. And today, today, he literally said, if you have hopper chips, the H100s, you cannot give them away. Give them away. Once Blackwell is fully up to value. Yeah, he said 40x, I think, was the performance. He said per watt was 40x the performance. So, you know, man, you've gone back and forth on this quite a bit. Because H100s are the workhorse. They literally are. People, and I'm having trouble wrapping my head around just how much better Blackwell is.
Because people are doing amazing things with H100s still. And to get 40x the performance, I think one of the first questions is, is the demand there to get to fill up all of those tokens? I see it as paramount. The H series was so powerful, both from its token throughput, but its thermal throughput. It is literally a power hog from a heat and dissipation perspective. If we're the same wattage, you could get 40x more power. Means a typical enterprise with a typical data center, even with a high power output, can go from barely able to get enough tokens and enough processing to start to transform, to now getting more than enough power, literally both tokens and efficiency, put into that particular facility.
And that's a paradigm change I think people are going to find out. And I think one of the things that was really interesting, he gave a contrast on how much investing is needed for traditional foundational models. I think the example he gave was 400 and something tokens for a basic request. They returned a wrong result versus a reasoning model, one of the more modern models that R1, for whatever reason, has made famous. 8,000 tokens. And again, to put this in context, the human, I think, reads at roughly about 100 tokens.
No, no. It's like 25. Yeah. So 25 tokens a second. So that's a lot of data. Oh, yeah, yeah, yeah. So one of the other things I wanted to kind of point on before we get too off into the AI world, I'm trying to wrap my head around another comment that he made. The one megawatt AI factory, data center, is a small AI factory. Now, both me and you have traditional data center history. A one megawatt data center used to be big.
There used to be a big old facility. Yeah. And it's going completely against the grain that everything is getting smaller and smaller. So multi-year investment to build a multi, one megawatt data center, now multi-megawatt data center, is a significant investment and takes quite a bit of time. But H100 is obsolete in 18 months. As people are thinking through their AI factory and keeping the lights on in other areas, how should they be thinking about consuming? Let's pick on your product, Kamawaza, the without TLDR.
You guys have an on-prem plate. In private cloud. How do you plan for that when I can't even keep up with the GPU pace? I think the market conflates training and inferencing a lot. Yes. And at one megawatt, the amount of inferencing, the world's largest company could automate probably all of its operations really easily. We're talking a set of four black wells, eight GPUs per server, four of them. You're probably in at what's called FP4. So the lower context and lower capability from a floating point that you can do on those NVIDIAs, you're probably talking somewhere in the north of a million tokens a second on a reasonable size model.
That is a tremendous amount of generation, of knowledge generation, knowledge processing, knowledge capability. And a company the size of like, I'm just throwing out there, FedEx of like 500,000 people, figure 100,000 people in that 500,000 are doing knowledge work. You're probably talking that level of knowledge work is probably that million tokens a second. But imagine 24-7, seven days a week, always running, never getting tired, like it's crazy. So that actually is super helpful context because a lot of these large SaaS providers are trying to sell me AI, not agents, but AI assistants, yesterday's technology for $35 a user.
So the math to get to transformation becomes actually pretty easy. And I made this argument yesterday that the math around AI inferencing and the investment in AI and the cycle, the refresh cycle, the depreciation schedule is pretty easy. That's pretty easy math. It is, especially, and sadly, but especially when you start mapping it, and it doesn't have to be directly human, what's a FTE mapping? And if you're average employee, FTE, 125,000, and an equivalent work unit is about 125,000 on the compute, you could say, all right, well, if it costs me a million dollars, that's eight FTE equivalent, is there an ROI on there?
And it does make sort of the math easy to process out. But you've touched on it multiple times. On the inferencing side, the feasibility, the output of this and the capability to get massive ROI is not a megawatt. It definitely starts at 10 kilowatts, maybe 50 kilowatts is a nice little sort of starting, still big, actually, back in the day, that would have been quite a few five kilowatt racks and 10 kilowatt racks. But nonetheless, the increment to start is reasonable.
Now, if you're an organization, and there are a few of them, but if you're an organization that needs training, that's a whole different story. And I mean, a megawatt data center, they're talking about 670 kilowatt racks. So nearly a megawatt of rack. Almost a megawatt of rack. We were at Super Compute in 24, and they were talking about the 500 kilowatt rack and the 600 kilowatt rack. It is here, especially with Blackwell. So as companies are thinking through kind of this plan obsolescence, how should they be thinking about the platform layer of this equation?
Like, how do I normalize the interface when the underlay is constantly changing? Yeah, I think that's the key. So obviously, that's our big bet. We think abstracting out the software, and in our case, we call it the orchestration layer, that all of these agents, all of these apps are actually going to plug into. They're going to plug in in three key ways. They're going to plug into an API, they're going to plug into an SDK from our programming, and they're going to plug into MCP, model content protocol, the ability to reach out and receive tool and function calls.
And when you can abstract that, and you can attach the data into that abstraction, and then you can attach the inferencing to accomplish that via that abstraction. So now those third-party apps and agents can just call in. You now have an orchestration layer at the software, and then the hardware can be swapped out. At the moment, absolutely, bang for buck, most tokens per second, NVIDIA has the market. And they're probably going to own the market on training for a while, but let's set that aside.
On inference, this is just amazing. But I wouldn't count out Intel, I wouldn't count out Qualcomm, I wouldn't count out AMD. We were both at an AMD conference yesterday. From a watt, from a thermal load, from a tokens per second, from a cost, as you get into those more reasonable 50-kilowatt, 100-kilowatt deployments for the large enterprise, there's going to be a lot of competitors. And you're going to want to be able to swap that out. And after you left, there was a few more, not even companies, hardware companies, people doing kernels, people doing compilers that are not CUDA-based.
And I think in the noise of all of GTC, NVIDIA's $3 trillion stock price, we lose sight that there are still plenty of companies doing inferencing on CPUs. Yeah, absolutely. It really does depend on the use case, and the importance is to develop the muscle to move processes to agents, and normalize that development. I'll break a little news right now. We just finished testing R1 on a Xeon 6, and it runs just fine. I've heard that just as long as there's enough RAM.
So you load up about a little over 500, so you can do it with 512 easy. If you get above 512 gigs of RAM, you can load it right in. We're getting something like 50 to 100 tokens per second. So that's a full reasoning model, full capability, processing at four to five times what a human can do at the edge on a standard Xeon 6. So you imagine the performance you get on a low-end GPU on a Jetson or something on one Blackwell.
Yeah, they have that. That's that new DGX station they just announced in the mini. So I'm super excited about the future. This podcast will probably be stale as soon as we release it. It probably will. Talk to me about what is your most exciting customer doing? My most exciting customer. So there's one that we both know about. They are transforming 30 or 40 enterprise data feeds. The biggest sort of highlight for a podcast. They're transforming 30 to 40 data feeds, and they moved the system of record.
The main system, these larger enterprises, especially manufacturing enterprises, utilize that base system of record that they have other apps on, that they have to get all their data and they have to get that data normalized so that it is the system of record. They've announced that they no longer want that system of record. That AI, since they're now having all their data feeds streamed right through AI, the need for that one system of record doesn't make any sense because as the data is going through the AI, if a CFO wants a report, they can generate the report on demand.
If the CFO wants a Monte Carlo simulation of cost going down, goods going this way, failure, they can do that absolutely on demand, and they don't need to first get it into the system of record to have another app and service to do that. And that's a game changer for sure. You know, my mind, you know, I'm a former SAP infrastructure architect guy, so my first mindset of like, oh, that system of record sounds a lot like SAP. Sure does, doesn't it?
And all of this transformation that was promised by HANA and that whole application suite sounds like a lot of SaaS companies are looking at agentic AI and trying to figure it out. And this is no secret. ServiceNow, Salesforce, everyone is going after it. Yeah, it's all those systems of records. Everyone is going after it. So it's not, I don't want to just pick on SAP. If folks want to find out more about Kamawaza, that stuff is in the link.
We'll try and put as much about GTC in the description below. com. Kamawaza has actually done an awful lot of Tech Field Day videos. I sure have. Matt is doing his magic of where he's, you know, running the latest model on CPUs and NVIDIA and Gaudi and all the good stuff. You can find that below. And full disclosure, you guys have partnered with us on a lot of testing. We really appreciate that. It's been a great partnership on the tests.
Talk to you next episode of the CTO Advisor Podcast. All right, sweet, man. Thank you. That was easy. That was fun.