AI's Future Unveiled: Nvidia's Blackwell, Planned Obsolescence
Transcript
All right. You can't tell it, but we're just a couple of blocks through GTC. Literally, Luke, and I've Luke, I've had you on a few times. You're now uh Kamawaza. First, two congratulations. First, congrats on the uh funding. That That was a big deal. That was a really big deal, actually. Uh Colorado, it's not the easiest, man. No, it's not. It's not a Colorado AI company. A Colorado AI company. People are looking at me sideways, let me tell you.
like the like the first thing like are you guys seriously in Colorado the part of the pitch deck like serious are you seriously in Colorado that's kind of how it started every time yeah so congrats on on the funny that is a big deal especially even with AI uh everyone has really great ideas I actually talked to a guy at the beyond cuda event yesterday that absolutely is following in your trails and his idea and uh the challenges so I just saw Matt and will probably post a picture somewhere here with his Apple uh uh goggles on.
Uh more than likely coding. Uh he was absolutely coding. Yes, more than likely coding. But the second piece of congratulations is playing obsolescence. You uh you got scared what about a year or so ago? tell you what you said. Nvidia is out just making sure that you cannot use the 100,000 GPU $100,000 GPUs you're buying today. You won't be able to use them a year a year year and a half from now. They had to do this.
They Yes. Thank you, by the way, on both of those. And if you just ran the math for the the TAM that they're filling into, the fact that they were a $3 trillion company made it so you weren't buying this on a normal depreciation cycle. You couldn't buy an H100 and expect five years, six years, seven years like you would try to ek out of a hard drive. To to have that level of volume, to have that level of revenue, they have to force their customers to buy typically it looks like a 12-mon cycle, 18 months on the long end, the most expensive capex ever in it.
And they're doing it and everybody screwed me. And I mean everybody like uh um I think that one post on LinkedIn got to almost 400,000 impressions. like it was my one of my whitest pushed out ones ever. And it's because everybody's thinking it's just a hardware game. It's a hardware matched to a software. ARM processors, unlike x86 processors, have a reductive set. So from one generation to the next, if you want a feature to be added into it, you have to have a net new chip to have that net new software they've written now be able to access it.
That means the previous version literally can't utilize that software for that new feature. and they're marching that obsolescence. And today today he he literally said if you have hopper chips the H100s you cannot give them away. Give them away. Once Blackwell is fully up to value. Yeah. He said 40x I think was the performance. He said uh 40 for per watt uh was 40x the performance. So, you know, me and you've gone back and forth on this quite a bit.
Uh, because H100s are the workhorse. They literally are. People and and this I'm I'm having trouble wrapping my head around just how much better Blackwell is because people are doing amazing things with H100's steel and to get 40x the performance. I think one of the first questions is is the demand there to get to to fill up all of those tokens. I see it as um I I see it as paramount. The H series was so powerful both from uh its token throughput but its thermal throughput.
It it is literally a power hog from a heat and dissipation perspective. If we're the same wattage, you could get 40x more power. Means a typical enterprise with a typical data center, even with a high power output, can go from barely able to get enough tokens and enough processing to start to transform to now getting more than enough power, literally both tokens and efficiency put into that particular facility. And that's a paradigm change I think people are going to find out. And I think one of the things that was really interesting, he gave a contrast on how much uh interesting is needed for traditional foundational models.
You know, I think the example he gave was like 400 and something tokens for basic request that returned a wrong result versus a reasoning model, one of the more modern models that you know, R1 for whatever reason has made famous. uh 8,000 tokens and uh you know and we to again to put this in context uh uh the a human I think reads at a at a roughly about a 100 tokens. No, it's like 25. Yeah. So 25 tokens uh a uh uh a second.
So that is quite the bit. That's a lot of data. Oh yeah. Yeah. Yeah. So, one of the the the other thing I wanted to kind of point on before we get too off into the AI world, I'm trying to wrap my head around another comment that he made. The one megawatt AI factory, data center is a small AI factory. Yeah. The now both of both me and you have traditional data center uh uh history. one megawatt a one megawatt data center used to be big.
That used to be a big older facility. Yeah. And uh it's going completely against the grain that everything is getting smaller and smaller. So multi-year investment to build to build a multi- one megawatt data center now multi-megawatt data center is a significant investment and takes quite a bit of time. Uh but H100's obsolete in 18 in 18 months. Yep. as people are thinking through their AI factory and keeping the lights on in other areas, how should they be thinking about consuming, you know, let's let's let's pick on your product Kamawaza the the without you know TL TLDDR you guys have a onrem plate how how in private cloud and how do you how do you plan for that when I can't even you know keep up with the GP PU play uh piece I I think um I think the market conflates training and inferencing a lot.
Yes. And at one megawatt the amount of inferencing the world's largest company could automate probably all of its operations like really easily. I mean, we're talking um we're talking a a a a set of four black wells, eight eight eight GPUs per server, four of them, you're probably in at at what's called FP4, so the lower context uh at lower capability um from a floating point that you can do on those NVIDIA, you're probably talking somewhere in the north of a million tokens a second on a on a reasonable size model.
That is a tremendous amount of generation of knowledge generation, knowledge processing, knowledge capability. And you know, a company the size of like I'm just throwing out there FedEx of like 500,000 people. Figure 100,000 people in that 500,000 are doing knowledge work. You're probably talking that level of knowledge work is probably that million tokens a second. But imagine 247, seven days a week, always running, never getting tired. Like it's crazy. So that that that actually is super helpful context because a lot of these large SAS providers are trying to sell me uh AI not agents but AI assistance yesterday's technology.
Yeah. Literally uh for $35 a user. So the math to get to transformation becomes actually pretty easy. And I I made this argument yesterday that the math around AI inferencing and the investment in AI and the cycle the refresh cycle the depreciation schedule is pretty easy to that's that's pretty easy math. It it is especially and sadly but especially when you start mapping it to and it doesn't have to be directly human but what's a FTE mapping and if you're average employee FTE1 125,000 and an equivalent work unit is about 125,000 on a compute you could say all right well if it cost me a million dollars you know that's eight FTE equivalent is there an ROI on there right uh and it does make sort of the math easy to process out but you you you've touched on it multiple times on the inferencing side the feasibility the output of this and the capability to get massive ROI is not a megawatt.
It definitely starts at 10 kilowatts maybe 50 kilowatts is a nice little sort of starting still big actually back in the day that would have been you know quite a few 5 kilowatt racks or 10 kilowatt racks. Um but but but nonetheless the the increment to start is reasonable. Um, now if you're an organization, and there are a few of them, but if you're an organization that needs training, that's a whole different story. And I mean a megawatt data center, they're talking about 670 kilowatt racks.
So, so nearly a megawatt of rack. Almost megawatt of rack wreck. Yeah, that is uh we we're at super computing 24 and they were talking about, you know, the 500 uh me uh let me call uh kilowatt rack and the 600 kilowatt rack. it is here uh especially with Blackwell. So as companies are thinking through kind of this this plan obsoles lessons how should they be thinking about the platform layer of this equation like how do I normalize the interface when the underlay is constantly changing?
Yeah, I think that's the key. So u obviously that's our big bet. We think abstracting out the software and in our case we call it the orchestration layer that all of these agents all of these apps are actually going to plug into. They're going to plug in in three key ways. They're going to plug into an API. They're going to plug into an SDK from a programming and they're going to plug into MCP model context protocols the ability to reach out and receive tool and function calls.
And when you can abstract that and you can attach the data into that abstraction and then you can attach the inferencing to accomplish that via that abstraction. So now those third party apps and agents can just call in you. You now have uh an orchestration layer at the software and then the hardware can be swapped out. At the moment absolutely bang forbuck most tokens per second Nvidia has the market uh and they're probably going to own the market on training for a while but let's set that aside on inference.
This is just amazing. But I wouldn't count out Intel. I wouldn't count out Qualcomm. I wouldn't count out AMD. We're both at an AMD conference yesterday. Uh from a watt, from a thermal load, from a tokens per second from a cost. As you get into those more reasonable 50 kilowatt, 100 kilowatt deployments for the large enterprise, there going to be a lot of competitors and you're going to want to be able to swap that out. the after you left there talked through there was a few more uh not even companies uh doing hardware companies people doing kernels people doing compilers that are not CUDA based and I think in the noise of all of GTC Nvidia's $3 trillion stock price we lose sight that there are Still plenty of companies doing emphas on CPUs.
Yeah, absolutely. Like it really does depend on the use case and the importance is to develop the muscle to move processes to agents and normalize that development. Uh I I'll break a little news right now. We just finished testing R1 on a Zeon 6 and it runs just fine. I've heard that. Just as long as there's enough RAM. So you load up about a little over 500. So you can do it with 512 easy. If you get above 512 gigs of RAM, you can load it right in.
We're getting something like 50 to 100 tokens per second. So that's a full reasoning model, full capability, processing at four to five times what a human can do at the edge on a standard Zeon 6. So you imagine the performance you get on a low-end GPU on on a on a Jetson or something on on one Blackwell. Yeah. Uh they have that. That's that new DGX station they just announced in the Mini. Yeah. So, I'm I'm super excited about the future.
The This podcast will probably be stale as soon as we re uh release it. It probably will. What are Talk to me about what what is your most exciting customer doing? Um my most exciting customer. So, uh there's one that we both know about. They are transforming 30 or 40 enterprise data feeds. uh the the the biggest sort of a a highlight for a podcast. They're transforming 30 to 40 data feeds and they've moved the system of record the the main system these larger enterprise especially manufacturing enterprises utilize that base system of record that they have other apps on that they have to get all their data and they have to to get that data norm normalified so that it is the system of record.
They've announced that they no longer want that system of record. That AI since they're now having all their data feeds stream right through AI, the need for that one system of record doesn't make any sense because as the data is going through the AI, if a CFO wants a report, they can generate the report on demand. If the CFO wants a Monte Carlo simulation of uh uh uh cost going down, goods going this way, failure, they can do that absolutely