Your GPUs Aren't Slow. Your Data Path Is
Transcript
All right, last video in our three-part sponsored series by NVIDIA and HPE to take a look back at HPE Discover and what did I find interesting? First video recap, just to overview the special relationship between HPE and NVIDIA and AI factories and how the AI factory approach that HPE takes is based around their Unleash AI and private cloud AI stack and how that partnership and series of partnerships is differentiated between HPE's competitors. Video number two, we talked about the importance of CPU to agentic AI, making sure that the AI is recommending and some type of deterministic process is approving or using judgment.
Today, we're going to take a final look at IO. I hinted to video number one of how I had my mind changed on the importance of specifically networking and inferencing. We know the importance in training, but video two, we talked about how GPU utilization in agentic workflows isn't the bottleneck. The bottleneck is getting data in and out of the GPU. Specifically, what's unique about inferencing versus training is agentic AI. Our ability to do these turns around the data. There's a process in which a decision or data is collected, a prompt is sent to the GPU and then the GPU turns on it, sends a response.
This creates a KV cache bottleneck. The amount of RAM that we need to have in a generative AI and V systems explode versus regular everything. So, we know the KB cash problem. How is HPE and Nvidia solving this problem together? The obvious turn is to increase network bandwidth and reduce latency. The lower the latency and the higher the bandwidth, the the more KB cash will have at the storage layer. However, there's still the gap of protocols and being able to get that data in and out of the storage system as quickly as possible.
HPE demoed on stage the second day of the conference a combination of technologies that they partnered with with Kamizawa from the AI platform perspective validated by Signal 65 this up to 20x improvement in KB cash efficiency and performance. That means for every one factor you go up in capability or performance, that's how much less GPU you need. What's the unique HPE sauce? The unique HPE sauce is the Alletra 10,000 I mean X10000 and the combination of these protocol improvements that Kamizawa and Signal 65 demonstrated that enables this performance increase in KB cash efficiency in use.
This is unique to HPE and Nvidia. Again, all AI factories are not created equally. These systems and these OEM relationships are critical as you go through and meet with your account teams and talk through your unique challenges and your applications, you really do have to dig deep into the technologies and the requirements. You want to learn more? com where I put some of these technologies to test. I haven't gotten specifically the HPE technologies in the lab, but you'll see the patterns and be able to understand how you yourself can run these POCs.
Until then, talk to you next time on the CTO Advisor CTO Dose.