17x Faster AI Inference? Kamiwaza Explains the HPE + NVIDIA Demo

6:41 · Watch on YouTube ↗

Transcript 1,259 words · about 8 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

All right, we're at HPE Discover and I am with my good friend Matt Wallace from Kamiwaza. Matt, good to good to see you again. >> Good to see you, too. Strange places, good times. >> You know what? Uh we had planned to get together anyway, but I was watching Fidelma's uh presentation for the CTO keynote day three of HPE Discover and I saw a familiar demo in the background. >> Yeah. " >> Yeah, it was it was easy to miss uh who was behind the curtain on that one a little bit.

>> So, what was behind the curtain? What what what? Give us the the the hero numbers because what you guys did with the HPE storage team the uh in a combination of this is sponsored by NVIDIA. So, NVIDIA compute in the background, but talk about what the hero numbers were. >> Yeah, I mean, the the big ones are total throughput on inference and then the time to first token speed, right? The amount of time between when you send the prompt and you get an answer.

And uh 17x and 20x and I forget which one is paired with which, but they're both, you know, incredible numbers. And that was something we got leveraging the combination of um HPE's X1000 um all flash object array along with NVIDIA and their GPUDirect capability, right? Which is already in May letting you move uh you know, bits of data directly in and out of the GPU at super fast speed. >> You guys have worked a lot with HPE, other vendors, uh with your platform to do a lot of AI agentic AI testing.

>> Yeah. >> Connect the dots for me. >> Yeah. Well, I mean, you actually get to see it. So, when you saw that grid of agents and you were seeing that kind of fintech uh you know, example where you see certain number of agent conversations, right? And it's trickling along. And on the side that's enabled by these technologies like Dex 10K, you see literally hundreds of agent conversations finishing, right? And it's those are real you know, they're real use cases.

So, what's driving that? Um you know, it it's certainly this integration where Kamiwaza, you know, orchestrates this you know, we have a superpowered uh high-performance middleware we wrote that sits between the X 10,000 and its API for object and then the GPU direct API on Nvidia. So, we can marshal those that conversational data, right? They call it KV cache. Right. It's everything the GPU or the AI remembers about the conversation you had so far. But people probably notice if you're talking with somebody like a chat you could hear a Claude and it seems fast and you walk away and come back to that conversation the next day, the first thing you say in a long conversation seems pretty slow until it kicks off.

It's warming back up cuz it has to redo the math to compute all of that previous conversation. That literally translates into you know, lost time and then lost GPUs, right? And so, you think about what's going on in the GPU market today, right? You just can't get that. There's too much of a shortage. And so, anything that saves you dollars on that uh you know, GPU spend, anything that saves your users time and gives you much more throughput is hugely valuable.

Now, do you want me to let you in on a little secret on why this matters more now? Because this is actually really timely. >> Right. >> So, if you go back in time like a year, right? These long conversations, if you were taking any open model that you would run in your data center, they just were not as good at dealing with that long context. The Frontier Labs had a really big advantage. Today now though, the open models have gotten much much better at dealing with that long conversation.

That means that this ability to take that KV cash and move it in and out and restore it and have those long conversations has become much more valuable cuz it's realistic. Developers aren't having to take those conversations and break it down into tight little chunks to be coherent anymore. And so now you have developers who are feeding 100,000 tokens of code in and they can actually expect it to reason and give them a real answer, right? Or you have people who are having this long conversation something like an open claw agent, right?

And they come back to say something new to it. Do they want to wait 2 minutes for it to respond? They do not, right? And this solves that problem. That's your 17x, your 20x. Let's say you have a generous memory pool and you can hold 20 long conversations or 50 long conversations. You're You're paying a lot for that, you know, GPU VRAM, right? But if you're serving hundreds or thousands of users, there's zero chance, right, that somebody's conversation is going to be there 2 hours or 4 hours later.

Look at what happens with the frontier providers, right? OpenAI, you know, Claude. You have You're going to pay money for that cash or it's going to be limited to 5 minutes or less. Anthropic, if you pay a lot extra, you can get to hold it for an hour. That's the absolute upper bound until you have to refresh it, right? That's pretty realistic, right? 2 hours, 4 hours, 8 hours. But this But this means is you can come back to the conversation the next day and instead of waiting 40 seconds or 60 seconds for that to kick off, you're waiting a second or two.

>> And those numbers may not seem like a lot, but scale breaks everything. And especially as we're thinking about uh agentic AI where this communication is going back and forth and we're scaling the conversations, uh the that 40-second um uh the that 40-second hydration time >> Yep. >> matters at scale cuz if we're talking about conversation after conversation, prompt after prompt, and we're feeding the data back and forth in a real workflow, that's where we start to see the uh not just the slow down, but the increased cost with trying to combat the problem with more RAM.

>> Yeah, you got it cuz you just hit it. There's two two aspects here. On one side, you have the user experience, right? Waiting 40 seconds, no one wants to do that. >> Right. >> But the flip side of this is it's not just about that. That 40 seconds isn't We're not just waiting in the ether for something to happen. You're grinding up that GPU to to get back to that state. And don't you have better things for your GPUs to do?

Like they're not cheap. >> So Matt, I really appreciate you taking time out of your busy schedule because you guys were at the keynote this morning. At least if if they didn't mention you, I wanted to to at least uh unveil the curtain a little bit. >> I appreciate that. >> You're having plenty of conversation with customers. I'll let you go back to uh kind of doing your day job. Thanks for uh joining the CTO Advisor yet again. >> Always a pleasure.