AI Meets High-Performance Computing (HPC)
Guy Currier, the Futurum Group Contributor, joins the podcast from Super Compute 23. Keith and Guy discuss the news out of Super Computer 2023 and how the intersection of AI infrastructure meets and relates to HPC infrastructure.
Transcript
Hey, how's it going? It's Keith Townsend, Principal of the CTO Advisor. I'm joined with a new guest. New to me in a couple of ways. One, new to the podcast and new coworker, Guy Currier, a contributor here at Futurum Group. Guy, welcome to the podcast. Hi, thanks, Keith. It's great to be here. So we're here at Super Compute 23. This conference has been going on for 35 years. This is my first Super Compute. What about you?
Third. Third. And when was the last time you came to the show? Well, I was here last year. And then before that, it was maybe six years ago. So you missed the big disappointing show in St. Louis, I think it was? Yeah. But other than that, this is like 15,000, 16,000 people here somewhere. I haven't gotten the official numbers, but it's a nice crowd. Yeah, they were telling us this might be the biggest ever. And really, it's no surprise because, you know, everything all over is all AI right now.
And AI systems are supercomputing systems. So let's talk about this. You know, let's just jump into the deep end. Obviously, we're in the beginning of the AI hype cycle. Everything has AI in it. And I don't know if supercompute has ever been considered AI. So machine language, definitely. Machine learning, definitely. A lot of scientific compute, a lot of use of GPUs, nothing new to this group. AI, you know, we kind of pooh-poohed on the term AI years previous because AI meant something specifically versus machine learning.
So, you know, you're a veteran of this show. How do you think the AI conversation overlaps with the supercompute conversation? Well, so I've been doing HPC, high-performance computing, for quite a long time. And AI first started really cropping up seriously within HPC discussions three or four years ago. Because efficiency is really important in HPC, meaning efficient use of the systems because they are so much in demand. And the performance needs so far outstrip even the very advanced and advancing capabilities we have today, right?
So AI, not generative AI, mind you, AI, just speaking really broadly, has been one way to start to more efficiently queue up jobs, for example. And to more efficiently run and rerun jobs with differing parameters to get to results more quickly. Now that said, those are what you might call minor subsystems of the overall. They might not even run on the particular supercomputing cluster we're talking about, right? But that has, like, it's such an academic field, HPC and supercomputing is so academic.
So naturally, there have been a great deal of studies and research done just around linking AI to HPC. So the thing is now, of course, like you said, the HPC world and supercomputing world is very familiar with the use of accelerators, data center GPUs, and very advanced high-speed, low-latency computing infrastructure. And a lot of those elements are what you need for successful AI. We've all known this, but with Chats GPT and generative AI, all of a sudden, everybody wants to get in on this and understand it better.
The key thing for me, and I think for everybody to understand, is the critical differences in what an AI supercluster or cluster, sorry, should look like compared to what an HPC one is. But those differences are your second step. Your first step is understanding that getting great performance, incredible performance, game-changing performance out of infrastructure is something these folks have been doing for a really long time. Yeah, and it kind of became real for me the first day on Monday. I got in a little bit early.
The conference has been running since Sunday. This is Tuesday afternoon. And I got in Sunday, I mean, Monday morning, and I went to a session about best practices for HPC in the public cloud put on by AWS. I was blown away by the velocity of the course. It's eight hours and you're trying to get in best practices of HPC in cloud. Two things when you put them together, you know, are kind of life, life, career long learning pursuits in themselves.
So you're mashing them together. We went through kind of the fundamentals of AWS all the way to creating a swarm cluster within AWS, like an hour into the presentation. So, you know, it's it's an amazing velocity, but I think it's representative of what CTOs, CIOs and their teams are experiencing as the business is waking up to the possibilities of AI. And they're asked to implement these highly complicated systems, whether, you know, you're talking about on premises and a shared cold facility, a SAS provider or even in the public cloud.
This is not these are not simple systems, especially given that most of us don't even know what we want out of them. Yeah, no, exactly. And by the way, I'm really jealous that you got to attend that session because that's where our visions and the visionary nature of the aspirations that we have for AI and for really the use of supercomputing. That's where it collides with reality. Right. Understanding the details and you get caught up. The devil really is in the details here, really is in the details.
There are so many bottlenecks. There are so many potential sources of especially when you're doing cloud based deployments, cloud and HPC, they don't mix really well. Right. When you mix them properly, you got a delicious salad dressing just like oil. Right. But on their own, it just sort of they don't necessarily combine because of the high or I should say low tolerance that whether it's an AI training system or implementation or, you know, a simulation of some kind, the low tolerance that they have for hiccups, difficulties, latencies, missing data, or even for security and access concerns on the other end.
Right. I think that that what, you know, I would look at, you know, I have a sort of quasi CTO role looking at the technology that we're using at Futurum Group and my part of Futurum Group, which is called Visible Impact. I think that that for this from the CTO standpoint, you have to understand a lot about the lifecycle of these applications. You can get immediate benefit from cloud based HPC or AI or I wouldn't say immediate, but very quick benefit. But that just like everything else, how tailored is that going to be to your particular mission, to your organization, to your goals?
As you move in your journey towards adoption, you're going to progress further and further in developing that lifecycle to suit your own particular needs. Yeah. As I as I delved into many of the sessions here, much of the low level technology, I found it extremely important, you know, just fundamentals of engineering gather the requirements first. And we're not talking about performance or maxims. Just basically, what are you looking to do? I remember what was AWS reinvented circa 2017 when they announced their deep vision and it was hot dog, not hot dog.
Yeah. It showed the power of AI and machine learning. And now that we have generative AI. That's a completely different model in a set of capabilities and a different infrastructure and a different it's sort of a different training cycle of its own. Those who've delved a little bit into MLOps, you know, machine learning MLOps sort of corollary to DevOps, you can see how different you actually have to production flows in AI and ML. So these paradigms are different. And your starting point, like you said, what is it that you're trying to accomplish?
What is it that you're trying to do? Doesn't just define how you take that journey, but which pathways you're going to choose, which ones are likely to be more productive. There's so much unknown right now. And I should add one other thing, this sort of cloud paradigm, sort of matching with this sort of agile paradigm and all these sort of things about fast failure and trying things and minimum viable, you know, product and that sort of stuff can apply just as well in the AI world.
It really doesn't apply in the HPC world. And so that's another case in which we want to apply the learnings of HPC. But we need to bring, you know, our let's say our more creative, HPC is creative, but our more expansive enterprise business orientation, even in the academic world. Yeah, I've run HPC in the enterprise before, and thou shall not mix data center operations with HPC. They're not the same. The requirements are not the same. The operations are not the same.
The security profiles are not the same. It's just the overall concerns are not the same. And as we're approaching AI projects and AI infrastructure, it's the same challenges or set of differing challenges. You know, if you need to do some facial recognition, you may not need to retrain or create a whole other train. You might be able to find a SAS type solution that will do that for you versus net new things that you need to, which will move to the next phase of our conversation.
The next phase of our conversation, using some of these high end GPU based systems. I've heard a new term here that I had not heard for accelerated computing or accelerated focused computing. The idea that we're using primarily accelerated computing next to CPUs or GPUs, basically, essentially next to CPUs where the primary compute is the GPUs. Now we're seeing these all in one systems and they kind of overtake our traditional view of a server. You know, NVIDIA just announced the H200. You know, that's the big boy.
AMD has the 300 series. Intel has their GPU. These are data center GPU accelerators. Yeah, GPU data center GPU accelerators. Is NVIDIA the only game in town or are we, you know, help me put some context around this? Well, let me tell you, one of the largest, and I'm embarrassed I don't remember the name of it, but one of the largest supercomputers in the world right now is actually 100% Intel. Intel has. That was the Aurora, the two.
Aurora, the Aurora. Thank you. I just heard that word five minutes ago and I forgot it already. So, yeah, the Aurora. NVIDIA is by far, I mean, I would say NVIDIA's dominance of this industry right now is greater than Intel's was in CPU, in X86, well in CPUs generally, let's say 20 years ago. They are not the only game in town, but they are right now the sun, moon and stars of accelerators. You mentioned accelerated, accelerated computing, focused computing.
I think you might also be talking about GPGPU, which is general purpose computing on GPUs. NVIDIA, this is something NVIDIA has pioneered and it's a way to a way to process to host workloads using primarily GPU systems and not CPUs so much anymore. And as you know, one of the trends lately in these architectures is to use the CPUs almost as a traffic cop or a manager or a supervising chip where most of the work is done by the GPUs. And the CPUs are really handling in and out and workflows and like all that other stuff.
I mean, the essential difference between or the essential addition that GPUs provide to a system is to be able to perform computations or processing in a massively parallel form and a massively repetitive form. So, you know, typical computing for general purpose use does not lend itself to that. But when you're talking about operating neural nets or or processing graphics, computer vision, image recognition, a great deal, but not all. By far from all of high performance computing applications where you're doing simulations, modeling, forecasting, that sort of thing.
Many of the calculations, however complex they may be, however large the data sets are, high precision, the measurements being used, they are highly repetitive and can be highly parallelized. And that's where something like a GPU can be helpful. So I'll put some context around this. I had a. Great conversation with a buddy of mine, Ray Lucchi, who hosts the great beers and podcasts, I mean, great beers with storage podcast alongside me. A couple of folks has been researching AI since like nineteen seventy six.
He told me and he recently took a course. And at the end of the course, they had a hackathon and the hackathon was to take a to develop a. Machine learning algorithm for pong to play, teach a computer to play pong. So for basically four lines of code or algorithm, basically, if if the dot comes at the top of the screen, move the paddle up middle screen, move it up a quarter, you know, all the way down to the fourth. Those were the four lines of logic.
So this repetitive process. The machines will play a game, they will play five hundred games. The neural net would then learn off the five hundred games. And he will run the the entire learning would be two thousand of these packages. So you're talking about five hundred to the two thousandth degree of repetitive processes. And this is that that that process you're talking about on a video that he said, 30, 70 or 30, 60. It takes days to run. You add more GPUs, you add more coup de corps, you reduce that time.
And now you take this very, very simple with just basically four parameters algorithm. It takes days to run on a fairly beefy desktop system with accelerated compute. My mind starts to break was what happens if you add additional parameters? Like how much more difficult that problem becomes for a CPU to solve versus a GPU. If the instruction set or the sorry, the the the computations required fall within the capabilities of the GPU. It's more than just add plus, you know, multiply, divide, but not honestly not a hell of a lot more.
But it does it very, very well. It does it very, very fast and it does it on a mass basis. And furthermore, the software and the ability to code these systems is extremely mature at this point. So there are lots of what you might call tricks to get more out of the same number of compute units. When you're thinking GPUs, talent and energy that used to go into CPUs and a more larger set of calculations. Now we're going to the smaller one to get more out of the same amount or exponentially more out of a larger amount.
And they actually talked about that when they talked about the supercompute 500 green list, which was that we're getting to the end of Moore's law. We're getting to the end of getting efficiency out of hardware directly. And we're getting to a point where the efficiencies have to come through better code, more efficient code. Well, yeah, I'm sorry. I've been doing this for a while, doing this business for a while. And it seems like we're on we're on a three to five year cycle where everything's going to be solved by software and then everything's going to be solved by hardware.
We are definitely in a software development phase. The market has been pushed very hard towards its limitations, its physical limitations with current technology and current capabilities. Not that I mean, you know, I mentioned NVIDIA AMD itself has done remarkable things with its micro architecture to bring memory processing and that CPU orchestration component or processing component as processing to. There's a lot of HPC that can't really use a GPU because of the complexity of the calculations required and the interdependency of them. But in any case, it's certainly there is a lot of fruit to be found in these kind of tricks.
There's a lot being done right now with data structures. A lot of people may not know AI uses much smaller data structures like much lower precision, four bits worth of data at a time versus 64 bits worth of data, so to speak. Numbers that are encoded at a lower precision. And there's a lot of creativity around that right now for AI purposes, all to get more out of the same amount of silica. Yeah, I was quite amazed. And I think I referenced it out if I had not put it in the show notes.
Basically, the basic concepts of LLM and training of LLM extremely the base concept is extremely simple. Right. And I forget who it was I was having a conversation with and I have to check the data. But they were saying the overall data set that a typical LLM is trained on is like 80 gigs. But the the repetitiveness and the predictions that 80 gigs might as well be and, you know, kind of translating into the data center world. So several petabytes of of traditional data center.
Analytics and searches and capability because of the repetitiveness and the complexity, the resulting complexity of the of the jobs we're running on GPUs. I'm having a lovely time coming up to speed on AI. I still feel like a kindergarten nerd and an advanced algebra course. And when I'm coming to these events, I love it. Guy that folks are like you are here to break down these topics for me. Do I know your own future group just general as a contributor?
Can people find you anywhere other than a future group? No, that's that's the place to find me right now. OK, so future group dot com. You can look in the analyst directory. You'll see me there. You can look at in our insight section to see some of the articles I've been publishing there. And hopefully you can find me on a few more podcasts, maybe a video or two now and then. Yeah, we'll make sure to have you. I did not bring the camera to this event, but I'll make sure to have you on a CTO advisor webinar soon as we talk through not just AI.
You have a pretty good knowledge across the spectrum of enterprise technology. I'm looking forward to learning more from. Thanks, I'm looking forward to having those discussions with you. You will learn more about the CTO advisor. You can find us on the web, the CTO advisor dot com. We're not quite integrated into the future groups website. You are not going to find me under insights and analysts yet on the future group site. You can still find all of my musings on the CTO advisor dot com.
You're going to find out you want to engage with me. I am still on X dot com at CTO advisor and of course, always on LinkedIn. Talk to you next CTO advisor podcast.