Unlocking AI Accelerators: How CloudFlare Optimizes AI Workloads & Saves Power
Transcript
(dramatic music) >> All right, the smile on my face is not because of the horrible job I did of drawing the, it kind of looks like an Airstream. It is because we've been talking about industry challenges, family, and Rebecca Weekly, VP of cloud infrastructure? >> Hardware systems engineering. >> Hardware systems engineering. >> In the Infrastructure group at CloudFlare. >> At CloudFlare, one, I don't know how she finds time to do things like this, so I appreciate that. You know, she's in a band.
>> It's for you. >> She does a lot of stuff with the Open Hardware Community, just fabulous resource. We're going to talk accelerators, and one of the use cases that we've been hearing a lot has been generative AI, large language models. I think there's a lot of confusion. Because when I purview the market and what customers are actually doing, very few customers actually need GPU accelerators for the type of AI learning, or not learning, AI that they're doing. Can you give us kind of a overview of accelerators and the use case for accelerators and kind of the pros and cons?
>> Oh my golly. Okay how long do we have? So when I think about accelerators, you know, you really have to look at the fundamental architecture of a general purpose processor and what we can attach to it. So over, I would argue the last 60 years, we've seen repeatedly a dichotomy between what we put next to a CPU, a co-processor model, and what we integrate into a CPU to actually improve performance on specific workloads. So a classic example is floating point math, right?
We used to not have that in the X86 architecture. We used to have a co-processor. My very first project, not to date myself, out of college was working on an SA brick which was an FPGA based acceleration brick on the NUMAlink architecture for SGI. And that was exactly because there was a series of, particularly for technical computing problems that could not be solved in a reasonable timeframe, like I would've had grand babies before the CP was done with it. And so we used an accelerated environment to be able to actually process that.
Fast forward 20 years later, and you will find we're in this AI cycle. This is not a new thing. AI has been with us for a very long time. What changed? Why are we in this current AI hype cycle? And it's exactly what you said, right? We're in this generative AI model, where what is happening with artificial intelligence is actually something that's generalized enough that it's changing how we look at the paradigm of interaction with computers, right? People are just talking to devices and having real life interactive answering of questions, and we can, you know debate and discuss whether that's a good thing for us as a society, if it's actually helping us question ourselves.
You know, there's lots of great commentary on that that we will not discuss here. But from an architecture perspective, the challenge of these models is that they're training on trillions of tokens, right? They're huge models. When a new ChatGPT 3-4 version drops every six months, that's the training timeline for these massive trillions of parameters that are being trained from for these 8x effects. But that's not actually how most people use AI. So most of us are never going to train ChatGPT type models.
>> I can't think of the number of billion for just long have a billion parameters. >> Yes. >> Right? >> Yes, it's just that, you know, again, back to the grandbabies comment, I hope I will have grandbabies by the time anything I could actually build or anything that could fit on your Airstream could actually generate this. But that's not what most of us are doing with AI. Most of us are using pre-trained models and integrating them for predictive analytics for typical approaches to say, is that a cat?
Is that in my case, a threat on my network? Is that something, is someone trying to DDoS these people? If that's the case, I'm using a trained model and if I start to observe drift, I may retrain a spot train scenario of that model, but I'm not going to go back to the beginning. And so it's really inferring, right? It's using the model on a new dataset, on a unique situation to determine is this what I think it is or not?
That is a very different workflow and product and even data model than training. And so I always, whenever we start the AI conversation I always want to separate, are we talking about training, which is a very different problem. This is a power, you know, how are we going to actually make sure we have enough data centers that are 100-200 megawatt campuses like it is what we talked about earlier, nuclear everything that we need to think about in that domain. >> So with this power concern, why drew the Airstream?
My wife and I, Melissa, we went on a 3000-mile trip a couple of years ago where we took the Airstream literally around the country, and Melissa came up with this wonderful idea, as executives normally do. , you know, we could put a camera on the back or we can use the builtin camera to collect imagery as we go through the country. So, you know, cat, not cat, you know how many dogs do we see inside of someone's car? Whatever clever piece of technology we're collecting that data, we're inputting that at that data.
But a serious problem is, you know, a typical data center rack is 8K. I don't know eight kilowatts of power. I don't know how I'm going to get that size of rack into my data, let alone if we're thinking about like large language models and the types of CPUs and computers and accelerators needed to run that. 7 to 10 kilowatts in power consumption constantly in that. So obviously this is not going to work. You know, that might be this much of a rack.
>> One four used- >> One four used and I can, in a top of rack switch. That's not going to job in my RV. It's not going to do the job in the typical data center. >> Yeah in fact your typical data center is designed totally and efficiently for that. >> Exactly. >> And they rebuild every 10 years maybe. These are buildings, it's not like, you know we complain if we're not getting a doubling of compute power every 18 months, Moore's law.
But when you talk about building facilities, this is a 10, 20 year cycle. >> So I think what Intel wanted us to talk about in sponsoring this content was when is it appropriate to kind of use PCIe offload versus one of their other competitors built in on CPU accelerators. And I thought, why not talk to CloudFlare, who obviously has all kind of new cases, whether they're you know, it's internal operations and tools or simply network acceleration. There are so many different instances, where you can save power requirement, save space, save cooling by simply using what's in your existing processor.
>> Yeah, so when I think about CloudFlare, right? We have 700 physical sites that we operate, over 700 physical sites, in 300 different cities over a hundred countries directly peered with over 12,000 networks. What I can find in Hawaii isn't going to match what I can find in Singapore, isn't going to match what I can find in Nevada, isn't going to match what I can find in India, right? All of that spread means I've got data centers where I have half of a four kilowatt rack.
Eight would be- delightful. >> A dream. >> Or I maybe have 15 kilowatts. You're not going to find 25, 40 kilowatts. I mean Google Next is happening right now. They're talking about their TPU versions. These are 45, 50 kilowatt racks that they're working on. If you own and operate your own data center and you're building it from the ground up, you can make a lot of these decisions. You're thinking about the utilities footprint that you need to associate yourself with.
If you're co-located, because you care about being 50 milliseconds from any user anywhere in the world, you've got a really different problem set, right? We've got a very different spread for power for network capabilities, for availability. So we're big on CPUs. CPUs are everywhere. They're everywhere in our network. They are our network. Fundamentally, the cloud is just somebody else's server and our network is absolutely a vastly interconnected domain of servers distributed all over the world. So for me, it's a massive architectural shift to take something off of CPU and to put it onto an accelerator.
So where do we use these things? It's where it makes sense. So the first domain where we've really looked at acceleration is within the NIC, right? So if we're running quick protocol every single time, we want to see good throughput and capabilities for that on our network interface card. It is in line to the traffic itself. I could do this processing on the CPU or I could do it in line to the network interface card and ensure that it's running that much more quickly.
So this is, no that was not a pun, although it was kind of pun. It was not intended to be a pun, you know? So this was the first domain where we really started to see enough A, standardization in network interface cards, enough standardization in the network protocols that you could get consistent acceleration leveraging an external accelerator. This has not been true in what I would argue the application acceleration space. You know, the big one that we all like to talk about are GPUs.
GPUs are the most common and certainly have been around a long time. My first company was Silicon Graphics also was responsible for doing graphics acceleration. And this is a domain where we, I would argue, first sought acceleration in the ecosystem because graphics tons of vector mathematics, which CPUs were absolutely not good at, and a deep parallel pipeline of pixels that you're operating against. So an alternate architecture emerged that was optimized for that kind of vector processing at scale. The fact that we now can use general purpose GPUs beyond graphics in the domain of vector processing for you know, what is effectively neural network mathematics, right?
That is the transition that has happened in this domain. It's very exciting. I would argue it's incredibly exciting what we've been able to do with this exponential increase in compute capacity. But the challenge is it's sitting on an external bus. It has different memory attached to it, usually significantly less memory 'cause it's optimized for bandwidth versus being optimized for capacity. So, you know, the smallest GPU is out there, an L4 GPU right? It has 24 gigabytes of GDDR memory versus like- >> Couple of terabytes or four terabytes of memory in a CPU.
>> Exactly. >> That's a big difference. >> So if you have a model with trillions of parameters, you're not going to be fitting that into... >> A GPU. Now, there's lots of really cool things that are happening within model quantization to be able to use into four, into eight. So we can fit some of these models into these small footprints, particularly for inferencing. But just to give you a sense of, in our network, when I use an accelerator for this kind of AI workload, it takes 55 times longer to run the model in the accelerator the first time versus subsequent inferences using that accelerator.
Where's that 55 times longer come from? It comes from loading the model from the CPU's memory, out through a PCIe add-in card interface into the GPU's memory right? It's that memory transfer overhead, you know, you're sucking up Niagara Falls through straw. It's not a very quick process and that can be a challenge, right? So the challenge in all of these things is working with your teams. And so I very much feel, and when I say teams, I mean the software developers.
Whenever we can have a model that is running inference at scale on CPUs, that model is going to run everywhere. We're going to benefit from latency reduction, because we are very close to every edge user everywhere with CPUs. If it has to run on an accelerated node, we're going to have to be in a place where we've got the power and the space to efficiently have our eight kilowatt rack, take an extra add-in card with at least a third as much power as just your standard CPU was taking just for the accelerator card, if not 2x, right?
There's edge 100s or over 700 watts. So we've got to build a system that can take that, that can fit into an infrastructure, and then we still have to pay that burden tax of loading that puppy up to run. Ideally, it'll never be inference. >> Ideally, so what you're seeing in here is and kind of accumulation of the conversations we've had with Intel, Dell Technologies, Google Cloud, from TLS acceleration and allowing your developers to just developing an objective way and then they get advantage of the accelerators and not needing to actually program for TLS acceleration.
It's just the CPU was handling the additional performance requirements without developers needing to tune their systems for the protocol to AI and machine learning inference in application. It's about the application. You can't simply look at whether it's a GPU CPU, overall system, and say how many transactions does it do a second. It is not that simple. It is taking the design elements. The reason why we went out to Dell Technologies, CloudFlare, and Google Cloud is to help give you an idea of who do you need to talk to to really understand your workload.
Who do I need to talk to to create the application Google Cloud in this instance? Who do I need to talk to to bring that to the far edge, Dell Technologies? Who do I need to talk to to actually do the networking? CloudFlare and who's powering a good portion of this with their processors and their accelerators and their DPUs? Intel, an example of this sponsor content. If you want to learn more about any of the companies we've mentioned below, I mean mentioned, you can follow them in below.
If you haven't seen those other videos, links to the other Lightboard sessions are below. Thank you Intel for sponsoring. Rebecca, as usual, thank you for just blowing my mind on just how big the entry is, how complex these problems are, some of the tools we have to address 'em. >> Thank you.