AI Acceleration: Dell Technologies, GPUs, & the Future of Enterprise Computing
Transcript
foreign Townsend principal of the CTO advisor I'm joined with Robert McNeil product manager product marketing manager product marketing manager for advanced accelerated Computing at Dell Technologies accelerated Computing at Dell Technologies which translates to AI in these days all the advanced Computing that most Enterprises are doing is AI related we're at VMware Explorer at the end of the show it's been a great show and surprisingly or not a lot of the content was AI focused we hear the big stories about companies investing 10 million dollars to
do large language models Etc needing these huge Nvidia gpus that take up 5K to 10k of power per server which in perspective you know my experience uh typical Colo facility we have maybe five to 10k of power going to an entire rack we'll get into power in a bit but the reality is Robert that most companies and most use cases don't require kind of these big honking gpus that we can't get get hold of and then when we do do we have the right power requirements
for it so we're going to talk about the Dell Technologies overall product portfolio especially as it relates to Intel and Intel's emx we're not going to get into speeds and fees of AMX but the ideal is that Xeon based acceleration is going to be good enough for some use cases probably the vast majority of use cases at some point we're going to need some type of GPU acceleration and then we're going to talk about kind of the maxims of where we can get to right so
Rob thanks for joining sure give me an overview what what are what what is Dale seeing in the market so uh you know obviously in February of this year when uh the open AI chat GPT project went live there was a a huge level of Interest across the Enterprise not not just companies that have already been active in the AI space but you know became a readily apparent I think the general population that almost every aspect of the workplace can be some way augmented to by
by gen AI I mean as a marketing person it's almost impossible not to take advantage of it to come up with the the framework for a blog or a white paper because then it while it does require human intervention to make sure that it's accurate because it's all beginning stages of a new uh of a new branch of computer science um a the impact that it proposes to the Enterprise for uh accelerated decision making and more accurate decision making and uh and and just for basic
functions like uh software development we're talking about we've got a variety of different ways to accelerate applications uh possibly across multiple Frameworks there's there been some evidence that gen AI code generation while there are risks and concerns about releasing sensitive IP to the public versions of these large language models are some of the best possible tools for quickly developing highly portable code that could run on a multi multitude of uh of CPU or GPU platforms which would typically require require very specialized skill sets that wouldn't
exist in a single human being it would be across a large team of people and despite the cost for some of these higher end you know eight way or four-way gpu-powered servers the the benefits are are huge uh when you look at what the equivalent would be to get to that level of acceleration in code development or some sort of competitive angle within Enterprise such as document Discovery so uh uh we made an announcement um at Dell tech World here in in Las Vegas in uh
at the that we had partnered uh with Nvidia and other companies to uh build the um uh uh thing called project Helix we brought that to Market as a Dollar General of AI Solutions at the end of uh July July 31st we launched our first uh dell daily design for generative Ai and um uh here at uh VMware Explorer uh We've supported VMware with the announcement of VMware private AI Foundation again using a lot of the same tools and uh same architecture but then running on
top of VMware Cloud Foundation um and uh you know like you said a lot of there's a lot of interest in getting started today and uh the availability of gpus does not have to be a barrier to entry there are there are options such as Amex so any intel scalar CPU has the ml or machine learning acceleration via the AmEx chipset it's supported within tensorflow Cafe 2 pi torch we have a demo on dell.com that illustrates using the AMX chipset for smart cities application doing real-time
traffic pattern analysis looking for safety hazards and identifying traffic you know congested in in the example we show just using the native CPU the number of streams that can be decoded real time then how many more can be decoded and encoded and decoded and inferenced in real time with the native CPU and AMX chipset and then taking a step further combining either just the native in Intel CPU with the Intel pcie Flex GPU and then you know the the the obvious winner in this uh in
this is going with the combination of the AMX chipset and a flex GPU so that I think you highlighted a trend that I'm seeing in Data Center Enterprise architecture in general which is there's this I there's this desire to bring AI in as just another application within the infrastructure but what we see in the media is mainly this use case where you know where we're seeing that there has to be some GPU involved these are some pretty interesting use cases that that have nothing to do
with GPU and just using regular onboard accelerators right right yeah well I mean what can't belie the fact that when it talked when you're talking about uh very um compute intensive tasks uh so inferencing is is uh is less demanding certainly uh uh inferencing for non-uh large language model applications like uh like uh traffic analysis for smart cities um uh but to get the you know the density and in some cases the large memory spaces that are required for not just doing the initial uh training
and customization for an llm but for an Enterprise that sees their own version or their their uh their branch of an llm that has been meticulously uh trained to have higher accuracy for the type of of uh work that they're doing uh that is a competitive asset in their portfolio so um the uh they want the infrastructure to continuously train and retrain as they get uh results back from inferencing and and can you know verify the accuracy of the model that's when we uh go from
so we've got here CPU pcie which is what I think most people think of when they think about a GPU and then going up to the the latest and current generation for four-way all the way up to you know eight-way gpus on the uh um Power Edge 96.80 or oh sorry XE 9680 is the open compute platform accelerator module interface and the um and one example of that is the Intel data center Max GPU that is one option on the XE 9640 which is also interesting
it's the first in this high end lineup from Dell Technologies in the power Edge line that supports direct liquid cooling for both the CPU and GPU so if you look at the the 9680 or the uh or the um four-way equivalent of the 9680 the um the 8640 as both are air cooled systems that uh you know are six Ru which you know uh um the 9640 by using direct liquid cooling and uh um you know some specialized rack infrastructure from for uh data center Cooling
and power efficiency from Adele can shrink down that footprint to 2ru so as as we see a trend because of the power high power requirements of gpus that uh there is a trend that is not yet taking off in HPC but starting to take off in companies that are looking at revamping infrastructure for for Gen Ai and GPU acceleration to go up to 45 kilowatt from 10 kilowatts to 45 kilowatts all the way up to 100 kilowatts per rack and the 9640 using electrical cooling that
that high density of uh of GPU acceleration can complete get it get a completely full rack 42 Ru rack with uh um you know up to 45 or 100 kilowatts so you're not don't have an inefficiency of utilization per rack because of uh um running out of rack space because of air cooling yeah this is something that I've been asked a lot about one the one the availability of facilities to provide you know 100 kilowatts of power or into Iraq is that that's a slim line
ironically you know kind of bringing in another Persona of my CTO advisor Flying Cloud RV we've been thinking a lot about how do we bypass the local uh Electric System right of the RV to power a full rack even in that uh environment so as we're thinking about Edge deployments Etc Power Solutions for power will come right the problem then becomes cooling right and that is that second level challenge that I've been talking to customers about how they cool this how do they get liquid cooling
right into Data Centers that were not designed to do or Edge case is a good example where maybe uh looking at your your use case for inferencing uh CPU might be the right answer right because well it cost um at you know at a scale that's right for you know disaggregated uh inferencing that that uh you know maybe in an inhospitable location and then you know the obviously the power and cooling when you're when you're at the edge yeah so Robert I really appreciate you stopping
by the CTO advisor studio and sharing deals not just generative ai ai story because I love the smart cities example this is stuff that people have worked been working on for years optimizing tight resources making sure that we have the capacity to support future operations Robert has been one of my more astute product marketing specialist he already provided a link below if you want to learn more about what Dale is doing around AI if you want to learn more about what the CTO advisor has gone
on and more about the CTO advisor Studio you can follow us on the web the ctoadvisor.com my DMs are open I am drinking from not just the accelerator fire hose but the AI fire holes and all of this is new learning to me come along and learn along with me DMS are open on most social platforms at CTO advisor talk to you next CTO advisor Studio