Using Private Data Center for AI Inferencing
Transcript
welcome to another CTO advisor video or continuation of CTO advisor video if you've watched my earlier section I'm Alis toook and I'm joined for this video by the world famous Ryan Shrout I I still won't say world famous but I appreciate it it's good to be here you know I've done hundreds of episodes of podcasts and interviews so this feels like home it's good to good to be talking with you and we're talking about initially at least something that has absolutely been your bread and butter
for many years this high performance accelerated workstation yeah this this is really where uh you know I kind of cut my teeth on Enthusiast components and high performance you know PCS so seeing a a Dell Precision workstation like this is is always exciting you got 24 core processors and RTX 5000 GPU is 16 gigs of RAM and this system has 128 gigs of M in it so a a great workstation and development house for sure but of course that's no high-end super powerful server sitting in
a data center true but it's an awesome machine for single use and what I demonstrated in the previous video was running an entire large language model on this laptop as well as augmenting that with that with some documentation that provides domain specific knowledge so I could ask very specific very useful questions of this AI running on this particular machine I thought that was pretty awesome I think I think there's some really compelling use cases for individuals that want to go uh create AI models on their
own or augment or fine-tune whatever you want to call it uh of AI models for things like that I know I have significant numbers of written documents over the last 15 years that I love to be able to search or emulate and uh really kind of process and ask questions of so that kind of demonstration that you did is is going to be applicable to a pretty wide range of audience then just go you don't have to have uh massive amounts of of servers or gpus
to go run it you can run it on on some pretty reasonable Hardware yeah you don't need the massive scalability of a cloud but I think there's a lot more use cases where that that Corpus of data that bunch of data that we use to do the augmented generation is not going to be just a single person but it's going to be more of a corporate level or application Level there's going to be use cases where this is not just making Ryan more efficient it's making
his entire team his entire organization more efficient and I'm not sure that works just on a laptop I think are then probably going to see that running in data centers on real servers yeah and I think that's that's an important point that we're that we're trying to make here with this right is that you can do your development you can do your research you can uh find some of those smaller use cases here but the point is you can show that this can scale up reasonably
I guess I would say scale up without having to go massive on things uh and apply it to larger user bases um and when you start to think about things of I think you used a seven billion parameter version of Lama on this but if you want to go to one of the larger models gotta have more memory gotta have larger subsystems for for some of that and that gets to be a much more efficient way of delivering the resources when you share it amongst a
group of Staff rather than having to buy that resourcing for a single so I think we will continue to see this but I think it's important to see that this is something you can put into an actual Enterprise data center you don't have to push this all out to the cloud particularly if that body of data that you're using for the retrieval augmentation is very private or regulated data you might have very good reasons for keeping that in your data center and still want to get
all of this value and and I think the value you start to look at things through a small medium business perspective and uh you look at companies like ours where you might have you might have 10 or 20 employees but that data set's pretty large maybe we're working on some private you know uh NDA type material and things like that that we don't want out into the public cloud and so we can keep that information local uh still run these types of queries and and get
the value of a large language model or any kind of AI model like that um uh and and and utilize it in our in our workflow and you can see use cases like um medical information that's that's got to remain on premises at the the clinic that is dealing with the patient uh yet being able to perform natural language queries against the contents of a maybe a single patient's medical history uh or the whole genetic connection to them that could be incredibly valuable it could it
could and and I think you know you're you're demo that you walk through shows that even if you're if you're a small company and you're working with in-house resources you could probably develop an application like this relatively easily right or you could go to uh you know a independent software company that's going to develop these types of things and allow you to do that rag kind of retrieval uh AI work on its own as well so yeah I think applicable to a broad range of Industry
verticals that are going to use these language models to kind of improve their workflows and and throughput uh a bunch of employees I think over time we'll see more efficient resource usage but what normally happens is as the software gets more efficient using the hardware and the hardware improves a little the software then leaps onwards to do more so we will see more of those 70 billion um models that that open AI are delivering uh we will see more of the larger amounts of data required
the very large number of tokens required for that context information because I I want to have a conversation with this AI that remembers what I was talking about yesterday or at least if I'm a clinician last time I saw this patient I was having a conversation with the AI and I want that history to come through so there's a whole lot of places where we'll see increasing and decreasing amounts of resource you y so I think a a great demo to show what you can do
on a on a local machine and it should start to spark people uh in their minds and ideas about what can you do with with lest scale up of this across across data center Ryan thanks very much for joining me on on this video your insights are always valuable thank you and thank you for joining us stay tuned to the CTO advisor we will have a lot more content from GTC as well as from the awesome things that Dell is doing around AI