Elastic Inference - Sharing Finite Resources at AWS Scale

3:17 · Watch on YouTube ↗

Transcript 417 words · about 3 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

all right it's very fitting that the 14 year old's math problem is still on the light board we're not going to talk about math Investments but we are going to talk about today's AWS everyday product which is elastic inference this is a common problem we see in the data center we have this finite resource that's a graphics card and we have that connected to a physical health we may have several of these graphics cards and we're digging out VMS based on this specialized resource what happens

when we take and look at the analysis of the use of these machines we may discover that utilization may be somewhere in the 10 percent to 30 percent utilization of each dedicated machine or the overall resource may only be used 25 percent and a data center this is a problem because we can't guess what instances will use the CPU or the specialized resource and what resources won't and the public cloud is even more so it's according to any jassy back in 2018 dedicated P3 instances which

are the big boys these are the ones we're going to do our big inferences on is only utilize this 10 to 30 percent if your Amazon this is extremely wasteful you want to divvy out these resources as much as possible to as many customers as possible what do you do hmm you find a way to virtualize that resource now I don't know if they're using cxl on the back end I'll link to my cxl video how they're physically doing this but they're pulling all of their

inference engines and allowing you to provision elastic inference at the beginning of your provisioning host provisioning and as your application needs to do inference it'll then pull those resources from the poor resources at a much lower rate and then return them Amazon gets a much higher utilization rate for these resources increases their overall revenue and you as a customer spend less money on these big P3 interested which can then be reprovisioned for customers who are going to have higher utilization this is a interesting solution and

the data center we're constantly figuring out how to do this exact thing Amazon has done it at scale and have has been doing it for roughly four years of skill got any questions you can hit me up on the web at CTO advisor on Twitter AWS everyday.com to have this in your inbox every day