Enterprise AI Is Missing a Control Plane (Equinix Distributed AI Hub Explained)

20:56 · Watch on YouTube ↗

Transcript 3,430 words · about 23 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

This exposes the security leakage that I've asked vendor after vendor after vendor about and you folks are the first ones that's brave enough to actually talk about it. >> [music] >> All right. We are ending up our road trip here in San Jose talking with the fine folks at Equinix. I'm told it's Equinix and not Equinix. I've been calling it Equinix my ever. I have with me senior technologists Cal and Lily, your solutions architect. >> Yes. The problem, Cal.

What is the problem that you folks are trying to solve? Sure. So you know, we're entering the era of agents and AI. Increasingly every company that talking to the the CIOs and CISOs they're they're all trying to basically help their employee productivity by basically deploying sidekicks, assistants, and co-pilots. They're trying to improve customer experience and they're also trying to come up with new creative IP and new solutions. But the problem that is occurring is that increasingly uh multiple groups in the company are trying to do these things and they're doing these in isolation.

So the developers for one app, they're going and developing it and they're setting their guardrail policies for that model in isolation compared to another group that is doing like a a card model for coding or uh marketing team using open AI for another reason. So there there are a bunch of uh you know, different organizations. So in one organization, for example they would use 180 different IDs to go ask the cloud model. So, that's that's like I would call it like you know, major challenge because now they need to make sure there there is no data leakage, that basically the you know, cost-wise also the cash if you use a common cash.

A lot of the queries across all these groups, you don't need to go and and and and you know, go to the model and pay again. Cost-wise, it's a problem right now. Consistent setting of policies is a is a problem right now. And then basically also the the pipes the because every query employees are adding context now as part of the prompts. And many times it's multimodal. And so the the amount of bandwidth that you need to move all this data in and out of the clouds and you know, your you know, clusters private clusters, it is increasing.

And the current networks have to be rethought. So, performance, price, and privacy/security are becoming a real nightmare. So, you're heading on. I'm going to do my impersonation of a chat GTP. No one talks about the the thing that no one is talking about is that this exposes the security leakage that I've asked vendor after vendor after vendor about and you folks are the first ones that's brave enough to actually talk about it, which is what happens when I have a security analyst with co-pilot and they have access to the online data set that they, you know, write security posture.

But when you put AI on top of that, they now can get insights that were not intended. How do I protect against that? That is the problem. It's a problem because you see, in most organizations, humans have access to employees have access to a lot of data. They just don't know that they have access to that data. Agents are relentless. They are now going and getting you every piece of data that correctly the the access control was set. You will go get that data, but as you said correctly, the AI is now stitching all these different pieces of data and suddenly you're getting insights that you're not supposed to be getting.

So, you simple question, where is Keith's Airstream? Like, I don't want you know where my Airstream is at, but it's it is the agent will stitch together public information, private data, and find the answer to that question. Absolutely. So, Lily, you're on the hot seat. Cal has promised me that your announcement, you folks are starting to tackle this problem. You kind of, you know, prepped us with uh picture what you just described. This is this is pretty scary looking.

Yes. So, Lily, and what Lily is showing here is basically all these different teams, and they're all accessing different services, like, you know, like different services from different providers, right? And those major problems that I mentioned, cost, uh privacy, data leakage, and and also performance. So, you see, what happens in many of these cases is that for a lot of these agents, they're going to access models. Because agents basically are taking intent in human terms, passing that to the models, and then basically getting the APIs they need to call to access all the different services.

The models are telling them. Right. So, the latency now increasingly when agents are talking to other agents, they cannot they want agent-level latency, not human-level latency. So, they want quickly to go to that model, get the response, go to the another agent, go to the And this is all happening in microseconds. It's like, micro fact seconds or fast as fast as my network connectivity or connectivity from agents. Matter of fact, it kind of speaks to what you folks do. You're kind of a airport.

Explain to me like what's the model here? Right. So, our uh lineage of a heritage for Equinix very interconnection hub. So, just like in an airport, you have United Airlines, American Airlines, Delta, and passengers are going between them or between, you know, different airlines or same airlines, different planes. We are like a digital airport where all the clouds, the networks, the finance companies, the the merchant uh the uh you know, uh OTT companies, uh the retail companies, you name it, pharma companies, they all come and inter- inter-exchange data between each other.

So, we have these kind of digital airports in about uh 70 plus metros globally in like about 35 plus countries. So, we are the place where today, like more than 90% of the data from the edge, from your, you know, Tesla cars, your phones, it is going through Equinix and landing in all these places. So, we are the place where that data is getting exchanged. And so, guess what? In the AI world, now the internets are talking to each other and that data is also going through Equinix.

You know, I drink my own champagne. I am a Equinix customer. The CTO of Hybrid Infrastructure is here. I advise com- companies and customers to do the exact same thing because of that airport analogy. You can get on and off public cloud as fast as possible. Uh public Some of these providers actually have facilities at a Equinix facility. So, it makes this luxury of data transfer possible, but it also exposes this problem. So, can I'm going to bring you I'm going to bring you guys for this problem.

You created the problem. How are you helping me solve it? Right. So, there are some key ways we are helping you. So, let me ask Lily to now draw what we call as a distributed AI hub. Distributed AI hub literally let's go ahead and draw. Oh, yeah. Yeah. Yeah. So, distributed AI hub is a framework where in this airport now you're going to now deploy some AI services which will help you to protect the actual data that it is adhering to the guardrail policies that you have set.

It's going to help you to observe uh that all the agents that suddenly they're not going rogue on you, that they're actually meeting their SLAs. You're going to basically it will help you to get low latency connectivity to all the models so that your you know, microsecond level that you mentioned, right? So, that you can get performance that you want. And also another important thing is that flexibility. So, now what I mean by that is we have customers today they're using some of the super intelligent models.

But then they're realizing I really don't need that. I want to leverage a much more focused open model. And then like, you know, 1/10 the cost. Yeah, a great example of that a deep project where all of, you know, we know our friends over at Tech Field Day I broke down 230,000 segments of conversations over the past 15 years. I needed to classify that data. If I sent that 14 gigs of data up to OpenAI, for example, to classify I would have been stuck with a pretty good pretty big bill.

I can use a 3 billion parameter model to do the same classification at a fraction of the price on my laptop. All right. See, so what we are finding is that customers will use the OpenAI type of models for certain tasks. But they also, as you said, for certain focused well-understood tasks they'll use open models. So, they'll use we we are increasingly seeing that it is like a hybrid/multi-cloud architecture and distributed data. And that's why we're calling it a distributed AI hub.

Because you will be leveraging models from the the hyperscalers, from the model providers. You will be doing what you did, which is run the model on a private cluster. And and so, you need to basically stitch together all of these different flow of models being called in the tool chain. And all that traffic is going through Equinix and through this hub right now. So, I call this the layer 2C. If I have a 4 + 1 AI infrastructure model, this is 2C.

My My audience is familiar with this. They're shaking their head, yes, Keith. This is exactly the type of solution that solves the the layer 2C problem we talk about reasoning. Where Where should certain AI run? What model should I select? How do I figure it out? You folks are addressing my security concerns with fan the 2C model. But, I guess the question is how. This seems like what we call in my community the black is a box.

Help open the science on the box. >> Box. Okay, so let's How does this work? Yeah, so Lily is going to now double-click on that box. And and we're going to explain exactly what are the different pieces, layers of the cake in that box in that cloud there, right? And what do each of these things do? All right, so let's let's open up the black is a box boxes. Lily, let What what's going on here? Yeah, so like for example, so in this side, we have all the user app.

For example, I want marketing member asking about a question about a specific product. So, the prompt will go inside here. So, the first step will be user prompt. And then, the traffic will go to the AI hub. And then, this hub act as a gateway that will choose which model are going to use based on the prompt. If that is a complicated question, it will go to big model. If that is like a easy question, it will go to some cheaper model.

So, but before every query goes through the model, it will go to the guardrail, for example, Palo Alto Prisma AI to pre- do a pre-check to check if there any uh improper content, like asking about like people's private information or doing something dangerous. And if the pre-check is passed, that Okay, that is proper. And then uh traffic will go back and then go to the model. So, so Lily, but you said gateway. So, can you give some example of a gateway?

Like so that people understand what is that gateway? Like Yeah, so like what it mean for a gateway is that we have like one entry point for multiple provider. So, like if we don't use this gateway, what's what we have if we want to connect to multiple model? You need to go to AWS console to set a card and create API key. You need to go to Azure console to set our Azure AI AI foundry and create API key there. And you need to go to Google Cloud as well.

So, you will have like three API key. And then you need to put that in your application. But with this gateway, we can connect to like a visual model, we can connect to more models, and then control all those like connection and the API key within this unified gateway, and that will like give the users like a clear dashboard, clear entry point for all the AI component. So, this is if as described for is absolutely layer two CML AI framework, this reasoning model.

So, this idea that a user prompt enters the system, the system decide decides what is the best model based on SLAs, uh context, etc. to run or provide the answer, return the model return the uh the results with using our system or policy guardrails. I guess that the question that that answers that I uh want to ask is what about autonomous system? So, that that's an example of a user prompt. What happens when I wanted to classify 205,000 conversations, which is a constant set of prompts.

I'm not, you know, maybe use VLLM against or uh some or I send multiple prompts at one time. How does the gateway handle that type of autonomous solution? So, I will start off and then I'll ask I'll let you do it. So, when you First of all, I just want to highlight that when we think of this distributed AI hub, think of these as all different services potentially from different vendors. Mhm. So, this distributed AI hub is a framework. So, for example, this guardrail service, we can get it from the cloud, right?

Uh this gateway could be a VM running on a server at Equinix in this uh digital airport. Uh this semantic cache uh could be a physical uh caching appliance that is hosted at Equinix. So, each of these components can come from different vendors and they can be either consumed through it as physical appliance that is hosted as a colo at Equinix or as a virtual VNF on a network edge platform that Equinix provides or as a cloud service that our fabric intelligent fabric can help you connect to that.

So, so basically, each of these things, the the distributed AI hub, multiple people and services from multiple vendors are working together. Is all this all that? And this distributed hub, is that a single API call? Like what what am I getting off my my marketing app on the left-hand side? What am I calling if I'm calling distributed so I would like let me go explain that. Okay, so that's that like a way back to the configuration. So when you are building this app when you are doing the configuration that in the past you need to like I mentioned you need to get the API key, but if you use the distributed AI hub you need to one get one key from this distributed AI hub and so like one key key here.

And this key allow you to use multiple model even tool even guardrail. And what is that what's this key can use depend on the policy. Yeah, so we can customize this. So if the this key for example marketing team need to use some internet search service so we can enable internet for this key but not for other. And for the maybe like there's a key for financial team and we need can apply the guardrail with those finance compliance on that key.

So key becomes my security object. I can apply my guardrails against specific keys. Similar to how it would work in the cloud like when I go to Google Cloud today I I can get a key and assign it rights and capability. Same concept. Yeah, but the benefit of using this like gateway is that you can have like a single place you need to you only need to apply your policy once and you can do like >> across multiple so all of this model if I wanted so Claude Gemini really doesn't matter.

If I I set the policy for the key for each platform within my distributed hub and then that key inherits that right for that application. Yeah, and you can also set policy for some like other for other base guy, you can dictate some keyword in your prompt and apply policy or you can apply policy on the employee ID. Uh you can customize the policy here like before go to the model and you can also do the policy on model routing. So for example, you want to like one thing only use I one model model here and maybe the thing from the United States using the model in >> So I get the so if I have uh distributed cloud or some sovereign cloud capability that I want to implement, this is a control point for that.

Yeah, it's like a unified control plan for all your AI component. And what we see a lot here is actually for the NCP tool because I different team want to have the access for different tool. So what we can do here is guy for different team or different app, we can open the accessibility for different NCP tool here. All right, guys. This is a bit overwhelming. And I think it is absolutely where I'm seeing a lot of AI projects stall.

I actually see end users within companies say, "Hold up, IT. " And this absolutely I think puts language to that problem and a opinionated framework for solving that problem. If people wanted to engage Equinix on this, let's say that they're they're not even hosting, they're not an Equinix customer, they don't have any equipment, but they want to have distributed AI, how do they reach out to you folks? Yeah, so in the distributed AI hub announcements, we have a email ID and a link they can go to that, click on that.

And then basically what we will do is we'll first engage with them and give it briefing with that customer about how what is what are the key reasons why you want to do this in a distributed AI hub at Equinix. What is the cost benefit? What is the ease of policy or consistent policy management benefit? What is the you know sovereignty now is becoming a big issue. What are the sovereignty things Equinix can do because they're all over the world, right?

And what are the latency benefits that you can get because you're accessing all these models, right? At you know machine speeds, right? We'll give what is known as first step a briefing. And we know yeah, and if they like it and they want to go to the next step, then we'll bring these partners in a workshop. So for example, Palo Alto and Equinix will do a joint workshop to show you how you can set these guardrails and actually architect this. And and actually our team like and Larry's part of my team with Lee, Andre, Brandon, and Angela will actually do this demo of this thing really working.

And if you're interested, we can actually do a POC. Joint POC with our partner. So you can actually test it out and then everything is good, then depending on which deployment model you want to use. If you want to use like a VNF, you can just click, click, click and deploy it. If you want a physical appliance because of bandwidth or control reasons, then you need to get a cage and get that appliance deployed. If you want a cloud, again click, click, click you can go and get these things set up in in Equinix.

So that's the process that we follow. And for your convenience, I'll put a link to that announcement below. com. Equinix, thanks for hosting us on this part of the CTO Advisor road trip. Talk to you next CTO Advisor live board.