Running Batch Processes Across Clouds - Multicloud

11:17 · Watch on YouTube ↗

Transcript 2,001 words · about 13 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

(air whooshing) (ethereal music) >> Hey, you're watching another episode of the CTO Advisor, CTO Dose. You may be wondering, hey, Keith one there's someone sharing the scene with you. And then two, you're not in your normal location. This is part of the CTO Advisor road trip. You'll see activity in the background we're in beautiful Colorado Granby. And we're at a RV park of all places, Luke. >> Absolutely crazy. We actually got here together randomly. >> Randomly? We planted this for about a month and then, you know, we're trying to coordinate and you were like, well, when in Granby you are going to be?

I'm like, I'm going to be at this RV park. You're like, I'm going to be at the RV park too. >> Kid drug me here for a vacation. So yeah. >> It is a beautiful place, but we're here to talk multicloud. >> Yeah. >> Not multiple clouds. >> Definitely not multiple. >> So I'm going to draw for you a multiple cloud picture. This, you know, we got our AWS, we have GCP. We'll just start out with these two- >> Yeah, that's good enough.

>> Clouds. And I have services and GCP that I want to run batch processes against data in AWS. I've been told time and time again that this isn't feasible because there's not a solution on the market that does this. Faction provides such a solution. >> That's correct. >> Let's talk this out for something like batch processing for ML, AI, batch processing, for example. >> And actually I think it's feasible technically but not business feasible by other vendors because of that egress fee.

For all that data going to GCP would actually have to be pushed out of GCP. >> Right. >> Incurring a massive egress fee to get the data into eight- >> Let's say I have, you know I got a one petabyte data store, which is a small. >> That would be on the small side. >> That would be on a small side. >> Yeah. >> You know, I could probably even do most of that in Google or in AWS. But when I get into the tens of petabytes pushing that much data out of GCP into AWS is just going to be cost prohibited.

>> Absolutely. Let alone that data's now in AWS. And once AWS writes to it, it's no longer the same data that you have in GCP. So now you have to actually egress the fee out of AWS, back to GCP, to synchronize the two data sets. >> Exactly. So- >> That's time and cost cuz it's infeasible. So that's what I meant by business wise, little infeasible, technically sure you could possibly do it. >> And this is why we don't do it.

>> That's why nobody does it. >> But you're the, the argument that you've made with me on several other videos is that Faction makes this possible in a business friendly matter. Draw it out for me. I got to sure see. >> Sure, sure. You mind if I erase the one petabyte? >> Oh no, go ahead. >> So we'll do the bigger one. We'll do a hundred petabytes. >> Oh, okay then. >> Stepping it up. Yeah, there you go.

In this right here. We actually call the FIX the Faction Internetwork eXchange. And what that is is it's our own layer two it's our own fiber. It's our own connectivity in that region. And there's two important pieces of that. 10. 10 in GCP. 10 in AWS. So the application's the same mount point the same access, same name space if you want to put a layer three on it but also more importantly, transitive routing. So GCP now talks to AWS.

AWS now talks to GCP as if they're on the exact same network. >> Okay. That's not something that I've saw from a storage provider. This isn't just storage. >> No, this is a true data fabric. We have to actually stitch this together in every region so that the storage is usable in the network and then applications and services between the various public clouds can now be intertwined. And this is true multiple, sorry, multicloud not multiple clouds of data and silos.

>> So I have my application. I'm just making, you know, I have a route 53 entry for my storage target. Now I'm just making a, a DNS call or? 10. Once again one millisecond of latency we can achieve almost I think an AWS 400 gigabits a second of access. So this is true physically and logically as if it's in that cloud, same thing at GCP, I think we can do 800 gigabits a second, one millisecond of access to that data and GCP.

>> So I got to ask the question cuz people want to know where's my data? If it's not in AWS, if it's not in GCP, where is it? >> Yeah, that is a little complicated. So it has to be in that region. So let's call this the Reston, Virginia region. We'll physically map out the fiber routes in the locations of the various public cloud providers and either select a colo provider. If they're in seven to 11 kilometers of the various public clouds or we'll actually have one built for us.

And that's where the data actually resides. >> So I've tried to build this myself like in my- >> You have. >> Yeah, yeah. >> Several tens of hundreds of million- >> No, not to this level. I mean, just from my application perspective. My data center is co-located in a facility that has AWS locality. I have a HPE storage array. >> Yeah. >> And then I have workloads in AWS in a region that I believe is close enough.

I don't have the technical talent or skill to do, go as far as what you guys have done. So I get mismatch performance. How do you guarantee that 400 gigabits per second of performance? >> So we have some really cool technology actually to accomplish this. So once again, this is our own fiber connecting into the on-ramps of that. >> Right. >> So we can actually then span as many direct connects or express route circuits as we need for the various providers.

But one thing we do to actually guarantee we actually call it Canary we'll turn on a micro VM, very, very small VM in every availability zone and let's call another availability zone. So this one's one, this one's two of GCP. AWS, I believe has 16 availability zones. So we'll turn on a micro VM in every one of those availability zones. And then we'll turn on both a TCP and a UDP packet stream. And we actually monitor the TCP and UDP packet flows as if they were data flows all the way through to our storage systems from every availability zone in every public cloud in every region ensuring we can achieve not only the latency, but the throughput provided for the SLO the service guarantee that we give our customers to the storage.

This is all proprietary systems. We turn these on every five seconds, monitor that data flow, shut it back down. >> Wow. So, all right, this is sounding before I sounded too good to be true, but now I'm starting to understand the tech behind it. So I have other questions around like the data services that reside on top of your data fabric. Great, you can get me the block storage super fast at a relatively low latency. >> Block file object, you name it, we can provide it.

>> That was the next question. And then what about advanced services? We're talking about enterprises that's used to, you know De-dupe and malware protection, like the physical capabilities of a storage array that are advanced and we're not quite seeing in cloud services are you guys offering some of those advanced capabilities? >> So once again, this data fabric in our multicloud data service is far more than storage. So yes, deduplication compression because it's actually dedicated arrays per customer that you scale up and scale down like public cloud infrastructure.

But this goes far further than that. We actually do search across clouds. So we're doing all the metadata tagging in the object service, out floating it so that an AWS search for an object is also the same search that you would provide in GCP, same metadata capabilities, the same functionality we provide that centrally. We're also providing identity, access, and management. We can actually reflect identity, access, and management whether it's on-prem, whether you're using AWSs you're using GCPS, you're using Azures or anyone else.

So that's one common bucket, one common file one common name space for all identity access management. If you can combine cross cloud routing, same IP large bucket of storage, low latency, unlimited speed with search and identity. Now you can turn on applications in AWS the same application in GCP, same application in any of the other cloud net region. And you're actually making what I would call jokingly a super mulitcloud. (laughing) >> All right. So I'm, this I'm actually pretty impressed with this.

>> Hey thanks man. >> IM part, kind of, I was going to ask you about the IM part because this is one of the biggest challenges- >> Oh, absolutely. >> Application development. >> Absolutely. >> When I have, this is the power of being in one cloud, like AWS, I can take a, IM identifier and I can assign rights to a object and AWS to another IM object and AWS, and it's what gives power and makes applications built in AWS, sticky. >> And seamless.

And you feel like you're almost locked to AWS because you're using their features and functions. >> And so the data piece of it, when I can have a unified IM across multiple clouds for the same data sets I talked to, actually my last video on super cloud was about this exact topic was, you know what, sure, so what if Snowflake gives me data in multiple clouds. When I want to share it outside of Snowflake, IM is broken. >> Yep. >> Identity and access management is broken.

Now I have to figure out identity asset management a second time. You folks are giving me that capability built into the underlay of the storage platform, data fabric. >> Yep. >> So it's really is a data fabric. >> It truly is. And we believe actually, I'll be honest, this is the only non-marketing data fabric. This is a true functional we have literally 25 of the fortune 500s. Some of the world's largest data sets hundred petabytes is actually pretty small in our vernacular.

Exabytes starts to get around an average customer size because once we built this you start to take advantage of true multicloud. >> Wow. So Luke I've really enjoyed this conversation. I'm glad we broke out the lightboard. >> Yeah. Same here it's fun. >> This is about as deep as I've ever gone with you guys talking about this topic, your data fabric and understanding the engineering, the capabilities. com. If you have questions that I didn't ask Luke, you can ask them, Luke is actually he's on Twitter.

>> Oh yeah. >> Twitter handle below, or you can DM me at CTO Advisor and I'll ask them on your behalf. com. Talk to you next CTO Advisor, CTO Dose.