Getting Real with Multi-Cloud: How Hybrid Enterprises are Evolving to Compete
Transcript
>> Hey everyone, Matt Wallace from faction here. I'm really excited to present to you today at the CTO advisor, virtual conference. We're going to to talk about multi cloud today, like factions got some unique perspective in this space, we've been involved in some really interesting multi cloud projects. So let's just dive right in just a tiny bit about myself. I have been involved in technology now for 20 years or so, and I've been at places like VMwere,ViaWest, level three, and so on.
I am a technologist at heart, and so the innovation we are talking about is definitely near and dear to me. You know, a little bit about faction, faction started off its life as a private cloud provider. And we built private clouds and cross nine locations and have thousands of hosts and so on. But, I think what's really interesting is what's been going on in the last few years where we evolved and taken, you know what we did from a CPU, compute storage network standpoint, and actually, pivoted a lot more towards multi cloud use cases.
And we got focused on how do we help enable people to take advantage of cloud innovation. And we sort of, built up this set of cloud data services with our cloud control volumes, which is a multi-cloud storage, offering central to that, and really get a lot focused on that. We still continue a huge partnership with VMwere, and are one of the leading VMwere cloud native us partners, you know, partner with Amazon as a consulting partner and of course, with Dell, we're helping power the Dell cloud storage services.
But I think what's really interesting is this multi-cloud focus, and what we're trying to do around innovation and data in these multi cloud environments. So, when we kind of start off just talking about multi-cloud, you kind of, have to define it. And in this case, we're going to say that it's the use of more than one public cloud for the purposes of getting advantages you couldn't attain with only a single cloud. Now, you know, some folks might say, oh, hey, if you take the same application and spread it into multiple clouds, you're effectively multi cloud.
In, there's nothing really special about what you're doing in any cloud and, you know, you know, it's not to quibble and argue about that definition. But I will say, that we are focused very specifically on the advantages, that you can only get from going multi-cloud, right? And I think, sure, there are some things like saving money, right? Just being able to, for example, scale spot instances, across multiple clouds is a big deal. When you know, prices, capacity can vary for those sorts of things.
You know, being resilient factors in right surviving, you know, failures and so on and some people say avoiding lock-in, at the end of the day, of course, to a certain extent, you are always choosing lock-in when you choose cloud services unless you're going for the absolute lowest common denominator, which case, you know, what are the real advantages you're getting. But, but most notably, and this is something I don't see talked a lot about, but I think it's really critical to talk about with respect to, you know, the motivations that people have around multi cloud is innovation.
Because I think people go to cloud in the first place, because they are interested in faster innovation, more business agility, being able to solve problems, and if you give up on all the other clouds, and you stick with purely one cloud, obviously, you're not going to get all of the innovation, if you assume you know, cloud a is 50% of the innovation, and cloud B is 35%, and cloud C is 15%, you're locking yourself out of at least 50% of all, innovation, if you choose specifically to be isolated into a single cloud.
So powerful there and I think you know, when you think about how many services there are in public cloud today and how rapidly they are growing, I think that's where you have to really consider any strategy, that is not oriented around this multi-cloud reality. So you think about the 10's of billions of dollars that is now cumulatively being collected by public cloud providers and you imagine that they are out there hiring as many of the best and brightest as they can, and they are furiously trying to build services to compete for this share of wallet, and that these services are sticky that people once they're there, they stay, they put their data in, they stick around for long time, very powerful motivation to win this battle.
And so this competition between public clouds is quickly going to drive innovation and that innovation means new services and the early days of public cloud, I think you could have said there was a lot of similarity between the clouds right and instances and instance, and network as a network and block storage and file storage and object storage and something like manage database server, there're going to be differences, but the differences are nowhere near as large as the similarities for services like that when you start getting into certain areas, more recent in services, things that are much more innovative, there's not a light for like you're not seeing libraries to abstract between, you know, some of these higher level services between the clouds, because it doesn't make any sense to do that, because they have much more different than they have similar.
So, your ability to go and select say best of breed set of services across multiple clouds, depends on your ability to do multi-cloud, and it's certainly not the only motivation for going multi-cloud, but it is a significant motivation, for going multi cloud. You can't afford to settle for only some of the innovation if one of the best reasons for being in public cloud in the first place is innovation. And so with that in mind, we're going to kind of talk about practical multi-cloud, what does it mean and where are we seeing it leverage, you know, both types of businesses from vertical standpoint and a little bit more about their use cases and draw from, you know, our real world work in this space.
So, one thing you know implementing multi-cloud, we have this idea of, you know that it can be practical today, you know, fashion and you know, the Dell cloud storage services that we deliver in partnership with Dell actually give us, you know, this great platform that's a springboard into a multi-cloud world. And we have this idea of tying together the on prime environment into what you do with our service, and then that being connected to multiple public cloud. So we create this data continuum for you.
Right, out of the box compatible, hopefully, with your on prem-investments, and hardware and operational expertise there, once we've got your data, you know, we are using our secure, highly isolated network backed by our patent portfolio to connect you into multiple clouds at the same time, giving you the ability to access the same copy of data that you'll only have to store one time making it accessible from multiple clouds. On top of that, we have this idea this concept of cloud data services, where we want to be able to do additional things above this, at the storage layer, in order to kind of transform data, so things like our storage gateway handler, things like what we call Luna, which is something that can take, for example, a VMwere VM off a block volume, and actually extract it and send it into the public cloud as a machine image, things like that, are allowing us to kind of, begin to layer on additional functionality on top of the just pure storage platform.
But then the day I think it's powerful, just thinking about it from the standpoint of what are the only ways that you can very easily deliver the exact same data, and I want to say synchronize, but it's really not synchronized, right? You can have multiple copies of data in multiple clouds. But even if you do it yourself, you know, you really have this challenge of what's the lag time to do that. We're talking about one copy of your data that is stored centrally, and so when you write to that data, it basically immediately becomes accessible to the other clouds as well.
So this month, proven exceptionally powerful for folks that want to implement multi -loud strategies in their applications because of what it does. You know, some of the things that are changing around multi-cloud we'll just talk a little bit about trends real quick to understand where some of these things are coming from, containerization is a really big deal, could be seems to like it's taking over the world. containerization is interesting, because there's definitely a drive to separate the persistent data from the rest of the application, containers want to scale horizontally.
Consequently, you certainly do not want your horizontally scaling container to have persistent data because that it has to be unique coz it's got unique persistent data on it. What is advantageous about the way we do things from a multi-cloud perspective with containerization is that we can make the same persistent volume data on our service available to multiple clouds. So imagine being able to move a container from Amazon to Azure or in the blink of a microsecond and then the persistent data remains the sama between those two instantiations.
That's pretty neat. Of course, people are getting more and more interested in cloud diversity because, as cloud providers provide more and more unique services, they are finding developer teams, really want to tap into that. And often it's times, it's not like the same team or the same do going, what we have to have Amazon and Azure at the same time, right? Instead, you've collected a big data set, and the first team goes, we're going to build whatever this is, right? This application that does XYZ, they go, they build it, Amazon's great for them, they are keys to it, and it's got a good set of services.
Team B comes along, you have to we have to do this team B. And Team B goes to implement, they go well, this would be twice as fast or five times as fast we could just use this service and Azure for it. Okay, well, but if the date is in Amazon, how are you going to deal with, you know, connecting those two things. And one of those things that we kind of avoid from you know, this is a real lock in the state of gravity.
It's very difficult to move petabytes of data, even with a really fast network to say nothing of egress charges, right? So we help solve for that problem and I think we're seeing more demand for variety of clouds across enterprises. I think, you know, the other thing here is just data gravity is becoming so significant because the data volume is enormous. Some of the projects that we have worked on recently are in the, you know, couple hundred petabytes, growing by multiple petabytes regularly sort of range, and it's just so much data and trying to send someone to move that data on demand between clouds is like borderline preposterous, so we're solving for some of that, and this are some of the things we see.
From a vertical standpoint, couple places we see it, oil and gas, and we're talking about, you know, especially petrochemical seismic analysis, you know, people who are essentially part of that production of oil, usually from a prospecting standpoint, healthcare, this can be anything from the big data, stuff like genomic analysis to things like disaster recovery for like, healthcare applications like, like an epic, for example, where they used to store patient records. There's things we see across financial services, you know, now we're talking usually about use cases or even from high performance computing type use cases where they need a lot of horsepower to hurry, media and entertainment that's really changed the world, in the sense that folks like Netflix and Hulu and HBO, who're all you know, Amazon even who are producing just massive amounts of content now, they're starting to do it a lot in local markets.
So, if you have a show that is, you know, produced in Thailand for a Thai audience, you know, we start thinking about that times all of the people in the world and all sorts of different production companies that have to do that and then some of them get converted to another language and some of them get subtitled and dubbed and some don't. There's all this variety, there's so many production houses that are involved in this and there's just period not enough to keep up with demand for producing more shows right now coz there's such a battle going on to make content and attract those viewers.
But, what's additionally pretty interesting about this is just that how multi-cloud enables collaboration. So you may have scenarios where people can easily move, you know, between one production house may have its desktop instance in Amazon and another may have an Azure and being available to swing the data that represents that video between those makes it a lot faster and easier. Another place we see a tone of work is, the auto manufacturer. So, it seems like even if you just look at the news, like every you know auto manufacturer today has a self driving car story.
And I'm not surprised. It's pretty amazing even to see something like Tesla, it navigate out on the freeway and be able to do things like change lanes independently, and we're really just at the beginning of this evolution, obviously. But these things being powered by machine learning means there's just massive amounts of data, and massive amounts of horsepower needed to retrain those models and new data coming in all the time. Now, since there's probably hundreds of billions of dollars per year of total economic value from that transformation to self driving cars, it's going to be worth in the long run but in the meantime, really challenging use case to solve for that we we do a better job at with multi cloud.
And then you know, it's not really a vertical, but a use case that we see consistency consistently across verticals is the business resiliency need, and often, we are addressing disaster recovery. So let me just dive in, and we'll go through a few of these and cover them specifically, right? So,one of the oil and gas projects that we've worked on, we've seen this pattern of wanting to virtually prospect for oil, using seismic data scan from the sea floor. And you know, being able to do that while simultaneously saving on operations.
You can potentially have this sort of, if you will try factor of wins it's right, where you save money on DR you're able to operationally leverage native cloud services with that secondary copy. And you're able to essentially go multi-cloud to enable what you do from an analytic standpoint to begin stretching across multiple clouds, using more than one clouds, cloud native services. So, you know, that's a mouthful, but you're really leveraging different things to kind of drive this use case, and we'll talk specifically about how this worked for a customer of ours, right.
They wanted to consolidate data centers, so there's clear cost saving motivation here by just getting rid of the Dr data Center. It was really important to handle unstructured data though they had three and a half petabytes to begin with, still growing pretty rapidly, they needed to be able to handle that unstructured data because if the Vmwere or VMs failed over and the unstructured data is not there, they just can't be that useful, right? They're expecting those people they would've processed that unstructured data footprint.
Hybrid disaster recovery service allowed them to do the VM you know, storage failover and get them back up and running. We targeted VMwere cloud on AWS, because of course, it can pull up a VMwere workload in a replicated using VMware format. It's guaranteed pretty much to work on both sides. But we helped them set up the public cloud landing zone, at the end of the day, you know, they had a variety of good outcomes, but you know, the key thing to me is the data lake they were replicating from an unstructured data standpoint, is now accessible from all the major public clouds.
So we're talking about going from something where they could really only use software on-prem. Now they're empowering their development teams to be able to do work with that same data set, but across clouds, leveraging cloud native services. Now, did they save a bunch of money on DR? Yes. Do they change it so that the could automatically sync their data around from on prem? Absolutely, which obviously helps with things like secondary copies where they also save money. But the reality is today, the thing that probably moves the needle most for this company, which is, you know, a fortune 500 and has huge stakes in this game is really that that data there spent so much time ever collecting, and is so essential to their investment decisions on where to spend time and money, now it's accessible from multiple public clouds, which means developers who think that they can kind of improve on the status quo from a machine learning or analytics standpoint, can go actually use that without having to copy or move that data, which would have been impossible.
So, really a pretty interesting win. From an automotive standpoint, yeah, I mean, a lot of this really just is around being able to get, you know, the equivalent of, you know, 10s of thousands if not millions of GPU hours, basically, to retrain models, new data is coming in all the times, you know, clips of video, driver telemetry, still images, and so on. Now, we're not even talking about the operational use case, of course, because most of this is all still in the replanning implementing phase or in D Phase.
But there's a sort of preparation for what does it look like when this is real time? Right? Can you, for example, being able to process data rapidly take a set of video clips and images from a set of cars on a particular freeway, upload it, do some analysis on it so that other cars can receive a download, so they don't get involved in accidents, so they can be better at avoiding a pothole, like, you name it, lots of interesting things that become possible in this use case, right?
Even something like route adjustment, you know, it's interesting because the ability to turn up a lot of spot instance capacity for GPUs occasionally can be potentially powerful for a use case like this. If you realize like Tesla, for example, posted their specifications, talking about how many hours it took them to recover across 10s of thousands of GPUs, it's pretty interesting read, but you know, conceptually, since, cloud provider infinity taury and hardware types, and so on, all vary really rapidly. Having this all available is really actually very key.
Because you know, your developers may not have access to the quantity or the quality of hardware that they want without this. And of course, you can try to build something bespoke, but if you're not running your stack 24/7, it's kind of a waste of money. And if you think about something like retraining a model for self driving cars, collecting tones of data all the time, for sure, but,once a model is trained, the actual decision making process against that model is in a relatively steady state, and so it makes a lot of sense to be able to use the cloud can handle that analytics portion when it's relevant much more quickly, you know, in a real world example, where we designed you know, we had literally a 10 plus terabit backplane required because there's just so much requirement to process things.
So in this case, it was very specific, that we need to solve for some sort of initial ingest processing some scale up across, you know, people doing additional analytics work ad hoc across multiple sites. But we also, we had to provide a compute environment that could take 200 petabytes of imagery data and actually retrain a model with it over the course of four days. So being able to train 50 petabytes of data, and actually, just so little time, right in a mere 24 hours and then within four days, being able to retrain all of that.
So, it's a really substantial amount of throughput that's required for that. Now, the ability to scale that up across multiple clouds is really a powerful thing. (coughs) Another use case here, biotech, I got personally involved in this quite a bit and that was enjoyable. There's a particular use case where high performance instances are leveraged to drive genomic analysis, that genomic analysis is used to kind of do things like identify deltas in the population. There's a software that is manufactured or made by a company called Parabricks that was recently acquired by Nvidia.
That software, leverages GPUs to accelerate that genomic analysis. Let's take a quick look at kind of the pipeline here that gets used for that process and we'll kind of walk through what happens. (nasal clearing) See you have in biotechnology, this material that takes actually samples of DNA from your bloodstream or other source, at that point, you have to convert it into these, these binary alignment mapping files, right? So you're taking snippets of genetic data, you're essentially using a sort of big data process if you will, to go and rearrange all those tiny pieces almost like trying different arrangements of a jigsaw puzzle so you can get a final sequence between them.
From there, you go into this variant calling step where you can actually compare it to other genomes. (nasal clearing) Both speed as well as velocity or throughput can matter in this case, right? So if you have a doctor's order, because somebody is literally in the office, they've detected cancer, and you're wondering, will they respond to this drug with this particular type of cancer? You know, genetic analysis can tell you that, but you have to sequence the genome and then you have to compare it to existing known good genomes in order to know for sure if that's likely to be true, yes.
There's another version of this where research matters, and this is where throughput starts to get interesting, where the tertiary analysis, where for example, you may have done research on hundreds of people who had a particular condition and have their DNA, and now you want to take all of those genomes and say, okay, of our study, all these people had this common condition, how are we going to compare all of them against the sort of baseline population at a time so that becomes really important to, to a throughput down endpoint, if you're doing research, you don't have to wait for that.
But you also probably don't want to waste money, and so now this is one of those cases where we've actually been able to help by allowing people to scale across multiple public clouds, taking advantage of spot instance capacity in each. One thing that we learned in this process, by the way that is super interesting is that, GPU availability and GPU sort of generations and scale, definitely not the same across public clouds, you know, we kind of tend to think they do it all, and yet, you go try to figure out who has the newest Nvidia GPU and how many of them they have, we're definitely not seeing equal participation in that things across the cloud providers.
So really interesting use case obviously, moves the needle for humanity in a very real way, but it's one of those cases where multi-cloud actually really lends itself to this now. Good example to have where cloud provider specific things can also be quite interesting, although it's not hard and fast required, Google's done a lot of work in the specifically, in the area around this variant calling, and actually has a specific model that is built on their own sort of deep learning platform that does its variant analysis in a completely novel way.
Not that it produces completely different results, but it can be accurate in cases where other methods were not accurate. Now, it's their case, it's possible to pull them that model down and leverage it without actually having to do it in Google's cloud. But of course, as we see, public cloud providers provide more innovation like that, very reasonable to expect there were certain things just like that, that your enterprise might want to do, where it's only really possible in a particular cloud. Last one maybe, to touch on here media entertainment, super interesting just in the idea that hey, you look at a Netflix or somebody like that that's out there you have all these studios that are producing content for Netflix or for Hulu or for whoever, you know, the need to collaborate and and collaborate at scale and collaborate almost simultaneously is only growing, and for example, if I've got, you know, the footage for a particular film or series that I want to, you know, have special effects A be handled by one studio, special effect B be handled by another studio, a third studio was doing colour correction on pixels.
Now, how do I do that simultaneously? And the answer can be multi-cloud. Right? Now, suddenly, your professionals are mobile, they're allowed to work wherever they want to work, they can go to co-working spaces, they can work from home, if they're traveling, they can access a cloud based desktop with a cloud GPU, and the data is there. It allows collaboration and if you know the data is in place, for one use case, it means now it's accessible to another use case without having to actually copy and move it.
And that's a huge deal honestly, with the size of media files. We will aound out, this whole thought about multi-cloud just by talking about three interesting things we've seen three sort of novel things about multi-cloud access. The first is supercharging, a container strategy. So we've worked with several people, now several enterprises that are working to, essentially leverage persistent volume claims in Kubernetes, so that they can take the persistent data of a container, turn that container down in one cloud, or just add new instances, potentially, in another cloud and have all of those containers access that same data, leveraging a persistent volume claim and have that essentially, instant motion of data between clouds, but it's not really moving, because the data sits in the middle on our platform, really powerful use case.
Another thing is this idea of universal regional sync and one of the things that we found is, there's not always a great way across teams and across groups to provide the same data to multiple applications, across multiple clouds and keep it all sort of synchronized in an automated fashion. So one of the things our platform can do is, regardless of where your data come from, comes from, regardless of if you want one of the instances on our service to be canonical, and others to receive replicated or whatnot, we are able to take that data and synchronize it, so it's available simultaneously in Region A and Region B, and Region C, and we handle all of that underlying replication, making it essentially dead simple for the developers or other sort of stakeholders can access that data.
And the third one, and this is sort of trend, is to have a single data lake right, which is to click everything into the data lake, yes, it's going to have multiple copies of the data, it will have the raw data, it will have some extracted and transform data where a schema has been applied, etc. But this idea that you're going to want many applications across the enterprise you will be able to access all of that data at the same time in the same way, mind you, not necessarily the same protocol, right?
So that data access, has to support object storage, it has to support HDFS, it probably has to support safes and NFS as well, but being able to access it in a multi-protocol way to keep it all together as powerful. And yes, you can do that in object storage in a particular cloud. But you definitely can't do it in multi-cloud, and as we were talking about, if you're taking the way the ability for your developers to use services across clouds, because you're going to deny them the ability to use whatever service, is kind of, most appropriate for their use case in the cloud, you have a challenge, because you're kind of denying them that cloud agility.
So, with that in mind, we get to touch in on all these kind of real world use cases that we've seen, and hopefully it has changed your perspective on multi-cloud. I'm really happy to kind of talk to you today. Looking forward to the session coming up so that we can talk and take questions. Thanks so much. (background music)