Webinar: Backups are Dead

31:59 · Watch on YouTube ↗

Transcript 4,220 words · about 28 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

all right everyone thanks for joining the CTL advisor and our webinar backups are dead the has been a long month as I inter introduced this this topic I'm super excited about doing it share my research in this area and talking to whether it's vendor clients or with enterprises struggling with this concept of moving past backups a little introduction to myself I am the CTO advisor the principle of CTO advisor and CTO visor on Twitter you can reach me via email Keith at the CTO of Viacom

quick career highlights I've been an IT for I hate to put it this way for three decades started my team career in the early 90s some highlights I've spent a couple of years as a program chief architect for a fairly large program within Lockheed Martin Department of Housing and Urban Development over about 12,000 end-users uh managed IT infrastructure from our architecture perspective for that fairly large organization moved on to PwC where I was a management consultant and tackle problems from a very different lens than traditional

IT and then I spent a couple of years in a fairly large enterprise managing again the IT infrastructure for a multi-billion dollar portfolio as the sa P system for that fairly large fortune 200 you can find my content on tech be public calm the CTO advisor comm of course YouTube and tech target as well as virtualized geek comm so hopefully we won't take more than a half an hour of your time on a Friday I know it's the end of the week I really appreciate for

those who are either joining the recording or joining this session live feel free to next questions in the chat and I'll answer those questions at the end of the presentation but this whole concept got started I was at a of all things and NetApp NetApp conference NetApp insight and a keynote er from The Economist's talked about data being the new oil this is something that if you followed Facebook Twitter all of the large web skillers is not anything new we've heard this concept if you're not

the if you're not paying for the solution you are the product your data is the information so with these large web scholars we've come to accept that as a fact however that is not the only case the it's not just the web skills I've noted a couple of large brands that are using data in interesting ways to become data driven businesses Capital One actually had a booth at AWS reinvent GE has gone big into data while they've had some financial troubles in the past their commitment

to data has not changed Ford again is for the legacy I for the legacy Auto Builders are kind of moving the needle when it comes to the connected car and then Southwest Airlines of all people some interesting stats from some of these major brands a 737 max which is the next generation Southwest Airlines plane from Boeing Southwest has flown 737 707 737-800 for quite some time 737 max the new generation the turbines generate about 844 terabytes of data per day in a 12-hour operating period that's

over that's just one aircraft over several aircraft in their fleet so Southwest alone generate generates petabytes of data per day just off of their airline data off just off their airplanes so IOT is just generating this massive amount of data I worked for a fortune 200 farmer in which we were challenged with pre on storing data off of the genome scene sequencers the ability to sequence the human genome was broke some years ago average costs of you know some crazy number of $100,000 to to sequence

the genome now nets that numbers come closer down to a thousand dollars so you have these sequencers running non-stop all day long to sequence several genomes a day at the tool of 200 gigabytes per sequence and just to give you some scope of well how many people will get there or will benefit from getting their genome sequence as we talked about target medicine and custom medicine in theory the University of visitor University of Oregon where they would like to see every new cancer patient have their

genome sequence so they could be ended identified for specific trials and treatments there's 14 point 1 million new cancer payment patients a year that's just the cancer use case there's many other use cases for the human genome and having your genome sequence other than cancer so again just the scale of these data sets autonomous vehicles famously this is probably one of the number one use cases when you go to a 16z they say the cloud is gonna die because of IOT and the edge for terabytes

of data per autonomous vehicle per day and then we're still not talking about the more common use cases that we see in many enterprises so dirty cameras manufacturing equipment wearables door locks etc abundance of IOT data is generating a ton of data in our enterprises so this is not just the big whelps WIPs Keller's that are worried about massive amounts of data if you manage storage you are accustomed to constantly requesting new data because our new disk and data services because of the explosive growth in

this in this data sources chest from IOT alone so what's done with all of this data well I was just at AWS reinvent last week and AWS did a fabulous job promoting AWS services one of which was the deep lense and which they take this IOT connected camera use IAI to identify they had a really cool hack at the at the event where you put a you could identify a hot dog and hot dog not hot dog obviously simple use case but these services grow always

add again a farmer which used this methodology of collecting massive amounts of data out of their s AP system and their commercial sales in general to identify the most at-risk patients that need it the high touch in custom contact of a practitioner nurse so early on in the sell of this product they'd have every new patient contacted by a nurse as this cells are these pharmaceutical grew into the building and to the billions while the financial math worked out that they could continue to hire nurses

they just simply can't hired that many nurses to contact that many patients so they had this bottleneck in which they needed some level of intelligence to reach out and say what patient should be contacted using KPIs in this massive amount of data they had collected on their patients before in consultation with their doctors and providers they were able to target and create call list of employees to call so cloud services such as tensorflow Watson Azur ai AWS are have become the engine for drilling this data

oil some employee some employers or organizations go at it on their own they may create a GPU farm so they may take open-source tensorflow and actually built their own AI ml on Prem and process that data enlarge that data if they believe that's the differentiating service to build their own AI farm they may do that on Prem and then there's global employees people are smart no matter where they're at AWS was obviously a huge conference but Alibaba also had a huge cloud conference just a few

months ago talent is spread across the world Microsoft is making big investments in development houses in in India we're seeing other third world countries such as Russia roll Israel has a really great startup environment the V Brown Bag guys were just in Israel doing a Ravello systems build a employees are global and this is just not from a technology perspective you look at Rd there's people doing amazing things when it comes to all timers research in Germany there's a great deal of oncology expertise on the

west coast the East Coast specializes in a certain type of medical research and they all need the same genomic data that brings us to the problems that many of these new vendors are trying to solve a while back up itself is dead backup is just a feature of data protection the thing that you want to look for critically that actually adds business value in brings to life or brings to value the the oil that's with this buried deep into our data are the the the metadata

and the services are associated with that data we have to give the data in the hands of the researchers researchers we have to get the data to AWS we have to get the data to Germany from New Jersey we have to get the data to Google in their tensile flow we have to get it to our own cloud our own private cloud of HPC and GPUs and MI and MML and AI we can't do this the way we've done it in the past I I'm sure

many of you have gotten on a plane with a set of data whether it's on a some DLT tape drives or I'm sorry some DLT tapes or some portable drives and carry that to another site because it was faster than copying it over the network we can't do business like that any longer the speed of business dictates that we do things much faster so basic features we were looking for in our quote unquote backup solutions which are now called data protection solutions are simple things such

as copy data if I need to refresh Oracle database from production to QA and QA can be in AWS they can be in Ravello systems it can be in your dr facility it can be right next to your copy of your production data right next to your production systems and we need to make a copy of that data we can no longer afford to make a basically zero level backup that data and restore that data and take up all that disk space 1 we have the

simple problem of time it takes time to copy that much data to we have a challenge of disk space we can't continue to just simply buy more discs in order to accommodate new copies of data for text QA development and then a third issue is just the speed of replication we can't we can no longer afford to give up private data center space in the case of on premises resources for these ancillary copies of data we want to put that stuff in a cloud we want

to do use the elasticity of the cloud to say that you know what I want to spin up a QA or test environment for a few hours or a few days and then tear that down and not incur the cost of maintaining that environment all the time copy data allows us to do that caching getting data from one side of the world to another side of the world or in another data center basically smart the ability to smartly place data where we think that data is

going to be accessed replication 4dr and for backup and dr so you know one of the primary drivers are rare the other day one of the primary drivers for cloud is still disaster recovery we want to still run VMware on Prem focused on production replicate those data sets out to AWS do goo compute or Azure and bring up our D our environment as needed again without incurring the cost of running those CPU instances constantly and that Tara wants storage when we don't need it and all

of that is powered by we want to take advantage of this new concept of what relatively new concept of object storage and s3 how can we do this relatively inexpensive if you've you know we can spin up generically speaking we can take a copy of a system running in our production environment create a equivalent ec2 instance where some block storage some EBS block storage in AWS and take a simple product like double take it just replicate real time the data between environments simply spoken that's the

old way of doing things it's expensive and it's it's ineffective at some point we will run out of bandwidth and it doesn't scale so we're looking at ways to leverage object storage which is much cheaper a little bit less performance performance but less Chica cheaper than traditional block storage these are the challenges and these are the capabilities we want our backup solutions to provide then outside of just practical technology in the value add that comes from moving data around the world and within regions there's the

red Latorre concerns if you haven't heard GDP are it goes into effect May 2018 this is the European Union's efforts to make sure that each European Union citizen has the right to be forgotten if you're joining this webinar from the US and you think oh this isn't a concern of mines it is if you do business with a European Union citizen and if you do this is in Europe in general you are in scope of GDP are and the penalties are severe up to 4 percent

of global profits so if you think about some of the penalties that Microsoft and Google have suffered from the right to be forgotten each one of those in instances so if there's 10 as the way I understand it if there's 10 instances of violations of the GDP are for each instance the entity can be penalized 4% not just the European profits but of the global profits this is going to be tested for sure in court but the EU does not seem to be letting up and

saying you know what we're gonna give an additional grace period for everything that I've seen and read they fully expect to implement this and full or start to enforce it in full beginning in May and then there's simply data sovereignty within your organization you may have concerns about where data stored you may if you're in the EU you may because especially with bresson you may be concerned that is data stored in an EU Union data centers versus in the UK or even we're saying that the

Ireland may be treated differently so data locality knowing where your data is at not just your production data your archive and backup data or even your replication data for dr it's important to know to have tagged that data and no other that data are at this this secondary market i like to call it secondary storage all the players that we're gonna get into in the next slide Dylan to this idea of secondary storage how do you leverage what had been considered a backup feature as more

of a data management solution how do we manage the data that we have across the globe so let's talk about the players in the industry so just few weeks ago that CommVault go and they have a solution hyper scale it is the traditional convolution but we've seen in the especially in the data center space awful lot of vendors attack this traditional dil EMC whose on the slide their data domain business where you have these scale out data domains or you sell out scale up storage units

in which you can continually expand your backup solution to accommodate the growth of space you know we had this problem when we first started to backup the disk is it's an expensive and hard to manage solution what happens when you expand across a single direct attack set of storage well wanna a practical problem that you have is that though that solution isn't balanced so you may have 18 months of backup on one scale out node and you add a second one well that new scale out

no you might put new backups on but now you have this you you have this failure domain in which if you lost one node you lose 18 months of backup that was the old approach so data domain and all these new solutions automatically read factor your backups across multiple scale-out units convo has introduced a solution built on Red Hat cluster file system that allows for the expansion across nose very traditional solution taking on a new approach they also have a cloud version of this solution that

allows you to replicate to the cloud so exact same solution that you have on Prem you can have an AWS Dhruva e is a born in the cloud backup solution that's focused on the end users they have their original solution it was about 75 has a customer of 75,000 dot backup mean 75,000 in users so think about it traditional laptop backups how do I get that user specific data the OSTs the word documents that critical database on the end-users device into a backup solution that was

their foray into the industry they sense including on expanding into the data center and now I'm backing up VMs specifically vmware vsphere de dos aisle data file is pretty interesting solution and since there's that you know what they they make no qualms about it they are focused on on on relational databases the hadoop sand Cassandra's of the world how do you back up these distributed databases that may be on pram off pram and a cloud database focused backup Kohi City cohesive is an interesting solution I

thought talk to raphson River the field CTO of cohesive they don't consider themselves a backup solution they consider themselves a data management platform they are that you can use them natively to backup your environment so again traditional plow cloud or on-premise workloads data is backed up to a cohesive virtual storage device and you can use cohesive software to back it up natively so you can use cohesive as the interface or it can be a target for a product like convo so a little bit best of

both worlds a little bit you can use it as both a new school cloud-based solution no a bit you still have challenges with concepts such as service but you can back your ec2 instances up to cohesive or you can backup your traditional vm's on pram to cohesive the either on pram and replicate that out to cloud-based instances rubric is very similar in cohesion it's a scale out solution you can add bricks to the cohesive platform to expand me sorry to the rubric platform to expand capacity

but rubric is not claiming to be secondary storage platform in a way that cohesive is there's a lot of confusion between Kohi ste and rubric from my view rubric is focused laser focus on backup as this initial use case while you can use it for some of the replication and dr scenarios it's specifically focused on that backup part of that protecting data on pram in the cloud and they are not a backup target for a secondary piece of software so they have no partnership with convo

for example example to be a target a storage NFS target for rubric and then you have the father of this kind of whole market of all the of convo cohesive the rubric is data domain dil EMC is not sitting back and letting a cohesive rule rubric and call eat all of this lunch it has also launched a cloud-based data domain to do replication up to data domain up to the AWS cloud they have some work to do that will just leave it at that it there

is some work to do then a non-traditional backup vendor quote-unquote backup but data protection vendor de trium which is a open converged system they don't like to I like I like to sometimes call them HCI they don't like that name they are open converged system that are that's very similar in and substance to HCI their scale out x86 they have the concept of data nodes nodes and storage nodes that you can mix and max for performance and capacity and they recently came out with their cloud-based

dvx de triomphe has data domain routes and that a lot of the executives and developers come from data domain so they've inherently put a lot of that capability in there open converged platform doesn't go down to the cloud level for black backups so you you would still need a backup solution but they replicate out to AWS you can have a cloud-based dvx system they're open commercial system in the cloud and have that as a target of your arm Prem dvx open converged system to replicate data

to very interesting solution then of course we can't forget the gorilla in this whole VM data center space V V is a well-known solution VM based they integrate well with both VMware and hyper-v they've gone into cloud-based platforms V is a solid privately owned backup focused entity that provides a VM focus backup capability all these companies at varying layers of of degree offer the capabilities we saw in the previous slides when it came to defeating data gravity or trying to beat the speed of light you

can't beat the speed of light but each one of these companies offer some form a copy data some layer caching caching might be if he was some of the newer born in the cloud solutions definitely replication and every one of them support object storage to some degree I highly recommend you engaging your sales team for each one of these organizations if you want to find out more about them so as promised we finished within a half an hour let's look at if there's any questions from

the crowd so you deem ported along DeLeon I love I just I don't only say names within webinar but that's a cool name so he asks what is open converged there's actually if you go to the visit the CTO advisors YouTube page I have an interview with the date um folks in which we we talk about open converged it's really interestingly enough as far as I can tell date um was the first person company to use the term open converged however I've seen it otherwise but

if you just search for date you'll CTO advisor on there you'll see my interview with datums CTO you go Patterson and we talked through the whole concept of I challenge I'm pretty good on what is the concept to open converged a question do you mention EMC DD and not HP what HPE store lunch once you know what that's a fair that's a fair question HPE store lunch once is another option out on the field so blatantly all of your major from the vendors I didn't mention

everyone and my apologies for HPE for not including them on a slide HPE store wants it is definitely a software-based solution that has some compelling solutions thanks for that call out any any other questions you all right with that said I appreciate you guys spending time stopping off oh the question was does data domain have any extra software advantages you know what I'm going to hold my tongue a little bit on the data domain question I think one of the knocks against data domain is that

it has been the leader in the space for so long they revolutionized this concept of scale out this base storage they are still the market leader would I say they have software advantages I would say from the view of what of of the criteria laid out later laid out I think cohesively root and rubric is a little bit ahead of data domain and specifically I think that's a good challenging question 2x your dill EMC rep of what are the event why should I still stay with

data domain versus one of these new solutions from a technology perspective of course they're gonna give you the spiel about the power of the dill brand and the dill EMC support capability that's all well and good but ask about the innovation within the software why did why did rubric when best of show at vmworld and data domain didn't that is I think a great pointed question 2x well good challenging question all right any other questions all right and really appreciate the time I appreciate the questions

this will be available for the recording will be available so if you need to come back reference it I will send out a link to the SlideShare also uh I fail to mention or fill the note on the vendor slide that they trim is also a c2 advisor client so you know just to be full disclosure thanks a lot have a great weekend and a great holiday if I don't talk to you guys on Twitter you