CTO Advisor 062 - What's so exciting about data protection w/Druva

Data protection startup Druva is coming fresh off a $80M investment. Another data protection startup won Best of VMworld. What is it about this market that has investors and judges excited? What’s more to data protection than backing data up to some tapeless media. Hasn’t products such as Datadomain provided all the innovation needed in this space? In this sponsored podcast, Druva’s CMO Matthew Morgan joins the CTO Advisor live from VMworld 2017 in Las Vegas. Matt sheds light on this growing market. Keith and Mark ask the who, what and how of not just the market but the Druva platform. What’s so exciting about backup? Subscribe iTunes | RSS

Transcript 3,714 words · about 25 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Alright, welcome to VMworld 2017. The CTO Advisor from the show floor. This is pretty exciting, Mark. It is. It's the first one from VMworld live. Yeah, so we've done a bunch of Dell EMC Worlds. We have. And this is, hopefully this goes a little bit better than the Dell EMC Worlds podcast. Yes, hopefully. No hardware failures this time. No hardware failures this time. The bought an expensive SD card. Yes, yes. So this is a special sponsored podcast with Druva.

For those of you who don't know who Druva is, this is a really expanding market. Data protection. Not just data protection, but secondary data. We're going to get a great overview of the various use cases and where this fits, but very exciting segment, Mark. Yeah, I agree. You look at the showroom floor and there's a lot around this. That's really exciting, right? Data protection is changing, right? It's not what it used to be. You know, it's different.

So why don't you introduce yourself to our audience? Yeah, good morning, good afternoon, everyone. My name is Matthew Morgan. I'm the Chief Marketing Officer at Druva. It's really great to be here with you. We're glad to have you. Yeah, thanks, man. First off, huge funding round. Congratulations. What is it? What, like $90 million or something? It's a crazy number. Yeah, you know, Druva's been very, very fortunate. We've been pioneering a new approach to data protection.

It leverages an as-a-service model, leverages the cloud, and it's really caught fire. Total funding to date for the company actually is around $200 million, but this latest round is the biggest we've ever done. So, overall, the market has seen data protection in the cloud. It got introduced kind of as a consumer product. That's right. But that's not the case with the products that we've seen on the show floor. Many of the products take it to the enterprise level, add enterprise features to it.

As a matter of fact, I think a data protection solution won best of VMworld. They did. They did. So what's the excitement around this market? Why so much investor excitement, so much excitement on the show floor? What's the deal? Yeah, I think there's been this untold secret. Legacy data protection systems don't work in the cloud world. A lot of people say they don't work at all, but I'm really here to articulate they don't work in the cloud world.

And the reason why is that basically they were designed around building and securing data within a data center, and it's been designed around a really long, complicated process. I've got to buy software. I've got to buy servers. I've got hard drives that fill up. I've got to add more. Oh, wait, I've got to refresh those drives. I can't afford to buy all those drives. I'll buy a tape drive. Oh, wait, I can dedupe some of that. I'll buy a dedupe thingamajig.

And then I need to make sure that if my building burns down, I can still have data. I'll buy or rent a truck and drive the truck. I mean, this whole chain of pain, if you will, around data protection is the way it's been done. But the problem comes into play when advances like VMware cloud on AWS come to market. Okay, so let's look at what VMware said at this show. There's going to be physical servers. Yes. There's going to be virtualized infrastructure.

Yes. Well, now, increasingly, there's going to be VMware cloud on AWS. This is a hybrid reality. The data is not going to be anywhere near this chain of pain. So I can't take a truck and just back it up to AWS's data center and pull out my tapes? It does. Exactly. And the reality is a lot of people think like, oh, the next generation of data protection is some hyper-converged appliance thing. That doesn't work either. The reality is that you need to back up cloud to cloud.

And this is where Druva comes in. No one else is really having this conversation, backing up cloud to cloud, being in AWS as a service to wrap around your virtual machines, back up and secure that data within the same cloud. Druva's also been able to pioneer some real advances around the data itself because when it's in the cloud, it's centralized versus the siloed hell that we've all built in our data centers of hard drive after hard drive after appliance after appliance where you can't see across that data.

It's isolated to the silo. So let's dig into that in a second. But I think there's a myth or at least a perception. It's the cloud. Why do I need backup in the cloud? It's the cloud. Great point. I think you think that the cloud provider is backing up your data. They don't back up the data like you would own their data. They protect themselves. But if I have a file in S3 and I delete it, can I just call AWS and say restore it from this date?

Yeah, I think you could try, but I don't think you're going to get anything from that. To be totally honest, this is a point of view and a problem that people don't even truly understand. People get confused between a service level agreement and what data protection is. So I'll give you a great example. com. Great application. Everybody in the world uses it. It's in the cloud. They have the no software with their logo, and they were really rolling things out.

All right. If some nefarious employee goes in and blows away 1,200, 1,500, 2,000 records, they're gone. They're gone. Right. If you don't have a data protection strategy, you're not getting that information back, right? So that is super scary, the thought of that. One, we've gotten past this apprehension of putting data in the cloud because of security. We understand that cloud companies do security better than we can. But there's a sense of protection when we physically have something.

That's right. We can put a lock and key if an employee leaves. I can just deactivate his badge and he can't come back in. That's right. So this concept that I've abstracted away the security from that physical sense, how does Druva help bring back that sense of security? Okay. So first off, let's just talk about the compatibility of the solution. If you're using Office 365 or G Suite or I've got a Salesforce or Box folders, right, those all live and breathe in the cloud.

So we wrap those with an as-a-service solution. VMware Cloud on AWS will wrap those virtual instances with an as-a-service solution. No hardware. We'll just attach it to the service. Our service is called Druva Cloud Platform. We're going to back up all of that data. That data then gets cataloged and reduplicated many times over within the confines of the cloud system. Our system has a global dedupe, so we're trying to reduce your footprint so you don't have this spiraling S3 cost you can't control.

We make use of Amazon Glacier so we can bring that cost down for storage that, frankly, you're not going to need access to but you still want to know you have. So that's the compatibility conversation. The conversation about how do we add value comes into play of how we actually give you access to that data. Anybody, anywhere within an organization that has the credentials can log into a single pane of glass and see all of that information, all of it. Even if I'm collecting it from 100, 1,000 endpoints, 100, 1,000 servers, you've got one contiguous blob of data that you're able to actually see and work with.

This allows you to do things you can't do with legacy approaches like governing that data, be able to understand the chain of custody, manage it for legal compliance, deal with legal issues within your organization. And we also do intelligence on that data, use machine learning around the metadata that surrounds the data. Machine learning is incredibly helpful. I'll give you a great example where it controls risk. The last three major cyberattacks have been something called ransomware. Ransomware targets the data. The last one actually didn't give you any option to get your data back.

It said, hey, pay your Bitcoin, but they never bothered to send you the key. So it's more like whiteware. The point is we can see that happening early because we see the metadata, and we've already secured that data. And there's no way ransomware is going to crawl into the Drupal store. There's too much security around that, and the way it's stored isn't compatible with any file system ransomware of the past. So you're able to protect it across the board and get access and do things with that information.

So, Mark, one of the things that caught my ear, as Matt was talking about, was this global catalog, this immutable, basically immutable catalog. Immutable backup protection to protect me from CryptoLocker or whatever the next ransomware is. And not just ransomware. So a pain in my rear that has raised its head throughout my career is something as simple in concept as a legal hold. So I'm legally obligated to hold this data. However, I have backup tapes that span seven years. How do I know?

You've got to find it. Yeah, how am I going to find it? You've got to find it, and you have to still be using something that can actually read that LTO2 tape ten years later. And you're not. You're not going to find something that you probably have that can read that tape anymore. It's a big problem. So something like that, Matt. How do you guys help? Conceptually, I think I know how to do it. But practically, how do you guys help with something like a legal hold?

Oh, that's a great use case. Log into the system, select the user, hit legal hold. You're done. That's pretty easy. Are you sure it's that easy? It's that simple. It really is that simple. We're going to go through and isolate that data. We're going to provide confirmation that data is protected. It will be then placed in a store. You'll be able to use tools against that data if you need to, if you need to. So another point that I wanted to hit on that you talked about was this global de-do.

One of the challenges when you talk about cloud backup is this uncontrollable cost. I can, in theory, take my ArcServe backup application that I've had in my data center for 20 years, point it to a VM, and back it up over a VPN or a Direct Connect. However, I have data transfer costs. That data doesn't necessarily de-do on the other side. How do you guys help alleviate? How does global de-do? Let's dig into that a little bit. What does it mean to de-do the data globally and then back it up?

Yeah, that's a great point. All right, so I look at any IT professional that's been in the business for 10, 15, 20 years. They probably have their own home rigs, and they've been backing up their home PCs. And if you look across the last 20 years, they probably have a cabinet somewhere of hard drives that they've used to kind of store that data because it has gotten so big that they can't throw it away. Apply that to an enterprise with 10,000 employees plus.

If you try to use that approach, you're going to have not just storehouses. You're going to have warehouses of hard drives. Okay, we make that problem go away because we use this global de-duplication technology at the block level. We're going to be able to identify and ensure that only changes are backed up. This is going to reduce 80%, 90%, 95% of the volume on endpoints because a lot of the documents that are being shared with many employees, they're the same document.

You don't need 17 versions. But we will still maintain authenticity to the user. So you'll still be able to restore it. But we're not only going to copy it if we need to. This global de-dupe is a game changer because, number one, it's going to save you a ton of money. But, number two, it makes the cloud really fast. Because I'm backing up a terabyte, I may not actually be copying a terabyte to the cloud. I may be backing up 100 megabytes.

That's a totally different game in terms of your response time when dealing with wide area network connections. So before you back it up, you look to see if you have it already. That's right. At a block level. At a block level. And then you don't move it if you don't already have it. Actually, you have a great point. We don't copy it first, then de-dupe it. That kind of defeats a lot of the bandwidth advantages. We're going to evaluate it before the first bit is moved.

The fastest data transfer is the data you don't have to transfer. Absolutely. That's it. So we've talked a lot about data protection. However, my mind is starting to go places where the use cases can go beyond data protection. Well, any time you have a lot amount of data, you can do something with that data. Right. So how can we do something different other than protect it? You've got all this data. What else can I do with it?

Yeah, I have all this data. So let's take the multi-cloud scenario. I have my primary workloads in AWS. Azure comes out with this cool AI ML solution. Yep. How do I use this solution to enable taking advantage of that AI in Azure but leave my primary workloads and data set in AWS? So this is a great conversation. This is often the separation between primary data and secondary data. When you're in the primary world, the machine learning and the AI attributions of technology really come into play around business insights and intelligence or opportunities.

That's an area that is served by some really great people. You mentioned a couple. When you deal with the secondary storage world, it's about taking that data that would be under analysis, ensuring that it is protected, and then doing AI and machine learning on the metadata associated with that data. So the AI and machine learning that Dhruv is going to be focused on is on that side of the equation, allowing you to get further insights as to who's using the data, why they're using the data, what are the opportunities there.

Two different sides of the same coin, but it's an important distinction. So I think that's an interesting play because if I start thinking about data loss protection, I want to know who's accessing data. Say, for example, PII data. I want to be able to identify PII data, see who's accessing it. When people access it that they don't normally access it and things like that. So there's a lot of solutions around there. I can see a lot of opportunities there.

So talking about opportunities, a lot of solutions around it, let's talk about the ecosystem around Dhruv. Yeah, absolutely. Who do you guys have partnerships with? What are some of the things that your partners are really excited about when it comes to use cases that customers need to know about? Yeah, absolutely. So we have a platform. Dhruva Cloud Platform is an open and extensible platform. And we've got partnerships that run the gamut. You've got partnerships with hardware manufacturers that we can literally put what we call as cloud cache, which is our software appliance on-prem to help organizations get faster RTO and RPO times.

We also have partnerships at the cloud infrastructure layer. Obviously, our big one is with AWS, but we also have a partnership with Microsoft Azure. You can select any of those solutions to store your data if you use Dhruva. We have partnerships with extended value on top of secondary storage. So Extero is a great solution in the legal realm. We have partnerships that reach out into the machine area, like with a company that would do something along the lines of managing machine learning data.

com. So you can't get away from a big conference like this without talking about IoT, edge compute. A16z has said that edge computing is going to kill cloud computing. However, I still don't want to manage backups. That IoT data that I create on-premises has to get backed up somewhere, even if A16z is right. And more than just backed up, it has to be backed up on a different system than the system that it's stored on. Because otherwise, it's not really data protection.

If that system fails, your data is still gone. So you have to back it up externally. So for that edge use case, what does Dhruva play in the edge use case? So most of the edge use cases have some sort of central compute mechanism that processes the data. It's often in the cloud because it's closest to the edge. Typically, it's custom, so you see a lot of infrastructures and service routines. We can wrap around that infrastructure as a service solution and protect all of that information.

And since we manage all data types, all of it's 100% compatible. You're absolutely right. The data that's streaming off the sensors and cameras and all the IoT world is cool, critical information to a lot of AI projects. It's cool, critical information to business intelligence. It can be very, very important to protect. So we give our customers a variety of solutions on that. We have compatibility extensions to a various selection of data types that come off those types of equipment.

So, Mark, any closing questions? I think my biggest thing is going to be the thing I always like to ask is how do you actually consume it? How do I license it from an enterprise? How do I actually buy it? Have you worked with AWS? I assume you have. Of course. We have this credit philosophy. We have the identical model for the consumption. You pay for what you use, and that's it. So we designed it to model after AWS marketplace for the idea that you're used to consuming AWS in this fashion.

So you're going to have no learning curve when you move to Druva. What it means is you could buy six terabyte credits, and that can last you a long time if you don't use it. Or you could consume it quickly and kind of re-up yourself. What you won't be charged for is a fixed license and maintenance fee or any of the traditional approaches to software. And another thing you will not be charged for is access to some appliance. These converged appliance players will charge you $150,000 or more just for the box or get the service.

We don't require that. If you want to put something on site, we'll give you a software appliance, use commodity hardware. You're going to get the same benefits, but you don't have the cost factor. We're still going to have a credit conversation about protecting servers. So it's a per-terabyte credit license. Yes, per-terabyte. Is that pre-backup or post-de-do? No, it's all post-de-do. Okay. So we are returning the savings to you. We don't keep a dime of it, which is phenomenal because you know how Amazon can creep up on you, right?

Yeah, yeah. Oh, my God. You've got this giant bill and sticker shock. Yes, you're absolutely right. So we're trying to give you protection on that, too. So speaking of how much it is, how do I get it? So is this an appliance I can go to the AWS Marketplace and get, or do I have to contact the channel? How do I engage your sales team? Yeah, no problem. com. You can actually download a free trial agent to get started for endpoints, for servers, anything you need.

All right. With that, we'll close out. Mark, any last comments for Druva? Recommendations? You know, I think the biggest thing is understanding that when you put a service, put software in the cloud or put a service, your faith in a cloud provider, their definition of backup is vastly different than, oops, my executive accidentally deleted something. Right. I have users that have built-in dedup. They delete their own data. No, they have data on the H drive, and then they have redundant data on the G drive.

It's the same data. If they delete everything on the H drive, they dedup. Exactly. But I think it's key that in the data protection world, you don't put your protected data on the same thing that is the source of the data. And you can't get much different than that than multi-cloud or cloud in general. Exactly. So with that, Matt, do you tweet or LinkedIn? How do people find you? Yeah, my Twitter handle is AtForwardTension, and I'm on LinkedIn and all the social networks.

You can find me if you look for me. Of course, you can find Druva on Twitter. And Druva is very active on Twitter, actually. And you know me. I'm on Twitter. Yes, I'm on Twitter. What about you? Are you on Twitter, Keith? So you know what? Druva brought with him a fancy pair of socks. Nice. You can find me and my socks on CTO Daily Dose on Twitter. And then, of course, at CTO Advisor, the podcast itself is at The CTO Advisor.

So if you have feedback for this session, you want to get in contact to Druva via us at The CTO Advisor, talk to you guys next episode of The CTO Advisor.