Data protection and big data - CTO Advisor Podcast 093

How does Big Data impact your data protection strategy? Just as important, how does your data protection strategy fit into your data management strategy. Keith has mentioned the concept that compute has an opposing gravity to data. How does compute’s gravity impact big data? Mark and Keith have on the CTO of startup Imanis Data to discuss why the data protection segment has seen so much investment. What are investors seeing that someone who hasn’t looked at their backup strategy in two years doesn’t? The CTO Advisor Data protection and big data - CTO Advisor Podcast 093 Play Episode Pause Episode 1x 00:00 / Subscribe Share Apple Podcasts Spotify RSS Feed Share Link Embed <blockquote class="wp-embedded-content" data-secret="BtRbt3Vv12"><a href="https://thectoadvisor.com/podcasts/data-protection-and-big-data-cto-advisor-podcast-093/">Data protection and big data &#8211; CTO Advisor Podcast 093</a></blockquote><iframe sandbox="allow-scripts" security="restricted" src="https://thectoadvisor.com/podcasts/data-protection-and-big-data-cto-advisor-podcast-093/embed/#?secret=BtRbt3Vv12" width="500" height="350" title="&#8220;Data protection and big data &#8211; CTO Advisor Podcast 093&#8221; &#8212; The CTO Advisor" data-secret="BtRbt3Vv12" frameborder="0" marginwidth="0" marginheight="0" scrolling="no" class="wp-embedded-content"></iframe><script> /*! This file is auto-generated */ !function(d,l){"use strict";l.querySelector&&d.addEventListener&&"undefined"!=typeof URL&&(d.wp=d.wp||{

Transcript 4,578 words · about 31 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Hey, how's it going? It's Keith Townsend and you're listening to the CTO Advisor Podcast. Mark, how's it going, man? It's going pretty good. I just turned 40. So, you know, I'm getting I'm officially an old man. You are. Last time we talked, I was 39. Yeah, that was your you're officially grown. I am. So we just actually we just recorded a podcast earlier today with Kelsey Hightower. And I'm hoping that we can salvage that episode.

It's really great conversation. We've been getting some really interesting industry guests on the phone. That was a good one with, you know, it's always interesting with Kelsey. Can't go wrong. But it's been a busy week with podcast for us. It has. We were at IndieVMUG and we got five episodes of what I think we're going to leave all five of maybe four of the five. Yeah. As virtualized geek episodes. com and check those out by time.

Plus Amy's podcast. That was fantastic. All right. So today we're going to talk data management. I've been looking forward to this one we have. Hari, what's what's your title officially over at Imanis? I'm the CTO and founder at Imanis. Hari. And I don't want to butcher your last name. I know you told me. How do you pronounce your last name? It's Hari Mankade. Mankade. That's that's easy enough. So, Hari, the the.

I think this is a topic and I want to direct this to Mark, too. Data management in general is a hot mark. I mean, you look at the subset data protection. It's it's it's on fire. It is. And it's it's challenging. Right. Because you've got scale. You've got protecting the data. You've got recovering the data. R. O. and storage technologies. I mean, data is has been growing forever. And now the technology is pivoting so much.

It's even harder than ever. It's something enterprises really struggle with today. I know I felt that, Keith, I think you have to. And, you know, one of your former roles. And I think one of the things that we were anxious to get you guys on to talk about. Obviously, there's a lot of activity in a space. So, you know, 200, I mean, a 90 million dollar investment here, 250 million dollar investment there, billion dollar unicorns. What what did I mean?

There's a lot of solutions on a market today. What do you guys see as a challenge that you're thinking to yourself? You know what? We're going to start yet another data management, data protection company. What paint the big picture for us? Where's the gap? Sure, sure. That's a that's an awesome question to start off the podcast. So data management, in my opinion, is a lot more than data protection. Data protection is, of course, one of the things like backup and recovery, point in time recovery, et cetera, is a is a critical component, component of data management.

But data management is more a holistic, a 360 degree view of the data in the data center. So by this, what I mean is it's it's it's it's like a cluster or a platform that in that is put up in the data center, which can pull the data from your production clusters and for for for backups or and also kind of seed this data to to a test cluster or a research cluster or a dev cluster. And also kind of the test and the dev and research clusters might be in the cloud also, because people don't have sufficient resources on their on prem data centers.

So essentially, it's a it's it's it's a platform that manages the data. It does orchestration of the data rather than just doing the backup and recovery per se. So as we look at, you know, kind of the larger picture of data management, let's let's talk through why this is even important. I just recently got finished with a research. Report on data protection, which is, as you mentioned, a subset of data, data management data is the new oil, blah, blah, blah, asset class and all the reality of it.

The reality is that. Enterprises are trying to leverage this new asset class of data and looking at, you know, what it takes to asset access or take advantage of this class, you need net new sets of capability or scaling of capabilities that you can't do that you can't get in today's products. Mark, what are some of the practical challenges that you've seen outside of data protection when it comes to data management and taking advantage of the explosion of data in mining of that data?

Well, I think the biggest challenge people face operationally, not the biggest, but popular is ransomware. Right. We see a ton of that inside an enterprise organization and protecting yourself against that. And the other one, I think, is cloud migration. And that often is coupled with test dev. Right. You know, I see a lot of people trying to spin up cloud external in the public cloud for for test dev instances and moving and migrating and owning all that data is the biggest challenge and protecting it.

I mean, honestly, it's it's not easy. Yeah. And so there's two components of it. It's like data mobility and moving data to the cloud. And then once you get it there, protecting it. But I don't think people really understand the scale or not. Everyone, I think, gets the scale of that challenge. When I think part of that is, you know, the cloud washing that's going on in the industry is you move to the cloud and it's protected magically and it's accessible from everywhere.

We know that those you know, those aren't based in reality. And I think one of the practical problems is and I love the point that you mean that you made that, you know, once you get to the cloud, it's accessible everywhere. Like a real practical problem is let's say that you're you've been a legacy client of AWS in the easy to instance region of what of Virginia, that east region. And AWS comes out with a new machine learning service that is originally only available in Oregon.

So you have this real data locality challenge of the data is in the east coast of the US and the compute that I want to run it against is in the West Coast. Not only do I need to move the data, but sometimes it's just too big to move. And I have to identify, you know, kind of what parts of the data should be moved. And Harry, I want to go back to you and ask the question. Like, how do you see that challenge being practically solved that data mobility problem?

Yeah, so that's yeah, that's actually a biggest problem. One of the bigger problems that enterprises face these days, especially when it comes to big data, things like Hadoop or Cassandra or Mongo and so on, because the sheer volume of the data is orders of magnitude more than that you would find in Oracle database or MySQL database. So when it comes to it, a data management platform there would need to be able to say down sample the data in a meaningful way before it can be pushed to to a different location.

There could also be the other angle that you brought up. The security angle is also very, very critical. And we have a lot of customers who are mortally afraid of exposing their data in the cloud. And especially in the QA and research environments, these environments are not as secure as your production environments. So what they want is from a data management platform to be able to mask personally identifiable information, things like credit card numbers, social security numbers that is present on the production clusters before that can be pushed to their QA clusters, because the QA clusters are generally in the cloud and they're not as secure as their production clusters.

So these are some of the basic requirements that a data management platform need to kind of incorporate. Again, other than the backup and recovery. So other thing that also frequently comes up is the retention angle, because the QA clusters, people create QA clusters, deploy data and then go away, forget about it. And they want the data management platform to be able to kind of manage that data and delete the data after, say, a month and so on. So there are a lot of these different aspects regarding data flow that a data management platform needs to kind of address.

Somehow in an architectural perspective, it kind of looks like the data management cluster was like the hub for the data flow and all these different clusters, other clusters like the production, et cetera, production, QA, research clusters are all the spokes. And the data management platform orchestrates the flow of data from one cluster to another cluster. So, Mark, as I look at the kind of the landscape of what's out there, all of the new, quote unquote, data protection companies seem to identify these core challenges with data management.

Yeah, automated test data management and just copy data management and masking are things everybody's facing. And there's a ton of companies that identify and are trying to solve that problem. And have you noticed the trend that primarily is data protection companies that have gone after that market of saying, hey, you know what? We are the base for your data management platform. Yep. And, you know, what I find interesting about that is that's typically a sales pitch to the infrastructure team.

Right. Essentially is what that is. But I always struggle with how do you do a sales pitch to the infrastructure team, your backup team that really benefits your development team? I find that to be pretty challenging. That is the case because, you know, if it's just a pure and this is a challenge of silos. If it's just a pure case of backup, there's plenty of backup solutions on a market that backs data up. But the value, Terry alluded to, isn't in being able to back up and restore the data.

It's all the advanced data management pieces. So one of the things is that talk to your new CMO, Peter Smails. One of the things that intrigued me about Imanis data is that you guys have this machine learning, AI slant. Like what is the, you know, machine learning, AI? Everyone wants to have that those two buzzwords as part of their their product pitch it as part of their deck. You know, it's it's it's the it's the the term du jour. It's the new cloud.

Right. I mean, it's the new AI and ML is the new cloud when no one knows what it means. Or, you know, you know, I like to use the term cloud washing. I think we're in an era where a lot of people are calling, you know, rejects and if then's ML and AI sometimes. So where does that differentiate for you, Ari? Yeah. So, again, because it's a from a data management architecture perspective, because it's a central hub, it has great visibility on the different things that are happening from a data perspective in the entire data center.

Right. So, for example, one of the things that we have done at Imanis, the basic thing is we have built a data model based on the change rate that we would see on the on a customer table. So this is based on machine learning, because we because when we do a backup from, say, Cassandra table, we know what the change rate is throughout the month. So, for example, in the beginning of the month, there is a lot of data that is ingested into the stable.

And then then maybe Mondays, the table is very active. A lot of updates are happening. But as the week progresses, the rate of update goes down. Maybe Friday evening, it's not very high. And Saturday mornings, again, there is a lot of data is bulk loaded into the table. So we have a very good idea of the entire pattern of the change. And we have built, for example, a machine learning model based on that. And if there is any anomaly, then we can use this model to kind of flag this anomaly and tell the user that, hey, there is something weird happening on your cluster.

So it could be potentially a ransomware attack on your table. Right. So somebody might be encrypting all the data there. And this is one of the places where we are deploying the machine learning to kind of make it more intelligent than just assuming that it's data at rest. Similarly, there could be other scenarios where we could kind of change the RPO or we can use the customer's business requirements to change the policies dynamically. So instead of doing a backup every hour, we know that the change rate is very high on Monday mornings.

And instead of doing it every hour, we do the backup every 15 minutes so that we can meet their business SLAs. So, Mark, I guess I'm not a storage expert. And there you are. Come on. There is a bunch of of of recent product announcements. VMware is parent company. I'm now VMware employees. I have to disclose that Dell EMC came out with their power max, which has, you know, I don't know, some crazy number like twenty thousand machine learning.

It makes a day. Yeah. From a practical perspective, does that appeal to you? Storage types that the the backup dynamically changes based on business requirements or does that create unknown problems for you? I think, you know, a little from a little from comedy. Right. One of the things that as a as a hardcore former storage person, consistency is always very important. Right. You want consistency and balance. Knowing your workload is important so that you can make sure your workload can function and meet your business requirements.

But at the same time, we're seeing those business requirements shift, you know, like like never before. And we're seeing our data protection and our security requirements go along with that. So I think we have to adapt to doing things in those new ways. m. kickoff backup just doesn't work anymore. Our data changes too much. It's too important. It's too too much at stake for most enterprises to deal with that. That's what it really can appeal to us.

So. So, Harry, Harry, is that something that you guys also consider? Like, let's say the rate of change is so great that, you know, normally in an eight hour backup window, that traditional sense that I can get the entire backup in that in that window. And for whatever we had an outlier, we you know, we refreshed our environment and and whatever the case, the weather's deep or whatever. It's just too much net new data for mission critical workload. And it won't happen in a eight hour period.

Can you guys like predict that for us? Yeah. Yes, definitely. We can. So if it's if it's like a pattern that happens, say, for example, every Saturday where new data is ingested into the database. And then do you then a data management platform, a intelligent data management platform can easily kind of predict that, OK, this is going to happen every Saturday. This has been happening every Saturday and is going to continue going forward. And then take some corrective measures to ensure that the SLAs are met.

So it could be that you do the backup more frequently or dedicate more resources to the backup, more computer resources, more network resources, more storage resources to ensure that the business requirements are met. So, Mark, one of the challenges I have with. I think this whole space is. The rate of change. In products itself. Back up when you look at, you know, backup as a subset of data protection, which is a sub, which is yet another subset of data management.

Backup is something that, you know, I just simply didn't change much, whether it's, you know, going from one version of a software product or, God forbid, going from one software product to a completely different software product. I simply did not. Change my backup infrastructure that that much. Well, you got stuck in the trap of thinking that just because you were doing backup one way that it was good enough. Right. Right. That's that's a common trap that we fall into in enterprise, not just in and backup.

We think we think that it's working. Don't fix it. Ain't broke. Don't fix it. But in reality, if you're not fixing it, what that really means is you're not moving forward. You're not keeping a pace with everything else that is changing around you. So is that an indication? So if I looked up and I'm running ArcServe six point two, I literally don't know what version ArcServe is on. I think that's the last version I use, which dates me.

And I'm getting successful backups. You know, it's happening for whatever reason. The solution still backs up data. Am I looking at the wrong metric? Should I be worried if I haven't really let's talk about a serious scenario like in the past two or three years, if I haven't revisited or upgraded my backup solution in two or three years, am I putting myself at like an unknown risk? I would say you definitely are, because I would almost guarantee someone in your company is, you know, deploying no SQL and you're probably not backing that up safely or you've got data in an AWS or an Azure and you're not backing that up safely.

And that's just from a strict backup perspective. Right. You've got all these new no SQL style databases that are coming out of nowhere and backing those up isn't the same as backing up your server. Right. It's just it just works differently. Same with any type of data lifestyle. Backing that up and protecting that data is more complex than just throwing an agent on it and telling it to back it up. So let's listen for both of you. I'm going to peel back the onion on this.

It's more complex. I have a file system. I have a database on that file system. I back up that file system. Worst case scenario, I restore that file system and do a consistency check on the database. I'm back up and running. Life is good. What am I missing from that scenario? Well, the thing is that might work when the data sizes are very small. We're talking about maybe hundreds of gigabytes, maybe. But if you're if for like the like the no SQL databases that Mark mentioned, which run easily in hundreds of terabytes, just backing up the file system.

And it's not just a single file system. It's across. It's a horizontal cluster. Scale out cluster. Now, getting a consistent point in a horizontal cluster is also non-trivial and backing up hundreds of terabytes in some sort of a full backup and incremental mode. So the traditional backups happen to be I do a full backup, maybe two or three times a week and do an incremental in between and so on. But doing full backups of hundreds of terabytes, two or three times a week is just not.

We'll bring down the data center, essentially the networking aspect of the data center. So let's say that I'm going to throw a scenario out there and you guys help, you know, lead me to water in this sense. So I have, let's say, a research database that's distributed across four or five nodes across. And let's say that it's in the same. Let's keep it simple. Send the same data center. So I could, in theory, you know, point it to a disk based backup system, get a full backup.

But what you're telling me, what I'm hearing is the complexity is I have what's essentially a distributed database. I have the individual file systems. But what I don't have is that metadata needed to reconstruct the database or even get granular restores. That may be involving two or three nodes in this four or five node cluster. I wouldn't be able to have that context to reach to perform a granular restore with traditional solutions. Exactly, exactly. So it's the restore has to be granular also because the user could have just dropped the table and restoring the entire hundred terabytes doesn't make sense.

And not to mention just the technical limitations, right? Your file system is going to have to support journaling. That has to be disabled. We're not talking, you know, a dump and sweep backup here that you might do on your Oracle database. Imagine trying to dump and sweep your Hadoop environment. Exactly. You've got this massive pool of data and that just doesn't work anymore. And not to mention a lot of organizations, because they have this distributed database now, for whatever reason, feel like that protects the data.

They don't think about data corruption. So I've seen organizations make a choice. Oh, I don't need to back up my Hadoop data because it's on these 12 distributed systems. Yeah. But in reality, you've got to protect against data corruption just as you would on any other. Exactly. That's a common fallacy, especially in the world of big data, because not only it's distributed, people assume that most of the big data systems have three replicas. So it's like having a disk with three mirrors and automatically people assume that it's three replicas and it's distributed.

There is no way it can be corrupted. You could have a single admin who drops a table and doesn't matter how many replicas are present. All the replicas will get deleted and the mirror gets deleted. A. as two different things, which we're used to in the enterprise. We think about data protection and backup. , we think about redundant servers. So in any, you know, no sequel, big data architecture, you still have to think about both of those things.

And then there's the practical sense of DR, you know, chances are, even if you, you know, let's say that you're in three different AWS regions and from a DR perspective, there is still a concept that what happens if AWS becomes unavailable? Highly unlikely, but possible. But we know software bugs happen, right? I think we've all experienced the software bugs take something down that shouldn't go down. Hmm. So that's happened to all of us in our careers. A. setup.

A. isn't DR in most scenarios. A. isn't DR. If I lose a site and I need to recover a database, the the the tricks of recovering a no sequel database, as you guys have mentioned, is very different than recovering a traditional Oracle or SQL SQL database. And then the other thing, just to briefly touch on it, is that data mobility. Right. More and more. I think we've had this conversation on the podcast before. Keith is data gravity versus compute gravity.

Right. And deciding what to move and when to move. That's a real challenge that I think everyone who's got a large pool of data that is in two locations is facing today. Or will face, let's say, start their cloud native journey. You know, you know, it's one of those complex things. We can't move it all. We can't move all the data. We have to identify what data needs to be moved and how best to move that, because even once we identify the subset, that subset still could be terabytes of data.

And moving terabytes of data over a network is still a lot. We can't change physics yet. Can't change physics. So, Harry, we're going to give you the last word. Can you plug a minus data? What do you guys do and how can people find you and kind of what's the elevator pitch? Sure. Minus Data is a data management company that was built ground up to solve the data management requirements for big data. So we have, we solve the backup and recovery aspects of platforms like Hadoop, Hive, HBase, HDFS, then Cassandra, Mongo, Couchbase, Vertica, and many other NoSQL databases.

And we also do things like replication. We solve the sampling and masking aspects. We have machine learning components built into our platform so that we are, we can identify things like ransomware attacks and so on. And similarly, we are completely, we are cloud native. So we could be deployed in the cloud. We could be used for cloud migration where people migrate data from on-prem to cloud on a regular basis or it's a one-time migration. We work very well with cloud native Hadoop deployments, things like HDI, HDInsights in Azure or EMR clusters.

And we have hierarchical storage management built into our storage infrastructure so that we can move the data between block and object storage. And similarly, we have things like deduplication, a data-aware deduplication engine also built into a platform so that we understand the data, the different data formats that are coming into our platform and the deduplication algorithm is adaptive and changes automatically based on the data it's working on so that we get the best results from a storage reduction perspective. So we also, we use erasure coding for data durability so that we don't have to keep three copies.

And I-Manus is actually a cluster that can scale from three nodes all the way up to thousands of nodes. All right. And where can folks find you guys on the web? com. All right. We really appreciate you taking out some time to talk to us about data management. You can find my good friend, Mark May, on the Twitters. You can. At SensiStorch. You can find his blog at VirtualStorchZone. That's right. com. com. I actually went there the other day.

I want to see if you were still active. It's been a while, but there's a poster. One, two, maybe. com. You can find the podcast in your favorite pod catcher. Share it with your friends. They will enjoy talking about big data, Hadoop, NoSQL and all of the above, even if they are auto mechanic. It applies across the board. Seriously, until then, talk to you next CTO dose. Thanks.