CTO Advisor 058 - Rubrik's Chief Technologist Chris Wahl
Rubrik is the surprise unicorn that focuses on data protection. They’ve made many waves by challenging DellEMC Datadomain’s dominance in disk-based backup. Rubrik Chief Technologist, Chris Wahl is a frequent speaker on the VMUG circuit. The Rubrik marketing team has quietly hired a all-star list of influencers to work in tech marketing. Mark challenges Chris on Rubrik’s role in the cloud powered DC. In classic Chris-form, Chris makes a strong argument for Rubrik’s approach to data protection. Keith also trys the push the limits of Rubrik’s API with crazy use cases for cloud native applications using secondary data. Show Notes Wahl Network Datanauts Podcast Subscribe iTunes | RSS
Transcript
All right, welcome to yet another exciting episode of the CTO Advisor. I'd like to give some clever Star Trek reference because, you know, we have Chris Wall on today and he's always doing these clever, him and Ethan Banks are always doing these clever back and forth about Lieutenant Kirk. I don't know who these Star Trek people are, but they're amazing. Did you just demote Kirk to lieutenant? Yes, I did. On purpose, actually. So who is Captain Ahura? Yes, absolutely.
Yes. We'll we'll just leave it at that. All right. So obviously we have the well-known Chris Wall on the podcast. This is take two. Me and Chris tried to record this podcast about three months ago before Mark came on. I kind of added some discipline to the show. Chris, for those who don't know you, can you give a brief introduction? I'll try. So I'm Chris Wall, the chief technologist at Rubrik, a company out in Palo Alto focused on data protection and recovery and archive and all that jazz.
Personally, I'm at Chris Wall on Twitter. com that's been around for about seven years now and co-host a podcast with my friend Ethan Banks called The Datanauts, where we talk about silos and how to destroy them in enterprise IT. So unlike this podcast, the Datanauts is super geeky. You guys get like I just listened to one about I. What is the IOT? Is that the term they use? I don't know. We were not supposed to. Yeah, I get the joke.
You obviously did listen to the show. The Internet of Things. Yes. The Internet of Things. It was a great. Although that one actually wasn't yet. No, no. Actually, it was super technical because Ethan got into like the different types of low power when that you can use. And I'm like, oh, my God, networking. Someone shoot me. We need to talk about something more important. We're not going to talk about networking. We're going to talk about storage, Mark.
So yet again, you get another conversation to lead. You know me. I love talking about storage all the time. And I'm still pretty confident since you don't like networking, you must be a storage guy. No, actually, that's a joke. I actually really like networking. The Ethan's part when he even deep dive into all of that good stuff. That was some of the best material. I'm going to stand by the fact that you're a storage guy. So, Mark, take it away.
Why should we care about storage? You know, I'm not sure that we should care about storage as much as we should care about what actually has gravity. And that's data. I think data in most organizations is crucial to everything that we're going to be doing to making decisions, to looking at what we did right and looking at what we did wrong. Data, data has gravity, right? Data does have indeed have gravity. And I think one of the things that I'm really struggling with, not conceptually, but more practically is we're starting to see a movement where not just data has gravity, but compute has gravity.
I. or even cloud capability, compute cloud capability in my data center. Yeah, I think the cloud has given compute gravity, whereas before it was kind of the given. But now that we're talking about moving workloads, we have to understand that compute locality to data matters. T. guys, you know, we love our storage arrays. Our first mindset goes to, you know what, less leverage, traditional vendor equipment. Let's take, you know, our our tier one storage array, our V match or Hitachi HDS or pure storage.
Or what? It doesn't matter. And let's put two of them. Well, let's put multiple ones all throughout the physical locations and replicate data in between. I can't afford to do that, and that doesn't necessarily solve my data gravity problem. How do you guys at Rubrik view, you know, data and data gravity? That is that is the million dollar question, man, that's that's pretty loaded. I guess within the data center itself, you've got your primary systems that you've given examples of, you know, the VMAX and whatnot.
And those are great. O. and ultimately being responsible for making sure that the bits and the blocks that you write to it are there when your compute goes to access them and do interesting things. But there's been a few challenges in that arena over the over time, I'll say specifically when it comes to protecting that data. And the things that you reference, like dual storage arrays and redundancy and things that just really solves availability of the data to the data center.
But doesn't really have anything to do with protecting against, you know, anything that's applying encryption, you know, such as ransomware. It doesn't really solve any issues around making sure someone doesn't delete something or change it or revision history or making unstructured data searchable. And so that's really where we're focused is applying intelligence towards a solution that can first of first off, be aware of what all that data is, whether or not it's structured or not. And then catalog it, index it, make it searchable and put it in places that are advantageous to the business, which may be on prem using object store or, you know, file store, whatever.
Or more popularly, using all these giant blob and object stores that are in the cloud, such as S3 and Microsoft's Azure's blob store. Right. So those are all different ways that you can protect data, make it, you know, dedupe compressed all that jazz, but also searchable usable. You know, it can be put in a different place, et cetera, using software, which is kind of cool. Wait, so you're confusing me. The when I think of storage and I think of secondary storage, what you're describing sounds like secondary storage for the most part.
I've heard that term before, but that's a box. What you're describing kind of isn't a box. I mean, how do I put that box of, you know, secondary storage in the cloud? I mean, traditionally it was the PBBA, the purpose built backup appliance, which is just a I'm not using this in kind of a negative light, but dumb, quote, fingers box. That was a landing pad for ingesting data and then trying to dedupe it or compress it or fingerprint it, whatever.
And you're right, that's that's kind of birthed the secondary storage architecture term, which is basically any storage that's not running a primary workload. So certainly not anything new. You could even take you know, you could have a pair of net app arrays and one's running production. The other one's just backing up production. And then it's both a primary and a secondary storage array. So it's not a very useful term as far as putting that into the cloud. It's really just an idea of there's all these storage platforms out there.
If you can abstract the file object or block, you know, system underneath and make that relatively uninteresting, which which it already kind of has been progressing that way anyways. And really just treat the applications and the data that those applications generate as objects or things or are really just data as a general term. Then you can place that data anywhere you want. The key is that the PBBAs or the secondary storage arrays that exist are really, like I said, just kind of landing pads for data.
They aren't really aware of what they're doing or why they're doing it. They're just chugging along and trying to compute to figure out how to make that data efficient when it lands on these spindles with a flash. If you can make an approach where you understand what you're backing up, you know, you know that it's a Windows server running these three applications. And, you know, this is what the previous chain of data backups look like. Then you can very easily push them into public cloud.
And it all just looks like a logical namespace. So it's really, again, it kind of goes back to using software to abstract away all of the nuances at the file and block level to take. Can you use them at your use them for your bidding? You know, kind of like Conan the Barbarian, you know, drive drive forth those storage arrays. Do my bidding. So I think that begs the question, why? What value can we get out of doing that?
Yeah, there's there's quite a few. I'll take a step back there because that in and of itself, what I described is more of the tech piece. But I would couple that with an approach that uses policy engines and intelligent software fabrics to do that work. So you certainly could just push things into cloud. Like, for example, I've got a Synology box at home and it backs up my environment and it doesn't really know what it's doing. It's just writing data to its file system and then shuffles that off to Glacier in Amazon.
And that that process, that workflow doesn't have a lot of value. I'm just kind of shuffling stuff from one point to another. And if I want to find something, I have to kind of dig through the Glacier snapshots and go through my vault and find what I need and kind of hydrate an entire backup. And it's kind of it kind of sucks, to be honest with you. But it is my, you know, my fireproof vault, in a sense, in the cloud.
The value is that if you can make it so that the software fabric that's doing all that work is both aware of the topology that you're trying to protect, you know, the applications, the data itself, but also where it's running, what host, is it physical, is it virtual? And it's aware of itself. It understands, you know, hey, I'm a cloud version of myself or a physical version of myself. And here's the resources I have available and how many nodes are available. Then the major advantage there is that you have a system that's somewhat autonomous for the most part.
It's intelligent about how it's protecting resources versus just, you know, oh, we can go 21 gigabits per second ingest, which is largely not that interesting. The other advantage is that once that data has been placed into, we'll say, target areas such as cloud or secondary data centers or colos or whatnot, it's all still kind of live, so to speak. You can search it, you know, you can restore from it, you can put it in wherever you need to go, and you can even run new workloads from those backups, you know, make copies or recover from a disaster or something like that.
So basically a very long way of saying because the whole workflow from the point where the data is actually being inventoried and backed up to put onto a system on-prem or in the cloud to actually being recovered is entirely managed by a system that understands everything that's going on, you know, across the board. It's not, you know, kind of playing baton race where you keep passing along the data to someone else to sort of figure out what's going on. You know that whole story, so you can make it very actionable.
So when I have this data that's now in the cloud, can I use it? Let's say I write it to Amazon. Can I now fire up that data and use that from within an Amazon instance? Is it available there or is it in, you know, this special format? I like special format. That's interesting. No, it's available in multiple different ways. So, I mean, in the most generic way, a lot of customers are going to use some sort of bucket with an S3.
And, hey, it's now there to be recovered. So I can put it into any rubric environment that's running the rubric code, you know, whether or not that's a cloud cluster or another data center, in case the first one got hit by a comet or something. And that's kind of the most common checkbox use case is I just want to have a giant repository. More interestingly, though, at least to me, is that as you have an environment within Amazon, everybody does. Everybody has a VPC somewhere or is using some services within Amazon.
You could, in fact, point the rubric software to your VPC and the various security groups that you're running and say, yeah, I've got a backup from my on-prem environment or even another cloud. And then say, take this backup from yesterday or a month ago or whatnot and construct an AMI or an AMI out of that. You know, go ahead and run that in EC2. Here's the EBS for Block Store and go. The software is actually smart enough to figure out, okay, well, I protected a workload with this many gigs of RAM and CPU and whatnot.
I'll go ahead and suggest the instance type that will match that or maybe be slightly larger and then place it in your VPC for you and run it. Oh, that's cool. So there's a lot of different ways you can do it. So I really could spin up a test bed with my actual data once it's ingested into S3, for example. Yeah, and I think, Mark, I think that's the key is that it's not some kind of clone or something. It's like the literal actual data, you know, bit-for-bit copy running on the workload that you originally had it on.
So there's, you know, other than networking, there's really not anything you have to do. And even if you want to just run it within the VPC for some kind of internal test, then there's really not any networking to do either. And that's cool because I see a lot of storage arrays do native cloud, well, not non-native cloud snapshots, right, where they take a snapshot on a VMAX and it writes to Amazon, but then it's still a VMAX snapshot. I can't do anything with it and take advantage of all that Amazon ecosystem.
Yeah, and I think there's a lot of overlap in parts of that story, but to be able to go completely through the workflow, some of the folks that I work with are saying, all right, I want to spin up an entire copy of my prod or a subset of an application tier within Amazon. They can actually automatically do that, and then every morning they have a fresh set of instances that make up that application tier to code against, and then they just throw it away.
So they're not having to kind of build out a dev environment on-prem that may get used only sporadically. They can just use this whenever they need it and then pitch it when they're done. So let's talk about a slightly different use case, because you said something that piqued my interest earlier on in the conversation. You have this global distributed file system. I don't know if that's the exact term you use, but this this availability of I'm deep on this idea of availability of metadata throughout multiple environments, whether it's my on-prem environment or my cloud environment or my production data center in my data center.
There's an immense amount of data. I mean, there's an immense amount of value in having access to the file system across a geographic reason, region and across data centers. Is that a solution that Rupert provides is provide a foundation for a global file system? Yeah, that's an interesting way to put that, because historically I think we've had one but not the other, like tape. You get metadata, but not the data. You know, you get a catalog to search, but you can't get the data unless you go find it.
And in other instances, you have the data, but not the metadata. It's like here is this giant backup file, but you don't know what's inside of it. Right. So that's exactly what we're trying to solve. If you've ever had to import the tape, the metadata from like a bunch of LTO tapes, you know exactly what you need. That's not fun work. Yeah. Yeah. My favorite is when you have to go find the right tape that has the catalog to then find the data tape.
You've got to restore the catalog. I digress. Exactly. So, yes, the system is designed, you know, I think of it as a global namespace because file system is a little deeper, I think. We're largely abstracting the file system itself. It's more of the namespace level, and the namespace includes the data and the metadata. They're always coupled together. In fact, every time we put data into an archive, you know, let's say it's S3 for sake of argument. We also co-locate the metadata with it because you never know.
I might lose the on-prem environment, including the config and metadata on the system itself. So those things are treated with equal value because we're largely dealing with irrelevant bits if we don't have the data and the metadata at the same time. You don't know what to restore without the metadata, and you can't restore something without the data. Therefore, you need both. And so, yes, everywhere that the data has been stored in some way, shape, or form, you also have access to all the metadata, which to kind of de-nerdify that, just, you know, what is it?
What's the index look like? What are the files in there? You know, what are the timestamps, ACLs, all that kind of jazz? They're all there. So watching some of your previous stuff, I know that Rubrik is all API-driven. So I'm going to get really fancy from at least from my knowledge of applications and data and ask the question, can I take can I build an app, a distributed app that says, you know what, I collect genomics data sets, which are huge in nature and can be relatively not necessarily become stale, but become to a point where keeping it on premises becomes impractical.
So I send that up to the cloud, but I get a algorithm, a new formula or something that I need to run it against that large historical data set. Can I write an app that basically says, hey, find out where this data exists, spin up this type of workload and then process run my query against around my algorithm locally against that data set. Yeah, that's a that's a that's a curious way to do it. And I'll kind of break that into two thoughts.
The first is kind of the big data and HPC space, which is well served by, you know, Apache and Spark and Hadoop and those, you know, those kind of folks that are trying to take a lot of data and hopefully cram it into memory if they can or distributed across a bunch of spindles and literally just be a query field for AI questions and machine learning questions and things like that. So I think that I think that ecosystem is well served by those particular applications when they're trying to do massive data sets and answer questions.
And a lot of cases, those will be on prem just because storage is the greatest cost factor in a public cloud. So that's the observations I've I've made from dealing with customers and the tech industry at kind of the big data, ML, AL, you know, the type of ecosystem. Now, if you drill down a little deeper, though, and we talk about the data set that rubric contains, the use case that I think you're scratching at largely comes from the first time I saw it was from ransomware, actually, where a couple of years ago one of our customers was hit and the system noticed there was a bunch of file modifications occurring and, you know, kind of warn somebody.
And the query that he ran in this case was, OK, find everything that had a particular modification date. Or in this case, I think he had a string modification, such as the file extension changed. And so you can't just globally query and say, OK, find something within the data sets that match this particular pattern. And if you find something that matches that pattern, in this case, the action that the gentleman wanted to do was restore from the previous iteration, restore from the file before this modification happened.
So I think that's just one example use case where the application, you know, quote unquote, that was used was essentially just a little homegrown app written in Python to find some stuff. But certainly, because as you alluded to, the API that rubric uses is not just something that's on the front end in case you happen to want to automate, you know, with rubric. It's actually much deeper than that. The system itself is using it to talk to the various services and nodes that live in the system.
The entire suite of features that are exposed in the graphics interface and by the system itself are all coded at the API layer because that's how we do business for our own internal needs. There's not some like weird Java SDK under the covers or anything you feel like that. And to me, that's important because how the heck are you going to do public cloud at scale with any meaningful penetration if you're not based on APIs like that's how cloud works. So, yes, it goes pretty deep.
So we have all this value in secondary storage. Wouldn't I want to be able to get this same value integrated with my primary storage, be it vSAN, pure or whatever? Wouldn't I want to work against my primary data just the same I would my secondary data, though? I mean, potentially, there's certainly tradeoffs there. And it's funny, we had a guest on the DataNauts recently. They talked around data locality and data efficiency and things like that. And the conclusions that we came to were that there are certain things that you can absolutely do in the primary storage world, and they make sense, but they come with a cost.
I think everything under the sun from a storage perspective can de-dupe these days and has flash inside of it. And there's tradeoffs for that. You obviously have CPU and things that you can offload to do the de-dupe. What byte boundary are you going to do that at? What are you going to use from a compute perspective to do the search and the analytics? I think the folks at the company that was called DataGravity recently got snapped up by Hytrust, if I remember correctly, because they tried to do that at the storage array level and dedicated an entire controller of the two to do that.
And that in itself wasn't very palatable, but maybe as software it was more palatable. And there are vendors in the ecosystem that try to be like a bump in the wire or running potentially as a virtual machine or something doing that. The challenge is that I think if you're a vendor that's in the storage space trying to offer this, you're forever going to be kind of bound politically or technologically to your product. And really the goal is to work beyond a single cloud vendor, a single storage vendor, but work kind of everywhere and make all of the infrastructure, whether or not that's public cloud infrastructure or in your data center, irrelevant from a constraints perspective.
So that just means Rubrik needs to offer primary storage. Well, we have no plans to go in primary storage, to be crystal clear, because I think where you run the data doesn't have to be where you're necessarily protecting the data. And in some cases, actually in my mind as an architect, that's an advantage because you don't have the same code base and you've got your prod environments not really impacted by anything going on in the secondary. Yeah, there's definitely some advantages there, especially when we start hitting financially relevant and business impactful workloads.
So I just want to make sure we're totally clear that Rubrik is not going in the primary storage market. We don't even really consider ourselves storage at all. Two of our three offerings don't even have hardware underneath them. I'm writing it down. Rubrik offering primary storage. No, no, man. Career limiting move. So with that said, you know, by the way, you should go back and listen to the podcast. We have with Jeff Snow from Microsoft and, you know, obviously of of PowerShell fame.
The great conversation, the where can again people find you on the Twitters and online? You're plug the wall network again for me. Sure. So feel free to hit me up. I have open DMS at Chris Wall on Twitter. L. or you can reach me. Wall network dot com. LinkedIn, whatever it is you want. I try to be available. I don't respond right away, but I try to read everything. And my good friend Mark me.
Where can folks find you? As always, you can find me on the Twitter says at Sensi Storage or if you want my blog, virtual storage zone dot com. All right. And my daughter doesn't like for me to call it on the Twitters. So you can find me on Twitter at CTO advisor. And then, of course, you can find the podcast, the CTO advisor dot com. Talk to you guys next episode.