CTO Advisor 068: CommVault HyperScale

I’m joined by Scott Lowe of AcutalTech Media, Justin Warren from PivotNine and Ray Lucchesi from Silverton Consulting. We are on the show floor of CommVault Go. That’s right I said, CommVault Go. First off I got the amazing opportunity to meet Captain Sullenberger who famously landing a jetliner on the Hudson River. So, it was a cool event despite being a backup conference. Secondarily, don’t call it backup. Data protection and data management has become an hot area of investment. Look no further than our show sponsor Durva Inc . CommVault is looking to catch some of the magic from Rubrik and Cohesity eating into Data Domain sells. This all-star team of analysts holds nothing back in questioning if CommVault’s new flagship product, Hyperscale, adds value. Subscribe iTunes | RSS

Transcript 4,129 words · about 28 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Hey, how's it going? com. You're listening to episode 68 of the CTO Advisor Podcast. You know what? Unfortunately, no Mark May again this week, but that's okay. As okay gets, we are joined by three amazing guests from the community. Scott Lowe from Actral Tech Media. You know Actral Tech from the Gorilla Guides and their mega cast. Ray Lucchici from the Great Beers on Storage. You know him with Howard Marks, great podcast. And again, we're joined in a second straight week by Justin Warren from Pivot Nine.

So great conversation. This is going to be a very insightful conversation on Commvault's latest solution, Hyperscale. Interesting solution. We challenge Commvault a little bit about how is this different from just going out and buying bits and pieces by yourself and putting it together. Where's the value in their solution? Before we get started, again, we'd like to thank our sponsor, Druva Inc. Druva has sponsored the past few podcasts. And according to Druva, Druva delivers the industry's first data management as a service.

A single SaaS platform that unifies data protection and management for endpoints, infrastructure and cloud applications. Unlike traditional systems, Druva aggregates this business critical data for scalable backup and disaster recovery while also unlocking the true value of search and advanced analytics for the governance of that data. For administrators, Druva means simplification for IT leaders scalability. For CFO, Druva means saving a boat load of money. For InfoSec, it means security by design into the highest standards. For VMware and Nutanix environments, Druva means unified backup, disaster recovery and archival all powered by AWS.

Druva roughly translates to saving your butt from data loss, litigation mishaps and regulatory fines. Druva, the future of data protection. Now on to our conversation with Commvault. All right, this is Nikki Schnupp here with Tech Field Day. We are gonna do a co-hosted podcast with a bunch of famous people here at Commvault Go. All right, so we're gonna go around the room and introduce ourselves. My name is Ray Lucchese. I'm a co-host of Greybeards on Storage. I write Ray on Storage blog and I'm with Silverton Consulting.

Justin. And I'm Justin Warren from Vivid9 and host of The Eigencast. And I'm Keith Townsend from The CTO Advisor, host of The CTO Advisor Podcast. And I'm Scott Lowe with Azure Tech Media. Our podcast is 10 on Tech and you may know us for things like our megacasts and all that good stuff. Rajiv. Hi, I'm Rajiv Kottamteril. I'm the VP of Engineering and the back department of Commvault. Hi, I'm Rahul. I work for Rajiv and I have the cloud and virtualization team reporting to me.

So we're here today at Commvault Go. We've talked about the last couple of days, the hyperscaler solution. I just kind of try to understand what it is and why you've done that and where you're going with that. Because it seems like it can go from very small environments to extremely large systems. So we started on the hyperscale technology and at that time we didn't have a name called hyperscale. We started on a technology whereby customers could deploy storage easily in the year 2010.

And then we experimented with a lot of things and finally came to a conclusion that instead of trying to do something from scratch, LustreFS seemed like a good fit that met with all our customer requirements. The key requirements that we were trying to solve was be able to use commodity storage that is locally attached without the need of large storage arrays. Be able to scale out horizontally, which was always a key part of an architecture because we could always add media agents and the place where we were falling short was these media agents were all pointing to a common storage array.

So instead of having a common storage array, having local storage that could be shared across all these media agents. And so the LustreFS turned out to be a good choice and LustreFS scales from a very small three-node architecture to multi-node, 32-node is what we have tested against. Is there a limit to the Lustre node count? On paper, they say no limit, but I'm assuming at some point, yes. There is a limit, yeah. There is going to be some limit. There's always limits, but we haven't seen one yet.

Okay, okay. And so you have your own appliance kind of environment as well as a software-only solution, is that true? Yes, so the appliance is something that we put together which will be tested, shipped from Commvault in addition, and they use the Jitsu service for building the appliance. The reference architecture, we are trying to partner with various vendors who make servers so that the customers don't feel that they're locked in to a Commvault box because what we do not want to do is lock in our customers to our hardware and that is number one priority for us.

The reference architecture scales higher because it's more of a scale up and scale out architecture rather than just a scale out architecture and it better suits our customer base which are usually large enterprises. One of the things we heard this morning when I believe it was Al that was speaking on stage was that there's a, I'm not sure it's a hard limit, but to about 100 terabytes for the appliance that you're selling directly, and beyond 100 terabytes you recommend people go to a reference architecture?

Yes. Is that 100 terabyte limit a hard limit or can people continue to add nodes from Commvault if they want to, beyond 100 terabytes? So the software is the same. So we have one software for hyperscale which is the cluster that is packaged into our media agent and we are selling it in two form factors. We're selling it as the appliance and we are selling it as the reference architecture. The appliance will tend to get more expensive when you go beyond a certain set of nodes and it would be cheaper for the customers to go with the reference architecture where they scale up individual nodes.

So it's the cost benefits that is limiting the number of nodes on the appliance. Okay. So it's not a hard limit. It's a software feature. There is absolutely no software limit. You could if you wanted to. I could if I want to, but it won't be economical. Right. But yeah, definitely. And what Rajiv said is we always believe in openness, so that's why we went with vendor of your choice. So we have already done the reference architecture with six vendors and Cisco decided to bring out their own version of the appliance.

But we definitely want to support whatever vendor you want and use your hardware of choice instead of boxing you in. We believe completely out of the box architecture. So from a support perspective, obviously this is a software only solution or software first solution I like to call it. Where are the touch points when it comes to interoperability? Because obviously if I can start out with the reference, I can start either way. I can buy the appliance from Commvault. In theory, there's no reason why I can go out and buy a reference appliance and add that to my cluster.

From a support perspective, where does the support model kind of break or touch points for supporting a hybrid solution I think? So in a hybrid solution, if you are starting off with the appliance, our recommendation is to create separate pools. So you have a pool of the appliance and a separate pool for your reference architecture. If you do put them together, if the issue is on the hardware that is related to the appliance, the support call, the first call comes to Commvault and then we hand it over to the hardware vendor.

If in the reference architecture, if it is the call will first anyway come to Commvault because they are seeing a problem with the Commvault software so the call always comes to us. And if we identify that it is a hardware-related call and it's a hardware that the customer has procured through the hardware vendor, we would then work with the customer, with the hardware vendor to get that result. And the dependency on the software, on the hardware is minimal. We have minimal dependency on the hardware and that has always been our strength.

And then they are, so they go in pools of three. So if you decide to mix and match the hardware, you're gonna get limited by the lowest common denominator. So some sell it in 100 terabyte pack and let's say the third vendor sells it in 150 terabyte, we will only be able to address 100 terabyte because we want to replicate and keep the erasure going. So mixing and matching that way is not recommended. But in storage pools, under one console, you could have one thing from Cisco, one of our appliance and one thing of your own.

And we have always helped the customer. We never take a issue, even if it's an issue that is with one of our partners, we never send the customer on its own. We have always been their partner, resolving the issue, no matter where the issue is. I'm intrigued about your development environment and creating all of this. So the internal environment that you run yourselves, there's a lot of people who are looking at things like continuous integration and continuous deployment. So could you tell me more a bit about how you actually run that development process yourselves internally?

Okay, so the software development process that we have is a quarterly service pack model where we release software every quarter. The software gets- Can you repeat that, what you just said about the release pack or whatever it was? It is a quarterly service pack model, so which is- A quarterly service pack, okay. So quarterly service pack model where 15th of the quarter, the next service pack is December 15th. So we have a fixed release date of the service pack. We publish the timelines ahead of time and we never miss our timelines.

The way the software works is the service pack is roughly broken up into two halves. The first half is for software development and the second half is for the quality assurance. And the quality assurance goes through three phases of acceptance. So initial acceptance into QA, then we deploy the software into our internal test console environments, which are pretty large and representative of majority of our customer base. After that, it comes into my environment, the production backups of Commvault. So it is actual production backups of Commvault that we are deploying this on.

And then we use that for a period of days, roughly 10 to 15 days, it gets soaked in the production environment. Any issues that we find in the production environment gets addressed. com. com, that entire website runs on our platform. And engineering is the number one customer for what we deliver in our platform. So we have a lot of web applications which run on our platform, which many of our customers don't know of. com, that is all being served out by our technology, which an MSP can host on their own and use these technologies to serve their customers.

So from a integration point with RHEL, are you guys rolling in RHEL software updates as part of that service pack or are those two separate channels? So the RHEL updates will be rolled into our service packs. So when a service pack is deployed, we will take care of patching the RHEL binaries also. It is completely seamless to the customer. Yeah, it's a black box approach for the customer as part of our service pack. So given that three month cadence, how much of that environment do you have automated or is there a lot of manual steps still required or is it pretty much all just lights out, press a button, code commit happens and everything gets deployed?

Literally the entire software development process is highly automated. The build process is completely automated and I'm in charge of the build, so I do the builds. Builds are done weekly. We have software application, the software check-in and check-out process, of course, that is manual, right? Because you have developers writing the code and checking in the software. Once the software gets checked in, it goes through a smoke test environment, which is completely automated. And then the new features that we are developing, they have to be manually tested because the automation process has not been written for it.

Once the new feature is accepted and tested and released in a service pack, the next service pack, we get it automated. So always our cadence is new features manually tested first and then meanwhile start developing the automation suite for that. And as soon as the feature is released, you have the automation now continuously testing that feature on a daily basis. So with the hyperscale environment, where you can go from three nodes to 60 nodes, I mean, you must be in an environment where you're testing various configurations of this sort of thing.

I mean, so you've got, do you have the Cisco Scale Protect solution in your environment? So you are testing those? Yes. So if another vendor were to come along and say, I want to provide a Scale Protect-like solution based on Commvault, you'd bring that hardware in and test it as well, is that? Yes, absolutely. Because there is some amount of changes that we have to make per vendor to make the deployment completely automated. So there are two ways to deploy the hyperscale.

You can have a completely touchless methodology where you download the image, push it in, and it'll automatically detect all the storage disks that you have. Almost DevOps-like, yeah. Or you go through the manual Unix, Linux-like installation and select the volumes that you want to use for the deduplication databases versus the storage. So if you want a completely touchless, seamless deployment experience, we try to detect and work with the hardware vendor to find out what kind of configuration we want. And once we want to do that, we get that hardware into our labs and use it in our test labs as a test environment for testing that.

So we currently have Cisco, we have Dell, we have a couple of other vendors which I don't remember offhand that we have these set up sitting in our lab. In the process. So in that environment, let's say vendor Cisco, would you have multiple configurations in their solution or would you have like one max configuration or? Depends. Cisco, we ended up with two or three, I think, of Cisco configurations. Dell, we just have one set of configurations. So it varies based on, because these are all quite a lot of investment to make, but yes.

And do you expect customers to actually adopt the software every quarter? I mean. No, absolutely not, absolutely not. These, all the customers, what we have told, you should be switching service packs yearly. That is our recommendation. So you adopt the software changes yearly, and if you have a critical breakage between the time period, we have the ability to hotfix the service pack that you are on. So Commvault Engineering officially supports- Hotfix. Hotfix, we call it a hotfix. Hotfix if they need a fix.

So it is once a year patch management for customers and hotfixes for critical issues that they find. If they want a new feature or a new functionality, like hyperscale, if they say I want hyperscale that got released with a new service pack 9, they have to go to service pack 9 to get that feature. So if I need a new feature, you have to go into that service pack to get the feature. We are not going to bring that back to your service pack.

And we have stopped the major release completely. We are not going to have a major release. So more or less Symfonic 12 or 11 or whatever it is? 11 is the final. 11 is the final service pack. And we don't mean marketing, we don't even call it 11. We just call it Comfort Data Platform. So from now on, all your functionality will be released on a hyperscale? Will be released as service packs. Will be released as service packs.

Okay, so hyperscale or not? Yeah, hyperscale or not. So hyperscale is just another addition to the existing platform. That platform did not just evolve into it. So software is now incremental forever, not full planning with incremental. No, incremental forever. Actually, and there's a reason for it. We were actually involved in helping customers upgrade. The first 300 customers were upgraded by engineering team members assigned to them. And we have seen through it a lot of MSPs, government agencies, going through a change is a pain.

And we believe in relieving the pain, change of pain. So it is going to be a part of the service pack and people can upgrade whenever they want. So is the hyperscale functionally equivalent to every other platform that Tombo uses today? This is basically taking your solution and putting it into a box and making it easier for clients with that, is that? Yeah. So yeah, it is. Yes, it is absolutely. So now, it's the old one and the new one.

So the software that you're buying with this solution is what I bought yesterday, but it's delivered to me in a different way. Yes, we took the software. We made sure that any complexity in deployment that customers had. We listened to all the comments. They said, oh, this is difficult. We said, okay, let's simplify it so that you can just one-click deploy the software. All the policies are auto-created for you and you can just do the entire deployment in less than 35 minutes.

When I last timed it in 35 minutes, you're up and running. So you're in a unique situation where you're both the backup administrator and you're head of engineering. You're on a much quicker cadence than the rest of other Commvault customers by recommendation. What are some of the lessons learned from being on the engineering side where you guys have this super automated process for application development via deployment on the customer side? What are some of the, what we call the DevOps or infrastructure as code lessons learned that you've learned from your constant deployment in your administrative role?

Okay. Number one lesson that I learned, change is always painful. That is the number one lesson. So however hard you try, however much automation you have, with change, you are going to find breakages. There is nothing that can prevent it. So anybody who says that change, I can make a change without downtime, that's number two. You make a change, something is going to break, whether it is an iPhone or an Android or if it's Commvault, doesn't matter. It's the same story everywhere.

So what we do is to ease the pain. That is why we went with this quarterly model where we limit the amount of change that goes in each quarter. So that way you are easing the pain of change because you're bringing in changes in increments. And any other model where you're doing a release after two years, you're bringing in a ton load of changes and making one big change and the customer is all of a sudden unleashed with thousands of features and that is never helpful.

So lots of small changes instead of lots of changes. Another thing that he has realized and made us realize, Rajiv spends very minimal time on managing the whole console. So we are managing by exception. So no matter how beautiful our console looks, on a daily basis, when you start looking at something, you're going to not like it, no matter how beautiful it looks. So he's completely managing by exception. When things go wrong in the environment, the jobs are not where they're supposed to be or some MA is not performing at the level it should be, he gets an alert.

And all he looks at is that alert and pretty much that's where he spends time. Beyond that, he doesn't like to spend time and that is what, whether you call it machine learning and all of that, we have incorporated our language to learn what's the normal behavior for this console and we bring it to you and we expect you to manage by exception, not looking at any console. Even though the dials are beautiful, we don't want you to look at it. Let me ask you another question around, if you guys don't mind, going back to the licensing model.

Yesterday, when I bought Commvault, I bought software, typically, and then I deployed, I paid a licensing fee and I paid annual maintenance fee. Can I still do that or is this basically, we're now moving to hyperscale is the way we're deploying everything and we're gonna go to a pay-as-you-go sort of thing rather than a license plus maintenance? You will still have the option of doing license plus maintenance and not use hyperscale if you choose to. So you can still just buy software?

You could still just buy software, yes. Hyperscale is not mandated on you. If you, and we would love to still continue partner with all our other hardware storage areas. But if you wanna buy the hyperscaler, then you are buying into the subscription consumptive model that licensing. Yeah, hyperscale is just, yeah. You can't buy the appliances outright? It's a subscription model is what I've been told. Okay. It's what is being sold for the appliance. Buying an appliance outright, what happens is over a period of three years, the server gets old, you don't need it, you really don't want to.

So a subscription, what it's taking care of is every three years, it's a refresh. It's equivalent of your two-year iPhone contract because beyond two years, the iPhones. I can see a case, though, where if you've made it, now, can I ask this question? Hyperscale, you said, has made it super easy to deploy and all that good stuff. Is that ease of deployment now that part of the original product as well? Yes, yes. Okay, that's what I wanna get. Yes, yes, yes.

So the hyperscale is just the storage piece of it, right? Okay. So the ease of deployment piece of the platform itself is independent of how the hyperscale modes are configured. So if you are buying Commvault today, which you can go to the trial software and download the trial and deploy it, the deployment process there on your own hardware, which you have not purchased from Commvault, and it is not hyperscale, this can be completed in less than 20 minutes because you're not even configuring the storage part.

You're just configuring the server components and that, once deployed, will auto-configure all the policies for you. Only piece that would be missing is the storage element that is the hyperscale. So we have segregated it into two pieces and they both tied together come as part of the hyperscale, otherwise you just get one piece. It's one platform, one core base, one set of developers. And we apply lessons learned in one place to another. So a lot of automation and all of that is coming through.

And yet we have not taken away the ability to do really complex tasks. So customers, many of our customers, though we have this easy cookie-cutter policies pre-built, many of our customers want the ability to change it. And we still allow you to go back and change everything. We have REST APS for all the operations that you want to do. Everything is available via REST. So a lot of customers will use that to automate their end-to-end. Okay, well, it's been about 20 minutes and I really have enjoyed our time with Rajit and Rahul and appreciate the discussion.

Thank you. Thank you very much. Really appreciate it. Thank you.