Is the storage industry dying? CTO Advisor 072
Howard Marks joins us this week to talk about the current state of the storage industry. With challenges from HCI vendors and Public Cloud, Keith asks both Howard and Mark if the storage industry is dying. If it’s not dying when do data center managers select HCI vs. Cloud vs. SAN? Great conversation on how data gravity impacts design decisions. Subscribe iTunes | RSS
Transcript
Hey, welcome to episode 72 of the CTO Advisor. And we have the acclaimed Howard Marks as our guest. But before we get to Howard, Mark, how's things going? How's the Christmas season and holiday season going for you? You know, it's doing pretty well. I'm ready for Christmas to be over. Just got back from Chicago. That was nice seeing you up there. A little cold, though. It was. You guys brought the snow with you. I'm glad you left.
As you guys left, the snow left. But I did get portellos, so that's all that really matters. Yet, I had no portellos at all. You did have French toast. You sent me pictures of portellos, but no said portellos at my home. So next time. What is a portello? What is wrong with you? Oh, God. I'm from New York. You've been in Chicago way too many times, Howard, to not have had portellos. Think of it like a Philly cheesesteak, but like moist and dipped in gravy.
What about Italian beef? Yeah, I'm not a Philly cheesesteak fan from Philly. The Italian beef is a much better sandwich. But before we get into that debate afterwards, so what we can say is that neither Italian beef nor the Philly cheesesteak is dead. However, I'm going to put this premise out there that the storage industry itself is dying. The good old days of there being multiple analysts, Howard having any kind of competition for doing white papers, e-books, seminars, webinars, et cetera, is over.
I no longer have dreams of becoming the next Howard Marks for storage. Howard, when you go, there will be no more storage industry. I think dying is a bit of an exaggeration. It's middle age. It doesn't move quite as fast as it used to. And it's got a cardiologist and went in and got a sedentary. But that doesn't mean it's dying. It's not young and vigorous the way it was. It's old and busted. Well, it's not quite old and busted, but middle aged, kind of like me.
It can't dunk anymore. Oh, no, definitely not. This is not Dr. J. And this is especially not Dr. J with the nets. So I'm going to push back on that a little bit. When I talked to my good friend, Duraj, when I talked to the Open Converge guys who have sponsored CTO Advisor content in the past, Datrium, this whole idea of storage as a standalone product and solution is no longer, right? Well, I mean, DSSD is no longer a thing.
DSSD is no longer a thing, but E8 and Accelero are. You know, look, the external array market is declining at about 2% a year, and that's not going away. Because on the one hand, software-defined storage, whether it's in the same server as the compute in the hyperconverged model or in another pile of servers in its own rack in the disaggregated model, is eating away market share. And at the same time, workloads are going to the public cloud. And if there is one product in the enterprise data center that public cloud vendors do not buy for their public cloud service, it's the traditional disk array.
Yeah, that's why I think what it's better to think of is the storage array market is declining at a reasonably decent pace. But it has a heck of a tail on it. Yeah, no, we're talking about a 2% to 3% decline. It might accelerate to a 5% decline in a bad year. So let's put some raw kind of scale or help people. You know, let's put our handprint up to the virtual mic screen and compare stuff. I've done enough seminars at CIO events and et cetera to know that storage is probably still the number one or number two cost in the data center.
Depending on if you're buying a bunch of network gear that year, that network gear might surpass storage. But for the most part, storage buying disk is still like the number one cost in a data center. That isn't sustainable. But let's get off. Before we get off of that. Something's got to be number one, right? Well, the question is, again, this is to compare just so people can get size. I've come from shops that have had a few $100,000 a year IT budgets to shops that IT budgets go over a billion dollars a year.
And storage is almost always one of the top line items, at least within the top two or three. In really big shops, storage is obviously second to people cost. And then it varies between the two. So the storage industry, so Howard, when you say the storage industry is declining at 2% to 5%, the traditional rate market, that's still, you know, fast forward 10 years now, that's still an extremely big market. And it's not, well, there's only one way to stop paying for storage.
And that's to stop creating new data. And I've talked to a- We are drowning in data. It has become a cliche amongst those of us who listen to vendor pitches. But what's happening is the cost of storage comes down every year. But the cost of storage doesn't come down as fast as the quantity of data being created goes up. Well, let me think about it this way. Is anything going to change that's going to start increasing that market? As we get better at doing stuff with that data, as data analytics get better and real-time analytics get better, is that market going to grow?
Well, that's one of the drivers we've had for the past three or four years, that since the target sent your daughter a postcard and you didn't know she was pregnant incident, corporate America has recognized that the data they used to just throw away has value. Yeah, but how many of those corporate Americans still kind of suck at it, right, enough that they're not doing enough to really unlock that? There's three groups in corporate America. There's the few, the targets, who are good at it, who have moved on from Hadoop to Spark.
There's the group that read the article about Target and decided they were going to keep all their data forever just in case, but doesn't actually have any way to analyze it. And then there's the third group that just goes, yeah, we're not doing that because we're not smart. So do any of those change this, or are we still going to be declining as we're eaten by software-defined X? Well, it's mostly it's eaten by software-defined X. Because remember, we're still buying the disk drives.
We're just running vSAN and Nutanix, and they're going in the servers instead of being in their own device that we bought from Dell EMC. So I guess that's another question I want to tease out, because there's a couple of things. There's storing the bits onto the physical disk, because that happens somewhere. And we're all in agreement that the creation of data isn't slowing down. That's growing. So the need to store 011s is increasing at a rapid pace, faster than the cost of storing those 0s and 1s.
Then there's the data analytics, being able to actually put value. We've used this term data as the new asset class. How do you actually glean value from that? So let's tackle the first problem, the storing of the 0s and 1s. They're storing the 0s and 1s in my own data center, and then they're storing the 0s and 1s in AWS and Azure or whomever, who has the ability to apply the AI and the ML to that data set. Is it getting cheaper to store that data in the public cloud than it is in my own data center?
That's a question of duration. And of use, right? Storing it's cheap. Using it's a whole different problem. So if we're talking about storing it in EBS, where it's being actively used, cheaper to do it in my data center. If we're talking about storing it in S3 or in Glacier, where it's infrequently used, then the question is, how long am I storing it for? So Howard, at AWS re-invent, Andy Jassy got on stage and said, oh, SageMasher, SAS-based ML AI that leverages object-based storage, S3.
So you can have your cheap and deep storage, and then you can actually use it for ML AI against that. And then we've all seen the products that accelerate object to varying degrees of success. Well, there's this argument that object. Even EMR works out of S3. So when I'm looking at storage traces on the benchmark project, we get a trace. I load it up into Elastic MapReduce, and that goes through S3, not through EBS. I think there's a huge market not directly related to storage for organizations that want to do Hadoop Spark-style analytics, to an extent that doesn't justify having two guys who are Hadoop specialists and a rack full of gear that's a Hadoop cluster.
And sending that data up to S3 and firing up EMR to analyze it and then shutting EMR down is just the right way to do it. Yeah, I mean, it all goes back to a lot of people are good at that, and they're not going to be good at that. So we want someone else to do it. It's partially the not good at that. And it's partially the, and I only run it 40 hours a year. So let's take that use case, in that we only run it 40 hours a year.
And we have this few, let's say it's a petabyte of data. Petabyte is the new terabyte. Let's say it's a petabyte of data, and I'm going to fire up EMR once a year to do some analysis 40 hours or 40 days. Is it cheaper to store that object data, that petabyte of data on prem in my own data center or idle in the cloud for that 11 months of the year that it's idle? What are we doing with the data when we're done with it?
I think that's what it comes down to for me. Is it going away, or am I keeping it there still? No, I'm keeping it. Let's say that I add to it every year. That one year was one petabyte. Every year, I add a quarter petabyte a year to it. And I perform that net new analysis every year. So it's incremental data. So I'm going to say on prem. It will cost more to store it in S3.
But it won't cost as much more to store it in S3 as the effort of shipping it back and forth and the ingress and egress charges would add up to. Yep, exactly. It's all about your usage pattern more than actually storing the data. Your puts and gets, they get expensive. But if this isn't 40 hours once a year, but it's four hours once a month, then it's definitely cheaper to leave it in S3. Yeah, so what we're talking about basically is the data gravity problem and where that specialized compute is at.
This is the problem that enterprises are running into is that compute is becoming way heavier than it was in the past. We could easily say, you know what? If the data's in Germany, it's easier to ship a bunch of chassis to Germany, spin up a complete compute cluster. And even if we only process the data once a year, it's still easier and cheaper to move the compute than it is to move the data. Now that we're talking about specialized compute services that are only existing in AWS, it's not so easy to say that I can move the compute.
Now the data kind of has to be where the compute is at. And it's not even so much proprietary. I mean, the stuff I run Spark and Scala in EMR, it's standard stuff. But I spent three days trying to spin up a Hadoop cluster in the lab where I've got compute. It's not about whether it was cheaper, it was about much easier. And there's a lot of these applications like Hadoop where if you're gonna bring it into your data center, you probably have to hire two guys.
So Mark, this becomes a question that's not just about cost, but ability. We talked about the zeros and ones, where is it cheaper to store the zeros and ones? It is not that easy of a question. There's this capability question that has to be answered. As we talk about the capability question, and we look at our data center, practical question, there's Nutanix, there's Datrium, there's our mutual friends at Scale Computing. There's still EMC, Pure, Tentree, the name goes on. We have these investments in our data center that we have to decide that we're either gonna make a new investment and whether it's a distributed storage technology or a traditional SAN technology.
How does this capability decision weigh upon us as we're looking at a practical decision, do I expand the lease to my VMAX? Well, I think at least in that perspective, from a storage array perspective, they're easy now. What used to be a VMAX, it used to take forever just to carve up a LUN to make a data store took, I don't know, countless hours, is now something almost anybody can do. Now it's about the care and feeding and the protecting of your data more than it is about managing your array.
So when I start thinking about those capabilities, it's not the storage layer, it's at the analytics layer and finding the people who can be qualified to run that layer far more than running the storage array. So from a zeros and ones, as you're debating on where to put the zeros and ones at, you're no longer thinking at the whole, okay, how much does it cost me to run a VMAX? That's not necessarily a real consideration other than how much does zeros and ones cost on VMAX per IOP versus how much does it cost on Tantri versus how much does it cost on Nutanix?
Yeah, exactly. Well, for VMAX, it does a little. To run a VMAX, you still need somebody who's had training. No, it comes easy now. It comes with an easy button now, Howard. But if once you have an external disk array or an external flash array, and you've turned the VASA vVols channel on, the user interface for creating a new VM and provisioning data for it is exactly the same SPBM user interface as if you had vSAN. And even if you're not using vVols or VASA.
No, no, but the best, you know, it's not where the storage is. Physical architecture doesn't dictate operational. It can be exactly the same. Agreed. So I'm gonna sneak in one last question and I'm gonna be greedy. We talked about if the storage industry is dying, that's obviously clickbait. The storage industry isn't dying. We have to store the bits somewhere. So disk drive makers are not worried about going out of business. However, when we abstract away into the data center, we have to make these decisions about collapsed models of three-tier architecture versus open converged, hyper converged, et cetera.
What are the real decision points today? Is that really even a consideration? Howard, you've written about this an awful lot, that the effort that it is to manage a pure array, a tantry array, it doesn't matter, it's not that great anymore. It's not that big of a deal to manage a storage array anymore. No. Is there a real value in, and is there a real financial or technical value in three-tier architecture versus a hyper converged architecture? Well, the thing you have to remember, if you separate storage from compute, you actually get some advantages.
The first is that your compute goes back to being stateless. And I'm a storage guy, I'm paranoid. And so for me to do a firmware upgrade on a server, if that server is running a hyper converged software stack, now when I take that server offline to do a firmware upgrade, I've reduced the resiliency of the system because it's got N minus one nodes running right now. Having the storage separate changes the planning process. The biggest difference comes when, how much flexibility do I want?
Because if I just want to add compute and I buy it from an HCI vendor, I mean, even if we're talking, you know, it's vSAN. So I'm buying a DL380 from HPE, but I have to pay VMware a vSAN license to access the data that's on vSAN for that node at, you know, seven, eight, nine grand, as opposed to if it was accessing an external array, I'd just plug it in. So how... So I'm going to push back on that because all these HCI vendors, and I think you've created a nice caveat with vSAN, but all these storage vendors have these storage only nodes and these data only nodes.
We're getting to a point where the complaint of, oh, if I want to independently scale storage and compute. It helps you to scale storage independently. If I need capacity and that's all I need, I can either add a storage only node and that's less expensive than adding a compute and storage node. But most vendors won't let you add a compute only node. And that means that you're paying storage markups for your compute. So let's talk to a practitioner, Mark. You've dealt with both environments, pros and cons.
You know, to me, if you're on a modern storage system, from a day-to-day perspective, you're not going to notice a difference either way. When it comes to what Howard talks about, paranoia about your data, that's a design decision that you have to make upfront. You know, what protection scheme you protect and making sure that when you grow and add capacity, you keep that in mind. But I like a storage array, but just because I'm a storage array guy, right? In the end, we're getting close to it, but it's not mattering either way, right?
I think for many organizations, it's a political personal preference decision more than it is a technology and there'll be a significant impact in cost or operations. Yeah, I agree. We'd rather buy a Chevy than a Pontiac. A Firebird and a Camaro are the same car. They come off the same line, but I know guys who would never own a Firebird. Yeah, I've made this argument for the past few years, ever since we started talking about what Wikibon called server sand.
At this point, if the goal is to save money, you're probably barking up the wrong tree. This is a design decision. What it is, the unique advantages, whether it's performance or capability that you're looking from a storage services, data services, or compute services from a platform that you're looking to get versus, hey, can I eke out some operational cost savings or even some acquisition cost savings? I think the hyper-converged vendors can make a argument that there might be a lower barrier of entry for some HCI solutions than there are some storage array solutions, but there's so many storage options out on the market now.
I think that's even a hard argument to make. I will give you that in the four node and down category that most HCI solutions are preferable to most external array solutions. For the remote, the robo. I'm not gonna put a sand in the, I'm not gonna put a sand in my warehouse in Alabama. And I would say that's not even robo. I mean, a four node VMware cluster with four nodes with 512 memory, which isn't that expensive, that's a lot of computing power.
That's well beyond robo. Robo to me is, I've got a file server, a print server, and maybe one or two other things, not data center in a box. Well, frequently it's, I've got the thing that runs the cash register and the inventory. Yeah, because robo frequently is big box stuff. So that's, you know, our friends at scale make a pretty good living off of that four node in a smaller market. Oh yeah, no, and Jason McCartney tweeted just the other day that he's talking about two node vSAN all the time.
Small, good disk arrays are rare. And- All right, so we can talk about this topic all day, but we don't have all day. This is not a geek podcast. We got a little geeky, that's okay. No, I have to go to Star Wars later. No, you wait, hold on. Star Wars came out yesterday, Howard. I'm not understanding. Yeah, why are you so behind? What do you mean you have 15? Well, you had a lightsaber. I have several lightsabers.
I don't have the three-way lightsaber though, Yeah, I am waiting for it to come out on Apple TV. I'm not, you know, I'm not a fan of the HUD so much. So for those who've already seen Star Wars and they're ready to stalk you online, Howard, how can they find you? net. For the blog. And I'm currently on part eight or nine of what looks like it's gonna be a 25 part storage basics and intro to data protection series. And you guys, you and Ray do an excellent Greybeards on storage podcast.
Where can they find that? Oh, on iTunes and Stitcher and all of the usual places. It's Greybeards on storage. There is an at Greybeards store, Twitter account that we use for it, but basically it just tweets when a new episode's ready. All right. And Mark, as usual, where can folks find you? com. All right. com. Of course you can also find the podcast. There's RSS feeds and iTunes links that you can subscribe to the podcast.
And then you can also find me on Twitter at CTO Advisor. We'll talk to you next CTO Advisor podcast. Thanks.