EMC DSSD Shared Persistent Memory - CTO029
EMC made a big deal of their DSSD product announcement. But, EMC makes a big deal out of every product announcement. Chad Sakac wrote his normal verbose post on the topic. To help digest the announcement I invited industry expert Howard Marks. How helps break down the technology and potential use cases. We also talk about other persistent memory technologies and why they haven’t taken off. Show Notes DeepStorage LLC on the Web SpeakingInTech with Chad talking DSSD Sandisk persistent memory solution Subscribe iTunes | RSS
Transcript
You're listening to episode 29 of the CTO advisor podcast. I'm honored to have industry expert Howard Marks. He's the, I think his title is chief scientist at deep storage LLC. We're going to talk EMC DSSD. We're going to try and go through the fluff, understand exactly what it is that the product is offering. How's that different than other persistent memory solutions on the market today? I really enjoyed the conversation with Howard. It's really hard to keep these things to 20 plus minutes or less than 30 minutes.
I hope you enjoy the conversation as much as I did. So Howard, of course, you being the chief scientist at deep storage LLC, we're going to talk Cisco API, VMware, NSX. We're going to really get deep into network. Oh, all this stuff I know, all this stuff I know nothing about. I gave up on networking with IPV6. It was like, nope, this is it. I gave up on IPV6 when IPV6 was announced. Well, not if you talk to Greg.
That's a conversation for our friend Greg Farrell in a future podcast. But one of the reasons I wanted to have you on is this EMC announcement, which kind of led to me to thinking about other things outside of the EMC announcement. And we'll get to it. But EMC's DSSD product completely caught me by surprise. Other than getting a pre-brief on it a couple of weeks ago. I'm like, oh, that's interesting. And then Chad Sackett, the president of VCE was on the L on L Regis podcast.
Speaking of tech with Greg Nierman, he gave a, you know, Chad's version of a brief overview of the solution. Chad does have trouble with terfs. He's never been at a loss for words, either written or spoken. That's for sure. But I wanted to get more of it. And Chad did a pretty decent job of explaining it. But I wanted to get an outsider's view. Can you tell us, first off, what is DSSD? Sure. DSSD is a rack scale flash system that uses switched PCIe.
So you have a series of servers in a rack and the DSSD system. And essentially, the flash in the DSSD system appears on the PCIe bus of all of those servers. Okay, so that's a bit amazing. Well, it's the next... Because PCIe does not need to be a shared bus between systems. Well, it's switched. It's not a shared bus. And the specs for external PCIe cables have actually been around a few years. O. And the pitch there was fiber channel host adapters and 10 gig ethernets are too expensive.
So put our switched box out and share them across multiple servers. Now, by the time that technology came to market, it wasn't economical. The cost of the things you were going to share came down. But those companies drove the PCIe forum to develop the tech for... Okay, so what happens if we put PCIe on a cable? How do we switch it? And all of that stuff. So we solved the hard computer science problem. How do you switch PCIe?
Right. Now, there's still a latency problem because it's 186,000 miles a second. It's not just a good idea. It's the law. And that's basically one nanosecond per foot. And that's why this is rack scale, not data center scale. Because DSSD is the new tier zero. And I expect, and I've talked to a couple of companies that are in stealth that were planning on doing this and looking for funding. So whether they come out or not depends on the funding side.
But the idea of a switched PCIe complex for a cluster of servers is the new tier zero. Ten years ago, we had Violin and TMS doing, look, flash on fiber channel. It's tier zero. It's what you do for oil and gas and high-speed trading and the other stuff where the cost of time is enormous. And then over the next 10 years, flash became normal and became tier one. And now, you can buy an all-flash FAS or an all-flash PCIe. You can buy an all-flash FAS or an all-flash VMAX or a new design all-flash array, like an XtremeIO or a SolidFire or a Pure.
And those systems all have all the data services we were used to because it's now storage. It's not special stuff for go fast. But there's always demand from a small number of customers with very deep pockets for something five times faster. So, RecScale, do we have numbers on what the cluster size should be for a typical DSSD solution? I don't know offhand, but my guess is that it's like 32 ports. So, and that would make, you know, those basic, you know, binary numbers, 32 ports for a initial...
It, you know, it might be 24. You know, anywhere from 24. But I think most people are looking at this for a use case. I don't know how much Tier Zero is out there that Flash isn't meeting. We'll get into kind of that next conversation of Tier Zero. You know, I'm an SAP guy. I want to say begrudgingly. And I've been sorry about that for a while. You know, it's, you got to take the good with the bad.
It's, you know, the SAP group has the money, so they... Oh, yeah. Oh, no, no, no. They have the means to buy the toys. So, you know, obviously SAP HANA, and we look at these in-memory database pieces. Companies like Pure like to make a argument that all Flash-backed SAP HANA is a great solution because, one, you know, that five minutes that it takes to load the memory, your in-memory database at the beginning gets cut down to two minutes. Hey, you do that once every month, I guess.
The problem is, if you're using HANA for mission-critical data, then HANA has to checkpoint back to that block device. And what becomes really interesting from an architectural point of view is not something like DSSD that maps to a block device, but something like DSSD that maps to memory. And you set me up perfectly for my next question. So we attended Storage Field Day 5 together. It was actually my first time, I think, meeting you in person. And we attended Storage Field Day 5.
On the bus from the airport. And that was almost two years ago. And we talked to SanDisk. And they were, I think, showing the Diablo stuff. The persistent memory stuff. That was two years ago. I would think about now, you know, I'm looking at HANA at my day job. And the appeal of something like Diablo or MEMSIS, whatever the HP technology, I'll ask you about that in a minute. What's the holdup? Why don't I have this goodness in the enterprise data centers today?
Well, there's a couple of things. And first of all, we have to realize everything old is new again. When I was in college, we had a 36065. And it had two banks of memory. It had two megabytes of core. And four megabytes of what we called LCS, large core storage. Which literally meant the ferrite donuts were bigger. And therefore, slower. But core was non-volatile. And so, when you submitted a job, you decided, do I want to run this in fast memory or slow memory?
When we start looking at the Diablo stuff, or the Plexistor stuff, or the next generation of that in-server, dim socket, persistent memory. It's interesting for HANA. But the server itself is still an inherently unreliable device. One x86 server has multiple single points of failure. And so, that it's persistent cuts down on how often I have to checkpoint my in-memory database back to some external persistent disk so that I don't lose the data. But if I lose the server, I still lose the data.
And that, I think, becomes... But if we've got a DSSD type architecture, where multiple servers can map the same non-volatile memory into their memory space, then HANA just has to build in a server failure in the cluster and remap the memory that was in that server to this other server to recover the data. That gets really interesting. That does get really interesting. Because if you've had to build a HANA scale-out solution, or even a HANA scale-up solution, one of the challenges is keeping memory in sync between two instances for high availability.
Memory is really fast. Right. And the changes to that memory can happen really, really fast. So unlike a... Well, when you start... Oracle database where two servers connect to the same Oracle database and keeping the last committed write is trivial, that's not so much with memory. So DSSD is a potential, I think, problem solver for a couple... When we're talking about this in-memory database piece, it's a problem solver for a multitude of issues, one of which is high availability.
Yeah, and it starts making the very fast applicable to more use cases. You know, the first 10 customers for DSSD are going to be the typical high-performance computing customers. You know, look, I live in New Mexico. There's two, you know, Sandia and Los Alamos, I'm sure, are going to order them. But when we start thinking about, okay, so what happens 10 years from now? Because, look, there's system-level software that needs to accommodate this new hardware architecture. SAP has to do a HANA rewrite to say, oh, wait, memory is persistent and shared.
That's very different from what HANA does now. Very different. It's really based on that x86 architecture where memory is dedicated to one system. Where memory is fast and volatile and local. And we start seeing memory is not quite so fast. It only goes 200 miles an hour instead of 300 miles an hour. But it's non-local and non-volatile. So that I now take all of that checkpointing and HA stuff out. And all I need to do is build some message queue to act like the transaction logs in a more conventional database.
So that when there's a failure, I know how to bring it back to consistent. Then we can start going, oh, that database is really fast. And now, you know, the NSA can keep all of our data online. Yeah, those real-time searches from the NSA become a lot easier. Yeah. So then I think that goes into the next question, which is databases beyond tier zero. You know, some industries out there, not just finance, some industries just have the money. And as we look at, you know, the next...
Normally, that would include oil and gas. Although not for the next couple of years because oil prices are seriously depressed. But, you know, when oil prices go back up above $80 a barrel, all of a sudden exploration is going to be important to Shell and Exxon and those guys. And they've got deep pockets. And if they can say, well, you know, we could run this application that does four passes on the seismic data and gives us a 60% chance of hitting a hot well.
Or if we can run it faster, we can raise that to 80%. Well, it kind of doesn't matter how much it costs. Compared to drilling a dry well, it's not much money. And then, you know, if you look in my industry, pharma, you know, if you're doing cancer research, the ability to load that data into a persistent layer of really fast memory and process it, and even have parallel activity on it, that's appealing. But I think one of the things that I haven't wrapped my head around, which is the conversation that Chad talked...
One of the topics that Chad talked about on speaking to Tech was that this isn't... This is something that looks different than a storage array. So I'm not presenting LUNs to a group of servers, but something else. What is this something else? What does it look like from an OS perspective? Well, from what I can remember, there is a block driver. So you can treat it like a LUN. There is a key value store driver. So you can store values with arbitrary keys and use it for the back end for that kind of application.
And I believe there's a memory map driver. Now, remember, we're talking about... It's a PCIe connection. So the personality is determined by the software. I think one of the things that Chad mentioned is that the... I don't think there's a Windows block driver for it yet, but I think there's a Linux block driver. Yeah, no, he said there's no Windows drivers at all. It's all Linux. But if you think about who the first round customers are, these are not people who are trying to run Exchange.
I mean, Puma would love to have one of these, but it would be overkill. Yeah, I don't remember much HPC happening in Windows lately. I mean, Linux is just a much better fit for this type of technology. So as we see the industry progressing over the next two to three years and this tier zero becomes more widely, this type of DSD type of technology in tier zero becomes more widely available, are we going to see that same mapping? So am I going to start seeing hyperconverged solutions with DSD type options in the next three to five years?
I would not expect it in the next three to five years. If you look at the hyperconverged solutions, they're not about storage performance. They're about... Our friends at Nutanix and SimpliVity would vigorously argue against what you're saying. Of course they would. But the truth is that the very nature of the hyperconverged model where I'm running storage management and compute on the same processors means that as the designer of the storage layer, CPU efficiency is high on my list of issues.
Where if I'm... So all the hyperconverged systems are shared nothing scale out storage. That happens to run in the hypervisor. But something like SolidFire is pure scale out shared nothing storage. If I'm working at SolidFire, then I can say I have all of the CPU in this box to provide the same services that an all-flash vSAN has to provide and run VMs. And so if I'm writing vSAN, I make compromises around that reduce my storage performance to reduce my CPU load so you can run more VMs.
And that's the right trade-off to make. But if I'm designing a storage system that has the same resources for the same amount of flash and doesn't have to run VMs, then I can run more aggressive compression algorithms. I can do all sorts of things because I'm not worrying about CPU utilization. Yeah, a couple of episodes ago, I had our good friend Punching Clouds on the podcast. He gave a really great breakdown and some of the compromises on from a design perspective that the engineers had to make between...
2. Why is this just coming now? What were the kind of lessons learned from previous versions of vSAN? I think you validated a lot of what he had to say, just the compromises that you have to do when you look at providing x86, basically hyper-converged. When you look at providing x86 compute and storage on the same platform, there are compromises that you have to make. Right, and there are use cases where those are great compromises. I mean, if I was running a state farm insurance or Home Depot or any of these other organizations that have 5,000 branch offices, then hyper-convergence in those branch offices is just the right way to solve that problem.
So let's talk about deep storage LLC for just a minute. So you have this great lab in New Mexico. Tell us what secrets you have running in there. Well, there are secrets. Okay, all right, not the secrets, but what's some of the more interesting stuff that you've performed in your lab over the past, let's say, year, year and a half? Well, we did some interesting work with Pivot3, who does higher level erasure codes with hyper-convergence for colder data. They started off in the video surveillance business.
And so disk efficiency was very important. And so they let you get into the kind of erasure codes that AmpliData or Cleversafe does, where it's like, I can lose a node and two disk drives across other nodes and still retain all my data. We just finished a JetStress project for Atlantis, where we can say, look, exchange on hyper-converged, why not? But the big thing is we're currently building a new storage benchmark suite. It's not so much that we're replacing VDBench. In fact, in Rev1, we're using VDBench.
The problem is that most people today test a storage array for one workload. But most people today don't use a storage array for one workload. And so how fast that storage array will run SAP is important, especially to you. But if I'm going to run SAP and 50 other VMs on it, then how well it's going to run the combination matters. And the data that we use has to dedupe and compress like real data. And we have to be able to test for things like, how reactive is the system?
If the demand changes, if of the 50 VMs, this one goes from demand of one to demand of 10, how quickly does the system respond to that and deliver the performance that it wants? So with that, we'll wrap up and end with that. Where can people find more information about deep storage LLC if they want? net. com because I've been writing there since the 20th century, because back when that was print and I miss being an ink-stained wretch. net. I really appreciate you jumping on on short notice.
We most definitely have to have you on in the future. There's a bunch of topics that I wanted to expand upon. But this is a reasonably live podcast. We'll talk to you anytime, my friend. Talk to you later.