VSAN 6.2 What you need to know - CTO027
Wikibon has predicted that Server SAN will overtake traditional storage arrays. The latest version of VSAN 6.2 moves a step closer to displacing legacy storage arrays. I’m joined by VMware Principle Architect Rawlinson Rivera (@punchinglcloud). We hit some critical questions about the latest version. Here’s a short list of topics we touched upon. Value proposition of VSAN What’s new Mission critical application support Policy based management for application such as SAP Erasure Coding Why no deduplication for magnetic disks Subscribe iTunes | RSS
Transcript
OK, I'm really excited about this podcast. If you haven't heard it in my voice, I'm joined by VMware's Rolison Rivera, aka Punching Clouds on Twitter. He's working in the office of the CTO of VMware. He is probably the most passionate cheerleader for VMware vSAN. We have a very wide ranging conversation about vSAN, about its capabilities, about what type of workloads should be placed on vSAN. A lot of the new features. We go over this whole erasure coding, RAID 5, 6 thing that we've been saying on Twitter.
A disclaimer, when I originally recorded the podcast, I used the wrong mic. So I recorded many of my own parts. I didn't do it verbatim, but I wanted to keep the content and make sure that I didn't lose any of the fidelity. And most definitely deliver a high quality product to you guys. Hope you enjoy the conversation. Rolison, go ahead and introduce yourself. My name is Rolison Rivera. I'm a principal architect in the office of the CTO for the Storage and Availability Business Unit at VMware.
Office of the CTO, what is it and what do you guys do? I work with some pretty amazing people, sort of responsible and in charge of leading some of the evangelism across the different products and technologies that we have. Being a customer and advisor to some of our customers, larger customers, and bringing forward some of the new technologies that we're bringing, making sure that we're putting the right things into the product and also working on some things that are not seen and won't be seen for a couple of years down the road, which is pretty exciting, pretty cool stuff.
So VMware Virtual SAN has been out for a couple of years. A lot of us really still don't know exactly what is a vSAN or a virtual SAN. Can you spend some time giving an introductory to the concept of virtual SAN and kind of the history? Well Keith, virtual SAN is a software-defined storage solution that is part of VMware's newly packaged hyper-converged software stack. Basically vSAN is a product that lives within the hypervisor and is able to manage the storage as an extension of the platform in the way that we take locally attached storage devices in many media types and forms, flash-based, magnetic devices as well.
It pulls them together in a way that they can now become a shared resource in an infrastructure where, if you look at some of the things we did in the past, that a lot of the VMware enterprise features, vMotion, DRS, and all these things, were designed depending on the capabilities of a storage infrastructure that was shared. Now we're dealing with this new sort of approach where you can consume locally attached storage devices and deliver the same benefits that you have in the infrastructure and still be able to leverage the entire stack of products and features that reside within the VMware portfolio.
Basically x86 storage, you know what, a bunch of us are a little bit hesitant about putting our mission-critical workloads on x86 storage. Is vSAN a replacement for tier one storage? What exactly is the target workload? I mean, this is a topic of discussion that, you know, people talk about a lot, and in all honesty, as much as I am a cheerleader and I believe in this product so much, I now sincerely believe I can do that because of the amount of performance, the capabilities, and the reliability of the product it is and how it's delivered.
It is truly at the point where we can do that. We have all the necessary components to actually, you know, rightfully state, yes, you can. We can scale to the numbers that are needed. You know, we can deliver the performance and capacity and all the most extraneous demands from any sort of business-critical application today with the way the product is today. Okay, we'll get into the details of the resiliency a little bit later on in the discussion, but what's the high-level value prop and the big changes in this version of vSAN?
Well, Keith, for anyone who doesn't know much, or even for those who know about vSAN, the beauty of this release is that we're bringing a lot of new technology to us, not necessarily new technology in the storage market, but in the way in which we're delivering it and what we're doing with it is incredible because we're dealing now with things that are affecting directly the financials and the performance of the product. So, you know, there's a lot of things that are being put out there and situations and the things that we have to plan for in today's world of the data center and some of the requirements that are being put out there, and a lot of this stuff is being done around data efficiency.
So, you know, per se, we talk about some of the things that are being delivered, and they're not entirely new in the industry, but in the way in which we're doing it for vSAN specifically, what we're enabling is pretty incredible on what we're doing in that space there. So, you know, when you think about, you know, we talk about our data efficiency technology capabilities, and there's a few things in there. They go from core features that are done or applied into the product themselves, or there are things that are just added new to the product which weren't there before, such as deduplication and compression.
But we also expand on other capabilities of the platform, and we enhance things like much lower level details, such as, you know, using data efficiency technologies and capabilities on the actual swap files themselves. So, all of this translates into cost savings, better results overall in terms of what can be done in the data center in terms of managing and dealing with storage. And overall, it translates into a lot of cost savings, which is at the end of the day, it's what a lot of our, you know, our people, our customers are trying to look at ways they can do a lot more with spending less.
And that's some of the things that we're actually trying to get into right now. Okay, looking over the feature list, let's get back to this mission critical application piece. One of the things that caught my eye, support for SAP? Yes, sir. And I know that's very, you know, very close to your heart right there. Very much so. Me being in my day job as an SAP infrastructure architect, the fear of SAP going down, the thought of putting SAP on a hyper-converged infrastructure bring my anxiety level down a notch.
Talk to me about what support and performance looks like on vSAN. Well, Keith, from a performance perspective, you know, now that, you know, one of the things that are new in vSAN in this release is that we're basically targeting and focusing as the all-flash architectures being the default deployment for it because of all the performance that it delivers and all the features that are specifically attached to it. When you think about, you know, what an application such as SAP and all of its components, what they need, you know, these are things that are relying on heavy performance.
But a lot of times we tend to focus on performance in the reference of IOPS. But one of the things that we are very keen on delivering with the all-flash architecture in vSAN, it's a sub-millisecond latency that's really what's important here. Overall, you have the availability, which is not only driven entirely in a software model where you can, you know, as some of the things I've been saying lately, software-defined availability and software-defined performance, where you can define it very specifically on the individual components of that application that you may want to run.
And be able to guarantee that in a way that from a performance and availability standpoint on a platform that can allow you to scale at a cost-effective, in a cost-effective way, which I think is, you know, obviously very important. These applications are, when it comes to mission critical, these are the ones that actually fit that model specifically. And in order to satisfy that, we have to be able to deliver all of the goods, the performance, the availability. And actually, when you look at it, the cost efficiency as well.
There's debate about what's software-defined storage and what's not software-defined storage. From a high level, I don't question whether or not vSAN is software-defined. One of the things that you guys have done well out of the gate is able to manage the underlying infrastructure based on policy. I'd like to hear in a little bit of detail how vSAN manages applications and specifically infrastructure with policies. Give you an example. Let's say that I have SAP because we're picking on SAP today and we have a backup application and there's contention.
A backup application runs overnight. Then there are some SAP batch jobs that need to run. How do I ensure that SAP gets the reliability and the resources from a performance perspective that it needs? This is right here. It's one of those things where this is why I get so pumped up because you get right to the things where we're probably the best at. And this all comes from, you know, not just so if you think about it from a more logical approach, you know, you need to provide all of this, all of these capabilities and services to this application for multiple reasons, right?
Just like you said, you know, performance-wise, but also let's think about, you know, the operational efficiencies and the risk mitigation strategies that you can take with this approach. And here's how vSAN lays it out. Today, one of the things that the product introduces is this so-called quality of service capability, which ties into your particular question right now. If today you wanted to guarantee that your SAP workload gets the, let's say, for example, the right number of IOPS in a point of contention, let's say month end comes in and obviously there'll be a lot of things contending for these resources that are located in that solution.
Within virtual SAN today, you can actually, through a policy, which is defined specifically on individual objects, not the entire VM, but maybe on a set of specific disks, right, which are represented in the form of objects. And you can actually enter the maximum amount of IOPS that each one of these devices can get. And you can use that to basically fix a sort of contention mitigation that happens as a contention mitigation strategy for what's happening in a certain time in the month where your application in a backup product may be contending for the resources that your application should be getting.
That is sustained through the entire lifecycle of the application itself and is monitored automatically from the system. So when you talk about and you look at being able to deliver guarantee service levels, which are systematically done without having some sort of human intervention in that sort of level of guarantee, this is what vSAN delivers today. And this is why in that particular aspect, we're pretty happy and satisfied with what we're able to do there. Enterprise storage administrators have their own culture and let's call them idiosyncrasies where they're very particular.
And one of the knocks against hyper conversion in general has been that there's no built in protections for let's say let's just outright and say it's stupid. You know, human error is one of the primary things that defeats great system design. And storage tier one storage has protected against tier just human error for quite some time. How does vSAN protect against human error of a administrator shutting down a vSphere server that was part of the resiliency of application? Well, this is one of the things where, you know, it's a it's you kind of cut a heart like they say in a rock in a hard place, right?
Whatever that term is practically used. But here's the deal. We are very mindful of those scenarios, which is why even in the very beginning, when vSAN first came out, we didn't really expose certain capabilities which were originally planned for, such as the one you see today with with the whole IOPS thing. Because these are things that could actually hurt a customer. They can hurt themselves when they're not have at least a certain level of knowledge. So in the aspect of availability at a very basic level, if you deploy any given application on to vSAN today out of the box, you have an availability ratio of plus one.
So if you were to say anything happens within the system, anything at any level, any sort of failure, whether it's a physical device, whether it's a network, including even a site these days. If you make a mistake of a single failure, in that sense, you're still protected, right, where your data will still be available. Now, the only aspect here that changes, which is in a case that we're able to leverage some of the existing technology that live within vSphere, in this case, HA.
So not only do you have hardware network and all this availability from vSAN natively itself that protects the admin from doing some of those things. But also, if the mistake is made, let's say particularly on a host where it's taken down, any application, any VM that is running there will come up on the other side. Will come up on another host as part of that cluster where that's a function of HA when it's treated in such fashion. The data will still be available because it's placed in multiple places depending on the availability requirements that are placed in that particular application.
And by default, the system gives you empty failures to tolerate equals one. And you can expand upon that as you need to now in the system as well. Let's stay on this availability topic a little bit longer. 2 was erasure coding. And then there was something in there about wait, rate five, rate six. Help set that up for me. Help explain this. What is erasure coding and how does it relate to rate five and six? Well, yeah, so basically, so when you look at it, erasure coding, in a sense, is just a way, a scheme of sort of encoding data into fragments that would allow a system to recover in the case that something is lost and it fails.
Particularly in our case, what we do is that we implement it in two different ways. We do in the traditional rate five, which most folks are particularly familiar with, and also in rate six, depending on the implementation that and the availability you're trying to have. So when you look at it, one of the things that we're doing here is that one of the things I like to say is that we deliver space efficiency without the risk on availability. But when it comes to leveraging this sort of distributed way of not only delivering space efficiency, but also keeping your data safe, it comes, obviously, in a way that you have to now have another host, potentially, depending on the type of erasure coding you may want to implement.
And this comes in the form of the number of the availability for the applications that you're trying to choose. So, for example, if you look at when we traditionally, without using this particular policy or this capability, we would do replication, right? Where replication won't be as efficient as it would be in the sense of space savings because we would do an exact replica copy however many times we need to in order to guarantee data availability and accessibility. Here, we obviously take a bit of a more efficient approach where we distribute the parity across all the different nodes that are within the system.
So in order that we can not only leverage the ability to use less space, but also now be able to implement and have the same availability requirements you would have for that particular object or the amount of data that you're trying to cover. In a sense, there's a lot of different types of implementations when it comes to erasure coding. Obviously, RAID 5 is one of them, RAID 6 is another one, all basically based on Reed-Solomon code. And some of these things are basically just ways of argumenting how the data is placed and some other calculations, which actually get pretty complex.
And I think that now that we have this feature, this capability, we're going to be doing a lot of education around this. As folks are asking already, what is it that I'm getting from this? How is this really being done? And it has to do with the way the data is placed and how we're recovering in the event of failures. But overall, keep in mind that this is just another way of us delivering space efficiency across the entire cluster and some of the new ways we do that to deliver a much better cost and value for vSAN itself.
Okay, so let me walk this back a little bit and break this down just a little bit because I think it's worth the time. So before, I would protect the application based on the number of replicas. So I could say, you know what, these six VMs make up the application and I need to have at least one or I need to have at least two copies or two additional replicas of this data. So I can survive up to two failures of vSAN nodes.
So obviously, I have a great more deal of redundancy. However, I'm taking one to one copy. So I'm basically using three times as much, at least three times as much physical disk space to protect the application. So with erasure coding, I can say, you know what, here's the level of availability I want. But instead of using replication as a mechanism for protection, I'm using rate like capability or parity data to protect the application. So I don't have the inefficient use of disk space.
Correct. So what's really amazing about that, what's impressive to me about that is that that's on an application level and not on a disk volume level. So I'm used to seeing the erasure coding based on volumes, but this is based on application. So I can, you know, I can get as granular or efficient as I want when it comes to my replication or is that one as I want when it comes to my protection scheme based on that subset of data that I need.
Absolutely. And just to get a little bit deeper on that. So, you know, we say about an application, but when you think about an application, what makes up that application? How many data points, how many repositories of data is it kind of pulling from a particular VM? You have the ability of doing that on a per disk basis within AVM, meaning that if depending on how valuable on the type of availability you want to get from that particular application. That VM, you can have different settings individually configured to each one of those disks at the same time.
You know, something that is relevant to the type of to the amount of data and the consumption that is being done. So when you think about, you know, a lot of these applications, let's talk about SAP for a second. You know, these things are really mission critical when it comes to that. As I said before, that's the king of the crop right there. So in the example of SAP, I can have my what we call SAP data, my SAP data protected.
I can take up that protection level, but the app server, which is a more of a ephemeral workload. If that goes away, if the disk for that goes away, I really don't care. I can I can spin up another SAP application server and life goes on. So I can get a granular saying, you know what, my SAP data specifically within this V app, protect that to the nth degree. Whereas the application servers, which are not as critical, you know what, I can have a protection level one or two or whatever the increments.
And that is exactly it. Right. And even in that particular case, when you get into these ephemeral type of scenarios, you have the ability of saying failures to tolerate zero. So that you don't have to have an extra copy, because remember, to use the razor coding, that capability, there's a couple of new requirements. So at a very minimum, you'll have to have four nodes to do rate five and to do rate six, you have to have six nodes within the cluster.
You don't have to expand in magnitudes of six, but within the cluster, at least that we can do that. But in the case when you're going to do a traditional deployment, have the failures to tolerate one, you have a replica copy of the same. But in this particular scenario here, if you like to, if it's something that fits within your standards and you're in the deployment of that application for support from a supportability standpoint, you could even change it to FTT zero. Which means you will not have a replica copy of that data because it's an ephemeral type of workload or an ephemeral type of amount of data.
Whereas on your data itself, you can go as far as failures to tolerate to any two failures simultaneously within the system. Right. You will be totally safe, which basically in a traditional deployment of vSAN without using this new capability. You basically have three times the amount of you consume three times the amount of data because we would have to have three replicas to protect that particular workload. When you do that now with a razor coding and you choose, for example, failures to tolerate two, which would be the highest level you can do it.
You have about a 50 percent less consumption of storage on that particular deployment, which now delivers the extreme levels of high availability that you want at a much more reduced amount of data capacity consumption. And that's where some of these things that we're doing are very, very proficient and very valuable to today's application. OK, so last question is on deduplication. Why only flash only with deduplication? That's an excellent question. 0 and folks were asking, yeah, we could do a flash and we want to have deduplication and everywhere.
And the reality is that, look, I take a pragmatic approach to this. To do the duplication, we all know that these type of data services from anybody and everybody is applicable to everyone. These are these are data services that are pretty heavy and expensive in terms of resources and what they do and what they consume. In a software defined model like the one we're deploying and implementing here, you know, if we were to deliver the same type of architecture and the same implementation as we do that in a in a traditional legacy or type of the traditional array world where these services are offloaded and dedicated to a piece of hardware which don't introduce overhead contention to your to your virtual machines or your applications.
In a hyperconverged infrastructure, this is totally different because we have to be conscious about about how we do this. So how do we tie that into what we do today? Only in O-Flash, because if you think about it, if we do this on a low performance type of hybrid configuration and still when I say low performance, because obviously you can't compare hybrid to an O-Flash configuration, that's obviously going to perform much lower. If you are a customer and you're thinking about a solution that you can scale, expand and expand as you need to in a cost efficient way, what is the most expensive way for you to scale?
Either you do that by buying another box because you're consuming all of your all of your resources because you're doing all these data efficiency services and the dupe and compression, or you want to scale by the disks. When you compare the cost there, obviously there's a very different impact there. Now, because in O-Flash, you are able to take the performance hit in the overhead because we have now the ability to deliver more performance at the right cost in terms of resource contention. The way in which we do that today, compression and duplication, we do that in an effective way, which doesn't increase, it doesn't have the impact and the amount of resources that it consumes in overhead so that a customer and obviously everyone can have their applications consuming the majority of the resources available in that converged infrastructure, that converged hardware.
Those are some of the things that we're looking at. When we're talking about do we do it on hybrid, do we do it on O-Flash, there is a pain that our customers will experience. In our case, we chose the path of going with O-Flash. We make it available so that now flash becomes the default way of going with vSAN because of the cost. You have the ability now that you can build these environments with two different types of devices and the cost of these devices are dropping exponentially.
They're going to get a lot larger in terms of capacity. When you do the math and you compare, what's the most valuable potential implementation for this? It makes the most sense in a hyper-converged solution for our customers to be in an O-Flash. Now, that yet again could be also questionable where I want to have it on hybrid. What happens to the customers that have it on hybrid? You know for a fact that if you run high compression into duplication in any sort of solution, particularly when it runs in software and it runs competing with VMs and all that kind of stuff, there's going to be some contention.
You're going to have to deal with that. Now, the question is which one makes the most sense. It's up to you. You can use it. For us, we think we deliver a much better value for our customer when we have it in O-Flash. So thanks for that. I think it was worth going in that detail because I've seen that question asked pretty often on the Twitter and you did a great job explaining the rationale on the design considerations and approaches.
Speaking of the Twitters, let's wrap up the call. Where can people find you on Twitter? I'm at Punching Clouds or you can find me also at my blog at Punching Clouds. And if people want to know more about Virtual SAN, where can they find out? You can follow Cormac Hogan at Cormac Hogan. That's his Twitter. Duncan YB, which is Duncan Eppin. Myself. You can visit the Virtual Blocks blog that we have for the VMware official blog. You'll find a lot of the technical marketing guys are blogging on there.
They're providing a lot of information and there's a lot more information to come out on Virtual SAN. And all these different topics. And we're going to make sure that everyone gets the most out of it. Not just from a technology standpoint, but from a cost perspective and a value proposition perspective. Because it's more than just the technology. It's what we do with it, how we implement it, which actually delivers a bigger value to our customers. Rolison, thanks so much for taking out the time to chat with me.
com. On Twitter, I'm at CTO Advisor. Share it with your friends. Like us on iTunes. Spread the word. Thanks.