CTO Advisor 061 - Lift and Shift Cloud Migrations
Packet Pushers Co-Founder Greg Ferro joins The CTO Advisor for an extended conversation on lift and shift migrations to the cloud. Keith and Greg shares the risks (People, Process and Technology) associated with traditional data center migrations that have similar infrastructure designs and footprints. The conversation moves on to the associated risks of migrating unmodified applications to the cloud. Is lift and shift a less expensive model than on-premises datacenter? What about sunk cost around licensing, hardware and facilities? Greg once said that only poor companies go to the cloud. Does that opinion still hold true? What about VMware Cloud on AWS? Show Notes Packet Pushers Human Infrastructure Newsletter Keith’s blog post on Lift and Shift two VMware Cloud on AWS
Transcript
Hey, you're listening to episode 61 of the CTO Advisor podcast. So we have more than a year of these things, at least weekly. So if you need to go back and find something to do for the next year, you know what, listen to a podcast a week and you'll be about two years back on technology. So we have a returning guest, Greg Farrell from the Packet Pushers. Greg, how's it going? I've been a long week this week. I've actually been out visiting a customer client.
I'd only have one or two of them left now, but just so educational to get out from behind the microphone and stand in the middle of a real data center and talk about real problems. Just a total experience. Yes, smelling the AC from the hot owl just does something to you energy wise. I don't know if you do that often enough, things happen. Trust me. All right. And that voice you heard in the background is our weekly co-host, Mark.
I gave Mark a second because, you know, he just ran from the car and he needed to catch his breath. Mark, how's it going? How is the week? I'm doing well. It's been a long week. Lots of traffic today. We have a big festival in town. So hopefully it's a festival that involves plenty of beer. It does. It's Oktoberfest. So, you know what? We really don't do Oktoberfest well here in the US. You know, in Cincinnati, we do it OK.
We're very German and we had the first Hofbrauhaus outside of Munich. So we like our German stuff, including our Oktoberfest. My hometown's having a beer week just because they're actually doing beer marathons. They've actually created London tube style maps of participating pubs. And you can actually walk 10 miles and visit 10 pubs. We do walk 10 miles. What would be wrong with that? We do chug and runs here. You chug a beer and run to the next one and then chug another beer and see when you pass out.
That's very American. In the UK, it's like, let's have a walk and drink all 10 beers. Nice. Yeah. Instead of trying to chug five beers and collapsing in five minutes. You got to get it done. You know what? I have a nice American style ale waiting for me after we're done with this podcast. And you guys are are in my way. So let's talk about let's talk about what we came to talk about, which is lift and shift.
So we've done you know, all of us have way more years in the data center than we care to admit. And we've done these lift and shift migrations time and time again, where we take a data center, move the data center down the street, fix the networking, fix all the broken static IP addresses, get the apps back up and running. We can probably do these migrations in our sleep, right? Still a lot of work, still a lot of work done. There's a lot of risk.
Yes, there's a lot of fear because you're not at all certain. We end up doing things like the last data center migration that I did. The customer wanted to keep everything the same. And the only thing that would be different is the actual data center. And they had a range of, you know, vendor A servers and vendor Y networking and vendor X networking services and vendors P storage array and so on and so on. And even though we were moving to the new data center, they insisted that we buy exactly the same brand of everything.
And when we got there, we found exactly the same as if we bought the brands that we wanted to buy, because we actually wanted to use different gear, because the new models didn't work anything like the old models. So even though they were the same vendor, and that was theoretically reducing risk, it didn't actually happen. And we had the usual massive problems where, you know, hard-coded IPs in the apps, you know, the storage functions didn't work correctly because the new storage array operated substantially differently.
When we migrated the app to the new data center and changed the public IP addresses, all sorts of things went wrong, firewall rules, load balances. And it's just, you know, this whole idea of migrating data centers, even when you work as hard. I mean, we spent millions in planning for this migration to move just from two data centers from the same to the same, as much the same as possible. And it could not be classed as a success in any way. Yeah, and I did to echo that, I did a SAP infrastructure migration for a fairly large company.
And Greg, we went a step further. We actually leased, we did a lease and trade of the exact same hardware. So if we needed G7, HP G7 servers. Now, granted, HP is now on generation 10 of their server. And at the time, the latest generation was nine. We went back and leased G7 servers, traded in the old stuff and migrated the new stuff. But SAP's migration strategy is that you have to reinstall the application and do a proper data migration.
So even with that, it was still a ton of risk, a ton of stuff that didn't work. So lifting the shifts, in theory, pretty simple practice, not so much. I think it's also important to note that we've all done these so much. We at least know the risks. We understand and we've quantified the risk when moving physical data centers. Because we've done it a ton. Well, we imagine we can, but we can't. Clearly, because Keith and I have just told stories of exactly the opposite.
Yeah, you know what? Exactly. We think we know the risk. And every data center move that I've done, there's been a completely different issue that I couldn't have foreseen after Greg, as you mentioned, spending millions of dollars on planning. Yeah, there was over 20 of us. You know, developers, infrastructure geeks, project managers, you know, and let's not count the vendor reps who are slavering at the trough, who couldn't wait to help themselves to sell us more gear. But, you know, there was just no shortage of resources to do this migration.
And it didn't just migrate over. There was so much going on and it just didn't work afterwards, really. You know, the bulk of it was there, most of it. But in terms of what the customers saw, it just didn't work. So we think we understand the risk. We think we can define the risks and mitigate them through, you know, design and research and discussions and meetings. And but the reality is we never do. But the flip side of this, of course, is that nobody gets punished when it doesn't work.
This is true, because, yeah, hey, we took the least risky route, which brings us to the conversation on cloud migrations. Every time I read one of these, I get I have to I have to calm down because I feel my blood pressure rising as I as I introduce this concept that a simple way to move to the cloud is to do a lift and shift of your data center to the cloud. We we just explained that just even doing a, quote, unquote, like for like migration is tough, but doing a lift and shift saying I'm going to run the stuff that I have in my data center, as is I'm not going to refactor the applications.
I'm just going to move a VM is a VM. It doesn't matter if it runs in my data center or if it runs in the in Amazon, easy to or Azure or Google compute a VM VM. Right. But it isn't right. It's not. The VMs are often connected on a very fast LAN network directly adjacent to each other. So. So when you want to do the motion around your data center for DRS, the latency tolerances are very tight when you move into the cloud, the latencies between any two VMs could be many tens of milliseconds greater.
They those data centers are physically quite large, often, you know, sometimes like a half a kilometer square, and if you're one VM is on one side of the data center and the other VM just happens to be instantiated on the other side, it's not uncommon to see them getting like 20 or 30 milliseconds delays because they actually have to traverse some sort of tree or they might be on a congested or in some cases, the VM might actually be running in a noisy on a noisy server where there's lots of neighbors.
You've got noisy neighbors and they're burning up the CPU power and you can't get. So then you go and buy, you know, you can mitigate some of this by buying guaranteed instances, but it still doesn't give you guarantees on performance, which is absolutely key to a lift and shift this, you know, the ability to sync VM so you can be motion them in the case of VMware. And the same thing applies to KVM is absolutely key. And your latency has to be not only small enough to make this happen.
It also has to be jitter free so that you actually get a consistent transition so that the chance of control happens at a time you can control. So, Mark, you're at Sissy Storage on Twitter. You're a storage guy. Yep. When we just talked about networking challenges, what are some of the storage challenges when you're talking about lift and shift? I think it's all the same challenges. I think anytime you're talking about lift and shift to the cloud, it's about latency between where your data or compute is and where it needs to go.
The challenges aren't going to change just because we're talking about storage. Now, you do have some specific ones. You may be tied to a certain type of array based snapshotting or something like that in your on premises workload. And that's going to change. But I think we know those things and can adjust accordingly. Yeah. And then there's just the change and change in this. And I think this is probably executives, I think, can understand this part of it.
The risk of changing and want a relationship and then to operating model going from not even talking about the financials, just going from having a team that racks and stacks, having a team that does break fix, having a team that does heating and cooling all the way up to the OS layer. The that that team just changes and that work changes, that workflow changes. Mark, you've been in vendor negotiations before. How important is retraining of staff when it comes to making a decision to change vendors?
I think it can be drastically important, but we've also mitigated that with things like virtualization and the commoditization of some things. But if you're a large enterprise shop, you've done things a certain way forever and your people don't know how to do things a different way often. So you've got to you've already got to work on that mentality outside of a lift and shift. And so it's extremely important to do that prior to that. But when it comes to lift and shift, it's even more important because you're going to have what was heating and cooling data center specialists.
And they're either going to get rift or you're going to retrain them to do something else. That's where TCO models come in, often talking about headcount reduction. You know, I don't see a lot of people actually doing that. I think they take a person and shove them into a new role and they either succeed or fail. I think more often than not, they fail because we don't give them the proper training, skill set and time to get to where we need them to be.
Greg, what's your thoughts on the risk associated the people risk associated with lift and shift to the cloud? I think there's a few things. One of the things is people lack training. In fact, what they lack is is a consistent ability to change and transition. We talk about technology being a change driven industry where we're all used to changing and retraining. But in fact, that's not so true. Although technology professionals do tend to transition quicker and adapt their careers and their outlooks fast, it's still enterprise IT is on a decade long transition cycle.
So if you suddenly transition from one major mode of operating your IT infrastructure, and we're just talking about infrastructure here, to another one, you're talking about a one in 10 year transition. If at the same time you're actually asking them to abandon everything they know about the way they run a storage array or the way they operate firewalls or the way they operate the KVMs. So you're sincerely asking them to transition everything about their knowledge and capabilities. And what I often find is that most IT organizations don't place emphasis on the impact of that.
Not only technically in terms of giving people the technical skills to cope with this, but also sociologically or emotionally. How do you take people who are literally watching their skills walk away from them, and they have to transition at a social level and an emotional level. The response to that is I now have to go and learn something new. Now, learning something new does happen, but not everybody is ready at a given point in their lives to go and throw away everything they knew and start it again.
So we underestimate the social impacts, I think, a lot of the times in IT. And IT managers are just so busy facing upwards towards their superiors that they rarely take time to have concern about the people involved, or they don't – they're not often able to adapt to that well. I think they like to play homage to the fact that IT organizations and enterprise organizations as a whole like to say people are our most valuable resource. They say that, but I don't think they act like that by and large.
They expect them to adapt instantly without ramifications to what they're asking of them. Yeah, so this is a plug for the Packet Pusher's human infrastructure newsletter. Greg, you guys talk about these topics all the time, kind of the human side of IT and infrastructure. This is where most of the time we succeed and fail with the people that we lead or the people that we work with or the people that lead us. Without that – without taking in that human factor in information technology, you're kind of – you're already starting from a position of weakness, so to speak.
Yeah, I do think people underestimate the humans in infrastructure. We're just taken to look like storage arrays or networks or computers, if you know what I mean, and we're just expected to intuitively – like I just read an article today talking about data center operators who are complaining that they can't find sufficient skilled staff. They can't find staff in the locations where they have data centers, and it costs too much – they want too much money to work for them. Like, A, these companies place data centers in remote parts of the world that nobody wants to live in.
No surprise they can't find staff, right? Yeah. If you've spent time and often money, but certainly time, skilling yourself up in a technology, you reasonably expect to get paid more. So why would an employer suddenly say it costs too much to buy skilled staff? It just – like the whole results of this survey sort of points to this lack of sort of awareness of the reality of this industry that everything about IT might be made of hardware components and so forth, but there is a key component here, which is the humans in infrastructure that nobody seems to be evaluating in a transition.
Sure, it's not fun emptying diesel out of your tanks once a year and rotating the diesel around so you're ready to go when the generators kick in, but it's also not fun being told by your boss, you know, we're transitioning from here to here, and your new skill set is this. Go. We treat it like a line item in a budget, right, literally. As a frontline manager, you don't see a lot of difference between that salary line item budget and all the equipment budget, so they get treated the same way.
And then usually there's one or two people in the team who embrace the change, who are comfortable with it or have experience from a previous role or willing to spend time or are in a position in their lives where they've got spare time, and then they embrace it and go and own it, and then they become the key point of failure or the single point because those are the people who then take ownership of the challenge and lead other people forward. Now, the question is, is that enough?
Quite often, those people carry the weight of the entire project on their shoulders, and it's usually not project managers. It's usually technologists or engineers, and they're the people who'll take the project to success with everybody following along behind them. But if you don't have enough of those people, then your project is almost definitely going to fail because your risks increase as your key dependencies reduce, like as you focus key dependencies in just a few people. OK, and then the last point in lift and ship is kind of the financial impact.
So, Mark, ma'am, you have started to kind of munch the numbers in the back end. It's really hard to make the argument that and I've had I've gotten to these arguments on social media with like AWS employees and maybe not so much Azure and Google Cloud Compute, but expect the AWS employees that it is cheaper to run a data center in the cloud with my existing workloads than it is to do so on premises. And, Mark, I don't I don't I don't see the numbers.
I don't buy it unless I'm getting brand new workload from nothing. Right. I'm a new company. I'm starting from scratch. I have nothing else. I don't have that sunk cost in a building, in a skill set, in licenses, in hardware and all of those things that factor in to a total cost of ownership of a package. I just don't see it. I mean, the math doesn't work. And Greg, you famously said that only poor companies go to the cloud.
Do you still do you still feel that way? Yeah, I do. I think if you can't invest in your IT or you can't see your way to doing the full numbers correctly. So it's very seductive. What I mean by poor companies is if you are so capital poor that a couple of million dollars spent on your IT infrastructure is a big number to you and a big scary number. Or if having 20 or 30 IT employees is a scary thing to you because you're so incompetent at managing people, then sure, going to the public cloud might make sense.
By the way, that summary just summarizes every startup out of Silicon Valley, by the way. Yes. And the thing to remember about a lot of the buzz about the cloud, a lot of the hype about the public cloud as such is that Silicon Valley optimizes for failure. That is when it comes to infrastructure for startups like Instagram, the way that they started in the public cloud was if you're a Silicon Valley investor, 99 out of 100 of your investments are going to fail.
You want to find one in 100 that's going to become a billion dollar plus success. There might be 5 or 10 others that get an exit and you get your money back but roughly 80% of them are going to fail. 80 out of every 100 are just going to be pack it up, walk away. So if you're a venture capitalist and you're fielding a portfolio, what are you optimizing for? Failure. Failure for shutting it down. So what you want is a month to month rental on an office block in Silicon Valley and a handful of Macs, some chairs and some desks which you probably bought second hand from the last failure.
And then if everything in the cloud – and then if the company fails in its second year of operation, just turn off the cloud, walk away. No servers, no lease contracts to terminate, no colo, no 5 plus 5 – so if you're poor, the public cloud makes perfect sense. And when I say poor, I mean poor as in time, poor as in capital to spend and you want to hold it back so you dribble out your opex so you've got a longer runway to survive.
Or maybe you're an established company that's actually going broke. Your market's transitioning into an end stage and you're going broke and nobody wants to spend $5 million on a refresh data center or buying some new IT infrastructure or even buying the right staff. Instead of paying good money for the right people, they want to cut corners, get second rate people and then bemoan the fact that their IT is miserable. Sure. So if your business looks like that, then also moving to the public cloud might look exciting.
If you're a normal enterprise with a good profitable business, a vibrant internal culture with strong leadership, then you'll tend to find that – and this is what I've experienced from several companies – that the cost of being in a private cloud is usually about one third, maybe about 40% of the cost of doing it in the public cloud. Literally, because you know what your workloads look like, you know how much capacity you need, you have people who can operate the infrastructure smart, who can extract value from those storage arrays and those networks and compute and those infrastructures, they can work with the apps that are in your business.
So we talk a lot about cloud-ready apps that are able to be 12-factor, written by DevOps, scale horizontal. Well, that's not what enterprises run. Most of them run SAP or Oracle or have middleware that's 20 years old. I once worked in a financial institution that's still running a specific type of minicomputer from the 1970s. It doesn't even speak TCP IP. Wow. Right? So I'd like to see you take that into the cloud. You have to replace that system.
Replacing that system was measured in the multi, multi millions of dollars. That's not a one year or a migrate to the cloud by next thing year type thing. That's a 20-year project to do that transition. So if you're a normal company with an established infrastructure, you've got some skills, but you've invested in your people whether to buy the right people or to train them and hold on to them and retain them, give them challenges. And if your business is profitable and you've got the money on the balance sheet to invest in your infrastructure, you'll generally find it half or less price.
If your management is generally dysfunctional, lacking skills, and they think that managing a data center and diesels and power facilities and racks and all that stuff, all that messy, messy stuff that we have, if you think doing that is hard, then you'll find a way to make those numbers say what you want them to say. So just to piggyback on this whole concept of it's not that simple. You know, I did the keynote at the Australian VMA back in early 2016. And that APAC region in general has embraced virtualization.
I think they I saw a stat that they're most of the companies are 90, 95 percent virtualized. So they're on X86, whether it's VMware, Hyper-V, KVM, Xen, doesn't matter. A VM will run in their private data center. It'll run in a cloud. And for whatever reasons, it's not just cost. It's not just one thing. They've on the whole have kind of ignored cloud. You know, it's yeah, it's similar, but not the same thing. There are some unique situations in Asia Pacific, and that is bandwidth related.
Yes. All assumptions around the public cloud assume that there's sufficient bandwidth to do the things that you want to do. And Asia Pacific is a heavily dominated by government controlled or semi monopolistic carriers who can, you know, do have very strong control over what's happening in the world out there. S. or in Europe where there's a lot of fat pipes going from place to place. Really, you can't move into the cloud no matter how hard you want because it's just not enough bandwidth.
Moving 50 terabytes of data around takes weeks over an Internet connection, which is 100 megs in size. So with all of this said. , VMware announced VMware Cloud on AWS. It's vSphere, the vSphere that in theory, the vSphere that I run in my data center is now vSphere on physical nodes inside of the AWS data center. Does it make sense to lift and shift from you guys see a scenario where financially, operationally, just kind of from a smart IT manager perspective, that it makes sense to lift and shift my vSphere from on premises to AWS or VMware's AWS based solution?
I'm going with a solid on the financial bit. It's going to depend. You have to look at your particular cost. But as I see it, no, I don't think it's going to make sense to ever do that, to lift and shift anyway. I think VMware on AWS is going to be almost the perfect DR strategy. That's a little controversial, but if I had to say there's a sweet spot, it's I can set it up as a DR. I can start setting up the, you know, the synchronization and the backup capabilities that I want there.
And then I've got a situation where over the next two years, I build a DR facility in AWS. It would not be a lift and shift. I think the sweet spot is literally a step-by-step migration from my existing DR facility, maybe shut down your second data center and then put this new one in the AWS. And then you can go to people and say, yeah, we're doing cloud. Look at us. Aren't we? That's the greatest thing ever. Yeah, Greg, I'm going to echo your opinion and have done.
I've done outsourced DR. I've done secondary DR and I'll look at the AWS solution with VMware. And it's very tempting. The it is. And, you know, I'll just throw names out there. I've done HPE's Helium service for DR. I'll just leave it at this. I don't recommend it. It is a tough solution to, you know, it's kind of virtualization, but it's not really VMware. I can't keep my workloads in sync from a config management perspective. Tempting, but not great.
What if your COLO contract's up? You've had a COLO contract for five years and you're up for renewal. And you're looking at a re-sign, say a five plus five, and the business has got something coming up and you don't know what the future looks like because your market's in transition or maybe your company's in a terminal stage of a collapse. You know, as an employee, it's always worth understanding whether your employer is actually doing well or not, by the way, just in case you think I'm being – you need to think – and, you know, so this could be, you know, you get into VMware AWS, you're committing to – what did we say?
I think I saw the numbers somewhere between 200 and 500,000 a month, which is roughly, you know – no, sorry, not – No, it's 200,000 a year for a one-year commitment. 500,000 a year. Most of the entry-level solutions are going to come in around a half a million. Well, a half a million a year for a DR that includes hardware, networking, and VMware licenses isn't that scary. No, it isn't. No, but it's all about how does that work? I see a lot of enterprise ITs, they're still stuck in reusing IP space in DR, you know, recovering a VLAN that has the same IP space.
How you do that there is going to matter. Mark, I've ran into – and these are the problems we ran into, like, Hylian. We wanted to just say, you know what? Our DR, our data center in one location is now our data center in another location. We bring up workloads, we bring up IP addresses and make it simple. It worked, but it was a hard – it's hard. Yeah, it's not an easy thing to do. It's a disaster.
It's not supposed to be easy. Well, it should be easy, but it's not. Like, the thing is that if you're – I think in terms of VMware and AWS, you are predicated on the fact that you have to have an SDDC. So that assumes that your primary data center or your primary facility, your VMware all the way from left to right, everything starts with the. So you have to be vSAN, you have to be NSX, you have to be running VMware ZSX, you have to be running vRealize, and that VMware and AWS will only work as a DR facility if you've got that in your primary facility.
And then how many people is that? Not that many. Exactly. Like, seven? There'd be a few hundred. Yeah, yeah, not many. I think VMware's had plenty of time to – VMware has had quite a bit of time to sell this out, and they talk about 4,000 customers of NSX now. So there's at least 4,000 customers out there who could have the potential with the networking, and you would have to be using vSAN, and you would have to be operating everything on vSAN.
Okay, well, there's 2,000 of them. So it's not a huge market in that sense. So if you're willing to make that transition, okay, I am totally with you, if you know what I'm saying. Yeah, I think the numbers were like 10,000 vSAN customers. But still, out of 500,000 VMware customers, that's not a significant number. No. All right, so wrapping this up. And then there's a few other things too. You're obviously in AWS. You then have to find a way to do a contract with AWS.
It's not ready. It's not fit for purpose yet. It's only for beta purpose. So you've got another year until this is really sort of mainstream. Yeah, this is one of my biggest beaten horses with VMware Cloud on AWS is that, you know, they still don't have direct connect options out the gate. For me, minimally, even in a DR scenario, I can't use the service today because as of today, only IP and IPsec is the only connectivity option. I'm not going to IP VPN into my secondary data center.
That's not an option for me. Well, I think it should be because as a networking professional, I think we're going to see the Internet used as a public WAN. So today you have a private WAN. Increasingly, you won't have a private WAN going forward. The only reason people use private WAN today is because they're used to using private WAN, and that's all they know. Now, there are some key advantages to using the direct connection services to Azure GCP or AWS.
But generally, it is wrong to consider that as the default solution. It should be an exceptional solution because everybody in the world is moving to using the Internet as a public WAN, and that's an industry-wide trend. So don't discount that. That will be an interesting conversation, follow-on conversation. Okay. So what I do think is much more interesting was VMware's HCX, was it the one that they announced? Yes, in Europe. In Europe. And that is fundamentally your vSphere as a DR facility.
So instead of putting it in VMware and AWS, you're talking about going into IBM SoftLayer or into OVH VCAN, the vCloud Air Network. So that means not NSX necessarily, not vSAN necessarily. It's basically just if I want to duplicate my ESX hosts into somebody else's data center. Now, of course, a lot of customers won't see IBM SoftLayer or OVH VCAN as cloud, and it's not. It's an ES colo facility, but it is in a way. So you could say that I am moving to the cloud.
But to do this, you have to deploy three appliances, the hybrid interconnect, the WAN optimization, and the network service extension appliance, which gives you the functionality to extend your data center, your ESX vSphere, into someone else's data center. That changes the conversation a lot. I still believe VMware's HCX is sort of like, I can't believe it's not cloud, cloud. It looks like a cloud. It smells like a cloud, but, buddy, it ain't cloud. All right. So let's close out.
We went a little long, and that's okay because this is a great conversation. I think it's super valuable. On the whole, we all are in violent agreement that lift and shift is hard to start with. Lift and shift to the cloud, that's a really tough sell, at least with this cloud. I haven't been able to find a guest who kind of says, hey, yes, I'm willing to come on your podcast and talk about why lift and shift to the cloud is such a great idea, and they didn't work for a cloud provider.
So there's that. I wanted to get this independent voice. So with that said, Greg, where can folks find more about you in the Packet Pushers as if, you know, they don't know where to find you? net. I'm on the Twitter as at EtherealMind. com. com. And, of course, you can find me on the Twitter, and my daughter hates it when I call it on the Twitter, but on the Twitter at CTO Advisor. com, where you'll find this and other delicious content.
Talk to you guys next episode of The CTO Advisor.