Selecting a Data Center Co-location

31:13 · Watch on YouTube ↗

Transcript 4,205 words · about 28 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

>> Alright, if you've been with us the whole day and you're watching this live, thank you for participating in our first-ever virtual event. It has been an incredible journey bringing this together up to this point and we figured we'd get ourselves on the agenda a little bit. Not too much CTO Advisor itself. I hope you've enjoyed the curriculum and agenda up to the point. And, I think, John Freymann of Freymann IT Services will bring is home for us. If you watched the Cloud Repatriation Panel, we ended off with kind of this cliffhanger.

What do you do if you need to repatriate workloads you've got now in your data center, and you no longer have a private data center to repatriate to? Well, John is going to help us understand some of the criteria for selecting a colocation provider if you decide to go that route. Take it over, John. >> Thank you, Pete. Yeah, we're going to be talking about how we went through the process of selecting a colocation for the CTO Advisor. A few months ago, Keith and I were talking about this new data center idea and it came to mind that he was planning on building something from the ground up, and I was just, frankly, a little bit appalled and concerned.

So I suggested we should look into colocation, and got in touch with a few vendors, and here's kind of what we went through in order to decide, is that the right solution for us, and what did we finally choose. We're going to talk about, initially, first, Keith's requirements. He'll be talking a little bit about what's unique about our environment, what he was hoping we would get out of this and how it would apply to the project at hand of hybrid infrastructure, and the migrations and the trials and tribulations of actually going live in that environment and maintaining and running it.

I did a little research and found some interesting things about some interesting industry observations, so we'll be talking about that. And it's a section I like to call, (speaking in foreign language) or Know Thy Self. Understand what you are capable of. What is your organization? What resources? What capabilities? What capacities do you have? And where are the capacity requirements within the computing world, and then, how can we cope with some that we may not have? Then, we're going to take a little bit of a look at the colocation providers, the critical factors, and finally, why did we settle on and actually start deploying, we're in the midst of that, with QTS.

At the end of the presentation, we won't be going over this today, but it will be in the presentation deck, is some appendices of additional details that we discovered in the process. So, Keith, I'd like you to talk about your requirements that we encountered here. >> So I had the initial idea. John talked about it. And I talked about this on social media if you follow me on social media. We were going to build a data center. Literally, we're going to take a card out from a retail office space in Silicon Valley.

We were going to record live video, pre-recorded video, kind of do the CTO Advisor thing right next to the data center. So, if you think about it, an executive briefing center. In the background, you'd have this glass partition and, in the background of that, you'd have the data center. And then, in the forefront, we'd have the studios, the CTO Advisor studios. With that, our business model was to take in sponsor equipment and be able to talk about HPE, and right next to HPE may be some Dell EMC equipment.

So we needed to be able to isolate those two environments from one another visually, not just physically isolate them, but visually isolate them. So I have some pretty custom and unique requirements when it comes to a cloud provider or a data center. Think about it. You know what, that's kind of a multi-tenancy thing, but I didn't have the scale of a multi-tenant infrastructure. So let's talk about the origin of why we're even building a data center and what drove the conversation and the requirements.

First, this is all about tell better stories. When I went to John and I said, "John, I want to build a data center", he was thinking traditional colocation services, which we, you know, look at a colocation provider, but specifically, managed services. So build VM infrastructure, and we'd go out and we'd sell services, VMs. Tim Crawford, in the Cloud Repatriation Session, talked about how he rented VMs from cloud providers before and this was just another one of those things. But it's not.

We're building an infrastructure, specifically, to talk about the enterprise journey to the hybrid cloud, this cloud repatriation conversation that we had where we ended up with a hybrid solution where you start out in the private data center, you move to the public cloud. Then, you discover, wow, we probably should not have completely left the private data center, and we need a balance. So we wanted to tell that story, the CTO Advisor data center journey. We wanted to take our audience on this journey.

So this is not just a one-month, two-month lab. This is going to be a year, multi-year journey, and it was all to just, again, enhance our brand and tell a better story. So how does that drive requirements? It needed to look and feel like a real customer's environment, meaning that I needed the production-like capabilities to at least look like production. I needed, you know, dual power sources. I needed multiple paths to the internet. But, at the same time, it couldn't be too new.

I did need modern hardware, because we're talking about the journey, not the destination. So I'm not going to have Intel scalable processor day one. So we went with some used Dell equipment. On top of that, it needed to be cloud-connected, not already connected to the cloud, but the ability to have cloud connectivity. Whether that was provisioning circuits, some type of cross-connect, high-speed one gigabit ethernet. And then, I needed this special physical set of requirements. One, I still need to be able to create the CTO Dose on-premise.

I needed to be able to go from the data center to my studio seamlessly. I needed to be have vendors in to have conversations. I needed a studio feel, but with data center capabilities. That included physically separating my sponsors who may be competitors. While the entire data center is sponsored by Intel, that doesn't mean that, when I go to get sponsorship from Dell, they want to see the three-par array in their video, in the videos we create as part of sponsored content with Dell versus the same thing with HPE.

So a pretty unique set of requirements with the whole purpose of telling a story, versus running production gear. So with that said, what's the different between production and a CTO Advisor hybrid infrastructure sponsored by Intel, we've got to make sure I get that in. Well, you could think about the, our environment, it's much more closer to a test lab, or non-prod environment. So if you're familiar with SAP landscapes, you know, you have the prod landscape, which is, you know, production-ready. It's high levels of uptime.

Downtime must be scheduled. The general public will be accessing the front end. I'm providing some type of service, or customers will be coming in and consuming their environment. So there was a bunch of redundancy required in productivity. In my environment, I just need to mimic that. It needs to look like I can have high-availability, but I can take it down at any time. We're not servicing customers. Other than my son's development projects, which can go up or down here and there, it really doesn't matter.

So lower uptimes were extremely accessible, or extremely permissible, or even, sometimes, even welcome, because I want to be able to do changes within the day when customers can come in and shoot video and we can make vBrownBag lives during business hours, versus using an environment after-hours. And then, there was just no public access, again, other than using the environment for my own pet projects, security and access is just something that looks different. >> So that led me to do a little bit of research, and we talked to a number of different partners and colleagues, and I found some interesting data center survey information from the Uptime Institute, their 2019 data center survey.

Surprisingly, 50% of services still remain in internal enterprise data centers. 34% of the respondents had a major outage in the last year. 50% had a major outage in the last three years. So an outage is a high probability. It's a flip of the coin over the three-year period. And of those who did, 10% incurred a million dollar loss as a result, or more, of those outages. The most interesting thing that I saw in the statistics were the following two, that 33% of those outages were power loss-driven and 31% were network access-driven.

So we're looking at 60% of your outages, of your downtime, is driven by power infrastructure and network infrastructure. So these are really critical. Also, the users said, 60%, if they'd done a better job of planning, configuring, and managing their environments, it wouldn't have happened. And kind of a side note that applied to our assessing, should we do this full cloud, do we want to do this co-lo, there's a lack of visibility and transparency of public cloud infrastructure. We can't tell who are the other tenants on the platform that we are and how their workloads may affect our throughput and our capabilities.

So those were all things, but the key thing here's that 60%, over 60% of the outages were either power or network issues. So what can we do to address that? Now, we want to move into what do we need to know to be successful? Know yourself. Know what you can do. Know what your organization can do. So let's start out by asking, what is the objective here? The bullseye is to provide your product, your data and intelligence, to your customers.

That's what's valuable. That's what your customers want. They don't really care about anything other than that. That's what they're buying from you. So what does this mean? Well, there are several layers here. At the top, the bullseye is your data intelligence, your product. Your product runs on some sort of computing infrastructure. Around that infrastructure is some form of security. Around that infrastructure is some form of networking and data communications in and out of your environment to your customer, to your trading partner, to your business partners.

Outside of that is yet another layer of security. Finally, all this stuff resides in some facility that needs to be secure. And that facility needs to be able to provide highly-reliable power, highly-reliable environmental or HVAC controls, fire suppression, et cetera. And you want security on that. You don't want just anybody getting into your equipment. You want to control who gets there, and you need logs, et cetera, to do that. Your business may require very intense levels of auditing on that if you have SOX compliance, HIPAA compliance, or any of those types of things.

So what does that mean? What people don't realize, that each of these seven areas have a full life cycle of services required to keep them up and running, and you're going to be investing in engineering, design, deployment, operations, maintenance and enhancement, migration, decommissioning of your product and how its delivered to your customers. But do you want the expertise needed to do the same levels of refreshing, migrations, your compute and storage, or your compute and storage security? Do you want to invest your resources and the similar engineering requirements for networking and network security?

Do you have people in your organization? Now, very large organizations probably do. They may have hired departments for managing facilities and facilities security. But smaller organizations won't, and they may fall behind and become noncompliant with some of their clients' requirements. So the key here is, you know, there are only seven areas, but there's 35 levels of expertise needed in terms of this. What can we do to reduce that? We'd like to go, taking a look at these, as I call them, rings of responsibility, data intelligence, the reality is, you are going to respond to that.

You have to take on ownership of your product and management and maintenance of that product, whether you implement, you know, an on-premise situation, colocation situation, or the cloud. You're going to own that. Now, when we move to compute and storage, we have some choices. The cloud takes on all of that hardware responsibility for you. You no longer have a responsibility for it. They may even take on some of your operating system level and you're only dealing with it at an application level.

In a colocation, you're still going to have that responsibility. There are co-locators that will give you bare-bones hardware, where you can, potentially, rent that on a, you know, per-week, per-month basis, or part of your leasing agreement, but you're still going to have the responsibility for that, and similarly for the compute and storage security. When we move into networking, some co-lo providers will provide you a turnkey network and security solution so you don't have to have that resource on your staff. Now, depending upon how you use your application, how your deliver it to your customer, it may be wise for you to have it in-house.

But if you're doing it on-prem, you absolutely have to have those. You do not have an option. And similarly with the facilities. In a facility environment, the co-lo provider in the cloud, they're going to provide and take care of all of that so you don't have to worry about being current with the latest biometric reader devices, tracking logs, keeping background checks on the people who are coming in and out of the facility. They will take care of that for you. But, if it's on-premise, you have all of those liabilities and responsibilities.

So when we were looking at what we need for our hybrid infrastructure data center, the CTOA needed to stand up a data center quickly. We don't have the skills and experience, or the, more accurately, we do have the skills and the experience, but we don't want to expend time and energy acquiring a facility, developing it, managing it, building it out, and maintaining it. We'd just as soon hand that off to someone else to do for us. We want to outsource the responsibilities for all the public power, public networking, and building environmental services.

We don't want to have to worry about fire suppression, so we need to get rid of this, you know, ancient, old Halon system, or something along those lines. We want someone else to take care of that for us, and we want to leave the implicit engineering understanding of all of the redundancies and, you know, do we have multiple physical entries into the building from different areas of the build for power, for communications, et cetera? We want to leave that to reliability, through experts that do that every single day.

And, by going to a co-lo provider, we have offloaded our responsibility for that and we have contracted, effectively, with them to do that work. So let's take a look at the colocation providers. Let's look at some of the critical factors and trade-offs that we considered when we made our selection on our co-lo provider. One of the things that you'll get into when you're looking at colocation is what is the Uptime Institute's tiering system, where Tier One is the least-reliable platform and environment and Tier Four is the most.

Not going to spend a lot of time on this. Many of you are very familiar with this, but the bottom line is, the question you want to ask, can I live with 28 hours of downtime a year, and that's what you're looking at on a Tier One environment, or do I need something more reliable which I know will only be down a maximum of 26 minutes per year, versus 28 hours. So be aware that every colocation provider and, in fact, cloud provider will offer different levels of tiering, and this can impact the pricing.

The pricing of their services, more typically, is regionally driven by their competitors. One of the huge surprises we had, at least, I had, I'll be frank with you, is don't assume all of the different data centers within a particular co-lo provider are of the same tier or offer the same services. The national brands like QTS are pretty consistent across all of their data centers in terms of the levels of service and the reliability and the availability. However, more regional types of co-lo providers, some locations may only be Tier One, versus near a major metropolitan area, they may be Tier Three or Tier Four.

So you need to be very cautious about drilling down, specifically, on your location. Also, I saw variability on the levels of services they offered with just a carrier hotel, in which case, you need to negotiate with each individual carrier for your redundant services for communications and networking. Or, were they are offering a combined service like we have with QTS where they're using two different global Tier One providers and offering us a port, and they take on the responsibility of the reliability of that network connection to the internet.

So be sure to drill down when you're looking into this with all of your clients, or with all of your data centers that you're looking at to locate. Also, if you think you may be growing into additional data centers in distributing your services, you want to be consistent across the centers you anticipate. Some other more global risk and access issues to consider. How close are you and your IT staff or the MSPs that you are using to the data center itself? You're going to have to be going there.

There are going to be times when you've got to swap out a disk. You've got to go in and put in new equipment, take out obsolete equipment, et cetera, or do physical audits and that type of thing. So you want some place that's reasonably close to your staff and your support organizations. Other things is worry about geologic and weather risks. There was one client who had built their data center and put out a colocation, as their disaster recovery plan right on the heart of the Madrid Fault in the Midwest here.

So if there was a major earthquake event in the Midwest, likely odds were, their disaster recovery center would've been taken offline. So worry about things like that. Or, is that data center on a flood plain? Is it in Tornado Alley? Is it in an area that's, you know, likely to be hit by hurricanes? Is it in the desert where temperature extremes are going to potentially cause issues with their HVAC? So consider these outside risks and make sure you look into those, also.

In addition, they need to have redundant access. Are they on redundant power grids? In our situation here, yeah. All the data centers in the Chicago area, we found, were on a loop. In fact, when we selected QTS, one of the things was, they are on a very high-powered loop, on the high-tension loop, and have their own power substation onsite on their property, run and managed by the local utility. Are there diverse redundant communication hubs? Are they on the metropolitan key loop into all the critical data centers, or data carrier hotels within the area?

Or, are they just on a tributary lane where there's a single point of failure? And then, what is the latency to your cloud providers? These are all significant, important things to consider. Another thing that surprised us is the concept of remote hands. Nearly all colocation providers are going to provide you with remote hands. Be cautious. Remote hands are just that. And I think, the best analysis is they are IT puppets. They are there to be your eyes, ears, and hands, but they're going to do strictly what you tell them to do.

MSPs bring in experienced engineers. They're going to be observing, evaluating, recommending actions. The distinction is, MSPs will provide you critical intelligent IT thinking and assessment and recommendations. Remote hands are strictly remote hands, someone there to type the keyboard for you, to plug a cable in for you, to power something up or down, to swap out a piece of hardware if you have spares onsite. So there's a very big distinction and you need to lower your expectation of remote hands. Now, some colocation providers may have a higher quality of services but, in general, you need to recognize, remote hands is not an experienced IT engineer.

So finally, looking at critical vendor areas, what were the things that were important to us? They need to meet our immediate goals. We were looking for them to be agile enough to respond to any future evolution of our requirements, or future requirements that we hadn't anticipated. We looked at vendor responsiveness. How quickly did their sales staff get back to us? How quickly did bids get resolved? How thoroughly did they express their knowledge of their engineering environment and their technology and their capabilities?

What was the breadth of their services. Were they just really providing a place for our rack with reliable power and reliable HVAC, or were there additional services that they could provide in terms of enhanced network access, and enhanced internet access, direct high-speed connections to the various cloud providers? Were they aligned to our mission in terms of what your company is doing? That's important on a couple of levels. If you're an organization that's highly agile, entrepreneurial, dealing with an organization that's going to be very, let me say, bureaucratic, will be a challenge.

It will general friction. However, if you are a bureaucratic organization and you're dealing with a company that is providing the services, it's more entrepreneurial and doesn't have the processes and the structures in place that your organization requires, you need to take a look at those kinds of things. Also, what's the organizational cohesion within the vendor? Are they all singing off the same song sheet? Do they all understand the capabilities of their environment? Is everybody telling you the same thing, or are you getting a lot of different stories depending upon who you talk to?

And finally, how flexible are they? We had some unique one-off requirements in terms of our video requirements and our recording requirements, and frankly speaking, a lot of colocation providers do not want cameras in their data hubs, and for appropriate reasons. We were able to establish policies and procedures that allowed us to meet those one-offs. So the vendor proved to be highly flexible. So here's why we ended up with QTS. They were rapid in their bidding and sales process. They provided unique flexibility for us in that, we may need racks out in Silicon Valley since many of our partners are in Silicon Valley, and we would be able to put them on our LAN.

We would be able to extend our LAN from Chicago to there. So that was a useful thing. The communication within QTS, from the marketing level, through the sales level, to the engineering level, to the implementation and deployment staff and operations has been coherent. There's been no finger-pointing between those groups. They have been good at communicating with each other. And they offered a range of services including L2, Layer 2 level networking between their facilities with a higher level of uptime than we could get from other vendors, and there was actually no cost differential even though we were getting a higher tier of service.

So attached to this document is an appendix that goes into more details on the pros and cons of on-premise, colocation, and cloud, some of the elements that were in our colocation bidding process, and some of the bullets of what we went through with the contracting process, and some of the surprises that we saw there. And that's, basically, what we went through. So I thank you for your time and attention. I know that this is the end of the day. You're anxious to get our of the room.

Keith, I'll let you finish with any parting words. >> It's been a long day. Follow us on Twitter, @CTOAdvisor. com is the website. Thanks for attending the event. (gentle tones)