CTO Adovisor 100: Bare Metal Cloud

RackN CEO Rob Hirschfeld ( @zehicle ), joins the CTO Advisor for episode 100!!! (Sorry Chad, we’ll get you on for a ceremonial episode 100). With the growth in interest for VMware Cloud on AWS and AWS bare metal cloud, what is the market for bare metal cloud and automaton? Rob attempts to argue the value of bare metal automation. Rob makes the argument for Fry Cooks racking and stack at the edge. The conversation is wide ranging to why OpenStack didn’t work and Kubernetes at the edge. The CTO Advisor CTO Adovisor 100: Bare Metal Cloud Play Episode Pause Episode 1x 00:00 / Subscribe Share Apple Podcasts Spotify RSS Feed Share Link Embed <blockquote class="wp-embedded-content" data-secret="B73jSC8e4i"><a href="https://thectoadvisor.com/podcasts/cto-adovisor-100-bare-metal-cloud/">CTO Adovisor 100: Bare Metal Cloud</a></blockquote><iframe sandbox="allow-scripts" security="restricted" src="https://thectoadvisor.com/podcasts/cto-adovisor-100-bare-metal-cloud/embed/#?secret=B73jSC8e4i" width="500" height="350" title="&#8220;CTO Adovisor 100: Bare Metal Cloud&#8221; &#8212; The CTO Advisor" data-secret="B73jSC8e4i" frameborder="0" marginwidth="0" marginheight="0" scrolling="no" class="wp-embedded-content"></iframe><script> /*! This file is auto-generated */ !function(d,l){"use strict";l.querySelector&&d.addEventListener&&"undefined"!=typeof URL&&(d.wp=d.wp||{},d.wp.receiveEmbedMessage||(d.wp.receiveEmbedMessage=function(e){var t=e.data;if((t||t.secret||t.message||t.value)&&!/[^a-zA-

Transcript 4,095 words · about 27 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Hey, it took almost, oh man, I think three months, Mark, but we're finally at episode 100 of the CTO Advisor. Don't say it, we can't be at 100. I think we're still at 99. You know what, we'll have our ceremonial 100 with Chad at some point, but we are at, this is actually episode 100 of the CTO Advisor, but we did not disappoint with the guest, Mark. No, we did not, I'm excited. We got Rob Hirschfield, co-founders, CEO of Racket.

Rob, welcome to the podcast. Keith, Mark, thanks for having me on, this is exciting, and being able to break into triple digits with you. Yeah, Rob got me in a little old fight. Rob got me in trouble with Dave McCrory because he had me on his podcast, Shiny Latest Pot. What is it, Shiny Latest? The Latest Shiny, L-A-T-I-S-T, yep. Yeah, he got me in trouble with Dave before later. Mr. Data Gravity himself. Mr. Data Gravity himself.

But it was a very, I thoroughly enjoyed the conversation, and I didn't realize that I haven't reciprocated and had Rob on our podcast. It's always super fun to talk to Rob. Rob, first off, introduction, tell the folks who don't know about you, who are you and what do you do at Racket? And tell me, because I don't know you, Rob. This is the first time I've talked to Rob, so I'm really excited to hear and have this conversation. It's fun to make new friends.

Boy, my name's Rob Hirschfeld, I'm CEO and co-founder of RackEnt. We specialize in data center automation, and which is confusing to people, we are on the physical layer of data center automation. So my past with, actually, Dave McCrory, I was co-founder of my first startup, where we actually built the first clouds, literally, 2001 ESX, the first betas. And I've been sort of in the data center operations on automation space since. About 10 years ago, we were in a team in Dell, where we were building, trying to ship clouds from the factory, got involved deeply in OpenStack, where I was on the board for a long time, in Hadoop, lately in Kubernetes, you know, sort of trying to figure out what made it so hard for people to bring up infrastructure and manage data centers.

And RackEnt came out of that. So four years ago, we left Dell and started trying to make this work as an independent software product. And that's what RackEnt does. We're a software-only company that, you know, sells software that helps people automate the physical layer of their data center, and beyond. We actually go way up the stack, too, for people. If you get the bottom right, everything else is easy. So, maybe back in 2008, 2009, 2010, I'm getting old and it's hard for me to recall back, Mark.

But Mark, do you remember the days, like, of early bare metal orchestration? I do. You know, there was a lot of hubbub around bare metal orchestration for a while. But it seemed to kind of go away in favor of more, you know, virtual infrastructure automation. Because I think that's the way the typical organization went by. But I'm curious, you know, what the market for bare metal orchestration looks like. Who uses it and who does it? Phew, boy. Yeah, the 2008 was an interesting turnaround because they figured out the secret recipe for VMware.

And BMC figured out that the secret recipe included a lot of storage sales. And at that point, bare metal automation became a wasteland because people just sold a lot of VMware infrastructure and vMotion and some really, really expensive data center infrastructure got built. And the vendors really squashed the bare metal automation side, to tell you the truth. Boy, and so what does it look like today? Because, you know, before you go into what it looks like from your perspective, like, Mark, from our perspective, like, you know, we're traditional infrastructure folks.

So if we have a requirement beyond kind of ESX and vSphere and such, you know, we'll look towards AWS and there's, AWS has kind of come full circle with their bare metal service, but the market is way bigger than that, I would assume. Yeah, you know, so there was a time when I was all, oh, bare metal orchestration was a key element. And interestingly enough, I kind of feel like we're on a precipice of that becoming important again, as people start to move more into that hybrid cloud model and understand that there's a different way of doing things.

I'm not quite convinced of that, but I definitely could see the case that we might be on the doorsteps of that revolution. That's why I think it's cool to see where we're at today, kind of walk backwards a little bit to where we came from and then look forwards. That's how my brain kind of works. I am ready to take on that challenge and see if we can convince you in the next 15 minutes. Well, let's go, Rob. I want to hear it.

Let's rumble. So, I mean, bare metal is really hard because the landscape for bare metal is incredibly heterogeneous. So the idea that I'm just going to buy, like the hyperscalers typically have been able to do so far is they just buy the same server over and over again, and then they automate that and you never see it. The real marketplace for bare metal, especially in enterprise, is much more heterogeneous. It has different vendors, different specs, different server classes, right? All those things, and the protocols are ancient.

We're dealing with protocols that were written 20 years ago and are still the lifeblood of the systems, even down to, and I'm going to go minute on you, there's traditional BIOS, which has been around forever, and they're trying to replace it with UEFI, this new, this enhanced firmware. And that is all over the map, flaky, not reliable, and most of the people we know, even though it's been out for years, still turn it off and go back to the even older stuff for management.

And that's what makes this environment so hard, right? Everybody's network topology is different, their naming convention is different, they put different NICs in their systems, and it changes the enumeration, and so the problem with bare metal and automating bare metal is it's not a single, thing, it's every variation under the sun. That's why it's hard. And that's why solutions like VMware became super popular. When you abstract away environments, you can go into, in the early days, it would be uncommon to have a set of DALE servers and a set of HP servers in the same physical cluster.

You could separate that out. Nowadays, you know what, it may not be best practice, but it's not unusual to say, you know, I have servers from different generations, different manufacturers in the same cluster, right? The VMware has abstracted that layer. But I can definitely see the use case. You know, Rob, I had you speak at Interop last year about composable infrastructure, this concept that HPE has, that you take a blade chassis and you're able to reprovision the blade chassis from being a ESXi host to a bare metal container host to a bare metal exchange server, whatever the magical use case is.

But I get the sense that that's kind of not the market you're going after when you say, you know, the heterogeneous use cases we're looking for in the enterprise. The dirty secret for a lot of this stuff is even within the same vendor line. Servers are very different week to week, month to month, they change components, they change all sorts of things. You change the firmware and you can change the behavior of a system and you have to keep up with the firmware or you might be vulnerable to security issues.

So, I mean, it's just, it's the reality of the environment. And I've just said enough that everybody's like, I'm never gonna own a data center again, never, get off my lawn. You know, that's, that's, you know, we see that a lot. You know, and Amazon's taken a ton of that business. Our premise at Rackant is that that is a symptom of not doing it right. And that if we can fix those problems, the cost of running infrastructure is gonna go down and it has to go down dramatically because we're in this emerging market with Edge where people are gonna be spinning up small IT infrastructures all over the planet, right?

In little locations, because they have to run Edge IoT workloads. And those are gonna be lights out, thousands of data centers, you can't touch them. And they're gonna be even more heterogeneous because the environments are gonna be even more, you're gonna have Nooks and Raspberry Pis and actual servers and desktops. And, you know, who knows what, you know, people have cobbled together for quantum, an arm. And it's just, it's just, that's the world. So you have to build something that automates against this moving target.

And believe it or not, we figured out how to do that. That's sort of the rack and magic. So that's interesting, because I think my first, when I first thought of this, the first thing that I thought of was how big of a pain in the butt it was to upgrade firmware on bare metal servers. And I remember like Xcat and Razer and things like that. And they didn't work that well. And then I remember an EMC world, maybe I'm gonna say five years ago, they had a rack HD, I think is what it was called, or something like that, that was supposed to, you know, make that easier.

At the time, it was all about bare metal provisioning. And the use that I saw then with where I was, it was a long time ago, was being able to take a physical thing, say, you know, a storage array node, and turn it into something different and move it around, you know, in a better way. I hadn't thought about the use cases of the edge. I think that's very different place than I was thinking. I like that. Well, it's coming.

We've been building technology that's really well suited for that, because it's small and lightweight and designed to be autonomous. But that's not why we built it initially, right? What we built was something that could handle all the automation tasks, and we just had to make things small and simple. And so from that perspective, it's a pretty broad range of components. The thing that you just named is something we thought was really important, which was RAID and BIOS configuration. Turns out it's really not the first thing on people's list.

And if you'll bear with me for a second, I'll try and explain this, because we ran around, you know, saying, oh yeah, we're gonna fix RAID and BIOS problems for people, and you need to. But it's not what really makes people have trouble with bare metal. What people have trouble with in bare metal is that there's actually a lot of workflow that you have to get done. There's a lot of integration that you have to do, so when I'm doing a bare metal boot sequence, I actually, the machines boot off of a network through three different, at least three different, sometimes more operating environments, a pre-execution or pixie environment, then they move to another shell, then, you know, an iPixie shell, then they move into a grub shell potentially, then they move into the actual operating system pre-execution environment, and then install an operating system.

It's literally moving through a sequence of boot environments, plus you have to interact with DHCP to get an IP address and get on the network. You have to actually get the right template files. The challenge with provisioning is that it isn't just one thing. It's actually this blend of synchronized actions that you have to do in exactly the right order with exactly the right protocols based on the architecture of the system. Literally, ARM servers and Macs and Intel servers all boot differently with different sequences.

They need different seed files. Once again, I'm making it sound incredibly scary. You're giving me flashbacks and nightmares. But here's the thing, and this is what, you know, if you can put all those things together and then start doing, you know, you build something that assumes you're gonna have flexible workflows, then you can actually take a machine through this process in a very repeatable way. And it is possible to do it. We just have neglected that space because we've been papering over it with virtualization for a long time.

It wasn't until, you know, I got deep into OpenStack and was watching everybody, you know, pull their hair out on operating it, that I realized that, you know, wait, the problem isn't OpenStack. The problem is everybody's data center is different and they don't automate that. And so if we fix that fundamental level, then we actually can solve this problem in a really repeatable way, which is what we do. So building on top of, once you get the system baselined, correctly set up and discovered, once you have that system set up, running a scripts to install stuff is much, much more reliable.

It's much easier. It's much more cloud-like if you want to think about it that way. And so, yeah. I'm sorry. So let's take a practical look at this. Like where we would, as we talk about the edge, the who edge. I'm like, I'm hearing background music. Yes. I was hearing DuckTales for some reason when you said that. Who we? Edge computing. That's a winner. In reality, let's take a look at like someone who's actually published something that says they're going to the edge.

And the traditional data center model of managing the edge either runs as a problems or not necessarily the path that they take. Chick-fil-A, 2,000 endpoints over with managing, you know, like 6,000 containers at the store level. Like what's the practical example of automating, when you take an example like that. So basically retail store, when you take an example where you want to manage containers all the way down to the edge or this bare metal app all the way to the edge because it could be some OT app that requires bare metal.

Right. And a lot of these incoming AI infrastructures where we want to do, you know, GPU or TPU processing in situ and field are going to require that type of controls. But what you're talking about for something like Chick-fil-A, so Chick-fil-A famously is building Kubernetes clusters on Nooks, which are next unit of compute. These little desktop video stream type servers. And so they're deploying Kubernetes in all their stores. And the reason they're doing that is because they can use the Kubernetes APIs to redeploy store automation really easily.

And so the Nooks are sort of pre-configured to build a cluster. They didn't spend a lot of, I haven't seen a lot of discussion about what it took to bring up the Nooks. They really talk about the benefit of Kubernetes as a way for them to distribute applications to the stores. We would take that a level further and actually deploy, be able to deploy the Nooks in a standard consistent way because the operating environment in that system needs to be standard too.

And one of the things that we would be able to do there is actually deploy the operating systems as a unit. So immutable based deployment, right against those systems. There's a ton of interesting caveats about network connectivity, standalones, stores need, you know, what, you know, can somebody plug in a new Nook and have it automatically install and join a cluster? Those are very real concerns for the IT operations of even a distributed environment like what you're talking about for Chick-fil-A because they're not gonna have an IT pro on site.

You know, the fry cook might need to plug in that new Nook or replace it. Wow, fry cooks are taking our job, Mark. This is Edge. So for Edge, you, the scenarios that we paint for Edge are, I'm at, I've got a data center sitting on a cell tower and the, you know, a delivery driver shows up, gets, you know, remote, you remote unlock the cabinet, they pull out a server, they put in a server and they close the cabinet. And that's the type of IT support that we're looking at for the Edge.

And if we can do that at the Edge, think of how powerful that becomes for enterprises with hundreds or thousands of servers, right? If you can be at a place, and we are, I mean, it's not an if, we're actually demonstrating these capabilities where you can literally just drop servers in, power them on and walk away and they will do everything they need to do completely autonomously. BIOS raid, firmware, OS discovery, burn in, join clusters, right? Tell your backend systems that it's done.

That's cloud behavior. We're just doing it for physical gear. Hmm. Mark, any last closing questions? Man, I'm trying to think of one that won't give us another hour of conversation. Yeah, this is the tip of the iceberg of the conversation. Yeah, I mean, there's a lot of conversations. I wanna know a lot. I wanna know how, how long it took. I wanna know, you know, what implementations, like a lot of those things that I wanna, you know, drill in on.

Unfortunately, I went on a 45 minute run with Rob. I think it was Interop. That's right, it's coming up. We had this detailed conversation. You know, we talked through, we also talked about state management, configuration management above the actual, you know, as you start to get into the application, and the application requires a solid and consistent underlay. Yeah. How do you do configuration and CICD and integration of CICD with the underlying infrastructure? So if you need to roll out a brand new image for the base host, how do you do that in an AD test type of manner?

There's like, you know, reading my mind. And you're actually naming the journey we take people through when we talk to them, right? They need to start with control, which is just, can I boot and install an operating system on demand? And then we bring them into a continuous provisioning type of flow where it's integrated into other parts of their system. And then from there, they can get into, is my BIOS updated? Do I have compliance and conformance reports, right? Is my system actually what I think it is?

And then from there, they can start to scale. But it's just like any IT process, right? You're walking somebody through, you know, this incrementally more robust lifecycle. Yeah. Metal's no different. To get to the point where the fry cook can just plug in and walk away is a lot of work. You know, so I think that, you know, I think I've kind of narrowed down one question is how can somebody start with this? Can they play? Is there a demo?

Can I just do this? I want to see it. You can. And actually, that's the only way we deliver. So everything we build, and we work really, really, really hard at this to have this five-minute experience. You know, you can down, the stuff we do is all Golang-based. It's a single executable. You can download it, run it on a laptop, and boot VMs into the system and prove that the whole thing works on your laptop, even before you touch a physical server.

And that's actually a lot of times how people get started, because they're playing with this and going into sort of headbanding to think, okay, my first experience with real bare metal provisioning might be in VMs on my laptop. But there's no reason it shouldn't work like that. We're big believers that you don't have to have, you know, a rack of servers sitting next to you to prove that the system works or to build good automation. We certainly don't want to do that.

That's loud and hot. So it doesn't take much to get started. What we've also done is we've made the system composable. So the same thing that makes us resilient as far as having different pieces. You know, if you have Dell servers, you don't want the HP server stuff laying around. Or if you don't care about BIOS, you don't want servers to go through BIOS phases if you don't need it. So the system, you get to pick and choose what parts of the system you want.

It's pluggable and composable. And that also means that when you start, Mark, when you take that first hour of doing this, you're not going to be distracted by all sorts of advanced functions. You're not going to try and do image-based deployments or, you know, pooling and cloud APIs. You know, all that stuff you can sort of ignore until you've got the basics. And then, you know, crawl, walk, run in infrastructure. And that's important. That's how, you know, all IT really needs to be.

So where can somebody find that? If they want more information, if they want to dig in on this, they want to learn more, where do they go? io actually is our portal. And it's, and that has the quick start guide right there on it. That's the easiest place to get started. And it's important to note, we, you know, I said this at the beginning and I'll reiterate it. Hardware is an on-premises thing. It's, it runs behind your firewall.

We are not a managed service. We are, you know, we have a portal that we host to make it easy for people to get started. But what we do and what people need in this stuff is self-managed infrastructure. It's not our hands or our services. You run the software, it runs your hardware. We do not connect to it. We don't mess with it. We don't, we don't want to. This is right, we're, we're, our customers really have a lot of expertise and we just make them faster.

We don't try and take their jobs. All right. That's, that's some good stuff. Mark, where can people find you and stalk you online? You can come to my house. I live in Northern Kentucky now. Or if you want to just find me online, Twitter's at SensiStorage or pretty much anywhere you want to be on SensiStorage. com. And I doubt if anybody wants to come to the South suburbs of Chicago to stalk me. Even as we're recording this on March 1st, because I think it's a nice 14 degrees out.

You can find me on the webs at ctoadvisor on Twitter. com is where you can find the podcast in the show notes and all of the blogging stuff. Rob, what about you? Where can folks stalk you on the Twitters? On Twitter, I am zehicle, Z-E-H-I-C-L-E. And I am very active and outspoken there. So if you want to poke the bear, you are welcome to tweet at me and make it go. And then we do host the latest shiny.

So you're welcome to follow the latest shiny or check us out too. If the edge topics are interesting, we've been diving deep, deep, deep into edge topics. And so there's a lot of content there. And I'll- L8ISTSH9Y, I'll speak on people. And I'll make sure to post a link to the show on which you got me in trouble with Dave McCoy. That's okay. Okay, I'll watch. It was a great conversation. It actually was really, that and the follow-on conversation that you guys had as you saw was some good stuff.

Until then, you can talk to us on the Twitters. Talk to your next CTO advisor.