Gene Kim: DevOps Evolution, AI Leadership, and Enterprise Transformations

In this episode of the CTO Advisor Podcast, Keith Townsend interviews Gene Kim, the renowned author of "The Phoenix Project" and "The Unicorn Project." Gene shares his experiences and insights from his extensive DevOps and IT leadership career, exploring how these fictional narratives reflect real-world challenges and triumphs in technology transformations. Gene discusses the role [...]

Transcript 3,588 words · about 24 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

All right, you're listening to and watching for some of you another episode of the CTO advisor podcast. I have the distinct pleasure of interviewing Gene Kim, author of famously the Phoenix Project, the Unicorn Project, these seminal pieces of, in some cases, fictional work. Gene, I got to admit, some of these stories, they didn't sound fictional to me, but we'll get into that. He has been a huge proponent of DevOps for years. He's been the CTO of TripWire for 13 years. His books have sold well over a million copies.

He's a Wall Street Journal bestselling author. Gene, welcome to the podcast. Ah, Keith, I'm delighted to be here. Thanks for having me on. So many of your guests are friends of mine, and I've just recently heard the interview you did on the Deloitte podcast, hearing about how easy it is to use GPUs some days. Great to be on. Thank you. Yeah, I've had the pleasure of talking to some cinema folks. The Deloitte Project was one of my favorite, because I'm from historically PWC, so I never thought that I'd be on a podcast with Deloitte, my number one competitor and enemy when I was at PWC, but it was a really, really great time.

I've been looking forward to this conversation. We've been working for a couple of weeks to get this scheduled. We got it scheduled. No specific topic, but I'm going to start the conversation now with some of the pre-chatter that we had before we hit record, and that is, you have two options. You can put some of your smartest people on the challenge of moving legacy applications out of VMs and putting them into the public cloud, or you can get them started on Gen AI projects and helping the organization get started.

I think DevOps, from the high-level umbrella, these folks that we give the challenge of implementing platform groups, DevOps programs, they're the same folks. What's your experience in having the conversation around this precious resource of my smartest people? Oh, my goodness. Yeah. What a great question. Boy, can I just sort of mentally rewind my head and just piece together an answer, because I suspect it resonates with your own journey. Yeah. I remember going to my first DevOps event in 2010, I got an email from John Willis and Damon Edwards.

They were running the first DevOps Days in Mountain View, California, and that was right a couple of months after the first DevOps Days that was in Ghent that Patrick Duvall ran. It was incredible. I mean, it was just, I immediately knew I had found my tribe. I'd always been, I've been studying high-performing technology organizations for 25 years, and it was typically focused on the operations and infrastructure folks and information security. And I sort of knew that there was this kind of a dev component out there that was a big piece of the puzzle, but I could never find the people who were sort of like-minded, kindred spirits, and then show up there.

I would say probably 50-50, developers, operations, and infrastructure. And it was just, I knew from like the first five minutes, like, oh my gosh, I've found my tribe. It was Patrick Duvall with a kickoff address, and they were showing the I Love Lucy episode, you know, eating the chocolates off the assembly line, you know, just when the assembly line was faster than what Lucy could handle. Anyways, it was just so fun chronicling and capturing, you know, just being able to hear all these talks about like these people doing 10 deploys a day, right, doing crazy things that were unthinkable, right.

And so it was my observation that it was really the kind of the best technical people, the best leaders who were kind of driving the charge. And that was in the tech giants, you know, Facebook, Amazon, Netflix, Google, later Microsoft, you know, and there was all the startups, Yelp, and, you know, all these kind of exciting startups, I mean, Lyft and Uber at the time. And driving these, you know, they were sort of pioneering these patterns. And you're right, it's also been my observation that, you know, of these people I've studied now for, you know, 10-15 years, it was really surprising to me that those people are still driving large, you know, these transformations in large complex organizations, but they're also, a quarter of them are leading, you know, these Gen AI pilots.

A friend of mine, Brian Scott, he's a principal engineer, a principal architect at Adobe. He co-leads the Gen AI rollout and governance programs. And so, you know, imagine there's like 5,000, 10,000 creators at Adobe, right, and their job is to sort of maximize liberty, but also maximize responsibility, you know, go as fast as the business needs, but make sure that, you know, the risks are managed. And so, they're working across this vast cross-functional team, you know, legal, compliance, information security, dev, ops, infrastructure, platform teams, trying to figure out like, all right, how do we, you know, take, you know, appropriate risks, you know, and let developers and creators do what they want, and yet make sure that they're not, you know, injecting existential risk into the organization.

You know, at Vanguard, right, there's a, you know, they have this enterprise-wide now called distributed experimentation, you know, at scale, right, it's not one person, you know, trying to decide what to do. It's like, you know, hundreds, if not thousands of people sort of like trying to figure out like what, you know, given these kind of exotic new technologies, how do we actually make these work for us? I did share with you one other, a friend of mine, Dr. Topal Pal, he's at now Fidelity.

And he was, you know, like many organizations, everyone's driving these pilots, trying to figure out like, what are these good for? And what aren't these good for? A friend of mine, he said, their first pilot with these like coding co-pilots, like did not achieve its goals, and they had to shut it down. Because on the one hand, you know, junior developers are able to be more productive and check in more code. But it put this huge burden on the senior developers who had to review it and say, oh, my gosh, this is not fit to deploy and run in production.

So that's not sustainable, that's not tenable. And so I think we're just still at the very early stages of learning kind of like, okay, how do we make this work, you know, not just for side projects, but you know, for, you know, teams, you know, 10s, 100s, maybe even 1000 people, it's just a fun time to be in the game. Keith, does this resonate with you? Yeah, it resonates with me. And, you know, you, you made the point before we hit the recording, you know, I wish we would have hit record, there was insights in that, you know, five to seven minute conversation.

But this, this, you know, tangent to the topic of our smartest people being put on some of the hardest problems in the enterprise. What we're noticing, I think, collectively, is that this isn't, this isn't a level one, level two problem. You know, whether we're talking about DevOps, gen AI, anything that touches or transforms an organization, these folks need to be skilled in, in negotiating, compromising, coercing groups that are outside of their direct sphere of influence. So you know, we're talking about security, compliance, product teams, etc.

And one of the things that interests me about some of your speakers for your upcoming show is that these speakers are not concentrated in the metas, the Googles, the hyperscalers. Talk to me about like, the DevOps journey outside of the large hyperscalers. Yeah, I mean, like, like you, I have nothing but admiration and respect and appreciation for how the hyperscalers, you know, really pioneered these DevOps principles and patterns. I learned so much from them. But for me, my area of passion for the last 10 years is not studying them.

It's really how are these principles and patterns being adopted in large complex organizations that have been around for decades or even centuries. And so I'm going into year 10 of running this conference called the DevOps Enterprise Summit. And we renamed it the Enterprise Technology Leadership Summit. And we originally called it a conference for horses, by horses, no unicorns allowed. And so it's been exciting. It's been over 1500 leaders, over 700 enterprises across almost every industry vertical. And I was telling you beforehand, right, these are the most amazing heroic technology journeys and transformations I've ever seen.

And so, like the oldest organization I presented was Barclays, a bank founded in the year 1695, which actually predates the invention of paper cash in the West. So but the absolute oldest organization was UK HMRC, His Majesty's Revenue and Custom Service founded in the year 1200. It's just, you know, there's no code that goes that far back, but there's certainly traditions and values and certainly processes that go back centuries. So it's just been so fun to help chronicle these journeys, technology leaders giving experience reports on in a very standard form, which I just love, right?

Here's our industry. Here's how we compete in it. Here's a business problem we set out to solve. Here's what we did. Here's what happened. Here's what we learned. Here's a problem that still remain. And I just love that because as adult leaders, as adult learners, we don't learn from people telling us what we think we should do, right? We learn from like how other people solve problems, right? And then, you know, we can, often that's the best way to, oh, and then we make a judgment.

Did it work for them? Do I like what happened to them? Do they get fired? Do they get promoted? Do they get more budget, right? And we will, you know, we'll make an informed judgment about how it applies to us. And so that's been super fun. So yeah, we have, you know, it's always been about experience reports and it's not just Dev and Ops. It's product and technology. It's about, you know, information security, compliance, it's about boundary spanning.

As you say, it's kind of that layer three organizational wiring. Yeah, I think that's what, you know, when people say DevOps is dead, it's like, what? No, DevOps is not dead. You know, I don't care about, you know, platform engineering, you know, SRE, those are all ways that we can help, you know, organizations wire themselves, you know, to better achieve goals as opposed to like silos that are at adversary, you know, have an adversarial relationship against each other. And, you know, I think the job of the technology leader is not getting any easier, right?

One of the things I learned working with Dr. Steven Speer, who is famous for his work studying Toyota, he wrote the most widely downloaded and read Harvard Business Review article all the time called Decoding the DNA of the Toyota Production System, and that came out in 1999. So we worked, we did a book that we worked on for four years, it came out last year called Wiring the Winning Organization. But one of the aha moments that he shared that just blew me away was, you know, the more functional specialties you have, you know, Dev, Ops, security, compliance, legal, like the more sophisticated your organizational wiring has to be.

And ours just got a little bit more complicated. We now have prompt engineers, MLOps, you know, we have, you know, we have now yet two, three more functional specialties that we got to figure out how to wire into organizations, right? We got to shift these data scientists way right, maybe even put them on pager duty, which probably may or may not, they may be a big fan of, but then we know what we learned in DevOps is that that's part of the pattern, right?

You build it, you run it. Anyway, so I'm just excited that in a couple weeks, we'll have these kind of great experience reports, but also these, you know, experience reports from the same technology leaders sharing like what works, what doesn't work in Gen AI. Just as you said, is it isn't it curious how some of the toughest challenges are always being given to the these kind of leaders who prove themselves, you know, solving the last problem that came up, which is like, how do we do this DevOps thing or platform engineering?

So I like to drill into that a little bit because I spent a lot of my time talking and advising folks on patterns that I see, and these patterns can be, you know, witnessed and leveraged, whether you're talking about rolling out DevOps, Gen AI, or any, again, technology that impacts the entire organization. I never thought of the idea that data scientists would have to wear pagers to be on pager duty, but it makes sense, right? When the LLM suggests to put glue on your pizza, someone is getting paged, and, you know, the SRE may not be able to necessarily solve that problem.

It may need to be a data, a deep functional area like a data scientist. So talk to me about some of these patterns that you witnessed time and time again, and these complex organizations that have to be addressed. Yeah, yeah. I think there's kind of three. Yeah. So working with Steve Spear on this, the Winning the Winning Organization book was such an eye opener because, you know, the question was, what's in common between DevOps and Lean and Agile and the Toyota production system and Lean?

And I think I could have given you a kind of a pretty good answer four years ago. But I think now, you know, working with Steve, like one of the most prominent Toyota researchers there's ever been, you know, we can say concretely, yeah, all those things are incomplete expressions of a far greater but simpler whole. And whenever you look at these transformations and frameworks, there's really three mechanisms of performance. You know, one is you have to slow down to speed up, we call that slowification.

The second is you have to sort of make the partition the problem so they're easier to solve so that you have, you know, the right amounts of coupling and coherence. And then the third is you have to sort of amplify even weak signals of failure so you can, you know, detect and correct faster, but ideally prevent, you know, and make sure that you have a culture that, you know, enables the accurate and quick transmission of important signals. And so I mentioned this just as a way to sort of ground my answer is that, yeah, whenever you have these silos, you know, sometimes silos are perfectly appropriate if they don't need to, if they don't have a lot to talk about.

But in the case of Dev and Ops, holy cow, like, you know, if you're doing big deployments that when things go wrong, cause chaos and disruption and catastrophe, well, then, you know, you do need a lot, you need a different orientation for that, you know, it can't be just JIRA tickets and handoffs, right? There has to be co-creation, joint problem solving. But there's other times where you don't want a lot of communication coordination. So you see that in the software architecture, you see that with platform teams, right?

It's like, I want to get an environment. I don't want to have to, you know, create a ticket. I don't have to get it approved. Like, I want to do it in one click or in the command line. And I think we're still kind of what we're learning is that, you know, we don't know where these new Gen AI pieces fit in, right? What parts can be put into a platform where, you know, I don't want to talk to data scientists, or which parts do you actually need them in the team?

So like when you said, like when someone starts recommending glue on your pizza, it's like, okay, is that a, like, yeah, what do we do? In fact, I saw a talk by the engineering lead for the person who owned the chat GPT team at OpenAI, and he just gave this amazing talk. And, you know, he was talking about, like, having to go deep into not just a Kubernetes cluster, but understand what was happening in GPUs, and how do you ramp capacity, you know, is there on this rocket ship to 100 million users?

I mean, you talk about, you know, the domain of full stack just getting even a bit bigger. Like, I don't know, like, who signed up for, you know, having to go down to the FTC, look at, you know, GPU memory utilization, right? But, you know, I don't think most people would say that's a reasonable place for most developers to spend time in. So how do we sort of configure our organization so that, you know, we're focusing on the business problem? But I was sharing with you the story that I tried to get a GPU instance in the cloud, and it took me two hours.

I couldn't even get Htop to run. I screwed up the driver installation. Like, that is not the way I want to spend my time. So it's an exciting, exciting place to be in the game. It is an exciting place to be in the game. So last question, let's, you know, kind of get a plug in for your show happening in actually just a couple of weeks. Who is this for? Like, we've talked a lot about, you know, we've talked all the way from GPU installation, driver installation, all the way up to big, moving organization leadership.

Where's the sweet spot? Where is these, who will be there and what conversations are going to be had? Yeah, that's a great question. Yeah, so I would say the sweet spot for the Enterprise Technology Leadership Summit are technology leaders. So, you know, they typically are, you know, third line leaders or more and stuff like the directors, you know, dev, infrastructure and ops, information security. And it's for anyone trying to figure out how to get better outcomes, right, that feel like DevOps-y like problems or platform engineering or SRE.

And one thing that I'm just super excited about is that date is a three-day conference, August 20th to August 22nd. But day two is almost like a single track, one day, ultimate learning day for Gen AI. And so we're going to have experience reports. Half the talks are from technology leaders from, you know, large, complex enterprises and half are from, so like, for example, you know, Vanguard and Adobe and Cisco and Adidas, they're all be sharing kind of like, what have we been doing and what have we learned?

What worked, what didn't work? And then the other half are, you know, subject matter experts from Gen AI companies. So we have Adam Seligman. He's VP of developer experience at AWS, who owns the Gen AI tooling. Paige Bailey, she was one of the senior product leads for Google Gemini and before that Palm 2. And so that's super cool. Joe Buechler is from OpenAI. And I'll be talking about like, what do technology leaders need to know, right? How do they make informed decisions on, you know, Gen AI is exciting, but what can they share from the field experiences about what's working, what's not working with their clients?

So I think it's just, I am super, super excited for this. And website, where can folks go register for the conference? com and just go to events. com slash events. And I'll give you a link of like the, what to expect. And it's a big blog post I wrote that shared what I'm most excited about for the conference. And I'll make sure that's in the show notes. com for the mothership, but you know, CTO Advisor is where you go get this deep technical content that I still love the platform.

And I love the audience. You can find me and engage with me on the Twitters. com these days at CTO Advisor. com? Yeah, absolutely. I'm on a, I still call it Twitter. I'm real Gene Kim. And yeah, I still, man, talking about Gen AI, it's probably one of the best places to learn about that. I mean, it's a ridiculous just how much you can learn just by, just through it. Absolutely. All right. Talk to you next CTO Advisor podcast.

Jane, thanks for joining us. Keith, thank you for having me on and keep up all the great work.