You Can’t Ride the Ticket Train to Your Hybrid Cloud Destination

30:41 · Watch on YouTube ↗

Transcript 4,985 words · about 33 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

hello thank you for joining this session I'm going to take a quick survey of the virtual room here how many of you like putting in tickets for things don't all raise your hands at once I can't see you anyways my guess is that you don't really enjoy having to put in tickets for things you wish you had them now I can pretty much guarantee that your developers don't enjoy putting in tickets for things but that's not the only reason why you can't ride the ticket train

to your Cloud destination my name is Dorman drewitz I'm with pagerduty and we're going to be talking today about some ways to change some of the practices that you have in order to achieve some of the cloud outcomes that you're hoping for so to start I'll go out there with a bit of perhaps an unpopular opinion which is that tickets shouldn't be the atomic unit of work uh in it and so what do I mean by an atomic unit of work which is there's always a

lot of jobs to be done and how do we manage and track all the jobs that need to be done uh and they have been often reduced down to units that fit into a ticket pre-defined units of work however what we'll talk about is how when the nature of work has to change or when it's unpredictable or uh when you have unstructured work that doesn't fit nicely into a ticket then this really breaks that Paradigm so it can't be the atomic base unit because not everything

fits into it now let's talk a little bit about um and give some room for the fact that ticketing systems this is not really about those systems this is really about the practice of using tickets and cues in order to organize work um lots of good intentions trying to rein in chaos instead of requests and ideas and projects flying around how do we translate and track all these however this has led to a few unintended consequences and this is the first of a couple of 90s

hip-hop uh Jam references I'll be making throughout the talk today um and the late Notorious B.I.G if you were talking about ticketing systems and practices probably would have said Mo tickets Mo Problems let's talk about why so first there's the kind of basic notion of all right let's use tickets to manage and tracks track the jobs to be done makes total sense seems reasonable however when you step back and look at what does that accomplish what does that unlock you haven't really changed the capacity of

your team so you end up with sort of same capacity different day day in day out you haven't really unlocked any sort of 10x factor that could help your team do more you've simply just reigned in any chaos and tracked what the jobs need to be for today that's not necessarily going to help you for tomorrow the second is this is a way for us to track our team productivity um and it's true you can use tickets as a way to understand how different teams or

different individuals and teams um how much work is getting done but you end up with tickets for ticket's sake it's the same kind of problem you run into in the development world with things like velocity points where once people kind of know that that's what they're measured on then sometimes you have unnatural behaviors in order to inflate the thing that is being measured so you get a sort of distorted view of productivity someone could be cranking through tickets and yet the amount of business value that

they've delivered and how much Innovation or progress on business projects may not be reflected in that they've simply worked through a lot of tickets um another sort of you know good intention that folks often have with ticketing practices is well we could use these tickets to capture a lot of information about our environment and sort of gain some situational awareness with what's happening on where are the issues coming from um what what are the different needs of the organization but what you can also end up

with is a lot of garbage tickets where the more fields that get entered in as oh we could capture this we could capture that folks who are under a lot of pressure to just crank through those tickets don't really want to take the time to fill out all of those fields and they'll put in the minimum viable response in order to be able to close that ticket out as an example of how garbage tickets come about and you haven't really been able to extract and gain

anything useful from that and then finally this is I think a really interesting one is how to organize work across team boundaries ticketing sort of practices generally aren't within a team it's between teams and so if other teams are dependent on another team in order to get work done uh Team a b and c for example all have requests into Team D team D getting flooded with all these requests says everybody's got to go through our or put in a ticket and go through our ticket

queue um that kind of Reigns in the chaos but then there's not necessary necessarily an inherent sense of was the project from Team a b or c were any of those of business priority and urgency that needed to be handled differently and none of those projects really go faster than the slowest ticket and so while in theory this way of managing work between boundaries you haven't created a sort of shared understanding of what the business priorities are in order to define the work that's going to

get done with an agreement between those teams and those dependencies that exist so this isn't to say that there aren't times and places for using tickets but what has happened in many cases is that folks have built up so many practices around filing things into tickets funneling work into cues that and measuring everything based on that that your your teams and more importantly your leadership teams are focused on something that is not inherently delivering business value and uh and so there's there's training to to fit

everything into um something that they're being measured on and they're missing the forest from the trees and this is all kind of challenges that can emerge even when things are sort of standing still but the reality is things aren't standing still and we see increased complexity um and forces of complexity sort of entering into the work that we need to do cloud is certainly one of them and we'll get to that in a minute but on the one hand we have expectations for responsiveness and reliability

as high as they've ever been teams and consumers want things and they want them now they expect things to be available at all times and highly performant and yet the operational complexity is also increasing and so we have everything from more processes having been digitized systems distributed Cloud being one of those distributed system forces but other things like microservices patterns certainly fit into that and then of course we have now distributed workers in a post-pandemic world where folks are working from anywhere adding new layers of

complexity to how teams coordinate collaborate and prioritize work that needs to be done so now let's sort of spend a little bit of a moment on the cloud and how this kind of fits in as I mentioned you know this is are we riding the ticket train to the cloud or not and hybrid is kind of silent because once you've introduced uh some element of cloud even if you're still managing on-premises environments you really now have your uh least common denominator includes Cloud environments so that

is what you're dealing with so our second sort of hip-hop reference here uh to quote the the great Will Smith we're going to get cloudy with it um and so the first thing is uh The increased change rates in the cloud some of this is the same kind of say cloud native patterns that you could be applying on premises and practices like continuous delivery but something to keep in mind is that even if you're early in your adoption of those practices the cloud providers aren't and

so you're inheriting an environment that is continuously changing and so that's something that you need to be prepared to handle is it continuously changing environment another important thing which is actually a really good reason to kind of get on board with some of those practices is that security best practices increasingly require the ability to change frequently whether it's rotating credentials or repaving environments and rolling out updated patching systems and libraries and and what have you in order to deal with the increased volume of vulnerability abilities

that are being uncovered so change is a constant Factor even if you're just looking at keeping up with basic security practices let alone all the new code and Innovation you want to be shipping on top of it now another sort of compounding factor in the cloud is the fact that you're dealing with a lot of ephemeral resources and and this again this can actually come up in your on-premises environments if you're managing kubernetes for example um as that tends to be creating and destroying uh systems

um within a kubernetes environment continuously you look at things like lambdas and now when something goes wrong you have a different sort of set of circumstances under which you are trying to go back and diagnose and troubleshoot and understand what is at the root of this failure in order for us to both fix it and perhaps uh build in our our changes in order to make sure that it doesn't happen again so recreating the problem um when you're trying to deal with a an outage or

a failure of some kind becomes different when you're dealing with ephemeral resources and then the last just point maybe the obvious one is that you're introducing more points of failure many of those are out of your control so when you think about reliability practices so site reliability practices and you're managing your error budgets you have a certain amount of that error budget that is going to be consumed by any of the other services that you're dependent on and so um this notion of resiliency engineering becomes

really important as you think about how do we maintain that performance and reliability experience that our customers and our our users demand so Cloud adds some some compounding issues to that complexity so if we come back to operating kind of from a ticket-based world when you've got increasing complexity and I'm going to come back to the notion of speed and change um all of these different points introducing potential uh failures into the system um as well as just the the improvements that you're trying to make

the projects that your teams are trying to support in order to deliver new business capabilities maybe for a competitive Advantage um maybe to improve a process all of those changes that you're trying to introduce how quickly can you move when something goes wrong how quickly can you move and how much of that work do you want sitting in a ticket queue which is sort of the slowest end of the spectrum when you can get it to the humans you can at least get it to human

speed but really thinking about how much work can you push along the Spectrum to machine speed if it's wrote if you've seen it before how much of that needs to be sitting in a ticket queue let alone even processed by a human and that's a sort of Paradigm I want you to think about in terms of instead of same amount of work different day that's being processed through a ticket queue how are you step function changing where your teams are spending their time and being able

to do more with the same so this is where we start to get into thinking about okay how do we think differently about work so now our third and final 90s hip-hop reference um coming from the great Missy Elliott is it worth it let me work it how do you determine what work is worth it and so it starts with being able to detect and diagnose and how can you automate that so that that doesn't become manual work sifting through tickets in a queue in order

to identify okay these are really the appropriately classified as P1 or sav1 severity issues um let alone the priority Project work that needs to get done but how do you catch those things that should not be sitting in a ticket Queue at the slowest end of the speed spectrum and how do you take steps to do that that next level of work for example with an issue how do you capture the diagnostic data right there at the point of the crime um so that you have

all the the forensic data in order to be able to diagnose and understand um instead of interrupting six different teams and ask them to each go capture and report back with their diagnostic information the same tests that they're always going to be running the same logs that they're going to be going through how can you relieve that additional burden off of those other teams so that they're not interrupted and then thinking about okay once you've detected and diagnosed an issue how do you coordinate and mobilize

in a more effective way um not everything needs to be in all hands on deck interrupt everybody until we can figure out uh who doesn't need to be involved situation uh you want to be able to do a pinpoint escalation to find exactly the right person on the right team to handle something that is urgent um and and high impact to your business so you can't lose that somewhere in a ticket queue that's gotta happen before something finds itself sitting in a ticket queue and beyond

that you have other people who want to be informed of that situation so that they can do their jobs well and I'll pick on customer service here where if there's an issue that's urgent and again as you depend more on your ephemeral distributed Cloud resources you're introducing new points of failure things are bound to go wrong but as that's impacting customers and they're calling in to customer service how are those customer service agents given visibility that yes this is an issue and it's being worked on

um do they have that information are they going to be going through ticket cues to try to figure that out how do you surface that to them in an effective way business stakeholders sales people who might be trying to close deals and business how do they know if there's something A system that they depend on is down and that's going to affect their ability to try to close business at the end of the quarter um and then another thing is thinking about okay another way that

we can start to deflect tickets so that instead of having this sitting in a queue what if we could just automate the resolution what if we could make that self-service to teams instead of team A B and C all waiting on a huge pile of work in a ticket queue to Team D what if 80 of that was self-service and now we're just left with the sort of tricky Oddball things that do require the teams to get involved and figure out how are we going to

deliver that it's a lot more interesting work for those teams and it means that the other teams are getting what they need a lot faster so those are some things to think about um how you can take things that are repetitive we always want to follow these practices in these scenarios that's a workflow that we want to be able to just run automatically and work yourself down the number of tickets that you your teams have to manage and figure out is this one worth it that's

what I want to work on and everything else has been automated it's been sent to the right team and uh it's been prioritized appropriately so a little bit of story time I want to give you a couple of examples of folks that have been on this journey um and some of the technologies that they've brought to bear and so aiops is a really hot topic and so that's one in particular I I want to call out what does that mean because I think a lot of

folks get distracted with the AI in aiops and they kind of they're thinking like robots and lasers and this is going to take me 12 to 18 months and I've got to have to hire some data scientists and data engineers and start to manage machine learning models and train them on on our operational data um actually it doesn't have to take that long you can point um a lot of that the sort of event data directly to a service via an API and have them they've

already trained models process and understand what's happening with all of that event data and start to do that correlation and compression that's the first part of the the sort of AI and AI Ops and is really just sort of dealing with the event data after that you want to be able to take action more intelligently and here one of the Frameworks I've heard is thinking about how you automate accelerate or augment what the humans are doing um and so Schneider National um is an organization I've

got a URL to the full case study if you want to watch and there's some great things to see in their presentation because uh some some practices around how you actually get your teams on board with this which is a really interesting question but they were sending everything into their ticketing system everything from requests including event data in order for that to get processed in order to determine where are the issues what needs to be dealt with and then taking steps from there getting a hold

of folks who can work on those issues but it was a lot of event data to deal with and was all sort of Landing in their ticketing systems and so um they adopted an aop solution and used that to actually eliminate about 5 000 plus uh manual notifications per month and um they also kind of wired this in so that even as uh uh event data would get sort of written into their itsm platform they could also Auto resolve those as the AI op systems recognized

no no that's actually something we've already seen or that's not really an issue so the notion of things like flapping alerts um uh that you know set off an alert but it resolves itself in a very short amount of time and isn't really a real issue but you have a ton of duplication where you know a critical database failure can can send alarm Bells off from 60 different systems because they're all dependent on that database and so how do you reduce all that noise so you

don't have a 60 UH 60 different alerts that need to be sorted through and understood and correlated you just have one that you say okay this is the critical issue and now we know exactly which team we need to bring in to work on that and as a result this is part of how they they realized a mean time to reduction mean time to resolution reduction of 24 and by doing this they were able to find after or some number of months that they were able

to actually shift some of their operational expense back into the development teams um which is really great news for the business uh and so reducing a lot of the noise out of the system but also reducing a lot of that manual effort so this is where kind of ai ai Ops fit in and I think something really important to understand is is kind of where this sits they didn't eliminate their ticketing system by any means um but this was sitting between their their monitoring all the

sort of uh systems that were generating that event data and that itsm sort of intercepting if you will and processing that data a lot more efficiently in an automated way and that just reduced a lot of that noise in the system um and really helped them become a lot more efficient and effective so that's one kind of solution uh an approach to think about the other one that I want to put out there is sort of less related to sort of machine driven events and more

related to when humans are requesting things so this is a self-service automation so how do you look at more of the requests that come in across an organization and turn that into a a self-service function for teams and so the example here um ample Organics again link a URL here for the the case study if you want to read it at rundeck.com and there um in their their it operators uh were supporting contractors and so again this is where you know the the dollar signs start

to add up when your contractors having to wait for environments they weren't allowed to provision those environments themselves for security and compliance reasons but when they had to put in tickets and then wait they're on the clock um but they're they're not really able to be very productive um meanwhile the the operations team was getting interrupted uh several times a day in order to handle various types of requests um and so by creating a self-service uh environment so that for example these contractors could automatically provision

on their own uh an environment they didn't have to go in they didn't have to access anything directly all of that was kind of secured over that that team boundary but they were able to get that the capability to provision on their own that reduced the the interruptions into the Ops Team at an increased the speed at which those contractors were able to do their QA work which meant that they were able to ship capabilities to customers faster and overall just increase the capacity for for

work on the team and so think about you know what if you could give your operations team 20 more time each week what would your teams do with that um so hopefully those are some uh you know useful and inspiring stories I mentioned in the uh the Schneider National um uh case study that uh that I put up there uh where they talk about their their AI Ops work they actually won an internal prize what I really love is you they actually show a video like

hype video that they shared out in their company um as a way to kind of get people excited and on board with a different way of working so file that away because that's kind of a very useful change management technique um but just to close out here let's think about sort of what are some steps to follow to kind of work your way out of the ticketing business um so we're not writing the ticket train to the cloud you shouldn't be in the ticketing business you

should be um or the ticket processing business I guess you should be in whatever business it is financial services retail what have you and so how do you how do you sort of shift the focus there um first is sort of there needs to be a new goal around fewer tickets right so you Whittle your way down to what really needs to be sitting in a ticket queue and everything else should should not be going there um and one of my favorite quotes from Adrian cocroft

uh former CTO at Netflix um was if you want something to get done figure out how many meetings and tickets it takes to get it done publish the number and make it go down so start to do a little bit of Investigation into where are these tickets coming from uh what what is all of this work uh how much of this is repeat and so um and if you want everybody to be on board with the fact that like hey our goal isn't how many tickets

are we processing which we talked about how that kind of promotes the you know creation of unnecessary tickets and we'll put in tickets for all the things and other unnatural behaviors and garbage tickets instead the goal is you know more like the golf score how can we kind of manage things with fewer tickets um now I'll say that this is a temporary goal because at the end of the day I think we don't really want ticket numbers to be the primary goal um and it really

should be about business value but to sort of trigger a little little bit of a change how can you make that number go down and look for ways to to get things out of ticket cues and be managed automatically outside of that so make that goal visible so folks are on board um then you're thinking about okay what are the steps before things end up as a ticket if we're trying to get that number to go down we're trying to get work dealt with automatically at

machine speed where possible before it lands in a queue where is it coming from follow that that flow of disruption data and request data to know where does this work originate from and how can I intercept that and automate what's going to happen next and so maybe you still write something into a ticket management system for record keeping but you really want to handle the work first so how much of that follow kind of follow the the flow and and intercept and handle from there and

then third I mentioned a little bit of kind of you know the change management and hyping um you know producing the hype video If you will um which is you have to sort of sell this Vision internally and really focusing on the outcomes that this is going to achieve um both the business outcomes in terms of increased capacity and output the ability to deliver on those transformative business projects faster the Phoenix projects of the world if you will but also like what the day in the

life of people in your teams it's going to be like um I asked at the beginning how many of you like putting in tickets well of your teams how many of them like processing tickets and so that's something that um you know you want to get people excited about the change that's going to be and introducing automation which sometimes people can feel threatened by but at the end of the day the work that you're trying to automate is the least interesting human work how do you

onboard automation as a team member great story from results CX on how they actually gave their automation uh sort of Bot collection a name um and and really you know a profile picture on the team and they really truly onboarded it and kept teaching the bot new things to do just as you would onboard a new Junior employees like this week you're going to learn how to do this and after you've finished learning how to process that type of work I'll teach you something else so

hopefully these have been some useful ideas for you to think about okay as we're on this journey to cloud and our goal is to deliver more business value we gotta get out from under this mountain of tickets which aren't really leading us anywhere and and so hopefully this has given you some ideas of how to do just that thank you again and enjoy the rest of the conference as promised wasn't that a great session if this is your first session of the day and you're wondering

what's next well you have options you can hang out in the chat room now and you can see there's a conversation going on in the chat room and the session that you're in this will be open until the end of this hour and the top of the next hour the next session will start if you want to go into the breakout room there's a general chat session you leave here go into the lobby and then enter General's chat you can chat with your regular attendees I'll

be floating between the two chat environments love your feedback if you have any questions please you can DM me within the platform or on Twitter at CTO advisor enjoy the rest of the conference