> "What happens when requirements change inside the launch window?" > "The alert fired but..."
In the running example I think it is clear that you could stop at any point and address the problem from that point rather than continuing. Questions such as "Why do we permit last minute changes at all?" are given no space.
In a complex system sometimes there is no root cause (cf. aviation crash investigations and the Swiss Cheese nature of risk). By allowing one person to shoulder the entire context of the incident and for that person be the one to suggest lasting corrective actions it leaves little need for the VP to even be involved. How would they know if the proposed corrections are worth the time and money to implement? I don't see any leadership in the example described in the article - I just see a knee-jerk reaction rather than any kind of proactive reflection.
> "I thought we fixed this last time!?"
> "This time was different in a way that our corrective actions didn't address."
I sympathise with this, and a lot of teams do operate that way, but I fundamentally disagree with it.
Part of the reason Amazon's CoE-driven engineer culture had such operation excellence is that the responsibility was driven the whole way up the management chain. If your manager didn't dig down to the root cause of a sev2, the director was damn well going to, and if the director didn't, Andy Jassy was going to call them on it.
That sounds like a lot of busy work, and a bunch of it undoubtedly is, but the flip side of it is that the actual root causes were addressed, and if the root cause needed serious cross-functional resources to address, the escalation would get you what you needed. Need approval to bounce 3 months of work off your roadmap to address? You have a VP on the line to make that call. Need a couple of other teams to fix their shit? Here's a principal engineer who outranks their directors to drive the work. And so on...
I'm here for the sentiment expressed by the SVP in this article: "I already believe we reached this point through rational choices. Let's talk about how to change the system so it doesn't happen again."
I do think his choice of words was a bit suboptimal. But, it got the point across.
> To drive change in your organization, don't ask "why did this happen?".
> Instead, ask:
> What are we changing so that the same class of failure is less likely next time?
Why are these framed as being almost mutually exclusive? I don't get it. "How/why did this happen?" informs how you answer "What are we changing?" Even the following section, "Reasonable people", functions in this way with a "Why did this happen?" followed by "What are we changing?".
Just feels like this article has some internal conflict with itself in order to achieve a "don't do that, do this" type of style.
One thing that's often left out that tends to grind organizations to a halt is that continuous improvement processes tend to be additive.
You always add rules, and alerts, and so on. You need to revisit existing processes to and see what can be removed or replaced.
My other nit with these is a term that's abused a lot "Root Cause", most complex issues have many contributing factors, not a single root cause, and if you force the teams to find one, they will.
I dislike RCA (Root Cause Analysis) it puts you in the wrong mindset, CFA would be better (Contributing Factor Analysis).
There's still some problem with that approach. OK, I will tell what I'm going change to prevent recurrence of the issue. How does that makes sense to the audience? Unless they just want hear that "some" change will be there and don't care about how that change would make any sense.
It boils down to what exactly is the ownership or accountability of your SVP around the issue. Why are they even bothered to ensure that there will be some change? If they are responsible for ensuring that the issue doesn't happen again, then they do need to say "I want the details".
"Okay, everyone is overworked, and we need you to hire more people for this team. We have too much tech debt, and we can't predict which bit of it will explode next week. The change needs to start with the attitudes of our executive leadership."
Oops, now I'm fired.
A good explanation is not a fix but is a good starting point to a good diagnosis which is the foundation of a good fix.
The issue is when the calls and urgency are raised before making some mental space to even verify that there is enough clarity on what happen and these interrupt the clarification and remediation planning process.
Saying you shouldn't ask "Why did this happen?" and you should ask "What we're changing?" instead is wanting to jump the gun and pretend a required step doesn't exist.
It does exists, is just the executive anxiety reasonable in terms of wanting to be protective of the business but as protective as a family member of a patient wanting a good outcome from the surgeon during the surgery.
Only worst because the surgeon's boss is this anxious family member, and she has a laboral gun pointing to the doctor's neck.
Strongly disagree with this.
In order to figure out a good solution, you have to have mastery over the problem. That takes a bunch of work.
> Understanding an issue is not the same as fixing it. A good explanation can make things worse. Once everyone agrees that the behaviour was reasonable, the urgency to change anything disappears.
This is a non-sequitor. We're not here to evaluate whether the behavior was reasonable, we're here to figure out how to avoid bad outcomes. If all the behavior was reasonable, then the problem lies somewhere else, and that's an important finding when crafting a proper solution.
The author is literally saying that asking why is the wrong question, and instead asking "What will we change?" is the right question. I get where this comes from, but you can't really answer what will change until you deeply understand why it happened.
I suspect the author still misunderstands the SVP. If I were the SVP of engineering saying this to someone, especially someone who is in a non-core engineering role like product, I probably recognize that they have a tendency to verbally spew and I want them to focus on the risk mitigation and response plan. I’ve already talked to my engineering leader because they informed me immediately, and I’ve already read the incident details.
But also the post reads like it comes from an inexperienced/immature service org, so maybe they’re just sorting their operational response process out.
"Okay, then I'm going to rewomble the dinglehop, which should un-galvatrate the percapitator." Why is the SVP in the meeting at all? If they don't want the details, they can't contribute to the change, or even evaluate whether the change makes sense at all. The SVP could be completely removed from this situation and the outcome would be exactly the same. Which shouldn't be surprising, since that's almost always true of all executives anywhere.
> Everyone nods their head, says "that makes sense", and we all move on with our day.
Why? Why do you need to be intimidated into actually making effective change? Why is everyone walking around with their hands tied, unable or afraid to enact change? Probably because of the terrible leadership at this company.
>> "We missed it because Alice was on holiday and Bob thought the Widgets team owned it".
> Okay. How do we make ownership unambiguous when someone is unavailable?
You cannot take this step without details. That Bob-Alice story is the details. You do actually need to know that story before you can take the step of knowing that "unambiguous ownership" would be a worthwhile change to make. What exactly is this SVP doing again?
This entire post could be summarized as "executives are useless" and "do the 5-whys", both of which everyone already knows.
This article may be more appropriate for LinkedIn.
The post-mortem process can cover all of these things:
- what happened
- impact
- root causes
- what will be changed (commitments)
You can't really talk about what to change w/o talking about root causes, and you can't talk about that without talking about what happened. You also can't consider the costs of commitments to be made w/o discussing actual and potential impact because resources are limited and the cost of opportunity is real.It's fine for executives to get only a summary of impact + commitments. But the whole post-mortem process has to be followed, and some people have to be aware of all the details.
I can't help but notice that a climate of fear around being deemed non-essential or otherwise getting on an authority figure's bad side tends to create significant mismatches in priorities and breakdowns of trust that could be upstream of significant organizational inefficiency in a variety of situations. Perhaps solutions to meta-problems of this nature might be high-leverage, as the thought leaders like to say
The executive did not display empathy.
People need to tell their stories. They need to be heard and have their work and its difficulty respected.
That's what went wrong. Everything else is an abuse victim rationalizing their abuse.
Every post-mortem I've done had:
1. timeline of events
2. impact
3. 5 whys — the details of how and why things happend
4. actionable items
I've never seen a post-mortem without actionable items.
Some of the best life advice I learned during conscription as a squad leader: do not explain if an explanation is not asked for. If you (or, importantly, your subordinates!) mess something up and are confronted by a superior, the only correct reply is "Yes sir I screwed up sir I’m sorry sir it won’t happen again".
I think the problem is when the execs aren’t looking at change-for-the-better. Instead they’re looking to appoint blame. Moving to a position of assumed competency in your staff and taking the perspective of how-to-improve-next-time is a real step forward in maturity. My guess is it’s rare and blame is easier: most management doesn’t reach this level, instead it’s immature.
I think that's a problem domain in itself. Identify when failures that occur even when skilled people doing their jobs properly, adjust the environment so that they no longer happen.
It is understandably a completely different task compared to what those skilled people are specialised to do. You probably need a dedicated role to do it.
Perhaps you could call them a manager. Their job is to see the multiple parts of the system. They should ask for the Details of what happened so they can determine why the problem occurred.
Consider one of the problems listed in the article
>"The alert fired, but the on-call engineer had already dealt with twenty low-value alerts that evening".
The engineer can say they were busy, they didn't see the alert, that they are swamped with things they think are low-value. Someone else can say the alert fired. Each person involved may have their own perspective, with different ideas as to what the problem actually is.
It's easy when you see problem described in terms of what the solution is. Someone needs to figure that out, to do that they need the details.
An SVP of engineering should care about the details, this is the essence of leadership. When you don’t care about the details you are just a manager and worthless i.e. you should be replaced with a leader.
Lots of folks are commenting on the internal dynamics, but the other part that didn't resonate with me was actually the what happens next part. This may be an overreaction on my part, stemming from watching lots of engineers suggest expensive solutions to relatively minor issues. And it's a mistake I've made myself before as well, I started in telecom with very strict standards for availability, so it's hard to beat out of me the desire to look through the most minor alarm for a potential problem brewing or get to a root cause of the most minor incidents.
And maybe this is just a part that the article dodges, but for something described as non-catastrophic, I would suggest that there shouldn't be a presumption of changes. In my view there should be an assessment of the risk of re-occurrence, and if it was a near miss, what is the risk had it not been a near miss. And then even should changes be part of the outcomes, how expensive are those to implement compared to whatever the issue was, and if they're too expensive compared to the risk, don't put resources into it.
My version of this is: "This is the kind of mistake that we get to make once. How do we prevent it from happening again."
I usually do want the details, but I also want to communicate my expectations up top and unambiguously.
“If your corrective action depends on people remembering a conversation from six months ago, you don't have a corrective action. You have organizational folklore.”
I think i’m going to call our ticket support system “folklore” from now on. It sure is used that way.
A Senior Vice President of ENGINEERING doesn't want the technical details?
This is a really really smart SVP. The answer to “how do we prevent this in the future” is to identify decisions that reduce that problem from happening again. This doesn’t have to be perfect. You can say “we are going to do X next time” and also say “but we don’t know how far X will work”. When the failure happens again, you can retire process X and move on to Y. The aim is to make a series of decisions that eventually get to the heart of the issue.
Everything else is a conversation that takes up space on a post-mortem or runbook.
> What they didn't want was for empathy to become the mechanism by which the organisation absolved itself of having to change.
That line is absolute gold
The assumption that something always has to change is the culture that leads startups to knee jerk their way into miles of red tape and performative bureaucracy
So in summary, instead of details propose how to change the system so that instead of hearing perfectly reasonable details about why the current system failed, go straight for changing the systems to prevent the chance of failure again.
I like this plan of action, it removes focus from what happened in the past to how can we prevent it from happening in the future.
There's a reason 5 Why's Analysis is a thing. If you just say "here's a bunch of stuff we're changing," but you haven't said enough about the root cause, then no one will understand whether we're doing enough (or too much; or the wrong things). I like the framing of being empathetic - everyone is smart, doing the best they can with their knowledge and experience at the time - but sometimes we didn't do all the right things, and getting feedback from more senior engineers helps drive that improvement.
I've never seen an incident report or postmortem that didn't include action items to cover the core issue noted in this writing: "can we / how do we prevent this failure mode from recurring?"
I'm not saying these moments of discovery / epiphany aren't valuable, they are and this retelling is enjoyably written.
I am saying that action items borne of incidents should be de rigueur.
What can be done, when writing an article, to not stumble during reading when an unknown abbreviation such as SVP occurs?
What if you work at an organisation that doesn't ask "How do we prevent this from happening again?"
Interesting. At our org we always do both. Root cause, what are we changing. I assumed that was standard practice.
If you just explain what happened and why, that's fine but...how are you going to make sure it doesn't happen again?
The problem I found more vexing as a manager was: how can you prevent this KIND of problem from occurring?
I'd have someone on my team make a technical error and xyz wouldn't work. Wed talk thru it, and they wouldn't make that exact mistake again. But there's literally 10k things that can go wrong in our system, so then a related mistake would happen later.
What they needed was improved pattern recognition vs if this / then that which comes from post mortems.
This reads like a LinkedIn post. I know kagi translate[1] supports "LinkedIn" language - could someone with an account there copy that article and translate from "LinkedIn" to "English"? I wonder if we'd uncover some simpler meaning this way.
Edit: I was wrong. It still reads like LinkedIn to me but “translating” doesn’t help. Thanks in any case.
SVP sounds like a dingle to me. Author is justifying it with some magical thinking, "Oh he's not being rude, he's being smart"
This is similar to the approach in healthcare, at least my experience in pharmacy.
When an error occurs we find the root cause but in 99% of cases a change to the environment/process is required to avoid it in the future.
Humans are all imperfect, their competence will change hour to hour let alone day to day. But you can control a system and put checks in place (admittedly a human can still do the process wrong, but then you need to think how the process can be clearer).
No-blame culture is very effective at providing an open environment to share mistakes, learn, but most importantly avoid reoccurance.
If I recall correctly, in The Pragmatic Engineer, this was described simply as "provide options, not excuses" and is a critical way of building credibility.
Best article I’ve read this month. The SVP expressed unqualified trust in the team. Asking the team for the next steps emphasizes this trust.
My father who was also involved in IT tells a similar story where the executive cut him off and said "I asked you for the baby, not the labor pains."
Love this idea; won't go over well at all with the other accountants I work with. MY CFO doesn't like graphs because they "hide too many details' and insists on everything presented to her being in a table. MY VP pulls out a calculator and re-does the math on tables in decks by hand to help him 'understand'.
A lot of leaders like to hide in the details because the bigger picture is a harder problem.
Just give me the solution to the problem of Lack of care to understand the details by people who can actually make a difference.
>If your corrective action depends on people remembering a conversation from six months ago, you don't have a corrective action. You have organizational folklore.
I appreciate this.
We call this incident postmortem, part of a company process. After serious problems and firefighting, when the patches and hacks and people working 24 hour and it admins wake on a Sunday morning to do work, we sit down and try to figure out how this will _never_ happen again.
> At first I thought "I don't want the details" sounded dismissive.
Well that is because it was delivered by someone who was actually being dismissive!
If they want to know what "we are changing" but don't want to know any detail as to why, even presented at an appropriate level, they are very likely looking for you to take all the risk of the decision and they are dismissing your need to have the necessary informed consent for it to become the we.
Now, if you're presenting that information at a level of detail that is truly irrelevant or deep, that is your problem to fix.
But otherwise be very wary of this pattern of communication where "we" might mean "you". Anyone with any sense of risk management should know not to do this.
If you have not seen it, watch Margin Call. When you get to the boardroom scene [0] you'll see this play out, but you will also see the (pretty rapacious, direct) boss understand properly that the information he is getting from a low ranking employee is nevertheless his problem to understand, even if only at a 10,000 feet level.
Beyond the drama of the way the scene ends, it's actually a very good model of how that interaction should play out for everyone — making it clear that everyone in the room shares in the "we" of it, including the decisionmaker.
I have been in rooms where it played out at least something like that, where the boss wants to make an attempt to share the burden of the detail, and I've been able to leave them feeling somewhat supported, or at least without feeling like I was being shafted.
And I've been in rooms with people who don't want to know the details but just want to know what "we" are going to fix — those guys have never been as good to work for. If they say "we" but leave you with the impression they may mean "you", it's bad.
[0] said scene, but really, really, watch the film if you never have:
https://www.americanrhetoric.com/MovieSpeeches/moviespeechma...
Thats funny. When I try to only give the conclusion, I'm often asked for the details.
> [The Manager] continued:
>> I know that if we get into the details, the reasons will be perfectly reasonable. You'll explain what happened, I'll understand why everyone made the decisions they made, and I'll empathise with you.
>> Then it'll happen again.
>> So I don't want the details. I want to know what we're changing.
Even if the manager's approach was correct, this is just a hurtful way to frame it to your team. It's great that the author could reverse-engineer what the manager was doing, but it shouldn't come to that.
Why couldn't the manager say something like:
> You are all great engineers, so I already know that everyone made the best decision available with the information they had at the time. What are we going to do differently moving forward?
I'm still not sure I agree with the strategy, because I don't see how you can understand what happens next without understanding what went wrong. But I can see how the framing might help steer the discussion.
If you don’t know what happened how would you be able to tell if the changes are addressing the problem? If you don’t care about the problem being addressed why jump in a call in the first place?
Post-mortem, 5-why, retrospection - all this to understand what and how it happened so the org may introduce change so that it doesnt happen again. Then another edge case happens and the cycle repeats. Humans do this cycle all the time - big orgs invented the above structure for this process to be visible and collaborative and so that the rest of the org can learn too. Then you put execs on top, they will come, ask some questions so they can sleep better at night. Corporate playbook, some people hate it, the rest understands it… and plays the game.
Learned this doing incident reviews. If the exec asks for details it usually means they don't believe you yet, so "skip to what changes" is actually the good outcome.
> Then I realised that "I don't want the details" wasn't being dismissive. The executive assumed that we were competent, and was saying "I already believe you. Now let's talk about what happens next".
Something about this feels wrong:
- If you trust them completely, you don't need to know what happens next either. Just trust them to do the right things, your leadership isn't required.
- If you don't trust them completely, how can you know if "what happens next" is appropriate without knowing the details?
Isn't this just reinforcing the idea that leadership doesn't need to have their feet on the ground?
The best leaders I've worked with were paying attention to things from the bottom-up as well as top-down.