Every engineer remembers certain incidents. The service that went down during peak traffic. The deployment that broke a critical workflow. The database change that had unexpected consequences. The alert that arrived at 2 a.m.
The immediate response is usually the same. Restore service. Reduce customer impact. Do a root cause analysis, get the system working again. These priorities are entirely appropriate. But once the incident is over, another question emerges.
What did we learn from what happened?
Many teams focus exclusively on identifying the technical cause. Good engineers go further. They try to understand the decisions, assumptions, and trade-offs that created the conditions for the failure.
Fixing the Problem Is Only the Beginning
The first responsibility after an incident is recovery. Customers need working systems. Teams need stability. Operations need confidence. But restoring service and understanding the incident are not the same thing.
A fix explains how the failure was resolved. It does not necessarily explain why it occurred. Good engineers know that technical recovery is only the first step. The deeper learning often begins afterward.
Cause vs Story
Every incident has a technical cause. Most also have a story.
Eventually someone discovers the immediate cause. A configuration error. An unexpected dependency. A capacity issue. A software defect. This explanation is important. But it is rarely the complete story.
Behind every technical failure is usually a sequence of decisions. Assumptions that were made. Risks that were accepted. Trade-offs that were chosen. Constraints that influenced behaviour. The incident is often the final chapter of a much longer story.
Good engineers learn to examine that story.
Look Beyond the Trigger
Imagine a deployment introduces a production outage. The technical cause may be straightforward. A bug reached production. Case closed. Or perhaps not.
How did the bug escape testing?
Why did the team believe the change was safe?
Were there warning signs that were ignored?
Did deadlines influence the decision?
Did previous successes create overconfidence?
The deployment may have triggered the incident. The conditions that allowed it to happen may have existed long before. Examining these conditions is often where the most valuable learning resides.
Avoid the Search for a Villain
One of the least useful questions after an incident is:
“Who caused this?”
The question feels natural. People want explanations. Organizations want accountability. But blame rarely improves understanding.
When teams focus on finding a person to blame, learning often stops. People become defensive. Information is withheld. Important details remain unexplored.
Good engineers ask a different question:
“What made this decision seem reasonable at the time?”
That question often reveals far more than assigning responsibility ever could.
Decisions Make Sense in Context
Looking backward creates a powerful illusion. Once the outcome is known, the correct choice often appears obvious. But the people involved did not have access to the future. They acted based on the information available at the time.
The goal of learning is not to judge past decisions using today’s knowledge. The goal is to understand how those decisions were made. What information was available? What pressures existed? What alternatives were considered? What assumptions seemed reasonable?
These questions help teams understand their decision-making process rather than simply evaluating the outcome.
Incidents Reveal Hidden Assumptions
Many assumptions remain invisible until something breaks.
A system is assumed to handle expected traffic. A dependency is assumed to be reliable. A process is assumed to catch errors. A team is assumed to have sufficient knowledge. Most of the time these assumptions remain untested. Incidents expose them.
This is one reason production failures can be valuable teachers. They reveal weaknesses that success often hides.
The Goal Is Better Decisions
The purpose of an incident review is not perfection. No team eliminates every failure. No engineer predicts every outcome. The goal is to improve the quality of future decisions.
Perhaps a hidden risk becomes visible. Perhaps a testing gap is discovered. Perhaps communication improves. Perhaps assumptions are challenged earlier.
The lesson is rarely: “Never make mistakes.”
The lesson is usually: “How can we make better decisions next time?”
Why This Matters
Engineering judgement becomes visible when things do not go according to plan. Incidents force teams to confront reality. They reveal assumptions, expose trade-offs, test decisions, and highlight weaknesses in systems and processes.
The technical cause is important, but it is only part of the story. Because every incident has a technical cause. Most also have a judgement story.
