All articles

What Is Root Cause Analysis (and Why Most Investigations Never Reach It)

Ask ten people what "root cause analysis" means and you'll get ten answers that are mostly a description of what happened, followed by whoever's name came up in the meeting. That's not root cause analysis. It's an incident report with a scapegoat.

Root cause analysis is a method: you trace an outcome back through the conditions and events that produced it, until you reach a cause that, if you fix it, actually prevents recurrence — not just this specific failure, but the class of failure it belongs to. The difference matters more than it sounds like it should, because almost every investigation that goes wrong goes wrong by stopping one or two steps too early.

The three places investigations stop short

Symptom-level. "The pump failed" is a symptom. Why it failed — bearing wear, a missed lubrication interval, a maintenance schedule that assumed a duty cycle the pump was never actually running — is the investigation. A report that describes the failure in more detail isn't going deeper; it's describing the same symptom more precisely.

Blame-level. "The operator missed a step" explains what happened, but it's rarely where the causal chain ends. Why did the operator miss it — was the step buried in a procedure nobody had revised in six years, was the interface ambiguous, was the operator covering two stations because the site was short-staffed that shift? Stopping at "operator error" closes the investigation exactly where it becomes useful.

Single-cause fixation. Real incidents are rarely one thing. They're usually two or three conditions that individually wouldn't have caused a problem, combining at the wrong moment. An investigation that finds "the" cause and stops looking has usually just found the most visible one.

What it actually takes to get past these

None of this requires more effort from investigators — most teams doing shallow RCA are working hard, just without a process that keeps pushing. What it requires is structure:

  • A method that keeps asking why instead of stopping at the first plausible answer, with somewhere to record the causes that were considered and ruled out — not just the one that was accepted.
  • Evidence discipline — each step in the chain should be something you can point to, not something that sounds right in the room.
  • Separation of cause types — an immediate cause (what triggered the event), a contributing factor (what made it worse or more likely), and a systemic or root cause (what, if fixed, prevents the class of problem) are different things, and conflating them is how "operator error" gets filed as a root cause.

This is the gap between an investigation that produces a defensible finding and one that produces a document. A guided questionnaire — sections and why-options that force the chain past the first answer, with each cause tagged by type and rated for risk — is one way to build that structure into the process instead of relying on the facilitator to remember it under pressure. It's how the questionnaire in GuidedRCA is built, but the underlying discipline applies whether or not you're using a tool for it.