Root-cause analysis: testing explanations against physical evidence

Turn a machinery failure story into competing causal hypotheses, discriminating checks and corrective actions whose effectiveness can be verified.

On this page

A plausible account of a machinery failure is not yet a demonstrated cause. Root-cause analysis asks which explanation fits the sequence, physical mechanism and available evidence, what alternative explanations remain and which intervention would reduce recurrence. The purpose is to improve the system that produced the event. A chain of “why” questions can organize inquiry, but the answers still need evidence.

Preserve the as-found condition before it disappears

Safe stabilization and any applicable reporting obligations come first. Within those boundaries, preserve logs, settings, photographs, removed parts, debris and the configuration found after the event. Cleaning a fracture surface, discarding a filter or overwriting a controller log can remove evidence needed to distinguish the initiating failure from consequential damage.

MAIB’s explanation of its work emphasizes timely evidence gathering before material is lost or changed. A company maintenance investigation does not replace an official casualty investigation or authorize interference with it. Record who collected an item, where it came from, its orientation and any changes made. A box of unlabelled fragments has much less explanatory value than a documented component history.

Write an event statement that avoids assuming the cause

“The pump lost discharge pressure while running at the recorded duty” is an event description. “The operator caused cavitation” already embeds an explanation and a person-focused conclusion. Start with the failed function, operating conditions, observed sequence and consequence. Separate direct observations from witness recollection and later interpretation.

Build a time line from multiple sources and retain uncertainty. A control-system timestamp, a handwritten log and a witness’s estimated time may have different precision. Mark when a signal was measured, when it was received and when it was recorded. Agreement between two reports copied from the same original is not independent corroboration. Trace the origin of evidence rather than simply counting how many documents repeat a statement.

Test whether the chronology is actually resolved

In an original example, an alarm is stamped 10:00:00 with ±2 s timing uncertainty and a trip is stamped 10:00:01 with ±2 s uncertainty. The intervals overlap: 09:59:58–10:00:02 and 09:59:59–10:00:03. The nominal one-second ordering does not establish that the physical alarm condition preceded the trip. A causal claim relying on that ordering needs better timing evidence.

This does not make the records useless. They constrain the events to a short period and can be combined with faster recorder data, relay sequence information or physical evidence. State which orderings are established and which remain possible. A diagram with precise arrows should not conceal uncertainty that is larger than the time separation being interpreted.

Alarm nominal time 0 with bounds−2 to+2 seconds and trip nominal time 1 with bounds−1 to+3 share one time axis. The trip-minus-alarm difference can range from−3 to+5 seconds, permitting either order.
Original bounded-time diagram at 48 drawing units/second. The records alone permit either order; nominal ordering does not establish physical causation. Difference bounds assume no additional constraint linking the two timing errors. A known shared clock offset could cancel in the difference, so correlated errors need separate treatment. Bars are bounds, not confidence intervals.

Develop competing mechanisms and predicted observations

For a damaged bearing, candidate explanations might include lubricant starvation, contamination, misalignment, overload or damage introduced during fitting. Each should predict observations: surface morphology, debris, oil-path condition, dimensions, temperature history or assembly evidence. Choose checks that distinguish the candidates rather than collecting more information that all candidates equally explain.

HSE’s HSG245 investigation workbook emphasizes systematic evidence review and immediate, underlying and root causes. The useful discipline is to remain open to alternatives and connect explanation to evidence. A “five whys” chain that ends in “lack of training” without testing the physical sequence has only moved the assumption to a more general level. Several contributing causes can coexist.

Use dimensional checks as plausibility tests

Assume a one-metre steel member with thermal-expansion coefficient 12 × 10⁻⁶ per kelvin experiences a 60 K temperature increase. Its free expansion would be αLΔT = 0.00072 m, or 0.72 mm. If a hypothetical assembly has only 0.30 mm free movement, thermal constraint is a plausible question to investigate. The expansion estimate is an original dimensional example, not a reconstruction of an actual casualty.

The calculation does not prove excessive force or the failure cause. Real restraint stiffness, differential temperatures, geometry and sliding behaviour determine the response. Measure the relevant clearances and temperature history and inspect the predicted contact or deformation evidence. A model that produces a large number can direct inquiry, but it must survive comparison with the physical configuration and alternative explanations.

Use negative evidence only within detection limits

“No debris was found” is meaningful only if the relevant location was inspected with a method capable of finding the predicted debris and before it could be removed. An empty strainer after cleaning cannot refute an earlier contamination event. Likewise, absence of an alarm is weak evidence if the sensor was unavailable, the threshold inappropriate or the recorder missing data.

For each hypothesis, distinguish evidence supporting it, evidence contradicting it and observations that are simply unavailable. Do not convert “not recorded” into “did not happen.” Confidence should reflect the quality and independence of the evidence, not the investigator’s familiarity with the explanation. A common failure mechanism is a reasonable starting candidate but not a substitute for examining this particular event.

Ask whether the proposed cause changes the outcome

A counterfactual asks whether removing the alleged cause, while keeping other relevant conditions comparable, would plausibly have prevented the event or reduced its likelihood. It helps distinguish a background condition from an effective causal contributor. The question is analytical; it does not justify recreating a hazardous failure on operating machinery.

Test the proposal through safe evidence, calculations, controlled engineering tests or comparison with relevant unaffected units. If both damaged and undamaged units share the alleged cause, investigate what additional condition distinguishes them. That does not automatically disprove the cause, because exposure and susceptibility may differ, but it challenges an incomplete explanation. Avoid choosing a cause merely because it is easy to change.

Separate local repair from recurrence prevention

Replacing the damaged part restores capability; preventing recurrence addresses the demonstrated contributors. If incorrect material entered through an ambiguous substitution process, another identical replacement may repeat the problem. A corrective action should identify the causal link it changes, the responsible owner, completion evidence and a way to assess effectiveness.

An instruction to “be more careful” is difficult to verify and may leave an error-prone interface untouched. A clearer specification, controlled part identity or redesigned connection can be more testable when supported by the investigation. Training may be appropriate for a demonstrated knowledge gap, but should not be the default answer to every event. Match the action to the evidence and the actual work conditions.

Judge absence of recurrence against opportunity

If a purely illustrative unchanged event process has a rate of one per 1,000 operating hours, the probability of no event in the next 100 hours is exp(−0.1) = 90.5%. Observing no recurrence over that short period is therefore quite compatible with no improvement at all. The calculation assumes a constant Poisson process; real mechanisms may require a different exposure model.

Combine follow-up exposure with direct verification that the intended causal link changed. A corrected drawing, confirmed part identity or measured clearance can provide earlier specific evidence than waiting for another failure. Keep both implementation and effectiveness evidence visible.

Close with confidence, unresolved questions and verification

A strong report distinguishes established facts, supported causal conclusions and unresolved possibilities. Explain why alternatives were rejected or retained and which evidence would change the conclusion. There is no requirement that every investigation discover one unique deepest cause. Several interacting technical and organizational conditions may be necessary to explain the event.

After implementing actions, verify the intended change and monitor a relevant opportunity for recurrence. “No repeat failure” over very little exposure is weak evidence; absence of the targeted defect during a meaningful inspection can be more direct. Retain the investigation so later events can test it. Root-cause analysis remains useful when its claims are traceable and revisable, not when the final diagram merely looks complete.

Sources