Knowledge / Risk and reliability
Common-cause failures: the hidden limit of redundancy
Why redundant equipment can fail together, and how dependency mapping, conditional probability and maintenance evidence expose shared weaknesses.
On this page
Adding a second component can improve reliability dramatically when failures are sufficiently independent. It can improve it much less when both components share a vulnerability. The central question is therefore not simply how many pumps, sensors or power sources exist, but which events can defeat the required function across all of them. This educational guide uses original hypothetical examples to explain that distinction. It does not prescribe a ship’s redundancy architecture, common-cause parameter or acceptable risk level.
Define the function before counting equipment
Suppose a hypothetical essential cooling service requires one of two pumps to deliver the specified flow for a stated mission. Two physically installed pumps provide redundancy only if either can actually support that requirement under the relevant conditions. If both are needed at maximum load, the arrangement has a different success criterion. If one cannot start after a loss of power, its nameplate capacity does not establish available redundancy for that initiating event.
Describe the mission, flow requirement, operating environment and restoration assumptions before drawing a reliability diagram. A system that tolerates one running-pump failure may still lose function during changeover, because the standby start, valve alignment or electrical supply is not available. Include these interfaces in the model boundary. Otherwise the analysis can report excellent component reliability while omitting the very path that determines whether the ship retains the function.
Distinguish several kinds of dependence
A common support-system failure directly removes something both components need, such as electrical power, cooling, instrument air or suction inventory. A cascading failure occurs when one failed component damages or overloads another. A common-cause failure concerns multiple component failures associated with a shared cause and a coupling mechanism. In practice these categories can overlap, but separating the mechanisms helps decide what to model explicitly and what remains in a residual common-cause term.
The NRC’s 1999 regulatory issue summary on common-cause failures identifies coupling characteristics such as similar hardware, maintenance, operating arrangements and environment. It is historical nuclear-sector guidance, not a maritime failure-rate source. The transferable question is what makes the redundant components susceptible together. “They are separate units” answers only a small part of that question.
Map dependencies across the whole functional chain
For the hypothetical cooling service, map each pump’s suction path, discharge path, driver, starter, control signal and supporting utilities. Then map what the two paths share. A common strainer, a single valve upstream of both branches or one switchboard can be more important than the pumps’ individual failure rates. Record normally open interconnections and temporary configurations as well as the design’s preferred line-up.
Also map organizational dependencies: identical maintenance instructions, a common calibration reference, the same spare-part batch or a shared configuration file. A well-intentioned maintenance campaign can introduce the same error into both trains. Physical separation does not prevent that mechanism. The purpose of the map is not to distrust every shared feature; it is to identify which features can defeat the claimed independence and what evidence or protection addresses them.
See why independence changes the arithmetic
Assume that, in the absence of a common-cause event, each hypothetical pump has a 0.01 probability of failing the mission. If those residual failures are independent and either pump suffices, the probability that both fail is 0.01 × 0.01 = 0.0001. This multiplication is valid only for the stated conditional model. It does not follow merely from the fact that two pumps have different equipment tags.
Now introduce a separate common-cause event with probability 0.0005 per mission that disables both. Assume the residual independent failures occur with the stated 0.01 probabilities conditional on that common event not occurring. The total failure probability is then 0.0005 + (1 − 0.0005) × 0.01² = 0.00059995. The result is approximately six times the independent-only estimate. All values are invented to demonstrate the model, not estimates for marine pumps.
Avoid double counting the same shared event
The calculation above makes the common event and the residual branch mutually exclusive by conditioning on whether the common event occurred. If a fault tree already contains failure of a shared electrical supply, that event should not also be embedded uncritically in the pumps’ independent rates and added again as generic common cause. Define the data boundaries: do recorded pump failures include external support failures, or only failures inside the pump assembly?
This bookkeeping can affect the answer as much as numerical sophistication. A model can be conservative in one place and non-conservative elsewhere when definitions drift. Keep an event dictionary identifying failure mode, component boundary, exposure basis and treatment of shared causes. Review data coding before fitting parameters. Two databases can report different failure rates because they count different events, even when both describe their equipment as “pumps.”
Use common-cause models without turning parameters into facts
A beta-factor model assigns a specified fraction of component failure probability or rate to a common-cause contribution under its model convention. It can be useful for screening simple redundant arrangements, but it compresses many mechanisms into one parameter. More detailed models can distinguish how many components in a group fail. Neither approach removes the need to identify shared vulnerabilities or establish that the data apply to the present design and operating context.
The NRC’s NUREG/CR-6268 description explains the collection, evaluation and coding of common-cause event data for nuclear probabilistic assessment. Its relevance here is methodological: parameter estimates depend on how events and groups are defined. No nuclear common-cause value is imported into this article’s marine example. A borrowed number should never conceal differences in maintenance, environment, failure mode or demand exposure.
Challenge diversity as carefully as duplication
Different technology can reduce some shared vulnerabilities. A distinct sensing principle may avoid one common measurement fault; separated locations may reduce exposure to a local fire or flood. But diversity can create new interfaces, training needs and maintenance complexity. Different manufacturers may still use the same internal component or rely on a common external utility. The relevant question is which specific failure mechanism the diversity removes and which remain.
For the hypothetical service, replacing one pump brand does not address a single obstructed suction source. Moving a control panel does not separate two power feeds routed through the same vulnerable space. Evaluate the entire path against named events. Record the evidence for separation, functional diversity and environmental survivability rather than assigning a generic “diverse” credit. Diversity is a design property to demonstrate, not an adjective that guarantees independence.
Treat a single observed defect as information about peers
If one redundant component is found with an incorrect setting, the other should not automatically retain its previous assumed failure probability. Ask whether the same procedure, tool, person or batch could have affected it. The observation may provide evidence of a shared susceptibility even before a second complete failure occurs. Preserve the distinction between confirmed failure, degradation and potential exposure while investigating the common mechanism.
The NRC’s NUREG-2225 addresses how observed deficiencies can change common-cause potential in retrospective nuclear risk assessment. The general lesson is conditional updating, not adoption of nuclear regulatory calculations aboard ships. In the cooling example, finding the same incorrect maintenance instruction applicable to both pumps changes the question immediately: the remaining pump’s apparent availability needs relevant verification, not reassurance based only on its last routine running indication.
Maintenance states can temporarily erase redundancy
Redundancy is a property of the current configuration as well as the design. If one train is isolated, failure of the remaining train may directly cause loss of function. If both are serviced with the same faulty method, the danger may persist after both work orders are marked complete. Staggering tasks can reduce some shared exposure, but it is not universally sufficient and may introduce other operational constraints.
Analyze the sequence of isolation, work, testing and restoration using approved procedures and the actual success criterion. Identify when redundancy is unavailable and who must know. A successful post-maintenance test should address the affected function and failure mechanism. Simply rotating equipment duty does not establish independence or eliminate hidden common defects. Do not conduct live fault-injection tests on essential ship services based on this educational discussion; verification requires competent planning and authorization.
Report what the calculation leaves out
The original example ignores repair during the mission, changing load, imperfect switching and partial failures. These omissions are acceptable only because the example illustrates one probability distinction. An actual model should include mechanisms material to the decision. A small common-cause term can dominate a highly redundant design, while repair and timing can be important for long missions. State what changes if the function requires continuous availability rather than success at one demand.
Sensitivity analysis should vary uncertain common-cause assumptions and examine the underlying mechanisms, not merely print more decimal places. A very low estimated common-cause probability may be unsupported by the available exposure. Absence of recorded simultaneous failures is weak evidence if reporting has missed degradation, equipment has rarely been demanded, or the fleet is small. Document those limitations and identify which additional observations would most improve the decision.
Produce a dependency register and a justified response
A useful deliverable combines the success criterion, functional diagram, shared-dependency register, event definitions, data basis and uncertainty assessment. For each material shared vulnerability, specify whether it is removed, explicitly modelled, controlled or unresolved. Link proposed actions to the mechanism: separate a supply only if that supply is the issue; improve maintenance verification when the coupling is procedural. Keep technical proposals subject to the vessel’s normal approval process.
The final review should be able to answer what survives each relevant failure and why. It should also explain how a new defect, modified interconnection or changed maintenance method would trigger reconsideration. Redundancy is not defeated only by spectacular simultaneous accidents. It can be weakened quietly by a shared assumption that no one checked. Making that assumption visible is often the most valuable result of common-cause analysis.
Sources
- Resolution of Generic Issue 145, Actions to Reduce Common-Cause Failures, 1999 · US Nuclear Regulatory Commission · Source check date: 2026-10-06
- NUREG/CR-6268 Revision 1, Common-Cause Failure Database and Analysis System, September 2007 · US Nuclear Regulatory Commission / Idaho National Laboratory · Source check date: 2026-10-06
- NUREG-2225, Basis for the Treatment of Potential Common-Cause Failure in the Significance Determination Process, September 2018 · US Nuclear Regulatory Commission · Source check date: 2026-10-06