Fault tree analysis practical guide

Fault tree analysis, or FTA, starts with one precisely defined unwanted outcome and works backward through the conditions that can produce it. It is useful when redundancy, shared resources or combinations of failures determine whether a system function is lost.

On this page

A fault tree can support a qualitative review without any probabilities. Quantification is a further step that requires suitable evidence, a consistent event basis and explicit treatment of dependencies. IEC 61025:2006 is the general FTA standard; its public description covers assumptions, events, failure modes and modeling guidance.

Define the top event before drawing

“Cooling failure” leaves important questions unanswered. Is the concern failure to start, inadequate performance during operation, or loss of a protective response? These are different top events and can require different trees.

Specify the function, failure criterion, configuration and relevant interval or demand. Record environmental assumptions, initial states, repair treatment and exclusions. Distinguish the physical event from its indication: an alarm that fails to report loss of airflow is not itself a loss of airflow.

A useful event statement is understandable without the analyst standing beside the diagram. It should also have a clear true-or-false meaning under the analysis conditions.

Gates express logic before arithmetic

An OR gate means at least one input event is sufficient. An AND gate means all input events are required. An AND gate does not, by itself, establish independence or a particular order in time.

For two events A and B:

  • P(A OR B) = P(A) + P(B) − P(A AND B)
  • P(A AND B) = P(A) × P(B given A)

Only with independence does the second expression become P(A) × P(B). “OR means add; AND means multiply” therefore omits essential conditions. NASA's fault-tree handbook, Chapter 6 explains these distinctions.

Basic events mark the chosen stopping points for decomposition. They are not necessarily the deepest physical causes. An undeveloped event is a disclosed gap, not an event whose probability can silently be set to zero.

An authored startup example

Consider an invented demonstration cabinet with two parallel fans. Either fan alone can provide the required airflow. Both receive the same valid start command and share a power supply. The air path is assumed clear. Continuous-running failures, repairs, command failures and other shared environmental causes are excluded from this small teaching model.

  1. Set the top event T: required airflow is not established on one start demand.

  2. Define C: the shared supply is unavailable on that demand. Define A and B as the fans' intrinsic failure-to-start conditions when adequate power is present. Keep supply loss out of their intrinsic-failure data.

  3. Write the logic: T = C OR (A AND B). Supply loss alone is sufficient. With supply available, both fan paths must fail for the top event to occur.

  4. Identify the minimal cut sets: {C} and {A, B}. A cut set is a combination sufficient to produce the top event. It is minimal when removing any member makes that combination insufficient. “Minimal” does not mean all cut sets have the same number of members.

  5. Introduce invented numbers solely to demonstrate calculation. Let P(C)=0.002. Given that C has not occurred, let each fan have failure-to-start probability 0.01, and assume these two intrinsic failures are conditionally independent.

  6. Split the calculation into the mutually exclusive supply-unavailable and supply-available cases:

P(T) = 0.002 + (1 − 0.002) × 0.01 × 0.01 = 0.0020998.

The result is approximately 0.210% for this defined demand. It is not a measured performance claim. Ignoring the supply would give 0.0001, or 0.010%, for the two-fan branch alone. The difference illustrates why an apparently strong redundant arrangement can be dominated by shared support.

Check what is actually independent

Different part numbers, separate rows and separate branches are not evidence of independence. Examine shared power, control, cooling, environment, installation, maintenance and design defects.

In the example, shared power is an explicit functional dependency. Other common-cause mechanisms may require separate modeling. Do not count the same contribution both in an all-causes fan probability and in an added shared-cause event.

Repeated occurrences of one event must retain one identity. For example, C OR C equals C. Squaring its probability because it appears twice would create an imaginary independent event. Overlapping cut sets likewise cannot generally be added as though mutually exclusive. Model structure and event identity matter as much as numerical precision.

Probability, rate and unavailability are different outputs

A probability is dimensionless and lies between zero and one. It needs context: failure on a particular demand, failure before a specified time, or being unavailable at a specified instant. A failure rate has inverse-time units, such as per hour. An event frequency also has an event-per-time basis but is not automatically the hazard rate of a non-repairable item.

For an initially functioning, non-repaired item with constant failure rate λ, the exponential model gives P(failure by t) = 1 − exp(−λt). This relationship and its assumptions are described in the NIST reliability handbook. It is not a universal conversion for every fault-tree input.

For example, λ = 0.00002 per hour and t = 50 hours give about 0.001, or 0.10%. The product λt is dimensionless. Multiplying two raw rates at an AND gate would instead produce inverse-hours squared, not a top-event probability. Repairable availability models, latent failures and testing intervals require the appropriate additional model.

A cut set is a structural result before it is a number

The example’s single-event cut set {C} reveals a shared-supply vulnerability regardless of the numerical probability assigned to C. Its importance to a decision still depends on the function, consequences and evidence. A very small unsupported probability should not make a single-point dependency disappear from the review. Conversely, a single-event cut set does not by itself prove that the actual system violates a requirement.

Keep cut-set identity separate from ranking. Two cut sets can overlap through a repeated basic event, and their probabilities then cannot generally be summed without accounting for that overlap. The NASA fault-tree handbook develops qualitative structure and quantification as distinct tasks. The analyst should preserve event definitions and the calculation method so that a numerical output can be traced back to the logical claim.

An original sensitivity comparison

Hold the example’s conditional fan probabilities at 0.01, but reduce the invented shared-supply probability from 0.002 to 0.0002. The same model gives 0.0002 + 0.9998 × 0.0001 = 0.00029998. Alternatively, retain supply probability 0.002 and halve both fan probabilities to 0.005. The result is 0.002 + 0.998 × 0.000025 = 0.00202495.

The first hypothetical change produces a larger reduction in this particular top-event estimate. This is a comparison inside the stated model, not a recommendation to purchase or redesign equipment. Costs, feasibility, new dependencies, other failure modes and uncertainty are omitted. Sensitivity identifies where assumptions influence a result; it does not prove that the corresponding real-world improvement is available or that its claimed probability can be achieved.

Match evidence to the event being quantified

A dataset labeled “pump failures” may combine failure to start, loss while running, reduced output and maintenance removals. The appropriate input depends on the tree’s event definition and demand or time basis. Record population, exposure, operating conditions, observation period and any exclusions before transferring a rate or probability. A number with a reputable source can still be the wrong number for the modeled event.

Zero observed failures does not imply zero probability. A small number of tests may provide little information about a rare event, and repeated tests of one unchanged item need not represent a diverse fleet. Missing observations, reporting thresholds and common environments can also affect the evidence. Where suitable data are unavailable, a qualitative conclusion or explicitly bounded sensitivity analysis can be more honest than a precise unsupported estimate.

Test the boundary with counterexamples

Ask whether a plausible condition can make the top event true while every modeled basic event is false. In the fan example, a blocked common air path would do so, because clear airflow was an explicit assumption. That observation does not invalidate the teaching tree; it identifies what the tree excludes. A real assessment must decide whether the assumption is justified or the event needs inclusion.

Also ask the reverse question: could the modeled combination occur without the top event under the actual operating state? If so, the supposed sufficient condition may be incomplete. These counterexamples examine logic before quantification. They are particularly useful when a solver produces a reassuring result that masks an ambiguous top-event definition.

Review the model and communicate its limits

Test simple cases before trusting a solver: in this example, C alone must cause T, A alone must not, and A together with B must cause T. Trace each branch to the system description. Review the exclusions, data applicability and dependence assumptions with people who understand the design.

Deliver the tree, event dictionary, minimal cut sets, assumptions, data provenance and unresolved gaps. Where quantified, include the output's basis, uncertainty and sensitivity to influential assumptions. A point estimate with many digits is not evidence of certainty.

Ordinary static AND/OR trees do not automatically represent ordering, switching, repair or changing operating phases. When these affect the answer, use an appropriate dynamic or state-based model. The NRC SAPHIRE technical reference illustrates why mission time, repairability and uncertainty need distinct mathematical treatment. An FTA remains conditional on its selected top event and modeled causes; it cannot certify that every possible hazard has been covered.