FMEA practical guide

Failure modes and effects analysis, or FMEA, asks a practical question: how could a function fail, and what would follow? Its value is a traceable explanation of failure and a set of justified actions. A completed spreadsheet is only the record of that reasoning.

On this page

This guide explains how to scope an FMEA, distinguish causes from effects, interpret prioritization, and review the result. The examples are invented teaching cases, not assessments of actual equipment.

What FMEA means and where FMECA differs

FMEA systematically examines ways an item or process can fail and their consequences. It can cover hardware, software, human actions and interfaces. FMECA adds a criticality assessment: consequences are ranked by severity, often alongside other measures of importance. FMECA does not necessarily mean a fully quantitative probability model. These distinctions follow the published scope of IEC 60812:2018.

A design FMEA might examine a fan assembly; a process FMEA might examine whether that assembly is connected correctly. They need different functions, causes and controls. Copying the design worksheet into a process review does not establish equivalent coverage.

Begin with a function and a boundary

Write what successful performance means before listing failures. “Cooling system” is an item name. “Provide the airflow needed by a demonstration cabinet during a start request” is a function that can be challenged.

Record the configuration, operating phase, interfaces, assumed environment, analysis depth and exclusions. Ask whether the analysis includes startup, continuous operation, maintenance and shutdown. These phases can have different failure modes and consequences.

Assemble evidence such as a functional description, drawings, interface definitions, requirements, test findings and relevant failure history. Identify which claims are measured, inferred or unknown. Agree how concerns will be escalated before seeing the scores, so the team does not change its rules to obtain a convenient result.

Keep mode, cause and effect separate

Consider one fan in a two-fan cabinet:

  • Function: establish airflow when commanded
  • Failure mode: fan does not start
  • Possible cause: an open electrical connection
  • Local effect: that fan supplies no airflow
  • System effect: airflow may remain adequate if the other fan performs its assumed function; redundancy is lost

“Fan failure” is too broad. Failure to start, stopping after startup, and inadequate airflow are different modes. “Overheating” may be a later effect, but it is not automatic: the heat load, remaining airflow and time must support that conclusion. A cause at one level can become a mode at a lower level; record the analysis level consistently.

An authored example in six steps

Assume a classroom demonstration cabinet has two parallel fans, either of which can meet the specified airflow alone. They share a power supply. A separate airflow indicator reports status but does not control either fan. This example addresses startup only and claims no real failure rates.

  1. Define success. At least one fan establishes sufficient airflow following a valid command, with the specified air path available. The actual acceptance threshold belongs in a real requirement; it is deliberately unspecified here.

  2. Examine each function. Fan A failing to start loses one path. Repeat for fan B, but do not assume the second row is an independent event merely because the component has a different label.

  3. Examine shared support. A power supply that provides no output can disable both fans. Record the shared consequence rather than describing each fan as an unrelated failure.

  4. Examine detection. An indicator stuck at “normal” can conceal loss of airflow. It does not itself remove airflow in this configuration. Link it to the relevant diagnostic concern instead of confusing indication with physical performance.

  5. Identify candidate changes and evidence. A separate supply, improved connection design or a diagnostic check might address different concerns. Each needs an owner, intended mechanism, acceptance criterion and suitable verification. Listing a proposal does not make it an effective control.

  6. Review the remaining concern. Confirm that implemented changes work under the stated assumptions. Revisit the worksheet when configuration, requirements or evidence changes. NASA's GSFC FMECA handbook explicitly treats the analysis as a maintained record throughout development and operation.

The example exposes a useful distinction: two fans provide potential redundancy, but that statement alone says nothing about the shared supply or the reliability of detection.

What the numbers do and do not mean

Some procedures use severity, occurrence and detection ratings, often written S, O and D, and multiply them into a risk priority number: RPN = S × O × D. Other procedures use different prioritization rules. There is no universal scale or universally valid action threshold.

Severity describes the assessed consequence. Occurrence concerns the specified cause or mode under the chosen procedure. Detection concerns the defined opportunity for existing controls to discover the problem. State that opportunity and the direction of the scale. A monitoring display alone does not establish detection at the required stage. Any credited mitigation also needs a timely, capable response.

With ordinal ratings, a larger number indicates a higher category; a score of 4 does not establish twice the probability or consequence of 2. Bowles identifies the problems of treating rankings as numerical quantities and of different profiles producing the same RPN in his research paper.

For a deliberately invented five-level example, suppose higher D means more difficult detection. The profiles S=5, O=2, D=2 and S=2, O=5, D=2 both produce 20. One emphasizes severity, the other occurrence. The equal product does not make them interchangeable. Keep the individual ratings and their rationale visible, and apply the project's severity and escalation rules.

An RPN has no physical unit and is not a failure probability, failure rate or expected annual loss. Dividing it by its maximum does not create a probability. A procedure may define occurrence categories using probability bands, but those bands still need a time or demand basis and supporting evidence. Their existence does not turn the complete RPN into a probability.

Review coverage by operating state and interface

A large number of rows is not evidence that the important functions were covered. Create a coverage view using operating states and interfaces: storage, start, operation, shutdown, maintenance and return to service where relevant. For each state, ask what the item must do, what it must not do and which external services it requires. A failure mode can be absent in one state and important in another.

Extend the invented fan example to maintenance. If fan B is deliberately unavailable, failure of fan A to start now removes the assumed airflow function by itself. The physical failure mode has not changed, but the system effect has. Keep the maintenance configuration explicit rather than silently using the two-fan consequence in every row. This is an original analytical extension, not a prescribed maintenance authorization.

Detection is a chain with a time requirement

“Alarm fitted” is an incomplete detection claim. Identify the relevant failure symptom, sensing path, indication, recipient, time available and response that the analysis credits. The alarm can detect a symptom without identifying its cause. A valid warning arriving after the consequence has developed may be too late for the claimed protective effect. Detection and mitigation therefore need separate descriptions even when they occur close together.

In a hypothetical teaching case, assume a dangerous condition develops after 40 seconds, an indication becomes available after 15 seconds, and the assumed response takes another 30 seconds. The combined 45 seconds misses the stipulated window. These arbitrary times demonstrate why the presence of a detector alone is insufficient; they do not establish human-performance data or a real safe response time. Uncertainty in each interval can matter as much as its nominal value.

Distinguish prevention from consequence reduction

An action that removes the cause of a failure mode differs from one that limits the resulting effect. Improving a connector can reduce a particular initiating mechanism; independent airflow detection can reveal a loss; a suitable protective response can limit a later consequence. They address different parts of the argument and need different verification. Do not change every worksheet rating merely because one action was completed.

For each proposal, state the mechanism it addresses and the evidence needed to show that mechanism changed. If a design modification introduces new interfaces or maintenance needs, examine those as well. An action is not automatically beneficial in every state. The useful question is what has improved under which assumptions, and whether an important concern moved elsewhere.

Close actions with evidence that matches the claim

NASA’s GSFC FMECA guidance treats analysis as a maintained engineering record. A practical action record should distinguish proposed, implemented and verified states. Verification should reference a drawing, inspection, analysis or test result appropriate to the specific claim, with its configuration and limitations. Merely replacing “open” with “closed” does not establish residual performance.

Before accepting a revised assessment, compare the new evidence with the original reason for action. Confirm that the affected rows and interfaces were updated and that the consequence description still matches the operating state. Keep unresolved questions visible even when their numerical priority is low. The result should explain what the team now knows, not only what it has entered into a form.

What a useful deliverable contains

Retain the function and requirement, mode, cause, local and final effects, existing controls, supporting evidence, assumptions, prioritization rationale, recommended action, owner and verification status. Separate current controls from proposed changes. Keep residual assessments linked to actual implementation evidence.

Check for missing interfaces, vague “operator error” labels, unsupported detection claims, mitigation claims without response capability, and low scores being used to dismiss severe consequences. Do not sum RPNs across unrelated rows as though the total measured system risk.

FMEA is useful for breadth and traceability, but a conventional single-mode review does not by itself quantify interacting failures. Where combinations determine the outcome, develop an appropriate system model such as a fault tree. Neither a long worksheet nor a low score proves completeness, safety or compliance.