Software FMEA: incorrect logic, stale data and systematic failure modes

Build software failure modes around functions, data age and state transitions, then connect a concrete timing budget and unit-conversion example to testable controls without borrowing hardware failure rates.

On this page

Software can deliver the wrong result while its processor, network and program loop appear healthy. The cause may be an incorrect requirement, a missed state transition, stale input or a valid-looking value interpreted with the wrong units. Software FMEA asks how these deviations propagate through the controlled system, which conditions reveal them, and what evidence supports the proposed detection and response.

Analyse the function and interface, not just the executable

NASA describes software FMEA as a bottom-up examination of software functions, interfaces and their system consequences. A useful boundary might be “accept a sensor update”, “convert engineering units”, “evaluate a command permissive” or “enter a fallback mode”. A row called merely “software fails” is too broad to expose the mechanism.

For each boundary, retain the required output, timing, valid input domain and relevant operating states. Startup, normal operation, degraded operation, restart and maintenance may have different contracts. The same output can be correct in one state and hazardous in another. Include hardware/software interactions while keeping the software mode distinct from its downstream physical effect.

Separate a systematic defect from its activation

A design or coding defect can persist in a released version and be activated by a particular input, sequence or configuration. It does not need the program to physically wear out. Its manifestation can nevertheless appear intermittent because the triggering conditions are intermittent. Resource exhaustion, concurrency and corrupted hardware inputs also deserve explicit mechanisms rather than a single undifferentiated software label.

A hardware component’s constant random failure rate cannot simply be assigned to that logic defect. Operational-profile or statistical software evidence can inform an assessment when its assumptions are justified, but an FMEA worksheet by itself does not create such a probability model. Prioritize credible effects and evidence gaps without inventing a numerical occurrence rate for each code path.

A useful row links mode, propagation, detection and evidence

GSFC-HDBK-8004 separates task and interface failure modes from mechanisms, causes and effects. For an input handler, one mode is “accepts an old acquisition as current”. A possible mechanism is refreshing the validity timer on every message receipt. The local effect is a falsely fresh data record; the next effect is a decision based on obsolete process information.

The proposed control must interrupt that chain: preserve a trustworthy acquisition age and validity state, define the handling of repeated or out-of-order updates, and specify what the consumer does when freshness is lost. The verification case freezes acquisition while transport messages continue. A network heartbeat alone would pass that test while leaving the actual stale-data mode undetected.

Worked age example: recent receipt is not recent measurement

Use hypothetical timestamps on a common corrected time basis. The last valid acquisition occurred at 0 ms, the latest message arrived at 240 ms, and the consumer evaluates at 250 ms. Receipt age is only 10 ms, but the underlying sample is 250 ms old. Resetting a freshness timer to message arrival therefore hides 240 ms of age in this example.

If producer and consumer clocks differ, the age comparison needs a bounded synchronization error or a protocol that supplies an equivalent trustworthy age. An untrusted timestamp should not gain validity merely because it exists. Record the source epoch, sequence and quality meaning, and distinguish a retransmission from a new acquisition. These are interface requirements, not assumptions to leave to a display label.

Worked response budget: include detection and physical completion

Stipulate a requirement that the defined fallback response is completed within 150 ms of the last valid acquisition. Let the age threshold H be 100 ms, the maximum age underestimation from relative clock error 5 ms, the consumer scan delay 30 ms, and output plus actuator completion 40 ms. Then the conservative bound is H + 5 + 30 + 40 = 175 ms, exceeding the requirement by 25 ms.

The algebraic maximum threshold under those assumptions is 150 − 5 − 30 − 40 = 75 ms. Choosing H = 70 ms would give 145 ms, leaving 5 ms of budget. This does not select a real timeout: the normal data-arrival envelope, jitter, clock faults and nuisance responses must also be checked. If those requirements conflict, redesign the timing or architecture instead of declaring the smaller number a complete fix.

Original stale-data example: acquisition at 0 ms, receipt at 240 ms and evaluation at 250 ms give receipt age 10 ms but sample age 250 ms. A hypothetical completed-response budget of 150 ms is exceeded by a 100+5+30+40=175 ms chain. A 70 ms age threshold gives 145 ms and 5 ms remaining budget, subject to the stated bounds and normal-update feasibility.
Original deterministic interface and timing examples. Timestamps share a corrected reference; the separate response budget includes a bounded clock error, scan delay and actuator completion. These figures are not software failure rates or a timeout recommendation for real equipment.

Wrong units can survive a range check

For a separate deterministic example, an input of 1000 kPa converts to 10 bar because 1 bar = 100 kPa. A faulty division by 1000 instead produces 1 bar. Against an illustrative 8 bar alarm threshold, the correct value exceeds the threshold and the incorrect value does not. Both can lie inside a broad numeric input range, so a simple “number is plausible” check may miss the defect.

The FMEA mode is wrong scaling or unit interpretation, not “pressure sensor inaccurate”. Its cause could be an interface contract mismatch or a conversion error; its effect depends on the consumer. Evidence should include unit-tagged requirements, boundary and representative conversion tests, and end-to-end confirmation of the displayed or used engineering quantity. The 8 bar value is a teaching threshold, not an equipment setting.

State and sequence errors need temporal test cases

A stale command retained across restart, an invalid mode value falling through a default branch, or an acknowledgement mistaken for physical completion can each violate the functional contract. Define the transition preconditions, allowed outputs and the state reached after a failed transition. There is no universally safe fallback value: stopping, holding or transferring control must follow the system hazard analysis.

A small sequence example also prevents a common mistake. For a 16-bit counter, the forward modular increment from 65535 to 0 is (0 − 65535) mod 65536 = 1; an unchanged counter gives 0. A plain numeric “new greater than old” check rejects the valid wrap. Modular arithmetic alone is still insufficient for restart epochs, replay or large gaps, so those cases need explicit protocol rules and tests.

Original command-state and interface review matrix

State and deviationRequired behaviour and discriminating test
RUN + stale valid dataA fresh receipt must not reset acquisition age. When the bounded freshness condition fails, request the specified protective response and verify its physical completion within the timing budget.
RUN + wrong unitsInterpret a pressure value using the agreed unit contract before evaluating the permissive. Inject 1000 kPa and check the 10 bar result; reject or handle a mismatched contract as explicitly required.
STARTUP → RUN + illegal transitionWith the required initialization or permissive absent, the command must not enable RUN. Test the rejected transition and the permitted transition separately, including restart after interruption.
RUN + omitted protective branchAssert the defined hazardous-input condition while the loop and watchdog remain healthy. The required trip request and resulting safe-state evidence must still appear; a requirement-to-test trace exposes a missing protective branch.

The four rows are original specification examples, not universal safe-state choices. NASA’s software-hazardous-requirements guidance supports making each analysis-derived protective behaviour traceable and verifiable; the actual permitted states, transition priorities and physical response must be defined for the system.

A watchdog proves only what it actually observes

A loop watchdog can detect that an execution path stopped updating it. It need not detect an incorrect result, stale source data or a semantic error that still executes on time. Similarly, a checksum can establish transmission integrity without establishing that the transmitted value is current or correctly scaled. Name the observable symptom each detection mechanism covers.

NASA’s hazardous-requirements guidance connects analysis-derived controls with verification evidence. For the examples above, useful tests deliberately retain old data under continuing communication, exercise the worst timing alignment, cross a counter wrap and restart, and verify units through the full consumer path. A test report should identify the build, configuration, conditions and observed response, not just say “fault handling tested”.

Keep the analysis tied to changes and unresolved assumptions

Each row should retain the function and mode, initiating condition, local and system effects, detection, compensating action, owner and evidence reference. When a requirement, data format, scheduler period or library changes, revisit affected rows and their interfaces. A correction can remove one failure path while creating another, especially around initialization and recovery.

Software FMEA provides a structured way to ask what the software can do incorrectly and how the system responds. It complements requirements review, static analysis, integration tests and system hazard analysis; it does not replace them or prove the absence of all defects. The practical outcome is a set of explicit failure mechanisms and verifiable controls, with timing and data semantics carried through to the physical result.

Sources

  1. NASA Software Engineering Handbook, Version D — 8.05 Software Failure Modes and Effects Analysis.
  2. NASA GSFC-HDBK-8004 — Guideline for Failure Modes and Effects Analysis and Risk Assessment (2024).
  3. NASA Software Engineering Handbook, Version D — SWE-192 Software Hazardous Requirements.