Knowledge / Risk and reliability
When failure order matters: the limits of a static fault tree
Compare simultaneous failure logic with cold standby, transfer coverage and response-time requirements using an explicit two-state transition calculation.
On this page
Two items marked failed at the end of a mission do not reveal which failed first, whether a spare started in time or how long the process was unprotected. Those distinctions matter when cooling, steering support or essential electrical service must remain available through a transition. A static fault tree can still be useful, but its event definitions must not silently replace a sequence-dependent requirement with an end-state snapshot.
Identify the fact that an AND gate forgets
The ordinary expression A AND B records that both conditions are true. It gives the same result for A then B, B then A, and simultaneous occurrence. That is appropriate when the top event depends only on their coexistence. It is insufficient when A changes B's operating exposure, initiates a transfer or prevents an action that was needed earlier.
The NASA dynamic-fault-tree chapter discusses sequence-dependent events, spares and state-model solutions. The general lesson is to preserve the relevant history. Adding the word standby to a static component label does not supply the missing dormant-state, switching or timing model.
Define a deliberately ideal cold-standby model
Assume one active unit and one identical cold spare. While inactive the spare cannot fail in this ideal model; when active it has the same constant failure rate λ as the first unit. Transfer is immediate and initially perfect. There is no repair, common cause, load interaction or maximum interruption constraint. Failure means losing service before mission time t.
These assumptions are strong. A real stored or idle unit can degrade, lose utilities or fail on demand. Keeping the ideal case explicit is useful because it isolates one mechanism: the spare accumulates operating exposure only after the active unit fails. A calculation for two units operating from time zero answers a different reliability question.
Derive survival from two mutually exclusive paths
The first success path is that the original unit survives the full mission, with probability e^(−λt). The second is that it fails at time s and the spare survives the remaining t−s. Integrating λe^(−λs)e^(−λ(t−s)) from s=0 to t gives λt e^(−λt). The paths are exclusive because the first unit either survives or fails before the end.
Therefore reliability is R(t)=e^(−λt)(1+λt), and mission failure probability is Q=1−R. The units check is important: λ has units h⁻¹ and t has units h, so λt is dimensionless. This is an original derivation for the stated ideal model, not a formula for every standby arrangement.
Compare the two exposure assumptions numerically
Take λ=0.002 h⁻¹ and t=20 h, giving λt=0.04. The cold-standby model gives Q≈0.000779. If two independent active units instead operate for all twenty hours and service is lost only after both fail, the static result is [1−e^(−0.04)]²≈0.001537.
The difference is not a contest between correct and incorrect mathematics. The models describe different exposure histories. The static active-parallel case exposes both units throughout the mission, while the ideal spare waits without accumulating failure risk. Applying either result to actual ship machinery requires evidence about idle failure, activation, repair and the functional response required during transfer.
Introduce imperfect transfer as its own event
Let c now denote the conditional probability that transfer succeeds when the first unit fails. Assume it is independent of failure time and equals 0.95. Only the second success path changes, so R(t)=e^(−λt)(1+cλt). With the same λ and t, Q≈0.002701. A five-percent transfer-failure probability has a visible effect even though the spare's active reliability is unchanged.
This c is transfer coverage, not a basic-event common-cause parameter and not proof-test coverage. Reusing one letter in separate models is harmless only when definitions remain explicit. A failed-start probability, a delay distribution and incomplete fault detection may need separate terms rather than being compressed into one undocumented coverage number.
A successful start can still be too late
Suppose an invented process can tolerate no more than 20 s without the required service. Detection takes 8 s and starting plus delivery establishment takes another 15 s, so the interruption is 23 s. The spare eventually runs, but that sequence fails the functional requirement by 3 s. A final-state inspection showing the spare running would miss the earlier violation.
The allowable interruption must come from process dynamics and the defined consequence, not from an assumed generic standby rule. Stored energy, fluid inventory or thermal capacity may bridge a delay, while changing load can shorten the margin. A time-dependent performance model is needed if interruption duration, partial output or recovery trajectory controls the hazard.
Choose a state description that preserves the question
A simple transition model can distinguish both units available, spare running after a successful transfer, and service failed. More realistic models may need dormant failure, detected versus hidden faults, repair, restoration and competing demands for a shared spare. Each added state should represent a distinction that affects a transition rate, success condition or consequence.
Continuous-time Markov models usually use memoryless holding times in their basic form. A fixed start delay or age-dependent degradation may require supplementary states, a semi-Markov formulation or simulation. Naming the method does not remove these assumptions. Record why its timing representation is adequate for the decision and test special cases whose answers are known.
Keep common causes and recovery visible
A fire that disables both units, a failed common suction or a shared start-air problem is not repaired by modeling transfer order more precisely. Dynamic sophistication can coexist with an omitted dominant cause. Conversely, assuming instantaneous repair may greatly overstate performance when access, spares or safe isolation are unavailable during the mission.
Treat recovery as a defined function with timing and evidence, not a generic success branch added after every failure. If the same crew must diagnose, isolate and start several systems, their actions may compete for time. The model should reflect the capacity of the response arrangement rather than multiplying nominally separate success probabilities.
Replace ideal dormancy with a stated dormant failure rate
Let the active rate remain λA while the spare can fail during dormancy at rate λD. With perfect transfer, no repair and independent failure mechanisms, the spare-success path must include survival during the waiting period s. Its integrand becomes λA e^(−λA s)e^(−λD s)e^(−λA(t−s)). The original unit's survival path remains e^(−λA t), even if the unused spare fails.
For λD>0, integration gives R=e^(−λA t)[1+λA(1−e^(−λD t))/λD]. As λD tends to zero, this approaches the ideal cold-spare result. When λD=λA, it becomes the independent active-parallel reliability 2e^(−λA t)−e^(−2λA t). Those limits are useful model checks. They do not establish that a real inactive pump has a constant dormant rate or identical active and dormant mechanisms.
Keep mission survival separate from later availability
A repaired system can be available at the end of a mission despite having lost service earlier. If the requirement is uninterrupted cooling, that earlier loss remains a mission failure after repair. A state model that merges repaired states with states that never failed may therefore compute point availability rather than the required probability of continuous service.
Use a separate absorbing failure state when the first loss irrevocably violates the modeled mission requirement. If recovery before a consequence deadline is allowed, represent the allowable interruption and recovery conditions instead. The choice is driven by the engineering endpoint, not by which state diagram is easiest to solve. Report both availability and mission reliability only when each has a defined use, and keep their labels distinct in comparisons.
Use the simpler model when its boundary is adequate
A static tree is reasonable when sequence does not affect the defined outcome, or when carefully derived conditional events already capture the relevant sequence. A more detailed model is justified when changing order, delay, dormant exposure or repair materially changes the answer. Complexity should resolve a known deficiency rather than merely enlarge the diagram.
Present the static comparison, the added states or timing variables, and the numerical effect of the added mechanism. State which assumptions remain unverified for any real equipment. The calculations here are educational reliability models; they do not establish emergency-service compliance, acceptable interruption time or operating permission for a ship.
Sources
- Fault Tree Handbook with Aerospace Applications · NASA · Source check date: 2026-10-06
- SAPHIRE technical-reference public abstract · US Nuclear Regulatory Commission / Idaho National Laboratory · Source check date: 2026-10-06