Knowledge / Risk analysis methods
Event-tree recovery: success criteria and time windows
Define recovery by restored function before a physical deadline, then test timing, access, dependencies and the accounting of recovered sequences.
On this page
A failed automatic response does not always determine the final outcome. Another available function or a feasible human action may restore the needed service before damage occurs. Event-tree recovery credit represents that possibility, but only within a stated scenario and time window. The statement that a system can eventually be repaired is not enough. The analysis must establish what is restored, when it becomes useful and which conditions make the recovery unavailable.
Define the function that must return
A recovery-success statement should describe the required output at the receiving system. For a cooling scenario, starting a pump may be an intermediate action while useful recovery requires sufficient flow through the affected consumer at an acceptable inlet condition. For an isolation scenario, reaching an actuator may be intermediate while useful recovery requires the relevant inflow to stop before a boundary is exceeded.
The endpoint determines the clock. Loss of a service for a few seconds, a reversible temperature excursion and irreversible damage are different consequences. A tree that uses one generic recovered label for all three can conceal that the action succeeded mechanically but failed to prevent the particular loss being counted. Define the endpoint before estimating a recovery probability.
Distinguish recovery from repair
IAEA SSG-3 (Rev. 1), paragraph 5.108 distinguishes functional recovery from repair and calls for feasibility justification. This is a nuclear PSA reference, not a maritime authorization. In a general engineering model, restoring service through an already available alternate path and replacing a damaged component are different scenarios with different information, access and resource demands.
A repair estimate derived from ordinary maintenance includes conditions that may disappear during the initiating event. A spare can exist ashore yet be unavailable within the scenario. The correct tool can exist on board yet be inaccessible. An analyst should not attach a routine mean repair time to a recovery branch without examining those differences and the distribution of completion times relevant to the deadline.
Build a consistent time budget
The public NRC IDHEAS report summary identifies structured task analysis and time uncertainty as parts of human reliability modelling. A maritime adaptation still needs its own task evidence. Separate the time until a usable cue appears, diagnosis and decision time, access or preparation time, execution time and the delay before the restored output reaches its success criterion.
These intervals should share one start reference. If travel and diagnosis overlap, adding their full durations may double count time; if they must occur in series, using only the longer one may be optimistic. A timeline can reveal that distinction. A short actuator stroke says little about a response whose dominant delay is recognizing the correct fault or obtaining access.
Work through a bounded timing example
Use an explicitly invented timing distribution, not a crew performance claim. Assume cue, diagnosis and access together take a fixed 45 s, with no overlap. Conditional on the recovery being accessible and otherwise feasible, let execution time T be uniform from 30 s to 90 s. The NIST uniform-distribution reference supplies the mathematical form; it does not establish that real response times are uniform.
If useful completion is required by 105 s after the initiator, execution must finish within 60 s. The conditional timing-success probability is (60 − 30)/(90 − 30) = 0.50. If the probability that the recovery is accessible and feasible in this sequence is assumed to be 0.80, and the stated T distribution is conditional on that state, overall recovery success is 0.80 × 0.50 = 0.40. No additional independence assumption is needed for that conditional product.
Compare deadline and delay changes
Keeping the other assumptions fixed, a deadline of 120 s permits 75 s of execution and gives timing success (75 − 30)/60 = 0.75, hence overall recovery probability 0.60. Alternatively, retain the 105 s deadline but increase the fixed delays to 60 s. Execution then has only 45 s, giving timing success 0.25 and overall recovery probability 0.20.
These sensitivities show that apparently modest timing changes can alter the credited branch substantially. They do not justify extending a physical deadline or assigning a more favourable distribution. The deadline should come from the relevant physical model and acceptance endpoint. Measurement uncertainty, changing loads and the time needed for the recovered service to become effective may make that deadline uncertain too.
Partition the failed path into disjoint outcomes
For a separate illustrative tree, suppose the initiating frequency is 0.20 per operating year and automatic response failure has conditional probability 0.050. The sequence entering the recovery question has frequency 0.010 per operating year. With conditional recovery success 0.40, the recovered part is 0.004 and the unrecovered part is 0.006 per operating year. They sum to the original 0.010.
Do not retain the entire 0.010 as an unrecovered sequence and then add 0.004 as recovered; that counts the same entering population twice. Do not subtract recovery from an unrelated path either. The recovery opportunity belongs to the particular state that supplies its cue, available equipment and time window. A different failed automatic response may demand a different recovery model.
Retain dependencies with the original failure
A recovery action can rely on the same signal, power or mistaken interpretation that defeated the original response. Calling it manual does not remove that link. If the first failure has removed the only indication, a recovery model that assumes immediate correct diagnosis contradicts the path. If the failed item physically prevents the alternate route, the latter cannot receive its normal availability credit.
Repeated attempts also require care. Two tries are not automatically independent trials. The first can consume the available time, exhaust a stored resource, change the equipment state or provide new diagnostic information. A formula such as one minus the square of a single-attempt failure probability is justified only by its actual probability assumptions, not by the presence of two boxes on a procedure.
Use evidence that resembles the scenario
Useful evidence can include relevant simulator observations, task walk-throughs, equipment response tests and physical calculations, each with its limitations stated. A drill with advance knowledge may measure execution under prepared conditions but provide little information about surprise diagnosis. Tests that omit access restrictions or allow unlimited retries cannot directly establish the probability of a time-bounded response in a damaged environment.
Record unsuccessful or incomplete observations as well as completed ones. If a test stops at its time limit, the completion time is not known merely because the record contains that limit. Distinguish late completion, incorrect action, unavailable resources and no attempt. Combining them into one average time can hide the very failure mechanisms the event-tree branch is meant to represent.
State what recovery credit leaves unresolved
A timely restored function can prevent one consequence while leaving another. A temperature excursion may have occurred before cooling returned, or isolation may stop further release while leaving material already outside its normal boundary. The end-state label should retain those residual conditions instead of treating recovered as an all-purpose healthy state.
A credible recovery branch therefore carries a function, deadline, scenario conditions, evidence basis and uncertainty description. It is not an instruction to attempt an action under unsafe conditions. The model should represent feasible capabilities within the actual operating framework, and it should expose the consequences of withholding uncertain recovery credit rather than conceal them behind an optimistic average.
Sources
- SSG-3 (Rev. 1), 2024 · IAEA · Source check date: 2026-10-07
- NUREG-2199 Volume 1 public summary · US NRC · Source check date: 2026-10-07
- Uniform Distribution · NIST/SEMATECH · Source check date: 2026-10-07