Knowledge / Risk and reliability
Uncertainty in risk analysis: ranges, sensitivity and decision robustness
Use uncertainty ranges, sensitivity tests and decision thresholds without confusing scenario bounds with statistical confidence.
On this page
A numerical risk estimate is a result of a model, data and assumptions. Its apparent precision does not reveal how firmly those ingredients are known. A useful uncertainty analysis asks whether the decision changes across credible alternatives, which uncertainties matter most and what evidence would improve the choice. This article develops original hypothetical maritime-equipment examples using only public methodological references. It is an educational guide. Its invented probabilities, costs and thresholds must not be used as operational acceptance criteria.
Define the decision before refining the estimate
Start with the actual choice: compare maintenance options, assess a proposed operating change or decide what evidence is needed before approval. State the population, time horizon, consequence measure and constraints. A fleet-average annual loss estimate cannot automatically answer whether one ship may sail with an impaired safety function. Those questions involve different exposure, authority and evidence. Mandatory requirements and non-negotiable safety constraints remain relevant even when a model estimates a small expected loss.
A decision statement might be: “Compare two hypothetical measures for reducing equipment-related repair and downtime costs over the next year, assuming both otherwise satisfy applicable obligations.” This deliberately narrow example avoids treating monetized property loss as a substitute for life safety or environmental protection. Define the baseline consistently. If one option is compared with current performance and another with an improved baseline, the resulting ranking has no common meaning.
Separate variability from lack of knowledge
Aleatory uncertainty refers to variation represented as random within the chosen model, such as the occurrence of demands or varying loads. Epistemic uncertainty concerns what is not sufficiently known about parameters, model structure or completeness. The distinction depends on the modelling purpose. A changing load might be represented as a random variable, while uncertainty about its distribution remains a separate knowledge problem. More data may reduce the latter without eliminating the former.
The NRC’s NUREG-1855 Revision 1 explicitly distinguishes parameter, model and completeness uncertainties in risk-informed decisions. Its regulatory application is nuclear, not maritime. The useful methodological lesson is to avoid treating uncertainty as only a range around one input. A very precise estimate for a model that omitted a major pathway can still be a poor decision basis.
Give every range a meaning
An interval could be an observed minimum and maximum, an engineering bound, a confidence interval, a Bayesian credible interval or a set of scenarios. These are different statements. A range selected by experts to explore plausible values is not automatically a 95% confidence interval. A confidence procedure describes repeated-sampling coverage under its assumptions; a Bayesian interval depends on a likelihood and prior model. Identify which interpretation is intended and why it is justified.
Record units, exposure, source population and relevance alongside the range. If failure data came from different duties or environments, a narrow statistical interval may understate transfer uncertainty. Conversely, a broad generic range may ignore useful local observations. Preserve both the original evidence and the adjustment made for the current application. This makes disagreement productive: reviewers can challenge a particular transfer assumption rather than debate an unexplained high and low value.
Propagate a simple model transparently
Suppose hypothetical annual expected repair-and-downtime loss is R = f × p × C. Let initiating frequency f be 0.05 per year, conditional probability of damaging progression p be 0.2 and mean cost per damaging event C be €1,000,000. The central calculation is €10,000 per year. This is an expectation over repeated comparable exposure, not a forecast that the owner will pay exactly €10,000 next year.
Now assume exploratory ranges f = 0.02–0.08 per year, p = 0.1–0.4 and C = €250,000–€1,000,000. Multiplying the low corners gives €500 per year; the high corners give €32,000. These are bounding scenario calculations under the assumed combinations, not probabilistic percentiles. Before interpreting them, ask whether the extreme inputs can occur together. Higher load may increase both damage probability and cost, making independent sampling inappropriate.
Find the threshold where the preferred choice changes
Assume hypothetical option A costs €4,000 annually and reduces the modelled loss to 40% of baseline. Option B costs €1,000 annually and reduces it to 70%. For this limited economic comparison, total expected annual costs are 4,000 + 0.4R and 1,000 + 0.7R. They are equal when R = €10,000. The central case therefore gives a tie, rather than a defensible claim that either option is decisively better.
At R = €2,000, totals are €4,800 for A and €2,400 for B. At R = €30,000, they are €16,000 and €22,000. The ranking reverses across the plausible range. This threshold is often more useful than adding decimal places to the central estimate. It identifies the question worth investigating: is the relevant baseline loss likely to lie below or above €10,000, and do the assumed reduction factors survive scrutiny?
Distinguish sensitivity from uncertainty contribution
A variable can have a strong mathematical effect but little uncertainty, or a modest effect over an extremely uncertain range. For the product R = f × p × C, doubling any one input doubles the output if the others are fixed. That observation alone does not identify which input deserves research. The answer also depends on the credible range, dependence and whether improved knowledge can change the decision.
One-at-a-time sensitivity is useful for explaining direction and detecting errors, but it can miss interactions. A model might show little effect from changing either of two inputs separately while their joint change crosses a failure threshold. Use combined scenarios where physical mechanisms suggest interactions. A tornado plot is a presentation device, not proof that the chosen ranges or independence assumptions are correct. Document how each variation relates to a credible operating situation.
Use simulation as a propagation tool, not an evidence generator
Monte Carlo simulation samples inputs from specified distributions and propagates them through a model. It can represent nonlinear outcomes and dependencies when these are modelled correctly. It cannot make an unsupported distribution credible. Increasing the number of samples reduces numerical sampling noise; it does not repair a missing scenario, a biased dataset or a mistaken assumption that correlated inputs are independent.
A reproducible simulation record should include distributions, parameter values, correlations or conditional structures, random-seed policy and convergence checks relevant to the reported outputs. Rare tail quantities may require much more evidence than a stable mean. Compare selected runs with hand calculations and limit cases. If a probability becomes negative or a physical capacity is exceeded, check the input model instead of trimming inconvenient samples silently. Such corrections change the model and must be documented.
Interpret zero observed failures carefully
Suppose an invented observation programme records zero failures over 30,000 comparable operating hours. Under a constant-rate Poisson model, the probability of zero failures is exp(−λT). Solving exp(−λT) = 0.05 gives a one-sided 95% upper rate bound of −ln(0.05)/30,000, approximately 9.99 × 10⁻⁵ per hour. The observed maximum-likelihood estimate is zero, but the uncertainty bound is not. Zero observations do not prove an impossible event.
The NIST reliability handbook provides the equivalent zero-failure bound through mean time between failures. The example here is newly constructed. It assumes suitable exposure accounting, complete detection and a constant rate; ageing, mixed equipment populations or dependence can invalidate that model. Do not read the frequentist bound as a 95% posterior probability about the true rate without introducing a Bayesian model. Model suitability matters as much as the arithmetic.
Challenge the model’s missing pathways
Parameter uncertainty asks how uncertain a value is within the model. Model uncertainty asks whether the selected relationships adequately describe the system. Completeness uncertainty asks what may have been omitted. A pump model that varies its internal failure rate but excludes loss of a shared suction source can give a narrow, reproducible and misleading result. A sensitivity exercise should therefore include alternative structures, not only alternative numerical inputs.
Practical challenges include reviewing operating transitions, maintenance configurations, common utilities, human response and external events. A workshop can identify candidates for analysis, but it cannot guarantee completeness. Record unresolved mechanisms and explain how they affect the decision. If a potentially dominant pathway is unquantified, say so explicitly. It may justify a conservative restriction, further engineering or additional evidence through the authorized decision process, rather than another round of parameter sampling.
Choose information that can change the decision
More data are useful when they reduce decision-relevant uncertainty at reasonable cost and within the available time. In the hypothetical option comparison, improving the estimate of baseline loss may help if it resolves the €10,000 crossover. But if the larger uncertainty concerns whether option A truly achieves a 60% reduction, collecting more baseline data alone may not settle the ranking. Research the assumption that controls the decision, not simply the one easiest to measure.
NIST’s discussion of sparse failure data cautions that highly reliable systems can yield little failure information despite substantial testing. Consider condition observations, relevant near misses, exposure quality and mechanism-specific evidence alongside event counts. Expert judgment can help structure uncertainty when data are scarce, but record the question, evidence, disagreement and calibration limits. Calling an estimate “expert” does not remove its uncertainty or make every expert opinion equally relevant.
Communicate a decision with conditions for revisiting it
A useful uncertainty statement gives the central result, the meaning of its range, the assumptions that drive the conclusion and any unresolved structural gaps. Separate the chosen action from the evidence supporting it. If the preference is robust across all credible cases, explain that. If it is conditional, state the condition plainly. Avoid a single coloured risk category that hides whether uncertainty crosses the organization’s decision boundary.
The deliverable should include an input register, reproducible calculation, sensitivity cases, dependency assumptions and a short decision note. Identify what new observation would trigger review: an increased demand rate, an unexpected shared failure, different operating exposure or evidence that a control performs worse than assumed. Uncertainty analysis is successful when it makes the next decision clearer and more honest. Its purpose is not to make every number certain before action becomes possible.
Sources
- NUREG-1855 Revision 1, Guidance on the Treatment of Uncertainties Associated with PRAs in Risk-Informed Decisionmaking, March 2017 · US Nuclear Regulatory Commission · Source check date: 2026-10-06
- NIST/SEMATECH e-Handbook, Constant repair rate (HPP/exponential) model · US National Institute of Standards and Technology · Source check date: 2026-10-06
- NIST/SEMATECH e-Handbook, Lack of failures · US National Institute of Standards and Technology · Source check date: 2026-10-06