Knowledge / Risk and reliability
LOPA proof testing: coverage, interval and hidden dangerous failures
Derive average hidden-failure probability, split partial-test coverage from full testing and expose the floor that more frequent partial testing cannot remove.
On this page
A protection layer may look healthy until a demand reveals that it cannot act. Proof testing is intended to discover relevant hidden failures before that demand, but a test only reduces the contribution of failure modes it can reveal and restore. Shortening a test interval is therefore not enough information to justify a better PFDavg. The coverage, full-test horizon, restoration and actual demand pattern must be defined together.
Identify which failure population is being averaged
Let λDU be the constant rate of dangerous undetected failures for a single nonredundant function. Dangerous means the failure prevents the specified protective action; undetected means ordinary diagnostics have not revealed it. Do not insert all failures, including safe trips, into this parameter. The function boundary includes the relevant sensing, logic and final action, not merely a component's catalogue label.
The HSE partial-proof-test example publicly illustrates splitting contributions by component and failure mode. Its indexed excerpt explicitly identifies simplifying assumptions. The present calculations use independently chosen values and a declared ideal model; they do not reproduce an approved proof-test interval or establish a safety-integrity claim.
Derive the complete-test approximation
Assume a perfect test and immediate restoration every T hours, no other downtime and random demands sampling the test cycle uniformly. At age t after restoration, hidden failure probability is 1−e^(−λDU t). Its exact cycle average is 1−[1−e^(−λDU T)]/(λDU T). For small λDU T, expansion of the exponential gives PFDavg≈λDU T/2.
The factor one-half comes from averaging the increasing hidden-failure exposure across the interval. It is not a universal property of every protective system. Detected failures awaiting repair, test downtime, incomplete restoration, redundant channels and common causes require additional modeling. The approximation applies to the stated random-failure contribution, not the entire safety function by default.
Check a complete-test numerical case
Take λDU=2×10⁻⁶ h⁻¹ and T=4,380 h. The dimensionless product is 0.00876, so the first-order estimate is 0.00438. The exact average is approximately 0.004367. The small difference confirms that the linear approximation is adequate for demonstrating this particular contribution, not that the underlying failure rate is known accurately.
At the end of the interval, the instantaneous hidden-failure probability is about 0.008722, roughly twice the cycle average. A demand concentrated just before testing does not sample the average uniformly. If operating practice links high-risk transfers to the end of a maintenance cycle, use the actual timing relationship rather than automatically applying PFDavg.
Split coverage into two failure-mode populations
Suppose a partial test every TP=4,380 h reveals 80% of λDU, while the remaining 20% is revealed only by a full test every TF=17,520 h. Here coverage C=0.8 is a fraction of the dangerous-undetected failure rate, not a fraction of hardware items or test steps. Assume each population is restored perfectly when its applicable test occurs.
To first order, PFDavg≈CλDU TP/2+(1−C)λDU TF/2. The contributions are 0.003504 and 0.003504, giving 0.007008. The uncovered twenty percent contributes as much as the covered eighty percent because it remains hidden for four times as long. Using only λDU TP/2 would erase that longer exposure.
Find the floor before increasing test frequency
Halve the partial interval to 2,190 h while leaving coverage and the full-test interval unchanged. The covered contribution halves to 0.001752; the uncovered contribution remains 0.003504. Total PFDavg becomes 0.005256, a 25% reduction from 0.007008 rather than 50%. Even an arbitrarily short partial interval cannot remove the residual 0.003504 in this first-order model.
That floor directs attention to failure modes the partial test misses, the full-test opportunity and whether the function can be redesigned or otherwise justified. It does not automatically mean more testing is better: testing can require bypasses, introduce errors or expose personnel. The full comparison must include those consequences and the actual restoration process.
Add bypass exposure with compatible demand weighting
For a separate simplified sensitivity, assume the function is wholly unavailable for 8 h in an 8,760 h year and demands are uniform in time. The bypass fraction is u=8/8,760≈0.000913. If q=0.007008 describes average hidden unavailability outside bypass, total demand unavailability is u+(1−u)q≈0.007915, not merely q.
This combination assumes q is conditional on not being bypassed and that demand exposure is compatible with the time fraction. If transfers are prohibited and actually prevented during testing, or if testing itself creates demands, a different weighting is needed. A bypass record is useful only when its dates, extent and operational consequences match the analysis.
Turn coverage into a failure-mode evidence map
For each failure mode, identify what the test stimulates, what it observes and what constitutes a pass. A simulated electronic input may test logic and output but leave the physical sensor connection unchallenged. Partial valve travel may demonstrate movement without demonstrating full closure or the required leakage performance. A displayed position is not automatically evidence of the complete physical barrier.
The HSE proof-test requirements excerpt calls for identifying which modes are and are not tested at each interval. Its value here is that explicit coverage question. A manufacturer's percentage should be tied to the installed configuration and the procedure actually performed before it is used as C.
Do not infer complete integrity from one PFD contribution
Random hardware contribution is only part of a protective function's performance. Systematic design errors, common causes, environmental limits, diagnostics, repair and human restoration can matter. A low result from λDU T/2 does not by itself establish a SIL or qualify a layer for every LOPA scenario. Suitability includes the specified action and response time.
The formula also assumes a low-demand interpretation; frequent or continuous demands may require a different reliability measure and model. The historical HSE appendix cites older standard editions, so its simplified example is used for the mechanism rather than as a current compliance rule. Applicable editions and vessel requirements must be checked in the actual engineering assessment.
Ask whether the full test really resets every modeled mode
The two-interval model assumes the less frequent full test reveals the entire remaining dangerous-undetected population. If a residual mode survives that test too, its clock does not reset at TF. It needs a justified replacement horizon, another effective test or a separate lifetime contribution. Labeling a procedure full cannot make this residual disappear.
This matters when a test simulates the sensor signal but never challenges the process connection, or moves a final element without testing the required sealing function. The overlooked mode can accumulate across several nominally completed test cycles. Repeated successful records then demonstrate repetition of the same limited challenge, not discovery of every relevant failure. State the reset assumption separately for each population, and revise the model when the implemented procedure cannot support it.
Use discoveries to question the assumed failure model
A proof-test finding should identify the failed function, failure mode, likely duration, diagnostic history and whether the specified test actually found it. Counting failed tests without considering exposure and coverage does not directly estimate λDU. Several discoveries can arise from one systematic cause, and one failed component can have remained unavailable through many demands.
Conversely, a long record with no discoveries may reflect good performance, low exposure, incomplete reporting or a test that does not challenge the important mode. Compare findings with the model's predictions and examine mismatches before adjusting the rate mechanically. Retain negative and positive observations with their test conditions. The evidence becomes useful when it can explain which modeled population was observed and which remained outside the test's reach.
Keep the test claim auditable through restoration
Record the failure-mode partition, rate basis, partial and full intervals, test steps, pass criteria, discovered faults, repair time and restoration evidence. A completed form without the actual observations cannot confirm coverage. Changes in sensor installation, valve duty or test method can invalidate the calculation even when the interval remains unchanged.
The worked model shows why incomplete coverage creates a residual contribution and why demand timing matters. It provides no instruction to extend an installed test interval or perform a hazardous test. Those decisions require the approved safety requirements, equipment documentation and a competent assessment of the whole protective function.
Sources
- OG54 Appendix 3: Partial Proof Testing Example · UK Health and Safety Executive · Source check date: 2026-10-06
- OG54 Appendix 1: Process for Defining SIS Proof Testing Requirements · UK Health and Safety Executive · Source check date: 2026-10-06