Standby equipment: hidden failures and functional-test coverage

Assess the complete standby response, from detecting a demand to sustaining the required service, and recognize what a start test leaves unproved.

On this page

Standby equipment can fail while appearing uneventful because normal operation does not call on its function. A stopped emergency pump or reserve generator may have no obvious symptom until demanded. Functional testing makes selected hidden failures visible. Its usefulness depends on the demand path, achieved duty, duration, supporting services and restoration after the test, not merely on whether a motor turned.

Define the demand and the required response

Write the function as a performance statement: what triggers the standby unit, what service it must provide, how quickly and for how long, and under which conditions. “Starts successfully” is narrower than “restores the required cooling flow following loss of the duty pump.” The latter also includes the detection logic, valve path, supply availability and actual hydraulic performance.

The public ISM Code text hosted by ClassNK, section 10.3, connects reliability measures with regular testing of standby arrangements and equipment not continuously used. The passage supports attention to hidden readiness; it does not prescribe one test interval or one test method for all equipment. The applicable vessel requirements, approved instructions and safety-management arrangements determine those details.

Trace the complete chain from condition to service

A low-pressure sensor may initiate a standby start through logic and a contactor, after which the motor drives the pump and flow passes through a non-return valve. A local manual start may bypass the sensor and logic. A motor run with a closed delivery path may not demonstrate flow. List the elements exercised by each test and the ones left outside it.

Include supporting energy and consumables: control power, starter batteries, starting air, fuel, lubrication, cooling and ventilation where applicable. A starting battery that passes a no-load voltage check may still fail under demand. A fuel tank with adequate indicated level may have an unavailable supply path. Each test should address the relevant failure mechanism and the real support condition rather than infer readiness from a single convenient indication.

Distinguish start, transfer, capacity and endurance

Starting demonstrates only part of readiness. A standby generator must also achieve suitable output, connect through the intended path where required, accept the relevant load and sustain operation. A pump must deliver the required flow and head under a meaningful system condition. Test duration matters when a cooling restriction or limited reserve causes delayed failure.

An original hypothetical pump test requires at least 45 m³/h at the specified system condition. The pump starts in 8 s but delivers only 36 m³/h, or 80% of the requirement. The start criterion may pass while the service criterion fails. Recording “pass” for the whole function would erase the important finding. The numbers are illustrative and must not be used as acceptance values for another pump.

A local-start test enters downstream of the automatic demand sensor and logic, so those are outside its coverage. The motor/pump starts in 8 seconds but delivers 36 cubic metres per hour against 45 required.80% capacity fails the stipulated service criterion even though starting occurs.
Original functional coverage diagram for a hypothetical test. Dashed path denotes unexercised automatic initiation, not known failed hardware. Solid path denotes the local-start and measured-delivery scope. The 8 s observation is not asserted to pass an unspecified response-time requirement; the explicit 45 m³/h criterion is not met. Figures are invented and give no live-test procedure, accepted interval, SIL or vessel-specific acceptance value.

Map coverage to failure modes

Coverage is the share of relevant hidden failure behaviour a test can reveal under its conditions. It is not the fraction of checklist boxes ticked. Ten easy checks of the starting circuit do not compensate for an untested delivery valve. A partial test can be valuable if its scope is explicit and a complementary method covers the remaining important modes.

HSE control-system guidance discusses end-to-end testing and the need to reveal failures under relevant demand conditions. Its process-safety context is used here for the testing principle, not to assign a safety-integrity level to marine standby equipment. Where a full demand test cannot safely be performed, document what the alternative proves and what uncertainty remains.

Understand why hidden exposure grows between tests

For an idealized dormant item with constant hidden-failure rate λ, perfect periodic tests every T hours and immediate repair, the average hidden unavailability is 1 − [1 − exp(−λT)]/(λT). When λT is small, this is approximately λT/2. The approximation describes time spent unknowingly failed; it is not the probability of an accident.

With invented λ = 0.00002 per hour and T = 720 h, the approximation is 0.0072, or 0.72%; the exact value is about 0.7166%. At T = 168 h, the approximation is 0.168%. This demonstrates interval sensitivity only. Imperfect test coverage, repair delay, test-induced faults, demand dependence and shared failures all violate the simplified model and require additional treatment. No universal interval follows from the example.

Include the risk created by testing

A test may temporarily remove protection, disturb a valve lineup or introduce wear and human error. Plan the test state, compensating arrangements, communication and abort conditions. Avoid testing redundant units simultaneously when this would remove the required service. A nominally harmless simulation can still activate downstream equipment if boundaries are misunderstood.

The test procedure should identify who controls the system and what indicates a safe return to normal. Test connections, forced signals and temporary overrides require tracking. Frequent testing is not automatically better if it repeatedly creates unmanaged impairment. The interval decision balances evidence of hidden failure with the actual test method and applicable requirements, rather than simply minimizing T in a formula.

Prove restoration to the intended standby state

A successful run can leave the selector in manual, a valve in test position or an automatic trigger inhibited. The machine may be mechanically healthy yet unavailable for the next real demand. Restoration is therefore part of the test, with independent confirmation where the consequence and procedure require it. Verify the actual normal configuration rather than relying on memory.

Record failures found, corrective work, retest outcome and the final state. A failed test followed by repair should remain in the history as a discovered failure, not be overwritten by the final pass. Otherwise the evidence used to choose intervals becomes falsely reassuring. When a test reveals a common defect, assess other units sharing the same design, maintenance history or supply dependency.

Do not infer certainty from repeated successful starts

Zero failures in a short test history does not establish a zero probability of failure on demand. For an original statistical illustration, 20 independent and identically distributed demands with zero failures give a one-sided 95% binomial upper bound of 1 − 0.05^(1/20), about 13.9%. This is a statement about limited evidence under a model, not a claim that the equipment actually has that failure probability.

Repeated starts under the same easy conditions may not be independent or representative of real demand. They can all omit a cold-start problem, transfer path or supporting utility. Preserve demand conditions and test coverage before using pass counts statistically. A large number of narrow tests cannot demonstrate an untested function; statistical confidence and functional completeness are separate questions.

Report readiness with its demonstrated limits

Useful test records contain the initiating condition, configuration, measured response time, achieved output, duration, instruments and acceptance criteria. Separate observed performance from assumptions about untested conditions. A test under calm harbour conditions may leave questions about a different environmental or load state; that limitation should guide the approved testing strategy.

The final readiness statement should explain which required function was demonstrated and which unresolved defect or limitation remains. Avoid converting the absence of recent demands into evidence of reliability. Standby readiness is an actively maintained condition, supported by relevant tests and correct restoration. The purpose is dependable service when needed, not an accumulating count of successful no-load starts.

Sources