Knowledge / Risk and reliability
Risk-informed maintenance: from failure mechanism to work order
Connect maintenance to failure mechanisms, hidden unavailability, test coverage, work-order evidence and verified restoration.
On this page
A maintenance programme is risk-informed when its tasks address the ways important functions can be lost, and when the evidence from those tasks changes subsequent decisions. A criticality score alone does not achieve this. The programme needs a connection between function, failure mechanism, consequence, task effectiveness and verified restoration. This educational guide uses original hypothetical shipboard examples. It does not authorize extending a survey interval, changing a manufacturer’s requirement or operating with a known impairment.
Begin with the service that must remain available
An equipment list is a useful inventory, but maintenance decisions begin with required functions. A pump may need to start on demand, deliver a stated flow, remain available during a degraded power condition and operate for a specified duration. Different failures challenge these requirements differently. A pump that rotates during a short test may still be unable to deliver the required flow through the actual system.
State the operating context, duty and consequence of functional loss. Two identical components can deserve different treatment if one serves a non-essential convenience and the other supports an essential function. Conversely, a small inexpensive component can be highly consequential if it is a common dependency for several systems. Do not equate purchase price with criticality. Include sensors, final elements, support services and the tasks needed to restore the function after maintenance.
Keep legal and approved requirements in the decision
The ISM Code’s maintenance provisions link conformity, inspections, corrective action, records and reliability of equipment whose sudden failure may create hazardous situations. A publicly available ClassNK-hosted ISM Code text, section 10 also addresses regular testing of standby arrangements. The applicable current flag, class and company requirements must still be checked for the vessel. An educational risk calculation is not authority to waive them.
Separate mandatory work from discretionary optimization before comparing options. A condition-monitoring proposal may supplement a required inspection without replacing it. If an alternative programme is permissible, its approval route and evidence requirements need to be established explicitly. Record the basis for each interval and task, including the document edition and applicable equipment. Otherwise a numerical optimization can quietly override an obligation that was never included in its objective function.
Match the task to the failure mechanism
Some deterioration gives a detectable warning before functional failure; other failures remain hidden until demand. Some mechanisms have a useful age relationship; others are dominated by random events or maintenance-induced errors. These distinctions affect whether condition monitoring, scheduled restoration, replacement, failure-finding or redesign is appropriate. “Preventive maintenance” is too broad a label to explain why a particular task should work.
NASA’s 2008 reliability-centred maintenance guide provides public engineering background on selecting maintenance around functions and failure behaviour. It concerns facilities and collateral equipment, not marine statutory intervals. The transferable reasoning is to justify task effectiveness. Measuring vibration will not detect every failure of a standby control circuit; replacing a mechanical seal at a fixed age does not resolve an incorrect operating alignment.
Separate criticality from immediate priority
Criticality describes the significance of losing a function under stated conditions. Immediate priority also depends on current condition, remaining protection, exposure and time available for intervention. A highly critical component with strong verified condition evidence may not be the most urgent work order today. A lower-ranked item can become urgent if its failure would coincide with another system already unavailable. The current configuration matters.
Use a risk matrix carefully. Ordinal scores are categories, not measured probabilities, and multiplying severity, occurrence and detectability numbers does not create a physical risk quantity. Keep the reasoning behind the ranking visible. A practical prioritization note should identify the credible failure, present evidence, consequence, affected barriers and deadline for decision. This supports a defensible work sequence rather than an unexplained list sorted by one composite number.
Understand a hidden failure and its test interval
Consider a hypothetical protective function with a constant dangerous undetected failure rate λDU = 2 × 10⁻⁶ per hour. Assume a perfect periodic test, immediate restoration, no diagnostic coverage, negligible test downtime and sufficiently rare independent demands. For small λDU × T, its average probability of failure on demand is approximately λDU × T/2. At T = 4,380 hours, the result is 0.00438; at 8,760 hours it is 0.00876.
These invented values illustrate why extending a test interval can increase hidden unavailability. The approximation is not a complete safety-function verification and does not establish acceptable intervals. It excludes common cause, imperfect coverage, repair time, systematic failure and failure introduced by testing. The model’s assumptions must be checked before using its result. HSE’s proof-testing requirements guidance specifically highlights the need to account for failures that a partial test cannot reveal.
Test coverage matters as much as the calendar
A functional test should demonstrate the relevant function and reveal the failure modes credited in the calculation. A lamp test is not an end-to-end trip demonstration. A partial valve movement may reveal some mechanical problems while leaving other failure modes untested. Record which parts of the chain are exercised, under what conditions and with what pass/fail criteria. The word “tested” is too vague for a reliability claim.
More frequent intrusive testing is not automatically safer. Testing may temporarily remove protection, expose personnel to hazards or introduce restoration errors. The optimum cannot be inferred from the simple λDU × T/2 relationship alone. Consider test risk, coverage, maintainability and the ability to verify the final configuration. Any change to the approved test regime requires the appropriate technical and organizational authorization. The educational calculation is a prompt to ask better questions, not a substitute for that review.
Check whether condition monitoring leaves time to act
Suppose an original hypothetical assessment finds that a particular deterioration mechanism can progress from a detectable warning to functional failure in as little as 30 days under the relevant conditions. The proposed programme samples every 7 days, requires up to 2 days for analysis, 12 days for a replacement part and 3 days for intervention and verification. A simple worst-phase allowance is 7 + 2 + 12 + 3 = 24 days, leaving 6 days.
This is only a screening timeline. Detection capability, progression variability, shipping delays and operational access can consume the margin. A mean warning interval is not a safe lower bound. If the condition signal does not reliably identify the mechanism, a short sampling interval will not repair that weakness. The output should identify the evidence needed for detection and response, then determine whether monitoring supports a useful decision before failure. No universal sampling ratio is prescribed.
Convert the analysis into an executable work order
A work order should identify the equipment and required function, the approved task reference, precautions and authorization, required competence, measurements, acceptance criteria and restoration verification. It should distinguish “inspect and record condition” from “adjust to a specified approved value.” Avoid instructions that invite technicians to improvise acceptance limits from historical habit. Where a measurement is required, include location, operating condition, instrument requirements and how an out-of-limit result is handled.
The work order should also preserve evidence useful for future analysis. “Completed, satisfactory” can hide the actual condition trend. Record relevant as-found and as-left values, defect mechanism, corrective action and whether the claimed functional test succeeded. Link unresolved defects to a decision owner rather than closing them administratively. The next analyst should be able to tell whether the task prevented deterioration, discovered a hidden failure or merely confirmed an already known condition.
Manage the risk created by maintenance itself
Isolation, dismantling, configuration changes and restoration alter the system’s protection. If two redundant trains are unavailable together, a maintenance plan can create a higher-risk state than either work order suggests individually. Review simultaneous tasks, shared utilities and the operating exposure during the maintenance window. This requires the vessel’s authorized safe-work and planning arrangements, not an informal trade-off made by a scheduling algorithm.
Restoration deserves explicit verification because the system may look complete while a bypass, drain or selector remains in the wrong state. Identify what confirms the final function and who receives the handover. HSE’s maintenance, inspection and testing overview emphasizes responsibilities, competent interval-setting, overdue work management and clear pass/fail criteria. These are general management principles; vessel-specific technical and statutory requirements remain the controlling basis.
Use backlog information without hiding exposure
A backlog count treats tasks as interchangeable when they are not. Five overdue low-consequence inspections may matter less than one overdue demonstration of a critical standby function. Record the function affected, the reason for delay, condition evidence, remaining protection and operating exposure. “Parts ordered” explains the delay but does not establish that the risk is acceptable until they arrive. An extension needs a decision with a basis and a review point.
Also investigate recurring causes of delay: inaccessible equipment, unclear instructions, inadequate spares, poor scheduling or a test that cannot be performed as written. Repeated deferral can reveal a programme that is not executable. Do not normalize overdue safety work by repeatedly changing the due date without technical review. A useful dashboard distinguishes completed work, demonstrated function, unresolved defects and temporary restrictions, rather than treating administrative closure as the only measure of success.
Update the programme with mechanism-specific evidence
Track failures and discoveries against relevant exposure: running hours, starts, demands, calendar ageing or environmental cycles. A rate per operating hour cannot automatically represent a failure that occurs mainly on start. Record changes in equipment, duty, parts and maintenance method so that trends are interpreted consistently. More detected defects after an improved inspection may indicate better detection rather than worse reliability; examine the mechanism and observation process before changing intervals.
A useful final package links each important function to failure modes, applicable obligations, chosen tasks, interval basis, acceptance criteria and feedback triggers. When evidence shows the task is ineffective, reconsider the strategy rather than merely increase its frequency. Risk-informed maintenance is a learning loop with controlled decisions. Its success is demonstrated by preserved function and credible evidence, not by maximizing the number of completed work orders or minimizing maintenance expenditure in isolation.
Sources
- International Safety Management Code, section 10, public ClassNK copy · IMO text hosted by ClassNK · Source check date: 2026-10-06
- NASA Reliability-Centered Maintenance Guide for Facilities and Collateral Equipment, September 2008 · National Aeronautics and Space Administration · Source check date: 2026-10-06
- Appendix 1: Process for Defining SIS Proof Testing Requirements · UK Health and Safety Executive · Source check date: 2026-10-06
- Maintenance, Inspection and Testing: Overview · UK Health and Safety Executive · Source check date: 2026-10-06