Evaluating collision-avoidance simulations: scenarios, metrics and validation

A practical framework for evaluating collision-avoidance simulations through scenario coverage, independent metrics, numerical checks and bounded statistical claims.

On this page

A collision-avoidance simulation can produce an impressive animation while leaving its most important claim untested. A successful-looking encounter may depend on perfect observations, instant steering or a cooperative target that the evaluator never questioned. Meaningful assessment asks what was tested, what could have failed, how outcomes were measured and which evidence connects the simulated system to its intended use. This article develops a public, non-proprietary evaluation framework with original numerical examples. It does not present an avoidance algorithm, certify a product or authorize autonomous operation.

1. Start with a claim narrow enough to test

An assertion that a system is safe is too broad to evaluate with a few scenarios. A useful claim specifies the function, operating conditions, available inputs and expected outcome. For example, a laboratory claim might concern avoiding geometric overlap in a defined set of two-vessel simulations with a particular response model and observation-error range. That is narrower and more testable than claiming general navigational safety.

List exclusions alongside the claim: shallow water, sensor outages, dense traffic, restricted visibility or other vessel behaviours may be outside the evidence. An exclusion is not automatically acceptable for an eventual operation, but it prevents a research result from silently expanding beyond what was tested. Define the acceptance logic before inspecting the results, and identify which failures require investigation even if an overall average looks good.

2. Separate verification from validation

Verification asks whether the implementation follows its specified equations and logic. Validation asks whether the representation is adequate for the intended real-world purpose. A simulator can integrate the chosen equations correctly while those equations omit a response delay that matters to the encounter. Conversely, a sound physical model can be implemented with a coordinate or unit error.

The ITTC 2024 manoeuvring-model validation procedure treats model documentation, comparison evidence and uncertainty as linked tasks. Its scope is manoeuvring models, not certification of a complete collision-avoidance system. Use that distinction to build a layered argument: physical response, sensing, decision logic, execution and outcome evaluation each need relevant evidence. Success in one layer should not be presented as proof for the layers it did not exercise.

3. Draw the system boundary explicitly

A complete simulation may contain a world model, ship dynamics, sensors, a tracking function, decision logic, an actuator model and an independent evaluator. State which of these are present and which are idealized. If the decision function receives exact target positions and future intentions, the test concerns a different information problem from one using delayed, uncertain observations.

The publicly available Scope excerpt from Sawada, Sato and Minami’s scenario-based safety-evaluation paper explicitly distinguishes its algorithm evaluation from sensor-detection performance. Only the publisher’s abstract and public scope snippets were accessible for this review; no full-paper method is reproduced here. The general lesson is to describe the tested boundary faithfully. A planner-only result should not acquire an end-to-end label merely because the demonstration includes a realistic-looking bridge display.

4. Build a scenario matrix with meaningful variation

Vary encounter geometry, relative speeds, initial range, available water, vessel response and information quality. Include ordinary cases, boundaries where a classification changes, and degraded conditions relevant to the claim. Geometric labels such as crossing or head-on do not by themselves establish which legal provisions apply; visibility and vessel status must be specified separately.

For an invented coverage exercise, three geometric families, two visibility contexts, two vessel-response conditions and three observation conditions produce 36 combinations. Running 25 stochastic realizations of each gives 900 runs. This arithmetic measures planned coverage, not completeness. The matrix may still omit another ship, a sensor bias or a relevant response interaction. Explain why each dimension matters and keep a list of known gaps rather than treating a large total as evidence that every important situation has been included.

5. Keep development and evaluation evidence separate

If scenarios are repeatedly used to adjust thresholds or tune a model, they become development evidence. Evaluating the final system on the same scenarios can overstate generalization. Preserve a separate evaluation set where feasible, and distinguish a previously unseen geometry from a familiar geometry with only a new random seed.

Data separation also matters for physical response models. A model fitted to one measured turn should not cite that same fitted curve as independent validation. Keep the provenance of parameters, training examples, calibration data and evaluation cases visible. When a failure is fixed, retain it as a regression case, but report that it is now known to the developers. A collection of remembered failures is useful without being a substitute for genuinely independent tests.

6. Measure outcomes beyond whether hulls overlap

Collision is an important outcome, but its absence alone does not describe the quality of a manoeuvre. Relevant outputs can include minimum hull clearance, time spent inside a declared screening domain, response delay, control saturation, route deviation and unresolved hazards at the end of the run. Each metric needs a definition, reference frame, unit and evaluation interval.

Do not silently mix predicted clearance used by the decision function with actual clearance in the simulated world. Their difference may be the central issue in a sensor-degradation test. Likewise, distinguish a soft performance preference from a hard failure criterion. A shorter route cannot compensate for a collision by improving a weighted average. Report safety-relevant failures individually and preserve the sequence that produced them; a single composite score can conceal incompatible outcomes.

7. Evaluate rules with context and explicit uncertainty

The International rules in the USCG compilation place collision-risk assessment and action within the circumstances and applicable encounter provisions. A simulator’s distance threshold or binary rule label is an interpretation for the test, not a substitute for the full rules. International and US Inland provisions must be kept distinct.

For each case, record the assumed visibility, vessel status, observations available at each time and the rationale for the expected behaviour. Some judgments, such as whether an action is sufficiently early or readily apparent, need a declared assessment method and competent review. Where reviewers disagree, retain the disagreement and its reason. Do not relabel every geometrically successful trajectory as compliant, or treat an uncertain classification as settled merely because software produces a single answer.

8. Verify geometry and numerical resolution

A discrete simulation can miss an event that occurs between samples. Imagine a point moving along a straight line from x equal to minus 15 metres to plus 15 metres in two seconds, with a fixed circular exclusion region of radius 5 metres centred at zero. Both sampled endpoints are outside, yet the path passes through the region. This is an original geometry example, not a ship model.

Checking only logged positions would falsely report no intrusion. Suitable event detection, swept-geometry checks or a justified temporal resolution are needed for the claimed metric. Reducing the step size and comparing outcomes can reveal sensitivity, but a stable result alone does not validate the physics. Test simple cases with independently known answers: straight motion, constant turn, stationary obstacles and coordinate rotations. Keep the independent evaluator from sharing every assumption and defect with the decision function.

9. Include delays and response limits in the evidence

A command can be feasible mathematically but unavailable physically. Steering rate, propulsion response, control limits and processing delays influence how quickly a planned trajectory can be achieved. Define where a delay is applied and distinguish sensing age, computation time, communication time and actuator response. Avoid adding or removing the same delay at multiple interfaces by accident.

At a hypothetical relative speed of 12 knots, a 60-second interval represents 0.20 nautical miles, or 370.4 metres, of relative travel. That conversion does not tell the evaluator what a safe decision delay is. It shows why an apparently small missing delay can matter. Compare the requested action with actual simulated execution and report any saturation or lag. An animation that draws the planned path while the simulated ship follows a different one is particularly misleading.

10. Use stochastic testing without inventing certainty

Randomized trials can explore specified distributions of initial conditions, disturbances or sensor errors. Record those distributions, correlations and random seeds. Independent seeds do not make trials representative of real operations, and repeating a narrow scenario does not create broad coverage. Systematic bias or an omitted behaviour can remain absent from every randomized run.

Suppose, purely for statistical illustration, 900 independent and identically distributed Bernoulli trials show zero failures. Solving the zero-failure binomial expression gives a one-sided 95% upper confidence bound of 1 minus 0.05 to the power 1/900, approximately 0.3323%. NIST’s exact-binomial documentation provides the statistical context. This is a bound for the assumed trial population, not a real-world collision probability. A deliberately stratified 36-cell matrix generally does not justify one common failure rate without additional assumptions.

11. Compare systems on the same information and constraints

A fair comparison gives alternatives the same scenario inputs, observation quality, physical response limits and evaluation rules. If one system sees perfect future trajectories and another sees delayed tracks, an outcome difference cannot be attributed solely to decision quality. Likewise, a computational budget or fallback policy can materially affect results and must be reported.

Report paired differences where cases are shared, including cases in which both systems fail. An average improvement can coexist with worse performance in a rare but important condition. Show results by relevant scenario group and examine the tail, not only the mean. If test cases were selected to challenge a known weakness, describe that purpose. Such stress testing can be valuable without representing the frequency of those cases in everyday traffic.

12. Investigate failures and retain reproducible evidence

For each important failure, preserve scenario configuration, software and model versions, seeds, observations, decisions, executed controls and outcome traces. Establish whether the issue began with unavailable information, an incorrect interpretation, an infeasible action or evaluation logic. The last visible symptom is not necessarily the initiating cause.

After a correction, rerun the failed case and relevant neighbouring cases. Then repeat the broader regression set to look for unintended effects. Record what changed and which claim the new evidence supports. Reproducibility does not require publishing private code or proprietary design details in an educational article. It requires enough controlled evidence within the responsible review process to distinguish a real improvement from a changed test or a favourable new seed.

13. Report the conclusion at the strength of the evidence

A useful conclusion states the tested function and conditions, the coverage achieved, observed failures, metric definitions, statistical assumptions and unresolved limitations. It distinguishes simulation findings from hardware tests, controlled vessel trials and operational evidence. The progression between those levels requires an appropriate safety and approval framework; simulation success alone does not authorize the next stage.

The examples, counts and calculations are original hypotheticals. They do not describe or establish the performance of a particular product. Public sources support only the stated context and scope. The aim is a transparent evaluation argument that can be challenged and improved, with no claim of publication, certification or demonstrated autonomous-navigation safety.

Sources