How the marine example becomes a learning problem
A transparent data generator
For each of 240 invented rows, independent uniform draws select Q, m and Tᵢ within the displayed ranges. The target is Tₒ = Tᵢ + Q/(4.18m) + ε, with ε uniform from −0.3 to +0.3 °C. This arbitrary additive noise is a teaching choice, not a calibrated sensor uncertainty. There is no time-series structure, vessel identifier, fouling history or operating restriction.
At Q = 418 kW, m = 20 kg/s and Tᵢ = 20 °C, the noiseless rise is 418/(4.18 × 20) = 5 °C and the outlet is 25 °C. Greater flow reduces the temperature rise at fixed duty. The quotient creates a smooth nonlinear regression task. It does not model cooler effectiveness, control valves or the available heat-transfer area.
Data seed 20261008 creates the fixed rows. A separate Fisher–Yates shuffle with seed 1701 assigns 144 training, 48 validation and 48 test rows before scaling. The weight seed changes only initial weights. All random draws use the documented Mulberry32 generator. The software version and seeds travel with the CSV.
Split first, scale on training only
For each input and the output, calculate the training mean μ and population standard deviation σ. Transform z = (x − μ)/σ. Reuse those exact training values for validation, test and new cases; a constant training column uses σ = 1. Inverse-transform the output with T̂ = σᵧŷ + μᵧ. Scaling all rows first would leak holdout information.
Only the 144 training rows determine gradient updates. Validation loss is monitored but never differentiated, and this implementation does not pick the best validation epoch: it keeps the last of the requested epochs. You may use validation to explore settings. Open the test report only when finished choosing. It is evaluated after training and never used for scaling or updates. Repeatedly choosing settings after seeing test scores turns that set into another validation set.
A random split is suitable for this independent synthetic generator only. Real ship signals often have temporal correlation and common vessel conditions. They may need separation by time, voyage or vessel plus independent prospective checks. A random row split cannot establish that a deployed model will generalize.
Forward pass, loss and backward pass
For standardized inputs zᵢ, hidden unit j computes sⱼ = Σᵢ wⱼᵢzᵢ + bⱼ and aⱼ = tanh(sⱼ). The linear output is ŷ = Σⱼ vⱼaⱼ + c. There are 5H + 1 trainable parameters: 3H input weights, H hidden biases, H output weights and one output bias. Hidden weights use a seeded uniform range ±√(6/(3+H)); output weights use ±√(6/(H+1)); biases start at zero.
The objective is L = (1/N)Σₙ(ŷₙ − yₙ)². For one row its output derivative is dₙ = 2(ŷₙ − yₙ)/N. Accumulate ∂L/∂vⱼ = Σₙdₙaₙⱼ and ∂L/∂c = Σₙdₙ. Propagate δₙⱼ = dₙvⱼ(1 − aₙⱼ²); then ∂L/∂wⱼᵢ = Σₙδₙⱼzₙᵢ and ∂L/∂bⱼ = Σₙδₙⱼ. All derivatives use the same pre-update weights.
One full-batch epoch updates every parameter simultaneously: θ ← θ − η∂L/∂θ. This lesson uses no momentum, mini-batches, regularization, early stopping or automatic search. A larger learning rate can worsen learning; more neurons or epochs do not guarantee better holdout performance. Training yields to the interface after at most four epochs, so cancellation can interrupt long runs.
Worked one-neuron pass and gradient
Use an illustrative standardized row z = (1, 0, −1), target y = 0.2 and H = 1. Set w = (0.2, −0.1, 0.3), b = 0, v = 0.5 and c = 0.1. These hand-chosen weights explain one step; they are not the default seeded initialization.
s = 0.2 × 1 − 0.1 × 0 + 0.3 × (−1) = −0.1. Thus a = tanh(−0.1) = −0.099667995, ŷ = 0.5a + 0.1 = 0.050166003 and L = (ŷ − 0.2)² = 0.022450227.
With N = 1, d = −0.299667995. The output-weight gradient is da = 0.029867308; the hidden delta is d × 0.5 × (1 − a²) = −0.148345590. Therefore ∂L/∂w₁ = −0.148345590, ∂L/∂w₂ = 0 and ∂L/∂w₃ = +0.148345590. At η = 0.1, w₁ becomes 0.214834559, v becomes 0.497013269 and c becomes 0.129966799. The automated tests independently check every gradient against central finite differences.
Read the result without overclaiming
RMSE is the square root of mean squared error in °C and emphasizes large errors. MAE averages absolute error in °C. R² = 1 − SSE/SST uses the test mean in SST and is not a probability or an accuracy percentage. If test targets have no variance, R² is undefined. The simple training-mean baseline gives the neural network a concrete comparator.
The preset experiment uses H = 6, 350 epochs, η = 0.03 and seed = 42. Its final training MSE is approximately 0.024146 and validation MSE is 0.031164 in standardized units. The held-out test values appear only after you finalize the run. These are lesson results, not real-ship measurements. Arithmetic may differ slightly across JavaScript runtimes.
A useful operational study would need traceable measurements, sensor quality controls, missing-data rules, a realistic split, credible baselines, uncertainty assessment, out-of-distribution checks and an independent validation plan. No alarm threshold, engine limit, failure probability or safety recommendation follows from this lesson.
Method sources
Primary backpropagation paper and official ML documentation checked 8 October 2026. The cooling generator and numerical example are original educational constructions, not source datasets. v1.0.0.