Calculations with context
Simple linear regression
Fit an unweighted OLS line with an intercept to two to thirty paired observations; inspect slope, residuals and error measures in their respective units.
STATISTICS · SENSOR CALIBRATION
Simple linear regression
Fit one straight line to paired measurements, then inspect what the line misses. This is an unweighted ordinary least-squares (OLS) teaching calculator with an intercept.
1. The question: how does a reference measurement relate to a sensor signal?
Imagine checking a pressure sensor from a vessel’s auxiliary system on a bench. Let x be trusted reference pressure (bar) and y be the electrical sensor output (V). The small data set below is entirely synthetic. It is not a calibration certificate, real vessel data or a safety limit. Each row is one paired measurement recorded in consistent units.
The aim is to describe y ≈ b₀ + b₁x. Least squares minimizes the sum of squared vertical residuals. Exchanging x and y solves a different optimization problem: algebraically inverting this line is not the same as regressing x on y.
2. Equations, symbols and units
ŷᵢ = b₀ + b₁xᵢ; eᵢ = yᵢ − ŷᵢ; SSE = Σ eᵢ² → minimum
x̄ = Σxᵢ/n; ȳ = Σyᵢ/n; Sxx = Σ(xᵢ − x̄)²; Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ)
b₁ = Sxy/Sxx; b₀ = ȳ − b₁x̄; SST = Σ(yᵢ − ȳ)²; R² = 1 − SSE/SST
RMSE = √(SSE/n); s = √[SSE/(n − 2)] (only when n > 2)
| Symbol | Meaning | Unit |
|---|---|---|
| n, i | Number of paired observations and row index | Unitless |
| xᵢ, yᵢ | Reference and measured output | x and y units |
| x̄, ȳ | Arithmetic means | x and y units |
| b₀; b₁ | Intercept at x = 0; output change per x unit | y; y/x |
| ŷᵢ; eᵢ | Fitted value; observed minus fitted | y |
| Sxx; Sxy | Centered sums of squares and cross-products | x²; x·y |
| SSE; SST | Residual sum of squares; sum of squared y deviations from its mean | y² |
| RMSE; s; R² | In-sample root mean square error; residual standard deviation; explained variation fraction | y; y; unitless |
3. A synthetic example you can follow by hand
For x = [0, 1, 2, 3, 4] bar and y = [1, 3, 4, 7, 10] V, n = 5, x̄ = 2 bar and ȳ = 5 V. Sxx = 4 + 1 + 0 + 1 + 4 = 10 bar²; Sxy = 8 + 2 + 0 + 2 + 10 = 22 bar·V.
b₁ = 22/10 = 2.2 V/bar; b₀ = 5 − 2.2 × 2 = 0.6 V; ŷ = 0.6 + 2.2x
The fitted values are [0.6, 2.8, 5, 7.2, 9.4] V and the residuals are [0.4, 0.2, −1, −0.2, 0.6] V. SSE = 0.16 + 0.04 + 1 + 0.04 + 0.36 = 1.6 V²; SST = 16 + 4 + 1 + 4 + 25 = 50 V².
RMSE = √(1.6/5) ≈ 0.565685 V; s = √(1.6/3) ≈ 0.730297 V; R² = 1 − 1.6/50 = 0.968
“Load exact worked example” restores precisely these five pairs and the bar/V labels; then click “Fit the line.” b₀ is the model’s value at x = 0. If zero is outside your measured range, do not interpret the intercept as a physical zero error.
4. Read residuals, not just R²
A positive residual puts the measured y above the line. Systematic curvature can suggest an inadequate straight-line model; widening scatter can suggest changing error scale; an isolated x value can have high influence. Measurements taken in time order may also have dependent errors. These plots do not establish a diagnosis: they are a starting point for checking. Investigate measurement conditions rather than deleting a suspected outlier without justification.
R² summarizes how much of the observed y variation is represented by this line in this data set. It does not establish causation, validate a physical mechanism or guarantee accuracy for new measurements. RMSE describes the fitted data’s error; s uses n − 2 to account for the two fitted parameters. They are different quantities. Constant y gives SST = 0, so R² is undefined. Constant x makes a unique slope unidentifiable.
5. Assumptions, numerical limits and responsible use
Every point has equal weight, and x is treated as a fixed, error-free explanatory variable. Material x measurement error, unequal sensor uncertainties, saturation or nonlinear behavior can require another method. Interpreting the residual standard deviation as a common error scale needs an adequate model, independent errors and roughly constant variance. No normal-distribution inference is made here: there are no p-values, confidence intervals or prediction intervals.
Do not claim reliable prediction outside the observed x range. Real calibration also requires a traceable reference, measurement uncertainty, repeated observations, temperature effects and the relevant technical procedures. A near-zero sum of residuals follows from fitting an intercept; it is not independent validation.
The implementation first centers both axes and scales them by their ranges, then uses compensated accumulation. This avoids the raw subtraction Σx² − (Σx)²/n at large offsets. Residuals are evaluated in centered coordinates; subtracting the rounded on-screen y and ŷ may differ in the last digits. Numerical limits are not physical acceptance criteria.
Sources and scope
NIST/SEMATECH: least-squares model fitting; NIST/SEMATECH: linear least-squares regression. This lesson is an original educational summary. Calculator version: ols-teaching-1.0.0.
Runs entirely in your browser. No upload, storage or network request. A local CSV is created only when you click Download, and includes every input, fitted value, assumption and calculator version.
Displayed values are rounded to 8 significant digits; the CSV retains JavaScript numeric precision. Centering and scaling protect against avoidable cancellation, but cannot recover precision already lost in entered numbers.
Enter paired values, then fit the line.
The calculation interface could not load. The example and explanation below remain readable; reload the page to recalculate.
Related context
The method explanation and worked example are on this page. The articles below provide additional context.
All calculators