Linear regression: fitting a line to data
Be able to draw a line that fits points by eye and use it to make simple predictions.
Prerequisites
- CLinear relationships in tablesrequired
Everyday explanation
Real measurements never lie exactly on a straight line.
y
| ●
| ●
| ● ●
| ●
| ●
|________________________ x
The points clearly follow a direction, but no straight line passes through all of them. The task is to find the line that fits best.
Do it by hand, with a ruler:
- Plot the points.
- Place the ruler so that roughly the same number of points fall above as below.
- Draw the line.
- Read off two points on the line (not on the data!) and calculate
kandm.
Then you can guess. If a new measurement is at x = 7, read what the line says about y there. That is a prediction — and in its simplest form, that is exactly what a machine learning model does.
Intuition
Calculate k and m from two points on the line:
Example: the line passes through (2; 14) and (6; 30).
So y = 4x + 6. Check with the other point: 4·6 + 6 = 30. ✓
How well does the line fit? Measure the residuals — the distance from each point to the line:
| x | Actual y | Line says | Residual |
|---|---|---|---|
| 1 | 11 | 10 | +1 |
| 2 | 13 | 14 | −1 |
| 3 | 19 | 18 | +1 |
| 4 | 21 | 22 | −1 |
Small residuals, half positive and half negative — a good line. If they are systematically positive at one end and negative at the other, the relationship is likely not straight.
Two warnings about predictions:
- Extrapolation is dangerous. The line is valid within the interval where you have data. An employee’s salary often increases linearly with experience in the first few years — but if you extend the line to 40 years, you get an unreasonably high salary that does not reflect reality.
- Correlation is not causation. The fact that sunglasses and ice cream sales follow the same line does not mean the glasses make people drink more ice cream. Both depend on the heat.
This is machine learning in miniature. The computer does the same thing, except it tries its way to the line that makes the squared residuals as small as possible — the least squares method.
Interactive
Read the line from a chart. Six measurements of how long a mobile phone lasts versus how much it is used:
Battery (%)
100 |●
| ●
80 | ●
| ●
60 | ●
| ●
40 |
|________________________
0 1 2 3 4 5 hours
The points: (0; 100), (1; 92), (2; 85), (3; 77), (4; 70), (5; 62).
Step 1 — the differences: −8, −7, −8, −7, −8. Almost constant → linear.
Step 2 — k: from (0; 100) to (5; 62): (62 − 100)/(5 − 0) = −7.6 percent per hour.
Step 3 — m: 100, because x = 0 is included.
The model: battery = 100 − 7.6·hours
Step 4 — use it:
- After 7 hours:
100 − 53.2 = 46.8 % - When is the battery dead?
0 = 100 − 7.6t→t ≈ 13.2hours
Step 5 — be sceptical. The second prediction lies far outside the measurement range (0–5 hours). In reality, the battery drains faster towards the end, and the phone shuts off at a few percent. The model is good for 0–6 hours and unreliable after that.
This is the most important habit to take away: a model is valid where it has data.
Mastery means
- Draws a line that fits the points
- Reads the slope (k) and intercept (m) from the line
- Uses the line to predict and identifies when it is not reliable
Sign in to do the exercises and build your mastery up.
Sources
- Matteboken (Mattecentrum) — free to read, non-profit association
- Khan Academy — matematik — CC BY-NC-SA 3.0
- scikit-learn User Guide (BSD-3) — BSD-3-Clause