Skip to content
AI-grafen
CBuilderDeep learning· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Turn the knobs: weights

Be able to adjust two knobs (k and m) to make the line fit, and see that training is turning knobs.

Prerequisites

Everyday explanation

Imagine an apparatus with two knobs:

   k (the slope)        m (the height)
      ◯                    ◯
   ↺     ↻              ↺     ↻

The knob k turns the line steeper or flatter. The knob m moves the whole line up or down.

Your task: turn them until the line fits the points as well as possible.

How do you know it has got better? You need an error number — a single number that says how bad it is right now.

The most common one is, for every point, to measure the distance up or down to the line, square it (so that minus becomes plus and large errors are punished extra), and take the mean.

error = the mean of (the actual y − the line's y)²

The smaller the error, the better the line. The goal is to find the knob settings that give the smallest error.

And that is what training is. A real model does not have two knobs but millions — but the principle is exactly the same.

Intuition

How do you know which way to turn?

You can try: turn k up a little, recompute the error. Did it get smaller? Carry on that way. Did it get larger? Turn the other way.

That works, but it is slow with two knobs and impossible with millions.

The clever method is to compute which way the error falls fastest, straight from the mathematics — that is called the gradient, and following it is called gradient descent.

The picture to have in your head:

  error
   |╲                    ╱
   | ╲                  ╱
   |  ╲___          ___╱
   |      ╲___  ___╱
   |          ╲╱          ← here the error is smallest
   |______________________ k

The error as a function of k is a bowl. Wherever you start there is a direction that goes downwards, and follow it long enough and you end up at the bottom.

The learning rate is how large the steps are:

Step lengthWhat happens
Too smallit takes for ever to get down
About rightit reaches the bottom in a reasonable time
Too largeit jumps back and forth across the bottom, or upwards

Those are the same three cases as in all training of neural networks — here with two knobs instead of millions.

Code

data = [(1, 11), (2, 13), (3, 19), (4, 21), (5, 27)]

def error(k, m):
    return sum((y - (k * x + m)) ** 2 for x, y in data) / len(data)

# 1. Trial and error — it works, but becomes unmanageable with more knobs
print("k=3 m=8 :", round(error(3, 8), 2))    # 3.20
print("k=4 m=6 :", round(error(4, 6), 2))    # 2.00
print("k=4 m=7 :", round(error(4, 7), 2))    # 1.80

# 2. Gradient descent — compute which way it slopes
def gradient(k, m):
    n = len(data)
    dk = sum(-2 * x * (y - (k * x + m)) for x, y in data) / n
    dm = sum(-2 * (y - (k * x + m)) for x, y in data) / n
    return dk, dm

k, m, lr = 0.0, 0.0, 0.01
for step in range(1, 2001):
    dk, dm = gradient(k, m)
    k -= lr * dk
    m -= lr * dm
    if step in (1, 10, 100, 500, 2000):
        print(f"step {step:>4}: k={k:.3f} m={m:.3f} error={error(k, m):.3f}")
# step    1: k=0.760 m=0.224 error=141.7
# step   10: k=4.093 m=1.184 error=8.54
# step  100: k=4.162 m=1.503 error=2.83
# step 2000: k=3.900 m=2.300 error=2.72   ← the bottom

# 3. Too large a learning rate: it goes off the rails
k, m, lr = 0.0, 0.0, 0.2
for step in range(5):
    dk, dm = gradient(k, m)
    k -= lr * dk; m -= lr * dm
    print(f"  lr=0.2 step {step}: k={k:.1f} error={error(k, m):.1f}")
#   lr=0.2 step 0: k=15.2 error=2244.5
#   lr=0.2 step 1: k=-70.5 error=61551.3    ← the wrong way, worse and worse

The last lines show the most important practical error in all training: with steps that are too large you miss the bottom and end up further away with every round.

Mastery means

  • Adjusts k and m to reduce the error
  • Explains what the error measures
  • Sees the connection to how a model is trained

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences