Skip to content
AI-grafen
DAI developerDeep learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Gradient descent

Be able to carry out gradient descent steps by hand on a simple loss, explain the role of the learning rate, and implement the loop in a few lines of Python.

Prerequisites

Intuition

We have a loss function — the error as a function of the weights. We want to reach the bottom. Gradient descent: work out the slope (the gradient) where we are standing, take a small step against the slope, repeat.

w ← w − η · L'(w)

η (eta) is the learning rate: the step length. Too small and it takes forever. Too large and you jump over the valley and can end up higher than you started.

Code

L(w) = (w − 3)², L'(w) = 2(w − 3). Starting at w = 0, η = 0.25:

stepwL'(w)new w
10−60 − 0.25·(−6) = 1.5
21.5−32.25
32.25−1.52.625

It is approaching 3. In Python:

w, eta = 0.0, 0.25
for step in range(20):
    grad = 2 * (w - 3)
    w = w - eta * grad
print(w)   # ≈ 3.0

With several weights the gradient is a vector of partial derivatives — the same rule, component by component. The stochastic variant of gradient descent (SGD) computes the gradient on a small batch of examples at a time, which is cheaper and often better.

Mastery means

  • Carries out two update steps by hand
  • Explains what happens with too large and too small a learning rate

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences