DAI developerLab· about 50 min· in the browser
Lab: gradient descent from scratch
Implement gradient descent for a simple loss and for linear regression, and see the effect of the learning rate.
Theory
w ← w − η·L'(w). For L(w) = (w−3)², L'(w) = 2(w−3). For linear regression with MSE the gradient with respect to k is (2/n)Σ(kxᵢ+m−yᵢ)xᵢ and with respect to m is (2/n)Σ(kxᵢ+m−yᵢ).
Sub-tasks
- steps —
gd_1d(w0, eta, steg)returns w aftersteg(steps) updates on L(w) = (w−3)². - gradient for the line —
grad_linje(k, m, xs, ys)returns (dL/dk, dL/dm) for MSE. - train the line —
trana_linje(xs, ys, eta, steg)returns (k, m) after gradient descent from (0, 0).
Passes when: mse <= 1
The starter code
runs in your browserdef gd_1d(w0, eta, steg):
# TODO: L(w) = (w-3)^2, L'(w) = 2(w-3)
...
def grad_linje(k, m, xs, ys):
# TODO: MSE-gradient: (2/n) sum(e*x), (2/n) sum(e) med e = k*x + m - y
...
def trana_linje(xs, ys, eta=0.002, steg=20000):
# TODO: starta i k = m = 0, uppdatera med grad_linje
...
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
gd_1d(0, 0.25, 3) ≈ 2.625; gd_1d(0, 0.5, 1) = 3.0; a line trained on the ice-cream data gives MSE < 1 (the answer k=3, m=−25 gives 0).
Common mistakes
- The sign: w − η·grad, not plus.
- Forgets the factor 2/n in the MSE gradient (still works with another η, but the test requires the exact gradient).
- Too large an η on the ice-cream data (x ≈ 20) explodes — normalise or use a small η (≈0.002).