Skip to content
AI-grafen
DAI developerLab· about 50 min· in the browser

Lab: gradient descent from scratch

Implement gradient descent for a simple loss and for linear regression, and see the effect of the learning rate.

Theory

w ← w − η·L'(w). For L(w) = (w−3)², L'(w) = 2(w−3). For linear regression with MSE the gradient with respect to k is (2/n)Σ(kxᵢ+m−yᵢ)xᵢ and with respect to m is (2/n)Σ(kxᵢ+m−yᵢ).

Sub-tasks

  1. steps — gd_1d(w0, eta, steg) returns w after steg (steps) updates on L(w) = (w−3)².
  2. gradient for the line — grad_linje(k, m, xs, ys) returns (dL/dk, dL/dm) for MSE.
  3. train the line — trana_linje(xs, ys, eta, steg) returns (k, m) after gradient descent from (0, 0).

Passes when: mse <= 1

The starter code

runs in your browser
def gd_1d(w0, eta, steg):
    # TODO: L(w) = (w-3)^2, L'(w) = 2(w-3)
    ...


def grad_linje(k, m, xs, ys):
    # TODO: MSE-gradient: (2/n) sum(e*x), (2/n) sum(e) med e = k*x + m - y
    ...


def trana_linje(xs, ys, eta=0.002, steg=20000):
    # TODO: starta i k = m = 0, uppdatera med grad_linje
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

gd_1d(0, 0.25, 3) ≈ 2.625; gd_1d(0, 0.5, 1) = 3.0; a line trained on the ice-cream data gives MSE < 1 (the answer k=3, m=−25 gives 0).

Common mistakes

  • The sign: w − η·grad, not plus.
  • Forgets the factor 2/n in the MSE gradient (still works with another η, but the test requires the exact gradient).
  • Too large an η on the ice-cream data (x ≈ 20) explodes — normalise or use a small η (≈0.002).