Skip to content
AI-grafen
DAI developerLab· about 75 min· server sandbox

Lab: a neural network in pure NumPy — forward, backprop, gradient check

Implement the forward pass, loss, backprop and training loop for a two-layer network, verify the gradients numerically and teach the network XOR.

Theory

z₁ = W₁x + b₁, a₁ = ReLU(z₁), z₂ = W₂a₁ + b₂, L = ½‖z₂ − y‖². Backwards: δ₂ = z₂ − y; ∂L/∂W₂ = δ₂a₁ᵀ; δ₁ = (W₂ᵀδ₂) ⊙ 1[z₁>0]; ∂L/∂W₁ = δ₁xᵀ. The numerical check (L(w+ε) − L(w−ε))/2ε exposes every mistake in the derivation.

Sub-tasks

  1. forward — forward(params, X) returns (z1, a1, z2) for a batch X (n×d).
  2. backward — backward(params, X, Y, cache) returns gradients for W1, b1, W2, b2 (mean over the batch).
  3. train — train(X, Y, hidden, lr, steps, seed) trains with gradient descent and returns params.

Passes when: loss <= 0.05

The starter code

runs in an isolated sandbox on the server
import numpy as np


def init(d, h, k, seed=0):
    rng = np.random.default_rng(seed)
    return {"W1": rng.normal(0, 0.5, (h, d)), "b1": np.zeros(h), "W2": rng.normal(0, 0.5, (k, h)), "b2": np.zeros(k)}


def forward(p, X):
    """X: (n, d). Returnerar (z1, a1, z2) med former (n, h), (n, h), (n, k)."""
    # TODO
    ...


def loss(z2, Y):
    return 0.5 * float(np.mean(np.sum((z2 - Y) ** 2, axis=1)))


def backward(p, X, Y, cache):
    """cache = (z1, a1, z2). Returnerar dict med dW1, db1, dW2, db2 (medel över batchen)."""
    # TODO
    ...


def train(X, Y, hidden=8, lr=0.5, steps=3000, seed=0):
    p = init(X.shape[1], hidden, Y.shape[1], seed)
    # TODO: loop: forward → backward → uppdatera alla parametrar
    ...
    return p

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

The gradient check gives a relative difference < 1e-5 for all parameters. The XOR loss after 3000 steps < 0.05.

Common mistakes

  • ReLU derivative forgotten → δ₁ = W₂ᵀδ₂ without masking (the gradient check for W1 fails).
  • Wrong transpose: W₁ is h×d, so the shape of ∂L/∂W₁ depends on your layout — stick to one.
  • The mean over the batch is missing → the gradient is n times too large.