DAI developerLab· about 75 min· server sandbox
Lab: a neural network in pure NumPy — forward, backprop, gradient check
Implement the forward pass, loss, backprop and training loop for a two-layer network, verify the gradients numerically and teach the network XOR.
Theory
z₁ = W₁x + b₁, a₁ = ReLU(z₁), z₂ = W₂a₁ + b₂, L = ½‖z₂ − y‖². Backwards: δ₂ = z₂ − y; ∂L/∂W₂ = δ₂a₁ᵀ; δ₁ = (W₂ᵀδ₂) ⊙ 1[z₁>0]; ∂L/∂W₁ = δ₁xᵀ. The numerical check (L(w+ε) − L(w−ε))/2ε exposes every mistake in the derivation.
Sub-tasks
- forward —
forward(params, X)returns (z1, a1, z2) for a batch X (n×d). - backward —
backward(params, X, Y, cache)returns gradients for W1, b1, W2, b2 (mean over the batch). - train —
train(X, Y, hidden, lr, steps, seed)trains with gradient descent and returns params.
Passes when: loss <= 0.05
The starter code
runs in an isolated sandbox on the serverimport numpy as np
def init(d, h, k, seed=0):
rng = np.random.default_rng(seed)
return {"W1": rng.normal(0, 0.5, (h, d)), "b1": np.zeros(h), "W2": rng.normal(0, 0.5, (k, h)), "b2": np.zeros(k)}
def forward(p, X):
"""X: (n, d). Returnerar (z1, a1, z2) med former (n, h), (n, h), (n, k)."""
# TODO
...
def loss(z2, Y):
return 0.5 * float(np.mean(np.sum((z2 - Y) ** 2, axis=1)))
def backward(p, X, Y, cache):
"""cache = (z1, a1, z2). Returnerar dict med dW1, db1, dW2, db2 (medel över batchen)."""
# TODO
...
def train(X, Y, hidden=8, lr=0.5, steps=3000, seed=0):
p = init(X.shape[1], hidden, Y.shape[1], seed)
# TODO: loop: forward → backward → uppdatera alla parametrar
...
return p
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
The gradient check gives a relative difference < 1e-5 for all parameters. The XOR loss after 3000 steps < 0.05.
Common mistakes
- ReLU derivative forgotten → δ₁ = W₂ᵀδ₂ without masking (the gradient check for W1 fails).
- Wrong transpose: W₁ is h×d, so the shape of ∂L/∂W₁ depends on your layout — stick to one.
- The mean over the batch is missing → the gradient is n times too large.