DAI developerProject· about 150 min· server sandbox
Project 2: Build a neural network from scratch
A complete neural network in NumPy — softmax output, cross-entropy, backprop, minibatches, training on 8×8 digits — with ≥ 90 % accuracy on held-out data and a training curve as an artifact.
Theory
Classification: logits → softmax → cross-entropy. The gradient with respect to the logits is conveniently (p − y_onehot)/n. Minibatches give cheaper, noisier steps. The curve (loss per epoch) is your most important diagnostic tool.
Sub-tasks
- softmax + cross-entropy —
softmax(Z)row-wise and stable;cross_entropy(P, y)mean −log p_y. - forward/backward —
forward(p, X)→ (z1, a1, logits);backward(p, X, y, cache)→ gradients (gradient check in the tests). - minibatch training —
train(X, y, hidden, lr, epochs, batch, seed)→ (params, loss_history). - evaluate + curve —
accuracy(p, X, y);plot_curve(hist, path)saves a PNG to out/ (artifact).
Passes when: accuracy >= 0.9
The starter code
runs in an isolated sandbox on the serverimport numpy as np
def init(d, h, k, seed=0):
rng = np.random.default_rng(seed)
return {"W1": rng.normal(0, np.sqrt(2 / d), (h, d)), "b1": np.zeros(h), "W2": rng.normal(0, np.sqrt(2 / h), (k, h)), "b2": np.zeros(k)}
def softmax(Z):
# TODO
...
def cross_entropy(P, y):
# TODO: medel av -log P[i, y[i]] (klipp till 1e-12)
...
def forward(p, X):
# TODO: z1 = X @ W1.T + b1; a1 = relu; logits = a1 @ W2.T + b2
...
def backward(p, X, y, cache):
# TODO: P = softmax(logits); dlogits = (P - onehot(y)) / n; sedan som i labben
...
def train(X, y, hidden=64, lr=0.1, epochs=30, batch=64, seed=0):
p = init(X.shape[1], hidden, int(y.max()) + 1, seed)
rng = np.random.default_rng(seed)
hist = []
# TODO: per epok: permutera, loopa batchar, uppdatera; hist.append(loss på hela X)
...
return p, hist
def accuracy(p, X, y):
# TODO
...
def plot_curve(hist, path="out/curve.png"):
import os
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
os.makedirs("out", exist_ok=True)
plt.figure(figsize=(5, 3)); plt.plot(hist); plt.xlabel("epok"); plt.ylabel("loss"); plt.tight_layout(); plt.savefig(path); plt.close()
return path
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
Gradient check < 1e-5; accuracy ≥ 0.90 on held-out data after ~30 epochs; a falling training curve in out/curve.png.
Common mistakes
- Softmax without subtracting the max → overflow.
- Cross-entropy with log(0) → −inf; clip p to ≥ 1e-12.
- Does not shuffle the data between epochs.