Skip to content
AI-grafen
DAI developerProject· about 150 min· server sandbox

Project 2: Build a neural network from scratch

A complete neural network in NumPy — softmax output, cross-entropy, backprop, minibatches, training on 8×8 digits — with ≥ 90 % accuracy on held-out data and a training curve as an artifact.

Theory

Classification: logits → softmax → cross-entropy. The gradient with respect to the logits is conveniently (p − y_onehot)/n. Minibatches give cheaper, noisier steps. The curve (loss per epoch) is your most important diagnostic tool.

Sub-tasks

  1. softmax + cross-entropy — softmax(Z) row-wise and stable; cross_entropy(P, y) mean −log p_y.
  2. forward/backward — forward(p, X) → (z1, a1, logits); backward(p, X, y, cache) → gradients (gradient check in the tests).
  3. minibatch training — train(X, y, hidden, lr, epochs, batch, seed) → (params, loss_history).
  4. evaluate + curve — accuracy(p, X, y); plot_curve(hist, path) saves a PNG to out/ (artifact).

Passes when: accuracy >= 0.9

The starter code

runs in an isolated sandbox on the server
import numpy as np


def init(d, h, k, seed=0):
    rng = np.random.default_rng(seed)
    return {"W1": rng.normal(0, np.sqrt(2 / d), (h, d)), "b1": np.zeros(h), "W2": rng.normal(0, np.sqrt(2 / h), (k, h)), "b2": np.zeros(k)}


def softmax(Z):
    # TODO
    ...


def cross_entropy(P, y):
    # TODO: medel av -log P[i, y[i]] (klipp till 1e-12)
    ...


def forward(p, X):
    # TODO: z1 = X @ W1.T + b1; a1 = relu; logits = a1 @ W2.T + b2
    ...


def backward(p, X, y, cache):
    # TODO: P = softmax(logits); dlogits = (P - onehot(y)) / n; sedan som i labben
    ...


def train(X, y, hidden=64, lr=0.1, epochs=30, batch=64, seed=0):
    p = init(X.shape[1], hidden, int(y.max()) + 1, seed)
    rng = np.random.default_rng(seed)
    hist = []
    # TODO: per epok: permutera, loopa batchar, uppdatera; hist.append(loss på hela X)
    ...
    return p, hist


def accuracy(p, X, y):
    # TODO
    ...


def plot_curve(hist, path="out/curve.png"):
    import os
    import matplotlib
    matplotlib.use("Agg")
    import matplotlib.pyplot as plt
    os.makedirs("out", exist_ok=True)
    plt.figure(figsize=(5, 3)); plt.plot(hist); plt.xlabel("epok"); plt.ylabel("loss"); plt.tight_layout(); plt.savefig(path); plt.close()
    return path

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

Gradient check < 1e-5; accuracy ≥ 0.90 on held-out data after ~30 epochs; a falling training curve in out/curve.png.

Common mistakes

  • Softmax without subtracting the max → overflow.
  • Cross-entropy with log(0) → −inf; clip p to ≥ 1e-12.
  • Does not shuffle the data between epochs.