Skip to content
AI-grafen
DAI developerDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Train a neural network in PyTorch

Be able to write a complete training loop — model, loss, optimiser, batches, evaluation — and read the training curve to see whether the model is learning, underfitting or overfitting.

Prerequisites

Code

import torch, torch.nn as nn

model = nn.Sequential(nn.Linear(784, 128), nn.ReLU(), nn.Linear(128, 10))
loss_fn = nn.CrossEntropyLoss()
opt = torch.optim.SGD(model.parameters(), lr=0.1)

for epoch in range(5):
    model.train()
    for X, y in train_loader:            # batches of (N, 784), (N,)
        opt.zero_grad()
        loss = loss_fn(model(X), y)
        loss.backward()
        opt.step()
    model.eval()
    with torch.no_grad():
        correct = sum((model(X).argmax(1) == y).sum().item() for X, y in val_loader)
    print(epoch, loss.item(), correct / len(val_loader.dataset))

Five lines that are the whole of deep learning: zero, forward, loss, backward, step.

Intuition

Read the curves:

Training lossValidation lossDiagnosis
fallingfallinglearning — carry on
fallingrisingoverfitting — memorising the training data
high, flathighunderfitting — too simple a model, or too low an lr
exploding / NaN—lr too high, or unnormalised data

CrossEntropyLoss takes the raw outputs (logits) and applies softmax internally — do not add a softmax yourself. Always validate on data the model has not been trained on.

Formal

An epoch = one pass over all the training data. The batch size controls how many examples go into each gradient step. Adam (torch.optim.Adam) is a popular optimiser that adapts the step length per parameter and often works without tuning the lr (start at 1e-3).

Mastery means

  • Writes a training loop with nn.Module, a loss and an optimizer
  • Interprets the training and validation loss over epochs

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences