Train a neural network in PyTorch
Be able to write a complete training loop — model, loss, optimiser, batches, evaluation — and read the training curve to see whether the model is learning, underfitting or overfitting.
Prerequisites
- DBackpropagationrequired
- DPyTorch — tensors and autogradrequired
Code
import torch, torch.nn as nn
model = nn.Sequential(nn.Linear(784, 128), nn.ReLU(), nn.Linear(128, 10))
loss_fn = nn.CrossEntropyLoss()
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for epoch in range(5):
model.train()
for X, y in train_loader: # batches of (N, 784), (N,)
opt.zero_grad()
loss = loss_fn(model(X), y)
loss.backward()
opt.step()
model.eval()
with torch.no_grad():
correct = sum((model(X).argmax(1) == y).sum().item() for X, y in val_loader)
print(epoch, loss.item(), correct / len(val_loader.dataset))
Five lines that are the whole of deep learning: zero, forward, loss, backward, step.
Intuition
Read the curves:
| Training loss | Validation loss | Diagnosis |
|---|---|---|
| falling | falling | learning — carry on |
| falling | rising | overfitting — memorising the training data |
| high, flat | high | underfitting — too simple a model, or too low an lr |
| exploding / NaN | — | lr too high, or unnormalised data |
CrossEntropyLoss takes the raw outputs (logits) and applies softmax internally — do not add
a softmax yourself. Always validate on data the model has not been trained on.
Formal
An epoch = one pass over all the training data. The batch size controls how many examples
go into each gradient step. Adam (torch.optim.Adam) is a popular optimiser that adapts
the step length per parameter and often works without tuning the lr (start at 1e-3).
Mastery means
- Writes a training loop with nn.Module, a loss and an optimizer
- Interprets the training and validation loss over epochs
Sign in to do the exercises and build your mastery up.
Sources
Leads to
- DMNIST from scratch
- DTransformers — the architecture
- EBatch and layer normalisation
- ECheckpoints, saving and restarting
- EFine-tuning language models
- EGPU memory, gradient accumulation and batches
- EOptimisers: momentum, Adam, scheduling
- ERegularisation: dropout, weight decay, early stopping
- EReproducibility
- ETransfer learning
- ETraining diagnostics
- FDistributed training
- FMixed precision training
Part of the goals (28)
- Train your first neural network
- Training neural networks for real
- Build a transformer from scratch
- Understand how generative AI works
- Image classification with convolutional networks
- Fine-tune and run your own models
- Build a voice interface
- Run models more cheaply: quantisation
- Frontier Lab — an independent research project
- Language models in practice
- AI safety in practice
- Fine-tune a model with LoRA
- Responsible AI in practice
- Build a RAG system you can trust
- Statistics for experiments
- Reproduce a paper
- Evals in practice
- Interpreting a language model
- Multimodal systems
- Deep reinforcement learning
- Build an agent you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- An AI service in operation
- Build a memory system for an agent
- Build an AI service that survives production
- AI, ethics and society