The goal E University
Training neural networks for real
Optimisers, regularisation, normalisation, initialisation, dataloaders and transfer learning — everything that separates a training run that works from one that does not.
- Knowledge nodes
- 86
- From zero
- about 65 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- BNeural networks: inspiration from the brain
- DActivation functions
- EVanishing and exploding gradients
- EGradient clipping
- EWeight initialisation
- DEmbeddings for categorical variables
- DThe perceptron
- EHow autograd works inside
- DDataloaders, batches and epochs
- EResidual connections
- ESequence models before transformers
- EAutoencoders
- EDeep learning on tabular data
- EBatch and layer normalisation
- ECheckpoints, saving and restarting
- EGPU memory, gradient accumulation and batches
- FMixed precision training
- EOptimisers: momentum, Adam, scheduling
- ELearning rate schedules and warm-up
- ERegularisation: dropout, weight decay, early stopping
- EDropout in detail
- ETransfer learning
- ETraining diagnostics
- EHyperparameters for deep networks
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: semantic search with embeddingsDa sandbox · about 45 minLab: tensors, autograd and your first training loopDa sandbox · about 50 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer4 knowledge nodes
BInvestigator10 knowledge nodes
- Data in everyday life
- Charts and tables
- Algorithmic thinking
- Rule-based systems and machine learning
- Binary numbers and bits
- Fractions, decimals, and percentages
- Negative numbers in everyday life and AI
- Training data, features and labels
- Classification: how a model sorts information
- Neural networks: inspiration from the brain
CBuilder12 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Python — strings and text processing
- Python — files, CSV and JSON
- Statistics — mean, median and spread
- Probability — the basics
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Variables and algebraic expressions
- Powers and roots
DAI developer31 knowledge nodes
- Derivatives and optimisation
- Time complexity and big-O notation
- Python — functions, scope and exceptions
- Python — modules, packages and virtual environments
- Git — version control
- Probability distributions
- The normal distribution and standardisation
- Gradient descent
- Logarithms
- Vectors
- Matrices and matrix multiplication
- Linear regression with several features
- Logistic regression and decision trees
- Decision trees
- Neural networks — the forward pass with matrices
- Activation functions
- Backpropagation
- Embeddings — words as vectors
- NumPy — arrays and vectorisation
- Overfitting and generalisation
- Pandas — tables in Python
- Data quality: missing values, duplicates, errors
- Feature engineering
- Embeddings for categorical variables
- The perceptron
- PyTorch — tensors and autograd
- Dataloaders, batches and epochs
- Training, validation and test
- Cross-validation
- Train a neural network in PyTorch
- Visualisation with Matplotlib
EUniversity27 knowledge nodes
- Numerical stability and floating point
- Ensembles: random forest and boosting
- Linear maps
- Eigenvalues and eigenvectors
- Matrix factorisation and low-rank approximation
- Vanishing and exploding gradients
- Gradient clipping
- Weight initialisation
- Parallelism and why GPUs
- How autograd works inside
- Residual connections
- Sequence models before transformers
- Singular value decomposition (SVD)
- PCA — principal component analysis
- Autoencoders
- Deep learning on tabular data
- Hyperparameter search
- Batch and layer normalisation
- Checkpoints, saving and restarting
- GPU memory, gradient accumulation and batches
- Optimisers: momentum, Adam, scheduling
- Learning rate schedules and warm-up
- Regularisation: dropout, weight decay, early stopping
- Dropout in detail
- Transfer learning
- Training diagnostics
- Hyperparameters for deep networks