Sequence models before transformers
Be able to explain RNNs/LSTMs, why long dependencies are hard, and why attention solved the problem.
Prerequisites
- DBackpropagationrequired
- DNeural networks — the forward pass with matricesrequired
Intuition
Before transformers, sequences were handled with RNNs: a hidden state h updated for every element.
hₜ = tanh(W·xₜ + U·hₜ₋₁ + b)
All the information about the past has to fit in h — a vector of fixed size. Two problems follow:
- Vanishing gradients. The gradient from step 50 back to step 1 is multiplied by U fifty times. If the factors are < 1 it dies; if they are > 1 it explodes. Long dependencies are never learnt.
- No parallelisation. hₜ requires hₜ₋₁. The whole sequence has to be processed in order — deadlock on a GPU.
LSTMs (1997) partly solved the first with gates: a cell state that information can pass through almost unchanged, plus forget, input and output gates that learn what to keep. That was enough for machine translation for several years.
But problem 2 remained. Attention solved both: every position reaches every other directly (one «hop»), and everything can be computed in parallel.
Code
import torch, torch.nn as nn
# An RNN: sequential, a hidden state of fixed size
rnn = nn.RNN(input_size=16, hidden_size=32, batch_first=True)
x = torch.randn(8, 100, 16) # (batch, time, features)
out, h = rnn(x)
print(out.shape, h.shape) # (8, 100, 32) (1, 8, 32)
# An LSTM: the same interface, but a cell state plus gates
lstm = nn.LSTM(16, 32, batch_first=True)
out, (h, c) = lstm(x)
# An illustration of the vanishing gradient: the product of 50 factors
import numpy as np
for factor in (0.9, 1.0, 1.1):
print(factor, round(factor ** 50, 6))
# 0.9 0.005154 ← the gradient has practically disappeared
# 1.0 1.0
# 1.1 117.39 ← it explodes
Where RNNs are still used: very long streams with small resources (embedded systems), and in modern state models (Mamba, RWKV) that revive the idea of a running state — but with constructions that can be parallelised during training.
Mastery means
- Explains RNNs and LSTMs
- Describes why long dependencies are hard
- Justifies why attention replaced them
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Attention Is All You Need — arXiv (open access; licence per article)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- Hochreiter & Schmidhuber — Long Short-Term Memory (1997) — open PDF