Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· fundamentals that rarely change· verified 2026-09-20· EN

Sequence models before transformers

Be able to explain RNNs/LSTMs, why long dependencies are hard, and why attention solved the problem.

Prerequisites

Intuition

Before transformers, sequences were handled with RNNs: a hidden state h updated for every element.

hₜ = tanh(W·xₜ + U·hₜ₋₁ + b)

All the information about the past has to fit in h — a vector of fixed size. Two problems follow:

  1. Vanishing gradients. The gradient from step 50 back to step 1 is multiplied by U fifty times. If the factors are < 1 it dies; if they are > 1 it explodes. Long dependencies are never learnt.
  2. No parallelisation. hₜ requires hₜ₋₁. The whole sequence has to be processed in order — deadlock on a GPU.

LSTMs (1997) partly solved the first with gates: a cell state that information can pass through almost unchanged, plus forget, input and output gates that learn what to keep. That was enough for machine translation for several years.

But problem 2 remained. Attention solved both: every position reaches every other directly (one «hop»), and everything can be computed in parallel.

Code

import torch, torch.nn as nn

# An RNN: sequential, a hidden state of fixed size
rnn = nn.RNN(input_size=16, hidden_size=32, batch_first=True)
x = torch.randn(8, 100, 16)          # (batch, time, features)
out, h = rnn(x)
print(out.shape, h.shape)             # (8, 100, 32) (1, 8, 32)

# An LSTM: the same interface, but a cell state plus gates
lstm = nn.LSTM(16, 32, batch_first=True)
out, (h, c) = lstm(x)

# An illustration of the vanishing gradient: the product of 50 factors
import numpy as np
for factor in (0.9, 1.0, 1.1):
    print(factor, round(factor ** 50, 6))
# 0.9  0.005154   ← the gradient has practically disappeared
# 1.0  1.0
# 1.1  117.39     ← it explodes

Where RNNs are still used: very long streams with small resources (embedded systems), and in modern state models (Mamba, RWKV) that revive the idea of a running state — but with constructions that can be parallelised during training.

Mastery means

  • Explains RNNs and LSTMs
  • Describes why long dependencies are hard
  • Justifies why attention replaced them

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences