Parallelism and why GPUs
Be able to explain why matrix multiplication parallelises well and what a GPU does differently from a CPU.
Prerequisites
Intuition
A matrix multiplication consists of millions of independent dot products. No result needs any other result — they can all be computed at the same time. That is called embarrassingly parallel, and it is exactly what a GPU is built for.
| CPU | GPU | |
|---|---|---|
| Cores | 8–64 complex ones | 5 000–20 000 simple ones |
| Optimised for | low latency per task | high throughput |
| Good at | branching, sequential logic | the same operation on a lot of data |
| Memory bandwidth | ~100 GB/s | 1–3 TB/s |
The bandwidth is often the real difference. In LLM decoding all the weights are read per token — then it is the memory bandwidth, not the computation, that sets the ceiling.
Formal
Amdahl's law sets the limit on what parallelisation can give. If the share of the work can be parallelised over units:
With and infinitely many units the ceiling is 20× — the serial five per cent dominates. That is why data loading, preprocessing and Python loops in the training loop can eat the whole GPU gain.
Arithmetic intensity (FLOPs per byte read) decides whether an operation is compute- or memory-bound:
| Operation | Intensity | Bound by |
|---|---|---|
| Element-wise addition | ~0.1 | memory |
| Matrix × vector (decoding) | ~2 | memory |
| Matrix × matrix (training, a large batch) | ~100+ | computation |
That is why training with large batches is compute-bound (the GPU is used well) while decoding one token at a time is memory-bound (the GPU stands waiting for the memory). It explains why batching helps so much in serving — it moves the work from memory-bound to compute-bound.
MFU (model FLOPs utilisation) measures what share of the GPU's theoretical capacity is actually used. 40–50 % is good in training; below 20 % means that something else is the bottleneck.
Code
import torch, time
def measure(n=4096, repeats=20, device="cuda"):
a = torch.randn(n, n, device=device, dtype=torch.bfloat16)
b = torch.randn(n, n, device=device, dtype=torch.bfloat16)
for _ in range(3): # warm-up
a @ b
torch.cuda.synchronize()
t0 = time.perf_counter()
for _ in range(repeats):
a @ b
torch.cuda.synchronize()
sec = (time.perf_counter() - t0) / repeats
return {"ms": round(sec * 1000, 2), "TFLOPs": round(2 * n**3 / sec / 1e12, 1)}
print(measure(device="cpu")) # {'ms': 2100.0, 'TFLOPs': 0.07}
print(measure(device="cuda")) # {'ms': 4.6, 'TFLOPs': 29.9}
def amdahl(p, N):
return 1 / ((1 - p) + p / N)
for p in (0.5, 0.9, 0.99):
print(p, [round(amdahl(p, n), 1) for n in (2, 8, 1000)])
# 0.5 [1.3, 1.8, 2.0] ← the ceiling is 2× whatever the number of units
# 0.9 [1.8, 4.7, 9.9]
# 0.99 [2.0, 7.5, 91.0]
When a GPU does not help: small models where the overhead dominates, heavily branching logic, data loading that cannot keep up (fix it with more workers and prefetching), and all the Python code in the loop.
Mastery means
- Explains why matrix operations parallelise well
- Describes the difference between a CPU and a GPU
- Identifies when a GPU does not help
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Amdahls lag (CC BY-SA 4.0) — CC BY-SA 4.0
- NVIDIA — GPU Performance Background — free to read