Skip to content
AI-grafen
EUniversityComputer science· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Parallelism and why GPUs

Be able to explain why matrix multiplication parallelises well and what a GPU does differently from a CPU.

Prerequisites

Intuition

A matrix multiplication consists of millions of independent dot products. No result needs any other result — they can all be computed at the same time. That is called embarrassingly parallel, and it is exactly what a GPU is built for.

CPUGPU
Cores8–64 complex ones5 000–20 000 simple ones
Optimised forlow latency per taskhigh throughput
Good atbranching, sequential logicthe same operation on a lot of data
Memory bandwidth~100 GB/s1–3 TB/s

The bandwidth is often the real difference. In LLM decoding all the weights are read per token — then it is the memory bandwidth, not the computation, that sets the ceiling.

Formal

Amdahl's law sets the limit on what parallelisation can give. If the share pp of the work can be parallelised over NN units: speedup=1(1−p)+p/N\text{speedup} = \frac{1}{(1-p) + p/N}

With p=0.95p = 0.95 and infinitely many units the ceiling is 20× — the serial five per cent dominates. That is why data loading, preprocessing and Python loops in the training loop can eat the whole GPU gain.

Arithmetic intensity (FLOPs per byte read) decides whether an operation is compute- or memory-bound:

OperationIntensityBound by
Element-wise addition~0.1memory
Matrix × vector (decoding)~2memory
Matrix × matrix (training, a large batch)~100+computation

That is why training with large batches is compute-bound (the GPU is used well) while decoding one token at a time is memory-bound (the GPU stands waiting for the memory). It explains why batching helps so much in serving — it moves the work from memory-bound to compute-bound.

MFU (model FLOPs utilisation) measures what share of the GPU's theoretical capacity is actually used. 40–50 % is good in training; below 20 % means that something else is the bottleneck.

Code

import torch, time

def measure(n=4096, repeats=20, device="cuda"):
    a = torch.randn(n, n, device=device, dtype=torch.bfloat16)
    b = torch.randn(n, n, device=device, dtype=torch.bfloat16)
    for _ in range(3):                      # warm-up
        a @ b
    torch.cuda.synchronize()
    t0 = time.perf_counter()
    for _ in range(repeats):
        a @ b
    torch.cuda.synchronize()
    sec = (time.perf_counter() - t0) / repeats
    return {"ms": round(sec * 1000, 2), "TFLOPs": round(2 * n**3 / sec / 1e12, 1)}

print(measure(device="cpu"))     # {'ms': 2100.0, 'TFLOPs': 0.07}
print(measure(device="cuda"))    # {'ms': 4.6,    'TFLOPs': 29.9}

def amdahl(p, N):
    return 1 / ((1 - p) + p / N)

for p in (0.5, 0.9, 0.99):
    print(p, [round(amdahl(p, n), 1) for n in (2, 8, 1000)])
# 0.5  [1.3, 1.8, 2.0]     ← the ceiling is 2× whatever the number of units
# 0.9  [1.8, 4.7, 9.9]
# 0.99 [2.0, 7.5, 91.0]

When a GPU does not help: small models where the overhead dominates, heavily branching logic, data loading that cannot keep up (fix it with more workers and prefetching), and all the Python code in the loop.

Mastery means

  • Explains why matrix operations parallelise well
  • Describes the difference between a CPU and a GPU
  • Identifies when a GPU does not help

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences