Quantisation
Be able to explain int8/int4 quantisation, measure the quality loss and the memory gain, and run a quantised model locally.
Prerequisites
Intuition
A 7B model in float16 takes 14 GB. In int4 it takes 3.5–4 GB and fits on a laptop. Quantisation = storing the weights with fewer bits.
The recipe (linear quantisation): for a group of weights, find the maximum absolute value → the scale s = max/127 (int8). Store q = round(w/s) as an integer; at computation time w ≈ q·s. With a zero point you can use the whole range asymmetrically. Finer groups (per channel, per 64–128 weights) → a smaller error, a little more metadata.
What is lost? The per-token perplexity rises somewhat: int8 nearly free, int4 often +1–3 % PPL, int3 and below do damage. Outliers (occasional huge weights or activations) are the main problem — hence methods such as GPTQ/AWQ that quantise «cleverly» with calibration data.
Quantisation also gives faster inference when memory bandwidth is the bottleneck (which it is at decoding): fewer bytes to read per token.
Code
import numpy as np
def quantize(w, bits=8, group=64):
"""Symmetric per-group quantisation. Returns (q int, the scales)."""
qmax = 2 ** (bits - 1) - 1
w = w.reshape(-1, group)
s = np.abs(w).max(axis=1, keepdims=True) / qmax
s[s == 0] = 1e-8
q = np.clip(np.round(w / s), -qmax - 1, qmax).astype(np.int8 if bits <= 8 else np.int16)
return q, s
def dequantize(q, s, shape):
return (q.astype(np.float32) * s).reshape(shape)
rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, (4096, 4096)).astype(np.float32); W[0, 0] = 1.0 # one outlier
for bits in (8, 4):
q, s = quantize(W, bits)
err = np.linalg.norm(W - dequantize(q, s, W.shape)) / np.linalg.norm(W)
mb = q.size * bits / 8 / 1e6 + s.size * 2 / 1e6
print(bits, "rel. error", round(float(err), 4), "MB", round(mb, 1), "against fp16", W.size * 2 / 1e6)
# 8 rel. error ~0.002 MB 16.9 against fp16 33.6
# 4 rel. error ~0.04 MB 8.5
Locally: llama.cpp with GGUF (Q4_K_M ≈ 4.5 bits/weight) or bitsandbytes in transformers (load_in_4bit=True).
Formal
Affine quantisation: , . The quantisation noise for uniform rounding has variance ; the relative error scales with in the group, which is why outliers ruin things and small groups help. GPTQ minimises layer by layer with Hessian information () from calibration data; AWQ scales the channels so that important (activation-heavy) weights get a smaller relative error. Activation quantisation (W8A8) requires handling activation outliers (SmoothQuant). Memory bandwidth: decoding a token reads all the weights once → the time ≈ bytes/bandwidth; halved bytes ≈ double the speed until the computation becomes the bottleneck.
Mastery means
- Explains int8/int4 quantisation: the scale, the zero point, per channel/group
- Measures the memory gain and the quality loss
- Runs a quantised model locally
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — arXiv (open access; licence per article)
- arXiv — AWQ: Activation-aware Weight Quantization — arXiv (open access; licence per article)
- llama.cpp (MIT) — MIT