Skip to content
AI-grafen
FAI engineeringInference and optimisation· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Quantisation

Be able to explain int8/int4 quantisation, measure the quality loss and the memory gain, and run a quantised model locally.

Prerequisites

Intuition

A 7B model in float16 takes 14 GB. In int4 it takes 3.5–4 GB and fits on a laptop. Quantisation = storing the weights with fewer bits.

The recipe (linear quantisation): for a group of weights, find the maximum absolute value → the scale s = max/127 (int8). Store q = round(w/s) as an integer; at computation time w ≈ q·s. With a zero point you can use the whole range asymmetrically. Finer groups (per channel, per 64–128 weights) → a smaller error, a little more metadata.

What is lost? The per-token perplexity rises somewhat: int8 nearly free, int4 often +1–3 % PPL, int3 and below do damage. Outliers (occasional huge weights or activations) are the main problem — hence methods such as GPTQ/AWQ that quantise «cleverly» with calibration data.

Quantisation also gives faster inference when memory bandwidth is the bottleneck (which it is at decoding): fewer bytes to read per token.

Code

import numpy as np

def quantize(w, bits=8, group=64):
    """Symmetric per-group quantisation. Returns (q int, the scales)."""
    qmax = 2 ** (bits - 1) - 1
    w = w.reshape(-1, group)
    s = np.abs(w).max(axis=1, keepdims=True) / qmax
    s[s == 0] = 1e-8
    q = np.clip(np.round(w / s), -qmax - 1, qmax).astype(np.int8 if bits <= 8 else np.int16)
    return q, s

def dequantize(q, s, shape):
    return (q.astype(np.float32) * s).reshape(shape)

rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, (4096, 4096)).astype(np.float32); W[0, 0] = 1.0   # one outlier
for bits in (8, 4):
    q, s = quantize(W, bits)
    err = np.linalg.norm(W - dequantize(q, s, W.shape)) / np.linalg.norm(W)
    mb = q.size * bits / 8 / 1e6 + s.size * 2 / 1e6
    print(bits, "rel. error", round(float(err), 4), "MB", round(mb, 1), "against fp16", W.size * 2 / 1e6)
# 8 rel. error ~0.002  MB 16.9  against fp16 33.6
# 4 rel. error ~0.04   MB 8.5

Locally: llama.cpp with GGUF (Q4_K_M ≈ 4.5 bits/weight) or bitsandbytes in transformers (load_in_4bit=True).

Formal

Affine quantisation: q=clamp(⌊w/s⌉+z, 0, 2b−1)q = \text{clamp}(\lfloor w/s \rceil + z,\, 0,\, 2^b-1), w^=s(q−z)\hat w = s(q - z). The quantisation noise for uniform rounding has variance s2/12s^2/12; the relative error scales with max⁡∣w∣/σw\max|w|/\sigma_w in the group, which is why outliers ruin things and small groups help. GPTQ minimises ∥WX−W^X∥2\|WX - \hat W X\|^2 layer by layer with Hessian information (H=XX⊤H = XX^\top) from calibration data; AWQ scales the channels so that important (activation-heavy) weights get a smaller relative error. Activation quantisation (W8A8) requires handling activation outliers (SmoothQuant). Memory bandwidth: decoding a token reads all the weights once → the time ≈ bytes/bandwidth; halved bytes ≈ double the speed until the computation becomes the bottleneck.

Mastery means

  • Explains int8/int4 quantisation: the scale, the zero point, per channel/group
  • Measures the memory gain and the quality loss
  • Runs a quantised model locally

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences