FAI engineeringInference and optimisation· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Quantisation in practice: GPTQ, AWQ, GGUF
Be able to quantise a model with different methods and measure the quality and the speed.
Prerequisites
- FQuantisationrequired
Intuition
| Method/format | What | When |
|---|---|---|
| bitsandbytes (nf4/int8) | quantises at load time, no calibration | quick to try, GPU, fine-tuning (QLoRA) |
| GPTQ | 4-bit with calibration data, layer-wise optimisation | GPU inference, the best quality per bit, takes 10–60 min to create |
| AWQ | protects the «important» weights via activation statistics | GPU, often a little better than GPTQ on small models |
| GGUF (llama.cpp) | K-quants (Q4_K_M, Q5_K_M, Q8_0), mixed precision | CPU/Apple Silicon/local, the best portability |
Always measure three things: the perplexity on a held-out text (WikiText or your own), your own eval (the task quality — that is the one that counts), and tokens/s at your batch size. A model that has lost 0.5 PPL but 8 pp on your eval is not «nearly as good».
A rule of thumb: Q8 ≈ lossless, Q5_K_M very good, Q4_K_M a good compromise, below Q4 only if you must.
Code
# GGUF via llama.cpp
python convert_hf_to_gguf.py models/qwen2.5-7b-instruct --outfile q7b-f16.gguf
./llama-quantize q7b-f16.gguf q7b-Q4_K_M.gguf Q4_K_M
./llama-perplexity -m q7b-Q4_K_M.gguf -f wikitext-2-raw/wiki.test.raw -c 2048 # PPL
./llama-bench -m q7b-Q4_K_M.gguf -p 512 -n 128 # tokens/s
# AWQ in Python
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
m = AutoAWQForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct"); tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
m.quantize(tok, quant_config={"w_bit": 4, "q_group_size": 128, "zero_point": True, "version": "GEMM"})
m.save_quantized("qwen7b-awq")
# after that: run your eval harness on f16, AWQ and Q4_K_M — the same cases, the same prompt, temperature 0
The report table: the format · GB · PPL · your own eval · tokens/s · the platform.
Mastery means
- Quantises a model with GPTQ/AWQ and to GGUF
- Measures perplexity, task quality and tokens/s before and after
- Chooses the format according to the runtime environment
Sign in to do the exercises and build your mastery up.
Sources
- llama.cpp (MIT) — MIT
- arXiv — AWQ: Activation-aware Weight Quantization — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0