Skip to content
AI-grafen
FAI engineeringInference and optimisation· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Quantisation in practice: GPTQ, AWQ, GGUF

Be able to quantise a model with different methods and measure the quality and the speed.

Prerequisites

Intuition

Method/formatWhatWhen
bitsandbytes (nf4/int8)quantises at load time, no calibrationquick to try, GPU, fine-tuning (QLoRA)
GPTQ4-bit with calibration data, layer-wise optimisationGPU inference, the best quality per bit, takes 10–60 min to create
AWQprotects the «important» weights via activation statisticsGPU, often a little better than GPTQ on small models
GGUF (llama.cpp)K-quants (Q4_K_M, Q5_K_M, Q8_0), mixed precisionCPU/Apple Silicon/local, the best portability

Always measure three things: the perplexity on a held-out text (WikiText or your own), your own eval (the task quality — that is the one that counts), and tokens/s at your batch size. A model that has lost 0.5 PPL but 8 pp on your eval is not «nearly as good».

A rule of thumb: Q8 ≈ lossless, Q5_K_M very good, Q4_K_M a good compromise, below Q4 only if you must.

Code

# GGUF via llama.cpp
python convert_hf_to_gguf.py models/qwen2.5-7b-instruct --outfile q7b-f16.gguf
./llama-quantize q7b-f16.gguf q7b-Q4_K_M.gguf Q4_K_M
./llama-perplexity -m q7b-Q4_K_M.gguf -f wikitext-2-raw/wiki.test.raw -c 2048     # PPL
./llama-bench -m q7b-Q4_K_M.gguf -p 512 -n 128                                    # tokens/s
# AWQ in Python
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
m = AutoAWQForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct"); tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
m.quantize(tok, quant_config={"w_bit": 4, "q_group_size": 128, "zero_point": True, "version": "GEMM"})
m.save_quantized("qwen7b-awq")

# after that: run your eval harness on f16, AWQ and Q4_K_M — the same cases, the same prompt, temperature 0

The report table: the format · GB · PPL · your own eval · tokens/s · the platform.

Mastery means

  • Quantises a model with GPTQ/AWQ and to GGUF
  • Measures perplexity, task quality and tokens/s before and after
  • Chooses the format according to the runtime environment

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences