QLoRA
Be able to fine-tune a quantised base model with LoRA adapters on a limited GPU and compare it with full LoRA.
Prerequisites
- FQuantisationrequired
- FLoRA — Low-Rank Adaptationrequired
Intuition
QLoRA = freeze the base model in 4 bits and train LoRA adapters in bf16 on top. The gradients are backpropagated through the quantised base, but only the adapters are updated.
The result: a 7B model fine-tuned on ~6 GB, a 65B on a single 48 GB card.
Three technical pieces that make it possible:
- NF4 (4-bit NormalFloat) — a quantisation format whose levels are information-theoretically optimal for normally distributed weights, which neural network weights roughly are.
- Double quantisation — the quantisation constants are quantised too. It saves ~0.4 bits per parameter.
- Paged optimisers — the optimiser state is moved to CPU memory at peaks, so that the training does not crash on temporary memory spikes.
The paper shows that QLoRA matches 16-bit full fine-tuning on their tasks.
Code
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # NF4, not ordinary int4
bnb_4bit_use_double_quant=True, # double quantisation
bnb_4bit_compute_dtype=torch.bfloat16, # the computations in bf16
)
m = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B", quantization_config=bnb, device_map="auto")
m = prepare_model_for_kbit_training(m, use_gradient_checkpointing=True)
m = get_peft_model(m, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
target_modules=["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"]))
opt = torch.optim.AdamW(m.parameters(), lr=2e-4) # a higher lr than for full fine-tuning
A memory comparison for 7B:
| The method | The base | The gradients + the optimiser | In total (approx.) |
|---|---|---|---|
| Full bf16 | 14 GB | 70 GB | 84 GB |
| LoRA bf16 | 14 GB | 0.5 GB | 15 GB |
| QLoRA nf4 | 3.5 GB | 0.5 GB | 6 GB |
The price: the training becomes 20–40 % slower per step (dequantisation and requantisation at every forward pass), and the adapter belongs to the quantised base — if you merge it with an fp16 base you get a small loss of quality. Run the evaluation in the same precision you intend to serve in.
Mastery means
- Fine-tunes a 4-bit quantised base model with LoRA
- Explains NF4, double quantisation and paged optimisers
- Compares QLoRA with LoRA in memory and quality
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — QLoRA: Efficient Finetuning of Quantized LLMs — arXiv (open access; licence per article)
- bitsandbytes (MIT) — MIT