Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

QLoRA

Be able to fine-tune a quantised base model with LoRA adapters on a limited GPU and compare it with full LoRA.

Prerequisites

Intuition

QLoRA = freeze the base model in 4 bits and train LoRA adapters in bf16 on top. The gradients are backpropagated through the quantised base, but only the adapters are updated.

The result: a 7B model fine-tuned on ~6 GB, a 65B on a single 48 GB card.

Three technical pieces that make it possible:

  1. NF4 (4-bit NormalFloat) — a quantisation format whose levels are information-theoretically optimal for normally distributed weights, which neural network weights roughly are.
  2. Double quantisation — the quantisation constants are quantised too. It saves ~0.4 bits per parameter.
  3. Paged optimisers — the optimiser state is moved to CPU memory at peaks, so that the training does not crash on temporary memory spikes.

The paper shows that QLoRA matches 16-bit full fine-tuning on their tasks.

Code

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",              # NF4, not ordinary int4
    bnb_4bit_use_double_quant=True,         # double quantisation
    bnb_4bit_compute_dtype=torch.bfloat16,  # the computations in bf16
)
m = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B", quantization_config=bnb, device_map="auto")
m = prepare_model_for_kbit_training(m, use_gradient_checkpointing=True)
m = get_peft_model(m, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
                                 target_modules=["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"]))

opt = torch.optim.AdamW(m.parameters(), lr=2e-4)    # a higher lr than for full fine-tuning

A memory comparison for 7B:

The methodThe baseThe gradients + the optimiserIn total (approx.)
Full bf1614 GB70 GB84 GB
LoRA bf1614 GB0.5 GB15 GB
QLoRA nf43.5 GB0.5 GB6 GB

The price: the training becomes 20–40 % slower per step (dequantisation and requantisation at every forward pass), and the adapter belongs to the quantised base — if you merge it with an fp16 base you get a small loss of quality. Run the evaluation in the same precision you intend to serve in.

Mastery means

  • Fine-tunes a 4-bit quantised base model with LoRA
  • Explains NF4, double quantisation and paged optimisers
  • Compares QLoRA with LoRA in memory and quality

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences