LoRA — Low-Rank Adaptation
Be able to explain why LoRA works (ΔW = BA), choose the target modules and the rank, fine-tune a model with LoRA and evaluate it with evals.
Prerequisites
- DBackpropagationrequired
- EFine-tuning language modelsrequired
- EMatrix factorisation and low-rank approximationrequired
- EModel evaluationrequired
Intuition
A full fine-tune of a 7B model updates 7 billion weights and needs optimizer state for all of them — ~100 GB. LoRA freezes the model and learns only a small low-rank update per matrix: W′ = W + ΔW with ΔW = B·A, where A is r × d and B is d × r, r ≪ d (8–64, say). For d = 4096, r = 16: 131 k parameters instead of 16.8 M per matrix.
Why does it work? Fine-tuning changes the model in few directions — ΔW has a low «intrinsic rank». The hypothesis is confirmed empirically: LoRA gets close to a full fine-tune on most tasks.
In practice: the target modules = which matrices (q, k, v, o, and often the MLP layers too — the latter helps most when you have the budget). α/r scales the update; keep α ≈ 2r as a start. After training, BA can be merged into W — zero extra inference cost. QLoRA: the base model in 4-bit, LoRA in bf16 → a 7B trains on an 8–10 GB GPU.
Measure against the base model and against a full fine-tune if you can; LoRA with too low an r underperforms on tasks that require new knowledge.
Code
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
import torch
name = "Qwen/Qwen2.5-1.5B"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16)
m = AutoModelForCausalLM.from_pretrained(name, quantization_config=bnb, device_map="auto")
m = prepare_model_for_kbit_training(m)
cfg = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"])
m = get_peft_model(m, cfg)
m.print_trainable_parameters() # ~1 % of the parameters
# …an ordinary training loop / Trainer on instruction data (loss on the answer only), lr ~2e-4, 1–3 epochs…
m.save_pretrained("adapter/") # just A and B — a few tens of MB
# inference: load the base plus the adapter, or m.merge_and_unload() to bake it in
The lab lora-tiny-labb implements the LoRA layer from scratch in NumPy/PyTorch and verifies it against a reference.
Derivation
Forward: . Parameters: against . Initialisation: , → at the start, so the model begins exactly as the base does. Gradients: , where — the same as a full fine-tune, but projected. The memory saving comes mainly from the optimizer state (Adam: 2 × the parameters in fp32) only being needed for . Choosing the rank: the effective rank of in a full fine-tune can be measured with an SVD; empirically is enough for style and format, more for new factual knowledge. The scaling means the learning rate does not have to be adjusted when changes (rsLoRA proposes ).
Mastery means
- Explains ΔW = BA, the rank and the α/r scaling
- Chooses the target modules and the rank with justification
- Fine-tunes with LoRA/QLoRA and evaluates against the base model and a full fine-tune
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — LoRA: Low-Rank Adaptation of Large Language Models — arXiv (open access; licence per article)
- arXiv — QLoRA: Efficient Finetuning of Quantized LLMs — arXiv (open access; licence per article)
- PEFT — dokumentation (Apache-2.0) — Apache-2.0