Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

LoRA — Low-Rank Adaptation

Be able to explain why LoRA works (ΔW = BA), choose the target modules and the rank, fine-tune a model with LoRA and evaluate it with evals.

Prerequisites

Intuition

A full fine-tune of a 7B model updates 7 billion weights and needs optimizer state for all of them — ~100 GB. LoRA freezes the model and learns only a small low-rank update per matrix: W′ = W + ΔW with ΔW = B·A, where A is r × d and B is d × r, r ≪ d (8–64, say). For d = 4096, r = 16: 131 k parameters instead of 16.8 M per matrix.

Why does it work? Fine-tuning changes the model in few directions — ΔW has a low «intrinsic rank». The hypothesis is confirmed empirically: LoRA gets close to a full fine-tune on most tasks.

In practice: the target modules = which matrices (q, k, v, o, and often the MLP layers too — the latter helps most when you have the budget). α/r scales the update; keep α ≈ 2r as a start. After training, BA can be merged into W — zero extra inference cost. QLoRA: the base model in 4-bit, LoRA in bf16 → a 7B trains on an 8–10 GB GPU.

Measure against the base model and against a full fine-tune if you can; LoRA with too low an r underperforms on tasks that require new knowledge.

Code

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
import torch

name = "Qwen/Qwen2.5-1.5B"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16)
m = AutoModelForCausalLM.from_pretrained(name, quantization_config=bnb, device_map="auto")
m = prepare_model_for_kbit_training(m)

cfg = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
                 target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"])
m = get_peft_model(m, cfg)
m.print_trainable_parameters()      # ~1 % of the parameters

# …an ordinary training loop / Trainer on instruction data (loss on the answer only), lr ~2e-4, 1–3 epochs…
m.save_pretrained("adapter/")       # just A and B — a few tens of MB

# inference: load the base plus the adapter, or m.merge_and_unload() to bake it in

The lab lora-tiny-labb implements the LoRA layer from scratch in NumPy/PyTorch and verifies it against a reference.

Derivation

Forward: h=Wx+αrBAxh = Wx + \tfrac{\alpha}{r}BAx. Parameters: r(din+dout)r(d_{in}+d_{out}) against dindoutd_{in}d_{out}. Initialisation: A∼N(0,σ2)A\sim\mathcal N(0,\sigma^2), B=0B = 0 → ΔW=0\Delta W = 0 at the start, so the model begins exactly as the base does. Gradients: ∂L/∂B=αr δ (Ax)⊤\partial L/\partial B = \tfrac{\alpha}{r}\,\delta\,(Ax)^\top, ∂L/∂A=αr B⊤δ x⊤\partial L/\partial A = \tfrac{\alpha}{r}\,B^\top\delta\,x^\top where δ=∂L/∂h\delta = \partial L/\partial h — the same δ\delta as a full fine-tune, but projected. The memory saving comes mainly from the optimizer state (Adam: 2 × the parameters in fp32) only being needed for A,BA, B. Choosing the rank: the effective rank of ΔW\Delta W in a full fine-tune can be measured with an SVD; empirically r∈[4,64]r\in[4, 64] is enough for style and format, more for new factual knowledge. The α/r\alpha/r scaling means the learning rate does not have to be adjusted when rr changes (rsLoRA proposes α/r\alpha/\sqrt r).

Mastery means

  • Explains ΔW = BA, the rank and the α/r scaling
  • Chooses the target modules and the rank with justification
  • Fine-tunes with LoRA/QLoRA and evaluates against the base model and a full fine-tune

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences