Target modules and rank in LoRA
Be able to choose which layers get adapters and which rank, and measure the effect.
Prerequisites
- FLoRA — Low-Rank Adaptationrequired
Intuition
Two choices decide the LoRA result: which matrices get adapters and which rank.
The target modules. The original paper put adapters only on q_proj and v_proj. Later work (QLoRA among others) shows that all the linear layers — including the MLP parts gate_proj, up_proj, down_proj — give better results when the budget allows. The MLP layers are roughly two thirds of the parameters; skipping them is leaving a lot unused.
The rank. r = 4–8 is enough for style and format. r = 16–64 for harder adaptation. A higher rank mainly helps when the task requires new knowledge, not just new behaviour.
Alpha. The scaling is α/r. Keep α ≈ 2r as a starting point, and you avoid having to adjust the learning rate when you change r.
Code
from peft import LoraConfig, get_peft_model
ALL_LINEAR = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
cfg = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=ALL_LINEAR, # not just q and v
task_type="CAUSAL_LM",
)
m = get_peft_model(base, cfg)
m.print_trainable_parameters()
# trainable params: 40,370,176 || all params: 7,281,... || trainable%: 0.55
# Counting the parameters by hand: r · (d_in + d_out) per matrix
def lora_params(layers=32, d=4096, d_ffn=11008, r=16, modules=ALL_LINEAR):
per_layer = 0
for name in modules:
d_in, d_out = (d, d) if name in ("q_proj", "k_proj", "v_proj", "o_proj") else \
((d, d_ffn) if name in ("gate_proj", "up_proj") else (d_ffn, d))
per_layer += r * (d_in + d_out)
return per_layer * layers
for r in (4, 16, 64):
print(r, f"{lora_params(r=r):,}")
# 4 15,335,424
# 16 61,341,696
# 64 245,366,784
How to choose without guessing: run a small sweep — r ∈ {8, 16, 32} × {attention only, all linear} — on a subset of the data, with three seeds, and measure on your own eval. It takes a few hours and is the only method that gives an answer for your particular task.
A diagnostic if you have access to a full fine-tuning: do an SVD on ΔW and look at the singular value spectrum. If it falls steeply after k values then r ≈ k is enough.
Mastery means
- Chooses the target modules deliberately
- Chooses the rank and the alpha with a justification
- Measures the effect of the choices instead of guessing
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — LoRA: Low-Rank Adaptation of Large Language Models — arXiv (open access; licence per article)
- arXiv — QLoRA: Efficient Finetuning of Quantized LLMs — arXiv (open access; licence per article)