Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Target modules and rank in LoRA

Be able to choose which layers get adapters and which rank, and measure the effect.

Prerequisites

Intuition

Two choices decide the LoRA result: which matrices get adapters and which rank.

The target modules. The original paper put adapters only on q_proj and v_proj. Later work (QLoRA among others) shows that all the linear layers — including the MLP parts gate_proj, up_proj, down_proj — give better results when the budget allows. The MLP layers are roughly two thirds of the parameters; skipping them is leaving a lot unused.

The rank. r = 4–8 is enough for style and format. r = 16–64 for harder adaptation. A higher rank mainly helps when the task requires new knowledge, not just new behaviour.

Alpha. The scaling is α/r. Keep α ≈ 2r as a starting point, and you avoid having to adjust the learning rate when you change r.

Code

from peft import LoraConfig, get_peft_model

ALL_LINEAR = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]

cfg = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05,
    target_modules=ALL_LINEAR,          # not just q and v
    task_type="CAUSAL_LM",
)
m = get_peft_model(base, cfg)
m.print_trainable_parameters()
# trainable params: 40,370,176 || all params: 7,281,... || trainable%: 0.55

# Counting the parameters by hand: r · (d_in + d_out) per matrix
def lora_params(layers=32, d=4096, d_ffn=11008, r=16, modules=ALL_LINEAR):
    per_layer = 0
    for name in modules:
        d_in, d_out = (d, d) if name in ("q_proj", "k_proj", "v_proj", "o_proj") else \
                      ((d, d_ffn) if name in ("gate_proj", "up_proj") else (d_ffn, d))
        per_layer += r * (d_in + d_out)
    return per_layer * layers

for r in (4, 16, 64):
    print(r, f"{lora_params(r=r):,}")
# 4   15,335,424
# 16  61,341,696
# 64  245,366,784

How to choose without guessing: run a small sweep — r ∈ {8, 16, 32} × {attention only, all linear} — on a subset of the data, with three seeds, and measure on your own eval. It takes a few hours and is the only method that gives an answer for your particular task.

A diagnostic if you have access to a full fine-tuning: do an SVD on ΔW and look at the singular value spectrum. If it falls steeply after k values then r ≈ k is enough.

Mastery means

  • Chooses the target modules deliberately
  • Chooses the rank and the alpha with a justification
  • Measures the effect of the choices instead of guessing

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences