Parameter-efficient fine-tuning: adapters, prefix, LoRA
Be able to compare PEFT methods and choose the right one for a given budget.
Prerequisites
- FLoRA — Low-Rank Adaptationrequired
Intuition
PEFT = train few parameters instead of all of them. The family has four main branches:
| The method | Where it sits | The parameters | The inference cost |
|---|---|---|---|
| LoRA | parallel to the weight matrices, ΔW = BA | 0.1–2 % | zero after a merge |
| Adapters | small bottleneck layers inserted in the blocks | 0.5–5 % | the extra layers remain |
| Prefix/prompt tuning | learnt «virtual tokens» first in the sequence | < 0.1 % | it eats context |
| (IA)³ | learnt scaling vectors on K, V and the FFN | ~0.01 % | near zero |
| BitFit | the bias terms only | ~0.08 % | zero |
LoRA won in practice for a single reason: the adapter can be baked into the base weights after training, so the inference becomes exactly as fast as the original. Adapters add layers that have to be run every time.
Formal
Why PEFT works at all: fine-tuning changes the model in a low-dimensional subspace. Aghajanyan et al. (2020) measured the «intrinsic dimension» of fine-tuning tasks and found that a few hundred to a few thousand parameters are enough to come close to a full fine-tuning on many tasks.
The memory budget is what decides in practice. A full fine-tuning of a 7B model in bf16 requires:
- the weights 14 GB + the gradients 14 GB + the Adam state 56 GB (fp32 m and v) ≈ 84 GB plus the activations.
With LoRA r=16 ~0.1 % of the parameters are trained:
- the weights 14 GB (frozen, no gradients) + the adapters and their optimiser state < 0.5 GB ≈ 15 GB.
With QLoRA (a 4-bit base) ≈ 6 GB — a 7B model fine-tuned on a consumer card.
Choosing in practice:
- Style, format, tone, domain terminology → LoRA with a low rank (4–16).
- New factual knowledge → a high rank or a full fine-tuning — or, usually better, RAG.
- Many tasks on the same base → LoRA adapters swapped in operation (one base, n adapters).
- Extremely memory-thrifty → (IA)³ or BitFit, at the price of a lower ceiling.
Mastery means
- Compares LoRA, adapters, prefix/prompt tuning and (IA)³
- Chooses a method to suit the memory budget and the task
- Knows what can be baked into the base model
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — LoRA: Low-Rank Adaptation of Large Language Models — arXiv (open access; licence per article)
- arXiv — Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning — arXiv (open access; licence per article)
- PEFT — dokumentation (Apache-2.0) — Apache-2.0