Project F: fine-tune and evaluate a language model
Be able to fine-tune a model with LoRA on your own dataset, measure against the base model and document it.
Prerequisites
- FDataset design for fine-tuningrequired
- FRegression tests for modelsrequired
- FTarget modules and rank in LoRArequired
Intuition
Project F: take an open model, adapt it to a task you care about, and show with numbers that it got better — without getting worse at anything else.
The deliverables:
| The part | The requirement |
|---|---|
| The dataset | ≥ 500 pairs, cleaned (duplicates, eval leakage), balanced, documented |
| The eval | ≥ 50 cases measuring the target task, built before the training |
| The durability suite | ≥ 50 cases for general abilities that are to be preserved |
| The baselines | the base model as it is, plus a prompted variant (few-shot) |
| The training | LoRA/QLoRA, the configuration logged, ≥ 2 seeds |
| The report | before and after on both suites, with the spread; the limitations |
The requirement that makes the project honest: the eval suite is built before the training. Otherwise it is unconsciously adapted to what the model happened to become good at.
Interactive
The result table you are to produce:
| The variant | The target eval | Durability | Trainable par. | GPU time |
|---|---|---|---|---|
| base, zero-shot | 0.41 ± 0.00 | 0.78 ± 0.00 | — | — |
| base, few-shot | 0.58 ± 0.01 | 0.78 ± 0.00 | — | — |
| LoRA r=8 | 0.79 ± 0.02 | 0.75 ± 0.01 | 15 M | 25 min |
| LoRA r=32 | 0.82 ± 0.02 | 0.71 ± 0.02 | 61 M | 31 min |
Read it: r=32 gives +3 pp on the target but −4 pp on the durability. Is that the right choice? It depends on the use — and that decision is the point of the project, not the highest figure.
Common mistakes:
- The few-shot baseline is skipped — and then nobody knows whether the fine-tuning was even needed.
- Only one seed.
- The durability suite is missing.
- The chat template differs between the training and the evaluation (a silent loss of quality).
- The eval cases written afterwards.
The timetable (about 5 h of active time): 1.5 h the dataset · 1 h the eval suites · 1 h the training (a sweep) · 1 h the evaluation · 0.5 h the report.
Mastery means
- Fine-tunes a model with LoRA on their own dataset
- Measures against the base model with their own eval and a durability suite
- Documents it reproducibly
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — QLoRA: Efficient Finetuning of Quantized LLMs — arXiv (open access; licence per article)
- arXiv — LIMA: Less Is More for Alignment — arXiv (open access; licence per article)
- PEFT — dokumentation (Apache-2.0) — Apache-2.0