Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 300 min· fast-moving, sources checked often· verified 2026-09-20· EN

Project F: fine-tune and evaluate a language model

Be able to fine-tune a model with LoRA on your own dataset, measure against the base model and document it.

Prerequisites

Intuition

Project F: take an open model, adapt it to a task you care about, and show with numbers that it got better — without getting worse at anything else.

The deliverables:

The partThe requirement
The dataset≥ 500 pairs, cleaned (duplicates, eval leakage), balanced, documented
The eval≥ 50 cases measuring the target task, built before the training
The durability suite≥ 50 cases for general abilities that are to be preserved
The baselinesthe base model as it is, plus a prompted variant (few-shot)
The trainingLoRA/QLoRA, the configuration logged, ≥ 2 seeds
The reportbefore and after on both suites, with the spread; the limitations

The requirement that makes the project honest: the eval suite is built before the training. Otherwise it is unconsciously adapted to what the model happened to become good at.

Interactive

The result table you are to produce:

The variantThe target evalDurabilityTrainable par.GPU time
base, zero-shot0.41 ± 0.000.78 ± 0.00——
base, few-shot0.58 ± 0.010.78 ± 0.00——
LoRA r=80.79 ± 0.020.75 ± 0.0115 M25 min
LoRA r=320.82 ± 0.020.71 ± 0.0261 M31 min

Read it: r=32 gives +3 pp on the target but −4 pp on the durability. Is that the right choice? It depends on the use — and that decision is the point of the project, not the highest figure.

Common mistakes:

  1. The few-shot baseline is skipped — and then nobody knows whether the fine-tuning was even needed.
  2. Only one seed.
  3. The durability suite is missing.
  4. The chat template differs between the training and the evaluation (a silent loss of quality).
  5. The eval cases written afterwards.

The timetable (about 5 h of active time): 1.5 h the dataset · 1 h the eval suites · 1 h the training (a sweep) · 1 h the evaluation · 0.5 h the report.

Mastery means

  • Fine-tunes a model with LoRA on their own dataset
  • Measures against the base model with their own eval and a durability suite
  • Documents it reproducibly

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences