EUniversityModel training and fine-tuning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Experiment tracking
Be able to log runs with their parameters and metrics and compare them systematically.
Prerequisites
- EReproducibilityrequired
Intuition
After 40 runs with different lr, batch, data and code changes nobody remembers which one gave 0.87. Experiment tracking = every run automatically logs:
- the configuration (all the hyperparameters, the model name, the dataset + version),
- the code version (the git commit, and whether the working tree was dirty),
- the metrics per epoch (loss, val F1 …),
- the artefacts (the model weights, the confusion matrix, examples of errors),
- the environment (package versions, GPU, seed).
Tools: MLflow, Weights & Biases, or a simple JSONL file + git. The tool is unimportant; the discipline of never running anything that is not logged is everything.
Code
import json, subprocess, time, uuid
from pathlib import Path
class Run:
def __init__(self, config: dict, log_dir="runs"):
self.id = time.strftime("%Y%m%d-%H%M%S-") + uuid.uuid4().hex[:6]
self.dir = Path(log_dir) / self.id; self.dir.mkdir(parents=True)
git = subprocess.run(["git", "rev-parse", "HEAD"], capture_output=True, text=True).stdout.strip()
dirty = bool(subprocess.run(["git", "status", "--porcelain"], capture_output=True, text=True).stdout)
(self.dir / "config.json").write_text(json.dumps({**config, "git": git, "dirty": dirty}, indent=2))
self.f = (self.dir / "metrics.jsonl").open("a")
def log(self, step: int, **m):
self.f.write(json.dumps({"step": step, "t": time.time(), **m}) + "\n"); self.f.flush()
run = Run({"lr": 2e-5, "epochs": 3, "data": "support-v3", "seed": 0})
for epoch in range(3):
run.log(epoch, train_loss=0.5 / (epoch + 1), val_f1=0.7 + 0.05 * epoch)
# comparing: read all runs/*/config.json + the last metrics line → a table (pandas)
With MLflow: mlflow.log_params(config); mlflow.log_metric("val_f1", f1, step=epoch); mlflow.log_artifact("cm.png").
Mastery means
- Logs runs with parameters, metrics and artefacts
- Compares runs systematically
- Ties every result to a code commit and a data version
Sign in to do the exercises and build your mastery up.
Sources
- MLflow — dokumentation (Apache-2.0) — Apache-2.0