Skip to content
AI-grafen
EUniversityLanguage models· about 240 min· evolving, reviewed regularly· verified 2026-09-20· EN

Project E: an NLP system end to end

Be able to build, evaluate and serve an NLP system with tests and evals.

Prerequisites

Intuition

Project E: a complete, runnable NLP system. A suggestion: classify incoming support tickets (category plus urgency) and expose it as an API.

Deliverables:

  1. Data: ≥ 800 labelled examples (your own, synthetic with review, or an open dataset), a guideline, κ on a sample, a stratified train/val/test split.
  2. Baselines: majority class plus TF-IDF/LR, tuned.
  3. Model: a fine-tuned encoder; macro-F1 with a bootstrap CI; a confusion matrix; an error analysis of 30 cases.
  4. Service: FastAPI with validation, /healthz, a 503 degradation (fall back to the TF-IDF model when the transformer is missing!), logging without the text.
  5. CI: tests plus the eval harness with a threshold and a regression list.
  6. Report: the hypothesis, the method, the results with uncertainty, the limitations, the reproducibility information (commit, data hash, seeds).

Interactive

An assessment matrix (apply it to yourself before you hand in):

PartPassStrong
Dataa split without leakageκ reported, edge cases in the guideline
Baselineexiststuned, with a CI
Modelbeats the baselinebeats it outside the CI, the error analysis leads to action
APIanswers correctlyvalidation, the 503 fallback tested
CItests greenan eval threshold plus regressions block the merge
Reportcompletenegative results and limitations reported

The most common shortcoming: the report has no uncertainty and the baseline is untuned. The second most common: the fallback exists in the code but has never been run.

Mastery means

  • Builds an NLP system end to end: data → fine-tuned model → API
  • Has tests, an eval harness in CI and a degraded mode
  • Reports the results with a baseline and uncertainty

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences