Project: build an AI service end to end
Be able to deliver a working AI service with retrieval, evals, quotas and observability.
Prerequisites
- EFrom prototype to productrequired
- EObservability for ML systemsrequired
- ERAG — retrieval-augmented generationrequired
Intuition
The project: a working AI service that somebody else can run and trust. A suggestion: a question service over a body of documents you have the right to use.
Deliverables:
| Part | Requirement |
|---|---|
| Retrieval | the chunking justified, hybrid or vector, recall@10 measured on ≥ 40 real questions |
| Generation | grounding with citations, «I do not know» behaviour, a structured answer |
| The API | validation, /healthz, error codes, a timeout, a quota per user |
| Degradation | a defined mode when the LLM or the vector DB is down — and tested |
| Evals | ≥ 30 cases in CI with a threshold and a regression list |
| Observability | a structured log, the p95 latency, the cost per call, the fallback share |
| Documentation | a README, an architecture sketch, the limitations, the running cost |
What separates a pass from a strong result is the last three rows — most people build the retrieval and the API and stop there.
Interactive
An assessment matrix — use it on yourself:
| Part | Pass | Strong |
|---|---|---|
| Retrieval | it works | the recall measured, the chunking strategy compared against alternatives |
| Grounding | citations exist | the grounding rate measured with a judge, «I do not know» works |
| The API | it answers | quotas, 503 degradation tested in CI |
| Evals | they exist | they block a merge, the regression list is reviewed |
| Observability | it logs | p95, the cost per call and drift alerts |
| The report | it describes | figures with uncertainty and honest limitations |
The most common shortcomings, in order: the degradation exists in the code but has never been run; the eval suite is written but is not run in CI; the cost per call is unknown; and the test questions were made up by whoever built the system.
A schedule (about 5 hours of active time, spread out): 1 h data and chunking · 1 h retrieval and measurement · 1 h the API and the degradation · 1 h evals in CI · 1 h observability and the report.
Mastery means
- Delivers a running AI service with retrieval and evals
- Has quotas, a degraded mode and observability
- Reports the measurements and the limitations honestly
Sign in to do the exercises and build your mastery up.
Sources
- FastAPI — dokumentation (MIT) — MIT
- arXiv — RAGAS: Automated Evaluation of RAG — arXiv (open access; licence per article)
- Google — Rules of Machine Learning — CC BY 4.0