Skip to content
AI-grafen
EUniversityAI product development· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Observability for ML systems

Be able to measure latency, cost, quality and drift in production.

Prerequisites

Intuition

An ML service can be up and at the same time wrong. Hence two kinds of measurement:

System health (as for any service): the error rate, the latency p50/p95/p99, the throughput, the queue, the resource use.

Model health (the particular part):

MetricCatches
The prediction distributionthe model starts answering differently
The input distribution (drift)the world has changed
Quality on a samplehuman review of 1 %
The abstention rate / low confidencethe model becomes uncertain
The cost per call and per task solvedthe budget
The share of fallbacks / degraded modethe dependencies are playing up

The point: a quality degradation shows in the distributions before it shows in complaints — if you measure them.

Code

import time, json, logging
log = logging.getLogger("ml")

def logged_prediction(model, x, request_id):
    t0 = time.perf_counter()
    p = model.predict_proba(x)[0]
    latency = (time.perf_counter() - t0) * 1000
    log.info(json.dumps({                    # a structured log — no personal data
        "request_id": request_id, "model": model.version,
        "latency_ms": round(latency, 1), "confidence": round(float(p.max()), 3),
        "class": int(p.argmax()), "features_hash": hash_features(x),
    }))
    return p

# Drift: compare the distribution over the last day with the reference period
from scipy.stats import ks_2samp
def drift_alert(reference, latest, alpha=0.01):
    stat, p = ks_2samp(reference, latest)
    return {"drift": p < alpha, "p": round(float(p), 5), "ks": round(float(stat), 3)}

Alerts that can be acted on — three rules: alert on symptoms the user notices (p95 latency, the error rate), not on every anomaly; every alert should have a runbook; and an alert that has not led to an action in three months should be removed. Alert fatigue is more dangerous than no alert.

Mastery means

  • Measures latency, cost, quality and drift
  • Sets alerts that can be acted on
  • Tells system health from model health

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences