Observability for ML systems
Be able to measure latency, cost, quality and drift in production.
Prerequisites
- EFrom prototype to productrequired
Intuition
An ML service can be up and at the same time wrong. Hence two kinds of measurement:
System health (as for any service): the error rate, the latency p50/p95/p99, the throughput, the queue, the resource use.
Model health (the particular part):
| Metric | Catches |
|---|---|
| The prediction distribution | the model starts answering differently |
| The input distribution (drift) | the world has changed |
| Quality on a sample | human review of 1 % |
| The abstention rate / low confidence | the model becomes uncertain |
| The cost per call and per task solved | the budget |
| The share of fallbacks / degraded mode | the dependencies are playing up |
The point: a quality degradation shows in the distributions before it shows in complaints — if you measure them.
Code
import time, json, logging
log = logging.getLogger("ml")
def logged_prediction(model, x, request_id):
t0 = time.perf_counter()
p = model.predict_proba(x)[0]
latency = (time.perf_counter() - t0) * 1000
log.info(json.dumps({ # a structured log — no personal data
"request_id": request_id, "model": model.version,
"latency_ms": round(latency, 1), "confidence": round(float(p.max()), 3),
"class": int(p.argmax()), "features_hash": hash_features(x),
}))
return p
# Drift: compare the distribution over the last day with the reference period
from scipy.stats import ks_2samp
def drift_alert(reference, latest, alpha=0.01):
stat, p = ks_2samp(reference, latest)
return {"drift": p < alpha, "p": round(float(p), 5), "ks": round(float(stat), 3)}
Alerts that can be acted on — three rules: alert on symptoms the user notices (p95 latency, the error rate), not on every anomaly; every alert should have a runbook; and an alert that has not led to an action in three months should be removed. Alert fatigue is more dangerous than no alert.
Mastery means
- Measures latency, cost, quality and drift
- Sets alerts that can be acted on
- Tells system health from model health
Sign in to do the exercises and build your mastery up.
Sources
- Google SRE Book — Monitoring Distributed Systems — CC BY-NC-ND 4.0 (reference only)
- Google — Rules of Machine Learning — CC BY 4.0