MLOps: deployment, versions, rollback
Be able to deploy a model version, compare it against the previous one and roll back.
Prerequisites
- EData versioningrequired
- EDocker — containersrequired
Intuition
A model in production is not a filename for an artefact but a version with a lifecycle.
Deployment strategies:
| Strategy | How | Risk |
|---|---|---|
| Big bang | swap everything at once | high |
| Shadow | the new model runs in parallel, the answers are not used | none — but no user signal |
| Canary | 1 % → 5 % → 25 % → 100 % | low, gradual |
| Blue-green | two full environments, swap the alias | low, a fast return |
| A/B | split traffic, measure the difference | low, gives causality |
Shadow first, then canary is the most common combination: the shadow shows that the new model does not crash and how its answers differ, the canary shows how the users actually react.
The most important requirement: the return should take seconds, not a redeploy. If a rollback is an alias swap you dare to deploy more often — and that in itself is the biggest safety improvement.
Formal
What has to be versioned together. A «model version» that is only the weight file is not enough:
| Part | Why |
|---|---|
| The weights | obviously |
| The preprocessing | the tokenizer, the scaling, the category order |
| The postprocessing | the thresholds, the formatting |
| The configuration | the temperature, the max tokens, the system prompt |
| The dependencies | the library versions |
| The data hash and the code commit | traceability to how it was trained |
Swapping only the weight file is a classic error: a new model with old preprocessing gives silently worse results, since nothing breaks.
Training/serving skew is the same phenomenon in another form: the preprocessing in training and in production is implemented in two different places and drifts apart. The only robust remedy is to share the code — the same function is used in both places, tested against the same fixture.
What is monitored in production:
| Category | Metric |
|---|---|
| Technical | latency p50/p95, error rate, throughput |
| Input | the feature distributions against the training data (drift) |
| Output | the prediction distribution, the share declined, the confidence |
| Quality | accuracy where ground truth arrives, often with a delay |
| Business | conversion, completed tasks |
| Cost | kronor per call and per user |
The output distribution is the best early warning. Quality metrics often require ground truth that takes days or weeks; that the share of positive predictions suddenly goes from 4 % to 11 % is noticed the same day.
An automatic rollback should be triggered by clear threshold breaches:
| Condition | Action |
|---|---|
| The error rate > 2× the baseline | roll back |
| The p95 latency > 2× the baseline | roll back |
| The output distribution deviates sharply | alert, and roll back on continued deviation |
| A quality metric falls below a floor | roll back |
A model registry. Every version has a stage: staging → production → archived. The transitions are logged with who decided and why. That is the same traceability requirement the AI Act places on high-risk systems, so it is work that has to be done anyway.
Code
import json, time
from dataclasses import dataclass, asdict, field
from pathlib import Path
@dataclass
class ModelVersion:
version: str
weights_sha: str
preprocessing_sha: str
config: dict
data_hash: str
git_commit: str
dependencies: dict
stage: str = "staging" # staging | production | archived
created: float = field(default_factory=time.time)
class Registry:
def __init__(self, path="registry.jsonl"):
self.path = Path(path)
def record(self, mv: ModelVersion, by: str, reason: str):
with self.path.open("a", encoding="utf-8") as f:
f.write(json.dumps({**asdict(mv), "by": by, "reason": reason},
ensure_ascii=False) + "\n")
def promote(self, version: str, to: str, by: str, reason: str):
self.record(ModelVersion(version=version, weights_sha="", preprocessing_sha="",
config={}, data_hash="", git_commit="", dependencies={},
stage=to), by, reason)
# Shadow: run the new model in parallel, do not use the answers
async def shadow(request, production, candidate, log):
answer = await production(request)
try:
candidate_answer = await candidate(request)
log.append({"equal": answer == candidate_answer,
"prod": answer, "candidate": candidate_answer})
except Exception as e:
log.append({"error": str(type(e).__name__)})
return answer # the user ALWAYS gets the production answer
# A canary with automatic rollback
class Canary:
def __init__(self, baseline, steps=(0.01, 0.05, 0.25, 1.0)):
self.baseline = baseline # {"error_rate": .., "p95_ms": .., "positive_share": ..}
self.steps = steps
self.index = 0
def share(self):
return self.steps[self.index]
def evaluate(self, current: dict) -> tuple[str, str]:
if current["error_rate"] > 2 * self.baseline["error_rate"]:
return "rollback", "the error rate has doubled"
if current["p95_ms"] > 2 * self.baseline["p95_ms"]:
return "rollback", "the p95 latency has doubled"
deviation = abs(current["positive_share"] - self.baseline["positive_share"])
if deviation > 0.5 * self.baseline["positive_share"]:
return "rollback", f"the output distribution deviates by {deviation:.3f}"
if self.index + 1 < len(self.steps):
self.index += 1
return "continue", f"raising to {self.steps[self.index]:.0%}"
return "done", "full deployment"
# A rollback should be an alias swap, not a redeploy
def roll_back(registry, to_version, by, reason):
registry.promote(to_version, "production", by, f"ROLLBACK: {reason}")
swap_alias("production", to_version) # seconds, not minutes
# Share the preprocessing code between training and serving
def preprocess(raw: dict) -> dict:
"""The ONLY implementation. Imported by both the training pipeline and the API."""
return {"age": float(raw["age"]),
"city": (raw.get("city") or "").strip().lower() or "<unknown>"}
def test_preprocessing_fixture():
"""The same fixture is run in both the training and the serving CI."""
assert preprocess({"age": "15", "city": " Malmö "}) == {"age": 15.0, "city": "malmö"}
assert preprocess({"age": 16, "city": None}) == {"age": 16.0, "city": "<unknown>"}
Mastery means
- Deploys a model version in a controlled way
- Compares it against the previous one in production
- Can roll back quickly
Sign in to do the exercises and build your mastery up.
Sources
- Google — Rules of Machine Learning — CC BY 4.0
- Sculley m.fl. — Hidden Technical Debt in Machine Learning Systems (NeurIPS 2015) — NeurIPS open access
- MLflow — dokumentation (Apache-2.0) — Apache-2.0