Skip to content
AI-grafen
FAI engineeringAI product development· about 90 min· fast-moving, sources checked often· verified 2026-09-21· EN

MLOps: deployment, versions, rollback

Be able to deploy a model version, compare it against the previous one and roll back.

Prerequisites

Intuition

A model in production is not a filename for an artefact but a version with a lifecycle.

Deployment strategies:

StrategyHowRisk
Big bangswap everything at oncehigh
Shadowthe new model runs in parallel, the answers are not usednone — but no user signal
Canary1 % → 5 % → 25 % → 100 %low, gradual
Blue-greentwo full environments, swap the aliaslow, a fast return
A/Bsplit traffic, measure the differencelow, gives causality

Shadow first, then canary is the most common combination: the shadow shows that the new model does not crash and how its answers differ, the canary shows how the users actually react.

The most important requirement: the return should take seconds, not a redeploy. If a rollback is an alias swap you dare to deploy more often — and that in itself is the biggest safety improvement.

Formal

What has to be versioned together. A «model version» that is only the weight file is not enough:

PartWhy
The weightsobviously
The preprocessingthe tokenizer, the scaling, the category order
The postprocessingthe thresholds, the formatting
The configurationthe temperature, the max tokens, the system prompt
The dependenciesthe library versions
The data hash and the code committraceability to how it was trained

Swapping only the weight file is a classic error: a new model with old preprocessing gives silently worse results, since nothing breaks.

Training/serving skew is the same phenomenon in another form: the preprocessing in training and in production is implemented in two different places and drifts apart. The only robust remedy is to share the code — the same function is used in both places, tested against the same fixture.

What is monitored in production:

CategoryMetric
Technicallatency p50/p95, error rate, throughput
Inputthe feature distributions against the training data (drift)
Outputthe prediction distribution, the share declined, the confidence
Qualityaccuracy where ground truth arrives, often with a delay
Businessconversion, completed tasks
Costkronor per call and per user

The output distribution is the best early warning. Quality metrics often require ground truth that takes days or weeks; that the share of positive predictions suddenly goes from 4 % to 11 % is noticed the same day.

An automatic rollback should be triggered by clear threshold breaches:

ConditionAction
The error rate > 2× the baselineroll back
The p95 latency > 2× the baselineroll back
The output distribution deviates sharplyalert, and roll back on continued deviation
A quality metric falls below a floorroll back

A model registry. Every version has a stage: staging → production → archived. The transitions are logged with who decided and why. That is the same traceability requirement the AI Act places on high-risk systems, so it is work that has to be done anyway.

Code

import json, time
from dataclasses import dataclass, asdict, field
from pathlib import Path

@dataclass
class ModelVersion:
    version: str
    weights_sha: str
    preprocessing_sha: str
    config: dict
    data_hash: str
    git_commit: str
    dependencies: dict
    stage: str = "staging"            # staging | production | archived
    created: float = field(default_factory=time.time)

class Registry:
    def __init__(self, path="registry.jsonl"):
        self.path = Path(path)

    def record(self, mv: ModelVersion, by: str, reason: str):
        with self.path.open("a", encoding="utf-8") as f:
            f.write(json.dumps({**asdict(mv), "by": by, "reason": reason},
                               ensure_ascii=False) + "\n")

    def promote(self, version: str, to: str, by: str, reason: str):
        self.record(ModelVersion(version=version, weights_sha="", preprocessing_sha="",
                                 config={}, data_hash="", git_commit="", dependencies={},
                                 stage=to), by, reason)

# Shadow: run the new model in parallel, do not use the answers
async def shadow(request, production, candidate, log):
    answer = await production(request)
    try:
        candidate_answer = await candidate(request)
        log.append({"equal": answer == candidate_answer,
                    "prod": answer, "candidate": candidate_answer})
    except Exception as e:
        log.append({"error": str(type(e).__name__)})
    return answer                       # the user ALWAYS gets the production answer

# A canary with automatic rollback
class Canary:
    def __init__(self, baseline, steps=(0.01, 0.05, 0.25, 1.0)):
        self.baseline = baseline        # {"error_rate": .., "p95_ms": .., "positive_share": ..}
        self.steps = steps
        self.index = 0

    def share(self):
        return self.steps[self.index]

    def evaluate(self, current: dict) -> tuple[str, str]:
        if current["error_rate"] > 2 * self.baseline["error_rate"]:
            return "rollback", "the error rate has doubled"
        if current["p95_ms"] > 2 * self.baseline["p95_ms"]:
            return "rollback", "the p95 latency has doubled"
        deviation = abs(current["positive_share"] - self.baseline["positive_share"])
        if deviation > 0.5 * self.baseline["positive_share"]:
            return "rollback", f"the output distribution deviates by {deviation:.3f}"
        if self.index + 1 < len(self.steps):
            self.index += 1
            return "continue", f"raising to {self.steps[self.index]:.0%}"
        return "done", "full deployment"

# A rollback should be an alias swap, not a redeploy
def roll_back(registry, to_version, by, reason):
    registry.promote(to_version, "production", by, f"ROLLBACK: {reason}")
    swap_alias("production", to_version)         # seconds, not minutes

# Share the preprocessing code between training and serving
def preprocess(raw: dict) -> dict:
    """The ONLY implementation. Imported by both the training pipeline and the API."""
    return {"age": float(raw["age"]),
            "city": (raw.get("city") or "").strip().lower() or "<unknown>"}

def test_preprocessing_fixture():
    """The same fixture is run in both the training and the serving CI."""
    assert preprocess({"age": "15", "city": " Malmö "}) == {"age": 15.0, "city": "malmö"}
    assert preprocess({"age": 16, "city": None}) == {"age": 16.0, "city": "<unknown>"}

Mastery means

  • Deploys a model version in a controlled way
  • Compares it against the previous one in production
  • Can roll back quickly

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences