Skip to content
AI-grafen
EUniversityEthics, law and society· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Responsibility and the human in the loop

Be able to design where a human makes the decision and how it is documented.

Prerequisites

Intuition

«A human in the loop» is not a state but a scale, and that difference is decisive:

DegreeMeansIs the human responsible?
Human-in-the-loopthe human makes the decision, the AI suggestsyes
Human-on-the-loopthe AI decides, the human monitors and can stop itpartly
Human-in-commandthe human sets the frame and reviews afterwardsat the system level
Full automationno human in the decisionthe organisation, not an individual

Writing «a human reviews it» in a policy is not enough. The review has to be meaningful, and that sets four requirements:

  1. The reviewer has time to actually review.
  2. The reviewer has enough information to be able to deviate.
  3. The reviewer has the authority to say no.
  4. Saying no is no more expensive for the reviewer than saying yes.

If any of the four is missing, the review is a formality, and the responsibility has in practice moved to the model.

Formal

Rubber-stamping is the most common failure: the reviewer approves nearly everything, because the AI is usually right and time is short. The human is then there legally but not practically.

Countermeasures that actually work:

CountermeasureEffect
Measure the deviation rateif it is below a per cent or so, it is not really being reviewed
Require a justification on approval, not just on rejectionforces a deliberation
Show the model's uncertainty and the underlying materialgives something to review
Rotate and spot-check the reviewersdetects the drift
Plant test cases with known errorsmeasures whether the review catches them

The last one is the most revealing and is used in safety-critical industries: at regular intervals submit a case where the AI is obviously wrong and measure how often it is caught.

Automation bias is the underlying mechanism: people systematically trust an automated suggestion more than their own judgement, especially under time pressure and when the system is usually right. It is well documented in aviation and medicine, and it does not go away because you are aware of it.

Where the human should be placed depends on the consequence:

ConsequencePlacement
Irreversible and large (a rejection, a dismissal, medical)in-the-loop — the human decides
Reversible but significanton-the-loop with the ability to undo
Small and reversiblein-command — spot checks afterwards

Documentation. For every decision that affects somebody, the following has to be reconstructable:

FieldWhy
The model version and configurationwhat actually ran
The input (or its hash)what the decision was built on
The model's output and confidencewhat was suggested
Who decided, and whenthe responsibility
Whether and why the human deviatedthe most valuable data in the whole log
The outcome, when it becomes knownto be able to measure whether the system works

The second to last row is worth gold: the deviations are exactly the cases where the model and a knowledgeable human disagree, and they are the best source of improvement there is.

Code

import hashlib, json, time
from dataclasses import dataclass, asdict
from enum import Enum

class Degree(Enum):
    IN_THE_LOOP = "the human decides"
    ON_THE_LOOP = "the AI decides, the human can stop it"
    IN_COMMAND = "spot checks afterwards"

def choose_degree(irreversible: bool, significant: bool) -> Degree:
    if irreversible and significant:
        return Degree.IN_THE_LOOP
    if significant:
        return Degree.ON_THE_LOOP
    return Degree.IN_COMMAND

@dataclass
class Decision:
    case_id: str
    model_version: str
    input_hash: str
    model_suggestion: str
    model_confidence: float
    decision: str
    decided_by: str
    deviated: bool
    justification: str         # REQUIRED even on approval
    timestamp: float
    outcome: str | None = None

def record(case, suggestion, confidence, decision, reviewer, justification, model_version):
    if not justification or len(justification) < 15:
        raise ValueError("a justification is required even when the suggestion is approved")
    return Decision(
        case_id=case["id"],
        model_version=model_version,
        input_hash=hashlib.sha256(
            json.dumps(case, sort_keys=True).encode()).hexdigest()[:16],
        model_suggestion=suggestion, model_confidence=confidence,
        decision=decision, decided_by=reviewer,
        deviated=(decision != suggestion), justification=justification, timestamp=time.time())

def review_quality(decisions: list[Decision]) -> dict:
    n = len(decisions)
    deviated = sum(d.deviated for d in decisions)
    per_reviewer = {}
    for d in decisions:
        e = per_reviewer.setdefault(d.decided_by, {"n": 0, "deviated": 0})
        e["n"] += 1; e["deviated"] += int(d.deviated)
    return {
        "deviation_rate": round(deviated / max(n, 1), 4),
        "warning": "below 1 % — probably not really being reviewed" if deviated / max(n, 1) < 0.01 else "ok",
        "per_reviewer": {k: round(v["deviated"] / v["n"], 3) for k, v in per_reviewer.items()},
    }

# Planted test cases: measure whether the review actually catches errors
def plant_test_cases(queue, test_cases, share=0.02):
    """Mix in cases where the model's suggestion is known to be wrong."""
    import random
    out = list(queue)
    for t in test_cases:
        if random.random() < share:
            out.insert(random.randrange(len(out) + 1), t | {"_test": True})
    return out

def catch_rate(decisions, test_case_ids):
    t = [d for d in decisions if d.case_id in test_case_ids]
    return round(sum(d.deviated for d in t) / max(len(t), 1), 3)
# < 0.8 means the review is missing errors it ought to catch

Mastery means

  • Distinguishes the degrees of human involvement
  • Places the human where they do some good
  • Documents the decisions traceably

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences