Responsibility and the human in the loop
Be able to design where a human makes the decision and how it is documented.
Prerequisites
- EThe EU AI Actrequired
Intuition
«A human in the loop» is not a state but a scale, and that difference is decisive:
| Degree | Means | Is the human responsible? |
|---|---|---|
| Human-in-the-loop | the human makes the decision, the AI suggests | yes |
| Human-on-the-loop | the AI decides, the human monitors and can stop it | partly |
| Human-in-command | the human sets the frame and reviews afterwards | at the system level |
| Full automation | no human in the decision | the organisation, not an individual |
Writing «a human reviews it» in a policy is not enough. The review has to be meaningful, and that sets four requirements:
- The reviewer has time to actually review.
- The reviewer has enough information to be able to deviate.
- The reviewer has the authority to say no.
- Saying no is no more expensive for the reviewer than saying yes.
If any of the four is missing, the review is a formality, and the responsibility has in practice moved to the model.
Formal
Rubber-stamping is the most common failure: the reviewer approves nearly everything, because the AI is usually right and time is short. The human is then there legally but not practically.
Countermeasures that actually work:
| Countermeasure | Effect |
|---|---|
| Measure the deviation rate | if it is below a per cent or so, it is not really being reviewed |
| Require a justification on approval, not just on rejection | forces a deliberation |
| Show the model's uncertainty and the underlying material | gives something to review |
| Rotate and spot-check the reviewers | detects the drift |
| Plant test cases with known errors | measures whether the review catches them |
The last one is the most revealing and is used in safety-critical industries: at regular intervals submit a case where the AI is obviously wrong and measure how often it is caught.
Automation bias is the underlying mechanism: people systematically trust an automated suggestion more than their own judgement, especially under time pressure and when the system is usually right. It is well documented in aviation and medicine, and it does not go away because you are aware of it.
Where the human should be placed depends on the consequence:
| Consequence | Placement |
|---|---|
| Irreversible and large (a rejection, a dismissal, medical) | in-the-loop — the human decides |
| Reversible but significant | on-the-loop with the ability to undo |
| Small and reversible | in-command — spot checks afterwards |
Documentation. For every decision that affects somebody, the following has to be reconstructable:
| Field | Why |
|---|---|
| The model version and configuration | what actually ran |
| The input (or its hash) | what the decision was built on |
| The model's output and confidence | what was suggested |
| Who decided, and when | the responsibility |
| Whether and why the human deviated | the most valuable data in the whole log |
| The outcome, when it becomes known | to be able to measure whether the system works |
The second to last row is worth gold: the deviations are exactly the cases where the model and a knowledgeable human disagree, and they are the best source of improvement there is.
Code
import hashlib, json, time
from dataclasses import dataclass, asdict
from enum import Enum
class Degree(Enum):
IN_THE_LOOP = "the human decides"
ON_THE_LOOP = "the AI decides, the human can stop it"
IN_COMMAND = "spot checks afterwards"
def choose_degree(irreversible: bool, significant: bool) -> Degree:
if irreversible and significant:
return Degree.IN_THE_LOOP
if significant:
return Degree.ON_THE_LOOP
return Degree.IN_COMMAND
@dataclass
class Decision:
case_id: str
model_version: str
input_hash: str
model_suggestion: str
model_confidence: float
decision: str
decided_by: str
deviated: bool
justification: str # REQUIRED even on approval
timestamp: float
outcome: str | None = None
def record(case, suggestion, confidence, decision, reviewer, justification, model_version):
if not justification or len(justification) < 15:
raise ValueError("a justification is required even when the suggestion is approved")
return Decision(
case_id=case["id"],
model_version=model_version,
input_hash=hashlib.sha256(
json.dumps(case, sort_keys=True).encode()).hexdigest()[:16],
model_suggestion=suggestion, model_confidence=confidence,
decision=decision, decided_by=reviewer,
deviated=(decision != suggestion), justification=justification, timestamp=time.time())
def review_quality(decisions: list[Decision]) -> dict:
n = len(decisions)
deviated = sum(d.deviated for d in decisions)
per_reviewer = {}
for d in decisions:
e = per_reviewer.setdefault(d.decided_by, {"n": 0, "deviated": 0})
e["n"] += 1; e["deviated"] += int(d.deviated)
return {
"deviation_rate": round(deviated / max(n, 1), 4),
"warning": "below 1 % — probably not really being reviewed" if deviated / max(n, 1) < 0.01 else "ok",
"per_reviewer": {k: round(v["deviated"] / v["n"], 3) for k, v in per_reviewer.items()},
}
# Planted test cases: measure whether the review actually catches errors
def plant_test_cases(queue, test_cases, share=0.02):
"""Mix in cases where the model's suggestion is known to be wrong."""
import random
out = list(queue)
for t in test_cases:
if random.random() < share:
out.insert(random.randrange(len(out) + 1), t | {"_test": True})
return out
def catch_rate(decisions, test_case_ids):
t = [d for d in decisions if d.case_id in test_case_ids]
return round(sum(d.deviated for d in t) / max(len(t), 1), 3)
# < 0.8 means the review is missing errors it ought to catch
Mastery means
- Distinguishes the degrees of human involvement
- Places the human where they do some good
- Documents the decisions traceably
Sign in to do the exercises and build your mastery up.
Sources
- EU AI Act (2024/1689) — EU legal act
- Google — People + AI Guidebook — free to read
- GDPR — Regulation (EU) 2016/679 — EU legal act