Product metrics for AI features
Be able to define metrics that capture value for the user, not just model quality.
Prerequisites
- EA/B tests and experiment designrequired
- EObservability for ML systemsrequired
Intuition
A model can get better without the product getting better — and the other way round. So two kinds of metric are needed:
| Model metrics | Product metrics | |
|---|---|---|
| Examples | accuracy, F1, the grounding rate, perplexity | tasks completed, the time to an answer, return visits, escalations |
| Changed by | the model | the whole experience |
| Measured in | the eval suite | production, A/B tests |
Connect them. If the grounding rate rises by 10 percentage points but no product metric moves, you have optimised something that did not matter — valuable information in itself.
A measurement framework that works: one north star metric (the value the product is to create), 3–5 supporting metrics that explain movements in it, and 2–3 guardrails that must not get worse.
Formal
An example for AI-grafen:
| The role | The metric | Why |
|---|---|---|
| North star | nodes mastered per hour of study | it captures actual learning per effort — not activity |
| Supporting | the share of started paths that are completed | |
| Supporting | the share of transfer tasks passed | it measures understanding, not recognition |
| Supporting | the time to the first «aha» (the first node mastered) | onboarding |
| Guardrail | the share of tutor answers that give the answer key away | must not increase |
| Guardrail | the cost per active pupil per month | must not exceed the budget |
| Guardrail | the share of pupils who leave mid-path | must not increase |
Metrics that are easy to game — and therefore bad north stars:
- Time in the app. Maximised by making things hard to find.
- The number of messages to the tutor. Maximised by giving unclear answers.
- Thumbs up. Maximised by giving the answer away.
The last is particularly important in education: research shows that pupils rate teaching lower when it is more demanding — even though they learn more. A satisfaction metric as the north star would steer the product the wrong way.
The rule: always ask «how would I maximise this metric if I did not care about the user?». If that is easy, the metric is wrong.
Code
from dataclasses import dataclass
@dataclass
class Metric:
name: str
role: str # north_star | supporting | guardrail
direction: str # up | down | stable
threshold: float | None = None
METRICS = [
Metric("nodes_mastered_per_hour", "north_star", "up"),
Metric("share_of_paths_completed", "supporting", "up"),
Metric("share_of_transfer_tasks_passed", "supporting", "up"),
Metric("share_of_answers_giving_the_key_away", "guardrail", "down", threshold=0.02),
Metric("cost_per_pupil_month", "guardrail", "stable", threshold=12.0),
]
def evaluate_release(before: dict, after: dict):
rows, blocks = [], False
for m in METRICS:
d = after[m.name] - before[m.name]
breach = (m.threshold is not None and
((m.direction == "down" and after[m.name] > m.threshold) or
(m.direction == "stable" and after[m.name] > m.threshold)))
blocks |= breach and m.role == "guardrail"
rows.append({"metric": m.name, "role": m.role, "delta": round(d, 4), "breach": breach})
return {"blocks": blocks, "rows": rows}
Mastery means
- Defines metrics that capture user value
- Tells model metrics from product metrics
- Avoids metrics that can be gamed
Sign in to do the exercises and build your mastery up.