Skip to content
AI-grafen
EUniversityAI product development· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Product metrics for AI features

Be able to define metrics that capture value for the user, not just model quality.

Prerequisites

Intuition

A model can get better without the product getting better — and the other way round. So two kinds of metric are needed:

Model metricsProduct metrics
Examplesaccuracy, F1, the grounding rate, perplexitytasks completed, the time to an answer, return visits, escalations
Changed bythe modelthe whole experience
Measured inthe eval suiteproduction, A/B tests

Connect them. If the grounding rate rises by 10 percentage points but no product metric moves, you have optimised something that did not matter — valuable information in itself.

A measurement framework that works: one north star metric (the value the product is to create), 3–5 supporting metrics that explain movements in it, and 2–3 guardrails that must not get worse.

Formal

An example for AI-grafen:

The roleThe metricWhy
North starnodes mastered per hour of studyit captures actual learning per effort — not activity
Supportingthe share of started paths that are completed
Supportingthe share of transfer tasks passedit measures understanding, not recognition
Supportingthe time to the first «aha» (the first node mastered)onboarding
Guardrailthe share of tutor answers that give the answer key awaymust not increase
Guardrailthe cost per active pupil per monthmust not exceed the budget
Guardrailthe share of pupils who leave mid-pathmust not increase

Metrics that are easy to game — and therefore bad north stars:

  • Time in the app. Maximised by making things hard to find.
  • The number of messages to the tutor. Maximised by giving unclear answers.
  • Thumbs up. Maximised by giving the answer away.

The last is particularly important in education: research shows that pupils rate teaching lower when it is more demanding — even though they learn more. A satisfaction metric as the north star would steer the product the wrong way.

The rule: always ask «how would I maximise this metric if I did not care about the user?». If that is easy, the metric is wrong.

Code

from dataclasses import dataclass

@dataclass
class Metric:
    name: str
    role: str                # north_star | supporting | guardrail
    direction: str           # up | down | stable
    threshold: float | None = None

METRICS = [
    Metric("nodes_mastered_per_hour", "north_star", "up"),
    Metric("share_of_paths_completed", "supporting", "up"),
    Metric("share_of_transfer_tasks_passed", "supporting", "up"),
    Metric("share_of_answers_giving_the_key_away", "guardrail", "down", threshold=0.02),
    Metric("cost_per_pupil_month", "guardrail", "stable", threshold=12.0),
]

def evaluate_release(before: dict, after: dict):
    rows, blocks = [], False
    for m in METRICS:
        d = after[m.name] - before[m.name]
        breach = (m.threshold is not None and
                  ((m.direction == "down" and after[m.name] > m.threshold) or
                   (m.direction == "stable" and after[m.name] > m.threshold)))
        blocks |= breach and m.role == "guardrail"
        rows.append({"metric": m.name, "role": m.role, "delta": round(d, 4), "breach": breach})
    return {"blocks": blocks, "rows": rows}

Mastery means

  • Defines metrics that capture user value
  • Tells model metrics from product metrics
  • Avoids metrics that can be gamed

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences