Skip to content
AI-grafen
EUniversityEthics, law and society· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Model cards and system documentation

Be able to write a model card with the intended use, the limitations and the evaluation.

Prerequisites

Intuition

A model card (Mitchell et al. 2019) is to a model what a datasheet is to a dataset. Nine headings:

  1. Model details — who, when, the version, the architecture, the licence.
  2. Intended use — what it is built for, and for whom.
  3. Out-of-scope use — what it should not be used for.
  4. Factors — which groups and conditions affect the performance.
  5. Metrics — which ones, and why those.
  6. Evaluation data — the source, the size, how it was chosen.
  7. Training data — the same, plus the known biases.
  8. Quantitative analyses — the results broken down by group, not just overall.
  9. Ethical considerations and recommendations.

Headings 3 and 8 do the most good — and are the ones most often missing.

Formal

The EU AI Act makes documentation a requirement, not a good habit. For high-risk systems (including education, recruitment, credit and justice) it requires technical documentation, logging, human oversight, risk management and a description of the data quality. For generative systems there are additional requirements to label synthetic content and to report the training data at a general level.

For AI-grafen the education connection is relevant: systems used to assess pupils or steer their education fall into the high-risk category. The platform's position — that a competence certificate is documented evidence and not a qualification, and that the model never sets grades — is partly an answer to exactly that.

What the documentation has to achieve in practice: somebody who did not build the system should be able to decide whether it suits their use, what risks it has, and what happens when it is wrong. If the documentation does not achieve that, it is for internal use, not transparency.

A practical tip: write the model card before the launch. It forces decisions about the intended use and the limitations while they can still be influenced.

Interactive

A model card — a short form to fill in:

# Model card: <name> v<version>

## Model details
Developed by … | Date … | Type … | Licence … | Contact …

## Intended use
Intended for: suggesting the next exercise for pupils in years 7 to upper secondary.
Intended users: pupils, teachers.

## Out of scope
Must NOT be used for grading, selection, or assessing individual
pupils' ability in formal contexts.

## Factors
Performance varies with: the pupil's level (A–G), the subject area, the language (Swedish only).

## Metrics
Mastery prediction: AUC. Recommendation quality: the share of steps completed.
Guard rail: the share of answers that reveal the answer key.

## Evaluation data
2 400 sessions, September 2026, every level. The test set was used once.

## Quantitative analyses
| Group | AUC | n |
|---|---|---|
| levels A–C | 0.81 | 900 |
| levels D–E | 0.78 | 1 100 |
| levels F–G | 0.71 | 400 |   ← worse, a small sample

## Ethical considerations
A risk of a self-fulfilling prediction if the recommendation steers too hard.
The countermeasure: the pupil can always choose freely; the suggestion is never a barrier.

## Recommendations
Use with a teacher in the loop. Re-evaluate every six months.

The table under «quantitative analyses» is the core: it shows where the model is weak, instead of hiding it in an average.

Mastery means

  • Writes a model card with the intended use and the limitations
  • Documents the evaluation broken down by group
  • Knows what the AI Act requires

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences