Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Summarisation and translation

Be able to use and evaluate generative models for summarisation and translation.

Prerequisites

Intuition

Two kinds of summarisation:

  • Extractive — pick sentences out of the original. It cannot hallucinate, but it comes out choppy.
  • Abstractive — write something new. It flows better, it can make things up.

For abstractive summarisation faithfulness is the real problem: the summary sounds good but contains claims that are not in the source, or reverses a negated claim.

The metrics and what they are good for:

MetricMeasuresThe weakness
ROUGEn-gram overlap with a referenceit rewards copying, blind to facts
BLEUthe precision of n-grams (translation)the same problem
chrFcharacter-based — better for Swedishstill superficial
BERTScoresemantic similarity via embeddingsit does not see factual errors
COMETa learnt quality metric for translationthe best correlation with people
An LLM judge against the sourcefaithfulness per claimit costs, it has to be calibrated

ROUGE alone says almost nothing about whether the summary is true.

Code

import json

FAITHFULNESS = """Split the SUMMARY into claims. For each one: is it supported by the SOURCE?
Answer in JSON: {"claims": [{"text": "...", "supported": true|false, "why": "..."}]}"""

def evaluate_summary(llm, source, summary):
    d = json.loads(llm(f"{FAITHFULNESS}\n\nSOURCE:\n{source}\n\nSUMMARY:\n{summary}",
                       temperature=0, json_mode=True))
    c = d["claims"]
    return {"faithfulness": sum(x["supported"] for x in c) / max(len(c), 1),
            "unsupported": [x["text"] for x in c if not x["supported"]],
            "compression": round(len(summary.split()) / max(len(source.split()), 1), 3)}

# Translation: chrF works better than BLEU for Swedish (morphologically rich)
from sacrebleu import CHRF, BLEU
ref = [["Katten sover på mattan."]]
hyp = ["Katten ligger och sover på mattan."]
print(round(BLEU().corpus_score(hyp, ref).score, 1),
      round(CHRF().corpus_score(hyp, ref).score, 1))     # 25.9  68.4

The example shows the problem: the translation is in practice correct, but BLEU gives 26 out of 100 because the words are not identical. chrF sees that the characters overlap; a person would have said «good».

A practical rule: use the automatic metrics to detect regressions between versions, and a faithfulness judge plus a human sample to decide quality.

Mastery means

  • Uses generative models for summarisation and translation
  • Evaluates with the right metrics and knows their weaknesses
  • Handles faithfulness and hallucination

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences