Summarisation and translation
Be able to use and evaluate generative models for summarisation and translation.
Prerequisites
- EEncoder–decoder transformersrequired
- ESampling: temperature, top-k, top-p, beamrequired
Intuition
Two kinds of summarisation:
- Extractive — pick sentences out of the original. It cannot hallucinate, but it comes out choppy.
- Abstractive — write something new. It flows better, it can make things up.
For abstractive summarisation faithfulness is the real problem: the summary sounds good but contains claims that are not in the source, or reverses a negated claim.
The metrics and what they are good for:
| Metric | Measures | The weakness |
|---|---|---|
| ROUGE | n-gram overlap with a reference | it rewards copying, blind to facts |
| BLEU | the precision of n-grams (translation) | the same problem |
| chrF | character-based — better for Swedish | still superficial |
| BERTScore | semantic similarity via embeddings | it does not see factual errors |
| COMET | a learnt quality metric for translation | the best correlation with people |
| An LLM judge against the source | faithfulness per claim | it costs, it has to be calibrated |
ROUGE alone says almost nothing about whether the summary is true.
Code
import json
FAITHFULNESS = """Split the SUMMARY into claims. For each one: is it supported by the SOURCE?
Answer in JSON: {"claims": [{"text": "...", "supported": true|false, "why": "..."}]}"""
def evaluate_summary(llm, source, summary):
d = json.loads(llm(f"{FAITHFULNESS}\n\nSOURCE:\n{source}\n\nSUMMARY:\n{summary}",
temperature=0, json_mode=True))
c = d["claims"]
return {"faithfulness": sum(x["supported"] for x in c) / max(len(c), 1),
"unsupported": [x["text"] for x in c if not x["supported"]],
"compression": round(len(summary.split()) / max(len(source.split()), 1), 3)}
# Translation: chrF works better than BLEU for Swedish (morphologically rich)
from sacrebleu import CHRF, BLEU
ref = [["Katten sover på mattan."]]
hyp = ["Katten ligger och sover på mattan."]
print(round(BLEU().corpus_score(hyp, ref).score, 1),
round(CHRF().corpus_score(hyp, ref).score, 1)) # 25.9 68.4
The example shows the problem: the translation is in practice correct, but BLEU gives 26 out of 100 because the words are not identical. chrF sees that the characters overlap; a person would have said «good».
A practical rule: use the automatic metrics to detect regressions between versions, and a faithfulness judge plus a human sample to decide quality.
Mastery means
- Uses generative models for summarisation and translation
- Evaluates with the right metrics and knows their weaknesses
- Handles faithfulness and hallucination
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — SummEval: Re-evaluating Summarization Evaluation — arXiv (open access; licence per article)
- sacreBLEU (Apache-2.0) — Apache-2.0