Multilinguality and Swedish models
Be able to judge how the tokenisation and the training data affect performance in Swedish.
Prerequisites
- DTokenisationrequired
- ELanguage models — training and generationrequired
Intuition
A model trained mostly on English performs worse in Swedish — but in specific, measurable ways:
- More tokens per word. Swedish compounds («kunskapsgraf», «arbetsmarknadsutbildning») are split into many pieces. 1.5–2× more tokens than the corresponding English text means a higher cost and less fitting into the context.
- Weaker factual knowledge about Swedish matters: laws, public agencies, history, geography.
- Linguistic slips: anglicisms, odd word order, directly translated expressions.
- Cross-lingual transfer, on the other hand, works surprisingly well for reasoning — the model can solve a problem in Swedish with abilities it learnt in English.
The conclusion: assume nothing. Measure in Swedish.
Formal
Evaluating in Swedish — a practical plan:
| What | How |
|---|---|
| The tokenisation efficiency | tokens per word on a representative Swedish text, compare models |
| Factual knowledge | your own questions about Swedish matters with an answer key |
| Language quality | a human judgement of fluency and idiom, blinded |
| Task performance | your actual task, on Swedish data |
The alternatives when Swedish is not good enough:
- Prompt in English, answer in Swedish — sometimes works surprisingly well.
- Fine-tune on Swedish domain data (LoRA is often enough for style and terminology).
- Nordic/Swedish models (GPT-SW3 from AI Sweden, KB-BERT/KB-Whisper from the National Library) — better tokenisation and Swedish data, but usually smaller models.
- RAG on Swedish sources — solves the factual knowledge without touching the model.
The last is usually the best first measure: the cheapest, the most controllable and it solves the most common problem (facts, not language).
Code
from transformers import AutoTokenizer
text_sv = "Arbetsmarknadsutbildningen på Kungliga Tekniska högskolan startade i höstas."
text_en = "The labour market training programme at the Royal Institute of Technology started last autumn."
for name in ["gpt2", "AI-Sweden-Models/gpt-sw3-356m"]:
tok = AutoTokenizer.from_pretrained(name)
sv, en = len(tok(text_sv)["input_ids"]), len(tok(text_en)["input_ids"])
print(f"{name:34s} sv {sv:3d} ({sv/len(text_sv.split()):.2f}/word) en {en:3d}")
# gpt2 sv 27 (2.70/word) en 17
# AI-Sweden-Models/gpt-sw3-356m sv 14 (1.40/word) en 19
The same sentence, nearly double the cost in an English-centric tokenizer. Over millions of calls that is money — and over a context window it is content that does not fit.
Mastery means
- Judges how the tokenisation and the training data affect Swedish
- Evaluates a model in Swedish instead of assuming
Sign in to do the exercises and build your mastery up.
Sources
- AI Sweden — GPT-SW3 — free to read
- Kungliga biblioteket — KB-lab modeller — open models