Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Multilinguality and Swedish models

Be able to judge how the tokenisation and the training data affect performance in Swedish.

Prerequisites

Intuition

A model trained mostly on English performs worse in Swedish — but in specific, measurable ways:

  1. More tokens per word. Swedish compounds («kunskapsgraf», «arbetsmarknadsutbildning») are split into many pieces. 1.5–2× more tokens than the corresponding English text means a higher cost and less fitting into the context.
  2. Weaker factual knowledge about Swedish matters: laws, public agencies, history, geography.
  3. Linguistic slips: anglicisms, odd word order, directly translated expressions.
  4. Cross-lingual transfer, on the other hand, works surprisingly well for reasoning — the model can solve a problem in Swedish with abilities it learnt in English.

The conclusion: assume nothing. Measure in Swedish.

Formal

Evaluating in Swedish — a practical plan:

WhatHow
The tokenisation efficiencytokens per word on a representative Swedish text, compare models
Factual knowledgeyour own questions about Swedish matters with an answer key
Language qualitya human judgement of fluency and idiom, blinded
Task performanceyour actual task, on Swedish data

The alternatives when Swedish is not good enough:

  • Prompt in English, answer in Swedish — sometimes works surprisingly well.
  • Fine-tune on Swedish domain data (LoRA is often enough for style and terminology).
  • Nordic/Swedish models (GPT-SW3 from AI Sweden, KB-BERT/KB-Whisper from the National Library) — better tokenisation and Swedish data, but usually smaller models.
  • RAG on Swedish sources — solves the factual knowledge without touching the model.

The last is usually the best first measure: the cheapest, the most controllable and it solves the most common problem (facts, not language).

Code

from transformers import AutoTokenizer

text_sv = "Arbetsmarknadsutbildningen på Kungliga Tekniska högskolan startade i höstas."
text_en = "The labour market training programme at the Royal Institute of Technology started last autumn."

for name in ["gpt2", "AI-Sweden-Models/gpt-sw3-356m"]:
    tok = AutoTokenizer.from_pretrained(name)
    sv, en = len(tok(text_sv)["input_ids"]), len(tok(text_en)["input_ids"])
    print(f"{name:34s} sv {sv:3d} ({sv/len(text_sv.split()):.2f}/word)  en {en:3d}")
# gpt2                               sv  27 (2.70/word)  en  17
# AI-Sweden-Models/gpt-sw3-356m      sv  14 (1.40/word)  en  19

The same sentence, nearly double the cost in an English-centric tokenizer. Over millions of calls that is money — and over a context window it is content that does not fit.

Mastery means

  • Judges how the tokenisation and the training data affect Swedish
  • Evaluates a model in Swedish instead of assuming

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences