NLP in Swedish: resources and pitfalls
Be able to use Swedish corpora and models and know their limitations.
Prerequisites
- DText preprocessingrequired
- EMultilinguality and Swedish modelsrequired
Intuition
Swedish resources to know about:
| Resource | What | The source |
|---|---|---|
| KB-BERT, KB-Whisper | an encoder and speech recognition | the National Library of Sweden |
| GPT-SW3 | generative models for the Nordic languages | AI Sweden |
| SUC | a balanced corpus with parts of speech | Språkbanken |
| Talbanken | a dependency treebank | Språkbanken |
| SweDN, SuperLim | benchmark suites for Swedish | Språkbanken |
| The Riksdag's open data | speeches and documents | the Riksdag |
| Wikipedia sv | ~2.6 million articles | CC BY-SA |
Three Swedish peculiarities that genuinely matter:
- Compounds. «arbetsmarknadsutbildning» (labour market training) is one word. The tokenizer splits it into pieces and search systems miss the parts.
- Rich morphology. The definite form, the genitive, the plural — character-based metrics (chrF) and lemmatisation help.
- A small amount of data compared with English, both in pretraining and in evaluation suites.
Formal
Measure instead of assuming. Four measurements that decide whether a model will do for Swedish:
- Tokens per word on representative Swedish text. An English-centric tokenizer gives 2.0–2.7; a Nordic one 1.3–1.6. That affects both the cost and the effective context length directly.
- Factual knowledge about Sweden — your own questions with an answer key (public agencies, laws, geography, history).
- Language quality — a blinded human judgement of fluency and idiom. Anglicisms and directly translated expressions are common.
- Your actual task on Swedish data. That is the only thing that decides.
Common errors in Swedish NLP projects:
- Using English stop-word lists or lemmatisers.
- Comparing against English benchmark figures and assuming they hold.
- Forgetting that compounds make keyword search unreliable — here subword indexing or a hybrid with embeddings helps.
- Treating «å», «ä», «ö» as special characters in normalisation.
When Swedish is not enough: RAG against Swedish sources solves factual knowledge most cheaply; fine-tuning solves style and terminology; switching to a Nordic model solves the tokenisation.
Mastery means
- Uses Swedish corpora and models
- Knows their limitations
- Evaluates in Swedish instead of assuming
Sign in to do the exercises and build your mastery up.
Sources
- Språkbanken Text (in Swedish) — open resources, licence per corpus
- AI Sweden — GPT-SW3 — free to read
- Kungliga biblioteket — KB-lab — open models