Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

NLP in Swedish: resources and pitfalls

Be able to use Swedish corpora and models and know their limitations.

Prerequisites

Intuition

Swedish resources to know about:

ResourceWhatThe source
KB-BERT, KB-Whisperan encoder and speech recognitionthe National Library of Sweden
GPT-SW3generative models for the Nordic languagesAI Sweden
SUCa balanced corpus with parts of speechSpråkbanken
Talbankena dependency treebankSpråkbanken
SweDN, SuperLimbenchmark suites for SwedishSpråkbanken
The Riksdag's open dataspeeches and documentsthe Riksdag
Wikipedia sv~2.6 million articlesCC BY-SA

Three Swedish peculiarities that genuinely matter:

  1. Compounds. «arbetsmarknadsutbildning» (labour market training) is one word. The tokenizer splits it into pieces and search systems miss the parts.
  2. Rich morphology. The definite form, the genitive, the plural — character-based metrics (chrF) and lemmatisation help.
  3. A small amount of data compared with English, both in pretraining and in evaluation suites.

Formal

Measure instead of assuming. Four measurements that decide whether a model will do for Swedish:

  1. Tokens per word on representative Swedish text. An English-centric tokenizer gives 2.0–2.7; a Nordic one 1.3–1.6. That affects both the cost and the effective context length directly.
  2. Factual knowledge about Sweden — your own questions with an answer key (public agencies, laws, geography, history).
  3. Language quality — a blinded human judgement of fluency and idiom. Anglicisms and directly translated expressions are common.
  4. Your actual task on Swedish data. That is the only thing that decides.

Common errors in Swedish NLP projects:

  • Using English stop-word lists or lemmatisers.
  • Comparing against English benchmark figures and assuming they hold.
  • Forgetting that compounds make keyword search unreliable — here subword indexing or a hybrid with embeddings helps.
  • Treating «å», «ä», «ö» as special characters in normalisation.

When Swedish is not enough: RAG against Swedish sources solves factual knowledge most cheaply; fine-tuning solves style and terminology; switching to a Nordic model solves the tokenisation.

Mastery means

  • Uses Swedish corpora and models
  • Knows their limitations
  • Evaluates in Swedish instead of assuming

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences