The goal E University
Language models in practice
Tokenisation, sampling, instruction models, structured output, hallucinations and Swedish — what you need in order to build with an LLM for real.
- Knowledge nodes
- 70
- From zero
- about 49 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- Dn-gram language models
- EByte-pair encoding
- EReasoning in language models
- EStructured output and schemas
- DThe transformer — an overview without formulas
- EGenerative models — an overview
- EEncoder–decoder transformers
- EAttention variants: cross, causal, sparse
- EThe KV cache
- EMultilinguality and Swedish models
- EHallucinations — causes and countermeasures
- EBase against instruction models
- EPerplexity
- ESampling: temperature, top-k, top-p, beam
- ESummarisation and translation
- ELanguage models for code
- ENLP in Swedish: resources and pitfalls
- EInformation extraction and NER
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: byte-pair encoding from scratchEa sandbox · about 60 minLab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: scaled dot-product attention with a causal maskDa sandbox · about 60 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer4 knowledge nodes
BInvestigator11 knowledge nodes
- Data in everyday life
- Algorithmic thinking
- Rule-based systems and machine learning
- Source criticism and responsibility in AI use
- Fractions, decimals, and percentages
- Coordinates: positions on a grid
- Negative numbers in everyday life and AI
- Training data, features and labels
- Classification: how a model sorts information
- Generative AI: how it creates content
- Language models and probabilities
CBuilder15 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Python — strings and text processing
- Python — files, CSV and JSON
- Statistics — mean, median and spread
- Probability — the basics
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Prompting — steering a language model
- Attention: how words influence each other
- Words as points: similar words close together
- Variables and algebraic expressions
- Powers and roots
DAI developer19 knowledge nodes
- Derivatives and optimisation
- Loss functions
- n-gram language models
- Tokenisation
- Text preprocessing
- Gradient descent
- The context window, system prompts and few-shot
- Precision, recall, F1 and ROC
- The transformer — an overview without formulas
- Logarithms
- Vectors
- Matrices and matrix multiplication
- Neural networks — the forward pass with matrices
- Backpropagation
- Embeddings — words as vectors
- Attention
- PyTorch — tensors and autograd
- Train a neural network in PyTorch
- Transformers — the architecture
EUniversity21 knowledge nodes
- Byte-pair encoding
- Reasoning in language models
- Structured output and schemas
- Information theory: entropy and KL divergence
- Generative models — an overview
- BERT and masked language modelling
- Encoder–decoder transformers
- Multi-head attention in detail
- Attention variants: cross, causal, sparse
- The KV cache
- Language models — training and generation
- Multilinguality and Swedish models
- Hallucinations — causes and countermeasures
- Base against instruction models
- Perplexity
- Sampling: temperature, top-k, top-p, beam
- Summarisation and translation
- Language models for code
- NLP in Swedish: resources and pitfalls
- Text classification with transformers
- Information extraction and NER