Language models and probabilities
Understand how a language model generates text by calculating the probability of the next word.
Prerequisites
Everyday explanation
Consider the sentence «I am going to send an email to …». Most people think of a colleague, a client, or a supplier. Few think of a robot or an asteroid.
You guess based on what you have experienced before. A language model does exactly that — except it has processed vast amounts of text and calculates: “colleague” 40%, “client” 20%, “boss” 10% … Then it selects a word, adds it, and calculates the probability of the next word.
An entire text is just a series of these probability calculations.
Interactive
Try it yourself: Write the beginning of a sentence, e.g. «The meeting started at …». What is the most likely answer? Perhaps “9:00” or “10:00”. That is the word’s “probability” in that context.
«I paid the invoice with …» → cash, card, check. «The weather in Stockholm is …» → rainy, sunny, windy.
Try a sentence where it is hard to guess: «My favourite colour is …» — here the answers are scattered, and a language model also becomes uncertain.
Mastery means
- Can explain why certain words are more likely than others in a given context.
- Can link the concept of probability to how a language model constructs sentences.
Sign in to do the exercises and build your mastery up.
Sources
- Swedish Wikipedia — Language model (CC BY-SA 4.0) — CC BY-SA 4.0
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
Part of the goals (13)
- Discoverer — build a game and train a machine
- AI, ethics and society
- An AI service in operation
- Language models in practice
- Build an agent you can trust
- Understand how generative AI works
- Computers that see and hear
- Build a RAG system you can trust
- Fine-tune and run your own models
- Build a memory system for an agent
- Builder — collect data, train a model and test AI critically
- Statistics for experiments
- Run models more cheaply: quantisation