Why does it take time for the AI to answer?
Be able to explain that the answer is built word by word and why long answers take longer.
Prerequisites
Everyday explanation
Have you noticed that the answer appears bit by bit, as if someone were typing it?
It is not an animation. That is exactly what is happening.
The model does one thing at a time:
"The capital"
"The capital of"
"The capital of Sweden"
"The capital of Sweden is"
"The capital of Sweden is Stockholm"
For each bit, the entire model runs once. An answer of 200 words means roughly 300 such runs, one after another.
Two wait times, and they feel different:
| What it is | Typical | |
|---|---|---|
| Time to first word | The model reads the question and starts | 0.3–2 s |
| Time per word after that | One run per bit | 10–50 ms |
That is why streaming answers feel faster even when they take the same total time — you see something happening instead of staring at an empty box.
Intuition
What affects the wait time:
| Factor | Effect | Why |
|---|---|---|
| Length of the answer | Large | One run per word |
| Model size | Large | More calculations per word |
| Length of the question | Affects first word | The whole question must be read in first |
| Queue at the provider | Variable | Many users at the same time |
| Distance to server | Small | Network latency |
The length of the question and the length of the answer affect different parts. A long question means it takes longer before the first word appears, but then it goes just as fast. A long answer makes the whole response longer.
That is why “answer briefly” is one of the most effective ways to make an AI feature faster — and it costs less too.
Why can’t it do it all at once? Because each word depends on the previous ones. The model must know that it wrote “The capital of Sweden is” in order to choose “Stockholm”. It cannot guess word 50 before it has decided on word 49.
But it can help several people at once. When many people ask at the same time, the server can run them in parallel because they do not depend on each other. That is why services can handle millions of users even though each individual answer is sequential.
One more thing that goes faster: if you have already asked a question with the same beginning, the server can reuse parts of the work. This is called caching and is the reason follow-up questions in a long conversation often start faster than you might think.
Interactive
Measure it yourself. Time it with your phone in a chatbot.
| # | Try | Measure |
|---|---|---|
| 1 | “What is the capital of Sweden?” | Time to first word, and total time |
| 2 | “Write a 500-word essay about Sweden” | The same two times |
| 3 | Paste a long text and ask “summarise in one sentence” | Same |
| 4 | Same question as 1, but ask it twice in the same conversation | Compare |
What you should see:
| Test | First word | Total | Why |
|---|---|---|---|
| 1 | Fast | Fast | Short question, short answer |
| 2 | Fast | Slow | The answer is long — many runs |
| 3 | Slow | Fast total | Long question to read in, short answer |
| 4 | Often faster | Parts of the work can be reused |
Tests 2 and 3 together are the point. They show that the two wait times are independent: what makes the start slow is the length of the question, what makes the whole thing slow is the length of the answer.
Practical conclusion. If you want an AI feature to feel fast:
- Ask for short answers.
- Do not send more text than necessary.
- Show the answer while it is being written, instead of waiting until everything is ready.
The third is the cheapest and does the most for how it feels.
Mastery means
- Explains that the answer is built token by token
- Knows why long answers take longer
- Understands what affects the wait time
Sign in to do the exercises and build your mastery up.
Sources
- Internetstiftelsen — Internetkunskap — free to read
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms