Skip to content
AI-grafen
CBuilderInference and optimisation· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Why does it take time for the AI to answer?

Be able to explain that the answer is built word by word and why long answers take longer.

Prerequisites

Everyday explanation

Have you noticed that the answer appears bit by bit, as if someone were typing it?

It is not an animation. That is exactly what is happening.

The model does one thing at a time:

"The capital"
"The capital of"
"The capital of Sweden"
"The capital of Sweden is"
"The capital of Sweden is Stockholm"

For each bit, the entire model runs once. An answer of 200 words means roughly 300 such runs, one after another.

Two wait times, and they feel different:

What it isTypical
Time to first wordThe model reads the question and starts0.3–2 s
Time per word after thatOne run per bit10–50 ms

That is why streaming answers feel faster even when they take the same total time — you see something happening instead of staring at an empty box.

Intuition

What affects the wait time:

FactorEffectWhy
Length of the answerLargeOne run per word
Model sizeLargeMore calculations per word
Length of the questionAffects first wordThe whole question must be read in first
Queue at the providerVariableMany users at the same time
Distance to serverSmallNetwork latency

The length of the question and the length of the answer affect different parts. A long question means it takes longer before the first word appears, but then it goes just as fast. A long answer makes the whole response longer.

That is why “answer briefly” is one of the most effective ways to make an AI feature faster — and it costs less too.

Why can’t it do it all at once? Because each word depends on the previous ones. The model must know that it wrote “The capital of Sweden is” in order to choose “Stockholm”. It cannot guess word 50 before it has decided on word 49.

But it can help several people at once. When many people ask at the same time, the server can run them in parallel because they do not depend on each other. That is why services can handle millions of users even though each individual answer is sequential.

One more thing that goes faster: if you have already asked a question with the same beginning, the server can reuse parts of the work. This is called caching and is the reason follow-up questions in a long conversation often start faster than you might think.

Interactive

Measure it yourself. Time it with your phone in a chatbot.

#TryMeasure
1“What is the capital of Sweden?”Time to first word, and total time
2“Write a 500-word essay about Sweden”The same two times
3Paste a long text and ask “summarise in one sentence”Same
4Same question as 1, but ask it twice in the same conversationCompare

What you should see:

TestFirst wordTotalWhy
1FastFastShort question, short answer
2FastSlowThe answer is long — many runs
3SlowFast totalLong question to read in, short answer
4Often fasterParts of the work can be reused

Tests 2 and 3 together are the point. They show that the two wait times are independent: what makes the start slow is the length of the question, what makes the whole thing slow is the length of the answer.

Practical conclusion. If you want an AI feature to feel fast:

  1. Ask for short answers.
  2. Do not send more text than necessary.
  3. Show the answer while it is being written, instead of waiting until everything is ready.

The third is the cheapest and does the most for how it feels.

Mastery means

  • Explains that the answer is built token by token
  • Knows why long answers take longer
  • Understands what affects the wait time

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences