Latency, throughput and batching in inference
Be able to measure and reason about latency against throughput and continuous batching.
Prerequisites
- EThe KV cacherequired
Intuition
Three numbers that are often mixed up:
| Metric | What | Who cares |
|---|---|---|
| TTFT (time to first token) | how quickly something starts happening | the user — this is what feels like «speed» |
| TPOT (time per output token) | how quickly the text rolls out | the user, on long answers |
| Throughput | tokens per second in total across all users | whoever pays for the hardware |
They are in conflict. A larger batch → higher throughput, worse TTFT for the individual. You have to choose which requirement governs — and measure both.
A rule of thumb that holds: a TTFT under 1 s and 20+ tokens/s is experienced as fast in a chat. Below 10 tokens/s it feels sluggish even if the answer is good.
Formal
Static batching waits for N requests, runs them together and releases them all when the longest is finished. A request that generates 20 tokens waits for one that generates 800. The utilisation becomes dreadful.
Continuous batching (Orca, vLLM) works at the iteration level: after every decoding step finished sequences are released and new ones are taken in in their place. The effect is 2–10× higher throughput at the same latency requirement — the single biggest gain in an LLM service.
Queueing theory is enough to set expectations: with an arrival rate λ and a service capacity μ the waiting time grows towards infinity as the utilisation ρ = λ/μ approaches 1. Plan for ρ ≈ 0.6–0.7 with a p95 requirement; run at 0.95 and you have no margin for traffic peaks.
A measurement protocol that gives comparable numbers: a fixed prompt length and a fixed number of generated tokens, a warm-up before measuring, the median and p95 (not the mean), and load from several simultaneous clients — otherwise you are only measuring the single-user case.
Code
import asyncio, time, numpy as np
async def one_request(client, prompt, max_tokens=128):
t0 = time.perf_counter(); first = None; n = 0
async for _tok in client.stream(prompt, max_tokens=max_tokens):
if first is None:
first = time.perf_counter() - t0
n += 1
total = time.perf_counter() - t0
return {"ttft": first, "tpot": (total - first) / max(n - 1, 1), "total": total, "tokens": n}
async def load_test(client, prompt, concurrent=16, per_client=5):
t0 = time.perf_counter()
res = await asyncio.gather(*[one_request(client, prompt)
for _ in range(concurrent * per_client)])
wall = time.perf_counter() - t0
ttft = np.array([r["ttft"] for r in res]); tpot = np.array([r["tpot"] for r in res])
return {"concurrent": concurrent,
"ttft_p50_ms": round(np.percentile(ttft, 50) * 1000),
"ttft_p95_ms": round(np.percentile(ttft, 95) * 1000),
"tokens_per_s_per_user": round(1 / np.median(tpot), 1),
"throughput_tokens_per_s": round(sum(r["tokens"] for r in res) / wall)}
# concurrent=1: ttft_p50 180 ms, 45 tok/s/user, throughput 45
# concurrent=16: ttft_p50 520 ms, 28 tok/s/user, throughput 448 ← 10× throughput, 3× TTFT
Mastery means
- Measures TTFT, time per token and throughput separately
- Explains continuous batching
- Chooses a batching strategy according to the requirements
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Efficient Memory Management for LLM Serving with PagedAttention (vLLM) — arXiv (open access; licence per article)
- arXiv — Orca: A Distributed Serving System for Transformer-Based Generative Models — arXiv (open access; licence per article)