Skip to content
AI-grafen
EUniversityInference and optimisation· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Latency, throughput and batching in inference

Be able to measure and reason about latency against throughput and continuous batching.

Prerequisites

Intuition

Three numbers that are often mixed up:

MetricWhatWho cares
TTFT (time to first token)how quickly something starts happeningthe user — this is what feels like «speed»
TPOT (time per output token)how quickly the text rolls outthe user, on long answers
Throughputtokens per second in total across all userswhoever pays for the hardware

They are in conflict. A larger batch → higher throughput, worse TTFT for the individual. You have to choose which requirement governs — and measure both.

A rule of thumb that holds: a TTFT under 1 s and 20+ tokens/s is experienced as fast in a chat. Below 10 tokens/s it feels sluggish even if the answer is good.

Formal

Static batching waits for N requests, runs them together and releases them all when the longest is finished. A request that generates 20 tokens waits for one that generates 800. The utilisation becomes dreadful.

Continuous batching (Orca, vLLM) works at the iteration level: after every decoding step finished sequences are released and new ones are taken in in their place. The effect is 2–10× higher throughput at the same latency requirement — the single biggest gain in an LLM service.

Queueing theory is enough to set expectations: with an arrival rate λ and a service capacity μ the waiting time grows towards infinity as the utilisation ρ = λ/μ approaches 1. Plan for ρ ≈ 0.6–0.7 with a p95 requirement; run at 0.95 and you have no margin for traffic peaks.

A measurement protocol that gives comparable numbers: a fixed prompt length and a fixed number of generated tokens, a warm-up before measuring, the median and p95 (not the mean), and load from several simultaneous clients — otherwise you are only measuring the single-user case.

Code

import asyncio, time, numpy as np

async def one_request(client, prompt, max_tokens=128):
    t0 = time.perf_counter(); first = None; n = 0
    async for _tok in client.stream(prompt, max_tokens=max_tokens):
        if first is None:
            first = time.perf_counter() - t0
        n += 1
    total = time.perf_counter() - t0
    return {"ttft": first, "tpot": (total - first) / max(n - 1, 1), "total": total, "tokens": n}

async def load_test(client, prompt, concurrent=16, per_client=5):
    t0 = time.perf_counter()
    res = await asyncio.gather(*[one_request(client, prompt)
                                 for _ in range(concurrent * per_client)])
    wall = time.perf_counter() - t0
    ttft = np.array([r["ttft"] for r in res]); tpot = np.array([r["tpot"] for r in res])
    return {"concurrent": concurrent,
            "ttft_p50_ms": round(np.percentile(ttft, 50) * 1000),
            "ttft_p95_ms": round(np.percentile(ttft, 95) * 1000),
            "tokens_per_s_per_user": round(1 / np.median(tpot), 1),
            "throughput_tokens_per_s": round(sum(r["tokens"] for r in res) / wall)}

# concurrent=1:  ttft_p50 180 ms, 45 tok/s/user, throughput 45
# concurrent=16: ttft_p50 520 ms, 28 tok/s/user, throughput 448   ← 10× throughput, 3× TTFT

Mastery means

  • Measures TTFT, time per token and throughput separately
  • Explains continuous batching
  • Chooses a batching strategy according to the requirements

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences