FAI engineeringInference and optimisation· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Serving and scaling models
Be able to set up an inference server with queues, timeouts and observability.
Prerequisites
Intuition
Serving a model is not starting model.generate() behind an API. Seven things have to be in place before it stands up to real traffic:
| Part | Why |
|---|---|
| Continuous batching | 2–10× throughput; vLLM/TGI do it for you |
| A queue with a cap | without a cap the latency grows without bound at a peak — better to reject with a 429 |
| A timeout per request | one that generates 4 000 tokens must not block the rest |
| A quota per user | protects against both abuse and bugs in clients |
| A health check | the orchestration has to know whether the node should get traffic |
| Measurement | TTFT, tokens/s, the queue length, GPU utilisation, cost |
| Degradation | a smaller model or a pre-written answer when the queue is full |
The second-to-last row is what makes debugging possible; the last is what stops the user meeting a blank page.
Code
import asyncio, time
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
app = FastAPI()
QUEUE = asyncio.Queue(maxsize=200)
CONCURRENT = asyncio.Semaphore(32)
class In(BaseModel):
prompt: str = Field(min_length=1, max_length=8000)
max_tokens: int = Field(default=512, ge=1, le=2048)
@app.get("/healthz")
async def healthz():
return {"ok": True, "queue": QUEUE.qsize(), "capacity": QUEUE.maxsize}
@app.post("/v1/generate")
async def generate(body: In):
if QUEUE.full():
raise HTTPException(429, "The system is overloaded — please try again shortly")
await QUEUE.put(1)
t0 = time.perf_counter()
try:
async with CONCURRENT:
try:
text = await asyncio.wait_for(engine.generate(body.prompt, body.max_tokens), timeout=60)
except asyncio.TimeoutError:
raise HTTPException(504, "The timeout was exceeded")
return {"text": text, "latency_ms": round((time.perf_counter() - t0) * 1000)}
finally:
QUEUE.get_nowait(); QUEUE.task_done()
Sizing: measure the throughput at the concurrency that gives an acceptable p95 latency (not the maximum throughput). Then run at ~60–70 % of that capacity in the normal case — the rest is margin for peaks. Autoscaling on GPUs takes minutes, so the margin has to exist in advance.
Mastery means
- Sets up an inference server with queues and timeouts
- Sizes it according to the p95 requirement
- Builds in observability and degradation
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Efficient Memory Management for LLM Serving with PagedAttention (vLLM) — arXiv (open access; licence per article)
- Text Generation Inference (HFOILv1.0/Apache-2.0) — Apache-2.0