Skip to content
AI-grafen
FAI engineeringInference and optimisation· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Serving and scaling models

Be able to set up an inference server with queues, timeouts and observability.

Prerequisites

Intuition

Serving a model is not starting model.generate() behind an API. Seven things have to be in place before it stands up to real traffic:

PartWhy
Continuous batching2–10× throughput; vLLM/TGI do it for you
A queue with a capwithout a cap the latency grows without bound at a peak — better to reject with a 429
A timeout per requestone that generates 4 000 tokens must not block the rest
A quota per userprotects against both abuse and bugs in clients
A health checkthe orchestration has to know whether the node should get traffic
MeasurementTTFT, tokens/s, the queue length, GPU utilisation, cost
Degradationa smaller model or a pre-written answer when the queue is full

The second-to-last row is what makes debugging possible; the last is what stops the user meeting a blank page.

Code

import asyncio, time
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field

app = FastAPI()
QUEUE = asyncio.Queue(maxsize=200)
CONCURRENT = asyncio.Semaphore(32)

class In(BaseModel):
    prompt: str = Field(min_length=1, max_length=8000)
    max_tokens: int = Field(default=512, ge=1, le=2048)

@app.get("/healthz")
async def healthz():
    return {"ok": True, "queue": QUEUE.qsize(), "capacity": QUEUE.maxsize}

@app.post("/v1/generate")
async def generate(body: In):
    if QUEUE.full():
        raise HTTPException(429, "The system is overloaded — please try again shortly")
    await QUEUE.put(1)
    t0 = time.perf_counter()
    try:
        async with CONCURRENT:
            try:
                text = await asyncio.wait_for(engine.generate(body.prompt, body.max_tokens), timeout=60)
            except asyncio.TimeoutError:
                raise HTTPException(504, "The timeout was exceeded")
        return {"text": text, "latency_ms": round((time.perf_counter() - t0) * 1000)}
    finally:
        QUEUE.get_nowait(); QUEUE.task_done()

Sizing: measure the throughput at the concurrency that gives an acceptable p95 latency (not the maximum throughput). Then run at ~60–70 % of that capacity in the normal case — the rest is margin for peaks. Autoscaling on GPUs takes minutes, so the margin has to exist in advance.

Mastery means

  • Sets up an inference server with queues and timeouts
  • Sizes it according to the p95 requirement
  • Builds in observability and degradation

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences