Skip to content
AI-grafen
EUniversityInference and optimisation· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Running models locally: llama.cpp, vLLM, Ollama

Be able to run a language model locally and measure tokens per second.

Prerequisites

Intuition

Running models locally has become reasonable. A 7–8B model quantised to 4 bits fits in about 5 GB and can be run on an ordinary laptop.

Reasons to do it:

ReasonComment
The data does not leave the buildingoften the decisive reason
No cost per callonly hardware and electricity
It works without the internet
Full control over the versionsthe model does not change under your feet
A low latency to the first tokenno network travel time

Reasons not to: the largest models do not fit, the quality of an 8B model is noticeably lower than a frontier model's, and you are responsible for the operations, the updates and the security yourself.

Three tools that cover most needs:

ToolFor
Ollamathe easiest to get started with, one user
llama.cppCPU and mixed CPU/GPU, the most portable
vLLMserving to many simultaneous users, high throughput

Formal

The memory requirement is governed by the quantisation:

QuantisationBytes/parameterAn 8B modelThe quality loss
fp162~16 GBnone
Q81~8 GBnegligible
Q5_K_M~0.7~5.6 GBsmall
Q4_K_M~0.6~4.8 GBnoticeable but acceptable
Q3~0.45~3.6 GBclear
Q2~0.3~2.4 GBoften unusable

Plus the KV cache, which grows with the context length: roughly 2⋅L⋅H⋅dhead⋅ntokens⋅2 \cdot L \cdot H \cdot d_{\text{head}} \cdot n_{\text{tokens}} \cdot bytes. For an 8B model with an 8 k context that is a few hundred megabytes to a couple of gigabytes depending on the precision.

A rule of thumb: Q4_K_M is the best trade-off for most people. Below Q4 the quality falls fast.

What actually limits the speed. Generation is memory-bandwidth-bound, not compute-bound: for every token all the model weights have to be read from memory.

tokens/s≲the memory bandwidththe model size in bytes\text{tokens/s} \lesssim \frac{\text{the memory bandwidth}}{\text{the model size in bytes}}

A 5 GB model on a system with 100 GB/s of bandwidth can therefore at best reach about 20 tokens/s. That explains why a faster processor hardly helps, while faster memory does — and why Apple's unified memory (high bandwidth) performs unexpectedly well at this task.

Two numbers to measure, and they are different:

MetricMeansBound by
Prefill (prompt processing)tokens/s while the prompt is readcomputation — parallelisable
Decode (generation)tokens/s while the answer is writtenbandwidth — sequential

Prefill can be ten times faster than decode. A tool that reports only one number hides which.

Batching only helps the throughput, not the latency for an individual user — but it helps a lot: the weights are read once for the whole batch. That is the whole point of vLLM and its PagedAttention, which makes the KV cache page-based so that many simultaneous sequences can share memory efficiently.

Code

# Ollama — the simplest
ollama pull llama3.1:8b-instruct-q4_K_M
ollama run llama3.1:8b-instruct-q4_K_M "Explain gradient descent in English."

# Measure via the API — Ollama reports prefill and decode separately
curl -s http://localhost:11434/api/generate -d '{
  "model": "llama3.1:8b-instruct-q4_K_M",
  "prompt": "Write 200 words about Fourier analysis.",
  "stream": false
}' | python3 -c '
import json, sys
d = json.load(sys.stdin)
prefill = d["prompt_eval_count"] / (d["prompt_eval_duration"] / 1e9)
decode  = d["eval_count"] / (d["eval_duration"] / 1e9)
print(f"prefill {prefill:.1f} tok/s   decode {decode:.1f} tok/s")
ttft = d["prompt_eval_duration"] / 1e9
print(f"the time to the first token: {ttft:.2f} s")
'
# Measure it yourself: the latency and the throughput differ
import time, requests

def measure(model, prompt, n=5, url="http://localhost:11434/api/generate"):
    ttft, tps = [], []
    for _ in range(n):
        t0 = time.perf_counter()
        r = requests.post(url, json={"model": model, "prompt": prompt, "stream": True},
                          stream=True)
        first = None
        tokens = 0
        for line in r.iter_lines():
            if not line:
                continue
            if first is None:
                first = time.perf_counter() - t0
            tokens += 1
        total = time.perf_counter() - t0
        ttft.append(first); tps.append(tokens / total)
    ttft.sort(); tps.sort()
    return {"ttft_p50": round(ttft[len(ttft)//2], 3),
            "ttft_p95": round(ttft[int(0.95*len(ttft))-1], 3),
            "tokens_per_s": round(sum(tps)/len(tps), 1)}

# The theoretical ceiling: the bandwidth divided by the model size
def ceiling(model_gb, bandwidth_gb_per_s):
    return round(bandwidth_gb_per_s / model_gb, 1)

for name, gb, bw in (("8B Q4 on a laptop", 4.8, 100), ("8B Q4 on a GPU", 4.8, 900),
                     ("70B Q4 on a GPU", 40.0, 900)):
    print(f"{name:<19} a ceiling of ~{ceiling(gb, bw):>5.1f} tokens/s")
# 8B Q4 on a laptop   a ceiling of ~ 20.8 tokens/s
# 8B Q4 on a GPU      a ceiling of ~187.5 tokens/s
# 70B Q4 on a GPU     a ceiling of ~ 22.5 tokens/s

The three rows at the end explain more about local inference than any benchmark: the speed follows the bandwidth divided by the model size, and that is why a large model on a fast card is roughly as fast as a small one on a laptop.

Mastery means

  • Runs a model locally
  • Measures throughput and latency
  • Chooses the tool and the quantisation according to the hardware

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences