Running models locally: llama.cpp, vLLM, Ollama
Be able to run a language model locally and measure tokens per second.
Prerequisites
- EDocker — containersrequired
- FQuantisationrequired
Intuition
Running models locally has become reasonable. A 7–8B model quantised to 4 bits fits in about 5 GB and can be run on an ordinary laptop.
Reasons to do it:
| Reason | Comment |
|---|---|
| The data does not leave the building | often the decisive reason |
| No cost per call | only hardware and electricity |
| It works without the internet | |
| Full control over the versions | the model does not change under your feet |
| A low latency to the first token | no network travel time |
Reasons not to: the largest models do not fit, the quality of an 8B model is noticeably lower than a frontier model's, and you are responsible for the operations, the updates and the security yourself.
Three tools that cover most needs:
| Tool | For |
|---|---|
| Ollama | the easiest to get started with, one user |
| llama.cpp | CPU and mixed CPU/GPU, the most portable |
| vLLM | serving to many simultaneous users, high throughput |
Formal
The memory requirement is governed by the quantisation:
| Quantisation | Bytes/parameter | An 8B model | The quality loss |
|---|---|---|---|
| fp16 | 2 | ~16 GB | none |
| Q8 | 1 | ~8 GB | negligible |
| Q5_K_M | ~0.7 | ~5.6 GB | small |
| Q4_K_M | ~0.6 | ~4.8 GB | noticeable but acceptable |
| Q3 | ~0.45 | ~3.6 GB | clear |
| Q2 | ~0.3 | ~2.4 GB | often unusable |
Plus the KV cache, which grows with the context length: roughly bytes. For an 8B model with an 8 k context that is a few hundred megabytes to a couple of gigabytes depending on the precision.
A rule of thumb: Q4_K_M is the best trade-off for most people. Below Q4 the quality falls fast.
What actually limits the speed. Generation is memory-bandwidth-bound, not compute-bound: for every token all the model weights have to be read from memory.
A 5 GB model on a system with 100 GB/s of bandwidth can therefore at best reach about 20 tokens/s. That explains why a faster processor hardly helps, while faster memory does — and why Apple's unified memory (high bandwidth) performs unexpectedly well at this task.
Two numbers to measure, and they are different:
| Metric | Means | Bound by |
|---|---|---|
| Prefill (prompt processing) | tokens/s while the prompt is read | computation — parallelisable |
| Decode (generation) | tokens/s while the answer is written | bandwidth — sequential |
Prefill can be ten times faster than decode. A tool that reports only one number hides which.
Batching only helps the throughput, not the latency for an individual user — but it helps a lot: the weights are read once for the whole batch. That is the whole point of vLLM and its PagedAttention, which makes the KV cache page-based so that many simultaneous sequences can share memory efficiently.
Code
# Ollama — the simplest
ollama pull llama3.1:8b-instruct-q4_K_M
ollama run llama3.1:8b-instruct-q4_K_M "Explain gradient descent in English."
# Measure via the API — Ollama reports prefill and decode separately
curl -s http://localhost:11434/api/generate -d '{
"model": "llama3.1:8b-instruct-q4_K_M",
"prompt": "Write 200 words about Fourier analysis.",
"stream": false
}' | python3 -c '
import json, sys
d = json.load(sys.stdin)
prefill = d["prompt_eval_count"] / (d["prompt_eval_duration"] / 1e9)
decode = d["eval_count"] / (d["eval_duration"] / 1e9)
print(f"prefill {prefill:.1f} tok/s decode {decode:.1f} tok/s")
ttft = d["prompt_eval_duration"] / 1e9
print(f"the time to the first token: {ttft:.2f} s")
'
# Measure it yourself: the latency and the throughput differ
import time, requests
def measure(model, prompt, n=5, url="http://localhost:11434/api/generate"):
ttft, tps = [], []
for _ in range(n):
t0 = time.perf_counter()
r = requests.post(url, json={"model": model, "prompt": prompt, "stream": True},
stream=True)
first = None
tokens = 0
for line in r.iter_lines():
if not line:
continue
if first is None:
first = time.perf_counter() - t0
tokens += 1
total = time.perf_counter() - t0
ttft.append(first); tps.append(tokens / total)
ttft.sort(); tps.sort()
return {"ttft_p50": round(ttft[len(ttft)//2], 3),
"ttft_p95": round(ttft[int(0.95*len(ttft))-1], 3),
"tokens_per_s": round(sum(tps)/len(tps), 1)}
# The theoretical ceiling: the bandwidth divided by the model size
def ceiling(model_gb, bandwidth_gb_per_s):
return round(bandwidth_gb_per_s / model_gb, 1)
for name, gb, bw in (("8B Q4 on a laptop", 4.8, 100), ("8B Q4 on a GPU", 4.8, 900),
("70B Q4 on a GPU", 40.0, 900)):
print(f"{name:<19} a ceiling of ~{ceiling(gb, bw):>5.1f} tokens/s")
# 8B Q4 on a laptop a ceiling of ~ 20.8 tokens/s
# 8B Q4 on a GPU a ceiling of ~187.5 tokens/s
# 70B Q4 on a GPU a ceiling of ~ 22.5 tokens/s
The three rows at the end explain more about local inference than any benchmark: the speed follows the bandwidth divided by the model size, and that is why a large model on a fast card is roughly as fast as a small one on a laptop.
Mastery means
- Runs a model locally
- Measures throughput and latency
- Chooses the tool and the quantisation according to the hardware
Sign in to do the exercises and build your mastery up.
Sources
- llama.cpp (MIT) — MIT
- vLLM — dokumentation (Apache-2.0) — Apache-2.0
- arXiv — Efficient Memory Management for Large Language Model Serving with PagedAttention — arXiv (open access; licence per article)