A small model or a large one?
Be able to weigh quality against cost and speed for a given task.
Prerequisites
Intuition
The largest model is nearly always the best on quality — and nearly always the wrong choice for every task.
| A small model | A large model | |
|---|---|---|
| The cost per call | 1× | 10–50× |
| The latency | 0.3–1 s | 2–10 s |
| Can be run locally | often | rarely |
| Simple tasks | just as good | unnecessarily expensive |
| Multi-step reasoning | worse | noticeably better |
The question is not «which model is best?» but «which is the cheapest one that manages this particular task?»
And the answer varies between the subtasks in the same system. Classifying a question, rephrasing a search term or summarising a paragraph a small model manages excellently. Reasoning its way to why a pupil made a particular error it does not.
Formal
Measure on your own task. Public benchmarks say surprisingly little about your particular case. Build a test set of 50–100 real examples with an answer key and run every candidate against it.
| Quantity | How you measure it |
|---|---|
| Quality | the accuracy or a judge score on your test set |
| Cost | kronor per 1 000 calls, with real prompt lengths |
| Latency | p50 and p95 — the mean hides the slow answers |
| Reliability | the share of valid answers (correct JSON, for instance) |
Routing — let a cheap model take what it can handle:
| Strategy | How |
|---|---|
| Rule-based | the question length, keywords, the task type |
| A classifier | a small model decides the difficulty |
| A cascade | the small model first; escalate at low confidence |
The cascade is the most profitable and the easiest to justify: run the small model, ask it to state whether it is uncertain, and send only the uncertain ones on. If 70 % get through on the small model you have cut the cost by roughly two thirds at nearly unchanged quality.
A worked example. A large model: 0.15 kr per call. A small one: 0.01 kr. 100 000 calls a month.
| The set-up | The cost per month |
|---|---|
| Everything to the large one | 15 000 kr |
| Everything to the small one | 1 000 kr (but worse on the hard ones) |
| A cascade, 70 % handled by the small one | 0.7·1 000 + 0.3·15 000 + 1 000 = 6 200 kr |
The last row counts the small model being run on every call (hence the closing 1 000) and 30 % going on to the large one.
Other routes to cutting the cost, often before you change model: a shorter context, prompt caching, fewer steps in the chain, and a cache over whole answers for recurring questions.
Code
import json, time
PRICE = {"small": {"in": 2.0, "out": 8.0}, "large": {"in": 30.0, "out": 150.0}} # kr/M tokens
def cost(model, in_tok, out_tok):
p = PRICE[model]
return (in_tok * p["in"] + out_tok * p["out"]) / 1e6
def compare(models, testset, run):
"""Run the same test set against several models and report quality, cost and latency."""
rows = []
for m in models:
right = kr = 0.0
times = []
for case in testset:
t0 = time.perf_counter()
answer, in_tok, out_tok = run(m, case["question"])
times.append(time.perf_counter() - t0)
right += float(answer.strip() == case["key"].strip())
kr += cost(m, in_tok, out_tok)
times.sort()
rows.append({"model": m,
"quality": round(right / len(testset), 3),
"kr_per_1000": round(kr / len(testset) * 1000, 2),
"p50_s": round(times[len(times) // 2], 2),
"p95_s": round(times[int(0.95 * len(times))], 2)})
return rows
# A cascade: the small model first, escalate on uncertainty
UNCERTAIN = "If you are not sure, answer exactly: UNCERTAIN"
def cascade(llm, question):
answer = llm("small", f"{UNCERTAIN}\n\n{question}")
if answer.strip() != "UNCERTAIN":
return {"answer": answer, "model": "small"}
return {"answer": llm("large", question), "model": "large"}
# A cost comparison per month
calls = 100_000
for name, large_share in (("all large", 1.0), ("all small", 0.0), ("a cascade", 0.30)):
small_kr = calls * 0.01 if name != "all large" else 0
large_kr = calls * large_share * 0.15
print(f"{name:<11} {small_kr + large_kr:>9,.0f} kr/month")
# all large 15,000 kr/month
# all small 1,000 kr/month
# a cascade 5,500 kr/month
Measure p95, not just the mean. A model with 1.2 s on average but 9 s at p95 feels slow to the user, since it is the slow answers you remember.
Mastery means
- Compares models on quality, cost and latency
- Chooses a model by the task, not by reputation
- Uses routing between models
Sign in to do the exercises and build your mastery up.
Sources
- Anthropic — Prompt caching — documentation, free to read
- arXiv — FrugalGPT: How to Use Large Language Models While Reducing Cost — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0