Skip to content
AI-grafen
DAI developerInference and optimisation· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

A small model or a large one?

Be able to weigh quality against cost and speed for a given task.

Prerequisites

Intuition

The largest model is nearly always the best on quality — and nearly always the wrong choice for every task.

A small modelA large model
The cost per call1×10–50×
The latency0.3–1 s2–10 s
Can be run locallyoftenrarely
Simple tasksjust as goodunnecessarily expensive
Multi-step reasoningworsenoticeably better

The question is not «which model is best?» but «which is the cheapest one that manages this particular task?»

And the answer varies between the subtasks in the same system. Classifying a question, rephrasing a search term or summarising a paragraph a small model manages excellently. Reasoning its way to why a pupil made a particular error it does not.

Formal

Measure on your own task. Public benchmarks say surprisingly little about your particular case. Build a test set of 50–100 real examples with an answer key and run every candidate against it.

QuantityHow you measure it
Qualitythe accuracy or a judge score on your test set
Costkronor per 1 000 calls, with real prompt lengths
Latencyp50 and p95 — the mean hides the slow answers
Reliabilitythe share of valid answers (correct JSON, for instance)

Routing — let a cheap model take what it can handle:

StrategyHow
Rule-basedthe question length, keywords, the task type
A classifiera small model decides the difficulty
A cascadethe small model first; escalate at low confidence

The cascade is the most profitable and the easiest to justify: run the small model, ask it to state whether it is uncertain, and send only the uncertain ones on. If 70 % get through on the small model you have cut the cost by roughly two thirds at nearly unchanged quality.

A worked example. A large model: 0.15 kr per call. A small one: 0.01 kr. 100 000 calls a month.

The set-upThe cost per month
Everything to the large one15 000 kr
Everything to the small one1 000 kr (but worse on the hard ones)
A cascade, 70 % handled by the small one0.7·1 000 + 0.3·15 000 + 1 000 = 6 200 kr

The last row counts the small model being run on every call (hence the closing 1 000) and 30 % going on to the large one.

Other routes to cutting the cost, often before you change model: a shorter context, prompt caching, fewer steps in the chain, and a cache over whole answers for recurring questions.

Code

import json, time

PRICE = {"small": {"in": 2.0, "out": 8.0}, "large": {"in": 30.0, "out": 150.0}}   # kr/M tokens

def cost(model, in_tok, out_tok):
    p = PRICE[model]
    return (in_tok * p["in"] + out_tok * p["out"]) / 1e6

def compare(models, testset, run):
    """Run the same test set against several models and report quality, cost and latency."""
    rows = []
    for m in models:
        right = kr = 0.0
        times = []
        for case in testset:
            t0 = time.perf_counter()
            answer, in_tok, out_tok = run(m, case["question"])
            times.append(time.perf_counter() - t0)
            right += float(answer.strip() == case["key"].strip())
            kr += cost(m, in_tok, out_tok)
        times.sort()
        rows.append({"model": m,
                     "quality": round(right / len(testset), 3),
                     "kr_per_1000": round(kr / len(testset) * 1000, 2),
                     "p50_s": round(times[len(times) // 2], 2),
                     "p95_s": round(times[int(0.95 * len(times))], 2)})
    return rows

# A cascade: the small model first, escalate on uncertainty
UNCERTAIN = "If you are not sure, answer exactly: UNCERTAIN"

def cascade(llm, question):
    answer = llm("small", f"{UNCERTAIN}\n\n{question}")
    if answer.strip() != "UNCERTAIN":
        return {"answer": answer, "model": "small"}
    return {"answer": llm("large", question), "model": "large"}

# A cost comparison per month
calls = 100_000
for name, large_share in (("all large", 1.0), ("all small", 0.0), ("a cascade", 0.30)):
    small_kr = calls * 0.01 if name != "all large" else 0
    large_kr = calls * large_share * 0.15
    print(f"{name:<11} {small_kr + large_kr:>9,.0f} kr/month")
# all large      15,000 kr/month
# all small       1,000 kr/month
# a cascade       5,500 kr/month

Measure p95, not just the mean. A model with 1.2 s on average but 9 s at p95 feels slow to the user, since it is the slow answers you remember.

Mastery means

  • Compares models on quality, cost and latency
  • Chooses a model by the task, not by reputation
  • Uses routing between models

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences