Cost modelling for LLM systems
Be able to compute the cost per call and per user and set budgets and quotas.
Prerequisites
Intuition
The basic formula:
the cost per call = (in_tokens × price_in + out_tokens × price_out) / 1 000 000
Output typically costs 3–5 times more than input. But in RAG systems the input often dominates anyway, since the context is large and the answer short.
Compute over the whole chain, not per call. One «conversation» can contain several LLM calls: rephrasing the question, an embedding, reranking, generation, and perhaps a judge. It is the sum that is the cost per user interaction.
Code
PRICES = { # kronor per million tokens — example values, check the current ones
"large": {"in": 30.0, "out": 150.0},
"small": {"in": 2.0, "out": 8.0},
"embed": {"in": 0.2, "out": 0.0},
}
def cost(model, in_tok, out_tok=0):
p = PRICES[model]
return (in_tok * p["in"] + out_tok * p["out"]) / 1e6
def rag_conversation(questions_per_session=6, context=4000, answer=300, cache_hit=0.8):
per_question = (
cost("embed", 50)
+ cost("small", 800, 100) # rephrasing and reranking
+ cost("large", int(context * (1 - cache_hit * 0.9)), answer)
)
return {"per_question_kr": round(per_question, 4),
"per_session_kr": round(per_question * questions_per_session, 3),
"per_1000_users_kr": round(per_question * questions_per_session * 1000, 1)}
print(rag_conversation())
# {'per_question_kr': 0.0919, 'per_session_kr': 0.551, 'per_1000_users_kr': 551.3}
print(rag_conversation(cache_hit=0.0))
# {'per_question_kr': 0.1699, 'per_session_kr': 1.019, 'per_1000_users_kr': 1019.5}
The difference between the rows is prompt caching — nearly a halving, with no change in quality.
The four biggest cost drivers, in order:
- The context size — send only what is needed. Halve the context and you halve the cost.
- The choice of model — a small model for simple substeps (rephrasing, classification, routine answers).
- Caching — the same prefix is reused.
- The number of calls per interaction — every extra step in the chain multiplies.
Quotas are not just economics but security: without a cap a buggy client or a looping agent can burn a month's budget overnight. Set a cap per user and day, per organisation and month, and a global kill switch — exactly like the platform's LLM_DAILY_BUDGET_SEK.
Mastery means
- Computes the cost per call and per user
- Sets budgets and quotas that hold
- Identifies the biggest cost drivers
Sign in to do the exercises and build your mastery up.
Sources
- Anthropic — Prompt caching — documentation, free to read
- Google — Rules of Machine Learning — CC BY 4.0