Cost, quotas and pricing
Be able to set quotas and prices that cover the inference cost.
Prerequisites
- ECost modelling for LLM systemsrequired
Intuition
AI products have a cost structure that traditional software does not: every use costs money.
| Ordinary SaaS | An AI product | |
|---|---|---|
| The marginal cost per user | near zero | substantial |
| Heavy users | free marketing | can run at a loss |
| A fixed price per month | works | dangerous without quotas |
That changes the pricing fundamentally. An unlimited plan at 99 kr a month works excellently until one user runs ten thousand calls a day.
Three pricing models:
| Model | The advantage | The disadvantage |
|---|---|---|
| A fixed price plus a quota | predictable for the customer | the quota has to be set right |
| Usage-based | it follows the cost exactly | unpredictable for the customer — many dislike it |
| Credits | it combines both | it needs explaining |
The first is the most common in consumer products, the second in developer tools.
Formal
Compute on the distribution, not the mean. Usage of AI products is extremely skewed: often a few per cent of the users account for a large share of the consumption.
Setting the price by the average user gives a loss on the tail. Compute instead:
| The step | The question |
|---|---|
| 1 | What does the median user cost? |
| 2 | What does the p95 user cost? |
| 3 | What does the most expensive per cent cost? |
| 4 | Which quota covers 95 % of the users without their noticing it? |
| 5 | Does the margin hold if everyone uses the quota up? |
Step 5 is the critical one. The quota defines your maximum cost per user, and it has to be paid for by the price. A quota of 500 calls at 0.15 kr is 75 kr — a price of 99 kr then gives 24 kr in gross margin before everything else, which is too thin.
The free tier is a marketing cost and should be treated as one:
| The question | Example |
|---|---|
| What does it cost per user per month? | 3 kr |
| How many free users does the budget tolerate? | 5 000 → 15 000 kr/month |
| What is the conversion rate? | 2 % |
| What is the acquisition cost per paying customer? | 15 000 / 100 = 150 kr |
If that last figure is lower than what a paying customer is worth over time, the free tier is profitable. Otherwise it is a leak.
Quotas are also a security mechanism, not just economics. Without a cap a buggy client, a loop or an abuse can burn a month's budget overnight. Three layers are needed:
| The layer | Protects against |
|---|---|
| Per user and day | individual abuse |
| Per organisation and month | going over budget |
| A global kill switch | a catastrophe |
Five ways of cutting the cost before raising the price:
- Prompt caching — often nearly a halving in RAG systems.
- A smaller model for routine answers, escalation when needed.
- A shorter context — send only what is needed.
- An answer cache for recurring questions.
- Fewer steps in the chain.
Be open about the quotas. A quota the user did not know about and hits in the middle of their work is a worse experience than a somewhat higher price. Show the consumption continuously.
Code
import numpy as np
def cost_distribution(calls_per_user, kr_per_call=0.15):
a = np.asarray(calls_per_user)
kr = a * kr_per_call
return {
"median": round(float(np.median(kr)), 2),
"mean": round(float(kr.mean()), 2),
"p95": round(float(np.percentile(kr, 95)), 2),
"p99": round(float(np.percentile(kr, 99)), 2),
"max": round(float(kr.max()), 2),
"share_from_the_top_5_per_cent": round(
float(np.sort(kr)[-max(1, len(kr) // 20):].sum() / kr.sum()), 3),
}
# A skewed distribution: a few account for most of it
rng = np.random.default_rng(0)
calls = rng.lognormal(mean=3.5, sigma=1.6, size=5000).astype(int)
f = cost_distribution(calls)
print(f)
def margin(price, quota_calls, kr_per_call=0.15, other_cost=8.0):
"""The margin in the worst case — everybody uses the quota up."""
worst = quota_calls * kr_per_call + other_cost
return {"price": price, "worst_case_cost": round(worst, 2),
"worst_case_margin": round(price - worst, 2),
"margin_per_cent": round((price - worst) / price, 3) if price else 0.0}
for price, quota in ((99, 500), (149, 800), (249, 2000)):
print(margin(price, quota))
# {'price': 99, 'worst_case_cost': 83.0, 'worst_case_margin': 16.0, ...}
# ↑ a 16 % margin in the worst case is too thin
def choose_quota(calls_per_user, coverage=0.95):
"""A quota that 95 % of the users do not notice."""
return int(np.percentile(calls_per_user, coverage * 100))
print("a quota covering 95 %:", choose_quota(calls))
# Three layers of protection
class BudgetGuard:
def __init__(self, day_per_user, month_per_org, global_day):
self.cap = {"user_day": day_per_user,
"org_month": month_per_org, "global_day": global_day}
self.consumption = {}
def may_run(self, user, org, cost):
c = self.consumption
if c.get(("u", user), 0) + cost > self.cap["user_day"]:
return False, "the daily quota is used up — it resets tomorrow"
if c.get(("o", org), 0) + cost > self.cap["org_month"]:
return False, "the organisation's monthly budget is used up"
if c.get(("g",), 0) + cost > self.cap["global_day"]:
return False, "the service is temporarily limited"
return True, "ok"
def record(self, user, org, cost):
for key in (("u", user), ("o", org), ("g",)):
self.consumption[key] = self.consumption.get(key, 0) + cost
The error messages in may_run are deliberately different. «The daily quota is used up — it resets tomorrow» is useful; «error 429» is not. Quotas the user understands are experienced as reasonable; quotas that merely stop are experienced as broken.
Mastery means
- Computes the cost per user
- Sets quotas that protect the margin
- Chooses a pricing model according to the cost structure
Sign in to do the exercises and build your mastery up.
Sources
- Anthropic — Prompt caching — documentation, free to read
- Google — Rules of Machine Learning — CC BY 4.0
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0