Prompt caching and prefix sharing
Be able to use prefix caching to cut the cost with repeated system prompts.
Prerequisites
- EThe KV cacherequired
Intuition
Every call with the same 2 000-token system prompt recomputes exactly the same KV cache. Prefix caching saves it and reuses it.
The requirement is that the prefix is identical token for token. So the prompt has to be structured correctly:
[the system prompt ] ← identical, cached
[the document/context ] ← identical per document, cached per document
[the conversation history] ← it grows but only at the end, the prefix is cached
[the user's question ] ← unique, always computed
What destroys the cache: a timestamp, a user name or a random id first in the prompt. A single character that differs and the whole prefix has to be recomputed.
Code
# Anthropic: explicit cache breakpoints
msgs = [{"role": "system", "content": [
{"type": "text", "text": SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"}},
{"type": "text", "text": DOCUMENT, "cache_control": {"type": "ephemeral"}}]},
{"role": "user", "content": question}] # the unique part last
# vLLM / llama.cpp: automatic prefix matching — nothing to configure,
# but the same requirement: an identical prefix.
def saving(system_tokens, question_tokens, calls, price_in=3.0, cache_discount=0.9):
"""the price per 1M tokens; cache_discount = the share of the price that disappears on a hit"""
without = (system_tokens + question_tokens) * calls
with_ = system_tokens + (system_tokens * (1 - cache_discount) + question_tokens) * (calls - 1)
return {"without": without, "with": round(with_),
"saved_pct": round(100 * (1 - with_ / without), 1),
"kr_saved": round((without - with_) / 1e6 * price_in, 2)}
print(saving(system_tokens=2000, question_tokens=100, calls=10_000))
# {'without': 21000000, 'with': 3200900, 'saved_pct': 84.8, 'kr_saved': 53.4}
Things to bear in mind: cache hits are usually cheaper but not free; the cache has a lifetime (minutes with most providers); and it does not affect the output quality at all — the result is identical. It is pure saving, provided the prompt is built in the right order.
Mastery means
- Uses prefix caching for repeated system prompts
- Structures the prompt so that the cache hits
- Computes the saving
Sign in to do the exercises and build your mastery up.
Sources
- Anthropic — Prompt caching — documentation, free to read
- arXiv — Efficient Memory Management for LLM Serving with PagedAttention (vLLM) — arXiv (open access; licence per article)