Skip to content
AI-grafen
FAI engineeringInference and optimisation· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Prompt caching and prefix sharing

Be able to use prefix caching to cut the cost with repeated system prompts.

Prerequisites

Intuition

Every call with the same 2 000-token system prompt recomputes exactly the same KV cache. Prefix caching saves it and reuses it.

The requirement is that the prefix is identical token for token. So the prompt has to be structured correctly:

[the system prompt      ]  ← identical, cached
[the document/context   ]  ← identical per document, cached per document
[the conversation history]  ← it grows but only at the end, the prefix is cached
[the user's question    ]  ← unique, always computed

What destroys the cache: a timestamp, a user name or a random id first in the prompt. A single character that differs and the whole prefix has to be recomputed.

Code

# Anthropic: explicit cache breakpoints
msgs = [{"role": "system", "content": [
            {"type": "text", "text": SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"}},
            {"type": "text", "text": DOCUMENT, "cache_control": {"type": "ephemeral"}}]},
        {"role": "user", "content": question}]   # the unique part last

# vLLM / llama.cpp: automatic prefix matching — nothing to configure,
# but the same requirement: an identical prefix.

def saving(system_tokens, question_tokens, calls, price_in=3.0, cache_discount=0.9):
    """the price per 1M tokens; cache_discount = the share of the price that disappears on a hit"""
    without = (system_tokens + question_tokens) * calls
    with_   = system_tokens + (system_tokens * (1 - cache_discount) + question_tokens) * (calls - 1)
    return {"without": without, "with": round(with_),
            "saved_pct": round(100 * (1 - with_ / without), 1),
            "kr_saved": round((without - with_) / 1e6 * price_in, 2)}

print(saving(system_tokens=2000, question_tokens=100, calls=10_000))
# {'without': 21000000, 'with': 3200900, 'saved_pct': 84.8, 'kr_saved': 53.4}

Things to bear in mind: cache hits are usually cheaper but not free; the cache has a lifetime (minutes with most providers); and it does not affect the output quality at all — the result is identical. It is pure saving, provided the prompt is built in the right order.

Mastery means

  • Uses prefix caching for repeated system prompts
  • Structures the prompt so that the cache hits
  • Computes the saving

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences