Long context versus retrieval
Be able to decide when long context and when retrieval is the right tool.
Prerequisites
Intuition
Models with a 128 k–1 M token context raise the question: is RAG still needed? The answer is «yes, usually» — for four reasons.
| Long context | Retrieval | |
|---|---|---|
| The cost | you pay for everything, every call | you pay for the 5 k relevant ones |
| The latency | the prefill grows quadratically | nearly constant |
| The quality | it falls in the middle of long contexts | a focused context |
| Updating | it has to be resent every time | swap the document in the index |
| Traceability | hard to know what was used | a citation per passage |
Lost in the middle (Liu et al. 2023): models find information reliably at the beginning and the end of the context but clearly worse in the middle. Having 200 000 tokens available does not mean the model uses them evenly.
Formal
A cost comparison for 1 000 questions against 500 k tokens of document material:
| The approach | Input tokens per question | In total | Relative |
|---|---|---|---|
| Everything in the context (if it fitted) | 500 000 | 500 M | 100× |
| Long context with a selection, 50 k | 50 000 | 50 M | 10× |
| RAG, the top 8 chunks at 400 | ~5 000 | 5 M | 1× |
With prompt caching the difference shrinks if the same long context is reused, but not when the material varies per question.
When long context wins anyway:
- The task requires the whole document (summarise a report, find contradictions between sections).
- The material is small enough (< 20 k tokens) that retrieval only adds complexity.
- The questions are aggregating («how many times is X mentioned?») where retrieval systematically misses.
The most common right solution is a hybrid: retrieval picks 20–50 k relevant tokens out of a large material, and the long context lets the model reason over them together. You do not have to choose.
Code
def cost_per_question(approach, document_tokens, k_chunks=8, chunk=400, context_ceiling=128_000):
if approach == "everything":
if document_tokens > context_ceiling:
return None # it does not fit
return document_tokens
if approach == "rag":
return k_chunks * chunk + 500 # + the system prompt and the question
if approach == "hybrid":
return min(50_000, document_tokens) # retrieval picks a large but focused selection
for doc in (20_000, 200_000, 2_000_000):
row = {a: cost_per_question(a, doc) for a in ("everything", "rag", "hybrid")}
print(f"{doc:>9,} tokens: {row}")
# 20,000 tokens: {'everything': 20000, 'rag': 3700, 'hybrid': 20000}
# 200,000 tokens: {'everything': None, 'rag': 3700, 'hybrid': 50000}
# 2,000,000 tokens: {'everything': None, 'rag': 3700, 'hybrid': 50000}
# Measure lost-in-the-middle on your own model: place the same fact at
# position 10 %, 50 % and 90 % in a long context and measure the recall.
Mastery means
- Compares long context and retrieval on cost and quality
- Knows about lost-in-the-middle
- Chooses and combines the right one for the task
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Lost in the Middle: How Language Models Use Long Contexts — arXiv (open access; licence per article)
- arXiv — Retrieval meets Long Context Large Language Models — arXiv (open access; licence per article)