Skip to content
AI-grafen
FAI engineeringLanguage models· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Long context versus retrieval

Be able to decide when long context and when retrieval is the right tool.

Prerequisites

Intuition

Models with a 128 k–1 M token context raise the question: is RAG still needed? The answer is «yes, usually» — for four reasons.

Long contextRetrieval
The costyou pay for everything, every callyou pay for the 5 k relevant ones
The latencythe prefill grows quadraticallynearly constant
The qualityit falls in the middle of long contextsa focused context
Updatingit has to be resent every timeswap the document in the index
Traceabilityhard to know what was useda citation per passage

Lost in the middle (Liu et al. 2023): models find information reliably at the beginning and the end of the context but clearly worse in the middle. Having 200 000 tokens available does not mean the model uses them evenly.

Formal

A cost comparison for 1 000 questions against 500 k tokens of document material:

The approachInput tokens per questionIn totalRelative
Everything in the context (if it fitted)500 000500 M100×
Long context with a selection, 50 k50 00050 M10×
RAG, the top 8 chunks at 400~5 0005 M1×

With prompt caching the difference shrinks if the same long context is reused, but not when the material varies per question.

When long context wins anyway:

  • The task requires the whole document (summarise a report, find contradictions between sections).
  • The material is small enough (< 20 k tokens) that retrieval only adds complexity.
  • The questions are aggregating («how many times is X mentioned?») where retrieval systematically misses.

The most common right solution is a hybrid: retrieval picks 20–50 k relevant tokens out of a large material, and the long context lets the model reason over them together. You do not have to choose.

Code

def cost_per_question(approach, document_tokens, k_chunks=8, chunk=400, context_ceiling=128_000):
    if approach == "everything":
        if document_tokens > context_ceiling:
            return None                      # it does not fit
        return document_tokens
    if approach == "rag":
        return k_chunks * chunk + 500        # + the system prompt and the question
    if approach == "hybrid":
        return min(50_000, document_tokens)  # retrieval picks a large but focused selection

for doc in (20_000, 200_000, 2_000_000):
    row = {a: cost_per_question(a, doc) for a in ("everything", "rag", "hybrid")}
    print(f"{doc:>9,} tokens: {row}")
#    20,000 tokens: {'everything': 20000, 'rag': 3700, 'hybrid': 20000}
#   200,000 tokens: {'everything': None,  'rag': 3700, 'hybrid': 50000}
# 2,000,000 tokens: {'everything': None,  'rag': 3700, 'hybrid': 50000}

# Measure lost-in-the-middle on your own model: place the same fact at
# position 10 %, 50 % and 90 % in a long context and measure the recall.

Mastery means

  • Compares long context and retrieval on cost and quality
  • Knows about lost-in-the-middle
  • Chooses and combines the right one for the task

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences