Skip to content
AI-grafen
EUniversityRAG and information retrieval· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Chunking documents

Be able to choose a chunk size and an overlap and measure the effect on retrieval.

Prerequisites

Intuition

Documents have to be split into pieces before they are embedded. The chunking decides more of the RAG quality than the choice of embedding model.

The trade-off:

Chunks too small (< 100 tokens)Too large (> 1 000 tokens)
they lack contextthe embedding becomes an averaged soup
many hits are neededirrelevant text takes up room in the context
high precision per hithard to know what produced the hit

The default choice: 200–500 tokens with a 10–20 % overlap. But the best is to split by structure — a heading, a paragraph, a section — instead of blindly by length. A chunk that is a whole section is nearly always better than one that is 400 tokens and stops in the middle of a sentence.

Code

import re

def chunk_by_heading(md, max_words=400, overlap=60):
    """Split on headings; long sections are split further with an overlap. Keeps the heading chain as metadata."""
    parts, heading = [], ""
    for block in re.split(r"\n(?=#{1,3} )", md):
        m = re.match(r"(#{1,3}) (.+)", block)
        if m:
            heading = m.group(2).strip()
        words = block.split()
        if len(words) <= max_words:
            parts.append({"heading": heading, "text": block.strip()})
        else:
            step = max_words - overlap
            for i in range(0, len(words), step):
                parts.append({"heading": heading, "text": " ".join(words[i:i + max_words])})
    return [p for p in parts if p["text"]]

# Put the heading in the text that is embedded — it carries a lot of context
for p in chunk_by_heading(document):
    p["embed_text"] = f"{p['heading']}\n\n{p['text']}"

Three things often forgotten:

  1. The metadata comes along: the source, the URL, the heading, the date, the licence — otherwise the answer cannot cite anything.
  2. Tables and code blocks should not be split down the middle.
  3. Measure: run your retrieval test set with three chunking strategies and compare recall@10. It takes an hour and is the most profitable hour in a RAG project.

Mastery means

  • Chooses a chunk size and an overlap with a justification
  • Preserves the metadata and the structure
  • Measures the effect on retrieval

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences