EUniversityRAG and information retrieval· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Chunking documents
Be able to choose a chunk size and an overlap and measure the effect on retrieval.
Prerequisites
- DRetrieval — finding the right textrequired
Intuition
Documents have to be split into pieces before they are embedded. The chunking decides more of the RAG quality than the choice of embedding model.
The trade-off:
| Chunks too small (< 100 tokens) | Too large (> 1 000 tokens) |
|---|---|
| they lack context | the embedding becomes an averaged soup |
| many hits are needed | irrelevant text takes up room in the context |
| high precision per hit | hard to know what produced the hit |
The default choice: 200–500 tokens with a 10–20 % overlap. But the best is to split by structure — a heading, a paragraph, a section — instead of blindly by length. A chunk that is a whole section is nearly always better than one that is 400 tokens and stops in the middle of a sentence.
Code
import re
def chunk_by_heading(md, max_words=400, overlap=60):
"""Split on headings; long sections are split further with an overlap. Keeps the heading chain as metadata."""
parts, heading = [], ""
for block in re.split(r"\n(?=#{1,3} )", md):
m = re.match(r"(#{1,3}) (.+)", block)
if m:
heading = m.group(2).strip()
words = block.split()
if len(words) <= max_words:
parts.append({"heading": heading, "text": block.strip()})
else:
step = max_words - overlap
for i in range(0, len(words), step):
parts.append({"heading": heading, "text": " ".join(words[i:i + max_words])})
return [p for p in parts if p["text"]]
# Put the heading in the text that is embedded — it carries a lot of context
for p in chunk_by_heading(document):
p["embed_text"] = f"{p['heading']}\n\n{p['text']}"
Three things often forgotten:
- The metadata comes along: the source, the URL, the heading, the date, the licence — otherwise the answer cannot cite anything.
- Tables and code blocks should not be split down the middle.
- Measure: run your retrieval test set with three chunking strategies and compare recall@10. It takes an hour and is the most profitable hour in a RAG project.
Mastery means
- Chooses a chunk size and an overlap with a justification
- Preserves the metadata and the structure
- Measures the effect on retrieval
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — arXiv (open access; licence per article)
- Qdrant — dokumentation (Apache-2.0) — Apache-2.0