Continued pre-training on domain data
Be able to adapt a base model to a domain corpus and evaluate it.
Prerequisites
- FPre-training in practicerequired
Intuition
Continued pre-training is next-token training on a domain corpus — not instruction pairs. It is used when the model needs to learn a language or a domain it has seen too little of: medical terminology, legal Swedish, a programming language, a company's internal conceptual world.
The difference from fine-tuning:
| Continued pre-training | Instruction fine-tuning | |
|---|---|---|
| The data | running text, no structure | (instruction, answer) pairs |
| The volume | 100 M–10 B tokens | 500–50 000 pairs |
| It learns | the domain's language and facts | to answer in a certain format |
| Followed by | instruction fine-tuning | — |
The order matters: continued pre-training first, then instruction fine-tuning. If you do it the other way round the pre-training erases the instruction behaviour.
Formal
Replay is not optional. If you train on domain data alone, catastrophic forgetting arises: the model loses its general ability, its instruction following and other languages. The practice is 1–10 % of the original distribution mixed into the corpus (Ibrahim et al. 2024).
The learning rate and the schedule: start from a low lr (for instance 10 % of the original pre-training's peak), with a short warm-up and a decaying schedule. A high lr «shakes loose» what the model already knows.
How much data is needed? Below ~100 M tokens the gain is rarely measurable and the risk of forgetting is significant — then LoRA or RAG are better choices. Above ~1 B tokens of real domain text it starts to pay off.
The evaluation in three suites, always:
- Domain perplexity on held-out domain text (it should fall).
- General durability (it should hold).
- The downstream task after instruction fine-tuning — that is the one that counts.
Perplexity as the only measure is misleading: it falls on the domain even when the model has become unusable for everything else.
Code
import random
def build_corpus(domain_documents, general_corpus, replay_share=0.05, seed=0):
rng = random.Random(seed)
n_replay = int(len(domain_documents) * replay_share / (1 - replay_share))
mixed = list(domain_documents) + rng.sample(general_corpus, min(n_replay, len(general_corpus)))
rng.shuffle(mixed)
return mixed
CONFIG = dict(
lr=3e-5, # ~10 % of the original pre-training peak
warmup_steps=500,
schedule="cosine",
epochs=1, # running text is normally run once
replay_share=0.05,
max_grad_norm=1.0,
)
def evaluate(model_before, model_after, suites):
return {name: {"before": round(f(model_before), 3), "after": round(f(model_after), 3)}
for name, f in suites.items()}
# {'domain_ppl': {'before': 14.2, 'after': 9.8}, ← the gain
# 'general_ppl': {'before': 11.5, 'after': 12.1}, ← an acceptable loss
# 'instruction': {'before': 0.71, 'after': 0.44}} ← expected: instruction fine-tune again afterwards
Mastery means
- Adapts a base model to a domain corpus
- Balances the domain data against replay
- Measures both the domain gain and the preserved general ability
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Simple and Scalable Strategies to Continually Pre-train Large Language Models — arXiv (open access; licence per article)
- arXiv — Don't Stop Pretraining: Adapt Language Models to Domains and Tasks — arXiv (open access; licence per article)