Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Continued pre-training on domain data

Be able to adapt a base model to a domain corpus and evaluate it.

Prerequisites

Intuition

Continued pre-training is next-token training on a domain corpus — not instruction pairs. It is used when the model needs to learn a language or a domain it has seen too little of: medical terminology, legal Swedish, a programming language, a company's internal conceptual world.

The difference from fine-tuning:

Continued pre-trainingInstruction fine-tuning
The datarunning text, no structure(instruction, answer) pairs
The volume100 M–10 B tokens500–50 000 pairs
It learnsthe domain's language and factsto answer in a certain format
Followed byinstruction fine-tuning—

The order matters: continued pre-training first, then instruction fine-tuning. If you do it the other way round the pre-training erases the instruction behaviour.

Formal

Replay is not optional. If you train on domain data alone, catastrophic forgetting arises: the model loses its general ability, its instruction following and other languages. The practice is 1–10 % of the original distribution mixed into the corpus (Ibrahim et al. 2024).

The learning rate and the schedule: start from a low lr (for instance 10 % of the original pre-training's peak), with a short warm-up and a decaying schedule. A high lr «shakes loose» what the model already knows.

How much data is needed? Below ~100 M tokens the gain is rarely measurable and the risk of forgetting is significant — then LoRA or RAG are better choices. Above ~1 B tokens of real domain text it starts to pay off.

The evaluation in three suites, always:

  1. Domain perplexity on held-out domain text (it should fall).
  2. General durability (it should hold).
  3. The downstream task after instruction fine-tuning — that is the one that counts.

Perplexity as the only measure is misleading: it falls on the domain even when the model has become unusable for everything else.

Code

import random

def build_corpus(domain_documents, general_corpus, replay_share=0.05, seed=0):
    rng = random.Random(seed)
    n_replay = int(len(domain_documents) * replay_share / (1 - replay_share))
    mixed = list(domain_documents) + rng.sample(general_corpus, min(n_replay, len(general_corpus)))
    rng.shuffle(mixed)
    return mixed

CONFIG = dict(
    lr=3e-5,                 # ~10 % of the original pre-training peak
    warmup_steps=500,
    schedule="cosine",
    epochs=1,                # running text is normally run once
    replay_share=0.05,
    max_grad_norm=1.0,
)

def evaluate(model_before, model_after, suites):
    return {name: {"before": round(f(model_before), 3), "after": round(f(model_after), 3)}
            for name, f in suites.items()}

# {'domain_ppl':  {'before': 14.2, 'after': 9.8},    ← the gain
#  'general_ppl': {'before': 11.5, 'after': 12.1},   ← an acceptable loss
#  'instruction': {'before': 0.71, 'after': 0.44}}   ← expected: instruction fine-tune again afterwards

Mastery means

  • Adapts a base model to a domain corpus
  • Balances the domain data against replay
  • Measures both the domain gain and the preserved general ability

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences