Ablation studies
Be able to design ablations that isolate a component's contribution, and interpret the results with uncertainty.
Prerequisites
- EReproducibilityrequired
- FEvals for language models and agentsrequired
Intuition
Your method has four parts and it beats the baseline. Which part is doing the work? Remove one at a time (an ablation) and measure. If «without part C» is as good as «everything», C is decoration. If «without A» falls to the baseline, A is carrying the result.
The rules:
- One thing at a time, everything else identical (the same data, seeds, budget, tuning).
- Several seeds per variant — ablation differences are often small and the noise large.
- Add as well, rather than only removing: baseline + A only, + B only. The order can reveal interactions (A only helps together with B).
- Substitute, not just remove: replace the component with a trivial variant (random reranking instead of none) to separate «it exists» from «it is good».
- Report in one table: variant · metric ± std · Δ against the full method.
It is the ablation that turns a result into an explanation. Project 6 (dropout) does this over p ∈ {0, 0.2, 0.5} and three seeds.
Code
import itertools, numpy as np, json
COMPONENTS = ["hybrid_retrieval", "reranking", "citation_requirement"]
def configurations(full=True, leave_one_out=True, add_one=True):
all_on = {k: True for k in COMPONENTS}
out = {"full": all_on, "baseline": {k: False for k in COMPONENTS}}
if leave_one_out:
for k in COMPONENTS: out[f"without_{k}"] = {**all_on, k: False}
if add_one:
for k in COMPONENTS: out[f"only_{k}"] = {**out["baseline"], k: True}
return out
def run_all(run_fn, seeds=(0, 1, 2)):
rows = []
for name, cfg in configurations().items():
m = np.array([run_fn(cfg, seed=s) for s in seeds])
rows.append((name, m.mean(), m.std(ddof=1)))
full = next(r for r in rows if r[0] == "full")[1]
print(f"{'variant':28s} {'metric':>8s} {'±':>6s} {'Δ':>7s}")
for name, mu, sd in rows:
print(f"{name:28s} {mu:8.3f} {sd:6.3f} {mu - full:+7.3f}")
return rows
The interpretation: a Δ smaller than about 2·std → the component has not been shown to matter. An interaction: «only_A» ≈ the baseline and «only_B» ≈ the baseline but «full» ≫ → A and B are needed together.
Mastery means
- Designs ablations that isolate a component's contribution
- Runs them with several seeds and reports the uncertainty
- Interprets the interactions between components
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Ablation Studies in Artificial Neural Networks — arXiv (open access; licence per article)
- arXiv — Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR) — arXiv (open access; licence per article)