Skip to content
AI-grafen
GFrontier LabScientific method· about 120 min· fundamentals that rarely change· verified 2026-09-20· EN

Ablation studies

Be able to design ablations that isolate a component's contribution, and interpret the results with uncertainty.

Prerequisites

Intuition

Your method has four parts and it beats the baseline. Which part is doing the work? Remove one at a time (an ablation) and measure. If «without part C» is as good as «everything», C is decoration. If «without A» falls to the baseline, A is carrying the result.

The rules:

  • One thing at a time, everything else identical (the same data, seeds, budget, tuning).
  • Several seeds per variant — ablation differences are often small and the noise large.
  • Add as well, rather than only removing: baseline + A only, + B only. The order can reveal interactions (A only helps together with B).
  • Substitute, not just remove: replace the component with a trivial variant (random reranking instead of none) to separate «it exists» from «it is good».
  • Report in one table: variant · metric ± std · Δ against the full method.

It is the ablation that turns a result into an explanation. Project 6 (dropout) does this over p ∈ {0, 0.2, 0.5} and three seeds.

Code

import itertools, numpy as np, json

COMPONENTS = ["hybrid_retrieval", "reranking", "citation_requirement"]

def configurations(full=True, leave_one_out=True, add_one=True):
    all_on = {k: True for k in COMPONENTS}
    out = {"full": all_on, "baseline": {k: False for k in COMPONENTS}}
    if leave_one_out:
        for k in COMPONENTS: out[f"without_{k}"] = {**all_on, k: False}
    if add_one:
        for k in COMPONENTS: out[f"only_{k}"] = {**out["baseline"], k: True}
    return out

def run_all(run_fn, seeds=(0, 1, 2)):
    rows = []
    for name, cfg in configurations().items():
        m = np.array([run_fn(cfg, seed=s) for s in seeds])
        rows.append((name, m.mean(), m.std(ddof=1)))
    full = next(r for r in rows if r[0] == "full")[1]
    print(f"{'variant':28s} {'metric':>8s} {'±':>6s} {'Δ':>7s}")
    for name, mu, sd in rows:
        print(f"{name:28s} {mu:8.3f} {sd:6.3f} {mu - full:+7.3f}")
    return rows

The interpretation: a Δ smaller than about 2·std → the component has not been shown to matter. An interaction: «only_A» ≈ the baseline and «only_B» ≈ the baseline but «full» ≫ → A and B are needed together.

Mastery means

  • Designs ablations that isolate a component's contribution
  • Runs them with several seeds and reports the uncertainty
  • Interprets the interactions between components

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences