Skip to content
AI-grafen
FAI engineeringStatistics and probability· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Causal inference

Be able to use DAGs, confounders and interventions to reason about causation in observational data.

Prerequisites

Intuition

Randomised experiments give causation for free. But often only observational data exists — and then the assumptions have to be made explicit.

Draw a DAG (a directed acyclic graph) of what affects what:

   experience
    ↙       ↘
tooling → productivity

Here experience is a confounder: it affects both whether the tool is used and how productive one is. Comparing tool users with non-users largely measures experience. The solution is to control for experience (stratify, match or include it in the model).

But do not control for everything. That leads to the next trap.

Formal

Three structures, three different decisions:

StructureThe pictureThe action
A confounder (a common cause)X←Z→YX \leftarrow Z \rightarrow Ycontrol for Z
A mediator (a chain)X→Z→YX \rightarrow Z \rightarrow Ydo not control if you want the total effect
A collider (a common effect)X→Z←YX \rightarrow Z \leftarrow Ynever control for Z — it creates a false relationship

Collider bias in practice: among admitted students the entrance test correlates negatively with grades, even though they are uncorrelated in the population — the admission is a collider. The same thing in ML: if you analyse only the users who have stayed in the service, «stayed» is a collider and every relationship within that group is distorted.

The backdoor criterion (Pearl) formalises the choice: control for a set SS that blocks every backdoor path from XX to YY without opening collider paths. Then P(Y∣do(X))=∑sP(Y∣X,s)P(s)P(Y\mid do(X)) = \sum_s P(Y\mid X, s)P(s).

Methods when there is no randomisation: matching, propensity scores, difference-in-differences, instrumental variables, regression discontinuity. All of them rest on assumptions that cannot be tested in the data — so they should be written out.

Code

import numpy as np
import statsmodels.api as sm

rng = np.random.default_rng(0); n = 5000
experience = rng.normal(0, 1, n)                               # the confounder
tooling    = (rng.normal(0, 1, n) + 0.8 * experience > 0).astype(float)
productivity = 0.3 * tooling + 1.0 * experience + rng.normal(0, 0.5, n)   # the true effect: 0.3

X1 = sm.add_constant(tooling)
print("without control:", sm.OLS(productivity, X1).fit().params[1].round(3))   # ≈ 0.95  ← 3× too high

X2 = sm.add_constant(np.column_stack([tooling, experience]))
print("with control:   ", sm.OLS(productivity, X2).fit().params[1].round(3))   # ≈ 0.30  ← right

# A collider: control for something BOTH cause → a false relationship arises
employed = (tooling + productivity + rng.normal(0, 0.3, n) > 1.2)
m = employed
print("among the employed:", np.corrcoef(tooling[m], productivity[m])[0, 1].round(3))   # negative!

Mastery means

  • Draws a DAG and identifies the confounders
  • Tells confounding from collider bias
  • Knows when observational data can give causal answers

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences