Ensembles: random forest and boosting
Be able to train a random forest and gradient boosting and explain why ensembles generalise better.
Prerequisites
- DDecision treesrequired
Intuition
A single decision tree is unstable: swap a few data points out and the tree looks entirely different. High variance, low bias.
Ensembles solve that in two completely different ways:
| Bagging (random forest) | Boosting (XGBoost and others) | |
|---|---|---|
| The trees are trained | independently, in parallel | sequentially |
| Each tree sees | a bootstrap sample plus randomly chosen columns | all the data, but weighted towards earlier errors |
| Combined by | voting or averaging | summation with a learning rate |
| Reduces | variance | bias (and variance) |
| Deep trees? | yes, fully grown | no, shallow (3–8 levels) |
| Risk of overfitting | low | high without control |
Bagging rests on the average of many noisy but unbiased estimates having a lower variance. That is why every tree should be deep and overfitted — the averaging tidies up.
Boosting instead builds each new tree to correct the previous tree's errors. That is why every tree should be shallow: it is to contribute a small correction, not a whole solution.
Formal
Why bagging works. With independent estimators of variance and pairwise correlation , the average has the variance
The second term disappears as grows — but the first does not. So it is not the number of trees that is the limitation but how correlated they are.
That is precisely why a random forest samples columns at every split, not just rows: without it every tree would choose the same strong feature at the top and be nearly identical. The column sampling (max_features) is the parameter that lowers , and therefore the most important one.
Gradient boosting is gradient descent in function space. Each new tree is fitted to the loss's negative gradient with respect to the current prediction:
The learning rate (often 0.05–0.1) and the number of trees trade against each other: halve and you need roughly twice as many trees.
Parameters that matter:
| Random forest | Gradient boosting |
|---|---|
n_estimators (more is never worse) | n_estimators (more can overfit) |
max_features — the most important | learning_rate — the most important |
min_samples_leaf | max_depth (3–8) |
subsample, colsample | |
| early stopping on a validation set |
The difference in the first row is decisive: in a random forest you can always add trees, in boosting the number has to be controlled.
The OOB estimate is the random forest's bonus: every tree sees only about 63 % of the data, so the remaining 37 % works as a free validation set. No separate split is needed.
Out-of-fold boosting in practice: always use early stopping against a validation set. Without it the number of trees is a guess, and if you guess too high the model overfits silently.
Code
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
X, y = make_classification(n_samples=4000, n_features=30, n_informative=8,
n_redundant=5, flip_y=0.05, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
for name, m in (("one tree", DecisionTreeClassifier(random_state=0)),
("random forest", RandomForestClassifier(n_estimators=300, random_state=0,
oob_score=True)),
("boosting", HistGradientBoostingClassifier(random_state=0,
early_stopping=True))):
m.fit(Xtr, ytr)
extra = f" OOB {m.oob_score_:.3f}" if hasattr(m, "oob_score_") else ""
print(f"{name:<14} training {m.score(Xtr, ytr):.3f} test {m.score(Xte, yte):.3f}{extra}")
# Why column sampling: the correlation between the trees decides
for mf in ("sqrt", 0.5, 1.0):
rf = RandomForestClassifier(n_estimators=200, max_features=mf, random_state=0)
print(f" max_features={str(mf):<5} CV {cross_val_score(rf, X, y, cv=5).mean():.4f}")
# max_features=1.0 means every tree sees every column → more correlated → worse
# The learning rate and the number of trees trade against each other
for lr, n in ((0.3, 50), (0.1, 150), (0.05, 300)):
gb = HistGradientBoostingClassifier(learning_rate=lr, max_iter=n,
early_stopping=False, random_state=0)
print(f" lr={lr:<5} n={n:<4} CV {cross_val_score(gb, X, y, cv=5).mean():.4f}")
# Boosting WITHOUT early stopping overfits silently
for n in (50, 200, 1000):
gb = HistGradientBoostingClassifier(max_iter=n, learning_rate=0.1,
early_stopping=False, random_state=0).fit(Xtr, ytr)
print(f" n={n:<5} training {gb.score(Xtr, ytr):.3f} test {gb.score(Xte, yte):.3f}")
The last block is the point: the training score keeps going up while the test score levels off or falls. In a random forest that does not happen — there you can add trees without risk.
Mastery means
- Explains bagging and boosting
- Trains and tunes both
- Knows why ensembles generalise better
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Hastie, Tibshirani & Friedman — The Elements of Statistical Learning — free to read online (authors' edition)
- arXiv — XGBoost: A Scalable Tree Boosting System — arXiv (open access; licence per article)