Skip to content
AI-grafen
EUniversityClassical machine learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Ensembles: random forest and boosting

Be able to train a random forest and gradient boosting and explain why ensembles generalise better.

Prerequisites

Intuition

A single decision tree is unstable: swap a few data points out and the tree looks entirely different. High variance, low bias.

Ensembles solve that in two completely different ways:

Bagging (random forest)Boosting (XGBoost and others)
The trees are trainedindependently, in parallelsequentially
Each tree seesa bootstrap sample plus randomly chosen columnsall the data, but weighted towards earlier errors
Combined byvoting or averagingsummation with a learning rate
Reducesvariancebias (and variance)
Deep trees?yes, fully grownno, shallow (3–8 levels)
Risk of overfittinglowhigh without control

Bagging rests on the average of many noisy but unbiased estimates having a lower variance. That is why every tree should be deep and overfitted — the averaging tidies up.

Boosting instead builds each new tree to correct the previous tree's errors. That is why every tree should be shallow: it is to contribute a small correction, not a whole solution.

Formal

Why bagging works. With BB independent estimators of variance σ2\sigma^2 and pairwise correlation ρ\rho, the average has the variance

ρσ2+1−ρBσ2\rho\sigma^2 + \frac{1-\rho}{B}\sigma^2

The second term disappears as BB grows — but the first does not. So it is not the number of trees that is the limitation but how correlated they are.

That is precisely why a random forest samples columns at every split, not just rows: without it every tree would choose the same strong feature at the top and be nearly identical. The column sampling (max_features) is the parameter that lowers ρ\rho, and therefore the most important one.

Gradient boosting is gradient descent in function space. Each new tree is fitted to the loss's negative gradient with respect to the current prediction:

Fm(x)=Fm−1(x)+η⋅hm(x),hm≈−∂L∂Fm−1F_{m}(x) = F_{m-1}(x) + \eta \cdot h_m(x), \qquad h_m \approx -\frac{\partial L}{\partial F_{m-1}}

The learning rate η\eta (often 0.05–0.1) and the number of trees trade against each other: halve η\eta and you need roughly twice as many trees.

Parameters that matter:

Random forestGradient boosting
n_estimators (more is never worse)n_estimators (more can overfit)
max_features — the most importantlearning_rate — the most important
min_samples_leafmax_depth (3–8)
subsample, colsample
early stopping on a validation set

The difference in the first row is decisive: in a random forest you can always add trees, in boosting the number has to be controlled.

The OOB estimate is the random forest's bonus: every tree sees only about 63 % of the data, so the remaining 37 % works as a free validation set. No separate split is needed.

Out-of-fold boosting in practice: always use early stopping against a validation set. Without it the number of trees is a guess, and if you guess too high the model overfits silently.

Code

import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier

X, y = make_classification(n_samples=4000, n_features=30, n_informative=8,
                           n_redundant=5, flip_y=0.05, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)

for name, m in (("one tree", DecisionTreeClassifier(random_state=0)),
                ("random forest", RandomForestClassifier(n_estimators=300, random_state=0,
                                                         oob_score=True)),
                ("boosting", HistGradientBoostingClassifier(random_state=0,
                                                            early_stopping=True))):
    m.fit(Xtr, ytr)
    extra = f"  OOB {m.oob_score_:.3f}" if hasattr(m, "oob_score_") else ""
    print(f"{name:<14} training {m.score(Xtr, ytr):.3f}  test {m.score(Xte, yte):.3f}{extra}")

# Why column sampling: the correlation between the trees decides
for mf in ("sqrt", 0.5, 1.0):
    rf = RandomForestClassifier(n_estimators=200, max_features=mf, random_state=0)
    print(f"  max_features={str(mf):<5} CV {cross_val_score(rf, X, y, cv=5).mean():.4f}")
#  max_features=1.0 means every tree sees every column → more correlated → worse

# The learning rate and the number of trees trade against each other
for lr, n in ((0.3, 50), (0.1, 150), (0.05, 300)):
    gb = HistGradientBoostingClassifier(learning_rate=lr, max_iter=n,
                                        early_stopping=False, random_state=0)
    print(f"  lr={lr:<5} n={n:<4} CV {cross_val_score(gb, X, y, cv=5).mean():.4f}")

# Boosting WITHOUT early stopping overfits silently
for n in (50, 200, 1000):
    gb = HistGradientBoostingClassifier(max_iter=n, learning_rate=0.1,
                                        early_stopping=False, random_state=0).fit(Xtr, ytr)
    print(f"  n={n:<5} training {gb.score(Xtr, ytr):.3f}  test {gb.score(Xte, yte):.3f}")

The last block is the point: the training score keeps going up while the test score levels off or falls. In a random forest that does not happen — there you can add trees without risk.

Mastery means

  • Explains bagging and boosting
  • Trains and tunes both
  • Knows why ensembles generalise better

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences