Time series
Be able to split time series correctly, build lag features and avoid leaking the future.
Prerequisites
Intuition
Time series break the assumption all other machine learning rests on: that the observations are independent and exchangeable. That has three consequences.
1. A random split is wrong. Train on randomly chosen days and test on the rest and the model has seen the future. Split chronologically:
|--------- training ---------|--- val ---|--- test ---|
2024 2025 2026
2. The baseline is harder than you would think. «Tomorrow will be like today» (the naive forecast) and «the same weekday last week» (the seasonal naive) are surprisingly strong. If your model does not beat them it has no value.
3. Leaking the future is easy to fall into. A rolling mean that happens to include the current or coming values gives fantastic results in the evaluation and worthless ones in production.
Formal
Features that do not leak:
| Feature | Construction | The trap |
|---|---|---|
| Lag | y.shift(k) with k ≥ the forecast horizon | shift(1) when you are forecasting 7 days ahead is leakage |
| Rolling mean | y.shift(1).rolling(w).mean() | rolling(w) without shift includes the current value |
| Difference | y.shift(1) - y.shift(2) | |
| Calendar | weekday, month, public holiday | no leakage — known in advance |
| External | weather, a campaign | has to be known at forecast time |
The rule: every feature has to be computable from only the information that existed at forecast time. Write the point in time down and ask for every column: did we know this then?
Cross-validation is done with an expanding or a sliding window, never with KFold:
fold 1: [train....] [test]
fold 2: [train.......] [test]
fold 3: [train..........] [test]
sklearn.model_selection.TimeSeriesSplit does this. Also put a gap between training and test if the target depends on several days back, otherwise it leaks across the boundary.
Metrics, and when they mislead:
| Metric | Good at | The trap |
|---|---|---|
| MAE | interpretable in the unit | |
| RMSE | punishes large errors | sensitive to outliers |
| MAPE | percentage-based, comparable | explodes when the true value is near 0 |
| sMAPE | symmetric | still unstable at small values |
| MASE | scaled against the naive forecast | the most honest — 1.0 means «as good as naive» |
MASE is preferable precisely because it has an obvious reference point: above 1.0 the model is worse than guessing yesterday's value.
Choosing a method: start with the classics. ARIMA, exponential smoothing and especially gradient boosting on lag features are hard to beat on most real series. Deep sequence models only win with many parallel series and a lot of data — and in the M competitions simple statistical methods have long stayed at the top.
Code
import numpy as np, pandas as pd
from sklearn.model_selection import TimeSeriesSplit
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.metrics import mean_absolute_error
rng = np.random.default_rng(0)
n = 730
idx = pd.date_range("2024-01-01", periods=n, freq="D")
trend = np.linspace(100, 140, n)
weekly = 12 * np.sin(2 * np.pi * np.arange(n) / 7)
y = trend + weekly + rng.normal(0, 5, n)
df = pd.DataFrame({"y": y}, index=idx)
HORIZON = 7 # we are forecasting 7 days ahead
def build_features(df, horizon=HORIZON):
d = df.copy()
# Every lag has to be at least `horizon` steps back — otherwise leakage
for lag in (horizon, horizon + 1, horizon + 7, horizon + 14):
d[f"lag_{lag}"] = d["y"].shift(lag)
for w in (7, 28):
d[f"mean_{w}"] = d["y"].shift(horizon).rolling(w).mean()
d[f"std_{w}"] = d["y"].shift(horizon).rolling(w).std()
d["weekday"] = d.index.dayofweek
d["month"] = d.index.month
return d.dropna()
d = build_features(df)
X, target = d.drop(columns="y"), d["y"]
# The baselines first — they have to be beaten
naive = d["lag_7"] # "the same day last week"
print(f"seasonal naive MAE: {mean_absolute_error(target, naive):.2f}")
# Chronological cross-validation with a gap
tscv = TimeSeriesSplit(n_splits=5, gap=HORIZON)
errors = []
for tr, te in tscv.split(X):
m = HistGradientBoostingRegressor(random_state=0).fit(X.iloc[tr], target.iloc[tr])
errors.append(mean_absolute_error(target.iloc[te], m.predict(X.iloc[te])))
print(f"model MAE: {np.mean(errors):.2f} ± {np.std(errors):.2f}")
# MASE: scaled against the naive forecast — 1.0 means "as good as naive"
def mase(true, pred, naive_pred):
return mean_absolute_error(true, pred) / mean_absolute_error(true, naive_pred)
# THIS IS WHAT LEAKAGE LOOKS LIKE — and how good it looks in the evaluation
leaked = df.copy()
leaked["mean_7_LEAKS"] = leaked["y"].rolling(7).mean() # no shift!
print("the correlation between the leaking feature and the target:",
round(float(leaked[["y", "mean_7_LEAKS"]].dropna().corr().iloc[0, 1]), 4))
# ~0.95 — the model effectively gets to see the answer, and the MAE becomes unrealistically low
The last block is worth running. A rolling mean feature without shift correlates nearly perfectly with the target, gives a fantastic validation result, and does not work at all in production — because the value it builds on does not exist when the forecast has to be made.
Mastery means
- Splits time series chronologically
- Builds lag and rolling features without leakage
- Chooses the baseline and the metric deliberately
Sign in to do the exercises and build your mastery up.
Sources
- Hyndman & Athanasopoulos — Forecasting: Principles and Practice — free to read online
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- pandas — User Guide (BSD-3) — BSD-3-Clause