Sequence models for time series
Be able to train a sequence model on time series and compare it with classical methods.
Prerequisites
- ESequence models before transformersrequired
- ETime seriesrequired
Intuition
A sequence model takes a window of history and produces a forecast.
input: [y_{t-63} ... y_{t-1}] + calendar, external
output: [y_t ... y_{t+13}] ← the whole horizon at once
Two ways of forecasting several steps:
| Way | How | The problem |
|---|---|---|
| Recursive | forecast one step, feed it back, repeat | the errors accumulate |
| Direct (multi-output) | forecast the whole horizon at once | no error accumulation |
Direct is nearly always better for longer horizons, and is the standard in modern models.
But do not start here. Sequence models only pay off under certain conditions, and there are fewer of them than you would think:
| Condition | Why |
|---|---|
| Many parallel series (hundreds or more) | the model can share patterns between them |
| A long history per series | otherwise there is nothing to learn |
| External variables that matter | the network can weigh them together |
| Non-linear relationships | otherwise linear methods suffice |
With one series and three years of data, gradient boosting on lag features nearly always wins.
Formal
Architectures, in rough order of how often they are the right choice:
| Model | Strength | Weakness |
|---|---|---|
| Boosting on lag features | a strong baseline, fast, interpretable | not a sequence model at all |
| N-BEATS / N-HiTS | strong on univariate series, interpretable blocks | less flexible with external variables |
| TFT (Temporal Fusion Transformer) | handles many covariates, gives attention interpretation | heavy, many hyperparameters |
| DeepAR | probabilistic forecasts, many series | autoregressive, slow inference |
| LSTM/GRU | simple, well understood | beaten by the above |
Scaling per series is necessary when the series are at different levels: a shop selling 10 units and one selling 10 000 cannot share a model without normalisation. The standard is to divide each series by its own mean in the training window — and to do that per window, not globally, in order to avoid leakage.
Probabilistic forecasts are often more valuable than point forecasts. A forecast of 100 units says nothing about whether you should stock 100 or 140. Two ways:
| Method | Gives |
|---|---|
| Quantile regression (pinball loss) | chosen quantiles, for instance 10 %, 50 %, 90 % |
| Parametric (negative binomial, normal) | the whole distribution |
The pinball loss for quantile : . It punishes under- and over-forecasting differently, which is exactly what stock-keeping requires.
A fair comparison requires the same things here as everywhere: the same split, the same horizon, baselines included, several seeds and the spread reported. The literature in the field has historically had problems with this — the M competitions were introduced partly as an answer, and in M4 (2018) a hybrid of statistics and neural networks won, while pure neural methods ended up below the statistical baselines.
Code
import numpy as np, torch, torch.nn as nn
WINDOW, HORIZON = 63, 14
def make_windows(y, window=WINDOW, horizon=HORIZON):
"""A sliding window. y: (T,) → X: (N, window), Y: (N, horizon)."""
X, Y = [], []
for i in range(len(y) - window - horizon + 1):
X.append(y[i:i + window])
Y.append(y[i + window:i + window + horizon])
return np.array(X, dtype="float32"), np.array(Y, dtype="float32")
class Forecaster(nn.Module):
"""Direct multi-output: the whole horizon at once, no error accumulation."""
def __init__(self, window=WINDOW, horizon=HORIZON, hidden=256, quantiles=(0.1, 0.5, 0.9)):
super().__init__()
self.quantiles = quantiles
self.f = nn.Sequential(
nn.Linear(window, hidden), nn.ReLU(), nn.Dropout(0.1),
nn.Linear(hidden, hidden), nn.ReLU(),
nn.Linear(hidden, horizon * len(quantiles)),
)
self.horizon = horizon
def forward(self, x):
return self.f(x).view(-1, len(self.quantiles), self.horizon)
def pinball(pred, target, quantiles):
"""pred: (B, Q, H), target: (B, H)"""
error = target.unsqueeze(1) - pred
q = torch.tensor(quantiles, device=pred.device).view(1, -1, 1)
return torch.maximum(q * error, (q - 1) * error).mean()
# Scaling PER WINDOW — not globally, or the future leaks in
def scale(X, Y):
m = X.mean(axis=1, keepdims=True)
s = X.std(axis=1, keepdims=True) + 1e-6
return (X - m) / s, (Y - m) / s, m, s
# A chronological split, never a random one
def split(X, Y, train_share=0.7, val_share=0.15):
n = len(X)
a, b = int(n * train_share), int(n * (train_share + val_share))
return (X[:a], Y[:a]), (X[a:b], Y[a:b]), (X[b:], Y[b:])
# A fair comparison: the same split, the same horizon, baselines included
def compare(y, models: dict):
X, Y = make_windows(y)
(_, _), (_, _), (Xte, Yte) = split(X, Y)
results = {"seasonal naive": float(np.abs(Yte - Xte[:, -7:][:, :1]).mean())}
for name, fn in models.items():
results[name] = float(np.abs(Yte - fn(Xte)).mean())
naive = results["seasonal naive"]
return {k: {"MAE": round(v, 3), "MASE": round(v / naive, 3)} for k, v in results.items()}
scale per window is the detail that is most often got wrong: normalise with the whole series' mean and you have used future values to scale the training data, and the validation becomes too optimistic.
Mastery means
- Shapes time series data for a sequence model
- Compares fairly against classical baselines
- Knows when deep models pay off
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — N-BEATS: Neural basis expansion analysis for interpretable time series forecasting — arXiv (open access; licence per article)
- arXiv — Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting — arXiv (open access; licence per article)
- Hyndman & Athanasopoulos — Forecasting: Principles and Practice — free to read online