Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Sequence models for time series

Be able to train a sequence model on time series and compare it with classical methods.

Prerequisites

Intuition

A sequence model takes a window of history and produces a forecast.

input:  [y_{t-63} ... y_{t-1}]  +  calendar, external
output: [y_t ... y_{t+13}]         ← the whole horizon at once

Two ways of forecasting several steps:

WayHowThe problem
Recursiveforecast one step, feed it back, repeatthe errors accumulate
Direct (multi-output)forecast the whole horizon at onceno error accumulation

Direct is nearly always better for longer horizons, and is the standard in modern models.

But do not start here. Sequence models only pay off under certain conditions, and there are fewer of them than you would think:

ConditionWhy
Many parallel series (hundreds or more)the model can share patterns between them
A long history per seriesotherwise there is nothing to learn
External variables that matterthe network can weigh them together
Non-linear relationshipsotherwise linear methods suffice

With one series and three years of data, gradient boosting on lag features nearly always wins.

Formal

Architectures, in rough order of how often they are the right choice:

ModelStrengthWeakness
Boosting on lag featuresa strong baseline, fast, interpretablenot a sequence model at all
N-BEATS / N-HiTSstrong on univariate series, interpretable blocksless flexible with external variables
TFT (Temporal Fusion Transformer)handles many covariates, gives attention interpretationheavy, many hyperparameters
DeepARprobabilistic forecasts, many seriesautoregressive, slow inference
LSTM/GRUsimple, well understoodbeaten by the above

Scaling per series is necessary when the series are at different levels: a shop selling 10 units and one selling 10 000 cannot share a model without normalisation. The standard is to divide each series by its own mean in the training window — and to do that per window, not globally, in order to avoid leakage.

Probabilistic forecasts are often more valuable than point forecasts. A forecast of 100 units says nothing about whether you should stock 100 or 140. Two ways:

MethodGives
Quantile regression (pinball loss)chosen quantiles, for instance 10 %, 50 %, 90 %
Parametric (negative binomial, normal)the whole distribution

The pinball loss for quantile qq: max⁡(q(y−y^), (q−1)(y−y^))\max(q(y-\hat{y}),\, (q-1)(y-\hat{y})). It punishes under- and over-forecasting differently, which is exactly what stock-keeping requires.

A fair comparison requires the same things here as everywhere: the same split, the same horizon, baselines included, several seeds and the spread reported. The literature in the field has historically had problems with this — the M competitions were introduced partly as an answer, and in M4 (2018) a hybrid of statistics and neural networks won, while pure neural methods ended up below the statistical baselines.

Code

import numpy as np, torch, torch.nn as nn

WINDOW, HORIZON = 63, 14

def make_windows(y, window=WINDOW, horizon=HORIZON):
    """A sliding window. y: (T,) → X: (N, window), Y: (N, horizon)."""
    X, Y = [], []
    for i in range(len(y) - window - horizon + 1):
        X.append(y[i:i + window])
        Y.append(y[i + window:i + window + horizon])
    return np.array(X, dtype="float32"), np.array(Y, dtype="float32")

class Forecaster(nn.Module):
    """Direct multi-output: the whole horizon at once, no error accumulation."""
    def __init__(self, window=WINDOW, horizon=HORIZON, hidden=256, quantiles=(0.1, 0.5, 0.9)):
        super().__init__()
        self.quantiles = quantiles
        self.f = nn.Sequential(
            nn.Linear(window, hidden), nn.ReLU(), nn.Dropout(0.1),
            nn.Linear(hidden, hidden), nn.ReLU(),
            nn.Linear(hidden, horizon * len(quantiles)),
        )
        self.horizon = horizon

    def forward(self, x):
        return self.f(x).view(-1, len(self.quantiles), self.horizon)

def pinball(pred, target, quantiles):
    """pred: (B, Q, H), target: (B, H)"""
    error = target.unsqueeze(1) - pred
    q = torch.tensor(quantiles, device=pred.device).view(1, -1, 1)
    return torch.maximum(q * error, (q - 1) * error).mean()

# Scaling PER WINDOW — not globally, or the future leaks in
def scale(X, Y):
    m = X.mean(axis=1, keepdims=True)
    s = X.std(axis=1, keepdims=True) + 1e-6
    return (X - m) / s, (Y - m) / s, m, s

# A chronological split, never a random one
def split(X, Y, train_share=0.7, val_share=0.15):
    n = len(X)
    a, b = int(n * train_share), int(n * (train_share + val_share))
    return (X[:a], Y[:a]), (X[a:b], Y[a:b]), (X[b:], Y[b:])

# A fair comparison: the same split, the same horizon, baselines included
def compare(y, models: dict):
    X, Y = make_windows(y)
    (_, _), (_, _), (Xte, Yte) = split(X, Y)
    results = {"seasonal naive": float(np.abs(Yte - Xte[:, -7:][:, :1]).mean())}
    for name, fn in models.items():
        results[name] = float(np.abs(Yte - fn(Xte)).mean())
    naive = results["seasonal naive"]
    return {k: {"MAE": round(v, 3), "MASE": round(v / naive, 3)} for k, v in results.items()}

scale per window is the detail that is most often got wrong: normalise with the whole series' mean and you have used future values to scale the training data, and the validation becomes too optimistic.

Mastery means

  • Shapes time series data for a sequence model
  • Compares fairly against classical baselines
  • Knows when deep models pay off

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences