Skip to content
AI-grafen
GFrontier LabTransformer architecture· about 120 min· fast-moving, sources checked often· verified 2026-09-20· EN

State space models and Mamba

Be able to compare SSM architectures with attention and read current research critically.

Prerequisites

Intuition

State space models (SSMs), like RNNs, carry a state forward through the sequence. That gives a completely different cost profile from attention:

AttentionSSM
TrainingO(T²), parallelO(T log T) or O(T), parallel through a scan
Generation per tokenO(T) with a KV cacheO(1) — constant
Memory at generationgrows with Tconstant
Exact recall of an arbitrary detailyeslimited by the size of the state

The last row is the decisive trade-off: attention can go back and look at any token at all; an SSM has to have saved it in its state.

Research

Mamba's contribution (Gu & Dao 2023) is selectivity: the SSM parameters are made dependent on the input, so the model can choose what goes into the state and what is forgotten. Earlier SSMs (S4) had fixed parameters and could not do content-based selection — which was exactly what was missing for language. The price is that the efficient convolutional form disappears, which is solved with a hardware-aware parallel scan.

What the evaluations actually show: pure SSMs match transformers on many language tasks at a comparable scale, but perform worse on tasks requiring exact copying or lookup in the context (induction tasks, «needle in a haystack»). Jelassi et al. (2024) showed this theoretically and empirically: a constant state means an information-theoretic limit on how much can be recalled.

That is why hybrid architectures are the standard in practice — Jamba, Zamba and the like mix a few attention layers in among many Mamba layers. A handful of attention layers is enough to restore the lookup ability, while the bulk of the cost becomes linear.

How to read such claims critically: «matches transformers» holds on which tasks, at which scale, with what tuning of the baseline? Check particularly whether the evaluation contains tasks requiring exact recall — that is where the difference shows, and that is why they are sometimes left out.

Formal

A linear SSM in continuous time: h′(t)=Ah(t)+Bx(t),y(t)=Ch(t)+Dx(t)h'(t) = A h(t) + B x(t), \qquad y(t) = C h(t) + D x(t)

Discretised with a step length Δ\Delta it becomes a recurrence: ht=Aˉht−1+Bˉxt,yt=Chth_t = \bar A h_{t-1} + \bar B x_t, \qquad y_t = C h_t

Since the recurrence is linear in hh the whole sequence can be computed with an associative scan at O(log⁡T)O(\log T) depth — unlike a non-linear RNN, which has to be run sequentially. That is what makes SSMs trainable at scale.

In S4 Aˉ\bar A is structured (HiPPO initialisation) and the parameters are independent of the input, which gives a global convolution. In Mamba Bˉ\bar B, CC and Δ\Delta are made functions of xtx_t — selectivity — which breaks the convolutional form but keeps the scan parallelisation.

Mastery means

  • Compares SSMs with attention in cost and ability
  • Explains selectivity in Mamba
  • Reads current architecture claims critically

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences