Encoder–decoder transformers
Be able to explain the T5/BART architecture and when encoder–decoder suits better than decoder-only.
Prerequisites
- DTransformers — the architecturerequired
Intuition
Three families of transformers, with different jobs:
| Type | Attention | Trained with | Examples | Good at |
|---|---|---|---|---|
| Encoder-only | full, bidirectional | masked tokens | BERT | understanding: classifying, extracting |
| Decoder-only | causal | the next token | GPT, Llama | generating |
| Encoder–decoder | the encoder full, the decoder causal plus cross | denoising / seq2seq | T5, BART, Whisper | transforming A → B |
Cross-attention is the key in the third: the decoder's queries fetch from the encoder's keys and values. The decoder therefore generates its answer while being able to «look at» the whole input the whole time — bidirectionally, in its entirety.
Formal
In an encoder–decoder block the decoder has three substeps:
- Causal self-attention over what has been generated so far: from the decoder's own states, with a triangular mask.
- Cross-attention: from the decoder, and from the encoder's output — no mask, the whole input is visible.
- A feedforward layer.
That means the input is processed bidirectionally (the encoder sees the whole source text in both directions) while the output is generated autoregressively.
Why decoder-only dominates anyway: a sufficiently large decoder-only model can solve transformation tasks by having both the input and the output in the same sequence («translate into English: … →»). That gives a single simple architecture, simpler scaling and simpler training on unstructured text. Encoder–decoder keeps its advantage when the input and the output have clearly different roles and the input is long — translation, summarisation, speech-to-text (Whisper).
Mastery means
- Explains the encoder–decoder architecture and cross-attention
- Chooses between encoder-only, decoder-only and encoder–decoder
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) — arXiv (open access; licence per article)
- arXiv — Robust Speech Recognition via Large-Scale Weak Supervision (Whisper) — arXiv (open access; licence per article)