Positional encoding and RoPE
Be able to explain sinusoidal, learnt and rotary positional encoding, and implement RoPE.
Prerequisites
- DAttentionrequired
- DTransformers — the architecturerequired
Intuition
Attention is permutation-invariant: shuffle the order of the tokens and you get the same set of outputs in a new order. «The dog bit the man» and «the man bit the dog» would look the same. The position therefore has to be added.
- Sinusoidal encoding (the original): a fixed vector pattern per position of sines and cosines at different frequencies. No learning, and it can in principle extrapolate.
- Learnt: one embedding per position, as for tokens. Simple, but it cannot go beyond the trained length.
- RoPE (rotary): instead of adding the position to x, you rotate Q and K by an angle proportional to the position. The dot product q·k then depends only on the difference in position — attention gets the relative position for free. The standard in modern LLMs (Llama, Qwen).
Code
import numpy as np
def rope(x, pos, base=10_000):
"""x: (d,) with d even. Rotates the pair (x[2i], x[2i+1]) by the angle pos * base^(-2i/d)."""
d = x.shape[0]
i = np.arange(d // 2)
theta = pos * base ** (-2 * i / d)
c, s = np.cos(theta), np.sin(theta)
x1, x2 = x[0::2], x[1::2]
out = np.empty_like(x)
out[0::2] = x1 * c - x2 * s
out[1::2] = x1 * s + x2 * c
return out
rng = np.random.default_rng(0); q, k = rng.normal(size=8), rng.normal(size=8)
a = rope(q, 3) @ rope(k, 1) # positions 3 and 1 (a distance of 2)
b = rope(q, 10) @ rope(k, 8) # positions 10 and 8 (a distance of 2)
print(np.isclose(a, b)) # True — only the relative position matters
print(np.isclose(a, rope(q, 5) @ rope(k, 1))) # False — a different distance
Sinusoidal encoding: PE[pos, 2i] = sin(pos / 10000^(2i/d)), PE[pos, 2i+1] = cos(...), added to the token embedding.
Formal
A rotation in plane by the angle : . RoPE sets , . Since rotation matrices are orthogonal and commute within the same plane: — a function of . The frequencies give slow rotations at high indices (a long range) and fast ones at low indices (local precision). Extending the context after training is done by scaling (positional interpolation, NTK-aware, YaRN).
Mastery means
- Explains why attention needs positional information
- Describes sinusoidal, learnt and rotary (RoPE) encoding
- Implements RoPE and shows that the relative position is preserved
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — RoFormer: Enhanced Transformer with Rotary Position Embedding — arXiv (open access; licence per article)
- arXiv — Attention Is All You Need — arXiv (open access; licence per article)