Skip to content
AI-grafen
EUniversityTransformer architecture· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Positional encoding and RoPE

Be able to explain sinusoidal, learnt and rotary positional encoding, and implement RoPE.

Prerequisites

Intuition

Attention is permutation-invariant: shuffle the order of the tokens and you get the same set of outputs in a new order. «The dog bit the man» and «the man bit the dog» would look the same. The position therefore has to be added.

  • Sinusoidal encoding (the original): a fixed vector pattern per position of sines and cosines at different frequencies. No learning, and it can in principle extrapolate.
  • Learnt: one embedding per position, as for tokens. Simple, but it cannot go beyond the trained length.
  • RoPE (rotary): instead of adding the position to x, you rotate Q and K by an angle proportional to the position. The dot product q·k then depends only on the difference in position — attention gets the relative position for free. The standard in modern LLMs (Llama, Qwen).

Code

import numpy as np

def rope(x, pos, base=10_000):
    """x: (d,) with d even. Rotates the pair (x[2i], x[2i+1]) by the angle pos * base^(-2i/d)."""
    d = x.shape[0]
    i = np.arange(d // 2)
    theta = pos * base ** (-2 * i / d)
    c, s = np.cos(theta), np.sin(theta)
    x1, x2 = x[0::2], x[1::2]
    out = np.empty_like(x)
    out[0::2] = x1 * c - x2 * s
    out[1::2] = x1 * s + x2 * c
    return out

rng = np.random.default_rng(0); q, k = rng.normal(size=8), rng.normal(size=8)
a = rope(q, 3) @ rope(k, 1)        # positions 3 and 1  (a distance of 2)
b = rope(q, 10) @ rope(k, 8)       # positions 10 and 8 (a distance of 2)
print(np.isclose(a, b))            # True — only the relative position matters
print(np.isclose(a, rope(q, 5) @ rope(k, 1)))   # False — a different distance

Sinusoidal encoding: PE[pos, 2i] = sin(pos / 10000^(2i/d)), PE[pos, 2i+1] = cos(...), added to the token embedding.

Formal

A rotation in plane ii by the angle mθim\theta_i: R(mθi)R(m\theta_i). RoPE sets qm=RΘ,mWQxmq_m = R_{\Theta,m} W^Q x_m, kn=RΘ,nWKxnk_n = R_{\Theta,n} W^K x_n. Since rotation matrices are orthogonal and commute within the same plane: qm⊤kn=(WQxm)⊤RΘ,m⊤RΘ,n(WKxn)=(WQxm)⊤RΘ, n−m(WKxn)q_m^\top k_n = (W^Q x_m)^\top R_{\Theta,m}^\top R_{\Theta,n} (W^K x_n) = (W^Q x_m)^\top R_{\Theta,\,n-m}(W^K x_n) — a function of n−mn-m. The frequencies θi=b−2i/d\theta_i = b^{-2i/d} give slow rotations at high indices (a long range) and fast ones at low indices (local precision). Extending the context after training is done by scaling θ\theta (positional interpolation, NTK-aware, YaRN).

Mastery means

  • Explains why attention needs positional information
  • Describes sinusoidal, learnt and rotary (RoPE) encoding
  • Implements RoPE and shows that the relative position is preserved

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences