Skip to content
AI-grafen
DAI developerLanguage models· about 40 min· fundamentals that rarely change· verified 2026-09-20· EN

Embeddings — words as vectors

Be able to explain what an embedding is, why similar words end up close together, calculate the similarity between embeddings with cosine, and know that embeddings are learnt as weights.

Prerequisites

Intuition

A computer cannot do arithmetic with the word "cat". An embedding gives every word (or token) a vector of, say, 300 numbers. The trick: the vectors are learnt so that words used in similar contexts get vectors pointing in similar directions.

Then similarity becomes something you can compute: cos(cat, dog) ≈ 0.8, cos(cat, tractor) ≈ 0.1. And relationships become directions: king − man + woman ≈ queen.

Formal

The embedding layer is a matrix E of size V×d (V words in the vocabulary, d dimensions). Word number i has the vector E[i] — the layer is a lookup table, mathematically the same as multiplying a one-hot vector by E.

E consists of weights trained with backprop just like everything else. In a language model they are shaped by the task "guess the next word"; in an embedding API (qwen3-embed, say) by the task "similar texts should get similar vectors".

Similarity: cos(u, v) = u·v / (‖u‖‖v‖). Search with embeddings = compute cos against all the documents and take the largest. That is the foundation of retrieval and RAG.

Code

import torch, torch.nn as nn
emb = nn.Embedding(num_embeddings=10_000, embedding_dim=64)
ids = torch.tensor([42, 7, 42])
v = emb(ids)              # (3, 64) — row 42 twice
cos = nn.functional.cosine_similarity(v[0], v[1], dim=0)

Mastery means

  • Calculates the cosine similarity between two embeddings
  • Explains how an embedding matrix works as a lookup table

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences