Skip to content
AI-grafen

What is a transformer? Explained from attention to language model

A transformer is an architecture where every token gets to look at all other tokens and weigh them by relevance — that is attention. Stack blocks of attention and small neural networks and you get the GPT, Llama and Qwen families.

To really understand it you need vectors, matrices, a neural network's forward pass and embeddings. AI-grafen checks which of these you already know and builds the path accordingly.

In the lab you implement scaled dot-product attention in NumPy, train a small GPT on a text and plot its attention maps.

Already a programmer? Go straight to AI for developers, RAG, agents or evals. How we know what you know.