Skip to content
AI-grafen

The goal G Frontier Lab

Multimodal systems

Shared vector spaces, vision–language models, document understanding, multimodal retrieval, video, and agents that see the screen — with hallucination and safety measurements that hold.

Knowledge nodes
49
From zero
about 39 h
Labs
6
See what you already know — no account

The diagnostic removes what you already know, so your path is usually much shorter.

What you can do afterwards

Labs along the way

You write the code. Tests you cannot see decide whether it holds up.

The whole path

Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map

AExplorer3 knowledge nodes
  1. Patterns and categories
  2. Sequences and precise instructions
  3. Integer arithmetic and order of operations
BInvestigator7 knowledge nodes
  1. Data in everyday life
  2. Algorithmic thinking
  3. Rule-based systems and machine learning
  4. Source criticism and responsibility in AI use
  5. Binary numbers and bits
  6. Training data, features and labels
  7. Classification: how a model sorts information
CBuilder10 knowledge nodes
  1. Functions and coordinate systems
  2. Programming logic — variables, conditions, loops
  3. Python — the basics
  4. Python — lists, loops and dictionaries
  5. Statistics — mean, median and spread
  6. Probability — the basics
  7. Digital representation: text, images and audio
  8. Linear regression: fitting a straight line to data
  9. Neural networks — the intuition
  10. Prompting — steering a language model
DAI developer15 knowledge nodes
  1. Derivatives and optimisation
  2. Loss functions
  3. Tokenisation
  4. Gradient descent
  5. Vectors
  6. Matrices and matrix multiplication
  7. Neural networks — the forward pass with matrices
  8. Backpropagation
  9. Embeddings — words as vectors
  10. Attention
  11. NumPy — arrays and vectorisation
  12. PyTorch — tensors and autograd
  13. Retrieval — finding the right text
  14. Train a neural network in PyTorch
  15. Transformers — the architecture
EUniversity8 knowledge nodes
  1. Image preprocessing and augmentation
  2. Sequence models before transformers
  3. Language models — training and generation
  4. RAG — retrieval-augmented generation
  5. Vision Transformer (ViT)
  6. Multimodal models — the basics
  7. Tool use
  8. Agents — plan, act, observe
FAI engineering4 knowledge nodes
  1. CLIP and contrastive image–text learning
  2. Image description and visual questions
  3. Document understanding: OCR, layout, tables
  4. Multimodal RAG
GFrontier Lab2 knowledge nodes
  1. Multimodal agents and computer control
  2. Video understanding