Skip to content
AI-grafen
CBuilderMultimodal models· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Images and text together

Be able to explain how an AI can describe an image in words.

Prerequisites

Intuition

If words can be points on a map — can images be too? Yes. A model is trained on millions of images with captions and learns to put the image and its caption close together on the same map.

Then: a new image → a point on the map → which words are nearby? «A cat asleep in a basket.» The model describes the image by finding nearby text.

And the other way round: write some text → find images near it. That is how searching for images with words works.

Interactive

Try an image describer (many AI assistants can take images): upload a photo you took yourself and ask for a description.

Check:

  • Did it count the objects correctly?
  • Did it guess at something that is not visible (say «on a beach» when it is a lake)?
  • Did it miss something small but important?

The model describes what is typical for images like that — not always what is actually there. Unusual images (a dog with three legs, a blue banana) are often described wrongly.

Mastery means

  • Explains that images and text can be placed on the same map of meaning
  • Describes how an image description can come out wrong

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences