Images and text together
Be able to explain how an AI can describe an image in words.
Prerequisites
Intuition
If words can be points on a map — can images be too? Yes. A model is trained on millions of images with captions and learns to put the image and its caption close together on the same map.
Then: a new image → a point on the map → which words are nearby? «A cat asleep in a basket.» The model describes the image by finding nearby text.
And the other way round: write some text → find images near it. That is how searching for images with words works.
Interactive
Try an image describer (many AI assistants can take images): upload a photo you took yourself and ask for a description.
Check:
- Did it count the objects correctly?
- Did it guess at something that is not visible (say «on a beach» when it is a lake)?
- Did it miss something small but important?
The model describes what is typical for images like that — not always what is actually there. Unusual images (a dog with three legs, a blue banana) are often described wrongly.
Mastery means
- Explains that images and text can be placed on the same map of meaning
- Describes how an image description can come out wrong
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Learning Transferable Visual Models From Natural Language Supervision (CLIP) — arXiv (open access; licence per article)
- Wikipedia — Multimodal learning (CC BY-SA 4.0) — CC BY-SA 4.0