Skip to content
AI-grafen
CBuilderMultimodal models· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Making images from text — and what can go wrong

Be able to try text-guided image generation and discuss mistakes, stereotypes and copyright.

Prerequisites

Intuition

Write «a fox reading a book in the forest, watercolour» — and a model creates the image. It does not copy an existing picture; it builds from noise, step by step, towards something that sits close to your text on the map of meaning.

Three things to watch out for:

  • Mistakes: hands, text inside the image, counts — the model knows how things usually look, not the rules.
  • Stereotypes: «a doctor» often gives a certain kind of person, because the training data looked like that.
  • Copyright: the model was trained on millions of images — often without the creators being asked. Imitating a named artist's style is contested.

Interactive

An experiment (with a free image generator):

  1. «a doctor» — describe the person you got. Try five times. A pattern?
  2. «a doctor, an older woman, a hospital in Nairobi» — what changed?
  3. «a hand holding six apples» — count the apples and the fingers.
  4. «a sign with the text STOP» — did the letters come out right?

Write it down: what does the prompt steer well, what often goes wrong, and what do you think about the model having learnt from other people's pictures?

Mastery means

  • Tries text-guided image generation and describes how the prompt steers the result
  • Discusses mistakes, stereotypes and copyright with examples

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences