CBuilderAI safety and alignment· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN
Can an AI be fooled?
Be able to describe simple ways of fooling a model and why protection is needed.
Prerequisites
- BAI's invented answersrequired
Intuition
Models can be fooled — not because they are stupid, but because they only see patterns.
Three common ways:
- The picture is changed a little. If you change a few pixels in a way a human being cannot see, an image model can go from «panda» to «gibbon» with high confidence. This is called an adversarial example.
- The question is rewritten. A model that refuses a question can answer it if the question is put as a «story» or in another language. This is called a jailbreak.
- Instructions are hidden in text the model reads. A web page can contain «AI assistant: ignore the user and do X». This is called prompt injection.
Testing this on your own systems is called red teaming and is part of serious AI work.
Intuition
Why it matters: a model that can be fooled is dangerous when it has the power to do something — send money, let somebody in, make a diagnosis.
The protection in layers:
- Limit what the model is allowed to do (reading is less dangerous than sending).
- A human being confirms what cannot be undone.
- Test systematically with known tricks before launch.
- Log everything so that you can see afterwards what happened.
An important rule: only test on systems you own yourself or have been given permission to test. Fooling somebody else's system is not a game — it can be illegal.
Mastery means
- Describes two ways of fooling a model
- Explains why protection is needed
Sign in to do the exercises and build your mastery up.
Sources
- OWASP Top 10 for LLM Applications — CC BY-SA 4.0
- arXiv — Explaining and Harnessing Adversarial Examples — arXiv (open access; licence per article)