Image description and visual questions
Be able to use a vision–language model and evaluate its answers.
Prerequisites
- EMultimodal models — the basicsrequired
- ELanguage models — training and generationrequired
Intuition
A vision–language model is a language model that has been given sight. The recipe is surprisingly simple:
- A pretrained image encoder (usually a ViT) splits the image into patches and gives one vector per patch.
- A small projection translates those vectors into the same space as the language model's token embeddings.
- The projected patches are put at the front of the token stream, as if they were words.
- The language model carries on as usual.
After that an image is just a handful of tokens, and everything the language model can do — follow instructions, reason step by step, answer in JSON — works on images.
What goes wrong is also predictable: the model describes what usually appears in images like that, rather than what is actually there. Ask «what colour is the car?» about an image with no car, and many models answer «blue».
Formal
Visual hallucinations come in three typical forms:
| Type | Example |
|---|---|
| Object | describes things that are not in the image |
| Attribute | the right object, the wrong colour, count or position |
| Relation | the right objects, the wrong relation between them |
The cause is that the language model's prior is strong. If it has seen hundreds of thousands of kitchen images with both a cooker and a fridge, it will mention the fridge even when it is not visible. The image signal competes with the statistics of the language, and often loses.
Countermeasures, in order of how much they help:
- Give the model a way out. «Answer "not visible in the image" if the information is missing» — the single most effective instruction.
- Ask narrowly. «Is there a car in the image? Yes or no.» beats «describe the image».
- Require localisation. Ask for the approximate position of every claimed object; hallucinated objects often get implausible coordinates.
- Resolution. Many errors are simply that the text or the detail is not visible — raise the resolution or crop before you blame the model.
- Two questions in opposite directions. «Is X there?» and «Is X missing?» should give consistent answers.
Evaluation. Avoid leaning on CIDEr and BLEU for image description — they measure word overlap with references and miss factual errors entirely. Better:
| Method | Measures |
|---|---|
| POPE-like yes/no questions | object hallucination, balanced |
| VQA accuracy on your own data | your actual task |
| Claim checking against the image | faithfulness per claim |
| A localisation requirement | whether the object was really seen |
The POPE trick is simple and good: for every image, ask both about objects that are there and about objects that are not but that usually occur alongside them. The share of false «yes» answers is a direct measure of the tendency to hallucinate.
Code
import json, torch
from transformers import AutoProcessor, AutoModelForVision2Seq
name = "HuggingFaceM4/idefics2-8b"
pro = AutoProcessor.from_pretrained(name)
m = AutoModelForVision2Seq.from_pretrained(name, torch_dtype=torch.bfloat16, device_map="auto")
STRICT = ("Answer ONLY from the image. If something cannot be seen, write \"not visible in the image\". "
"Answer as JSON: {\"objects\": [{\"name\": str, \"position\": \"top left|...\"}], "
"\"unsure_about\": [str]}")
def ask(image, text, max_tokens=300):
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": text}]}]
prompt = pro.apply_chat_template(messages, add_generation_prompt=True)
inp = pro(text=prompt, images=[image], return_tensors="pt").to(m.device)
out = m.generate(**inp, max_new_tokens=max_tokens, do_sample=False)
return pro.decode(out[0, inp["input_ids"].shape[1]:], skip_special_tokens=True)
# A POPE-like hallucination measurement
def pope(image, present: list[str], absent: list[str]):
def yes(obj):
answer = ask(image, f"Is there {obj} in the image? Answer only Yes or No.", 5)
return answer.strip().lower().startswith("yes")
tp = sum(yes(o) for o in present)
fp = sum(yes(o) for o in absent)
return {"recall": tp / max(len(present), 1),
"false_yes": fp / max(len(absent), 1)}
print(pope(image, present=["a chair", "a table"], absent=["a cat", "a lamp"]))
# {'recall': 1.0, 'false_yes': 0.5} ← half the non-existent objects were «seen»
Run pope over a couple of hundred images before you trust a VLM in production. false_yes is the number that decides whether the model will do, and it is never in the model card.
Mastery means
- Uses a VLM for description and visual questions
- Evaluates the answers against the image
- Recognises and handles visual hallucinations
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Visual Instruction Tuning (LLaVA) — arXiv (open access; licence per article)
- arXiv — Evaluating Object Hallucination in Large Vision-Language Models (POPE) — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0