Skip to content
AI-grafen
FAI engineeringMultimodal models· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Image description and visual questions

Be able to use a vision–language model and evaluate its answers.

Prerequisites

Intuition

A vision–language model is a language model that has been given sight. The recipe is surprisingly simple:

  1. A pretrained image encoder (usually a ViT) splits the image into patches and gives one vector per patch.
  2. A small projection translates those vectors into the same space as the language model's token embeddings.
  3. The projected patches are put at the front of the token stream, as if they were words.
  4. The language model carries on as usual.

After that an image is just a handful of tokens, and everything the language model can do — follow instructions, reason step by step, answer in JSON — works on images.

What goes wrong is also predictable: the model describes what usually appears in images like that, rather than what is actually there. Ask «what colour is the car?» about an image with no car, and many models answer «blue».

Formal

Visual hallucinations come in three typical forms:

TypeExample
Objectdescribes things that are not in the image
Attributethe right object, the wrong colour, count or position
Relationthe right objects, the wrong relation between them

The cause is that the language model's prior is strong. If it has seen hundreds of thousands of kitchen images with both a cooker and a fridge, it will mention the fridge even when it is not visible. The image signal competes with the statistics of the language, and often loses.

Countermeasures, in order of how much they help:

  1. Give the model a way out. «Answer "not visible in the image" if the information is missing» — the single most effective instruction.
  2. Ask narrowly. «Is there a car in the image? Yes or no.» beats «describe the image».
  3. Require localisation. Ask for the approximate position of every claimed object; hallucinated objects often get implausible coordinates.
  4. Resolution. Many errors are simply that the text or the detail is not visible — raise the resolution or crop before you blame the model.
  5. Two questions in opposite directions. «Is X there?» and «Is X missing?» should give consistent answers.

Evaluation. Avoid leaning on CIDEr and BLEU for image description — they measure word overlap with references and miss factual errors entirely. Better:

MethodMeasures
POPE-like yes/no questionsobject hallucination, balanced
VQA accuracy on your own datayour actual task
Claim checking against the imagefaithfulness per claim
A localisation requirementwhether the object was really seen

The POPE trick is simple and good: for every image, ask both about objects that are there and about objects that are not but that usually occur alongside them. The share of false «yes» answers is a direct measure of the tendency to hallucinate.

Code

import json, torch
from transformers import AutoProcessor, AutoModelForVision2Seq

name = "HuggingFaceM4/idefics2-8b"
pro = AutoProcessor.from_pretrained(name)
m = AutoModelForVision2Seq.from_pretrained(name, torch_dtype=torch.bfloat16, device_map="auto")

STRICT = ("Answer ONLY from the image. If something cannot be seen, write \"not visible in the image\". "
          "Answer as JSON: {\"objects\": [{\"name\": str, \"position\": \"top left|...\"}], "
          "\"unsure_about\": [str]}")

def ask(image, text, max_tokens=300):
    messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": text}]}]
    prompt = pro.apply_chat_template(messages, add_generation_prompt=True)
    inp = pro(text=prompt, images=[image], return_tensors="pt").to(m.device)
    out = m.generate(**inp, max_new_tokens=max_tokens, do_sample=False)
    return pro.decode(out[0, inp["input_ids"].shape[1]:], skip_special_tokens=True)

# A POPE-like hallucination measurement
def pope(image, present: list[str], absent: list[str]):
    def yes(obj):
        answer = ask(image, f"Is there {obj} in the image? Answer only Yes or No.", 5)
        return answer.strip().lower().startswith("yes")
    tp = sum(yes(o) for o in present)
    fp = sum(yes(o) for o in absent)
    return {"recall": tp / max(len(present), 1),
            "false_yes": fp / max(len(absent), 1)}

print(pope(image, present=["a chair", "a table"], absent=["a cat", "a lamp"]))
# {'recall': 1.0, 'false_yes': 0.5}   ← half the non-existent objects were «seen»

Run pope over a couple of hundred images before you trust a VLM in production. false_yes is the number that decides whether the model will do, and it is never in the model card.

Mastery means

  • Uses a VLM for description and visual questions
  • Evaluates the answers against the image
  • Recognises and handles visual hallucinations

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences