Skip to content
AI-grafen
GFrontier LabMultimodal models· about 120 min· fast-moving, sources checked often· verified 2026-09-20· EN

Video understanding

Be able to explain how the time dimension is handled in video models.

Prerequisites

Intuition

Video is images plus time — and time is expensive. One minute at 30 frames a second is 1 800 frames. Running every frame through an image model is both unsustainable and unnecessary: neighbouring frames are almost identical.

Four ways of handling the time dimension:

ApproachHowHandles
Sparse samplingtake 8–32 frames evenly spaced, encode each, pool them«what is the clip about»
3D convolutionconvolution kernels over (time, height, width)local motion, short actions
Spatio-temporal attentionthe tokens are (time, patch); attention over bothlong dependencies, the most expensive
Memory / streamingkeep a compressed state over timehour-long videos

The decisive question for every task: does the model have to see the order? «Which sport is being played?» can be answered from a single frame. «Did the person hit the ball before or after the referee blew the whistle?» cannot.

And that difference is exactly where most video benchmarks turn out to be weaker than they look.

Formal

An uncomfortable insight about evaluation. On many popular video benchmarks a model that sees a single random frame gets close to the top result. That means the tasks can be solved from appearance, not from the sequence of events — and that improved figures do not necessarily reflect better temporal understanding.

Two baselines you should always run, before drawing any conclusion about time:

  1. The single-frame baseline. The same model, one random frame. How close does it get?
  2. Shuffled order. The same frames in random order. Does the result fall?

If they are close to the original, your benchmark is not measuring time. That is among the most useful things you can know about a video model — and it takes an afternoon to find out.

Sampling strategies:

StrategySuits
Even sampling of N framesclassifying short clips
Key frames at scene changeslong videos with a clear structure
Dense sampling in windows of interestaction detection with time bounds
Frames + a transcript of the speechlectures, meetings, interviews — often the cheapest large win

The last row is underrated. For all content where somebody speaks, the transcript carries the bulk of the information, at a fraction of the cost. Combine ASR with sparse frame sampling before you reach for a heavy video model.

The cost order in practice: transcribe → sparse frame sampling → key frames → dense spatio-temporal attention. Move down the list only when the measurement shows that you need to.

Code

import numpy as np

def even_sampling(n_frames, n=16):
    return np.linspace(0, n_frames - 1, n).round().astype(int).tolist()

def key_frames(frames, threshold=0.3, max_n=32):
    """Pick the frames where the image changes noticeably — scene changes rather than even sampling."""
    chosen, previous = [0], grey(frames[0])
    for i in range(1, len(frames)):
        g = grey(frames[i])
        if float(np.abs(g - previous).mean()) > threshold:
            chosen.append(i); previous = g
        if len(chosen) >= max_n:
            break
    return chosen

# The two baselines that decide whether your evaluation actually measures time
def temporal_baselines(model, dataset, rng=np.random.default_rng(0)):
    full = np.mean([model(d["frames"]) == d["answer"] for d in dataset])
    one_frame = np.mean([model([d["frames"][rng.integers(len(d["frames"]))]]) == d["answer"]
                         for d in dataset])
    shuffled = np.mean([model(list(rng.permutation(d["frames"]))) == d["answer"] for d in dataset])
    return {"full": round(float(full), 3),
            "one_frame": round(float(one_frame), 3),
            "shuffled_order": round(float(shuffled), 3),
            "requires_time": bool(full - max(one_frame, shuffled) > 0.10)}

print(temporal_baselines(model, dataset))
# {'full': 0.71, 'one_frame': 0.68, 'shuffled_order': 0.70, 'requires_time': False}
# → the benchmark measures appearance, not the sequence of events

That output is the usual outcome the first time somebody runs the test. It does not mean the model is bad — it means the task is not measuring what people thought.

Mastery means

  • Describes how time is modelled in video models
  • Chooses a sampling strategy from the task
  • Evaluates with metrics that require temporal understanding

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences