Video understanding
Be able to explain how the time dimension is handled in video models.
Prerequisites
Intuition
Video is images plus time — and time is expensive. One minute at 30 frames a second is 1 800 frames. Running every frame through an image model is both unsustainable and unnecessary: neighbouring frames are almost identical.
Four ways of handling the time dimension:
| Approach | How | Handles |
|---|---|---|
| Sparse sampling | take 8–32 frames evenly spaced, encode each, pool them | «what is the clip about» |
| 3D convolution | convolution kernels over (time, height, width) | local motion, short actions |
| Spatio-temporal attention | the tokens are (time, patch); attention over both | long dependencies, the most expensive |
| Memory / streaming | keep a compressed state over time | hour-long videos |
The decisive question for every task: does the model have to see the order? «Which sport is being played?» can be answered from a single frame. «Did the person hit the ball before or after the referee blew the whistle?» cannot.
And that difference is exactly where most video benchmarks turn out to be weaker than they look.
Formal
An uncomfortable insight about evaluation. On many popular video benchmarks a model that sees a single random frame gets close to the top result. That means the tasks can be solved from appearance, not from the sequence of events — and that improved figures do not necessarily reflect better temporal understanding.
Two baselines you should always run, before drawing any conclusion about time:
- The single-frame baseline. The same model, one random frame. How close does it get?
- Shuffled order. The same frames in random order. Does the result fall?
If they are close to the original, your benchmark is not measuring time. That is among the most useful things you can know about a video model — and it takes an afternoon to find out.
Sampling strategies:
| Strategy | Suits |
|---|---|
| Even sampling of N frames | classifying short clips |
| Key frames at scene changes | long videos with a clear structure |
| Dense sampling in windows of interest | action detection with time bounds |
| Frames + a transcript of the speech | lectures, meetings, interviews — often the cheapest large win |
The last row is underrated. For all content where somebody speaks, the transcript carries the bulk of the information, at a fraction of the cost. Combine ASR with sparse frame sampling before you reach for a heavy video model.
The cost order in practice: transcribe → sparse frame sampling → key frames → dense spatio-temporal attention. Move down the list only when the measurement shows that you need to.
Code
import numpy as np
def even_sampling(n_frames, n=16):
return np.linspace(0, n_frames - 1, n).round().astype(int).tolist()
def key_frames(frames, threshold=0.3, max_n=32):
"""Pick the frames where the image changes noticeably — scene changes rather than even sampling."""
chosen, previous = [0], grey(frames[0])
for i in range(1, len(frames)):
g = grey(frames[i])
if float(np.abs(g - previous).mean()) > threshold:
chosen.append(i); previous = g
if len(chosen) >= max_n:
break
return chosen
# The two baselines that decide whether your evaluation actually measures time
def temporal_baselines(model, dataset, rng=np.random.default_rng(0)):
full = np.mean([model(d["frames"]) == d["answer"] for d in dataset])
one_frame = np.mean([model([d["frames"][rng.integers(len(d["frames"]))]]) == d["answer"]
for d in dataset])
shuffled = np.mean([model(list(rng.permutation(d["frames"]))) == d["answer"] for d in dataset])
return {"full": round(float(full), 3),
"one_frame": round(float(one_frame), 3),
"shuffled_order": round(float(shuffled), 3),
"requires_time": bool(full - max(one_frame, shuffled) > 0.10)}
print(temporal_baselines(model, dataset))
# {'full': 0.71, 'one_frame': 0.68, 'shuffled_order': 0.70, 'requires_time': False}
# → the benchmark measures appearance, not the sequence of events
That output is the usual outcome the first time somebody runs the test. It does not mean the model is bad — it means the task is not measuring what people thought.
Mastery means
- Describes how time is modelled in video models
- Chooses a sampling strategy from the task
- Evaluates with metrics that require temporal understanding
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — ViViT: A Video Vision Transformer — arXiv (open access; licence per article)
- arXiv — Revisiting the "Video" in Video-Language Understanding — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0