Skip to content
AI-grafen
CBuilderComputer science· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Compressing data

Be able to explain why repetition can be compressed and what gets lost.

Prerequisites

Intuition

The text AAAAAAAABBBB is 12 characters. But it can be written 8A4B — 4 characters. The same information, less room. That is called compression.

It works because the data has a pattern. Random data (A9x2QbZ) cannot be compressed — there is no pattern to exploit.

Lossless compression (ZIP, PNG): everything can be reconstructed exactly. Lossy compression (JPEG, MP3): information is thrown away — the kind the eye or ear barely notices. The image becomes much smaller but never exactly the same again.

Intuition

Try it yourself: compress AABBBBCCCCCCCCD with the «count + character» method.

→ 2A4B8C1D = 8 characters instead of 15. Nearly half.

But ABCABCABC becomes 1A1B1C1A1B1C1A1B1C = 18 characters — twice as long! The method only suits data with long runs. Real compression programs pick a method to suit the data.

The link to AI: a language model is in a sense an enormous compression of text — it has learnt the patterns in it, not the text itself. The better a model predicts the next word, the more efficiently it has «compressed» the language. That is why perplexity and compression belong together.

Mastery means

  • Explains why repetition can be compressed
  • Distinguishes lossless from lossy compression

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences