Compressing data
Be able to explain why repetition can be compressed and what gets lost.
Prerequisites
Intuition
The text AAAAAAAABBBB is 12 characters. But it can be written 8A4B — 4 characters. The same information, less room. That is called compression.
It works because the data has a pattern. Random data (A9x2QbZ) cannot be compressed — there is no pattern to exploit.
Lossless compression (ZIP, PNG): everything can be reconstructed exactly. Lossy compression (JPEG, MP3): information is thrown away — the kind the eye or ear barely notices. The image becomes much smaller but never exactly the same again.
Intuition
Try it yourself: compress AABBBBCCCCCCCCD with the «count + character» method.
→ 2A4B8C1D = 8 characters instead of 15. Nearly half.
But ABCABCABC becomes 1A1B1C1A1B1C1A1B1C = 18 characters — twice as long! The method only suits data with long runs. Real compression programs pick a method to suit the data.
The link to AI: a language model is in a sense an enormous compression of text — it has learnt the patterns in it, not the text itself. The better a model predicts the next word, the more efficiently it has «compressed» the language. That is why perplexity and compression belong together.
Mastery means
- Explains why repetition can be compressed
- Distinguishes lossless from lossy compression
Sign in to do the exercises and build your mastery up.
Sources
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Datakomprimering (CC BY-SA 4.0) — CC BY-SA 4.0