Digital representation: text, images and audio
Be able to explain how characters, pixels and audio samples are converted into numerical values, and thereby how data is fed into a model.
Prerequisites
- BBinary numbers and bitsrequired
Intuition
Everything fed into a model is numerical values. The only question is how the conversion happens.
- Text: every character has a unique number (Unicode).
A= 65,å= 229,😀= 128512. A text string is essentially a list of numbers. (Language models go one step further and group characters into tokens.) - Image: a grid of pixels; each pixel is represented by three numbers 0–255 (red, green, blue). An image of 1000 × 1000 pixels contains 3 million numbers.
- Audio: a microphone measures air pressure 44 100 times per second; each measurement is a number. One second of audio corresponds to 44 100 numbers.
The model does not interpret what the numbers mean in themselves — it identifies patterns in the numerical sequences.
Code
text = "Hej!"
print([ord(c) for c in text]) # [72, 101, 106, 33]
print(chr(229)) # å
# a 2×2 image in RGB: rows → pixels → (r, g, b)
bild = [[(255, 0, 0), (0, 255, 0)],
[(0, 0, 255), (0, 0, 0)]]
print(bild[0][1]) # (0, 255, 0) — green
# one second of silent audio at 8 kHz
ljud = [0] * 8000
Size: 1000 × 1000 pixels × 3 bytes = 3 MB uncompressed. Compression (JPEG, MP3) reduces file size by removing information that humans do not perceive.
Mastery means
- Explains how characters, pixels and audio samples become numbers
- Estimates the size of a file based on resolution and bit depth
Sign in to do the exercises and build your mastery up.
Sources
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Unicode (CC BY-SA 4.0) — CC BY-SA 4.0