Skip to content
AI-grafen
CBuilderComputer science· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Digital representation: text, images and audio

Be able to explain how characters, pixels and audio samples are converted into numerical values, and thereby how data is fed into a model.

Prerequisites

Intuition

Everything fed into a model is numerical values. The only question is how the conversion happens.

  • Text: every character has a unique number (Unicode). A = 65, å = 229, 😀 = 128512. A text string is essentially a list of numbers. (Language models go one step further and group characters into tokens.)
  • Image: a grid of pixels; each pixel is represented by three numbers 0–255 (red, green, blue). An image of 1000 × 1000 pixels contains 3 million numbers.
  • Audio: a microphone measures air pressure 44 100 times per second; each measurement is a number. One second of audio corresponds to 44 100 numbers.

The model does not interpret what the numbers mean in themselves — it identifies patterns in the numerical sequences.

Code

text = "Hej!"
print([ord(c) for c in text])        # [72, 101, 106, 33]
print(chr(229))                        # å

# a 2×2 image in RGB: rows → pixels → (r, g, b)
bild = [[(255, 0, 0), (0, 255, 0)],
        [(0, 0, 255), (0, 0, 0)]]
print(bild[0][1])                      # (0, 255, 0) — green

# one second of silent audio at 8 kHz
ljud = [0] * 8000

Size: 1000 × 1000 pixels × 3 bytes = 3 MB uncompressed. Compression (JPEG, MP3) reduces file size by removing information that humans do not perceive.

Mastery means

  • Explains how characters, pixels and audio samples become numbers
  • Estimates the size of a file based on resolution and bit depth

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences