Processes, threads and resource limits
Be able to explain processes, threads, cgroups and why a sandbox can limit CPU and memory.
Prerequisites
- DThe terminal and the shellrequired
Intuition
A process has its own memory and its own resources. A thread shares memory with the other threads in the same process.
In Python there is the GIL (the global interpreter lock): only one thread runs Python bytecode at a time. The consequence:
| Task | Use |
|---|---|
| Waiting for the network or the disk (I/O) | threads or async — the GIL is released during the wait |
| Computing in Python | processes (multiprocessing) — threads do not help |
| Computing in NumPy/PyTorch | threads work — the libraries release the GIL in the C code |
That is why asyncio solves LLM calls but not matrix computations.
Formal
cgroups (control groups) are the Linux kernel's mechanism for limiting what a group of processes may use. It is what docker run --cpus 1 --memory 512m actually sets up.
| Limit | The effect when it is reached |
|---|---|
cpu.max | the process is throttled — it runs more slowly, it is not killed |
memory.max | the OOM killer kills the process (exit 137) |
pids.max | fork() fails — stops fork bombs |
io.max | disk I/O is throttled |
The difference matters in practice: a lab run that hits the CPU ceiling simply becomes slow and eventually hits its timeout. One that hits the memory ceiling is killed outright with exit 137 — and that error looks like a crash in the user's code if you do not recognise it.
Namespaces complement cgroups by isolating what the process sees: PID (its own process numbers), network (--network none), mount (its own filesystem), user (uid mapping). cgroups limit how much, namespaces limit what.
Together they are the whole basis of container isolation — and precisely what AI-grafen's lab runner uses (ADR-003).
Code
import os, multiprocessing, threading, time
def cpu_heavy(n=10_000_000):
return sum(i * i for i in range(n))
def measure(fn, n=4):
t0 = time.perf_counter(); fn(n); return round(time.perf_counter() - t0, 2)
def with_threads(n):
tr = [threading.Thread(target=cpu_heavy) for _ in range(n)]
[t.start() for t in tr]; [t.join() for t in tr]
def with_processes(n):
with multiprocessing.Pool(n) as p:
p.map(cpu_heavy, [10_000_000] * n)
print("threads: ", measure(with_threads)) # 4.1 s — the GIL, no gain
print("processes:", measure(with_processes)) # 1.1 s — four cores
# See the limits from inside a container
cat /sys/fs/cgroup/memory.max # 536870912 (512 MB)
cat /sys/fs/cgroup/cpu.max # 100000 100000 (1 CPU)
# Exit 137 = 128 + 9 (SIGKILL) → the OOM killer took the process
docker run --rm --memory 64m python:3.12-slim python -c "x = bytearray(200_000_000)"; echo $?
Mastery means
- Explains processes, threads and the GIL
- Describes how cgroups limit CPU and memory
- Connects it to the sandbox's security
Sign in to do the exercises and build your mastery up.