Activation functions
Be able to compare ReLU, GELU, sigmoid and tanh, and explain dead neurons and vanishing gradients.
Prerequisites
Intuition
Without an activation function a deep network is just a long chain of matrix multiplications — which is a single linear function. The activation is what makes the depth meaningful.
| Function | Formula | Property |
|---|---|---|
| ReLU | max(0, x) | fast, the standard in hidden layers; can «die» |
| GELU | x·Φ(x) | a smooth ReLU, the standard in transformers |
| sigmoid | 1/(1+e⁻ˣ) | 0–1, for a binary probability in the last layer |
| tanh | (eˣ−e⁻ˣ)/(eˣ+e⁻ˣ) | −1–1, centred on zero |
| softmax | eˣⁱ/Σeˣʲ | a distribution over classes, the last layer |
Two classic problems: sigmoid and tanh have derivatives near zero when x is large or small → the gradient dies in deep networks (the vanishing gradient). ReLU has derivative 0 for negative x → a neuron that always gets a negative input stops updating entirely (a dead neuron).
Code
import numpy as np
relu = lambda x: np.maximum(0, x)
leaky_relu = lambda x, a=0.01: np.where(x > 0, x, a * x)
sigmoid = lambda x: 1 / (1 + np.exp(-x))
tanh = np.tanh
gelu = lambda x: 0.5 * x * (1 + np.tanh(np.sqrt(2/np.pi) * (x + 0.044715 * x**3)))
x = np.array([-3.0, -0.5, 0.0, 0.5, 3.0])
for name, f in [("relu", relu), ("gelu", gelu), ("sigmoid", sigmoid), ("tanh", tanh)]:
print(f"{name:8s}", np.round(f(x), 3))
# relu [0. 0. 0. 0.5 3. ]
# gelu [-0.004 -0.154 0. 0.346 2.996]
# sigmoid [0.047 0.378 0.5 0.622 0.953]
# tanh [-0.995 -0.462 0. 0.462 0.995]
# the derivative of sigmoid at x = ±5:
d = lambda x: sigmoid(x) * (1 - sigmoid(x))
print(np.round(d(np.array([-5.0, 0.0, 5.0])), 5)) # [0.00665 0.25 0.00665]
Hence the rule: ReLU/GELU in the hidden layers, sigmoid or softmax only at the output — and then usually built into the loss function (CrossEntropyLoss takes logits).
Mastery means
- Compares ReLU, GELU, sigmoid and tanh
- Explains vanishing gradients and dead neurons
- Chooses the activation according to the layer's role
Sign in to do the exercises and build your mastery up.
Sources
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — Gaussian Error Linear Units (GELUs) — arXiv (open access; licence per article)