Skip to content
AI-grafen
DAI developerDeep learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Activation functions

Be able to compare ReLU, GELU, sigmoid and tanh, and explain dead neurons and vanishing gradients.

Prerequisites

Intuition

Without an activation function a deep network is just a long chain of matrix multiplications — which is a single linear function. The activation is what makes the depth meaningful.

FunctionFormulaProperty
ReLUmax(0, x)fast, the standard in hidden layers; can «die»
GELUx·Φ(x)a smooth ReLU, the standard in transformers
sigmoid1/(1+e⁻ˣ)0–1, for a binary probability in the last layer
tanh(eˣ−e⁻ˣ)/(eˣ+e⁻ˣ)−1–1, centred on zero
softmaxeˣⁱ/Σeˣʲa distribution over classes, the last layer

Two classic problems: sigmoid and tanh have derivatives near zero when x is large or small → the gradient dies in deep networks (the vanishing gradient). ReLU has derivative 0 for negative x → a neuron that always gets a negative input stops updating entirely (a dead neuron).

Code

import numpy as np

relu       = lambda x: np.maximum(0, x)
leaky_relu = lambda x, a=0.01: np.where(x > 0, x, a * x)
sigmoid    = lambda x: 1 / (1 + np.exp(-x))
tanh       = np.tanh
gelu       = lambda x: 0.5 * x * (1 + np.tanh(np.sqrt(2/np.pi) * (x + 0.044715 * x**3)))

x = np.array([-3.0, -0.5, 0.0, 0.5, 3.0])
for name, f in [("relu", relu), ("gelu", gelu), ("sigmoid", sigmoid), ("tanh", tanh)]:
    print(f"{name:8s}", np.round(f(x), 3))
# relu     [0.    0.    0.    0.5   3.   ]
# gelu     [-0.004 -0.154  0.    0.346  2.996]
# sigmoid  [0.047  0.378  0.5   0.622  0.953]
# tanh     [-0.995 -0.462  0.    0.462  0.995]

# the derivative of sigmoid at x = ±5:
d = lambda x: sigmoid(x) * (1 - sigmoid(x))
print(np.round(d(np.array([-5.0, 0.0, 5.0])), 5))    # [0.00665 0.25 0.00665]

Hence the rule: ReLU/GELU in the hidden layers, sigmoid or softmax only at the output — and then usually built into the loss function (CrossEntropyLoss takes logits).

Mastery means

  • Compares ReLU, GELU, sigmoid and tanh
  • Explains vanishing gradients and dead neurons
  • Chooses the activation according to the layer's role

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences