Sigmoid and softmax as functions
Be able to derive sigmoid and softmax from the exponential function and explain their properties.
Prerequisites
Intuition
A neural network spits out arbitrary numbers — «logits» — which can be −13.7 or 42. Probabilities have to lie between 0 and 1 and sum to 1. Something has to translate.
Sigmoid translates one number into one probability:
| −4 | 0.018 |
| −1 | 0.269 |
| 0 | 0.5 |
| 1 | 0.731 |
| 4 | 0.982 |
S-shaped, symmetric about , flat at both ends.
Softmax translates several numbers into one distribution:
All of them become positive (thanks to ) and the sum becomes exactly 1 (thanks to the division by the sum).
A simple rule: a yes/no question → sigmoid. A question with several mutually exclusive alternatives → softmax. Several labels that can hold at the same time → sigmoid per label.
Derivation
Sigmoid from the odds. The odds of an event are . The log odds (the logit) are . Solve for :
Sigmoid is therefore the inverse of the logit function. It is not an arbitrary S-curve that somebody made up — it is the function that turns log odds into a probability, which is precisely what logistic regression needs.
The derivative is unusually elegant:
The derivation: , so .
The derivative can therefore be computed from the function value — no need to keep around. That is one of the reasons sigmoid was popular in early networks.
But: the maximum is , and at the derivative is near zero. That is vanishing gradients in their purest form, and the reason ReLU took over in deep networks.
Softmax properties:
- Translation invariance. for every constant — the factor cancels. It is used for numerical stability () and means that only the differences between the logits matter.
- Sigmoid is the two-class special case: .
- Temperature. : a low gives a sharp distribution, a high an even one. It is the same parameter that controls the randomness in text generation.
- The Jacobian is — a matrix, not a simple element-wise derivative, since every output depends on all the logits.
Code
import numpy as np
def sigmoid(z):
return 1.0 / (1.0 + np.exp(-z))
def softmax(z):
z = np.asarray(z, float)
e = np.exp(z - z.max()) # translation invariance → numerical stability
return e / e.sum()
for z in (-4, -1, 0, 1, 4):
s = sigmoid(z)
print(f"z={z:3d} σ={s:.4f} σ'={s * (1 - s):.4f}")
# z= -4 σ=0.0180 σ'=0.0177
# z= 0 σ=0.5000 σ'=0.2500 ← the maximal derivative
# z= 4 σ=0.9820 σ'=0.0177 ← almost flat: a vanishing gradient
print(softmax([2.0, 1.0, 0.1])) # [0.6590 0.2424 0.0986]
print(softmax([2.0, 1.0, 0.1]) - softmax([102.0, 101.0, 100.1])) # ~0 — invariant
# Sigmoid is softmax with two classes
z = 1.3
print(round(float(sigmoid(z)), 6), round(float(softmax([z, 0.0])[0]), 6)) # 0.785835 0.785835
# The temperature controls the sharpness
for T in (0.5, 1.0, 2.0, 5.0):
print(f"T={T}: {np.round(softmax(np.array([2.0, 1.0, 0.1]) / T), 3)}")
# T=0.5: [0.864 0.117 0.019] sharp
# T=1.0: [0.659 0.242 0.099]
# T=2.0: [0.502 0.304 0.194]
# T=5.0: [0.4 0.327 0.273] even
# The derivative can be computed from the function value — handy in the backward pass
s = sigmoid(np.array([-2.0, 0.0, 2.0]))
print(np.round(s * (1 - s), 4)) # [0.105 0.25 0.105]
Mastery means
- Derives sigmoid and computes its derivative
- Explains softmax and its properties
- Knows when to use which
Sign in to do the exercises and build your mastery up.
Sources
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- Matteboken (Mattecentrum) — free to read, non-profit association
- Khan Academy — matematik — CC BY-NC-SA 3.0