Skip to content
AI-grafen
DAI developerMathematics· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Sigmoid and softmax as functions

Be able to derive sigmoid and softmax from the exponential function and explain their properties.

Prerequisites

Intuition

A neural network spits out arbitrary numbers — «logits» — which can be −13.7 or 42. Probabilities have to lie between 0 and 1 and sum to 1. Something has to translate.

Sigmoid translates one number into one probability:

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

zzσ(z)\sigma(z)
−40.018
−10.269
00.5
10.731
40.982

S-shaped, symmetric about (0,0.5)(0, 0.5), flat at both ends.

Softmax translates several numbers into one distribution:

softmax(z)i=ezi∑jezj\mathrm{softmax}(z)_i = \frac{e^{z_i}}{\sum_j e^{z_j}}

All of them become positive (thanks to ex>0e^x > 0) and the sum becomes exactly 1 (thanks to the division by the sum).

A simple rule: a yes/no question → sigmoid. A question with several mutually exclusive alternatives → softmax. Several labels that can hold at the same time → sigmoid per label.

Derivation

Sigmoid from the odds. The odds of an event are p1−p\frac{p}{1-p}. The log odds (the logit) are z=ln⁡p1−pz = \ln\frac{p}{1-p}. Solve for pp:

ez=p1−p  ⇒  ez−pez=p  ⇒  p=ez1+ez=11+e−ze^z = \frac{p}{1-p} \;\Rightarrow\; e^z - p e^z = p \;\Rightarrow\; p = \frac{e^z}{1+e^z} = \frac{1}{1+e^{-z}}

Sigmoid is therefore the inverse of the logit function. It is not an arbitrary S-curve that somebody made up — it is the function that turns log odds into a probability, which is precisely what logistic regression needs.

The derivative is unusually elegant:

σ′(z)=σ(z) (1−σ(z))\sigma'(z) = \sigma(z)\,(1 - \sigma(z))

The derivation: σ(z)=(1+e−z)−1\sigma(z) = (1+e^{-z})^{-1}, so σ′(z)=−(1+e−z)−2⋅(−e−z)=e−z(1+e−z)2=σ(z)⋅e−z1+e−z=σ(z)(1−σ(z))\sigma'(z) = -(1+e^{-z})^{-2}\cdot(-e^{-z}) = \dfrac{e^{-z}}{(1+e^{-z})^2} = \sigma(z)\cdot\dfrac{e^{-z}}{1+e^{-z}} = \sigma(z)(1-\sigma(z)).

The derivative can therefore be computed from the function value — no need to keep zz around. That is one of the reasons sigmoid was popular in early networks.

But: the maximum is σ′(0)=0.25\sigma'(0) = 0.25, and at ∣z∣>5|z| > 5 the derivative is near zero. That is vanishing gradients in their purest form, and the reason ReLU took over in deep networks.

Softmax properties:

  1. Translation invariance. softmax(z+c)=softmax(z)\mathrm{softmax}(z + c) = \mathrm{softmax}(z) for every constant cc — the factor ece^c cancels. It is used for numerical stability (c=−max⁡zc = -\max z) and means that only the differences between the logits matter.
  2. Sigmoid is the two-class special case: softmax([z,0])0=ezez+1=σ(z)\mathrm{softmax}([z, 0])_0 = \frac{e^z}{e^z + 1} = \sigma(z).
  3. Temperature. softmax(z/T)\mathrm{softmax}(z/T): a low TT gives a sharp distribution, a high TT an even one. It is the same parameter that controls the randomness in text generation.
  4. The Jacobian is ∂pi∂zj=pi(δij−pj)\frac{\partial p_i}{\partial z_j} = p_i(\delta_{ij} - p_j) — a matrix, not a simple element-wise derivative, since every output depends on all the logits.

Code

import numpy as np

def sigmoid(z):
    return 1.0 / (1.0 + np.exp(-z))

def softmax(z):
    z = np.asarray(z, float)
    e = np.exp(z - z.max())          # translation invariance → numerical stability
    return e / e.sum()

for z in (-4, -1, 0, 1, 4):
    s = sigmoid(z)
    print(f"z={z:3d}  σ={s:.4f}  σ'={s * (1 - s):.4f}")
# z= -4  σ=0.0180  σ'=0.0177
# z=  0  σ=0.5000  σ'=0.2500   ← the maximal derivative
# z=  4  σ=0.9820  σ'=0.0177   ← almost flat: a vanishing gradient

print(softmax([2.0, 1.0, 0.1]))               # [0.6590 0.2424 0.0986]
print(softmax([2.0, 1.0, 0.1]) - softmax([102.0, 101.0, 100.1]))   # ~0 — invariant

# Sigmoid is softmax with two classes
z = 1.3
print(round(float(sigmoid(z)), 6), round(float(softmax([z, 0.0])[0]), 6))   # 0.785835 0.785835

# The temperature controls the sharpness
for T in (0.5, 1.0, 2.0, 5.0):
    print(f"T={T}: {np.round(softmax(np.array([2.0, 1.0, 0.1]) / T), 3)}")
# T=0.5: [0.864 0.117 0.019]   sharp
# T=1.0: [0.659 0.242 0.099]
# T=2.0: [0.502 0.304 0.194]
# T=5.0: [0.4   0.327 0.273]   even

# The derivative can be computed from the function value — handy in the backward pass
s = sigmoid(np.array([-2.0, 0.0, 2.0]))
print(np.round(s * (1 - s), 4))               # [0.105 0.25  0.105]

Mastery means

  • Derives sigmoid and computes its derivative
  • Explains softmax and its properties
  • Knows when to use which

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences