Why multiple layers are needed
Be able to explain why multiple layers can recognise more complex patterns.
Prerequisites
- CNeural networks — the intuitionrequired
- CTurn the knobs: weightsrequired
Everyday explanation
A single layer can draw one straight boundary. That goes a long way — but not always.
One layer handles this: But not this:
○ ○ │ ● ● ○ ● │ ● ○
○ ○ │ ● ● ● ○ │ ○ ●
↑ ↑
one straight line works no straight line works
The pattern on the right (known as XOR) requires two boundaries. A single layer cannot make two.
The solution is to stack layers. The first layer draws several straight boundaries. The second layer combines them: «inside boundary 1 AND outside boundary 2».
With multiple layers, the network can build any complex shape from straight segments — just as you can build a round shape from many small straight pieces.
Intuition
The hierarchy in a vision network is the clearest example:
| Layer | What it finds |
|---|---|
| 1 | edges and light transitions |
| 2 | corners and arcs (combinations of edges) |
| 3 | eyes, wheels, letters (combinations of corners) |
| 4 | faces, cars, words |
Each layer builds on the previous one. No one programmed that order — it emerges because each layer only sees the output from the previous one.
The crucial detail: the activation function.
Without it, stacked layers are meaningless. Two straight boundaries in a row just become... another straight boundary. Mathematically: a linear function of a linear function is still linear. Ten layers without an activation function can do exactly the same thing as one.
The activation function is a small bend between layers. The most common one, ReLU, is absurdly simple:
if the number is negative → 0
otherwise → the number itself
That single bend is all that is needed. With it, stacked layers can build any shape; without it, they can only draw a line.
Deep or wide? A single sufficiently wide layer can theoretically approximate any function. But in practice, it would need to be absurdly wide. Several narrower layers are almost always more efficient — just as it is easier to build something from reusable parts than to do it all in one piece.
Interactive
See it yourself in TensorFlow Playground (playground.tensorflow.org). Five experiments, ten minutes.
| # | Do | What happens |
|---|---|---|
| 1 | Select the circular dataset, 0 hidden layers | the error stays high — a straight line cannot enclose a circle |
| 2 | Add 1 hidden layer with 4 neurons | the error drops, the boundary becomes a polygon around the circle |
| 3 | Select the spiral dataset, keep 1 layer | it gets stuck again — the spiral is too complex |
| 4 | Add a second layer | now it handles it, but needs more epochs |
| 5 | Set activation = Linear on the spiral | the error gets stuck no matter how many layers you add |
Experiment 5 is the most important. It concretely shows that layers without an activation function add nothing — a hundred linear layers are still just a line.
Also look at the images inside the neurons while it trains. In the first layer, you see simple straight boundaries in different directions. In the second layer, you see how they are combined into corners and areas. The hierarchy is visible immediately.
Question to ponder: what happens if you add ten layers to the simple dataset? (Answer: it still works, but trains slower and becomes less stable. More is not always better.)
Mastery means
- Explains what an extra layer adds
- Knows why the activation function is needed
- Gives examples of patterns that require multiple layers
Sign in to do the exercises and build your mastery up.
Sources
- TensorFlow Playground (Apache-2.0) — Apache-2.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0