Skip to content
AI-grafen
CBuilderDeep learning· about 30 min· fundamentals that rarely change· verified 2026-09-20· EN

Why multiple layers are needed

Be able to explain why multiple layers can recognise more complex patterns.

Prerequisites

Everyday explanation

A single layer can draw one straight boundary. That goes a long way — but not always.

One layer handles this:        But not this:

   ○ ○ │ ● ●                     ○ ● │ ● ○
   ○ ○ │ ● ●                     ● ○ │ ○ ●
       ↑                              ↑
   one straight line works         no straight line works

The pattern on the right (known as XOR) requires two boundaries. A single layer cannot make two.

The solution is to stack layers. The first layer draws several straight boundaries. The second layer combines them: «inside boundary 1 AND outside boundary 2».

With multiple layers, the network can build any complex shape from straight segments — just as you can build a round shape from many small straight pieces.

Intuition

The hierarchy in a vision network is the clearest example:

LayerWhat it finds
1edges and light transitions
2corners and arcs (combinations of edges)
3eyes, wheels, letters (combinations of corners)
4faces, cars, words

Each layer builds on the previous one. No one programmed that order — it emerges because each layer only sees the output from the previous one.

The crucial detail: the activation function.

Without it, stacked layers are meaningless. Two straight boundaries in a row just become... another straight boundary. Mathematically: a linear function of a linear function is still linear. Ten layers without an activation function can do exactly the same thing as one.

The activation function is a small bend between layers. The most common one, ReLU, is absurdly simple:

if the number is negative → 0
otherwise              → the number itself

That single bend is all that is needed. With it, stacked layers can build any shape; without it, they can only draw a line.

Deep or wide? A single sufficiently wide layer can theoretically approximate any function. But in practice, it would need to be absurdly wide. Several narrower layers are almost always more efficient — just as it is easier to build something from reusable parts than to do it all in one piece.

Interactive

See it yourself in TensorFlow Playground (playground.tensorflow.org). Five experiments, ten minutes.

#DoWhat happens
1Select the circular dataset, 0 hidden layersthe error stays high — a straight line cannot enclose a circle
2Add 1 hidden layer with 4 neuronsthe error drops, the boundary becomes a polygon around the circle
3Select the spiral dataset, keep 1 layerit gets stuck again — the spiral is too complex
4Add a second layernow it handles it, but needs more epochs
5Set activation = Linear on the spiralthe error gets stuck no matter how many layers you add

Experiment 5 is the most important. It concretely shows that layers without an activation function add nothing — a hundred linear layers are still just a line.

Also look at the images inside the neurons while it trains. In the first layer, you see simple straight boundaries in different directions. In the second layer, you see how they are combined into corners and areas. The hierarchy is visible immediately.

Question to ponder: what happens if you add ten layers to the simple dataset? (Answer: it still works, but trains slower and becomes less stable. More is not always better.)

Mastery means

  • Explains what an extra layer adds
  • Knows why the activation function is needed
  • Gives examples of patterns that require multiple layers

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences