Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

How autograd works inside

Be able to explain the computation graph and backward functions, and write an autograd function of your own.

Prerequisites

Intuition

Every operation on a tensor with requires_grad=True is recorded in a computation graph: the nodes are tensors, the edges operations. The graph is built dynamically during the forward pass.

loss.backward() walks the graph backwards and applies the chain rule: every operation knows how to turn the gradient coming in into gradients for its inputs. The result ends up in .grad on the leaf nodes (the parameters).

Three things that often confuse people:

  • Gradients are accumulated — hence opt.zero_grad() every step.
  • The graph is freed after the backward pass (unless retain_graph=True).
  • torch.no_grad() builds no graph at all — use it at inference, it saves memory and time.

Code

import torch

x = torch.tensor([2.0], requires_grad=True)
y = x ** 3 + 2 * x            # y = x³ + 2x  →  dy/dx = 3x² + 2 = 14
y.backward()
print(x.grad)                 # tensor([14.])

x.grad.zero_()                # otherwise the next gradient accumulates on top

# Your own autograd function: the forward and the backward are defined explicitly
class Square(torch.autograd.Function):
    @staticmethod
    def forward(ctx, x):
        ctx.save_for_backward(x)
        return x ** 2
    @staticmethod
    def backward(ctx, grad_out):
        (x,) = ctx.saved_tensors
        return grad_out * 2 * x         # d(x²)/dx = 2x, the chain rule: multiply it in

z = torch.tensor([3.0], requires_grad=True)
Square.apply(z).backward()
print(z.grad)                 # tensor([6.])

# A gradient check — the standard test for your own backward implementations
print(torch.autograd.gradcheck(Square.apply, (torch.randn(4, dtype=torch.double, requires_grad=True),)))

gradcheck compares your analytical gradient with a numerical approximation. If it passes, the implementation is almost certainly right — it is the same trick used with hand-written backprop.

Mastery means

  • Explains the computation graph and the backward call
  • Writes an autograd function of their own
  • Knows when the graph should be switched off

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences