Skip to content
AI-grafen
DAI developerDeep learning· about 45 min· evolving, reviewed regularly· verified 2026-09-20· EN

PyTorch — tensors and autograd

Be able to create and compute with tensors, let autograd compute the gradients, and understand the connection between what you did by hand in backprop and what `.backward()` does.

Prerequisites

Code

import torch

x = torch.tensor([1.0, 2.0, 3.0])
W = torch.tensor([[1.0, 0.0, 2.0],
                  [0.5, 1.0, 0.0]], requires_grad=True)
y = W @ x                # tensor([7.0, 2.5])
L = (y ** 2).sum()       # a scalar loss
L.backward()             # autograd: the gradient of L with respect to W
print(W.grad)

W.grad contains ∂L/∂W = 2·y·xᵀ — the same thing you would have worked out by hand with the chain rule. Autograd records every operation and runs the chain rule backwards.

Intuition

A tensor is an n-dimensional array (0-dim = a number, 1-dim = a vector, 2-dim = a matrix) that can live on the GPU and that remembers how it was computed.

Three things to keep track of:

  • requires_grad=True says "track this one, I want the gradient".
  • .backward() fills in .grad on every tracked tensor.
  • Gradients accumulate — zero them (W.grad.zero_() or optimizer.zero_grad()) before every new step, otherwise the old gradients are summed in.

torch.no_grad() switches the tracking off when you just want to compute (during evaluation, say).

Formal

Common shapes: x.shape, x.view(2, 3), x.T, x.unsqueeze(0). The batch dimension is usually first: a batch of images is (N, C, H, W). Device: x.to("cuda") moves it to the GPU; every tensor in an operation has to be on the same device.

Mastery means

  • Computes a gradient with autograd and compares it with the hand calculation
  • Explains what requires_grad and .backward() do

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences