Energy and environmental impact
Be able to estimate the energy use of training and inference and compare the alternatives.
Prerequisites
- ECost modelling for LLM systemsrequired
Intuition
AI consumes electricity. The question is how much, compared with what, and which choices actually make a difference.
Two items:
| Training | Inference | |
|---|---|---|
| When | once | every call |
| The order of magnitude | MWh to GWh | watt-seconds |
| Dominates the total? | at the start | after enough calls |
A large model used by millions of people consumes more on inference over time than on training. The training is a one-off cost; the inference carries on.
Where the electricity comes from matters enormously. The same computation in Sweden and on a coal-dependent grid differs by roughly a power of ten in carbon dioxide emissions. That is a larger factor than nearly all the technical optimisations put together.
Formal
The basic formula for training:
| The quantity | Means | Typically |
|---|---|---|
| the power per GPU | 0.3–0.7 kW | |
| the time | hours | |
| the number of GPUs | ||
| PUE | the overhead for cooling and so on | 1.1–1.6 |
An example: 512 GPUs at 0.4 kW for 30 days, PUE 1.2:
That corresponds to roughly eight Swedish houses' annual consumption. With the Swedish electricity mix (about 30 g CO₂e/kWh) that becomes about 5 tonnes of CO₂e; with a coal-dependent mix (700 g/kWh) closer to 124 tonnes.
Inference per call is small but multiplied. A rough estimate for a medium-sized model: 0.3–3 Wh per answer. At 10 million answers a day that is 3–30 MWh a day — that is, in the same region as a whole training run, every week.
What actually matters, in order of size:
| The choice | The effect |
|---|---|
| The origin of the electricity | up to a 20× difference in CO₂e |
| The model size | roughly proportional to the computation |
| Quantisation | 2–4× less energy per call |
| Batching | better GPU utilisation, often 2–5× |
| Caching | eliminates the call entirely |
| Not retraining unnecessarily | the largest saving there is |
Be honest about the uncertainty. The figures above are orders of magnitude, not measurements — the real consumption depends on the hardware, the utilisation and the data centre. codecarbon and similar tools measure the actual consumption during a run and are to be preferred over estimates where possible.
And put it in context. A single LLM answer consumes less than streaming video for a few minutes. That does not make the question unimportant — the sum over billions of calls is significant, and the data centres' share of electricity consumption is growing fast — but the proportions are worth getting right when you argue.
Code
SWEDISH_MIX = 30 # g CO2e/kWh
EU_MIX = 250
COAL_MIX = 700
def training_energy(gpu_power_kw, gpu_count, days, pue=1.2):
return gpu_power_kw * 24 * days * gpu_count * pue # kWh
def co2(kwh, g_per_kwh):
return kwh * g_per_kwh / 1000 # kg CO2e
E = training_energy(0.4, 512, 30)
print(f"{E:,.0f} kWh") # 176,947 kWh
for name, mix in (("Swedish", SWEDISH_MIX), ("EU", EU_MIX), ("coal", COAL_MIX)):
print(f" {name:<7} {co2(E, mix) / 1000:>6.1f} tonnes CO2e")
# Swedish 5.3 tonnes CO2e
# EU 44.2 tonnes CO2e
# coal 123.9 tonnes CO2e ← a 23× difference, the same computation
# Compare with everyday references
HOUSE_YEAR = 20_000 # kWh
FLIGHT_STHLM_NY = 1_000 # kg CO2e per passenger, return (a rough figure)
print(f" = {E / HOUSE_YEAR:.1f} houses' annual consumption")
print(f" = {co2(E, SWEDISH_MIX) / FLIGHT_STHLM_NY:.1f} flights Stockholm-New York (Swedish mix)")
# Inference dominates over time
WH_PER_ANSWER = 1.0
for answers_per_day in (100_000, 10_000_000):
annual = answers_per_day * WH_PER_ANSWER * 365 / 1000
print(f"{answers_per_day:>12,} answers/day → {annual:>10,.0f} kWh/year "
f"({annual / E:.1f}× the training)")
# 100,000 answers/day → 36,500 kWh/year (0.2× the training)
# 10,000,000 answers/day → 3,650,000 kWh/year (20.6× the training)
# Measure instead of estimating where you can:
# from codecarbon import EmissionsTracker
# with EmissionsTracker() as t: train()
Mastery means
- Estimates the energy use of training and inference
- Compares it with everyday references
- Knows which choices actually make a difference
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Carbon Emissions and Large Neural Network Training — arXiv (open access; licence per article)
- CodeCarbon (MIT) — MIT
- Energimyndigheten — myndighetsmaterial