Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Reasoning in language models

Be able to use and evaluate chain-of-thought and understand its limitations.

Prerequisites

Intuition

Ask the model to «think step by step» before the answer and multi-step tasks improve markedly — mathematics, logic, planning.

Why? The model has a fixed compute budget per token. Writing intermediate steps gives it more tokens to compute in, and each step becomes a simpler prediction than jumping straight to the final answer.

Variants:

  • Zero-shot CoT: add «Let us think step by step.»
  • Few-shot CoT: show examples where you reason yourself.
  • Self-consistency: sample 5–20 reasonings at T > 0 and take the most common final answer. Substantially better at mathematics, but 5–20× more expensive.
  • Reasoning models (the o series, R1): trained to produce long internal traces before the answer.

Formal

The important warning: the reasoning trace is not necessarily the explanation for the answer.

Turpin et al. (2023) showed that models can be influenced by something in the prompt (that all the examples have the answer «A», say) and then produce a convincing reasoning that never mentions the real cause. The trace is post hoc rationalisation, exactly as people's sometimes are.

The consequences in practice:

  • Do not use the trace as audit evidence or as an explanation to a user.
  • Evaluate the final answer, not how plausible the reasoning sounds.
  • Do not show raw reasoning traces to end users — they create a confidence that is not warranted.

When CoT does not help: simple lookup questions (just more expensive), and classification where the answer is a label. Always measure with and without before making it the default.

Code

from collections import Counter
import re

def self_consistency(llm, question, n=5):
    answers = []
    for _ in range(n):
        out = llm(f"{question}\n\nThink step by step and finish with 'ANSWER: <number>'.", temperature=0.7)
        m = re.search(r"ANSWER:\s*(-?\d+(?:[.,]\d+)?)", out)
        if m:
            answers.append(m.group(1).replace(",", "."))
    if not answers:
        return None, 0.0
    most_common, count = Counter(answers).most_common(1)[0]
    return most_common, count / len(answers)     # the share = a rough confidence measure

answer, agreement = self_consistency(llm, "A shop sells 3 apples for 12 kr. What do 7 apples cost?")
print(answer, agreement)  # 28 0.8

The degree of agreement is useful: low agreement means the task is hard for the model — escalate or abstain.

Mastery means

  • Uses and evaluates step-by-step reasoning
  • Knows about self-consistency
  • Understands that the trace is not a true explanation

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences