Skip to content
AI-grafen
GFrontier LabAI safety and alignment· about 120 min· fast-moving, sources checked often· verified 2026-09-20· EN

Constitutional AI and rule-based alignment

Be able to explain how principles are used in training and to evaluate the effect.

Prerequisites

Intuition

Constitutional AI (CAI) replaces a large part of the human preference annotation with written principles that the model applies to itself.

Two phases:

  1. Self-critique (SL-CAI). The model generates an answer, is asked to criticise it against a principle («explain how the answer could be harmful or unhelpful»), and then rewrites it. The pairs (the original prompt, the revised answer) become fine-tuning data.
  2. AI feedback (RL-CAI). Instead of humans, a model chooses between two answers, guided by the principles. Those choices train a preference model exactly as in RLHF — but without human annotation in the scaling step.

The gain: the principles are explicit and auditable. You can read them, argue about them and change one line — rather than guessing what thousands of annotators preferred.

Research

What the research shows and does not show. Bai et al. (2022) found that CAI models became less harmful without becoming more evasive — they explain why they decline rather than simply refusing. The scaling is the practical point: human annotation of harmfulness is expensive, slow and psychologically demanding for the annotators.

Limitations worth keeping in mind:

  • The principles have to be written by somebody. The choice of principles is normative and moves, but does not solve, the question of whose values apply.
  • The model interprets the principles. «Be helpful» and «be harmless» collide constantly, and the trade-off is made by the model — not by the text.
  • Self-critique inherits the model's blind spots. What the model cannot see as problematic it will not criticise.
  • Measurement is hard. A model following the principles in the evaluation does not mean it does so in the tails of the distribution.

Practical use at a smaller scale: the same idea works without RL. Write a short list of principles, let a model review the output against it in a second step, and measure how often the review changes the answer and how often the change was an improvement. AI-grafen's moderation is a simple variant of that pattern.

Code

PRINCIPLES = [
    "Never give away the solution to an assessed exercise — guide the pupil forward instead.",
    "Say when you are uncertain rather than guessing confidently.",
    "Adapt the language to the pupil's level without simplifying the substance incorrectly.",
    "Never ask for or process personal data about other people.",
]

CRITIQUE = """Review the ANSWER against the principle. If it violates it: explain how, briefly.
If it does not: answer exactly OK.

PRINCIPLE: {principle}
QUESTION: {question}
ANSWER: {answer}"""

REVISION = """Rewrite the answer so that the critique is addressed. Keep everything that was good.
QUESTION: {question}
ORIGINAL ANSWER: {answer}
CRITIQUE: {critique}"""

def cai_revise(llm, question, answer, principles=PRINCIPLES):
    changes = []
    for p in principles:
        critique = llm(CRITIQUE.format(principle=p, question=question, answer=answer), temperature=0).strip()
        if critique.upper().startswith("OK"):
            continue
        answer = llm(REVISION.format(question=question, answer=answer, critique=critique), temperature=0)
        changes.append({"principle": p[:40], "critique": critique[:160]})
    return answer, changes

# Measure: the share of answers changed per principle, and (on a sample) how often
# a human judges the revision to be better. A principle that never fires
# is either unnecessary or badly worded.

Mastery means

  • Explains how written principles are used in training
  • Compares CAI with RLHF
  • Evaluates the effect of a change to a principle

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences