Skip to content
AI-grafen

Closed beta for adults · free

Evals that hold up between releases

LLM as judge, regression tests, statistical significance and multiple comparisons — so that «better» means something.

Sound familiar?

Two percentage points better — improvement or noise?
The LLM judge agrees with itself.
The regression is found by users, not by the tests.

What you can do afterwards

The goal covers 62 knowledge nodes from the basics up, roughly 54 hours if you start from zero. The diagnostic removes what you already know.

Labs along the way

You write the code in the browser or in a sandbox on the server. Hidden tests decide whether it holds up.

What you get

Diagnosis first

The questions follow the prerequisite chain back from the goal. You skip what you already know.

The shortest path

Only the knowledge you are missing, in the order it builds on itself. The path is recalculated as you learn.

Labs with hidden tests

You write the code. Tests you cannot see decide whether it holds up — not whether it looks right.

Proof of what you can do

Mastery requires several kinds of evidence. The certificate lists them and can be verified.

See where you stand — in ten minutes

No account. You see right away what you already know and where your path would start.

Evals in practiceF

LLM as judge, regression tests, statistical significance and how to value negative results — measurement that holds between releases.