Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evals for code generation

Be able to build pass@k evals with sandboxed execution.

Prerequisites

Intuition

Code evals have an answer key that can be run — but four pitfalls recur:

  1. Contamination. HumanEval and MBPP are in every training corpus. High numbers can be memorisation. Use your own tasks or rotating benchmarks (LiveCodeBench).
  2. Leakage in the task. If the solution is visible in the docstring or in an imported example you are measuring nothing.
  3. Weak tests. A task whose tests only check one normal case passes broken code. The tests should cover edge cases.
  4. Correctness only. Code can be right and still be insecure, slow or unreadable.

Always run in a sandbox. Generated code should be regarded as untrusted input — --network none, a timeout, resource caps.

Code

import numpy as np

def pass_at_k(n, c, k):
    if n - c < k:
        return 1.0
    return 1.0 - float(np.prod(1.0 - k / np.arange(n - c + 1, n + 1)))

def run_code_eval(model, tasks, sandbox, n=10, k_list=(1, 5)):
    rows = []
    for t in tasks:
        solutions = [model(t["prompt"], temperature=0.8) for _ in range(n)]
        c = 0
        for code in solutions:
            res = sandbox.run(files={**t["files"], t["target"]: code, **t["tests"]},
                              timeout_s=30, network=None, memory_mb=512)
            c += int(res["status"] == "passed")
        rows.append({"id": t["id"], "n": n, "c": c,
                     **{f"pass@{k}": round(pass_at_k(n, c, k), 3) for k in k_list}})
    overall = {f"pass@{k}": round(float(np.mean([r[f"pass@{k}"] for r in rows])), 3) for k in k_list}
    return overall, rows

# Complement it with:
#  - a security scanner (SAST) on the generated code
#  - a check that every imported package is on an allowed list
#  - the running time and memory per solution (working but unmanageably slow code)

Build your own tasks from your codebase: take real issues or functions, write the tests first, remove the implementation. That gives tasks that are impossible to have memorised and that measure something you actually care about.

Mastery means

  • Builds pass@k evals with sandboxed execution
  • Avoids contamination and test leakage
  • Measures more than functional correctness

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences