FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Evals for code generation
Be able to build pass@k evals with sandboxed execution.
Prerequisites
- EBuild an eval harnessrequired
- ELanguage models for coderequired
Intuition
Code evals have an answer key that can be run — but four pitfalls recur:
- Contamination. HumanEval and MBPP are in every training corpus. High numbers can be memorisation. Use your own tasks or rotating benchmarks (LiveCodeBench).
- Leakage in the task. If the solution is visible in the docstring or in an imported example you are measuring nothing.
- Weak tests. A task whose tests only check one normal case passes broken code. The tests should cover edge cases.
- Correctness only. Code can be right and still be insecure, slow or unreadable.
Always run in a sandbox. Generated code should be regarded as untrusted input — --network none, a timeout, resource caps.
Code
import numpy as np
def pass_at_k(n, c, k):
if n - c < k:
return 1.0
return 1.0 - float(np.prod(1.0 - k / np.arange(n - c + 1, n + 1)))
def run_code_eval(model, tasks, sandbox, n=10, k_list=(1, 5)):
rows = []
for t in tasks:
solutions = [model(t["prompt"], temperature=0.8) for _ in range(n)]
c = 0
for code in solutions:
res = sandbox.run(files={**t["files"], t["target"]: code, **t["tests"]},
timeout_s=30, network=None, memory_mb=512)
c += int(res["status"] == "passed")
rows.append({"id": t["id"], "n": n, "c": c,
**{f"pass@{k}": round(pass_at_k(n, c, k), 3) for k in k_list}})
overall = {f"pass@{k}": round(float(np.mean([r[f"pass@{k}"] for r in rows])), 3) for k in k_list}
return overall, rows
# Complement it with:
# - a security scanner (SAST) on the generated code
# - a check that every imported package is on an allowed list
# - the running time and memory per solution (working but unmanageably slow code)
Build your own tasks from your codebase: take real issues or functions, write the tests first, remove the implementation. That gives tasks that are impossible to have memorised and that measure something you actually care about.
Mastery means
- Builds pass@k evals with sandboxed execution
- Avoids contamination and test leakage
- Measures more than functional correctness
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Evaluating Large Language Models Trained on Code (HumanEval, pass@k) — arXiv (open access; licence per article)
- arXiv — LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code — arXiv (open access; licence per article)