Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Language models for code

Be able to use and evaluate code models and understand fill-in-the-middle.

Prerequisites

Intuition

Code models are trained on code plus text, but with two differences from ordinary language models:

1. Fill-in-the-middle (FIM). An ordinary autoregressive model can only continue forwards. But in an editor the cursor is in the middle of a file — there is code both before and after. FIM training rearranges examples into <prefix> <suffix> <middle> so that the model learns to fill gaps in with knowledge of both directions.

2. Executable evaluation. Code has an answer key that can be run. pass@k measures the probability that at least one of k generated suggestions passes the tests — a far harder and more honest measure than textual similarity.

Formal

pass@k is estimated unbiasedly from n generated solutions of which c pass (Chen et al. 2021): pass@k=E[1−(n−ck)(nk)]\text{pass@}k = \mathbb E\left[1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}\right] Generate n = 20, count c, compute it for k = 1, 5, 10. Generating k times and taking the share that succeeded is biased and overestimates.

Benchmarks and their status: HumanEval (164 tasks, saturated and contaminated), MBPP, SWE-bench (real GitHub issues, much harder), LiveCodeBench (rotating tasks published after the models' cutoff — the best protection against contamination).

Risks specific to generated code:

  • Security holes: models reproduce insecure patterns that are common in the training data (SQL injection, hard-coded secrets). Pearce et al. (2021) found that ~40 % of generated code in security-sensitive scenarios contained vulnerabilities.
  • Hallucinated dependencies: the model imports packages that do not exist — which opens the way to «slopsquatting», where somebody registers precisely that package name.
  • Licence contamination: code that verbatim reproduces copyleft-licensed training data.
  • Silent incorrectness: the code runs and looks reasonable but does the wrong thing. The tests are your only protection.

So the rule is simple: review and test generated code as if it came from an unknown contributor — because it does.

Code

import itertools, numpy as np

def pass_at_k(n: int, c: int, k: int) -> float:
    """An unbiased estimate. n generated, c correct."""
    if n - c < k:
        return 1.0
    return 1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1))

for c in (0, 1, 5, 20):
    print(c, [round(pass_at_k(20, c, k), 3) for k in (1, 5, 10)])
# 0  [0.0, 0.0, 0.0]
# 1  [0.05, 0.25, 0.5]
# 5  [0.25, 0.806, 0.984]
# 20 [1.0, 1.0, 1.0]

# The FIM format (StarCoder/CodeLlama style)
PROMPT = "<fim_prefix>{prefix}<fim_suffix>{suffix}<fim_middle>"

# Always run generated code in a sandbox — never in the process that generated it.

AI-grafen's own code exercises and labs build on exactly this principle: generated or submitted code is run in an isolated container without a network, and is judged by tests — not by how the code looks.

Mastery means

  • Explains fill-in-the-middle and why it is needed
  • Evaluates code models with executable tests
  • Knows the risks with generated code

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences