Language models for code
Be able to use and evaluate code models and understand fill-in-the-middle.
Prerequisites
- CPython — the basicsrequired
- ELanguage models — training and generationrequired
Intuition
Code models are trained on code plus text, but with two differences from ordinary language models:
1. Fill-in-the-middle (FIM). An ordinary autoregressive model can only continue forwards. But in an editor the cursor is in the middle of a file — there is code both before and after. FIM training rearranges examples into <prefix> <suffix> <middle> so that the model learns to fill gaps in with knowledge of both directions.
2. Executable evaluation. Code has an answer key that can be run. pass@k measures the probability that at least one of k generated suggestions passes the tests — a far harder and more honest measure than textual similarity.
Formal
pass@k is estimated unbiasedly from n generated solutions of which c pass (Chen et al. 2021): Generate n = 20, count c, compute it for k = 1, 5, 10. Generating k times and taking the share that succeeded is biased and overestimates.
Benchmarks and their status: HumanEval (164 tasks, saturated and contaminated), MBPP, SWE-bench (real GitHub issues, much harder), LiveCodeBench (rotating tasks published after the models' cutoff — the best protection against contamination).
Risks specific to generated code:
- Security holes: models reproduce insecure patterns that are common in the training data (SQL injection, hard-coded secrets). Pearce et al. (2021) found that ~40 % of generated code in security-sensitive scenarios contained vulnerabilities.
- Hallucinated dependencies: the model imports packages that do not exist — which opens the way to «slopsquatting», where somebody registers precisely that package name.
- Licence contamination: code that verbatim reproduces copyleft-licensed training data.
- Silent incorrectness: the code runs and looks reasonable but does the wrong thing. The tests are your only protection.
So the rule is simple: review and test generated code as if it came from an unknown contributor — because it does.
Code
import itertools, numpy as np
def pass_at_k(n: int, c: int, k: int) -> float:
"""An unbiased estimate. n generated, c correct."""
if n - c < k:
return 1.0
return 1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1))
for c in (0, 1, 5, 20):
print(c, [round(pass_at_k(20, c, k), 3) for k in (1, 5, 10)])
# 0 [0.0, 0.0, 0.0]
# 1 [0.05, 0.25, 0.5]
# 5 [0.25, 0.806, 0.984]
# 20 [1.0, 1.0, 1.0]
# The FIM format (StarCoder/CodeLlama style)
PROMPT = "<fim_prefix>{prefix}<fim_suffix>{suffix}<fim_middle>"
# Always run generated code in a sandbox — never in the process that generated it.
AI-grafen's own code exercises and labs build on exactly this principle: generated or submitted code is run in an isolated container without a network, and is judged by tests — not by how the code looks.
Mastery means
- Explains fill-in-the-middle and why it is needed
- Evaluates code models with executable tests
- Knows the risks with generated code
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Evaluating Large Language Models Trained on Code (HumanEval, pass@k) — arXiv (open access; licence per article)
- arXiv — Efficient Training of Language Models to Fill in the Middle — arXiv (open access; licence per article)
- arXiv — Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions — arXiv (open access; licence per article)