FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Code agents
Be able to build an agent that writes, runs and corrects code against tests.
Prerequisites
- DTesting with pytestrequired
- EThe ReAct loop: think, act, observerequired
- ELanguage models for coderequired
Intuition
Code has a property text lacks: it can be run. That makes code agents the agent case that works best — the test result is objective, automatic feedback in every round.
The loop:
read the task and the existing code
→ write/change the code
→ run the tests in the sandbox
→ tests green? done
→ otherwise: read the error output, change, try again (at most n rounds)
What decides the quality is not the model but the feedback: the whole stack trace, which tests failed, and which code was run. An agent that is only told «the tests failed» is groping.
Code
def code_agent(task, files, run_tests, llm, max_rounds=6):
"""run_tests is ALWAYS run in a sandbox without a network. Returns (ok, report)."""
history = []
for rnd in range(max_rounds):
ok, report = run_tests(files)
if ok:
return {"status": "done", "round": rnd, "files": files, "history": history}
if rnd and report == history[-1]["report"]:
history.append({"round": rnd, "note": "no change in the test result"})
return {"status": "stuck", "round": rnd, "files": files, "history": history}
change = llm.generate({
"task": task,
"files": files,
"test_report": report[:4000], # the whole error output, not just "failed"
"previous_attempts": [h.get("summary") for h in history[-2:]],
})
files = apply(files, change) # a patch, not the whole file
history.append({"round": rnd, "report": report, "summary": change["motivation"]})
return {"status": "budget_exhausted", "files": files, "history": history}
Four things that make the difference in practice:
- Patches, not whole files — less risk of the agent deleting unrelated code.
- Stop at stagnation — the same test result two rounds running means the strategy is not working.
- The agent must not be able to change the tests — otherwise it «solves» the task by removing the test.
- A sandbox without a network — generated code is untrusted by definition.
AI-grafen's lab runner is built for exactly this: files in, the tests run in isolation, a structured result out.
Mastery means
- Builds an agent that writes, runs and corrects code against tests
- Uses the test result as the control signal
- Runs all generated code in isolation
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — arXiv (open access; licence per article)
- arXiv — Evaluating Large Language Models Trained on Code — arXiv (open access; licence per article)