Skip to content
AI-grafen
FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Code agents

Be able to build an agent that writes, runs and corrects code against tests.

Prerequisites

Intuition

Code has a property text lacks: it can be run. That makes code agents the agent case that works best — the test result is objective, automatic feedback in every round.

The loop:

read the task and the existing code
→ write/change the code
→ run the tests in the sandbox
→ tests green? done
→ otherwise: read the error output, change, try again (at most n rounds)

What decides the quality is not the model but the feedback: the whole stack trace, which tests failed, and which code was run. An agent that is only told «the tests failed» is groping.

Code

def code_agent(task, files, run_tests, llm, max_rounds=6):
    """run_tests is ALWAYS run in a sandbox without a network. Returns (ok, report)."""
    history = []
    for rnd in range(max_rounds):
        ok, report = run_tests(files)
        if ok:
            return {"status": "done", "round": rnd, "files": files, "history": history}
        if rnd and report == history[-1]["report"]:
            history.append({"round": rnd, "note": "no change in the test result"})
            return {"status": "stuck", "round": rnd, "files": files, "history": history}

        change = llm.generate({
            "task": task,
            "files": files,
            "test_report": report[:4000],           # the whole error output, not just "failed"
            "previous_attempts": [h.get("summary") for h in history[-2:]],
        })
        files = apply(files, change)                 # a patch, not the whole file
        history.append({"round": rnd, "report": report, "summary": change["motivation"]})
    return {"status": "budget_exhausted", "files": files, "history": history}

Four things that make the difference in practice:

  1. Patches, not whole files — less risk of the agent deleting unrelated code.
  2. Stop at stagnation — the same test result two rounds running means the strategy is not working.
  3. The agent must not be able to change the tests — otherwise it «solves» the task by removing the test.
  4. A sandbox without a network — generated code is untrusted by definition.

AI-grafen's lab runner is built for exactly this: files in, the tests run in isolation, a structured result out.

Mastery means

  • Builds an agent that writes, runs and corrects code against tests
  • Uses the test result as the control signal
  • Runs all generated code in isolation

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences