Skip to content
AI-grafen
FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Observability for agents

Be able to log, trace and debug agent runs step by step.

Prerequisites

Intuition

An agent that fails without a trace is impossible to debug. You only see «it went wrong» — not where.

Log per run: a run_id, the task, the model and prompt version, the end state, the total cost and the time.

Log per step: the step number, the model's output, the tool and its arguments, the result (truncated), the tokens, the latency, and any error.

That gives a trace — a timeline that can be read from the top down. The standard for that is OpenTelemetry spans, but a JSONL file per run goes a long way and can be read with jq.

Code

import json, time, uuid
from contextlib import contextmanager

class Trace:
    def __init__(self, task, model, prompt_id, path="traces"):
        self.run_id = uuid.uuid4().hex[:12]
        self.f = open(f"{path}/{self.run_id}.jsonl", "w", encoding="utf-8")
        self.t0 = time.perf_counter()
        self._write("start", {"task": task, "model": model, "prompt_id": prompt_id})

    def _write(self, kind, data):
        self.f.write(json.dumps({"run_id": self.run_id, "kind": kind,
                                 "t_ms": round((time.perf_counter() - self.t0) * 1000), **data},
                                ensure_ascii=False) + "\n")
        self.f.flush()

    @contextmanager
    def step(self, n, tool=None, args=None):
        t = time.perf_counter()
        entry = {"step": n, "tool": tool, "args": str(args)[:200]}
        try:
            yield entry
            entry["error"] = None
        except Exception as e:
            entry["error"] = f"{type(e).__name__}: {e}"[:200]
            raise
        finally:
            entry["latency_ms"] = round((time.perf_counter() - t) * 1000)
            self._write("step", entry)

    def end(self, status, tokens):
        self._write("end", {"status": status, "tokens": tokens}); self.f.close()

Three questions the trace should be able to answer in ten seconds: Which step broke? What did we feed into it? How much did the run cost and where did the time go?

Privacy: log the tool arguments and results truncated, and mask fields that may contain personal data. An agent trace is otherwise one of the most sensitive logs in the system.

Mastery means

  • Logs every agent step traceably
  • Debugs a failed run from the trace
  • Measures the cost and the latency per step

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences