Q: What is a trace, and how do you debug an agent?
A trace is the complete recording of everything your AI system did in one run: every prompt, every model response, every tool call with its arguments and results, plus the tokens and time each step consumed.
You need traces because agents fail invisibly. In normal code, a failure gives you a stack trace pointing at a line. When an agent produces a wrong answer, the code ran perfectly. The failure is somewhere in the reasoning chain: a retrieval step pulled the wrong chunks, a tool returned an error the model ignored, or the model took a wrong turn at step three that poisoned everything after it. The only way to find the failure is to read the trace and see what the model saw at each step.
Debugging an agent looks like this: open the trace of a failed run, walk through it step by step, find where reality diverged from your intent, then fix that step (usually the prompt, the tool description, or the retrieval). Tools like LangSmith and Langfuse capture and display traces; the OpenAI platform has tracing built in.
Interviewers increasingly probe this, because it separates people who've actually built and debugged AI agents through trial and error from people who've copy-pasted code from a YouTube tutorial. A sample interview question looks like this: "your agent gives wrong answers, what do you do?" The answer starts with "I read the trace."