Agent Evaluation
Evaluating an agent means checking whether it chose the right tools and reached a correct outcome, not just whether one response looked reasonable.
Prerequisites
Overview
A single LLM call can be evaluated by comparing its output to an expected answer. An agent takes many steps — evaluating it means checking the whole trajectory: did it pick the right tools, in a reasonable order, and reach a correct final outcome.
Where It Fits
Full Agent Trace
Tool Choice Correctness
Step Ordering
Final Outcome
Key Points
- Trajectory evaluation
- Checking the sequence of tool calls and decisions an agent made, not only its final answer.
- Outcome correctness
- A correct-looking final answer reached via the wrong tools or a lucky path is still a fragile result worth catching.
- Reproducibility
- Non-deterministic agent behavior makes evaluation noisier — running each test case multiple times is often necessary.
Interview Question
Why is evaluating an agent harder than evaluating a single prompt?
A single prompt produces one output that can be checked directly. An agent produces a whole trajectory — a sequence of decisions and tool calls — so evaluation has to check whether it took a sound path to the answer, not just whether the final answer happens to look right, since a lucky or fragile path won’t generalize.
Explain It in 30 Seconds
Agent evaluation checks the entire trajectory of tool calls and decisions, not just the final output, since a correct-looking answer reached the wrong way is a fragile result that evaluation needs to catch.
Real-World Stack
Technologies commonly used to implement this in production.